Zhichun Zhu

dblp:40/1416 · DBLP profile ↗
← Back
26ranked-venue papers
2as first author
3since 2021 · last 2023
0000-0002-7928-9024ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 25 · 2 first-author · 3 since 2021Software engineering, systems software and programming languages · 5
YearPublicationVenuePosition
2023 Polling-Based Memory Interface
abstract
Non-volatile memory has been extensively researched as the alternative for a DRAM-based system; however, the traditional memory controller cannot efficiently track and schedule operations for all the memory devices in heterogeneous systems due to different timing requirements and complex architecture supports of various memory technologies. To address this issue, we propose a hybrid memory architecture framework called POMI (POlling-based Memory Interface). It uses a small buffer chip inserted on each DIMM (Dual In-line Memory Module) to decouple operation scheduling from the controller to enable the support for diverse memory technologies in the system. Unlike the conventional DRAM-based system, POMI uses a polling-based memory bus protocol for communication and to resolve any bus conflicts between memory modules. The buffer chip on each DIMM will provide feedback information to the main memory controller so that the polling overhead is trivial. We propose two unique designs. The first one adds additional bus lines for sending the feedback information, and the second one utilizes the Command/Address bus. The framework provides several benefits: a technology-independent memory system, higher parallelism, and better scalability. Our experimental results show that POMI can efficiently support both homogeneous and heterogeneous systems. Compared with the conventional DDR4-2400 implementation, our scheme improves the performance of memory-intensive workloads by 3.7% on average. Compared with an existing interface for hybrid memory systems, Twin-Load, it also improves performance by 22.0% on average for memory-intensive workloads.
Trung Le 0003, Zhao Zhang 0008, Zhichun Zhu
ACM Trans. Design Autom. Electr. Syst.3
2021 POMI: Polling-Based Memory Interface for Hybrid Memory System
abstract
Modern conventional DRAM main memory system will no longer satisfy the growing demand for capacity and bandwidth on today’s data-intensive applications. Non-Volatile Memory (NVM) has been extensively researched as the alter-native for DRAM-based system due to its higher density and non-volatile characteristics. Hybrid memory system benefits from both DRAM and NVM technologies, however traditional Memory Controller (MC) cannot efficiently track and schedule operations for all the memory devices in heterogeneous systems due to different timing requirements and complex architecture supports of various memory technologies. To address this issue, we propose a hybrid memory architecture framework called POMI. It uses a small buffer chip inserted on each DIMM to decouple operation scheduling from the controller to enable the support for diverse memory technologies in the system. Unlike the conventional DRAM-based system, which relies on the main MC to govern all DIMMs, POMI uses polling-based memory bus protocol for communication and to resolve any bus conflicts between memory modules. The buffer chip on each DIMM will provide feedback information to the main MC so that the polling overhead is trivial. This gives several benefits: technology-independent memory system, higher parallelism, and better scalability. Our experimental results using octa-core workloads show that POMI can efficiently support heterogeneous systems and it outperforms an existing interface for hybrid memory systems by 22.0% on average for memory-intensive workloads.
Trung Le 0003, Zhao Zhang 0008, Zhichun Zhu
ICCD3
2021 Memory-Side Prefetching Scheme Incorporating Dynamic Page Mode in 3D-Stacked DRAM
abstract
Modern multiprocessor systems running multiple applications concurrently exhibit irregular memory access pattern during different phases of execution. The principle of locality is hard to exploit in the presence of such irregular memory requests and may result in additional delays due to resource conflicts throughout memory hierarchy. Prefetching is a promising technique to reduce the memory access latency where data is speculatively fetched ahead of time and stored in a faster memory structure like cache or dedicated prefetch buffer. The emergence of 3D-stacked DRAM provides huge internal bandwidth that makes memory-side prefetching an effective approach to improving system performance. Leveraging the unique architecture of 3D-stacked DRAM, we introduce a memory-side prefetching scheme that works in conjunction with dynamic page mode to reduce memory access latency. We introduce a novel prefetch buffer management scheme that makes intelligent replacement decision based on the utilization and recency of the prefetched data, which also serves as a guidance for future prefetching. Simulation results indicate that our approach improves performance by 21.8 percent on average, compared to a baseline scheme that prefetches a whole row on consecutive hits and implements static open page policy. Our scheme also outperforms an existing memory-side prefetching scheme by 13.2 percent on average, which dynamically adjusts the prefetch degree based on the usefulness of prefetched data.
Muhammad M. Rafique, Zhichun Zhu
IEEE Trans. Parallel Distributed Syst.2
2020 DeepSwapper: A Deep Learning Based Page Swap Management Scheme for Hybrid Memory Systems
abstract
In this paper, we introduce DeepSwapper, a deep learning-based page swap management scheme that utilizes RNN to perform fast, energy-efficient, and temperature-aware page swapping in hybrid memory systems. DeepSwapper comprises of LSTM units of RNN model to predict the future memory accesses to guide its swap management scheme, a dynamic page swap management scheme that utilizes DRAM capacity efficiently by enabling hot pages in a swap group to be swapped with cold pages of another swap group, and a temperature-aware page swap management scheme, which first predicts the future writes to NVM pages and then, decides to migrate those pages with frequent writes in hot NVM banks to DRAM to enhance the NVM lifetime.
Majed Valad Beigi, Bahareh Pourshirazi, Gokhan Memik, Zhichun Zhu
PACT4
2019 Writeback-Aware LLC Management for PCM-Based Main Memory Systems
abstract
With the increase in the number of data-intensive applications on today's workloads, DRAM-based main memories are struggling to satisfy the growing data demand capacity. Phase Change Memory (PCM) is a type of non-volatile memory technology that has been explored as a promising alternative for DRAM-based main memories due to its better scalability and lower leakage energy. Despite its many advantages, PCM also has shortcomings such as long write latency, high write energy consumption, and limited write endurance, which are all related to the write operations. In this article, we propose a novel writeback-aware Last Level Cache (LLC) management scheme named WALL to reduce the number of LLC writebacks and consequently improve performance, energy efficiency, and lifetime of a PCM-based main memory system. First, we investigate the writeback behavior of LLC sets and show that writebacks are not uniformly distributed among sets; some sets observe much higher writeback rates than others. We then propose a writeback-aware set-balancing mechanism, which employs the underutilized LLC sets with few writebacks as an auxiliary storage for the evicted dirty lines from sets with frequent writebacks. We also propose a simple and effective writeback-aware replacement policy to avoid the eviction of the dirty blocks that are highly reused after being evicted from the cache. Our experimental results show that WALL achieves an average of 30.9% reduction in the total number of LLC writebacks, compared to the baseline scheme, which uses the LRU replacement policy. As a result, WALL can reduce the memory energy consumption by 23.1% and enhance PCM lifetime by 1.29×, on average, on an 8-core system with a 4GB PCM main memory, running memory-intensive applications.
Bahareh Pourshirazi, Majed Valad Beigi, Zhichun Zhu, Gokhan Memik
ACM Trans. Design Autom. Electr. Syst.3
2018 WALL: A writeback-aware LLC management for PCM-based main memory systems
abstract
In this paper, we propose WALL, a novel writeback-aware LLC management scheme to reduce the number of LLC writebacks and consequently improve performance, energy efficiency, and lifetime of a PCM-based main memory system. First, we investigate the writeback behavior of LLC sets and show that writebacks are not uniformly distributed among sets; some sets observe much higher writeback rates than others. We then propose a writeback-aware set-balancing mechanism, which employs the underutilized LLC sets with few writebacks as an auxiliary storage for storing the evicted dirty lines of sets with frequent writebacks. We also propose a simple and effective writeback-aware replacement policy to avoid the eviction of the writeback blocks that are highly reused after being evicted from the cache. Our experimental results show that WALL achieves an average of 26.6% reduction in the total number of LLC writebacks, compared to the baseline scheme, which uses the LRU replacement policy. As a result, WALL can reduce the memory energy consumption by 19.2% and enhance PCM lifetime by 1.25x, on average, on an 8-core system with a 4GB PCM main memory, running memory intensive applications.
Bahareh Pourshirazi, Majed Valad Beigi, Zhichun Zhu, Gokhan Memik
DATE3
2018 CAMPS: Conflict-Aware Memory-Side Prefetching Scheme for Hybrid Memory Cube
abstract
Prefetching is a well-studied technique where data is fetched from main memory ahead of time speculatively and stored in caches or dedicated prefetch buffer. With the introduction of Hybrid Memory Cube (HMC), a 3-D memory module with multiple memory layers stacked over a single logic layer using thousands of Through Silicon Vias (TSVs), huge internal bandwidth availability makes memory-side prefetching a more efficient approach to improving system performance. In this paper, we introduce a memory-side prefetching scheme for HMC based main memory system that utilizes its logic area and exploits the huge internal bandwidth provided by TSVs. Our scheme closely monitors the access pattern to memory banks and make intelligent prefetch decisions for rows with high utilization or causing row buffer conflicts. We also introduce a prefetch buffer management scheme that makes replacement decision within the prefetch buffer based on both the utilization and recency of the prefetched rows. Our simulation results indicate that our approach improves performance by 17.9% on average, compared to a baseline scheme that prefetches a whole row on every memory request. Our scheme also outperforms an existing memory-side prefetching scheme by 8.7% on average, which dynamically adjusts the prefetch degree based on the usefulness of prefetched data. In this sample-structured document, neither the cross-linking of float elements and bibliography nor metadata/copyright information is available. The sample document is provided in "Draft" mode and to view it in the final layout format, applying the required template is essential with some standard steps.
Muhammad M. Rafique, Zhichun Zhu
ICPP2
2016 Refree: A Refresh-Free Hybrid DRAM/PCM Main Memory System
abstract
As the number of concurrently running applications on chip multiprocessors and the size of each application's working set increase, so does the demand for larger capacity memories. DRAM-based memories can no longer satisfy this growing demand due to their scalability limit. Hybrid main memories, such as DRAM plus PCM, have been proposed as a solution to this issue. However, using a very small DRAM is not effective in a hybrid memory system as well. On the other hand, the larger the DRAM size, the higher the refresh operation cost. In this paper, we introduce Refree, a scheme to eliminate DRAM refresh operations in a hybrid DRAM/PCM main memory system. When it is time to refresh a row, Refree evicts the row from DRAM instead. This can be done since a recently accessed row has already been "refreshed" by the access, while a row that hasn't been accessed for a long time is very likely to hold obsolete data and does not need to be refreshed and kept in the DRAM. If an evicted row is dirty, it will be written back to the PCM. To alleviate the potential performance loss due to the long PCM write latency, we propose a scheme that distributes writebacks of a dirty DRAM row over an epoch time (i.e., 128ms) to prevent the long-time obstruction of other requests to the PCM devices. Our simulation results reveal that Refree achieves an average of 11.7% reduction in memory power consumption and 4.2% performance improvement on a quad-core system running NAS and PARSEC applications with 4GB DRAM and 32GB PCM, compared to the standard auto-refresh scheme. Compared with a recently proposed refresh-reduction scheme, Refree can also save memory power by 3.1% on average with a negligible 0.2% performance loss.
Bahareh Pourshirazi, Zhichun Zhu
IPDPS2
2014 Access-Aware Memory Thermal Management
abstract
Recently, thermal management of memory subsystem has become an issue for server platforms. An effective solution is to gate processor cores and thus stop sending memory requests when memory thermal emergency happens. Existing schemes usually treat all the programs running on a system equally upon memory thermal emergency. However, since the memory temperature is closely related to the memory throughput and each program generates different amount of memory traffic, the memory heat produced by each program can be very different. As a result, the existing schemes may cause severe performance loss with a low efficiency on memory temperature reduction. To address this issue, we propose a new memory thermal management scheme that halts memory-intensive applications longer than non-memory-intensive ones upon memory thermal emergency. Thus, the memory can be cooled down more quickly, which in turn brings down the performance loss. We implement and evaluate our scheme on a real system. Compared to an existing memory thermal management scheme that halts the running programs equally in a round-robin way, our scheme can improve performance by 14.3% on average (up to 21.4%).
Suyu Zhang, Zhichun Zhu
NAS2
2014 Mini-Rank: A Power-EfficientDDRx DRAM Memory Architecture
abstract
Memory power consumption has become a severe concern in multi-core computer platforms. As memory data rate, capacity and bandwidth are being pushed higher and higher, the power consumption of memory systems becomes a significant part in the overall system power profile. Conventional memory systems do not provide an efficient mechanism for managing its power and performance tradeoff. We propose a novel mini-rank architecture for DDRx memories to reduce memory power consumption by breaking each DRAM rank into multiple narrow mini-ranks and activating fewer devices for each request. We also propose a heterogeneous mini-rank design to further improve the performance-power tradeoff for each workload based on its memory access behavior and bandwidth requirement. The evaluation results show that homogeneous mini-rank significantly reduces memory power with small performance loss. For instance, using four-core multiprogramming workloads, a x32 mini-rank configuration reduces memory power by 19.5 percent with 1.3 percent performance loss on average for memory-intensive workloads. Heterogeneous mini-rank further improves the balance between the performance and power saving. For instance, it reduces the memory power by up to 38.0 percent with an average performance loss of 2.4 percent, compared with a conventional memory system. In comparison, the x32 homogeneous mini-rank reduces memory power by up to 25.4 percent; while the x8 homogeneous mini-rank incurs performance loss by up to 19.3 percent. Furthermore, heterogeneous mini-rank achieves consistently good performance-power tradeoff for workloads made by programs of diverse memory access behavior and bandwidth requirement.
Kun Fang 0005, Hongzhong Zheng, Jiang Lin, Zhao Zhang 0010, Zhichun Zhu
IEEE Trans. Computers5
2013 Conservative row activation to improve memory power efficiency
abstract
With the fast improvement on memory bandwidth and capacity, the memory power consumption has become a major contributor to the overall system power profile. Due to the increasing importance of memory-level parallelism at the multi-core era, most memory scheduling schemes eagerly exploit such parallelism to optimize performance. A common policy used by memory controllers today is, whenever possible, always trying to open memory banks for pending requests to maximize bank-level parallelism and throughput. However, we find that this is neither power optimal nor necessary for maintaining performance because usually many banks are open while waiting for the data bus ownership.
Kun Fang 0005, Zhichun Zhu
ICS2
2013 Thermal Modeling and Management of DRAM Systems
abstract
With increasing data rate and power density, high-performance memories have started to require dynamic thermal management (DTM), following the trend of processor and hard drive. There are also lack of a memory thermal model and simulation tools to facilitate the research of memory DTM. This study investigates the approach of coordinating processor, which is the source of memory access requests, and memory to improve system performance and/or power efficiency during memory thermal emergency. Two such schemes, namely adaptive core gating (DTM-ACG) and coordinated DVFS (DTM-CDVFS), are proposed and evaluated on a real server platform. DTM-ACG gates processor cores and DTM-CDVFS scales down the frequency and voltage level of processor cores according to memory thermal emergency level. Their combination, namely DTM-COMB, is also evaluated. The experimental results show that the two schemes, while successfully controlling memory activities and handling thermal emergencies, improve performance significantly under the given thermal envelope. The measurement results from an Intel SR1500AL server testbed show that on average, DTM-ACG and DTM-CDVFS improve performance by 6.7 and 15.3 percent, respectively, over a prior memory bandwidth throttling scheme. DTM-CDVFS also reduces the processor power rate by 15.5 percent and system (including processor and memory) energy by 22.7 percent. Additionally, we propose a DRAM thermal model and validate it with measurement on the instrumented server platform. We find that our proposed model faithfully catches the dynamic DRAM temperature changes; the average difference between the modeled and measured temperature is less than $(1^{\circ}{\rm C})$.
Jiang Lin, Hongzhong Zheng, Zhichun Zhu, Zhao Zhang 0010
IEEE Trans. Computers3
2011 Memory Architecture for Integrating Emerging Memory Technologies
abstract
Current main memory system design is severely limited by the decades-old synchronous DRAM architecture, which requires the memory controller to track the internal status of memory devices (chips) and schedule the timing of all device operations. This rigidity has become an obstacle of integrating emerging memory technologies such as PCM into existing memory systems, because their timing requirements are vastly different. Furthermore, with the trend of embedding memory controllers into processors, it is crucial to have interoperability among general-purpose processors and diverse memory modules. To address this issue, we propose a new memory architecture framework called universal memory architecture (UniMA). It enables the interoperability by decoupling the scheduling of device operations from memory controller, using a bridge chip at each memory module to perform local scheduling. The new architecture may also help improve memory scalability, power efficiency, and bandwidth as previously proposed decoupled memory organizations. A major focus of this study is to evaluate the performance impact of local scheduling of device operations. We present a prototype implementation of UniMA on top of DDRx memory bus, and then evaluate its efficiency with different workloads. The simulation results show that UniMA actually improves memory system efficiency for memory-intensive workloads due to increased parallelism among memory modules. The overall performance improvement over the conventional DDRx memory architecture is 3.1% on average. The performance of other workloads is reduced slightly, by 1.0% on average, due to a small increase of memory idle latency. In short, the prototype and evaluation demonstrate that it is possible to integrate diverse memory technologies into a single memory architecture with virtually no loss of overall performance.
Kun Fang 0005, Zhao Zhang 0010, Zhichun Zhu
PACT4
2010 Heterogeneous Mini-rank: Adaptive, Power-Efficient Memory Architecture
abstract
Memory power consumption has become a big concern in server platforms. A recently proposed mini-rank architecture reduces the memory power consumption by breaking each DRAM rank into multiple narrow mini-ranks and activating fewer devices for each request. However, its fixed and uniform configuration may degrade performance significantly or lose power saving opportunities on some workloads. We propose a heterogeneous mini-rank design that sets the near-optimal configuration for each workload based on its memory access behavior and its memory bandwidth requirement. Compared with the original, homogeneous mini-rank design, the heterogeneous mini-rank design can balance between the performance and power saving and avoid large performance loss. For instance, for multiprogramming workloads with SPEC2000 application running on a quad-core system with two-channel DDR3-1066 memory, on average, the heterogeneous mini-rank can reduce the memory power by 53.1% (up to 60.8%) with the performance loss of 4.6% (up to 11.1%), compared with a conventional memory system. In comparison, the x32 homogeneous mini-rank can only save memory power by up to 29.8%; and the x8 homogeneous mini-rank will cause performance loss by up to 22.8%. Compared with x16 homogeneous mini-rank configuration, it can further reduce the EDP (energy-delay product) by up to 15.5% (10.0% on average).
Kun Fang 0005, Hongzhong Zheng, Zhichun Zhu
ICPP3
2010 Power and Performance Trade-Offs in Contemporary DRAM System Designs for Multicore Processors
abstract
DRAM memory is playing an increasingly important role in the overall power profile of latest-generation servers with multicore processors. With many power saving techniques adopted into processor design, memory power consumption can now exceed processor power consumption when a system runs memory-intensive workloads. There is an urgent need to fully evaluate the memory power profile of contemporary DRAM memories and to re-investigate DRAM memory designs, configurations, and optimizations from both power and performance perspectives. This study fills the gap by studying the performance and power consumption of multicore systems with DDR3 memory under different configurations. It includes comprehensive results regarding memory power breakdown, including background, operation, read/write, and I/O power, as well as performance. Comparisons with DDR2 and FB-DIMM are also included. The results show clearly that DRAM system configurations, including page policy, power mode, device configuration, burst length, channel organization, and the selection of DRAM technology, affects the memory power consumption significantly besides the performance. The optimal choice of some configurations is application-dependent, suggesting that reconfigurable or hybrid configurations are worth further studies.
Hongzhong Zheng, Zhichun Zhu
IEEE Trans. Computers2
2009 Decoupled DIMM: building high-bandwidth memory system using low-speed DRAM devices
abstract
The widespread use of multicore processors has dramatically increased the demands on high bandwidth and large capacity from memory systems. In a conventional DDR2/DDR3 DRAM memory system, the memory bus and DRAM devices run at the same data rate. To improve memory bandwidth, we propose a new memory system design called decoupled DIMM that allows the memory bus to operate at a data rate much higher than that of the DRAM devices. In the design, a synchronization buffer is added to relay data between the slow DRAM devices and the fast memory bus; and memory access scheduling is revised to avoid access conflicts on memory ranks. The design not only improves memory bandwidth beyond what can be supported by current memory devices, but also improves reliability, power efficiency, and cost effectiveness by using relatively slow memory devices. The idea of decoupling, precisely the decoupling of bandwidth match between memory bus and a single rank of devices, can also be applied to other types of memory systems including FB-DIMM.
Hongzhong Zheng, Jiang Lin, Zhao Zhang 0010, Zhichun Zhu
ISCA4
2008 Memory Access Scheduling Schemes for Systems with Multi-Core Processors
abstract
On systems with multi-core processors, the memory access scheduling scheme plays an important role not only in utilizing the limited memory bandwidth but also in balancing the program execution on all cores. In this study, we propose a scheme, called ME-LREQ, which considers the utilization of both processor cores and memory subsystem. It takes into consideration both the long-term and short-term gains of serving a memory request by prioritizing requests hitting on the row buffers and from the cores that can utilize memory more efficiently and have fewer pending requests. We have also thoroughly evaluated a set of memory scheduling schemes that differentiate and prioritize requests from different cores. Our simulation results show that for memory-intensive, multiprogramming workloads, the new policy improves the overall performance by 10.7% on average and up to 17.7% on a four-core processor, when compared with scheme that serves row buffers hit memory requests first and allows memory reads bypassing writes; and by up to 9.2% (6.4% on average) when compared with the scheme that serves requests from the core with the fewest pending requests first.
Hongzhong Zheng, Jiang Lin, Zhao Zhang 0010, Zhichun Zhu
ICPP4
2008 Mini-rank: Adaptive DRAM architecture for improving memory power efficiency
abstract
The widespread use of multicore processors has dramatically increased the demand on high memory bandwidth and large memory capacity. As DRAM subsystem designs stretch to meet the demand, memory power consumption is now approaching that of processors. However, the conventional DRAM architecture prevents any meaningful power and performance trade-offs for memory-intensive workloads. We propose a novel idea called mini-rank for DDRx (DDR/DDR2/DDR3) DRAMs, which uses a small bridge chip on each DRAM DIMM to break a conventional DRAM rank into multiple smaller mini-ranks so as to reduce the number of devices involved in a single memory access. The design dramatically reduces the memory power consumption with only a slight increase on the memory idle latency. It does not change the DDRx bus protocol and its configuration can be adapted for the best performance-power trade-offs. Our experimental results using four-core multiprogramming workloads show that using x32 mini-ranks reduces memory power by 27.0% with 2.8% performance penalty and using x16 mini-ranks reduces memory power by 44.1% with 7.4% performance penalty on average for memory-intensive workloads, respectively.
Hongzhong Zheng, Jiang Lin, Zhao Zhang 0010, Eugene Gorbatov, Howard David, Zhichun Zhu
MICRO6
2008 Software thermal management of dram memory for multicore systems
abstract
Thermal management of DRAM memory has become a critical issue for server systems. We have done, to our best knowledge, the first study of software thermal management for memory subsystem on real machines. Two recently proposed DTM (Dynamic Thermal Management) policies have been improved and implemented in Linux OS and evaluated on two multicore servers, a Dell PowerEdge 1950 server and a customized Intel SR1500AL server testbed. The experimental results first confirm that a system-level memory DTM policy may significantly improve system performance and power efficiency, compared with existing memory bandwidth throttling scheme. A policy called DTM-ACG (Adaptive Core Gating) shows performance improvement comparable to that reported previously. The average performance improvements are 13.3% and 7.2% on the PowerEdge 1950 and the SR1500AL (vs. 16.3% from the previous simulation-based study), respectively. We also have surprising findings that reveal the weakness of the previous study: the CPU heat dissipation and its impact on DRAM memories, which were ignored, are significant factors. We have observed that the second policy, called DTM-CDVFS (Coordinated Dynamic Voltage and Frequency Scaling), has much better performance than previously reported for this reason. The average improvements are 10.8% and 15.3% on the two machines (vs. 3.4% from the previous study), respectively. It also significantly reduces the processor power by 15.5% and energy by 22.7% on average.
Jiang Lin, Hongzhong Zheng, Zhichun Zhu, Eugene Gorbatov, Howard David, Zhao Zhang 0010
SIGMETRICS3
2007 Thermal modeling and management of DRAM memory systems
abstract
With increasing speed and power density, high-performance memories, including FB-DIMM (Fully Buffered DIMM) and DDR2 DRAM, now begin to require dynamic thermal management(DTM) as processors and hard drives did. The DTM of memories, nevertheless, is different in that it should take the processor performance and power consumption into consideration. Existing schemes have ignored that. In this study, we investigate a new approach that controls the memory thermal issues from the source generating memory activities - the processor. It will smooth the program execution when compared with shutting down memory abruptly, and therefore improve the overall system performance and power efficiency. For multicore systems, we propose two schemes called adaptive core gating and coordinated DVFS. The first scheme activates clock gating on selected processor cores and the second one scales down the frequency and voltage levels of processor cores when the memory is to be over-heated. They can successfully control the memory activities and handle thermal emergency. More importantly, they improve performance significantly under the given thermal envelope. Our simulation results show that adaptive coregating improves performance by up to 23.3% (16.3% on average) on a four-core system with FB-DIMM when compared with DRAM thermal shutdown; and coordinated DVFS with control-theoretic methods improves the performance by up to 18.5% (8.3% on average).
Jiang Lin, Hongzhong Zheng, Zhichun Zhu, Howard David, Zhao Zhang 0010
ISCA3
2007 DRAM-Level Prefetching for Fully-Buffered DIMM: Design, Performance and Power Saving
abstract
We have studied DRAM-level prefetching for the fully buffered DIMM (FB-DIMM) designed for multi-core processors. FB-DIMM has a unique two-level interconnect structure, with FB-DIMM channels at the first-level connecting the memory controller and Advanced Memory Buffers (AMBs); and DDR2 buses at the second-level connecting the AMBs with DRAM chips. We propose an AMB prefetching method that prefetches memory blocks from DRAM chips to AMBs. It utilizes the redundant bandwidth between the DRAM chips and AMBs but does not consume the crucial channel bandwidth. The proposed method fetches K memory blocks of L2 cache block sizes around the demanded block, where K is a small value ranging from two to eight. The method may also reduce the DRAM power consumption by merging some DRAM precharges and activations. Our cycle-accurate simulation shows that the average performance improvement is 16% for single-core and multi-core workloads constructed from memory-intensive SPEC2000 programs with software cache prefetching enabled; and no workload has negative speedup. We have found that the performance gain comes from the reduction of idle memory latency and the improvement of channel bandwidth utilization. We have also found that there is only a small overlap between the performance gains from the AMB prefetching and the software cache prefetching. The average of estimated power saving is 15%.
Jiang Lin, Hongzhong Zheng, Zhichun Zhu, Zhao Zhang 0010, Howard David
ISPASS3
2005 A Performance Comparison of DRAM Memory System Optimizations for SMT Processors
abstract
Memory system optimizations have been well studied on single-threaded systems; however, the wide use of simultaneous multithreading (SMT) techniques raises questions over their effectiveness in the new context. In this study, we thoroughly evaluate contemporary multi-channel DDR SDRAM and Rambus DRAM systems in SMT systems, and search for new thread-aware DRAM optimization techniques. Our major findings are: (1) in general, increasing the number of threads tends to increase the memory concurrency and thus the pressure on DRAM systems, but some exceptions do exist; (2) the application performance is sensitive to memory channel organizations, e.g. independent channels may outperform ganged organizations by up to 90%; (3) the DRAM latency reduction through improving row buffer hit rates becomes less effective due to the increased bank contentions; and (4) thread-aware DRAM access scheduling schemes may improve performance by up to 30% on workload mixes of memory-intensive applications. In short, the use of SMT techniques has somewhat changed the context of DRAM optimizations but does not make them obsolete.
Zhichun Zhu, Zhao Zhang 0010
HPCA1
2004 Design and Optimization of Large Size and Low Overhead Off-Chip Caches
abstract
Large off-chip L3 caches can significantly improve the performance of memory-intensive applications. However, conventional L3 SRAM caches are facing two issues as those applications require increasingly large caches. First, an SRAM cache has a limited size due to the low density and high cost of SRAM and, thus, cannot hold the working sets of many memory-intensive applications. Second, since the tag checking overhead of large caches is nontrivial, the existence of L3 caches increases the cache miss penalty and may even harm the performance of some memory-intensive applications. To address these two issues, we present a new memory hierarchy design that uses cached DRAM to construct a large size and low overhead off-chip cache. The high density DRAM portion in the cached DRAM can hold large working sets, while the small SRAM portion exploits the spatial locality appearing in L2 miss streams to reduce the access latency. The L3 tag array is placed off-chip with the data array, minimizing the area overhead on the processor for L3 cache, while a small tag cache is placed on-chip, effectively removing the off-chip tag access overhead. A prediction technique accurately predicts the hit/miss status of an access to the cached DRAM, further reducing the access latency. Conducting execution-driven simulations for a 2 GHz 4-way issue processor and with 11 memory-intensive programs from the SPEC 2000 benchmark, we show that a system with a cached DRAM of 64 MB DRAM and 128 KB on-chip SRAM cache as the off-chip cache outperforms the same system with an 8 MB SRAM L3 off-chip cache by up to 78 percent measured by the total execution time. The average speedup of the system with the cached-DRAM off-chip cache is 25 percent over the system with the L3 SRAM cache.
Zhao Zhang 0010, Zhichun Zhu, Xiaodong Zhang 0001
IEEE Trans. Computers2
2002 Fine-Grain Priority Scheduling on Multi-Channel Memory Systems
abstract
Configurations of contemporary DRAM memory systems become increasingly complex. A recent study shows that the application performance is highly sensitive to choices of configurations. In this study we show that, by utilizing fine-grain priority access scheduling, we are able to find a workload independent configuration that achieves optimal performance on a multichannel memory system. Our approach can well utilize the available high concurrency and high bandwidth on such memory systems, and effectively reduce the memory stall time of memory-intensive applications. Conducting execution-driven simulation of a 4-way issue, a 2 GHz processor, we show that the average performance improvement for fifteen memory-intensive SPEC2000 programs by using an optimized fine-grain priority scheduling is about 13% and 8% for a 2-channel and a 4-channel Direct Rambus DRAM memory system, respectively, compared with gang scheduling. Compared with burst scheduling, the average performance improvement is 16% and 14% for the 2-channel and 4-channel memory systems, respectively.
Zhichun Zhu, Zhao Zhang 0010, Xiaodong Zhang 0001
HPCA1
2000 A permutation-based page interleaving scheme to reduce row-buffer conflicts and exploit data locality
abstract
DRAM row-buffer conflicts occur when a sequence of requests on different rows goes to the same memory bank, causing much higher memory access latency than requests to the same row or to different banks. We analyze the sources of row-buffer conflicts in the context of superscalar processors, and propose a permutation based page interleaving scheme to reduce row-buffer conflicts and to exploit data access locality in the row-buffer. Compared with several existing schemes, we show that the permutation based scheme dramatically increases the hit rates on DRAM row-buffers and reduces memory stall time of the SPEC95 and TPC-C workloads. The memory stall times of the workloads are reduced up to 68% and 50%, compared with the conventional cache line and page interleaving schemes, respectively.
Zhao Zhang 0010, Zhichun Zhu, Xiaodong Zhang 0001
MICRO2
2000 Memory Hierarchy Considerations for Cost-Effective Cluster Computing
abstract
Using off-the-shelf commodity workstations and PCs to build a cluster for parallel computing has become a common practice. The cost-effectiveness of a cluster computing platform for a given budget and for certain types of applications is mainly determined by its memory hierarchy and the interconnection network configurations of the cluster. Finding such a cost-effective solution from exhaustive simulations would be highly time-consuming and predictions from measurements on existing clusters would be impractical. We present an analytical model for evaluating the performance impact of memory hierarchies and networks on cluster computing. The model covers the memory hierarchy of a single SMP, a cluster of workstations/PCs, or a cluster of SMPs by changing various architectural parameters. Network variations covering both bus and switch networks are also included in the analysis. Different types of applications are characterized by parameterized workloads with different computation and communication requirements. The model has been validated by simulations and measurements. The workloads used for experiments are both scientific applications and commercial workloads. Our study shows that the depth of the memory hierarchy is the most sensitive factor affecting the execution time for many types of workloads. However, the interconnection network cost of a tightly coupled system with a short depth in memory hierarchy, such as an SMP, is significantly more expensive than a normal cluster network connecting independent computer nodes. Thus, the essential issue to be considered is the trade-off between the depth of the memory hierarchy and the system cost. Based on analyses and case studies, we present our quantitative recommendations for building cost-effective clusters for different workloads.
Xiaodong Zhang 0001, Zhichun Zhu
IEEE Trans. Computers3