EDBT 2026 Demo / reviewers in the wild / expert
Mengjie Mao
dblp:09/10802
· DBLP profile ↗
22ranked-venue papers
7as first author
2since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 20 · 7 first-author · 1 since 2021Software engineering, systems software and programming languages · 2Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
7 papers |
Memory systems · 52% GPUs and heterogeneous computing · 17% Hardware reliability and fault tolerance · 7% |
Topics — the 25 heaviest of 27, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
GPUs and heterogeneous computing
GPU architecture |
0.8 | 3 | 2017 | An Energy-Efficient GPGPU Register File Architecture Using Racetrack Memory · IEEE Trans. Computers 2017 TEMP: thread batch enabled memory partitioning for GPU · DAC 2016 VWS: a versatile warp scheduler for exploring diverse cache localities of GPGPU applications · DAC 2015 |
Memory systems
non-volatile memory |
0.7 | 3 | 2018 | TriZone: A Design of MLC STT-RAM Cache for Combined Performance, Energy, and Reliability Optimizations · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018 State-Restrict MLC STT-RAM Designs for High-Reliable High-Performance Memory System · DAC 2014 Exploration of GPGPU Register File Architecture Using Domain-wall-shift-write based Racetrack Memory · DAC 2014 |
Memory systems › cache
STT-RAM cache |
0.4 | 2 | 2018 | TriZone: A Design of MLC STT-RAM Cache for Combined Performance, Energy, and Reliability Optimizations · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018 State-Restrict MLC STT-RAM Designs for High-Reliable High-Performance Memory System · DAC 2014 |
Memory systems
cache |
0.3 | 1 | 2018 | TriZone: A Design of MLC STT-RAM Cache for Combined Performance, Energy, and Reliability Optimizations · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018 |
Memory systems › cache › STT-RAM cache
MLC STT-RAM cache |
0.3 | 1 | 2018 | TriZone: A Design of MLC STT-RAM Cache for Combined Performance, Energy, and Reliability Optimizations · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018 |
Memory systems › non-volatile memory
multi-level cell |
0.3 | 1 | 2018 | TriZone: A Design of MLC STT-RAM Cache for Combined Performance, Energy, and Reliability Optimizations · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018 |
Hardware reliability and fault tolerance
error correction |
0.3 | 2 | 2018 | State-Restrict MLC STT-RAM Designs for High-Reliable High-Performance Memory System · DAC 2014 TriZone: A Design of MLC STT-RAM Cache for Combined Performance, Energy, and Reliability Optimizations · IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. 2018 |
Memory systems
emerging memory technologies |
0.3 | 1 | 2017 | An Energy-Efficient GPGPU Register File Architecture Using Racetrack Memory · IEEE Trans. Computers 2017 |
Memory systems › emerging memory technologies › spintronic memory
racetrack memory |
0.3 | 1 | 2017 | An Energy-Efficient GPGPU Register File Architecture Using Racetrack Memory · IEEE Trans. Computers 2017 |
Processor architecture and microarchitecture › register file
register file design |
0.3 | 1 | 2017 | An Energy-Efficient GPGPU Register File Architecture Using Racetrack Memory · IEEE Trans. Computers 2017 |
Electronic design automation › high-level synthesis › memory synthesis
memory partitioning |
0.2 | 1 | 2016 | TEMP: thread batch enabled memory partitioning for GPU · DAC 2016 |
Memory systems › memory access
parallel memory access |
0.2 | 1 | 2016 | TEMP: thread batch enabled memory partitioning for GPU · DAC 2016 |
Memory systems › data locality
cache locality |
0.2 | 1 | 2015 | VWS: a versatile warp scheduler for exploring diverse cache localities of GPGPU applications · DAC 2015 |
Memory systems
cache management |
0.2 | 1 | 2015 | VWS: a versatile warp scheduler for exploring diverse cache localities of GPGPU applications · DAC 2015 |
Hardware accelerators and domain-specific architectures › machine learning accelerator › in-memory computing accelerator
memristive crossbar accelerator |
0.2 | 1 | 2015 | RENO: a high-efficient reconfigurable neuromorphic computing accelerator design · DAC 2015 |
Hardware accelerators and domain-specific architectures › machine learning accelerator
neural network accelerator |
0.2 | 1 | 2015 | RENO: a high-efficient reconfigurable neuromorphic computing accelerator design · DAC 2015 |
Emerging computing paradigms
neuromorphic computing |
0.2 | 1 | 2015 | RENO: a high-efficient reconfigurable neuromorphic computing accelerator design · DAC 2015 |
Reconfigurable computing and FPGAs › reconfigurable computing
reconfigurable accelerator |
0.2 | 1 | 2015 | RENO: a high-efficient reconfigurable neuromorphic computing accelerator design · DAC 2015 |
GPUs and heterogeneous computing › GPU scheduling
warp scheduling |
0.2 | 1 | 2015 | VWS: a versatile warp scheduler for exploring diverse cache localities of GPGPU applications · DAC 2015 |
Hardware reliability and fault tolerance
memory reliability |
0.2 | 1 | 2014 | State-Restrict MLC STT-RAM Designs for High-Reliable High-Performance Memory System · DAC 2014 |
Memory systems › non-volatile memory › magnetic random access memory › STT-MRAM
MLC STT-RAM |
0.2 | 1 | 2014 | State-Restrict MLC STT-RAM Designs for High-Reliable High-Performance Memory System · DAC 2014 |
Energy-efficient computing › memory energy efficiency
register file energy reduction |
0.1 | 2 | 2017 | An Energy-Efficient GPGPU Register File Architecture Using Racetrack Memory · IEEE Trans. Computers 2017 Exploration of GPGPU Register File Architecture Using Domain-wall-shift-write based Racetrack Memory · DAC 2014 |
Energy-efficient computing › energy-efficient architecture
GPGPU energy efficiency |
0.1 | 1 | 2017 | An Energy-Efficient GPGPU Register File Architecture Using Racetrack Memory · IEEE Trans. Computers 2017 |
Reconfigurable computing and FPGAs
reconfigurable architecture |
0.1 | 1 | 2015 | RENO: a high-efficient reconfigurable neuromorphic computing accelerator design · DAC 2015 |
Memory systems
cache design |
0.1 | 1 | 2014 | State-Restrict MLC STT-RAM Designs for High-Reliable High-Performance Memory System · DAC 2014 |
Methods — techniques the papers use, named apart from their topics
nonuniform strength ECC · 0.3dynamic cache partitioning · 0.3write buffer design · 0.3warp scheduling · 0.3register mapping · 0.3thread batch scheduling · 0.2OS memory management · 0.2mixed-signal computation · 0.2memristor crossbar · 0.2CTA-aware scheduling · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DPFT: A Framework for Decomposed Probabilistic Forecasting with Transformers
Yifang Wang, Boliang Yuan, Mengjie Mao, Yiheng Zhong, Xinkun Wang, Senwei Liang |
ICIC (26) | 3 |
| 2021 | Exploring Applications of STT-RAM in GPU ArchitecturesabstractUse of modern GPUs has been extended from traditional 3D graphic processing to computing acceleration of many scientific, engineering, and enterprise applications. In modern GPUs, on-chip memory capacity keeps increasing to support thousands of chip-resident threads. For example, a large register file is needed in order to efficiently process highly-parallel threads in single instruction multiple thread (SIMT) fashion, and a large shared memory is often implemented to allow data sharing among the threads on the chip. On-chip memory capacity of GPUs, however, is highly constrained by large memory cell area and high static power consumption of conventional SRAM implementation. In this work, we propose to utilize the emerging multi-level cell (MLC) spin-transfer torque RAM (STT-RAM) technology to implement register file and shared memory in GPUs. Compared to SRAM, MLC STT-RAM (or MLC-STT) has a much smaller cell area as well as ultra-low standby power, thanks to the non-volatility of MLC-STT technology. Hence, the footprint and leakage power of the implemented memory components are substantially reduced. Moreover, in light of asymmetric performance of soft and hard bits of a MLC-STT cell, we propose a dynamic data remapping strategy in register file and shared memory implementations that allows a flexible tradeoff between the memory access time and the available capacity: frequently-accessed data is always mapped to the fast rows built with the soft bits of the MLC-STT cells while the slow rows composed of the hard bits are used only when a larger capacity is critically needed. We also develop a novel rescheduling scheme to minimize the waiting time of the issued warps to access register banks in the register file, which is induced by the long writeback operations through the reordering of the issued warps. Finally, an early termination technology is also applied to save the write energy of the shared memory if the bits of the memory do not flip. Experimental results on benchmarks of ISPASS2009, Rodinia, Parboil, and CUDA show that on average, MLC-STT register file can achieve 3.28% system performance improvement, 9.48% energy reduction, and 38.9% energy efficiency improvement compared to conventional SRAM-based design. Meanwhile, MLC-STT shared memory leads to 3.45% system performance improvement, 49.3% energy reduction, and 116% energy efficiency improvement. Xiaoxiao Liu 0001, Mengjie Mao, Xiuyuan Bi, Hai Li 0001, Yiran Chen 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2019 | Thread Batching for High-performance Energy-efficient GPU Memory DesignabstractMassive multi-threading in GPU imposes tremendous pressure on memory subsystems. Due to rapid growth in thread-level parallelism of GPU and slowly improved peak memory bandwidth, memory becomes a bottleneck of GPU’s performance and energy efficiency. In this article, we propose an integrated architectural scheme to optimize the memory accesses and therefore boost the performance and energy efficiency of GPU. First, we propose a thread batch enabled memory partitioning (TEMP) to improve GPU memory access parallelism. In particular, TEMP groups multiple thread blocks that share the same set of pages into a thread batch and applies a page coloring mechanism to bound each stream multiprocessor (SM) to the dedicated memory banks. After that, TEMP dispatches the thread batch to an SM to ensure high-parallel memory-access streaming from the different thread blocks. Second, a thread batch-aware scheduling (TBAS) scheme is introduced to improve the GPU memory access locality and to reduce the contention on memory controllers and interconnection networks. Experimental results show that the integration of TEMP and TBAS can achieve up to 10.3% performance improvement and 11.3% DRAM energy reduction across diverse GPU applications. We also evaluate the performance interference of the mixed CPU+GPU workloads when they are run on a heterogeneous system that employs our proposed schemes. Our results show that a simple solution can effectively ensure the efficient execution of both GPU and CPU applications. Bing Li 0017, Mengjie Mao, Xiaoxiao Liu 0001, Tao Liu 0023, Zihao Liu 0015, Wujie Wen, Yiran Chen 0001, Hai Li 0001 |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2018 | TriZone: A Design of MLC STT-RAM Cache for Combined Performance, Energy, and Reliability OptimizationsabstractSpin-transfer torque random access memory (STT-RAM) is a promising technology for future nonvolatile caches and memories. To increase the storage density, multilevel cell (MLC) technique was recently introduced to STT-RAM designs at the cost of degraded access speed, reliability, and energy efficiency. Existing MLC STT-RAM cache architectures primarily focus on the performance and energy optimizations but ignore the crucial demand for reliability. In this paper, we propose “TriZone”-a holistic design scheme for MLC STT-RAM cache to simultaneously meet the requirements of performance, energy, and reliability. Three cache block configurations, namely hard, soft, and mixed, are constructed with the hard-bit, soft-bit, and both hard-bit and soft-bit of MLC STT-RAM, respectively. By observing the difference of these cache blocks, a nonuniform strength ECC (NUS-ECC) is developed to guarantee the operational reliability of a cache block with a variable decoding delay adapting to the needs of error correction (e.g., the number of the erroneous bits). The whole MLC STT-RAM cache is then partitioned into three regions, each of which is composed of different cache blocks. In order to achieve the best tradeoff among performance, energy, and reliability, we then introduce the dynamic cache partitioning to determine the partition of this tri-way MLC STT-RAM cache according to the runtime characteristic of various applications. Experiment results show that compared with conventional performance-driven MLC STT-RAM cache design with pessimistic ECC, TriZone can improve the system performance and energy by averagely 11.7% (10.0%) and 13.3% (15.7%), respectively, for single-threaded (multiprogram) applications. The additional area overhead associated with NUS-ECC is limited by ~ 3%. Zihao Liu 0015, Mengjie Mao, Tao Liu 0023, Wujie Wen, Yiran Chen 0001, Hai Li 0001, Danghui Wang, Yukui Pei, Ning Ge 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2017 | An Energy-Efficient GPGPU Register File Architecture Using Racetrack MemoryabstractExtreme multi-threading and fast thread switching in modern GPGPU require a large, power-hungry register file (RF), which quickly becomes one of major obstacles on the upscaling path of energy-efficient GPGPU computing. In this work, we propose to implement a power-efficient GPGPU RF built on the newly emerged racetrack memory. Racetrack memory has small cell area, low dynamic power, and nonvolatility. Its unique access mechanism, however, results in a long and location-dependent access latency, which offsets the energy saving benefit it introduces and probably harms the performance. In order to conquer the adverse impacts of racetrack memory based RF designs, we first propose a register mapping scheme to reduce the average access latency. Based on the register mapping, we develop a racetrack memory aware warp scheduling (RMWS) algorithm to further suppress the access latency. RMWS design includes a new write buffer structure that improves the scheduling efficiency as well as energy saving. We also investigate and optimize the design where multiple concurrent RMWS schedulers are employed. Experiment results show that our propose techniques can keep a GPGPU performance similar to the baseline with SRAM based RF while the RF energy is significantly reduced by 48.5 percent. Mengjie Mao, Wujie Wen, Yaojun Zhang, Yiran Chen 0001, Hai Li 0001 |
IEEE Trans. Computers | 1 |
| 2017 | Cross-Layer Optimization for Multilevel Cell STT-RAM CachesabstractSpin-transfer torque random access memory (STT-RAM), as an emerging nonvolatile memory technology, provides very dense array structure and extremely low leakage power consumption. It demonstrates a great potential in replacing conventional static random access memory technology to develop the next-generation on-chip cache memory of microprocessors and graphics processing units. The multilevel cell (MLC) design of STT-RAM that stores two or more bits in one cell potentially has higher storage capacity and faster system performance, attracting significant attention. In this paper, we first quantitatively evaluated the data storage density of the MLC STT-RAM. Our results revealed limited density improvement because of the large size of access transistor induced by high write current amplitude requirement and asymmetry of switching behavior. Moreover, the read and write accesses of existing MLC STT-RAM cache designs require two-step operation. The system level evaluation shows that the long access latency could amortize the performance speed brought by larger cache size, and even degrade the system performance for some applications. To unleash the potential of MLC STT-RAM cache, we proposed a new design through a cross-layer co-optimization. The memory cell structure integrated the reversed stacking of magnetic junction tunneling for a more balanced device and design tradeoff. In architecture development, we presented an adaptive mode switching mechanism: based on application's memory access behavior, the MLC STT-RAM cache can dynamically change between low latency single-level cell mode and high capacity MLC mode. Furthermore, we divided cache lines into fast and slow regions and investigated new data migration policies to allocate frequently access data to fast regions. Simulation results show that the proposed techniques can improve the system performance by 10.2% and reduce the energy consumption on cache by 9.5% compared with conventional MLC STT-RAM cache design. Xiuyuan Bi, Mengjie Mao, Danghui Wang, Hai Li 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2016 | TEMP: thread batch enabled memory partitioning for GPUabstractAs massive multi-threading in GPU imposes tremendous pressure on memory subsystems, efficient bandwidth utilization becomes a key factor affecting the GPU throughput. In this work, we propose thread batch enabled memory partitioning (TEMP), to improve GPU performance through the improvement of memory bandwidth utilization. In particular, TEMP clusters multiple thread blocks sharing the same set of pages into a thread batch and dispatches the entire thread batch to a stream multiprocessor. TEMP separates the memory access streams of different thread batches by OS memory management, preserving the intrinsic locality of thread batches and increasing the memory access parallelism. Experimental results show that TEMP can obtain up to 10.3% performance improvement and 14.6% DRAM energy reduction compared to a state-of-the-art scheduler without any memory-side optimizations. Mengjie Mao, Wujie Wen, Xiaoxiao Liu 0001, Jingtong Hu, Danghui Wang, Yiran Chen 0001, Hai Li 0001 |
DAC | 1 |
| 2016 | Sliding Basket: An adaptive ECC scheme for runtime write failure suppression of STT-RAM cache
Mengjie Mao, Enes Eken, Wujie Wen, Hai Li 0001, Yiran Chen 0001 |
DATE | 2 |
| 2016 | A holistic tri-region MLC STT-RAM design with combined performance, energy, and reliability optimizations
Wujie Wen, Mengjie Mao, Hai Li 0001, Yiran Chen 0001, Yukui Pei, Ning Ge 0001 |
DATE | 2 |
| 2016 | Heterogeneous systems with reconfigurable neuromorphic computing acceleratorsabstractDeveloping heterogeneous system with hardware accelerator is a promising solution to implement high performance applications where explicitly programmed, rule-based algorithms are either infeasible or inefficient. However, mapping a neural network model to a hardware representation is a complex process, where balancing computation resources and memory accesses is crucial. In this work, we present a systematic approach o optimize the heterogeneous system with a FPGA-based neuromorphic computing accelerator (NCA). For any applications, the neural network topology and computation flow of the accelerator can be configured through a NCA-aware compiler. The FPGA-based NCA contains a generic multi-layer neural network composed of a set of parallel neural processing elements. Such a scheme imitates the human cognition process and follows the hierarchy of neocortex. At architectural level, we decrease the computing resource requirement to enhance computation efficiency. The hardware implementation primarily targets at reducing data communication load: a multi-thread computation engine is utilized to mask the long memory latency. Such a combined solution can well accommodate the ever increasing complexity and scalability of machine learning applications and improve the system performance and efficiency. Through the evaluation across eight representative benchmarks, we observed on average 12.1× speedup and 45.8× energy reduction, with marginal accuracy loss comparing with CPU-only computation. Sicheng Li 0001, Xiaoxiao Liu 0001, Mengjie Mao, Hai Li 0001, Yiran Chen 0001, Boxun Li, Yu Wang 0002 |
ISCAS | 3 |
| 2015 | An efficient STT-RAM-based register file in GPU architecturesabstractModern GPGPUs employ a large register file (RF) to efficiently process heavily parallel threads in single instruction multiple thread (SIMT) fashion. The up-scaling of RF capacity, however, is greatly constrained by large cell area and high leakage power consumption of SRAM implementation. In this work, we propose a novel GPU RF design based on the emerging multi-level cell (MLC) spin-transfer torque RAM (STT-RAM) technology. Compared to SRAM, MLC STT-RAM (or MLC-STT) has much smaller cell area and almost zero standby power due to its non-volatility. Moreover, by leveraging the asymmetric performance of the soft and the hard bits of a MLC-STT cell, we propose a remapping strategy to perform a flexible tradeoff between the access time and the capacity of the RF based on run-time access patterns. A novel rescheduling scheme is also developed to minimize the waiting time of the issued warps to access register banks. Experimental results over ISPASS2009 and CUDA benchmarks show that on average, our proposed MLC-STT RF can achieve 3.28% performance improvement, 9.48% energy reduction, and 38.9% energy efficiency enhancement compared to conventional SRAM-based design. Xiaoxiao Liu 0001, Mengjie Mao, Xiuyuan Bi, Hai Li 0001, Yiran Chen 0001 |
ASP-DAC | 2 |
| 2015 | RENO: a high-efficient reconfigurable neuromorphic computing accelerator designabstractNeuromorphic computing is recently gaining significant attention as a promising candidate to conquer the well-known von Neumann bottleneck. In this work, we propose RENO -- a efficient reconfigurable neuromorphic computing accelerator. RENO leverages the extremely efficient mixed-signal computation capability of memristor-based crossbar (MBC) arrays to speedup the executions of artificial neural networks (ANNs). The hierarchically arranged MBC arrays can be configured to a variety of ANN topologies through a mixed-signal interconnection network (M-Net). Simulation results on seven ANN applications show that compared to the baseline general-purpose processor, RENO can achieve on average 178.4x (27.06x) performance speedup and 184.2x (25.23x) energy savings in high-efficient multilayer perception (high-accurate auto-associative memory) implementation. Moreover, in the comparison to a pure digital neural processing unit (D-NPU) and a design with MBC arrays co-operating through a digital interconnection network, RENO still achieves the fastest execution time and the lowest energy consumption with similar computation accuracy. Xiaoxiao Liu 0001, Mengjie Mao, Beiye Liu, Hai Li 0001, Yiran Chen 0001, Boxun Li, Yu Wang 0002, Hao Jiang 0014, Mark Barnell, Qing Wu 0002, J. Joshua Yang |
DAC | 2 |
| 2015 | VWS: a versatile warp scheduler for exploring diverse cache localities of GPGPU applicationsabstractMassive multi-threading of GPGPU demands for efficient usage of caches with limited capacity. In this work, we propose a versatile warp scheduler (VWS) to reduce the cache miss rate in GPGPU. VWS retains the intra-warp cache locality using an efficient per-warp working set estimator and enhances intra-/inter-cooperative thread array (CTA) cache locality through imposing a CTA-aware scheduling policy and a new CTA dispatching mechanism. The significantly improved hit rate of cache hierarchy enables VWS to achieve on average 38.4% and 9.3% IPC improvement across diverse GPGPU applications compared to a widely-used and a state-of-the-art warp schedulers, respectively. Mengjie Mao, Jingtong Hu, Yiran Chen 0001, Hai Li 0001 |
DAC | 1 |
| 2015 | The applications of memristor devices in next-generation cortical processor designsabstractDiscovery of memristor opened a new era of the research on universal memory thanks to many attractive properties demonstrated by this emerging device. In this paper, we switch our research focus to neuromorphic computing, which, same as memory technology, significantly benefits from the technical advances of memristor. Particularly, we present the implementation of cortical processor augmented with neuromorphic computing accelerators (NCAs) for cognitive applications, including: 1) the design details and basic operations of the NCA based on memristor crossbars; and 2) the integration between conventional pipeline and NCAs. At the end, we also discuss the scalability of our proposed NCA designs. Hai Li 0001, Beiye Liu, Xiaoxiao Liu 0001, Mengjie Mao, Yiran Chen 0001, Qing Wu 0002, Qinru Qiu |
ISCAS | 4 |
| 2014 | Prefetching techniques for STT-RAM based last-level cache in CMP systemsabstractPrefetching is widely used in modern computer systems to mitigate the impact of long memory access latency by paying extra cost in memory and cache accesses. However, the efficacy of prefetching significantly degrades in the memory hierarchy using the emerging spin-transfer torque random access memory (STT-RAM) as last-level cache (LLC) due to the long write access latency. In this work, we propose two orthogonal but complimentary techniques to improve the prefetching efficacy of STT-RAM based LLC in chip multi-processor (CMP) systems, namely, request prioritization (RP) and hybrid local-global prefetch control (HLGPC). Simulation results show that by combining these two techniques, we can achieve 6.5%~11% system performance improvement and 4.8%~7.3% LLC energy saving in a quadcore system with a 2MB~8MB STT-RAM based LLC, compared to the system with only basic prefetching. Mengjie Mao, Guangyu Sun 0003, Yong Li 0009, Alex K. Jones, Yiran Chen 0001 |
ASP-DAC | 1 |
| 2014 | Exploration of GPGPU Register File Architecture Using Domain-wall-shift-write based Racetrack MemoryabstractSRAM based register file (RF) is one of the major factors limiting the scaling of GPGPU. In this work, we propose to use the emerging nonvolatile domain-wall-shift-write based racetrack memory (DWSW-RM) to implement a power-efficient GPGPU RF, of which the power consumption is substantially reduced. A holistic technology set is developed to minimize the high access cost of DWSW-RW caused by the sequential access mechanism. Experiment results show that our proposed techniques can improve the GPGPU performance by 4.6% compared to the baseline with SRAM based RF. The RF energy efficiency is also significantly improved by 2.45×. Mengjie Mao, Wujie Wen, Yaojun Zhang, Yiran Chen 0001, Hai Li 0001 |
DAC | 1 |
| 2014 | State-Restrict MLC STT-RAM Designs for High-Reliable High-Performance Memory SystemabstractMulti-level Cell Spin-Transfer Torque Random Access Memory (MLC STT-RAM) is a promising nonvolatile memory technology for high-capacity and high-performance applications. However, the reliability concerns and the complicated access mechanism greatly hinder the application of MLC STT-RAM. In this work, we develop a holistic solution set, namely, state-restrict MLC STT-RAM (SR-MLC STT-RAM) to improve the data integrity and performance of MLC STT-RAM with the minimized information density degradation. Three techniques: state restriction (StatRes), error pattern removal (ErrPR), and ternary coding (TerCode) are proposed at circuit level to reduce the read and write errors of MLC STT-RAM cells. State pre-recovery (PreREC) technique is also developed at architecture level to improve the access performance of SR-MLC STT-RAM by eliminating unnecessary two-step write operations. Our simulations show that compared to conventional MLC STT-RAM, SR-MLC STT-RAM can enhance the write and read reliability of memory cells by 10--10000×, allowing the application of simple error correction code schemes. Compared to single-level-cell (SLC) STT-RAM, SR-MLC STT-RAM based cache design can boost the system performance by 6.2% on average by leveraging the increased cache capacity at the same area and the improved write latency. Wujie Wen, Yaojun Zhang, Mengjie Mao, Yiran Chen 0001 |
DAC | 3 |
| 2013 | Coordinating prefetching and STT-RAM based last-level cache management for multicore systemsabstractData prefetching is a common mechanism to mitigate the bottleneck of off-chip memory bandwidth in modern computing systems. Unfortunately, the side effects of prefetching are an additional burden on off-chip communication and increased cache write operations. With the proposal of spin-transfer torque random access memory (STT-RAM) based last-level caches (LLCs) for their high density and low power consumption, the increase of write pressure to the cache from prefetching coupled with the characteristically long write access compared with traditional SRAM caches exacerbates the performance cost of prefetching schemes. In this work, we propose two orthogonal techniques to reduce the negative performance impact induced by aggressive prefetching on multicore systems employing STT-RAM based LLC. First, basic priority assignment prioritizes the different types of access requests of LLC by their criticality and responds to them based on priority. Second, priority boosting differentiates requests by application and prioritizes the relatively few requests from applications with non-intensive accesses to the LLC, which usually creates the most severe performance degradation in multi-core systems. Combining these two prioritization policies can alleviate the negative effect induced by aggressive prefetching. Our results show that these techniques can achieve an 8.3 average application speedup compared to a baseline, prefetch only design without prioritization. Mengjie Mao, Hai Li 0001, Alex K. Jones, Yiran Chen 0001 |
ACM Great Lakes Symposium on VLSI | 1 |
| 2013 | Unleashing the potential of MLC STT-RAM cachesabstractIn this paper, we study the use of multi-level cell (MLC) spin-transfer torque RAM (STT-RAM) in cache design of embedded systems and microprocessors. Compared to the single level cell (SLC) design, a MLC STT-RAM cache is expected to offer higher density and faster system performance. However, the cell design constrains, such as the switching current requirement and asymmetry in write operations, severely limit the density benefit of the conventional MLC STT-RAM. The two-step read/write accesses and inflexible data mapping strategy in the existing MLC STT-RAM cache architecture may even result in system performance degradation. To unleash the real potential of MLC STT-RAM cache, we propose a cross-layer solution. First, we introduce the reverse magnetic junction tunneling (MTJ) into MLC cell design, which offers a more balanced device and design tradeoff and enables 2x storage density than SLC. At architectural level, we propose a cell split mapping method to divide cache lines into fast and slow regions and data migration policies to allocate the frequently-used data to fast regions. Furthermore, an application-aware speed enhancement mode is utilized to adaptively tradeoff cache capacity and speed, satisfying different requirements of various applications. Simulation results show that the proposed techniques can improve the system performance by 10.3% and reduce the energy consumption on cache by 26.0% compared with conventional MLC STT-RAM. Xiuyuan Bi, Mengjie Mao, Danghui Wang, Hai Li 0001 |
ICCAD | 2 |
| 2013 | CD-ECC: content-dependent error correction codes for combating asymmetric nonvolatile memory operation errorsabstractThe write operation asymmetry of many memory technologies causes different write failure rates at 0 →1 and 1 → 0 bit-flipping's. Conventional error correction codes (ECCs) spend the same efforts on both bit-flipping directions, leading to very unbalanced write reliability enchantment over different bit-flipping distributions of codewords (i.e., the number of 0 →1 or 1 → 0 bit-flipping's). In this work, we developed an analytic asymmetric write channel (AWC) model to analyze the asymmetric write errors in spin-transfer torque random access memory (STT-RAM) designs. A new ECC design concept, namely, content-dependent ECC (CD-ECC), is proposed to achieve balanced error correction at both bit-flipping directions. Two CD-ECC schemes - typical-corner-ECC (TCE) and worst-corner-ECC (WCE), are designed for the codewords with different bit-flipping distributions. Our simulation results show that compared to the common ECC schemes utilized in embedded applications like Hamming code, CD-ECCs can improve the STT-RAM write reliability by 10 - 30x with low hardware overhead and very marginal impact on system performance. Wujie Wen, Mengjie Mao, Xiaochun Zhu, Seung-Hyuk Kang, Danghui Wang, Yiran Chen 0001 |
ICCAD | 2 |
| 2012 | Distributed replay protocol for distributed uniprocessorsabstractData speculation technique has been heavily exploited in various scenarios of architecture design. It bridges the time or space gap between data producer and data consumer, which gives opportunities to processors to gain significant speedups. However, large instruction windows, deep pipeline and increasing latency of on-chip communication make data misspeculation very expensive in modern processors. Mengjie Mao, Hong An, Bobin Deng, Xuechao Wei, Wenting Han |
ICS | 1 |
| 2010 | FACRA: Flexible-Core Architecture Chip Resource AbstractorabstractA family of flexible-core chip multiprocessors (FCMPs) has been recently proposed to allow simple, identical physical cores to be aggregated dynamically to form larger and more powerful logical processors. However, such flexible-core architecture faces a new significant scheduling problem in the operating system, which traditionally assumes only fixed-number and fixed-granularity processors. This paper proposes a framework, called FACRA, that employs low-level runtime software to simplify OS resource allocation and process scheduling on FCMPs. Through exporting a simple, uniform processor abstraction on flexible-core chip resource, FACRA provides a set of functions with uniform interface for system-level scheduling on FCMPs. To verify the design, FACRA is built on TFlex (a typical FCMP) in our experiments, and two well known process schedulers, round-robin and dynamic-priority scheduler of Linux 2.6.11, are modified to schedule on TFlex. The evaluation results demonstrate that FACRA can efficiently simplify OS resource allocation and process scheduling on FCMPs with negligible performance loss. Hong An, Yongqing Ren, Mengjie Mao, Mu Xu, Qi Li 0034 |
PDCAT | 4 |