EDBT 2026 Demo / reviewers in the wild / expert
Danghui Wang
dblp:28/4366
· DBLP profile ↗
31ranked-venue papers
1as first author
15since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 29 · 1 first-author · 14 since 2021Artificial intelligence and machine learning · 1Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A co-optimization framework toward energy-efficient cloud-edge inference with stochastic computing and precision-compensating NAS
Mengchao Zhang, Huijuan Duan, Danghui Wang, Meikang Qiu |
J. Syst. Archit. | 5 |
| 2026 | A Transverse-Read-Assisted Fast Valid-Bits Collection in Stochastic Computing MACs for Energy-Efficient in-RTM DNNsabstractIt looks very attractive to coordinate racetrack-memory (RM) and stochastic-computing (SC) jointly to build an ultra-low power neuron-architecture. However, the above combination has always been questioned in a fatal weakness that the heavy valid-bits collection of Racetrack Memory-Magnetic Tunnel Junctions(RM-MTJ), a.k.a. accumulative parallel counters (APCs), cannot physically match the requirement for energy-efficient in-memory DNNs. Fortunately, a recently developed Transverse-Read (TR) provides a lightweight collection of valid-bits by detecting domain-wall resistance between a couple of MTJs on a single nanowire. In this work, we first propose a neuron-architecture that utilizes parallel TRs to build an ultra-fast valid-bits collection scheme specifically targeted at Multiply-Accumulate (MAC) units for in-RTM DNNs, where the multiplication operation inherently involves two operands. To solve the huge storage for full stochastic sequences caused by the limited TR banks, a hybrid coding, pseudo-fractal compression, is designed to generate stochastic sequences by segments. To overcome the misalignment by the parallel early-termination, an asynchronous schedule of TR is further designed to regularize the vectorization, in which the valid-bits from different lanes are merged in multiple RM-stacks for vector-level valid-bits collection. However, an inherent defect of TR, i.e., neighbor parts cannot be accessed simultaneously, could limit the throughput of the parallel vector multiplication, therefore, an interleaving data placement is used for full utilization of the memory bus among different vectors. The experimental results demonstrate that the SC-MAC architecture with TR achieves 2.88×-4.40× speedup over CORUSCANT (the state-of-the-art processin-RM architecture), while simultaneously reducing energy consumption by 1.26×-1.42×. Xingwu Dong, Danghui Wang |
IEEE Trans. Computers | 4 |
| 2026 | Accelerating Point Cloud Sampling by Parallel Structure Deconstruction
Hengzhe Chi, Jinzhe Zhang, Danghui Wang |
IEEE Trans. Parallel Distributed Syst. | 4 |
| 2025 | Can Short Hypervectors Drive Feature-Rich GNNs? Strengthening the Graph Representation of Hyperdimensional Computing for Memory-efficient GNNsabstractHyperdimensional computing (HDC) based GNNs are significantly advancing the brain-like cognition in terms of mathematical rigorousness and computational tractability. However, the researches in this field seem to have a “long vector consensus” that the length of HDC-hypervectors must be designed to mimic that of cerebellar cortex, i.e., ten thousands of bits, to express human’s feature-rich memory. To system architects, this choice presents a formidable challenge that the combination of numerous nodes and ultra-long hypervectors could create a new memory bottleneck that undermines the operational brevity from HDC. To overcome above problem, in this work, we shift our focus to rebuilding a set of more GNN-friendly HDC-operations, by which, short hypervectors are sufficient to encode rich features via enjoining the strong error tolerance of neural cognition. To achieve that, three behavioral incompatibilities of HDC with general GNNs, i.e., feature distortion, structural bias, and central-node vacancy, are found and successfully resolved for more efficient feature-extraction in graphs. Taken as a whole, a memory-efficient HDC-based GNN framework, called CiliaGraph, is designed to drive one-shot graph classifying tasks with only hundreds of bits in hypervector aggregation, which offers 1 to 2 orders of memory savings. The results show that, compared to the SOTA GNNs, CiliaGraph reduces the memory access and training latency by an average of $292 \times$ (up to $2341 \times$) and $103 \times$ (up to $313 \times$), respectively, while maintaining the competitively accuracy. Danghui Wang |
DAC | 3 |
| 2025 | SCArmor: Layer-Bit Joint Hardening with a Fast Genetic Optimization for Cost-Efficient and High-Reliable SC-DCNN CircuitsabstractRadiation-Induced single event upsets (SEUs) significantly damage the reliability of Deep Convolution Neural Networks (DCNNs) in harsh environments, i.e., deep space. Stochastic computing (SC) has exhibited the outstanding tolerance against SEUs via diluting a binary weight in a longer stochastic stream. However, the exponential expanding of stochastic streams and multi-layer neural networks together incur an unaffordable hardware consumption in their fully hardened SC-DCNN circuits. In this paper, SCArmor, a genetic evolution-based hardening framework, is proposed to select the critical bits and layers fast for the reliable and lightweight anti-radiation designs. First, a layer-wise error propagation model (LEPM) is built to evaluate the layer-sensitivity that indicates how many of a faulty-layer errors propagated to the downstream, serving as the reliability weights across the layers in SC-DCNNs. Second, the bit-sensitivity is modeled with sensitivity-driven fault injection (SDFI) to highlight those critical bits in both stochastic and unary numbers, which are further accumulated within a layer, i.e., convolution, activation, pooling, and full connection, to estimate the fault-tolerant abilities of the individual layers. Finally, by jointly using above two sensitivities as the fast emulation of inference accuracy, a re-training-free genetic algorithm is embedded into the hardening framework to locate the primary layers and SC-bits that should be hardened under the constraints of both inference accuracy and circuit cost. On the popular SC-DCNNs, the results show that, compared with the SOTA hardening methods, SCArmor reduces the area, power, and latency by 22.21%, 29.69%, and 43.15% respectively, only using 42.46% cost of full-hardening approaches. Regarding reliability, the accuracy of SCArmor achieves up to 98.63% of that of the radiation-free inference, outperforming the SOTAs by 3.13%. Yubin Zhang, Danghui Wang |
ICCAD | 4 |
| 2025 | SEER: A Slothful Encoding to Mitigate Read Error of Stop-in-Middle Domain-Wall for Reliable Racetrack Memory
Danghui Wang |
WASA (1) | 2 |
| 2024 | SC-GNN: A Communication-Efficient Semantic Compression for Distributed Training of GNNsabstractTraining big graph neural networks (GNNs) in distributed systems is quite time-consuming mainly because of the ubiquitous aggregate operations that involve a large amount of cross-partition communication for collecting embeddings/gradients during the forward and backward propagation. To reduce the volume of the communication, some recent approaches focused on decaying each of connections via sampling, quantifying, or delaying until satisfactory trade-off are obtained between volume and accuracy. However, when applied to popular GNNs, those approaches are found to be bounded by a common volume/accuracy Pareto frontier which shows that the decaying for individual connection cannot further accelerate the aggregate of training. In this work, SC-GNN, a semantic compression of the cross-partition communication, is proposed to concentrate a group of connections as a high-level semantics and transmit to a target partition. Since carrying the overall intent of a group, the semantics can keep transferring the interactions, i.e., embeddings/gradients, between a pair of remote partitions until GNN models converge. In addition, a connection-pattern based differential optimization is proposed to further prune those weak connections, while guaranteeing the training accuracy. The results show that, for multi-field datasets, the compression rate of SC-GNN is 40.8 × higher than SOTA methods and the epoch time is reduced to 31.8% on average. Danghui Wang |
DAC | 3 |
| 2023 | PseudoSC: A Binary Approximation to Stochastic Computing within Latent Operation-Space for Ultra-Lightweight on-Edge DNNsabstractRecently, stochastic computing (SC) is increasingly popular in constructing MAC for on-edge DNNs benefiting from its outstanding energy-efficiency, including its adequate precision and gate-level operation. However, current SC-DNN systems always include a lot of costly SNGs/APCs for inevitably switches between binary and stochastic domains, which mortgages incongruous resources to pay the "bill" of the domain-switches and impedes highly-concurrent deployments. In this work, PseudoSC, a binary approximation to low-discrepancy SC, is proposed to totally remove the domain-switch for SNG/APC-free SC-DNNs. Its basic idea is to virtually re-arrange a couple of stochastic operands into a 2-D latent op-space, in which, original Monte Carlo sampling can be partitioned into three sub-ops, i.e., two fixed binary-ops and a fractal recursion. In theory, the recursion forms an isomorphic partition of the sampling repeated in smaller scales until the binary base-case achieved, as a result, a SC-op is well approximated only with binary-ops. Based on above theory, a multi-lane micro-architecture is designed to unroll the recursion within a few cycles and its advantages on hardware saving is verified under popular DNNs. The evaluation shows that the DNN-models with our schemes achieve 98.7% accuracy of the fixed-point implementations, which significantly outperform other SOTA methods. In addition, its reduced structure improves the power efficiency by 3.67 times on average. Zhaoqing Wang, Danghui Wang |
DAC | 3 |
| 2023 | ExpoNAS: Using Exposure-based Candidate Exclusion to Reduce NAS Space for Heterogeneous Pipeline of CNNsabstractEfficiently deploying neural network models on heterogeneous edge systems faces unique challenges. Two major issues are underutilized pipelines caused by hardware heterogeneity and explosive growth of architecture search space. To address these challenges, we propose ExpoNAS, a framework enabling fast deployment-oriented neural architecture search on heterogeneous platforms. ExpoNAS evaluates model partitioning and deployability before search to meet accuracy and latency constraints. However, repeated partitioning leads to exponential growth in architecture search space. We first observe a strong correlation between operational exposure and accuracy improvement, and then we propose an exposure-based reduction technique that excludes the operations of low exposure. By removing these non-critical operations, the exposure-based pruning successfully reduces architecture search time. Our experiments on popular datasets show that the ExpoNAS pipelined design improves execution efficiency by a factor of 2 compared to traditional batch processing. Without compromising accuracy, ExpoNAS is able to reduce the architecture search time by 47.5% for discovering high-accuracy networks, compared to the SOTA NAS strategies. Huijuan Duan, Danghui Wang, Xingting Zhao |
ICPADS | 2 |
| 2023 | A Noise-Driven Heterogeneous Stochastic Computing Multiplier for Heuristic Precision Improvement in Energy-Efficient DNNsabstractStochastic computing (SC) has become a promising approximate computing solution by its negligible resource occupancy and ultralow energy consumption. As a potential replacement of accurate multiplication, SC can dramatically mitigate the problematic power consumption by DNNs. However, current SC-multipliers illustrate an extremely imbalanced accuracy across product space, i.e., neglectable noise with large products but significant noise for small ones, which is discordant to the distribution of products by the sparse matrix in neural computing. In this article, we present a heterogeneous SC-multiplier that heuristically performs three divergent approximating multiplication, including “set-to-0,” “look-up-table,” and “low-discrepancy-SC,” for appropriate precision-provision in the whole space of products. Due to those popular DNN models cannot achieve consensus on the boundaries of above operations, a training-involved method is proposed to determine the settings with limited overhead. In this way, those models successively learn the SC-operation characters and exhibit a definitely improvement on network precision. The experiment shows that, for single multiplication, the product noise can be restrained by 36.86% on average, and for multiplication in multiple network models, the accuracy improvement reaches to 5.5% on average. Furthermore, a group of proposed logic-reduction techniques can improve the energy efficiency by 65% in the system-level evaluation. Danghui Wang, Shengbing Zhang, Xiaoya Fan |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2022 | DCNN search and accelerator co-design: Improve the adaptability between NAS frameworks and embedded platforms
Chuxi Li, Xiaoya Fan, Shengbing Zhang, Zhao Yang 0005, Danghui Wang, Meng Zhang 0047 |
Integr. | 6 |
| 2022 | An Automatic-Addressing Architecture With Fully Serialized Access in Racetrack Memory for Energy-Efficient CNNsabstractRacetrack memory, an emerging low-power magnetic memory, promises a competitive replacement for traditional memory in the accelerators. However, random access in racetrack memory is time and energy expenditure for CNN accelerators because of its large amount of invalid-shifts. In this article, we propose an automatic-addressing architecture that builds a novel data layout to guarantee that the next round of memory access can be always satisfied at the in-situ or rigorously adjacent cells of current round, producing a fully serialized access footprint that can drive instant port-alignment without any invalid-shifts in racetrack memory. By this way, original address-based access degrades to the selections repeated among the three candidates, i.e., onein-situcell and two neighbor cells. Based on this simplification, a lightweight access management can generate the sequence of one-out-three selections according to the deterministic access behaviors defined by CNN hyper-parameters. The evaluation shows that, when deploying the five popular CNN applications to our architecture, the physical shifts of racetrack is curtailed by 74.64 percent over legacy layout, which achieves 54.2 and 42.1 percent energy reduction on read and write, respectively. A case study of YOLOv2 indicates that our architecture performs 6.503 GOp/J that achieves$18.5 \times$improvement to server-level GPUs. Danghui Wang, Jianfeng An, Xiaoya Fan |
IEEE Trans. Computers | 3 |
| 2022 | MemUnison: A Racetrack-ReRAM-Combined Pipeline Architecture for Energy-Efficient in-Memory CNNsabstractThough ReRAM has been greatly successful in reducing energy consumption of various neural networks, it still suffers write amplification in energy, which impedes ReRAM to provide efficient storage for the ubiquitous streaming data in CNNs, such as feature-maps. Racetrack memory, an emerging magnetic memory technique, is a proper candidate to hold streaming data since it enjoys fast sequential-access with ultra-low operating energy in read and write. In this work, we propose a hybrid processing-in-memory architecture, called MemUnison, that coordinates ReRAM and racetrack to overcome the expenditure storage of streaming data in ReRAM. By placing feature-maps in racetrack and leaving weights in ReRAM, a datapath is constructed between the two sides to form a fetch-process-writeback pipeline. As the invalid-shifts of the racetrack memory incurs a large amount of pipeline bubble, we propose a row-based access that can read and write a feature-map without any invalid-shifts. For the row-based operation, a cohesive controlling method is proposed to coordinate racetrack and ReRAM. In runtime, convolution kernels are scheduled in ReRAM banks for cross-channel calculations of one row, by which computing complexity of a convolutional layer can be reduced by 4 orders of magnitude, excessing the 2 order of reduction by traditional ReRAM. Danghui Wang, Shengbing Zhang, Xiaoya Fan |
IEEE Trans. Computers | 3 |
| 2021 | Hardware-Aware NAS Framework with Layer Adaptive Scheduling on Embedded SystemabstractNeural Architecture Search (NAS) has been proven to be an effective solution for building Deep Convolutional Neural Network (DCNN) models automatically. Subsequently, several hardware-aware NAS frameworks incorporate hardware latency into the search objectives to avoid the potential risk that the searched network cannot be deployed on target platforms. However, the mismatch between NAS and hardware persists due to the absent of rethinking the applicability of the searched network layer characteristics and hardware mapping. A convolution neural network layer can be executed on various dataflows of hardware with different performance, with which the characteristics of on-chip data using varies to fit the parallel structure. This mismatch also results in significant performance degradation for some maladaptive layers obtained from NAS, which might achieved a much better latency when the adopted dataflow changes. To address the issue that the network latency is insufficient to evaluate the deployment efficiency, this paper proposes a novel hardware-aware NAS framework in consideration of the adaptability between layers and dataflow patterns. Beside, we develop an optimized layer adaptive data scheduling strategy as well as a coarse-grained reconfigurable computing architecture so as to deploy the searched networks with high power-efficiency by selecting the most appropriate dataflow pattern layer-by-layer under limited resources. Evaluation results show that the proposed NAS framework can search DCNNs with the similar accuracy to the state-of-the-art ones as well as the low inference latency, and the proposed architecture provides both power-efficiency improvement and energy consumption saving. Chuxi Li, Xiaoya Fan, Shengbing Zhang, Zhao Yang 0005, Danghui Wang, Meng Zhang 0047 |
ASP-DAC | 6 |
| 2021 | Balancing memory-accessing and computing over sparse DNN accelerator via efficient data packaging
Xiaoya Fan, Tengteng Yao, Danghui Wang |
J. Syst. Archit. | 7 |
| 2020 | An Energy-Efficient AES Encryption Algorithm Based on Memristor Switch
Danghui Wang, Chen Yue, Ze Tian, Ru Han |
ICA3PP (3) | 1 |
| 2020 | A Hot/Cold Task Partition for Energy-Efficient Neural Network Deployment on Heterogeneous Edge Device
Jiaxiang Zhao, Danghui Wang |
ICA3PP (2) | 3 |
| 2020 | ENAS oriented layer adaptive data scheduling strategy for resource limited hardware
Chuxi Li, Xiaoya Fan, Yuling Geng, Danghui Wang |
Neurocomputing | 5 |
| 2019 | A low-power sensor polling for aggregated-task context on mobile devices
Danghui Wang |
Future Gener. Comput. Syst. | 2 |
| 2018 | A locality-aware shuffle optimization on fat-tree data centers
Danghui Wang, Meikang Qiu, Yao Chen 0008, Bing Guo 0003 |
Future Gener. Comput. Syst. | 2 |
| 2018 | TriZone: A Design of MLC STT-RAM Cache for Combined Performance, Energy, and Reliability OptimizationsabstractSpin-transfer torque random access memory (STT-RAM) is a promising technology for future nonvolatile caches and memories. To increase the storage density, multilevel cell (MLC) technique was recently introduced to STT-RAM designs at the cost of degraded access speed, reliability, and energy efficiency. Existing MLC STT-RAM cache architectures primarily focus on the performance and energy optimizations but ignore the crucial demand for reliability. In this paper, we propose “TriZone”-a holistic design scheme for MLC STT-RAM cache to simultaneously meet the requirements of performance, energy, and reliability. Three cache block configurations, namely hard, soft, and mixed, are constructed with the hard-bit, soft-bit, and both hard-bit and soft-bit of MLC STT-RAM, respectively. By observing the difference of these cache blocks, a nonuniform strength ECC (NUS-ECC) is developed to guarantee the operational reliability of a cache block with a variable decoding delay adapting to the needs of error correction (e.g., the number of the erroneous bits). The whole MLC STT-RAM cache is then partitioned into three regions, each of which is composed of different cache blocks. In order to achieve the best tradeoff among performance, energy, and reliability, we then introduce the dynamic cache partitioning to determine the partition of this tri-way MLC STT-RAM cache according to the runtime characteristic of various applications. Experiment results show that compared with conventional performance-driven MLC STT-RAM cache design with pessimistic ECC, TriZone can improve the system performance and energy by averagely 11.7% (10.0%) and 13.3% (15.7%), respectively, for single-threaded (multiprogram) applications. The additional area overhead associated with NUS-ECC is limited by ~ 3%. Zihao Liu 0015, Mengjie Mao, Tao Liu 0023, Wujie Wen, Yiran Chen 0001, Hai Li 0001, Danghui Wang, Yukui Pei, Ning Ge 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 8 |
| 2017 | XNOR-POP: A processing-in-memory architecture for binary Convolutional Neural Networks in Wide-IO2 DRAMsabstractIt is challenging to adopt computing-intensive and parameter-rich Convolutional Neural Networks (CNNs) in mobile devices due to limited hardware resources and low power budgets. To support multiple concurrently running applications, one mobile device needs to perform multiple CNN tests simultaneously in real-time. Previous solutions cannot guarantee a high enough frame rate when serving multiple applications with reasonable hardware and power cost. In this paper, we present a novel process-in-memory architecture to process emerging binary CNN tests in Wide-IO2 DRAMs. Compared to state-of-the-art accelerators, our design improves CNN test performance by 4× ~ 11× with small hardware and power overhead. Lei Jiang 0001, Wujie Wen, Danghui Wang |
ISLPED | 4 |
| 2017 | FlexLevel NAND Flash Storage System Design to Reduce LDPC LatencyabstractAggressive technology scaling and adoption of multilevel-cell technique lead to progressive increase of bit error rate (BER) of NAND flash memory. Consequently, conventional error correction code is not adequate to guarantee system reliability. As an alternative, low density parity check (LDPC) code is introduced to provide more powerful error correction capability. However, to achieve better performance, LDPC code demands extra memory sensing operations and more data transfer cycles, directly leading to longer read latency. To achieve both system reliability and read efficiency, we propose the FlexLevel NAND flash storage system design in this paper. FlexLevel consists of two levels of optimization: 1) LevelAdjust and 2) AccessEval. At device level, the LevelAdjust technique is proposed to reduce BER by broadening noise margin via threshold voltage level reduction. With LevelAdjust, BER is greatly reduced and no extra sensing levels are required to protect data integrity. Hence, read performance is improved. However, while LevelAdjust can improve system reliability and read performance, it causes density loss. To balance read performance improvement and density loss, we propose the AccessEval technique at system level. AccessEval identifies data with high LDPC overhead and only applies LevelAdjust technique to these data. The experimental results show that compared with the best existing works, the proposed design can achieve up to 11% read speedup with negligible density loss. Jie Guo 0002, Wujie Wen, Jingtong Hu, Danghui Wang, Hai Li 0001, Yiran Chen 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2017 | Cross-Layer Optimization for Multilevel Cell STT-RAM CachesabstractSpin-transfer torque random access memory (STT-RAM), as an emerging nonvolatile memory technology, provides very dense array structure and extremely low leakage power consumption. It demonstrates a great potential in replacing conventional static random access memory technology to develop the next-generation on-chip cache memory of microprocessors and graphics processing units. The multilevel cell (MLC) design of STT-RAM that stores two or more bits in one cell potentially has higher storage capacity and faster system performance, attracting significant attention. In this paper, we first quantitatively evaluated the data storage density of the MLC STT-RAM. Our results revealed limited density improvement because of the large size of access transistor induced by high write current amplitude requirement and asymmetry of switching behavior. Moreover, the read and write accesses of existing MLC STT-RAM cache designs require two-step operation. The system level evaluation shows that the long access latency could amortize the performance speed brought by larger cache size, and even degrade the system performance for some applications. To unleash the potential of MLC STT-RAM cache, we proposed a new design through a cross-layer co-optimization. The memory cell structure integrated the reversed stacking of magnetic junction tunneling for a more balanced device and design tradeoff. In architecture development, we presented an adaptive mode switching mechanism: based on application's memory access behavior, the MLC STT-RAM cache can dynamically change between low latency single-level cell mode and high capacity MLC mode. Furthermore, we divided cache lines into fast and slow regions and investigated new data migration policies to allocate frequently access data to fast regions. Simulation results show that the proposed techniques can improve the system performance by 10.2% and reduce the energy consumption on cache by 9.5% compared with conventional MLC STT-RAM cache design. Xiuyuan Bi, Mengjie Mao, Danghui Wang, Hai Li 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2017 | Data-Pattern-Aware Error Prevention Technique to Improve System ReliabilityabstractProgram disturb, read disturb, and retention time noise are identified as three major contributors to multilevel cell (MLC) NAND flash memory bit errors. With program/erase cycling and technology scaling, bit error rate (BER) of MLC NAND flash memory rapidly increases. Previous works revealed that BER is heavily dependent on data patterns. Based on this observation, we propose data-pattern-aware (DPA) error protection technique to extend the lifespan of NAND flash-based storage systems. DPA manipulates the ratios of 0's and 1's in the stored data to reduce the probability of the data patterns, which are susceptible to device noises. By minimizing the vulnerable data patterns, our scheme can effectively reduce the BER and improves the system endurance. Our DPA scheme also incorporates a data management scheme to minimize the redundancy-induced performance overhead. Simulation results show that our scheme can increase flash system life expectancy by up to 4×. Jie Guo 0002, Danghui Wang, Zili Shao, Yiran Chen 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2016 | Improving read performance of STT-MRAM based main memories through Smash Read and Flexible ReadabstractSpin Transfer Torque Magnetoresistive RAM (STT-MRAM) has been recently deemed as one promising main memory alternative for high-end mobile processors. With process technology scaling, the amplitude of write current approaches that of read current in deep sub-micrometer STT-MRAM arrays. As a result, read disturbance errors (RDEs) emerge. Both high current restore required (HCRR) reads and low current long latency (LCLL) reads can guarantee read reliability and utterly remove RDEs. However, both of them degrade system performance, because of extra restores or a longer read latency. And neither of them always achieves the better performance when running a wide variety of applications. In this paper, we present two architectural techniques to boost read performance for STT-MRAM based main memories in the presence of RDEs. We first propose Smash Read (S-RD) to shorten the latency of HCRR reads by injecting a larger read current. We further introduce Flexible Read (F-RD) to dynamically adopt different types of read schemes, S-RD and LCLL, to maximize main memory system performance. On average, our techniques improve system performance by 9~13% and reduces total energy by 4~8% over all existing read schemes including HCRR and LCLL. Lei Jiang 0001, Wujie Wen, Danghui Wang, Lide Duan |
ASP-DAC | 3 |
| 2016 | TEMP: thread batch enabled memory partitioning for GPUabstractAs massive multi-threading in GPU imposes tremendous pressure on memory subsystems, efficient bandwidth utilization becomes a key factor affecting the GPU throughput. In this work, we propose thread batch enabled memory partitioning (TEMP), to improve GPU performance through the improvement of memory bandwidth utilization. In particular, TEMP clusters multiple thread blocks sharing the same set of pages into a thread batch and dispatches the entire thread batch to a stream multiprocessor. TEMP separates the memory access streams of different thread batches by OS memory management, preserving the intrinsic locality of thread batches and increasing the memory access parallelism. Experimental results show that TEMP can obtain up to 10.3% performance improvement and 14.6% DRAM energy reduction compared to a state-of-the-art scheduler without any memory-side optimizations. Mengjie Mao, Wujie Wen, Xiaoxiao Liu 0001, Jingtong Hu, Danghui Wang, Yiran Chen 0001, Hai Li 0001 |
DAC | 5 |
| 2015 | FlexLevel: a novel NAND flash storage system design for LDPC latency reductionabstractLDPC code is introduced in NAND flash memory to handle high BER (bit error rate) incurred by technology scaling. Despite strong error correction capability, LDPC decoding induces long NAND flash read latency. In this work, we propose FlexLevel -- a robust NAND flash storage system design to improve data reliability and read efficiency affected by the LDPC operations. FlexLevel first reduces BER by enlarging noise margins via Vth (threshold voltage) level reduction. It reduces the sensing levels of LDPC but also causes loss of storage capacity. To compensate this capacity loss with minimum impact on read performance, FlexLevel identifies the data with high LDPC overhead and only applies the Vth level reduction technique to those data. Experimental results show that compared with state-of-the-art, FlexLevel can achieve up to 33% read speedup with very moderate capacity loss. Jie Guo 0002, Wujie Wen, Jingtong Hu, Danghui Wang, Hai Li 0001, Yiran Chen 0001 |
DAC | 4 |
| 2014 | DPA: A data pattern aware error prevention technique for NAND flash lifetime extensionabstractThe recent research reveals that the bit error rate of a NAND flash cell is highly dependent on the stored data patterns. In this work, we propose Data Pattern Aware (DPA) error protection technique to extend the lifespan of NAND flash based storage systems (NFSS). DPA manipulates the ratio of 1's and 0's in the stored data to minimize occurrence of the data patterns which are susceptible to bit error noise. Consequently, the NAND flash cell bit error rate is reduced, leading to system endurance extension. Our simulation result shows that, with marginal hardware and power overhead, DPA scheme can increase the NFSS lifetime by up to 4×, offering a complementing solution to other lifetime enhancement techniques like wear-leveling. Jie Guo 0002, Danghui Wang, Zili Shao, Yiran Chen 0001 |
ASP-DAC | 3 |
| 2013 | Unleashing the potential of MLC STT-RAM cachesabstractIn this paper, we study the use of multi-level cell (MLC) spin-transfer torque RAM (STT-RAM) in cache design of embedded systems and microprocessors. Compared to the single level cell (SLC) design, a MLC STT-RAM cache is expected to offer higher density and faster system performance. However, the cell design constrains, such as the switching current requirement and asymmetry in write operations, severely limit the density benefit of the conventional MLC STT-RAM. The two-step read/write accesses and inflexible data mapping strategy in the existing MLC STT-RAM cache architecture may even result in system performance degradation. To unleash the real potential of MLC STT-RAM cache, we propose a cross-layer solution. First, we introduce the reverse magnetic junction tunneling (MTJ) into MLC cell design, which offers a more balanced device and design tradeoff and enables 2x storage density than SLC. At architectural level, we propose a cell split mapping method to divide cache lines into fast and slow regions and data migration policies to allocate the frequently-used data to fast regions. Furthermore, an application-aware speed enhancement mode is utilized to adaptively tradeoff cache capacity and speed, satisfying different requirements of various applications. Simulation results show that the proposed techniques can improve the system performance by 10.3% and reduce the energy consumption on cache by 26.0% compared with conventional MLC STT-RAM. Xiuyuan Bi, Mengjie Mao, Danghui Wang, Hai Li 0001 |
ICCAD | 3 |
| 2013 | CD-ECC: content-dependent error correction codes for combating asymmetric nonvolatile memory operation errorsabstractThe write operation asymmetry of many memory technologies causes different write failure rates at 0 →1 and 1 → 0 bit-flipping's. Conventional error correction codes (ECCs) spend the same efforts on both bit-flipping directions, leading to very unbalanced write reliability enchantment over different bit-flipping distributions of codewords (i.e., the number of 0 →1 or 1 → 0 bit-flipping's). In this work, we developed an analytic asymmetric write channel (AWC) model to analyze the asymmetric write errors in spin-transfer torque random access memory (STT-RAM) designs. A new ECC design concept, namely, content-dependent ECC (CD-ECC), is proposed to achieve balanced error correction at both bit-flipping directions. Two CD-ECC schemes - typical-corner-ECC (TCE) and worst-corner-ECC (WCE), are designed for the codewords with different bit-flipping distributions. Our simulation results show that compared to the common ECC schemes utilized in embedded applications like Hamming code, CD-ECCs can improve the STT-RAM write reliability by 10 - 30x with low hardware overhead and very marginal impact on system performance. Wujie Wen, Mengjie Mao, Xiaochun Zhu, Seung-Hyuk Kang, Danghui Wang, Yiran Chen 0001 |
ICCAD | 5 |