VLDB 2026 Research / reviewers in the wild / expert
Zhigang Mao
dblp:37/5644 · also Zhi-Gang Mao
· DBLP profile ↗
83ranked-venue papers
0as first author
39since 2021 · last 2026
0000-0001-9431-9853ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 69 · 34 since 2021Software engineering, systems software and programming languages · 11 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 since 2021Security and privacy · 2Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Viper: An ILP-Based Vectorization Framework for Fully Homomorphic Encryption
Weidong Yang 0007, Xinmo Li, Xiangmin Guo, Jianfei Jiang 0001, Naifeng Jing, Qin Wang 0009, Zhigang Mao, Weiguang Sheng |
ASP-DAC | 7 |
| 2026 | CHIME: A Case for Efficient Long-Context Attention-FC Disaggregated Inference with DIMM-PIM
Yanning Yang, Dong Du 0003, Zhigang Mao, Naifeng Jing, Yubin Xia, Haibo Chen 0001 |
ISCA | 6 |
| 2026 | HARMONY: A Hardware-Aware Mapping and Optimizing Framework for Computing-in-Memory AcceleratorsabstractThe increasing adoption of artificial intelligence has spurred the development of specialized deep neural network (DNN) accelerators. Among them, computing-in-memory (CIM) architectures are promising for their in-situ computation capability, which alleviates the computation and data movement bottlenecks of modern DNNs. However, the diversity of models and hardware designs makes it challenging to fully exploit CIM accelerators. Existing approaches often rely on manual mapping or provide limited automation, struggling to integrate general-purpose optimizations with CIM-specific features. In this work, we present HARMONY, a hardware-aware compilation framework for CIM accelerators. At its core is a hardware intermediate representation (IR) that unifies computational and memory abstractions. Based on this IR, HARMONY introduces an automatic mapping algorithm that identifies offloadable operators and constructs a hybrid software–hardware IR. This enables systematic integration of general-purpose and CIM-specific scheduling primitives within a unified search space, which is efficiently explored using reinforcement learning (RL). Extensive evaluations show that HARMONY supports a broader set of operators than existing CIM compilers and consistently delivers substantial performance and energy improvements across diverse DNN workloads. These results demonstrate that HARMONY provides both generality and efficiency, making it a practical compilation solution for CIM accelerators. Xinmo Li, Weidong Yang 0007, Naifeng Jing, Qin Wang 0009, Zhigang Mao, Weiguang Sheng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2026 | Leveraging Tensor Dataflow for Improved Thermal Performance on 3D-Stacked SRAM ArchitectureabstractWhile 3D-stacked SRAM architectures have demonstrated prominent performance speedup for tensor computing by exploiting higher bandwidth, larger buffer and reduced latency, they suffer from thermal challenges owing to vertical stacking nature of chips. In this paper, we identify that tensor dataflow may further exacerbate the thermal issues, so we propose T3D, the first thermal-aware tensor framework for 3D-stacked SRAM architectures, leveraging tensor dataflow characteristics to significantly enhance thermal performance. Specifically, we first perform a quantitative formulation to identify the most energy-efficient tensor dataflow with given 3D constraints, effectively reducing heat generation without performance loss. Then, we develop a thermal-aware 3D architectural floorplan to improve heat spreading by optimizing the spatial arrangement of multiple SRAM macros with varying power overheads, which is caused by mismatched data access rates of tensor computing. Experimental results show that, our proposed T3D can reduce the peak chip temperature by 12.9°C on certain LLM and DNN workloads over the state-of-the-art 3D solutions. Pengyu Liu 0004, Zelong Yuan, Yingkun Liu, Jianfei Jiang 0001, Qin Wang 0009, Zhigang Mao, Naifeng Jing |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2026 | Glitch-Aware Optimization of 4-2 Compressor-Based Approximate MultipliersabstractApproximate multipliers have widely been used in error-resilient applications with hardware-efficient and relaxed precision requirements, such as multimedia signal processing and deep learning. Aiming to achieve improvements in power efficiency and performance, approximate multipliers are generally designed by simplifying the implementation circuits. However, the redundant switching activities (also known as glitches) due to unbalanced signal paths are seldom considered, leading to significant dynamic power. Moreover, prior approximate designs often overlook transistor sizing optimization, which limits their potential for power–delay efficiency. This article proposes a custom design and optimization framework for approximately 4-2 compressors at the structural and circuit levels, which effectively reduces glitches by balancing output delays. Based on the devised approximate 4-2 compressors, the constructed multipliers can then achieve significant reductions in the generation and propagation of spurious activities. In addition, to further reduce the dynamic power consumption, a delay-aware signal routing (DASR) strategy is introduced for interconnecting approximate compressors. The stability and efficiency of the proposed designs under varying conditions are verified by extensive simulations. The experimental results on HLMC 28-nm CMOS technology show that the approximate 4-2 compressors obtained by the proposed framework achieve 10.4%–52.1% power–delay product (PDP) reductions compared to existing approximate designs with the same accuracy, resulting in 5.20%–33.1% PDP improvements for multipliers. Moreover, the proposed optimization framework is generalizable to arbitrary approximate 4-2 compressor designs. Tongjing Wu, Honglan Jiang, Jianfei Jiang 0001, Qin Wang 0009, Zhigang Mao |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2025 | MDNMP: Metapath-Driven Software-Hardware Co-Design for HGNN Acceleration with Near-Memory ProcessingabstractHeterogeneous graph neural networks (HGNNs), which capture rich structural and semantic information by learning low-dimensional vertex representations based on metapath, have drawn considerable attention in recent years. Due to substantial memory consumption and unique irregular access patterns, its performance is hindered by memory-bound metapath instance matching and aggregation. To address this challenge, the recent proposal employs near-memory processing (NMP) and achieves impressive performance speedups. However, due to oversight of the intrinsic characteristics of metapath, it fails to fully exploit the potential of NMP. Jianfei Jiang 0001, Qin Wang 0009, Zhigang Mao, Naifeng Jing |
ASP-DAC | 4 |
| 2025 | AttenPIM: Accelerating LLM Attention with Dual-mode GEMV in Processing-in-MemoryabstractLarge Language Models (LLMs) have demonstrated unprecedented generative performance across a wide range of applications. While recent heterogeneous architectures attempt to address the memory-bound bottleneck from attention computations by processing-in-memory (PIM) offloading, they overlook two critical characteristics of attention GEMVs that distinguish them from traditional PIM scenarios: (1) dynamic matrix dimensions that scale with token length, and (2) distinct GEMV patterns between score computation ($Q \times K_{t}$) and context computation ($S \times V$). Existing PIM designs, employing either uniform or transposed computing modes, suffer from inefficiencies in newly generated element preparation or distinct GEMV execution. To address these limitations, we propose AttenPIM, a software-hardware co-design for efficient PIM-based attention acceleration. For bank-level execution, we propose dual-mode computing modes tailored for score and context computations with PIM-oriented data layouts and execution flows for KV storage, supported by a low-cost configurable per-bank PIM unit (PU). For system-level execution, we leverage token-level and head-level concurrency to ensure workload balance and maximize bank PU parallelism. Furthermore, dynamic allocation and kernel fusion methods are proposed to further minimize memory overhead. Experimental results demonstrate that AttenPIM achieves $1.13 \times-5.26 \times$ speedup and reduces energy consumption by 17 %-49 % compared to two state-of-the-art PIM baselines. Dongxu Lyu, Jianfei Jiang 0001, Qin Wang 0009, Zhigang Mao, Naifeng Jing |
DAC | 6 |
| 2025 | EPIC: Error PredIction and Correction for Power-Efficient Voltage Underscaling Multiply-Accumulate UnitabstractMatrix multiplication dominates the power consumption in compute-intensive applications such as deep neural networks (DNNs), spurring intensive investigations into power-efficient multiply-accumulate (MAC) units. Among the mainstream low-power design methodologies, voltage underscaling can achieve effective power savings yet induce timing errors that may lead to catastrophic accuracy loss. In this paper, we propose an error prediction and correction framework (denoted as EPIC) for arbitrary MAC unit under voltage underscaling, which predicts the timing errors and samples the correct output by using a delay-tunable clock. A prediction bits searching algorithm is proposed to enhance the prediction accuracy with low hardware cost, resulting in up to 100% accuracy. While preserving the accuracy, EPIC achieves up to 52% power savings over the corresponding MAC operating at nominal voltage. With transistor-level optimizations, EPIC incurs only 8% area and 1% power overheads, achieving 100% error correction under a voltage underscaling ratio of $\mathbf{0. 7 4}$. Compared to state-of-the-art error resilient circuit designs, EPIC consumes 60%-88% less area. Additionally, to achieve the accuracy performance of EPIC in error-resilient applications, we propose a simulation workflow involving precise timing features, enabling an accurate simulation of voltage underscaling MAC in large-scale applications. The experimental results show that, under voltage-underscaling, the MAC with EPIC consumes 11% less power than the one without EPIC, when a same accuracy as exact implementation is required in multi-layer perceptron (MLP). Tongjing Wu, Xiaolu Hu, Siting Liu 0001, Hui Wang 0023, Weifeng He, Zhigang Mao, Honglan Jiang |
DAC | 7 |
| 2025 | A Low-Power Mixed-Precision Integrated Multiply-Accumulate Architecture for Quantized Deep Neural NetworksabstractAs mixed-precision quantization techniques have been widely considered for balancing computational efficiency and flexibility in quantized deep neural networks (DNNs), mixed-precision multiply-accumulate (MAC) units are increasingly important in DNN accelerators. However, conventional mixed-precision MAC architectures support either signed × signed or unsigned ×unsigned multiplications. The signed ×unsigned multiplication enhancing the computing efficiency of DNNs with ReLU activations has never been considered in the design of mixed-precision MAC. Thus, this work proposes a mixed-precision MAC architecture supporting six operation modes, int8 × int8, int8 × uint8, two int4 × int4, two int4 × uint4, four int2 × int2, and four int2 × uint2. In this design, to balance the power and delay of different modes, the multiplication is implemented based on four precision-split 4×4 multipliers (PS4Ms). The accumulation is integrated into the partial product accumulation of the multiplication to eliminate redundant switching activities in separate compression. With 10% area reduction, the proposed MAC denoted as PS4MAC, reduces the power by over 35%, 42%, and 56% for 8-bit, 4-bit, and 2-bit operations, respectively, compared with the design based on the Synopsys DesignWare (DW) multipliers. Additionally, it achieves over 23% power savings for 8-bit operations compared to state-of-the-art (SotA) mixed-precision MAC designs. To save more power, an approximate computing mode for 8-bit multiplication is further designed, resulting in a MAC unit enabling eight operation modes, referred to as PS4MAC_AP. Finally, output-stationary systolic arrays (SAs) are explored using the above-mentioned MAC designs to implement DNNs operating under a 1 GHz clock. Our designs show the highest energy efficiency and outstanding area efficiency in all 8-bit, 4-bit, and 2-bit operation modes. Compared with the traditional SA with high-precision-split multipliers, PS4MAC_AP improves the energy efficiency for 8-bit operations by 0.6 TOPS/W, and PS4MAC achieves 0.4 TOPS/W - 0.7 TOPS/W improvement for all operation modes. Xiaolu Hu, Xinkuang Geng, Zhigang Mao, Jie Han 0001, Honglan Jiang |
DATE | 3 |
| 2025 | HEILP: An ILP-Based Scale Management Method for Homomorphic Encryption CompilerabstractRNS-CKKS, a fully homomorphic encryption (FHE) scheme, enabling secure computation on encrypted data, has widely be used in statistical analysis and data mining. However, developing RNS-CKKS programs requires substantial knowledge of cryptography, which is unfriendly to non-expert programmers. A critical obstacle is the scale management, which affects the complexity of programming and performance. Different FHE operations impose specific requirements on the scale and level, necessitating programmer intervention to ensure the recoverability of the results. Furthermore, operations at different levels have a significant impact on program performance. Existing methods rely on heuristic insights or iterative methods to manage the scales of ciphertexts. However, these methods lack a holistic understanding of the optimization space, leading to inefficient exploration and suboptimal performance. This work proposes HEILP, the first constrained-optimization-based approach for scale management in FHE. HEILP expresses node scale decision and scale management operation inserting as an integer linear programming model which can be solved with existing mathematical techniques in one shot. Our method creates a more comprehensive optimization space and enables a faster and more efficient exploration. Experimental results demonstrate that HEILP achieves an average performance improvement of 1.72 x over existing heuristic method, and outperforms a 1.19 x performance improvement with 48.65 x faster compilation time compared to the state-of-the-art iteration-based method. Weidong Yang 0007, Shuya Ji, Jianfei Jiang 0001, Naifeng Jing, Qin Wang 0009, Zhigang Mao, Weiguang Sheng |
DATE | 6 |
| 2025 | AsyncDIMM: Achieving Asynchronous Execution in DIMM-Based Near-Memory ProcessingabstractDIMM-based near-memory processing (NMP) architectures address the “memory wall” problem by incorporating near-memory accelerators (NMAs) into main memory devices for high memory bandwidth and low energy consumption. However, critical challenges prevent efficient asynchronous execution between host and NMAs in DIMM-NMP architectures. Memory controllers (MCs) distributed at the host side and the NMA side issue memory accesses independently without synchronization on memory states, which may lead to memory bus contention and DRAM errors. Therefore, most existing DIMM-NMP designs adopt synchronous execution to prevent concurrent memory accesses. However, this intervention wastes either the host or the NMA computation capability. In this work, we propose AsyncDIMM, a novel DIMM-NMP design with efficient asynchronous execution based on existing memory buses. It enables single access mode (host or NMA), concurrent access mode, and a seamless switch between them. First, we propose the offload-schedule-return mechanism with explicit and implicit synchronization to ensure memory access correctness for all memory modes. Second, to further improve bandwidth utilization and decrease access latency, we introduce optimized timing constraints for offloading, a locality-aware switch-recovery method for scheduling, and adaptive batch with timing-division multiplexing notification for returning. Finally, we present a detailed design with limited hardware modifications to conventional host and NMA MCs, which is extensively validated on the FPGA. Comprehensive experiments demonstrate that AsyncDIMM outperforms four NMP baselines by $1.19 \times-1.92 \times$, enabling efficient asynchronous execution with up to $2.25 \times$ bandwidth utilization uplift and 47% access latency reduction. Dongxu Lyu, Jianfei Jiang 0001, Qin Wang 0009, Zhigang Mao, Naifeng Jing |
HPCA | 5 |
| 2025 | Low-Power Multiplier Designs by Leveraging Correlations of 2$\times$×2 Encoded Partial ProductsabstractMultipliers, particularly those with small bit widths, are essential for modern neural network (NN) applications. In addition, multiple-precision multipliers are in high demand for efficient NN accelerators; therefore, recursive multipliers used in low-precision fusion schemes are gaining increasing attention. In this work, we design exact recursive multipliers based on customized approximate full adders (AFAs) for low-power purposes. Initially, the partial products (PPs) encoded by 2×2 multiplications are analyzed, which reveals the correlations among adjacent PPs. Based on these correlations, we propose 4×4 recursive multiplier architectures where certain full adders (FAs) can be simplified without affecting the correctness of the multiplication. Manually and synthesis tool-based FA simplifications are performed separately. The obtained 4×4 multipliers are then used to construct 8×8 multipliers based on a low-power recursive architecture. Finally, the proposed signed and unsigned 4×4 and 8×8 multipliers are evaluated using a 28nm CMOS technology. Compared with DesignWare (DW) multipliers, the proposed signed and unsigned 4×4 multipliers achieve power reductions of 16.5% and 11.6%, respectively, without compromising area or delay; alternatively, the delay can be reduced by 20.9% and 39.4%, respectively, without compromising power or area. For signed and unsigned 8×8 multipliers, the maximum power reductions are 9.7% and 13.7%, respectively, albeit with a trade-off in area. Siting Liu 0001, Hui Wang 0023, Qin Wang 0009, Fabrizio Lombardi, Zhigang Mao, Honglan Jiang |
IEEE Trans. Computers | 6 |
| 2025 | Bridge-NDP: Efficient Communication-Computation Overlap in Near Data Processing SystemabstractNear data processing (NDP), enabled by near data accelerators (NDAs) within DIMM-based main memory, enhances performance by providing more aggregated bandwidth and reducing long-distance data transfers. While the performance of NDAs has received widespread attention, the overhead of host-NDA communication has been overlooked, becoming a bottleneck in NDP systems. To alleviate performance degradation from communication, we propose Bridge-NDP, the first NDP architecture that implements a workflow with efficient communication-computation overlap. Bridge-NDP is built upon the conventional NDP architecture and can be easily applied to existing NDP designs, regardless of the memory level where NDAs are attached. Specifically, we introduce a novel direct host-NDA communication method that utilizes existing memory buses as bridge buses, avoiding the need for new interconnections. It enables seamless integration with other memory accesses while achieving high bandwidth utilization with minimal hardware overhead. For the system-level workflow design, we optimize and extend existing dataflow to achieve richer computing paradigms with fewer redundant memory accesses. Additionally, we provide programming support with efficient API designs and data management to hide low-level resource details and ensure correctness guarantees. Comprehensive experiments demonstrate that Bridge-NDP achieves significant performance improvements, with speedups of$1.8\times $–$3.1\times $and bandwidth utilization improvement of$2.0\times $–$2.9\times $over the state-of-the-art NDP solutions. Pengyu Liu 0004, Dongxu Lyu, Jianfei Jiang 0001, Qin Wang 0009, Zhigang Mao, Naifeng Jing |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2025 | MACS: A Multidomain Collaborative Adaptive Clock Scheme for Large-Scale Reconfigurable Dataflow AcceleratorsabstractTo guarantee reliability and correctness, VLSI circuits are designed with conservative margins to maintain timing and power integrity against process, voltage, and temperature (PVT) variations across diverse workloads. However, worst-case PVT and workload conditions rarely occur in practice, resulting in significant timing slack and hence performance and energy loss, especially in reconfigurable dataflow accelerator RDA due to their large-scale and configurable features. Previous studies have attempted to exploit workload or PVT slack, yet achieving limited benefits for reconfigurable dataflow accelerator (RDAs) with large-scale processing element PE arrays. The key issues come from restricted scaling ranges for the clock, insufficient representations for the workload, and unbalanced workloads within processing elementss (PEs). To address these challenges, this article proposes the first multidomain collaborative adaptive clock scheme (MACS) to efficiently exploit both the workload and PVT timing slack for large-scale reconfigurable dataflow acceleratorss (RDAs). MACS partitions the RDA into several clock domains and allows constrained clock domain crossing, which enhances the hardware efficiency with minimal overhead and supports timing validation using conventional static timing analysis (STA) tools. In each domain, an operand-aware workload detection unit is developed, using both static configurations and dynamic operands to assess workload. The detected workload, combined with the monitored PVT conditions, determines the subsequent clock period. Additionally, to enable the exploration of timing slack over a broader range, the period range of the adaptive clock is extended. Experimental results show that MACS achieves a performance improvement of 76.3% or an energy saving of 36.6% with a hardware cost of 3.5%. Shuya Ji, Weidong Yang 0007, Jianfei Jiang 0001, Naifeng Jing, Honglan Jiang, Zhigang Mao, Qin Wang 0009 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2025 | A Hierarchical 3-D Physical Design Method for Ultralarge-Scale Logic-on-Memory CGRA ChipabstractFace-to-face bonded 3-D (F2F 3D) technology, with the potential to significantly reduce chip area while enhancing performance, stands as one of the most promising ways to extend Moore’s Law. However, current 3-D physical design flows are often modifications of 2-D design flows and rely on technical personnel to manually modify technical files. Furthermore, existing research on 3-D design flow primarily focuses on module implementation, with very few studies addressing hierarchical design methods for large-scale chips. In this article, we first introduce a 3-D physical design flow which concurrently optimizes the timing of both the logic tier and the memory tier, achieving synchronized physical design for both tiers. Then, we develop a bottom-up hierarchical 3-D physical design flow to extend the 3-D design flow to large-scale chip design. Through coordinated power planning, clock tree design, and interconnect unit design, we enhance the power, performance, and area (PPA) metrics of the entire chip. Using our RTL-to-GDS physical design flow, we successfully implemented a 28-nm CMOS logic-on-memory (LoM) 3-D coarse-grained reconfigurable architecture (CGRA) chip with over 50 million gates. Experimental results demonstrate that our 3-D flow improves timing by 16.1% while reducing voltage drop by 38.6% compared to the 2-D design. In addition, the power-delay product (PDP) of the 3-D chip decreases by 10.2%, showcasing better performance. Zizheng Dong, Shuaipeng Li, Weijia Zhu, Ang Li 0045, Qin Wang 0009, Naifeng Jing, Weiguang Sheng, Jianfei Jiang 0001, Zhigang Mao |
IEEE Trans. Very Large Scale Integr. Syst. | 9 |
| 2024 | Bridge-NDP: Achieving Efficient Communication-Computation Overlap in Near Data Processing with Bridge ArchitectureabstractNear data accelerators (NDAs) enable near data processing (NDP) within main memory that benefits performance by providing more aggregated bandwidth and reducing long-distance data transfer. Most prior works focus on reaping higher internal bandwidth to improve performance of the NDA itself. However, the overhead of interactive communication between host and NDAs is overlooked, which has become the bottleneck of NDP systems. In this paper, we propose bridge-NDP, a novel NDP architecture that exploits existing memory buses serving as bridge buses to fully utilize bandwidth. With bridge access enabled by optimized bridge commands, bridge-NDP efficiently overlaps communication and computation. It can be applied to existing NDP systems regardless of the memory level NDAs are attached to. For a variety of key computing kernels from machine learning, data analytics, etc., our evaluation shows that bridge-NDP speeds up not only the NDA performance itself (1.13×-3.62×), but also the host-NDA collaboration performance (2.43×-4.21×), achieving more bandwidth utilization (1.12×-3.67× and 1.48×-4.13×) over the state-of-the-art NDP solution. Jianfei Jiang 0001, Qin Wang 0009, Zhigang Mao, Naifeng Jing |
ASPDAC | 4 |
| 2024 | SparGNN: Efficient Joint Feature-Model Sparsity Exploitation in Graph Neural Network AccelerationabstractWith the rapid explosion in both graph scale and model size, accelerating graph neural networks (GNNs) at scale encounters significant pressure on computation and memory footprint. Exploiting data sparsity with pruning, which exhibits remarkable effect in deep neural networks (DNNs), while still lags behind in GNN acceleration. This is because costly pruning overhead upon large graphs and inefficient hardware support will eclipse the benefit of GNN sparsification. To this end, this paper proposes SparGNN, an algorithm and accelerator co-design that can efficiently exploit data sparsity in both features and models to speedup GNN acceleration while reserving its accuracy. In algorithm, to reduce the overhead of iterative pruning, we distill a sparsified subgraph to substitute the original input graph for pruning, which can low-costly excavate the potential data sparsity in both features and models without accuracy compromise. In hardware, to improve data locality of the sparsified feature-weight multiplication, we design compressed row-/column-wise product dataflow for efficient feature updating. We then propose lightweight hardware changes to make our design applicable to conventional GNN accelerators. The experimental results show that compared to the state-of-the-art GNN accelerators, SparGNN reduces $1.5 \sim 4.3 \times$ computation and gains an average of 1.8 6.8 $\times$ speedup with $1.4 \sim 9.2 \times$ energy efficiency improvement. Jianfei Jiang 0001, Qin Wang 0009, Zhigang Mao, Naifeng Jing |
ASPDAC | 4 |
| 2024 | A novel vehicle collision detection system: Integrating audio-visual fusion for enhanced performance
Kunyue Li, Zhengji Zhao, Qixuan Cai, Qin Wang 0009, Naifeng Jing, Zhigang Mao, Jianfei Jiang 0001 |
Expert Syst. Appl. | 6 |
| 2024 | A Comprehensive Dataflow-Mapping Optimization for Fully Pipelined Execution in Spatial Programmable ArchitectureabstractAlthough spatial programmable architectures have demonstrated high-performance and programmability for a variety of applications, they suffer from the pipeline unbalancing issue which restricts resource utilization and degrades the performance. In this paper, we identify that spatial initiation interval (SpII) can quantitatively describe the impact of pipeline unbalancing on performance, so we formulate SpII for the first time in spatial architectures. To achieve an optimal SpII, we propose dataflow decomposing and integrated mapping to enable high performance dataflow-mapping on spatial architectures. Dataflow decomposing decomposes the application graph into subgraphs and runs them serially, so that it adapts the regular spatial architecture to various application dataflows, particularly for extremely unbalanced datapaths without incurring large buffering overhead. Based on the quantitative SpII, we propose integrated mapping to consider operator placing, operand routing and pipeline balancing at the same time that can find a better SpII for fully-pipelined execution on spatial architectures. The experiment results show that our proposal can gain an average of 2.1× performance speedup on a variety of application kernels over the state-of-the-art approaches. Pengyu Liu 0004, Ang Li 0045, Jianfei Jiang 0001, Qin Wang 0009, Zhigang Mao, Naifeng Jing |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2024 | RecPIM: Efficient In-Memory Processing for Personalized Recommendation Inference Using Near-Bank ArchitectureabstractDeep learning (DL)-based personalized recommendation systems consume the major resources in modern AI data centers. The embedding layers with large memory capacity requirement and high bandwidth demand have been identified as the bottleneck of personalized recommendation inference. To mitigate the memory bandwidth bottleneck, near-memory processing (NMP) would be an effective solution which utilizes the through-silicon via (TSV) bandwidth within 3D-stacked DRAMs. However, existing NMP architectures suffer from the limited memory bandwidth caused by hard-to-scale TSVs. To overcome this obstacle, integrating the compute-logic near memory banks becomes a promising but challenging solution, since large memory capacity requirement limits the use of 3D-stacked DRAMs and irregular memory accesses lead to poor data locality, heavy TSV data traffic and low bank-level bandwidth utilization. To address this problem, we propose RecPIM, the first in-memory processing system for personalized recommendation inference using near-bank architecture based on 3D-stacked memory. From the hardware perspective, we introduce a heterogeneous memory system combined with 3D-stacked DRAM and DIMMs to accommodate large embedding tables and provide high bandwidth. By integrating processing logic units near memory banks on DRAM dies, our architecture can exploit the enormous bank-level bandwidth which is much higher than TSV bandwidth. Then, we integrate a small scratchpad memory to exploit the unique data reusability of DL-based personalized recommendation systems. Furthermore, we adopt a unidirectional data communication scheme to avoid additional cross-vault data transfer. From the software perspective, we present a customized programming model to facilitate memory management and task offloading. To reduce the data communication through TSVs and enhance the utilization of bank-level bandwidth, we develop an efficient data mapping scheme by partitioning the vector into smaller subvectors. Experimental results show that RecPIM achieves up to 2.58× speedup and 49.8% energy saving for data movement over the state-of-the-art NMP solution. Weidong Yang 0007, Shuya Ji, Jianfei Jiang 0001, Naifeng Jing, Qin Wang 0009, Zhigang Mao, Weiguang Sheng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2024 | DeltaGNN: Accelerating Graph Neural Networks on Dynamic Graphs With Delta UpdatingabstractGraph neural network (GNN) accelerators have achieved prominent performance speedup on static graphs but fallen with inefficiency on dynamic graphs. The reason is that in dynamic graphs, updating on a few vertices will introduce enormous redundant neighbor reaggregation and feature reupdating. Moreover, evolving graph structure makes graph preprocessing impractical and incurs random memory accesses which can only be determined at runtime. In this article, we propose DeltaGNN, an algorithm and accelerator co-design for GNN acceleration on dynamic graphs. In algorithm, we first propose a delta updating algorithm, which identifies the sensitivity of vertices and reduces the aggregation and updating operations of insensitive vertices without accuracy compromise. In hardware, we propose a novel sensitivity remapping cache to satisfy the dissimilar reusability of vertices under different sensitivity without preprocessing requirement. To tackle the workload imbalance, we implement feature-disperse execution to support different feature updating between sensitive and insensitive vertices. Moreover, we introduce vertex feature coalescing to reduce the amount of feature vectors by exploiting the locality within vertex accesses. We then propose lightweight yet effective hardware optimizations to make our design applicable to conventional GNN accelerators. Compared to the state-of-the-art GNN accelerators, our DeltaGNN gains an average of$1.5\times $–$11.8\times $speedup and$1.3\times $–$8.6\times $energy efficiency improvement on dynamic graphs. Jianfei Jiang 0001, Qin Wang 0009, Zhigang Mao, Naifeng Jing |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2024 | 3A-ReRAM: Adaptive Activation Accumulation in ReRAM-Based CNN AcceleratorabstractReRAM-based computing is good at accelerating convolutional neural network (CNN) inference due to its high computing parallelism, but its rigid crossbar structure may become less efficient in the face of the random data sparsity abundant in CNNs. In this study, we propose$3A$-ReRAM, a novel crossbar architecture that can dynamically predict the accumulated results to enable adaptive activation accumulation, so that both zero and small values in feature map can be exploited in each matrix-vector multiplication (MVM) operation for speedup. To dynamically predict the results, we propose an efficient parallel predictor to find larger adapted boxes for increased computing parallelism without hurting accuracy. For a better scheduling between the dynamic predictions, we propose an efficient input window management with light-weight hardware support. With dynamic prediction and calculation,$3A$-ReRAM architecture naturally fits the ReRAM crossbar structure but enables a totally different way to dynamically exploit the sparsity and small values in feature maps. It greatly improves the performance by increasing the computing parallelism and saves energy consumption by much less analog-digital conversions. The evaluation results show that$3A$-ReRAM architecture can increase the performance by up to$13.03\times $,$16.31\times $,$2.46\times $, and$2.58\times $compared to ReRAM-based CNN accelerators ISAAC, PUMA (sparsity-unaware) and SRE, FORMS (sparsity-aware), and the total energy can be reduced by$8.93\times $,$10.07\times $,$2.97\times $, and$4.58\times $, respectively. Jianfei Jiang 0001, Qin Wang 0009, Zhigang Mao, Naifeng Jing |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2023 | An Efficient near-Bank Processing Architecture for Personalized Recommendation SystemabstractPersonalized recommendation systems consume the major resources in modern AI data centers. The memory-bound embedding layers with irregular memory access patterns have been identified as the bottleneck of recommendation systems. To overcome the memory challenges, near-memory processing (NMP) would be an effective solution which provides high bandwidth. Recent work proposes an NMP approach to accelerate the recommendation models by utilizing the through-silicon via (TSV) bandwidth in 3D-stacked DRAMs. However, the total bandwidth provided by TSVs is insufficient for a batch of embedding layers processed in parallel. In this paper, we propose a near-bank processing architecture to accelerate recommendation models. By integrating the compute-logic near memory banks on DRAM dies of the 3D-stacked DRAM, our architecture can exploit the enormous bank-level bandwidth which is much higher than TSV bandwidth. We also present a hardware/software interface for embedding layers offloading. Moreover, we propose an efficient mapping scheme to enhance the utilization of bank-level bandwidth. As a result, our architecture achieves up to 2.10X speedup and 31% energy saving for data movement over the state-of-the-art NMP solution for recommendation acceleration based on 3D-stacked memory. Weidong Yang 0007, Qin Wang 0009, Naifeng Jing, Jianfei Jiang 0001, Zhigang Mao, Weiguang Sheng |
ASP-DAC | 6 |
| 2023 | Pipeline Balancing for Integrated Mapping in High Performance Spatial Programmable ArchitectureabstractRecently, spatial programmable architectures have gained increasing popularity owing to their performance and programmability, while the achievable performance is highly related to how the operators are mapped onto a number of processing elements (PEs) in the spatial architectures. In this paper, we first identify that in the spatial mapping process, the pipeline balancing problem is essential by affecting the spatial initial interval (SpII). Hence, we formulate the SpII for the first time in spatial architecture. The quantitative formulation enables an integrated mapping algorithm which combines operator placement, operand routing and pipeline balancing at the same time. In addition, to reduce the balancing hardware cost, we propose a bridge-buffer structure to facilitate operand routing and buffering on demand. To reduce the mapping searching space, we propose three optimization techniques to trade off solution quality, mapping time and hardware overhead. The experiment results show that the proposed integrated mapping algorithm can reduce the SpII by 42.3%, which in turn improves the throughput and algorithm runtime up to 1.74× and 3.08× over the state-of-the-art heuristic spatial mapping. Pengyu Liu 0004, Jianfei Jiang 0001, Qin Wang 0009, Zhigang Mao, Naifeng Jing |
FPL | 7 |
| 2023 | RTMDet-R2: An Improved Real-Time Rotated Object Detector
Haifeng Xiang, Naifeng Jing, Jianfei Jiang 0001, Weiguang Sheng, Zhigang Mao, Qin Wang 0009 |
PRCV (12) | 6 |
| 2023 | Approximate Processing Element Design and Analysis for the Implementation of CNN Accelerators
Honglan Jiang, Hai Mo, Jie Han 0001, Leibo Liu, Zhigang Mao |
J. Comput. Sci. Technol. | 6 |
| 2023 | CDAR-DRAM: Enabling Runtime DRAM Performance and Energy Optimization via In-Situ Charge Detection and Adaptive Data RestorationabstractWith the increasing of dynamic random access memory’s (DRAM) capacity, the refresh operation rapidly becomes a major concern to the performance of the current computational system. Moreover, conservative timing parameters adopted for access operations make an increasing amount of negative impact on system performance and energy efficiency. In this article, we propose an in-situ charge detection and adaptive data restoration DRAM (CDAR-DRAM) architecture, which can dynamically adjust the refresh rate and relax the constraints on access timing by removing pessimistic timing margins for PVT variations. CDAR-DRAM employs a low-cost skewed-inverter-based detector to monitor the bitline voltage in runtime and estimate real-time timing parameters of cells. Based on the detector, an adaptive refresh and restore scheme (CDAR-ref) is presented, which progressively reduces the refresh rate and partially restores cells’ voltage just enough for cells with sufficient charge, thereby optimizing both refresh and restoration operations. Moreover, a supplementary adaptive access scheme (CDAR-acc) is presented, which detects the runtime charge level of recently accessed rows and reduces access latency aggressively, benefitting workloads in a single-core system and memory nonintensive workloads in a multicore system. CDAR’s flexibility allows the two schemes to be combined. The evaluation shows that in an eight-core system, the combined scheme improves performance and energy efficiency by 15.2% and 22.6%, respectively. Yuxuan Qin, Chuxiong Lin, Weifeng He, Yanan Sun 0003, Zhigang Mao, Mingoo Seok |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2023 | A Reschedulable Dataflow-SIMD Execution for Increased Utilization in CGRA Cross-Domain AccelerationabstractWhen a coarse-grained reconfigurable array (CGRA) architecture shifts toward cross-domain acceleration, control flow and memory accesses often degrade the processing elements (PEs) utilization and array efficiency by breaking the intact dataflow graph (DFG) into regions with mismatched pipelining rate and access–execution stages. In this article, we propose a reschedulable dataflow and SIMD execution, which decouples the DFG with mismatched dataflow into multiple independent subgraphs. We map only one subgraph at a time but with fully unrolling, and reschedule different subgraphs serially in the runtime. Therefore, each subgraph works in its own way without interfering with others. At the same time, an individual subgraph can execute its dataflow in stream for utilization improvement, while unrolled instances composing as SIMD facilitate request coalescing for efficient memory access. With lightweight hardware modification, our design can be integrated in a general CGRA architecture. The experimental results show that our proposal improves the performance and energy efficiency over stream-dataflow CGRA in static-scheduling (Plasticine) by$1.6\times $and$1.8\times $, over which in dynamic scheduling (TIA) by$1.5\times $and$2.7\times $, and outperforms Plasticine organized in vector-SIMD by$1.2\times $and$1.4\times $. Naifeng Jing, Jianfei Jiang 0001, Qin Wang 0009, Zhigang Mao |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2023 | BC-MVLiM: A Binary-Compatible Multi-Valued Logic-in-Memory Based on Memristive CrossbarsabstractLogic-in-memory with memristive crossbars is an attractive approach for realizing beyond von Neumann architectures. Multi-valued logic (MVL) containing more than two logic levels can enhance the computing speed with reduced number of logic operations. In this paper, a binary-compatible multi-valued logic-in-memory (BC-MVLiM) scheme is proposed with memristive dual-crossbars where both inputs and outputs are represented by the multi-level cells of memristors. Both of the binary and multiple-valued logic operations can be implemented in the proposed BC-MVLiM scheme depending on the radix of inputs. The proposed BC-MVLiM circuitry supports multiple row-wise and column-wise logic gates with multiple fan-ins and fan-outs for binary and ternary systems by leveraging both inter- and intra-crossbar operations. Experimental results show that the proposed BC-MVLiM-based multi-digit adder enhances the computation speed by up to 76.10% and 83.82%, for binary and ternary systems, respectively, as compared with the previously published memristive logic designs. By preventing the errors from propagating across multiple stages, the error rate of proposed BC-MVLiM is also reduced by up to 98.20% compared to the previous memristive logic designs in the presence of device variations. Yanan Sun 0003, Zhi Li 0058, Weifeng He, Qin Wang 0009, Zhigang Mao |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2022 | A Low Coupling and Lightweight Algorithm for Ship Detection in Optical Remote Sensing ImagesabstractIn recent years, many ship detection algorithms based on convolutional neural networks (CNNs) have been proposed to improve the performance of ship detection. However, with the increase in model complexity and size, it is challenging to deploy these models to resource-constrained edge platforms. In this letter, a low coupling algorithm that belongs to anchor-free methods is proposed for ship detection to reduce the model complexity and still obtain a competitive performance. The proposed low coupling network (LCNet) is easy to deploy and contributes to speeding up the inference and improving memory utilization. In addition, we propose a model compression process consisting of the quantization-aware training (QAT) method and a structural pruning method based on Taylor expansion, which can effectively reduce the model size according to hardware resource constraints. Comparative experimental results demonstrate that LCNet outperforms the state-of-the-art ship detection and natural object detection algorithms, with a 95.27% mAP and 88.91% F1 score on the HRSC2016 dataset. Our proposed model compression method also achieves a compression ratio of at least 80% with a negligible loss of performance. Guochao Deng, Qin Wang 0009, Jianfei Jiang 0001, Qirun Hong, Naifeng Jing, Weiguang Sheng, Zhigang Mao |
IEEE Geosci. Remote. Sens. Lett. | 7 |
| 2022 | A Hybrid-Grained Remapping Defense Scheme Against Hard Failures for Row-Column-NVMabstractRow-column-NVM (RC-NVM) is a new architecture for emerging nonvolatile memory (NVM), such as ReRAM, PCM, and STT-RAM. It leverages the symmetry of crossbar structure and supports both row and column memory accesses. The new architecture is well fit for the applications with different access patterns which suffer from low efficiency and high-energy consumption in traditional memory architecture. However, existing hard failure defensive techniques for emerging NVM are inappropriate for RC-NVM. Therefore, the limited device endurance makes RC-NVM ephemeral and unreliable. In order to overcome this challenge, we propose a hybrid-grained remapping defensive technique to fight against hard failures. By hybrid-grained remapping, we can effectively avoid multiple reads issue of fine-grained remapping and low utilization issue of coarse-grained remapping and significantly increase RC-NVM lifetime with little performance loss. Moreover, motion vector and remap-aware write optimizations are also proposed to further improve RC-NVM reliability and degrade performance loss and energy consumption of write operation. An evaluation shows that the hybrid-grained remapping can increase the lifetime by 61.1% compared to single coarse-grained remapping with only 5% performance loss and 7% energy increase when 10% pages are mapped out. Optimizations can further improve lifetime, performance, and energy by 5%, 11.4%, and 30.5%, respectively. Taozhong Li, Naifeng Jing, Zhigang Mao, Yiran Chen 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2022 | A Universal RRAM-Based DNN Accelerator With Programmable Crossbars Beyond MVM OperatorabstractResistive-RAM (RRAM)-based deep neural network (DNN) accelerator has shown a great potential as it is good at the matrix–vector multiplication (MVM) operator. However, it does not benefit non-MVM operators, such as transcendental activation or elementwise operations, which often require customized CMOS circuits in conventional DNN accelerator designs. In this article, we propose a new RRAM-based DNN inference accelerator, which leverages the proposed RRAM-CORDIC and RRAM-MLP algorithms to make the transcendental and elementwise operators calculable in the RRAM crossbar just like MVM. Both algorithms can exploit the higher multiply-and-accumulation (MAC) parallelism that is traditionally expensive in CMOS but now efficient in the RRAM crossbar. Then, we further propose an intercrossbar pipelining scheme, which can balance the number of crossbars for MVM and non-MVM operations and orchestrate them in pursuing higher DNN computing throughput. The experimental results show that both algorithms can sustain a high arithmetic accuracy and deliver less than 1% DNN accuracy loss on typical inference workloads. The elimination of expensive CMOS circuits, in turn, can trade more crossbar resources in the same area to speed up the performance by$1.16\times $to$2.33\times $. With the extended operators, the RRAM-based DNN accelerator can switch crossbar functions at will, and apply for a diverse of DNN models in a unified in-memory accelerator architecture. Jianfei Jiang 0001, Yongxin Zhu 0001, Qin Wang 0009, Zhigang Mao, Naifeng Jing |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2022 | MSLM-RF: A Spatial Feature Enhanced Random Forest for On-Board Hyperspectral Image ClassificationabstractHyperspectral imaging (HSI) greatly improves the capacity to identify and monitor ground objects due to the high spectral resolution. As the real-time remote sensing monitoring and warning tasks are getting more attention, new algorithms for low-power on-board classification are required to reduce the transmission time of satellite downlink. In this paper, we propose the Multi-Scale Local Maximum Random Forest (MSLM-RF) to significantly reduce the energy consumption while retaining high classification accuracy. The proposed MSLM-RF uses multi-scale maximum filters for spatial feature extraction and Random Forest for classification after spectral and spatial features fusion. The spatial features are efficiently extracted with low computational complexity by regarding the maximum light intensity values in different ranges of pixels as anchor points. MSLM-RF only consists of integer comparisons and a few additions, thereby eliminating the energy-hungry operations such as multiplication and exponentiation. According to experimental results on the HSI benchmark datasets, MSLM-RF delivers a better trade-off in accuracy and computational complexity than the state-of-the-art classification algorithms. Besides, MSLM-RF gets higher average classification accuracy and lower energy consumption than the previous on-board algorithms. The obtained results show the suitability of the proposed algorithm to accomplish practical real-time classification tasks on-board with low energy consumption. Shuai Yuan 0016, Yanan Sun 0003, Weifeng He, Qianrong Gu, Zhigang Mao, Shikui Tu |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2022 | A Novel Architecture Design for Output Significance Aligned Flow with Adaptive Control in ReRAM-based Neural Network AcceleratorabstractResistive-RAM-based (ReRAM-based) computing shows great potential on accelerating DNN inference by its highly parallel structure. Regrettably, computing accuracy in practical is much lower than expected due to the non-ideal ReRAM device. Conventional computing flow with fixed wordline activation scheme can effectively protect computing accuracy but at the cost of significant performance and energy savings reduction. For such embarrassment of accuracy, performance and energy, this article proposes a new Adaptive-Wordline-Activation control scheme ( AWA-control ) and combines it with a theoretical Output-Significance-Aligned computing flow ( OSA-flow ) to enable fine-grained control on output significance with distinct impact on final result. We demonstrate AWA-control -supported OSA-flow architecture with maximal compatibility to conventional crossbar by input retiming and weight remapping using shifting registers to enable the new flow. However, in contrast to the conventional computing architecture, the OSA-flow architecture shows the better capability to exploit data sparsity commonly seen in DNN models. So we also design a sparsity-aware OSA-flow architecture for further DNN speedup. Evaluation results show that OSA-flow architecture can provide significant performance improvement of 21.6×, and energy savings of 96.2% over conventional computing architecture with similar DNN accuracy. Taozhong Li, Naifeng Jing, Jianfei Jiang 0001, Qin Wang 0009, Zhigang Mao, Yiran Chen 0001 |
ACM Trans. Design Autom. Electr. Syst. | 5 |
| 2022 | An Efficient CNN Accelerator Using Inter-Frame Data Reuse of Videos on FPGAsabstractConvolutional neural networks (CNNs) have had great success when applied to computer vision technology, and many application-specific integrated circuit (ASIC) and field-programmable gate array (FPGA) CNN accelerators have been proposed. These accelerators primarily focus on the acceleration of a single input, and they are not particularly optimized for video applications. In this article, we focus on the similarities between continuous inputs in video, and we propose a YOLOv3-tiny CNN FPGA accelerator using incremental operation. The accelerator can skip the convolution operation of similar data between continuous inputs. We also use the Winograd algorithm to optimize the conv$3\times 3$operator in the YOLOv3-tiny network to further improve the accelerator’s efficiency. Experimental results show that our accelerator achieved 74.2 frames/s on ImageNet ILSVRC2015. Compared to the original network without Winograd algorithm and incremental operation, our design provides a$4.10\times $speedup. When compared with other YOLO network FPGA accelerators applied to video applications, our design provided a$3.13\times $–$18.34\times $normalized digital signal processor (DSP) efficiency and$1.10\times $–$14.2\times $energy efficiency. Shengzhao Li, Qin Wang 0009, Jianfei Jiang 0001, Weiguang Sheng, Naifeng Jing, Zhigang Mao |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2021 | CDAR-DRAM: An In-situ Charge Detection and Adaptive Data Restoration DRAM Architecture for Performance and Energy Efficiency ImprovementabstractAs the capacity of DRAM continues to grow, the refresh operation rapidly becomes the performance and power-efficiency bottleneck. Also, restore time, the time given for recharging cells post access, makes an increasingly large amount of negative impact on performance. To tackle these problems, in this paper, we propose an in-situ charge detection and adaptive data restoration DRAM (CDAR-DRAM) architecture, which can dynamically adjust the refresh rate and also relax the constraints on restore time. The proposed CDAR-DRAM employs a low-cost skewed-inverter-based detector, which can reduce the excessive timing margins that prior work added to guarantee the functionality of leaky DRAM cells under the worst-case temperature condition. Moreover, an adaptive DRAM refresh and restore scheme is proposed, which can switch automatically between two modes: (i) a refresh mode that supports adaptive refresh rate, and (ii) a restore mode that relaxes the constraints on restore time dynamically for cells having sufficient charge. With the transistor-and architecture-level simulations, we evaluate the CDAR-DRAM in an 8-core system across different workloads. Compared with the prior art, the proposed architecture achieves a 9.4% improvement in system performance and a 14.3% reduction in energy consumption, without requiring the time-consuming profiling process which many prior works employed. Chuxiong Lin, Weifeng He, Yanan Sun 0003, Zhigang Mao, Mingoo Seok |
DAC | 4 |
| 2021 | Reducing Memory Access Conflicts with Loop Transformation and Data Reuse on Coarse-grained Reconfigurable ArchitectureabstractCoarse-Grained Reconfigurable Arrays (CGRAs) are promising to have low power consumption and high energy-efficiency characteristics as accelerators. Recent years, many research works focus on improving the programmability of the CGRAs by enabling the fast reconfiguration during execution. The performance of these CGRAs critically hinges upon the scheduling power of the compiler. One of the critical challenges is to reduce memory access conflicts using static compilation techniques. Memory accessing conflict brings the synchronization overhead which causes the pipelining stall and reduces CGRA performance. Existing compilers usually tackle this challenge by orchestrating the data placement of the on-chip global memory (OGM) in CGRA to let the parallel memory accesses avoid the bank conflict. However, we find bank conflict is not the only reason that causes the memory access conflicts. In some CGRAs, the bandwidth of the data network between OGM and processing element array (PEA) is also limited due to the low power design principle. The unbalanced network bandwidth loads is another reason that causes memory access conflicts. Furthermore, the redundant data access across iterations is one of the primary causes of memory access conflicts. Based on these observations, we provide a comprehensive and generalized compilation flow to reduce the memory conflicts. Firstly, we develop a loop transformation model to maximize the inter-iteration data reuse of the loops to reduce the memory accessing operations under the software pipelining scheme. Secondly, we enhance the bandwidth utilization of the network between OGM and PEA and avoid the bank conflict by providing a conflict-aware spatial mapping algorithm which can be easily integrated into existing CGRA modulo scheduling compilation flow. Experimental results show our method is capable of improving performance by an average of 44% comparing with state-of-the-art CGRA compiling flow. Yuge Chen, Zhongyuan Zhao 0004, Jianfei Jiang 0001, Guanghui He 0002, Zhigang Mao, Weiguang Sheng |
DATE | 5 |
| 2021 | Subgraph Decoupling and Rescheduling for Increased Utilization in CGRA ArchitectureabstractWhen coarse-grained reconfigurable array (CGRA) architecture is shifting towards general-purpose, some complex control flows, such as nested loop, conditional branch and data dependence, may embarrass it and reduce the processing element (PE) array utilization by breaking the intact dataflow graph (DFG) into multiple regions with inconsistent control regions. This paper proposes subgraph decoupling and rescheduling, which decouples the inconsistent regions into control-independent subgraphs. Each subgraph can be rescheduled with zero-cost domino context switching and parallelized to fully utilize the PE resources. Then, we propose lightweight hardware changes based on general CGRA architecture to enable our design. The experiment results show that our proposal can improve the performance and energy efficiency by 1.35× and 1.18× over a static-mapped CGRA (Plasticine), and by 1.27× and 1.45× over an instruction-driven CGRA (TIA). Qin Wang 0009, Jianfei Jiang 0001, Weiguang Sheng, Guanghui He 0002, Zhigang Mao, Naifeng Jing |
DATE | 6 |
| 2021 | A 3.85-Gb/s 8 × 8 Soft-Output MIMO Detector With Lattice-Reduction-Aided Channel PreprocessingabstractThis article presents an 8 × 8 lattice-reduction-aided (LRA) soft-output multiple-input multiple-output (MIMO) detector for Chinese enhanced ultrahigh throughput (EUHT) wireless local area network (LAN) standard. The preprocessing algorithm combining simplified-sorting Cholesky decomposition and low-complexity decoupled lattice reduction (LDLR) is proposed to reduce computational complexity and latency with parallelism improvement. In addition, K-best detection adopts a sorting-reduced strategy utilizing approximate ordered sequence. Compared with other published LRA K-best detection algorithms, simulation results show that our proposed algorithm has performance improvement. In addition, in order to save hardware resources, a folded K-best architecture and an optimized intermediate storage strategy are introduced. Furthermore, a fully pipelined VLSI architecture is designed in Semiconductor Manufacturing International Corporation (SMIC) 40-nm 1P9M technology to support the 8 × 8.64 -QAM MIMO-OFDM system. The detector can achieve 3.85-Gb/s data throughput at 641-MHz clock frequency with 0.71-μs latency. The proposed detector is competitive in terms of latency, throughput, and area efficiency to state-of-the-art works and can meet the data-rate requirement of the EUHT standard. Zhuojun Liang, Dongxu Lv, Chao Cui, Haibao Chen, Weifeng He, Weiguang Sheng, Naifeng Jing, Zhigang Mao, Guanghui He 0002 |
IEEE Trans. Very Large Scale Integr. Syst. | 8 |
| 2020 | Enabling Resistive-RAM-based Activation Functions for Deep Neural Network AccelerationabstractThe Resistive-RAM (RRAM) based deep neural network (DNN) accelerators have shown great potential as they are good at solving matrix-vector multiplication (MVM). However, this computing paradigm does not benefit other NN operations like activation, which may be built upon various transcendental functions and require customized circuit as in current RRAM-based NN accelerators. In this paper, we propose the RRAM-CORDIC algorithm and crossbar design which enable various transcendental activation calculations on a RRAM crossbar just like MVM. By applying encoding and multi-iteration transformation, the RRAM-CORDIC can exploit higher MAC (multiply-and-accumulation) parallelism that is traditionally uneconomic in CMOS but now efficient in RRAM crossbar. In addition, it can work in a pipelined manner with high computing throughput. Experiment results show that the RRAM-CORDIC algorithm can sustain high accuracy on different transcendental functions, and deliver less than 0.5% NN accuracy loss on typical DNN inference. The elimination of CMOS circuit in turn can trade more computing resources for MVM in the same area budget that improves the performance up to 47% for different networks. Taozhong Li, Ning Guan, Qin Wang 0009, Guanghui He 0002, Weiguang Sheng, Zhigang Mao, Naifeng Jing |
ACM Great Lakes Symposium on VLSI | 7 |
| 2020 | Towards Higher Performance and Robust Compilation for CGRA Modulo SchedulingabstractCoarse-Grained Reconfigurable Architectures (CGRA) is a promising solution for accelerating computation intensive tasks due to its good trade-off in energy efficiency and flexibility. One of the challenging research topic is how to effectively deploy loops onto CGRAs within acceptable compilation time. Modulo scheduling (MS) has shown to be efficient on deploying loops onto CGRAs. Existing CGRA MS algorithms still suffer from the challenge of mapping loop with higher performance under acceptable compilation time, especially mapping large and irregular loops onto CGRAs with limited computational and routing resources. This is mainly due to the under utilization of the available buffer resources on CGRA, unawareness of critical mapping constraints and time consuming method of solving temporal and spatial mapping. This article focus on improving the performance and compilation robustness of the modulo scheduling mapping algorithm for CGRAs. We decomposes the CGRA MS problem into the temporal and spatial mapping problem and reorganize the processes inside these two problems. For the temporal mapping problem, we provide a comprehensive and systematic mapping flow that includes a powerful buffer allocation algorithm, and efficient interconnection & computational constraints solving algorithms. For the spatial mapping problem, we develop a fast and stable spatial mapping algorithm with backtracking and reordering mechanism. Our MS mapping algorithm is able to map loops onto CGRA with higher performance and faster compilation time. Experiment results show that given the same compilation time budget, our mapping algorithm generates higher compilation success rate. Among the successfully compiled loops, our approach can improve 5.4 to 14.2 percent performance and takes x24 to x1099 less compilation time in average comparing with state-of-the-art CGRA mapping algorithms. Zhongyuan Zhao 0004, Weiguang Sheng, Qin Wang 0009, Wenzhi Yin, Pengfei Ye, Jinchao Li, Zhigang Mao |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2019 | mRNA: Enabling Efficient Mapping Space Exploration for a Reconfiguration Neural AcceleratorabstractDeep learning accelerators have emerged to enable energy-efficient and high-throughput inference from edge devices such as self-driving cars and smartphones, to data centers for batch inference such as recommendation systems. However, the actual energy efficiency and throughput of a deep learning accelerator depends on the deep neural network (DNN) loop nest mapping on the processing element array of an accelerator. Moreover, the efficiency of a mapping dramatically changes by the target DNN layer dimensions and available hardware resources. Therefore, the optimal mapping search problem is a non-trivial high-dimensional optimization problem. Although several tools and frameworks exist for compiling to CPUs and GPUs, we lack similar tools for deep learning accelerators. To deal with the optimized mapping search problem in deep learning accelerators, we propose mRNA (mapper for reconfigurable neural accelerators), which automatically searches optimal mappings using heuristics based on domain knowledge about deep learning and an energy/runtime cost evaluation framework. mRNA targets MAERI, a recently proposed open-source deep learning accelerator that provides flexibility via reconfigurable interconnects, to run the unique mappings for each layer generated by mRNA. In realistic machine learning workloads from MLPerf, the optimal mappings identified by mRNA framework provides 15% to 26% lower runtime and 55% to 64% lower energy for convolutional layers and 24% to 67% lower runtime and maximum 67% lower energy for fully connected layers compared to simple reference mappings manually picked for each layer. Zhongyuan Zhao 0004, Hyoukjun Kwon, Sachit Kuhar, Weiguang Sheng, Zhigang Mao, Tushar Krishna |
ISPASS | 5 |
| 2019 | A Novel Resistive Memory-based Process-in-memory Architecture for Efficient Logic and Add OperationsabstractThe coming era of big data revives the Processing-in-memory (PIM) architecture to relieve the memory wall problem that embarrasses the modern computing system. However, most existing PIM designs just put computing units closer to memory, rather than a complete integration of them due to their incompatibility in CMOS manufacturing. Fortunately, the emerging Resistive-RAM (ReRAM) offers new hope to this dilemma owing to its inherent memory and computing capability using the same device. In this article, we propose a ReRAM memory structure with efficient PIM capability of both logic and add operations. It first leverages non-linearity to suppress sneak current and thus sustains high memory density. Using a differential bit cell, it also enables efficient processing of arbitrary logic functions using the same memory cells with non-destructive operations. Then, a novel PIM adder is proposed, which customizes a sneak current path as the carry-chain for fast carry propagation and improves adder performance significantly. In the experiment, the proposed PIM demonstrates higher efficiency in both computing area and performance for logic and addition, which greatly increases the ReRAM PIM applicability for future computable architectures. Taozhong Li, Qin Wang 0009, Yongxin Zhu 0001, Jianfei Jiang 0001, Guanghui He 0002, Jing Jin 0005, Zhigang Mao, Naifeng Jing |
ACM Trans. Design Autom. Electr. Syst. | 7 |
| 2018 | Optimizing the data placement and transformation for multi-bank CGRA computing systemabstractThis paper provides a data placement optimization approach for Coarse-Grained Reconfigurable Architecture (CGRA) based computing platform in order to simultaneously optimize the performance of CGRA execution and data transformation between main memory and multi-bank memory. To achieve this goal, we have developed a performance model to evaluate the efficiency of data transformation and CGRA execution. This model is used for comparing the performances difference when using different data placement strategies. We search for the optimal data placement method by firstly choosing the method which generates the best CGRA execution efficiency from the candidates who can generate the optimal data transformation efficiency. Then we choose the best data placement strategy by comparing the performance of the selected strategy with the one generated through existing multi-bank optimization algorithm. Evaluation shows our approach is capable of optimizing the performance to 2.76x of state-of-the-art method when considering both data-transformation and CGRA execution efficiency. Zhongyuan Zhao 0004, Yantao Liu, Weiguang Sheng, Tushar Krishna, Qin Wang 0009, Zhigang Mao |
DATE | 6 |
| 2018 | Design of Area-Efficient and Highly Reliable RHBD 10T Memory Cell for Aerospace Applications
Jing Guo 0004, Lei Zhu 0004, Huiliang Cao, Chunhua Qi, Xuebing Cao, Liyi Xiao, Zhigang Mao |
IEEE Trans. Very Large Scale Integr. Syst. | 11 |
| 2017 | Reliability analysis of memories suffering MBUs for the effect of negative bias temperature instabilityabstractIn this paper, the effect of negative bias temperature instability (NBTI) on MBUs sensitivity of 65 nm bulk technology memories is analyzed and simulated by Geant4. A MTTF reliability model including NBTI stress time is proposed for memories protected by error correction codes (ECCs). Both cases of scrubbing and nonscrubbing are considered. By using the proposed model, the predicted MTTF results align well with the simulation MTTF results in the radiation environment. Shanshan Liu 0001, Liyi Xiao, Xuebing Cao, Zhigang Mao |
ASP-DAC | 4 |
| 2017 | A static-placement, dynamic-issue framework for CGRA loop acceleratorabstractThis paper presents a static-placement, dynamic-issue (SPDI) framework for the coarse-grained reconfigurable architecture (CGRA) in order to tackle the inefficiencies of the static-issue, static-placement (SISP) CGRA. This framework includes the compiler that statically places the operations and hardware design, a SPDI CGRA, that automatically schedule the operations. We stress on introducing the SPDI CGRA in this paper. This newly designed hardware model adds the token buffer, which is capable of automatically scheduling the operations inside processing elements (PE), along with a router network that can effectively transform and control data flow among the PE array. This design lets the hardware share the responsibility for the compiler, making them cooperate to deal with the issuing, placement and routing problem. Evaluation of our study shows that our framework can reach on average 1.28, 1.30 and 1.33 higher than three state-of-the-art SISP CGRA using REGIMap, RS compile flow and the EPIMap approaches respectively. The area overhead is nearly 0.93% per token buffer entry for each PE relative to SISP CGRA. Zhongyuan Zhao 0004, Weiguang Sheng, Weifeng He, Zhigang Mao, Zhaoshi Li |
DATE | 4 |
| 2017 | A 12-bit 4928 × 3264 pixel CMOS image signal processor for digital still cameras
Wei Jin 0004, Guanghui He 0002, Weifeng He, Zhigang Mao |
Integr. | 4 |
| 2017 | Dynamic data split: A crosstalk suppression scheme in TSV-based 3D IC
Qin Wang 0009, Zhenyang Chen, Jianfei Jiang 0001, Zheng Guo 0001, Zhigang Mao |
Integr. | 5 |
| 2017 | Novel Radiation-Hardened-by-Design (RHBD) 12T Memory Cell for Aerospace Applications in Nanoscale CMOS TechnologyabstractIn this paper, a novel radiation-hardened-by-design (RHBD) 12T memory cell is proposed to tolerate single node upset and multiple-node upset based on upset physical mechanism behind soft errors together with reasonable layout-topology. The verification results obtained confirm that the proposed 12T cell can provide a good radiation robustness. Compared with 13T cell, the increased area, power, read/write access time overheads of the proposed 12T cell are -18.9%, -23.8%, and 171.6%/-50.0%, respectively. Moreover, its hold static noise margin is 986.2 mV which is higher than that of 13T cell. This means that the proposed 12T cell also has higher stability when it provides fault tolerance capability. Jing Guo 0004, Lei Zhu 0004, Shanshan Liu 0001, Liyi Xiao, Zhigang Mao |
IEEE Trans. Very Large Scale Integr. Syst. | 8 |
| 2017 | In Situ Error Detection Techniques in Ultralow Voltage Pipelines: Analysis and OptimizationsabstractIn order to achieve high tolerance against process, voltage, and temperature variations in the ultralow voltage (ULV) circuits, in situ error detection and correction (EDAC) techniques were presented. However, circuits adding the capability of error detection incur large hardware overhead, especially in ULV due to larger delay variability. In this paper, we analyze the hardware overhead of error detection techniques in pipelines based on three different sequential elements: flip-flops, two-phase latches, and pulsed latches. By exploiting the cycle-borrowing ability, we propose a technique called sparse insertion of error detecting registers on the two-phase latch-based and pulsed-latch-based pipelines to reduce the sequential logic area. Furthermore, we propose a delay-padding methodology using a multi-Vtcell library in ULV circuits to reduce EDAC hardware overhead. The proposed techniques are applied on a benchmark six-stage pipeline operating at 0.35 V in a 65-nm CMOS. The analysis results show that our proposed techniques can reduce the total area by 26%-33% and the error detecting register count by 2.9-4.3× compared with conventional EDAC techniques. Wei Jin 0004, Seongjong Kim, Weifeng He, Zhigang Mao, Mingoo Seok |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2016 | Enabling in-situ logic-in-memory capability using resistive-RAM crossbar memoryabstractRecently, logic-in-memory (LIM) is gaining growing interest because it eliminates the unnecessary data movement between the memory and logic components that embarrasses both performance and power dissipation in modern microprocessors. However, most of the existing LIM just puts the logic and memory closer rather than a true integration due to the incompatibility of logic and memory circuit structures. In this paper, we propose a in-situ LIM design by leveraging the emerging ReRAM memory in a crossbar structure. It performs in-situ logic processing using the same memory cells based on the resistive states of ReRAMs with non-destructive operations, and therefore can exploit the large internal bandwidth available in the array without data readout. The logic exploration exposes that the proposed LIM can support different logical functions and get in-situ results without moving data in and out of the memory. We believe that the proposed design provides a promising solution for a true logic processing capability within memory. Naifeng Jing, Taozhong Li, Zhongyuan Zhao 0004, Wei Jin 0004, Yanan Sun 0003, Weifeng He, Zhigang Mao |
FPT | 7 |
| 2016 | High performance parallel turbo decoder with configurable interleaving network for LTE application
Zhiting Yan, Guanghui He 0002, Weifeng He, Shuaijie Wang, Zhigang Mao |
Integr. | 5 |
| 2015 | Fault Secure Encoder and Decoder Designs for Matrix CodesabstractTransient multiple cell upsets (MCUs) are becoming major issues in the reliability of memories exposed to radiation environment. Error correction codes (ECCs) are commonly used to protect memories against MCUs. Among ECCs, matrix codes have obvious advantages due to the simplicity of the encoding and decoding algorithm that enables low overheads. However, an important issue is that when ECCs are used, the encoder and decoder circuits also suffer from errors which affect the reliability of the memory systems. In this paper, low overhead fault secure encoder and decoder designs for matrix codes are proposed to protect encoder and decoder. By using the properties of the parity check matrix of matrix codes, the proposed designs efficiently implement a parity prediction scheme with low overheads. They can detect all errors deriving from a single node in encoder and decoder circuits. A fault secure memory system is established and evaluated, and the obtained results show that the proposed scheme has lower area and power overheads. Shanshan Liu 0001, Liyi Xiao, Jing Guo 0004, Zhigang Mao |
CAD/Graphics | 4 |
| 2015 | Redundancy based Interconnect Duplication to Mitigate Soft Errors in SRAM-based FPGAsabstractSoft error induced reliability problem has already become a major concern for modern SRAM-based FPGAs (Field Programmable Gate Arrays) even at the ground level. In this paper, we propose a duplication-with-recovery (DWR) technique to recover the configuration bit faults on interconnects, which contribute to the majority of soft errors in FPGAs. Based on a study on the detailed routing structure in real FPGAs, DWR leverages redundant resources for interconnect duplication and enables fault recovery with lightweight circuit-level support. Compared with traditional fault tolerant techniques, DWR retains the fault recovering capability but eliminates expensive copies. The experimental results show that a large portion of the interconnects can be protected, which in consequence significantly reduces the vulnerable configuration bits. In addition, DWR does not alter the placement and routing from standard design flow, and therefore does not affect the design closure but greatly improves the design reliability in a cost-effective way. Naifeng Jing, Jianfei Jiang 0001, Weifeng He, Zhigang Mao |
ICCAD | 6 |
| 2015 | Parasitic Parameters Impacts Investigation on Soft Error Rate by a Circuit Level FrameworkabstractIn highly reliable CMOS integrated circuits, parasitic parameters have dramatic impacts on SER(soft error rate) estimation and affect the design decision. We proposed a circuit level SER characterization framework(ASSET-SPI) to evaluate the impacts by conducting statistical fault injection experiments automatically on the circuit spice netlist containing parasitic parameters. Experiments on ISCAS benchmark circuits(implemented in 180nm process) demonstrate ASSET-SPI is feasible for circuit level SER evaluation. While experiments on inverter chains(implemented in 180, 130 and 65nm process) show parasitic parameters introduce -2.95% to 19.82% variation on SER. The results remind us that parasitic parameters should be considered in design time SER evaluation to avoid over pessimistic/optimistic SER estimation and inappropriate design decision. Weiguang Sheng, Zhongyuan Zhao 0004, Zhigang Mao |
PRDC | 3 |
| 2015 | Soft Error Hardened Memory Design for Nanoscale Complementary Metal Oxide Semiconductor TechnologyabstractRadiation-induced single event upsets (SEUs), or soft errors, have become a dominant factor in the reliability degradation of nanoscale memories. In this paper, based on the SEU physics mechanism, and reasonable layout-topology, a novel soft error hardened memory cell is proposed in 65 nm Complementary Metal Oxide Semiconductor (CMOS) technology. The design comparisons for several hardened memory cells in terms of access time (read access time and write access time), power consumption, and layout area are also executed. The main advantage of the proposed cell is that it can provide 100% fault tolerance, which is very useful for memory applications in severe radiation environments. Furthermore, Monte Carlo simulations are carried out to evaluate the effects of process, voltage, and temperature (PVT) variations. From simulations, we confirmed that the proposed cell has exhibited a sufficient multiple-node upset tolerance capability even under PVT variations. Jing Guo 0004, Liyi Xiao, Shanshan Liu 0001, Zhigang Mao |
IEEE Trans. Reliab. | 6 |
| 2014 | Area and throughput efficient IDCT/IDST architecture for HEVC standardabstractHigh Efficiency Video Coding (HEVC) is new video coding standard beyond H.264/AVC. In this paper, an area and throughput efficient 2-D IDCT/IDST VLSI architecture for HEVC standard is presented. Adopting proposed data flow scheduling and shared constant multiplication structure, the architecture supports variable block size IDCT from 4×4 to 32×32 pixels as well as 4×4 pels IDST. Using 65nm technology, the synthesis results show that the maximum work frequency is 500MHz and the architecture hardware cost is about 145.4K gate count. Compared with previous work, our design achieves more than 50% reduction in hardware cost and 66% improvement in throughput efficiency. Experimental results show that the proposed architecture is able to deal with real-time HEVC IDCT/IDST of 4K×2K (4096×2048)@30 fps video sequence at 412MHz in average. In consequence, it offers a cost-effective solution for the future UHDTV applications. Ziyou Yao, Weifeng He, Guanghui He 0002, Zhigang Mao |
ISCAS | 5 |
| 2014 | Enhanced Memory Reliability Against Multiple Cell Upsets Using Decimal Matrix CodeabstractTransient multiple cell upsets (MCUs) are becoming major issues in the reliability of memories exposed to radiation environment. To prevent MCUs from causing data corruption, more complex error correction codes (ECCs) are widely used to protect memory, but the main problem is that they would require higher delay overhead. Recently, matrix codes (MCs) based on Hamming codes have been proposed for memory protection. The main issue is that they are double error correction codes and the error correction capabilities are not improved in all cases. In this paper, novel decimal matrix code (DMC) based on divide-symbol is proposed to enhance memory reliability with lower delay overhead. The proposed DMC utilizes decimal algorithm to obtain the maximum error detection capability. Moreover, the encoder-reuse technique (ERT) is proposed to minimize the area overhead of extra circuits without disturbing the whole encoding and decoding processes. ERT uses DMC encoder itself to be part of the decoder. The proposed DMC is compared to well-known codes such as the existing Hamming, MCs, and punctured difference set (PDS) codes. The obtained results show that the mean time to failure (MTTF) of the proposed scheme is 452.9%, 154.6%, and 122.6% of Hamming, MC, and PDS, respectively. At the same time, the delay overhead of the proposed scheme is 73.1%, 69.0%, and 26.2% of Hamming, MC, and PDS, respectively. The only drawback to the proposed scheme is that it requires more redundant bits for memory protection. Jing Guo 0004, Liyi Xiao, Zhigang Mao |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2013 | An energy-efficient and scalable eDRAM-based register file architecture for GPGPUabstractThe heavily-threaded data processing demands of streaming multiprocessors (SM) in a GPGPU require a large register file (RF). The fast increasing size of the RF makes the area cost and power consumption unaffordable for traditional SRAM designs in the future technologies. In this paper, we propose to use embedded-DRAM (eDRAM) as an alternative in future GPGPUs. Compared with SRAM, eDRAM provides higher density and lower leakage power. However, the limited data retention time in eDRAM poses new challenges. Periodic refresh operations are needed to maintain data integrity. This is exacerbated with the scaling of eDRAM density, process variations and temperature. Unlike conventional CPUs which make use of multi-ported RF, most of the RFs in modern GPGPU are heavily banked but not multi-ported to reduce the hardware cost. This provides a unique opportunity to hide the refresh overhead. We propose two different eDRAM implementations based on 3T1D and 1T1C memory cells. To mitigate the impact of periodic refresh, we propose two novel refresh solutions using bank bubble and bank walk-through. Plus, for the 1T1C RF, we design an interleaved bank organization together with an intelligent warp scheduling strategy to reduce the impact of the destructive reads. The analysis shows that our schemes present better energy efficiency, scalability and variation tolerance than traditional SRAM-based designs. Naifeng Jing, Shrikanth Ganapathy, Zhigang Mao, Minyi Guo, Ramon Canal, Xiaoyao Liang |
ISCA | 5 |
| 2012 | A pre-emphasis circuit design for high speed on-chip global interconnectabstractOn-chip global interconnects are speed and power bottleneck in state-of-the-art chips. Pre-emphasis technique is an efficient way to improve the performance of the global communication. This paper first performs delay analysis of a global wire to work with a pre-emphasis circuit in time domain. Based on the analysis, a new pre-emphasis circuit design is proposed. Simulation results show that the pre-emphasis circuit can increase the link bandwidth by more than 40% and 20% in capacitive and capacitive-resistive coupled 10mm global link respectively. The new pre-emphasis circuit design can be applied in high speed global communication. Jianfei Jiang 0001, Weiguang Sheng, Zhigang Mao, Weifeng He |
ISCAS | 3 |
| 2012 | Novel O-GEHL Based Hyperblock Predictor for EDGE ArchitecturesabstractControl flow speculation plays a pushing role to the performance of block-atomic EDGE architectures. Hyperblock predictors, which leverage the style of "Exit + Target" to predict the next hyperblock address, enable high efficient hyperblock-level control flow speculation for EDGE architectures. Recently, a series of binary prediction techniques have been studied and modified to adapt for the exit predictor in hyperblock predictors, including the O-GEHL prediction technique, which was first presented at 1st Championship Branch Prediction Competition. Our paper investigated different mispredict sources in O-GEHL based exit predictor in hyperblock predictors, and proposed two improved strategies: the O-GEHL based exit predictor without chooser and the O-GEHL based exit predictor employing binary O-GEHL prediction. Performance evaluation results showed that: the proposal without chooser outperformed previously published one by 0.7% with the hardware resource ranging from 16KB to 1MB; the proposal employing 8 binary O-GEHL predictor improved the performance by 3% with the largest hardware resource in this paper (1MB); the proposal employing 4 binary O-GEHL predictor for the first 4 exits averagely improved the performance by 2% with the hardware resource ranging from 16KB to 1MB. Pengfei Gou, Mingyan Yu, Zhigang Mao |
NAS | 4 |
| 2012 | SEU fault evaluation and characteristics for SRAM-based FPGA architectures and synthesis algorithmsabstractReliability has become an increasingly important concern for SRAM-based field programmable gate arrays (FPGAs). Targeting SEU (single event upset) in SRAM-based FPGAs, this article first develops an SEU evaluation framework that can quantify the failure sensitivity for each configuration bit during design time. This framework considers detailed fault behavior and logic masking on a post-layout FPGA application and performs logic simulation on various circuit elements for fault evaluation. Applying this framework on MCNC benchmark circuits, we first characterize SEUs with respect to different FPGA circuits and architectures, for example, bidirectional routing and unidirectional routing. We show that in both routing architectures, interconnects not only contribute to the lion's share of the SEU-induced functional failures, but also present higher failure rates per configuration bits than LUTs. Particularly, local interconnect multiplexers in logic blocks have the highest failure rate per configuration bit. Then, we evaluate three recently proposed SEU mitigation algorithms, IPD, IPF, and IPV, which are all logic resynthesis-based with little or no overhead on placement and routing. Different fault mitigating capabilities at the chip level are revealed, and it demonstrates that algorithms with explicit consideration for interconnect significantly mitigate the SEU at the chip level, for example, IPV achieves 61% failure rate reduction on average against IPF with about 15%. In addition, the combination of the three algorithms delivers over 70% failure rate reduction on average at the chip level. The experiments also reveal that in order to improve fault tolerance at the chip level, it is necessary for future fault mitigation algorithms to concern not only LUT or interconnect faults, but also their interactions. We envision that our framework can be used to cast more useful insights for more robust FPGA circuits, architectures, and better synthesis algorithms. Naifeng Jing, Ju-Yueh Lee, Zhe Feng 0002, Weifeng He, Zhigang Mao, Lei He 0001 |
ACM Trans. Design Autom. Electr. Syst. | 5 |
| 2011 | Fault modeling and characteristics of SRAM-based FPGAs (abstract only)abstractThe reliability of SRAM-based Field Programmable Gate Array (FPGA) is susceptible to Single Event Upset (SEU) fault. To investigate the fault impact, particular the fault in interconnects on FPGA functionality, this paper proposes a SEU fault analysis framework by evaluating the fault with a unified metric. This metric, termed as criticality, quantifies the sensitivity of FPGA functional failure to the SEU fault on logical and interconnect configuration bits. Considering the post layout information, our framework can characterize the SEU fault with respect to different FPGA architectures and CAD algorithms, such that the sensitivity of FPGA functional failure can be investigated in detail during design phase. The experiment result quantitatively shows that the configuration bits in interconnects dominate those in LUTs, several times both in bit number and criticality contribution. The ratio of their criticalities is even higher when LUT input size increases from 4 to 6. The higher criticality of interconnects than their LUT counterpart is due to their natural sensitivity to functional failure instead of their majority of bits. In addition, it is also shown that, among the three common types of switch boxes, the Subset switch box is less fault tolerant than Wilton and Universal. Naifeng Jing, Ju-Yueh Lee, Chun Zhang 0003, Jiarong Tong, Zhigang Mao, Lei He 0001 |
FPGA | 5 |
| 2011 | Quantitative SEU Fault Evaluation for SRAM-Based FPGA Architectures and Synthesis AlgorithmsabstractThis paper studies the SEU (Single Event Upset) fault for SRAM-based FPGAs. Considering detailed fault behavior on various circuit elements in a post-layout FPGA application, we develop a simulation-based SEU evaluation tool that quantifies fault contribution for each configuration bit. Using this tool and MCNC benchmark circuits, we study the fault characteristics of FPGA circuits and architectures. We show that interconnects not only contribute to the lion share of functional failures, but also have higher failure rate per configuration bit than LUTs. Particularly, multiplexers in local interconnects have the highest failure rate per bit. We find that tuning LUT and cluster sizes helps to reduce the rate (up to 38% in our experiments). In addition, we evaluate two recent fault mitigation algorithms IPD and IPF, which reduce LUT faults by an average of 74% and 15% respectively. But when interconnects are taken into account, the reduction via IPD which considers only LUT faults is merely 6% on chip level. Yet the reduction via IPF which implicitly considers interconnect faults is still around 15%. Therefore, synthesis algorithm should be evaluated with interconnect faults and future algorithms should be developed with consideration of interconnect faults explicitly. Naifeng Jing, Ju-Yueh Lee, Zhe Feng 0002, Weifeng He, Zhigang Mao, Shi-Jie Wen, Richard Wong, Lei He 0001 |
FPL | 5 |
| 2011 | Mitigating FPGA interconnect soft errors by in-place LUT inversionabstractModern SRAM-based FPGAs (Field Programmable Gate Arrays) use multiplexer-based unidirectional routing, and SRAM configuration cells in these multiplexers contribute to the majority of soft errors in FPGAs. In this paper, we formulate an In-Placed inVersion (IPV) on LUT (Look-Up Table) logic polarities to reduce the Soft Error Rate (SER) at chip level, and reveal a locality and NP-Hardness of the IPV problem. We then develop an exact algorithm based on the binary integer linear programming (ILP) and also a heuristic based on the simulated annealing (SA), both enabled by the locality. We report results for the 10 largest MCNC combinational benchmarks synthesized by ABC and then placed and routed by VPR. The results show that IPV obtains close to 4× chip level SER reduction on average and SA is highly effective by obtaining the same SER reduction as ILP does. A recent work IPD has the largest LUT level SER reduction of 2.7× in literature, but its chip level SER reduction is merely 7% due to the dominance of interconnects. In contrast, SA-based IPV obtains nearly 4× chip level SER reduction and runs 30× faster. Furthermore, combining IPV and IPD leads to a chip level SER reduction of 5.3×. This does not change placement and routing, and does not affect design closure. To the best of our knowledge, our work is the first in-depth study on SER reduction for modern multiplexer-based FPGA routing by in-placed logic re-synthesis. Naifeng Jing, Ju-Yueh Lee, Weifeng He, Zhigang Mao, Lei He 0001 |
ICCAD | 4 |
| 2011 | Effective multi-standard macroblock prediction VLSI design for reconfigurable multimedia systemsabstractReconfigurable computing arrays facilitate the flexibility with high performance for regular and computation-intensive algorithms in multimedia processing. However, the efficiency of the irregular and control-intensive algorithms becomes the performance bottleneck of reconfigurable multimedia systems. In this paper, we propose the design and VLSI implementation of a novel memory efficient macroblock prediction and boundary strength (Bs) calculation engine. The control-intensive algorithms, including intra mode prediction, motion vector prediction, and Bs calculation, are implemented with 4x4 block level pipeline to achieve real-time decoding for H.264/AVC high profile and Chinese AVS Jizhun profile. Compared with existing designs, our design achieves 60% registers reduction for neighboring block load and update. Implementation results indicate that the proposed architecture can support 1920×1088@30fps of H.264 and AVS decoding at 86 MHz. Yuliang Tao, Guanghui He 0002, Weifeng He, Qin Wang 0009, Jun Ma 0012, Zhigang Mao |
ISCAS | 6 |
| 2011 | A thermal-aware task mapping flow for coarse-grain dynamic reconfigurable processorabstractThis paper presents a task level mapping flow for coarse-grained dynamic reconfigurable array processor based on static thermal-aware mapping techniques. The flow is composed of front-end SUIF tool, temporal partitioning algorithm, thermal aware sub-graph mapping algorithms and back-end RAM compiler to compile HLL task into binary code for the processor automatically. Using compact thermal model, the temperature distribution of each task sub-graph on reconfigurable RC array is pre-estimated statically. The runtime sequence of all task sub-graphs is generated ultimately with the random searching algorithm to balance the reconfigurable array's temperature. Experimental results show that the average maximum temperature and temperature distribution range can be reduced about 6.3°C and 12°C, respectively. Weifeng He, Naifeng Jing, Zhigang Mao |
ISCAS | 4 |
| 2011 | A clock-less transceiver for global interconnectabstractHigh speed and low power transceivers start to be used for global interconnection in state-of-the-art System-on-Chips (SoCs). In traditional transceivers, the bandwidth is largely dependent on the clock rate. This paper presents a clock-less transceiver for global interconnect. The asynchronous transceiver makes the data rate only depend on the link delay and can be conveniently used with low swing scheme to create a high speed and low power communication system. The transceiver is demonstrated and simulated. The simulation results indicate that the transceiver can be used in high speed and low power global communications. Jianfei Jiang 0001, Weiguang Sheng, Weifeng He, Zhigang Mao |
VLSI-SoC | 5 |
| 2011 | A 230mV 8-bit sub-threshold microprocessor for wireless sensor networkabstractA customized design flow for ultra-low power cell library and an 8-bit ultra-low power microprocessor for wireless sensor network application are presented in this paper. According to the logic pre-synthesis results of the 8-bit microprocessor HDL code, frequently used standard cells are collected to develop a customized sub-threshold cell library through size and structure modifications. The ultimate transistor-level netlist of the processor is generated by the RTL code re-synthesis results through cell substitution under the sub-threshold cell library. Using the SPICE simulator, experimental results show that our 8-bit microprocessor can work at a supply voltage as low as 230mV at full temperature range of all technology corners, which has a power only 79nW and the frequency 10 KHz at 230mV and room temperature. As a result, the proposed microprocessor provides a feasible solution for emerging energy-constrained applications. Wei Jin 0004, Weifeng He, Zhigang Mao |
VLSI-SoC | 4 |
| 2011 | Robust design of sub-threshold flip-flop cells for wireless sensor networkabstractAs a major sequential logic element, D flip-flop is an indispensible cell in logic cell library. In this paper, we proposed two improved sub-threshold D flip-flop circuits (mTGMS and emC2MOS D flip-flop) after conducting robustness analysis of several typical flip-flop circuits. Using SMIC 0.18um CMOS technology, the simulation results show that the minimum work voltage of our proposed mTGMS and emC2MOS is 0.19V and 0.18V, the minimum average power is 13.2pW and 14.1pW, while the minimum Power Delay Product (PDP) is 13aJ and 4.35aJ respectively. Wei Jin 0004, Weifeng He, Zhigang Mao |
VLSI-SoC | 4 |
| 2011 | A general statistical estimation for application mapping in Network-on-ChipabstractDesign space exploration is crucial to an optimal application mapping in Network-on-Chip. However, the optimality evaluation of the explored solution has been neglected in previous studies. In this paper, we propose an efficient and credible statistical estimation approach to evaluate the optimality of explored solutions with respect to the mapped communication, which is directly related to power dissipation in the network. Our approach is motivated by a basic statistical property on the solution space, and we consider the diversities in different complex on-chip network designs to make it more applicable. The statistical estimation and the optimality evaluation are validated in experiments by real and synthetic applications. It demonstrates an estimating error around 6% on average, which tends to be even smaller when problem scales up. We envision that the fidelity of our statistical estimation approach will promote its applicability in the promising Network-on-Chip designs. Naifeng Jing, Weifeng He, Zhigang Mao |
VLSI-SoC | 3 |
| 2011 | On-chip structure and addressing scheme design for 2-D block data processing in a 64-core array systemabstractHigh throughput and high computation many core based on chip architecture is a development trend for high performance digital signal processing platform design. With frame and block data processing technology feature, the general linear memory architecture and on chip interconnection scheme may cause performance bottleneck for the application. In this paper work, we propose a 64-core array on chip architecture for high throughput and high performance 2-D block data processing, which with the features of hierachical architecture level, mirror symmetric interconnection structure and a Block level 2-D data addressing scheme. From the experiment show, the design can meet real time processing requirement of the 1080P data frame from H.264/AVC specification by use of block level matrix computation with average dynamic power report at 469.3mw. Jing Xie 0010, Huimin Xing, Zhigang Mao |
VLSI-SoC | 3 |
| 2010 | Statistical estimation and evaluation for communication mapping in Network-on-Chip
Naifeng Jing, Weifeng He, Yongxin Zhu 0001, Zhigang Mao |
Integr. | 4 |
| 2009 | Soft error optimization of standard cell circuits based on gate sizing and multi-objective genetic algorithmabstractA radiation harden technique based on gate sizing and multi-objective genetic algorithm (MOGA) is developed to optimize the soft error tolerance of standard cell circuits. Soft error rate (SER), chip area and longest path delay are selected as the optimization goals and fast fitness evaluation algorithms for the three goals are developed and embedded into the MOGA. All the three goals are optimized simultaneously by optimally sizing the gates in the circuit, which is a complex NP-Complete problem and resolved by MOGA through exploring the global design space of the circuit. Syntax analysis technique is also employed to make the proposed framework can optimize not only pure combinational logic circuit but also the combinational parts of sequential logic circuit. Optimizing experiments carried out on ISCAS'85 and ISCAS'89 standard benchmark circuits show that the proposed optimization algorithm can decrease the SER 74.25% with very limited delay overhead (0.28%). Furthermore, the algorithm can also reduce the area for most of the circuit under test by average 5.23%. The proposed technique is proved to be better than other works in delay and area overhead and suitable to direct the design of soft error tolerance integrated circuits in high reliability realms. Weiguang Sheng, Liyi Xiao, Zhigang Mao |
DAC | 3 |
| 2008 | A Hybrid Anti-Collision Algorithm for RFID with Enhanced Throughput and Reduced Memory ConsumptionabstractIn order to solve the transponder collision problem in a RFID system, this paper proposes a hybrid algorithm which combines strengths of existing high-performance algorithms while avoiding their major drawbacks. Experiments show this hybrid algorithm uses fewer time slots and less total communication time compared to adaptive slot-count algorithm based on ALOHA as well as enhanced anti-collision algorithm based on binary tree, both of which are superior and well accepted algorithms in literature. Meanwhile, high computational intensity of adaptive slot-count algorithm is significantly relieved, and memory consumption of enhanced anti-collision algorithm is reduced to a negligible level. Results from simulations under large number of transponders and long transponder IDs, which is the case in real RFID applications, suggest the advantage of this hybrid algorithm is prominent. Majun Zheng, Jing Xie 0010, Zhigang Mao, Yongxin Zhu 0001 |
EUC (1) | 3 |
| 2008 | Versatile and Efficient Techniques for Speeding-Up Circuit Level Simulated Fault-Injection CampaignsabstractFault injection in circuit level has proved to be cumbersome and time-consuming when employed to characterize the soft error sensitivity of digital circuits, hence new generation of CAD tool is required to automate the faults insertion and the validation of soft error mitigation mechanisms of the circuits. This paper outlines the characteristics of a new fault-injection platform HSECT-SPI (HIT Soft Error Characterization Toolkit-Spice Based) and its evaluation in some benchmark circuits implemented with distinct processes and soft error hardening techniques. It also details some techniques devised and implemented within the platform to automate and speed-up the circuit level fault-injection experiments. Experimental results are provided, showing that the platform is efficient, accurate and can direct the design of soft error immune circuits with at least three orders of magnitudes speed gain. Weiguang Sheng, Liyi Xiao, Zhigang Mao |
PRDC | 3 |
| 2007 | An Improved Frame-Level Pipelined Architecture for High Resolution Video Motion EstimationabstractFrame-level pipelined motion estimation structure achieves high throughput by exploiting the explicit parallelism among motion estimation blocks. In this paper, an improved frame-level pipelined architecture for FSBM motion estimation is proposed. The design efforts are focused on reducing the size of internal data buffers and the hardware overheads for high resolution video motion estimation. Compared with previous high performance architectures, the proposed architecture employs the smallest number of data buffers as well as removes the data broadcasting operations and keeps nearly 100% fully pipelined computation. As a result, this architecture offers a feasible solution for SHDTV video pictures. Weifeng He, Zhigang Mao |
ISCAS | 2 |
| 2004 | An adaptive motion estimation algorithm based on evolution strategiesabstractBased on evolution strategies (ESs), a novel adaptive motion estimation search algorithm (AESME) is presented. ESs consider evolutionary progress on the phenotype level. In contrast, genetic algorithms focus on heredity genetic mechanisms on the chromosome level. In ESs, the mutation operation accords with the normal distribution law. In the AESME algorithm, the (/spl mu/, /spl lambda/)-ES algorithm is adopted to block motion estimation, and the adaptive scheme is advanced to improve the convergence rate on the basis of the 1/5 success rule. Experimental results demonstrate that this algorithm has similar performance to that of the full-search (FS) algorithm, and owing to the inherent parallelism and low complexity of ESs, AESME is suitable for VLSI implementation. Hui Wang 0023, Zhigang Mao |
ICASSP (3) | 2 |
| 2004 | An adaptive motion estimation algorithm based on evolution strategies with correlated mutations
Hui Wang 0023, Zhigang Mao |
ICIP | 2 |
| 2000 | Implementation of Java Card Virtual Machine
Liu Songyan, Zhigang Mao, Yizheng Ye |
J. Comput. Sci. Technol. | 2 |
| 1999 | A New Algorithm for Retiming-Based Partial ScanabstractIn this paper, a new algorithm for retiming-based partial scan is presented. The proposed technique achieves a good-for-test configuration by moving registers to a specific set of edges that is selected for scan. As a part of our work, an algorithm of counting how many cycles that contain a specific edge is given. Moreover, the validity to place some registers on a definite set of edges by retiming is also discussed. Comparing with the existed retiming-based methods, our approach is characterized by less area overhead and comparable fault coverage. Experimental results of some ISCAS'89 benchmarks showed the effectiveness of our method. Zulan Huang, Yizheng Ye, Zhigang Mao |
Asian Test Symposium | 3 |
| 1998 | Test Pattern Generation for Column Compression MultiplierabstractWhen used as the building cell of a parallel multiplier, the (4,2) counter tree is better suited than a Wallace tree for a VLSI implementation because of its more regular structure. In this paper test pattern generation for the CC multipliers is presented following a brief introduction to the structure of the (4,2) counter and the column compression multiplier using a (4,2) counter as its building cell. In conclusion, less test patterns are enough to exhaustively test the CC multiplier. Pingying Zeng, Zhigang Mao, Yizheng Ye, Yuliang Deng |
Asian Test Symposium | 2 |