VLDB 2026 Research / reviewers in the wild / expert
Weifeng He
dblp:90/8030 · also Wei-Feng He
· DBLP profile ↗
56ranked-venue papers
1as first author
33since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 49 · 1 first-author · 26 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards disordered pick-and-place of aero-engine blades: A vision-guided method based on stable diffusion model and oriented object detection
Yizhen Yin, Weifeng He, Yuanhan Hou, Caizhi Li, Zhigao Wang, Qichun Hu, Senlin Zhu |
Eng. Appl. Artif. Intell. | 4 |
| 2026 | TPDA-DRAM: A Variation-Aware DRAM Improving System Performance via In-Situ Timing Margin Detection and Adaptive MitigationabstractDRAM latency remains a critical bottleneck in the performance of modern computing systems. However, the latency is excessively conservative due to the timing margins imposed by DRAM vendors to accommodate rare worst-case scenarios, such as weak cells and high temperatures. In this study, we introduce a temperature- and process-variation-aware timing detection and adaptation DRAM (TPDA-DRAM) architecture that dynamically mitigates timing margins at runtime. TPDA-DRAM leverages innovativein-situcross-coupled detectors to monitor voltage differences between bitline pairs inside DRAM arrays, ensuring precise detection of timing margins. Additionally, the proposed detector inherently accelerates the precharge operation of DRAM, thereby reducing the precharge latency by up to 62.5%. Building upon this architecture, we propose two variation-aware timing adaptation schemes: 1) a process-variation-aware adaptation (PVA) scheme that accelerates access to weak cells, mitigating process-induced timing margins, and 2) a temperature-variation-aware adaptation (TVA) scheme that leverages temperature information and the restoration truncation technique to reduce DRAM latency, mitigating temperature-induced timing margins. Evaluations on an eight-core computing system show that TPDA-DRAM improves average performance by 21.8% and energy efficiency by 18.2%. Yuxuan Qin, Chuxiong Lin, Guoming Rao, Weiguang Sheng, Weifeng He |
IEEE Trans. Computers | 6 |
| 2026 | DBP-CIM: Energy-Efficient 8T SRAM-Based Diagonal-Block Parallel Computing-in-Memory With Compact Data Layout for Arithmetic OperationsabstractIn this work, an energy-efficient bit-parallel static random-access memory (SRAM)-based computing-in-memory (SRAM-CIM) is proposed for general-purpose in-memory arithmetic operations to adapt diverse computing tasks. A compact diagonal-block parallel (DBP) mapping scheme and a novel arithmetic flow are proposed to address the hardware underutilization issue in conventional two-sided stationary bit-parallel CIM architectures. Specifically, the DBP mapping method is implemented to enhance the throughput by reorganizing the intermediate and final results into diagonal memory blocks, effectively reducing the vacant CIM cells caused by the dynamic bit width during computing. In addition, the proposed hardware-efficient arithmetic flows employ: 1) a pipelined ADD scheme to reduce the critical path latency in near-memory computing units; and 2) shift-based arithmetic operations that halve the hardware resources required for multiplication and division while reducing energy consumption. The post-layout simulations on 28-nm CMOS technology show that the proposed DBP-CIM achieves higher energy efficiency and throughput for general-purpose arithmetic operations, compared with state-of-the-art works. Furthermore, evaluations on the general-purpose benchmarks demonstrate that the DBP-CIM reduces energy consumption and computing cycles by up to 55.9% and 60.9%, compared to the conventional bit-parallel CIM. Dengfeng Wang, Chengjun Chang, Weifeng He, Guanghui He 0002, Yanan Sun 0003 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2026 | A Heterogeneous CIM Architecture With Splittable Nonvolatile Computing-in-SRAM Cell Pairs Enabling Efficient On-Chip Neural Network InferenceabstractHeterogeneous computing-in-memory (CIM) offers a promising solution for efficient neural network (NN) accelerations by leveraging the characteristics of different types of memories. However, the previous heterogeneous CIM either requires additional data transfer between different types of isolated memories or solely relies on in situ embedded single-level (SL) or three-level (TL) nonvolatile memories (NVMs), making it difficult to trade off between storage density and robustness benefits. In this article, a new heterogeneous CIM architecture (NVS-SPT) with enhanced storage density and inference robustness is proposed to enable full on-chip acceleration of practical-scale NNs while scalable to larger models. A splittable nonvolatile computing-in-static random access memory (SRAM) cell pairs (nvS2RAM-CIM) is proposed with hybrid in situ embedded SL and TL resistive random access memory (ReRAM) groups, allowing flexible configuration as split or linked state to enhance storage density and restore yield. A layerwise hybrid-coding search (LHCS) algorithm with bitwise and tritwise data-aware mapping (BTM) method is proposed to determine the optimal weight coding patterns with high array utilizations. In addition, a merged hybrid-coding block (MHCB) generation scheme is employed to enable high computing parallelism by merging the dense computing patterns. The proposed NVS-SPT demonstrates up to$4.2\times $higher storage density compared with previous heterogeneous CIM with pure SL-ReRAMs and achieves up to 44.7% enhanced NN accuracy, compared with previous unified ternary coding. Furthermore, the proposed NVS-SPT exhibits up to$1.72\times $and$1.44\times $enhanced energy efficiency with$3.10\times $and$1.42\times $higher computing density, compared with previous heterogeneous CIM based on pure SL- or TL-ReRAMs, respectively. Dengfeng Wang, Liukai Xu, Weifeng He, Guanghui He 0002, Xueqing Li 0002, Yanan Sun 0003 |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2025 | EPIC: Error PredIction and Correction for Power-Efficient Voltage Underscaling Multiply-Accumulate UnitabstractMatrix multiplication dominates the power consumption in compute-intensive applications such as deep neural networks (DNNs), spurring intensive investigations into power-efficient multiply-accumulate (MAC) units. Among the mainstream low-power design methodologies, voltage underscaling can achieve effective power savings yet induce timing errors that may lead to catastrophic accuracy loss. In this paper, we propose an error prediction and correction framework (denoted as EPIC) for arbitrary MAC unit under voltage underscaling, which predicts the timing errors and samples the correct output by using a delay-tunable clock. A prediction bits searching algorithm is proposed to enhance the prediction accuracy with low hardware cost, resulting in up to 100% accuracy. While preserving the accuracy, EPIC achieves up to 52% power savings over the corresponding MAC operating at nominal voltage. With transistor-level optimizations, EPIC incurs only 8% area and 1% power overheads, achieving 100% error correction under a voltage underscaling ratio of $\mathbf{0. 7 4}$. Compared to state-of-the-art error resilient circuit designs, EPIC consumes 60%-88% less area. Additionally, to achieve the accuracy performance of EPIC in error-resilient applications, we propose a simulation workflow involving precise timing features, enabling an accurate simulation of voltage underscaling MAC in large-scale applications. The experimental results show that, under voltage-underscaling, the MAC with EPIC consumes 11% less power than the one without EPIC, when a same accuracy as exact implementation is required in multi-layer perceptron (MLP). Tongjing Wu, Xiaolu Hu, Siting Liu 0001, Hui Wang 0023, Weifeng He, Zhigang Mao, Honglan Jiang |
DAC | 6 |
| 2025 | ASF-CRB: an Energy-Efficient Activity-Balanced Charge-Recycling Bus Architecture Based on Alternate Signal FlippingabstractCharge-recycling bus (CRB) achieves quadratic energy savings by stacking two data channels between VDD and VSS. However, it suffers from middle voltage (VMID) drift due to unbalanced data activities of the two channels, which increases signal propagation delay and causes timing violations. In this paper, we present ASF-CRB, a CRB architecture that addresses the VMID drift problem at minimum energy costs based on a novel data transmission scheme named Alternate Signal Flipping (ASF). By reconfiguring the signal transmission path every cycle through a multiplexer at each pipeline stage, ASF ensures all bits flip once in every two pipeline stages. This averages each channel's high and low activities to 50% and thus contributes to balanced data activities between stacked channels. Lightweight latch structures tailored to CRBs are presented to minimize the delay and energy overheads, and a complete ASF-CRB architecture is developed for applicability in VLSI systems. The proposed ASF-CRB is implemented at a 28nm CMOS process. Post-layout simulations show that ASF-CRB achieves up to 41.0% energy saving at similar VMID stability and ∼2.7× smaller VMID drifts at similar energy consumption compared with the state-of-the-art CRB structure. Xiangyu Ran, Chuxiong Lin, Weifeng He |
ACM Great Lakes Symposium on VLSI | 4 |
| 2025 | Incorporating machine learning in shot peening and laser peening: A review and beyond
Zhifen Zhang, James M. Griffin, Guangrui Wen, Weifeng He, Xuefeng Chen 0002 |
Adv. Eng. Informatics | 6 |
| 2025 | An Ultraspeed Middle Voltage and Timing Analyzer With Near-SPICE Accuracy for Charge-Recycling BusesabstractBy stacking two data channels between VDD and VSS, Charge Recycling Buses (CRBs) halve the voltage swing on interconnects, achieving significant power savings for energy-efficient on-chip data transmission. However, the middle voltage (VMID) between the two channels may fluctuate dynamically due to the diversity of input data, which significantly impacts data propagation delay and reliability. Unfortunately, existing SPICE-based simulators do not run fast enough to identify the worst-case VMID fluctuation and propagation delay, posing great challenges in CRB design. In this paper, we present a dedicated CRB Simulator for fast and accurate VMID and timing analysis. A highly condensed VMID fluctuation model, which integrates each cycles VMID changes into a single closed-form formula, is embedded in the simulator to predict VMID values at clock edges. In addition, a speed-monitoring algorithm is developed to track the continuously changing signal propagation speed under intra-cycle VMID fluctuations for accurate delay estimation. Both the VMID fluctuation model and the delay estimation algorithm involve only a small number of arithmetic operations, thus featuring remarkably low computational complexity. Compared with HSPICE, our CRB Simulator runs > 1.0×105 times faster on average across various CRB circuits, with a VMID prediction error of only 1.3mV and a delay estimation error as low as 0.6%. The significant speed improvement combined with high accuracy makes our CRB Simulator an efficient and reliable solution for VMID and timing analysis in CRB circuit design. Xiangyu Ran, Chuxiong Lin, Yuxuan Qin, Weifeng He |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2025 | HyCTor: A Hybrid CNN-Transformer Network Accelerator With Flexible Weight/Output Stationary Dataflow and Multicore ExtensionabstractHybrid convolutional neural network (CNN) and Transformer networks are emerging in computer vision, combining convolutional, linear, and attention layers to achieve high accuracies with moderate model sizes. Developing the accelerators for hybrid networks is pivotal to simultaneously optimize the static matrix multiplication (MM) in convolutional and linear layers, as well as dynamic MM in attention layers. However, the existing accelerators are primarily designed for either CNNs or Transformers, resulting in increased data movement to support dynamic MM and potential under-utilization of hardware for static MM. To enhance computational performance and energy efficiency for hybrid networks, we propose HyCTor, an accelerator featuring flexible output-stationary (OS) and weight-stationary (WS) dataflows, along with a multicore extension for higher throughput. The parallel array of HyCTor supports interlayer slicing and intralayer splicing to improve the utilization for static MM, and enables seamless switching between OS and WS dataflow to minimize the data movement in dynamic MM. By leveraging structured sparsity in OS dataflow and unstructured sparsity in WS dataflow, the computational efficiency is further boosted for each layer through flexible dataflow selection based on the sparsity ratio. Besides, a novel QuadLoop-mesh topology is proposed to address the complex data dependencies in hybrid networks and minimize data transmission distances in the multicore HyCTor. Experimental results on ResNet-18, ViT-B, and TransIAR-AF show that the proposed single-core HyCTor achieves$1.83\times $,$1.65\times $, and$2.41\times $speedup than state-of-the-art (SOTA) accelerators with 100% utilization rate in most layers, and$3.82\times $–$38.5\times $speedup than RTX4090 GPU. The energy efficiency of HyCTor is improved by$1.81\times $–$8.77\times $compared with SOTA accelerators. Moreover, the 4-core HyCTor achieves speedups of$3.32\times $,$2.58\times $, and$2.91\times $, while the 16-core HyCTor achieves speedups of$7.05\times $,$4.05\times $, and$9.64\times $compared to 1-core HyCTor on three networks. Shuai Yuan 0016, Weifeng He, Zhenhua Zhu 0002, Fangxin Liu, Zhuoran Song, Guohao Dai 0001, Guanghui He 0002, Yanan Sun 0003 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2025 | An Efficient Multi-View Cross-Attention Accelerator for Vision-Centric 3D Perception in Autonomous DrivingabstractVision-centric 3D perception has become a key mechanism in autonomous driving. It achieves exceptional perceptual performance mainly by introducing a novel attention,multi-view cross-attention(MVCA), for learnable feature extraction and fusion from surround-view cameras. Despite its superiority, MVCA encounters severe inefficiencies in sample, processing elements (PE), and pipelined processing, owing to the redundant and non-uniform sampling-aggregation and rigorous inter-operator dependencies. To address these issues, this article proposes a dedicated MVCA accelerator, MVAtor, with algorithm-architecture co-optimization for vision-centric 3D perception based on multi-view inputs flexibly. For sample inefficiency, a 3-tier hybrid static-dynamic sample and a sensitivity-aware feature pruning approach are proposed to eliminate the 86.03% sample overhead and 24.48% memory requirement, only incuring <1% accuracy loss with no need of fine-tuning. For PE inefficiency, a spatial pruner and sequential sampler collaboration strategy is proposed to improve the sampler utilization without compromising pruner’s throughput, which outperforms the previous design by 53.7~96.1% energy-delay product reduction. For pipeline inefficiency, a fine-grained-tiling assisted highly-pipelined architecture is constructed in MVAtor by exploiting the decoupling opportunities on inter-view sparsity, thereby saving 61.03% external memory access while boosting the overall throughputs by 1.83×. Extensively evaluated on representative benchmarks, MVAtor attains 1.38~7.67× and 1.67~11.15× improvement on energy and area efficiency respectively, compared to the state-of-the-art related accelerators. Dongxu Lyu, Gang Wang 0063, Wenjie Li 0003, Weifeng He, Guanghui He 0002 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 8 |
| 2025 | Robust Monolithic 3D Carbon-Based Computing-in-SRAM With Variation-Aware Bit-Wise Data-Mapping for High-Performance and Integration DensityabstractBit-serial computing-in memory with SRAM cells (SRAM-CIM) enables a full set of integer and floating-point arithmetic operations and various data-intensive computations. Carbon nanotube field-effect transistors (CN-MOSFETs) with high scalability, energy-efficiency, and low process thermal budget are attractive to realize high-dense monolithic three-dimensional (M3D) SRAM-CIM. However, CN-MOSFETs possess unique process variations with asymmetric spatial correlations which can significantly influence the performance and reliability of carbon-based SRAM-CIM. In this paper, new M3D-4N4P SRAM-CIM cells with CN-MOSFETs are proposed with optimized profiles for achieving ultra-high integration density while preserving robustness of data-access and computation. Furthermore, the variation-aware bit-wise data-mapping method is proposed for enhancing the performance of carbon-based SRAM-CIM by leveraging the spatial correlations of CN-MOSFETs. By minimizing the area skew of vertically-stacked layers, the areas of proposed M3D-4N4P SRAM-CIM cells are reduced by up to 50.32% compared to the previous 6N2P SRAM-CIM cells assuming carbon nanotube transistor technology. The proposed M3D-4N4P SRAM-CIM array also achieves by up to$2.17\times $higher throughput on arithmetic operations and 18.34% lower computing latency with 25.36% reduced energy consumptions on MAC-based benchmarks, respectively, compared to the previous 2D-6N2P SRAM-CIM array. Dengfeng Wang, Weifeng He, Qin Wang 0009, Hailong Jiao, Yanan Sun 0003 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2024 | Reducing DRAM Latency via In-situ Temperature- and Process-Variation-Aware Timing Detection and AdaptionabstractLong DRAM access latency has a significant impact on modern system performance. However, the improvement of DRAM access latency is limited, as the DRAM vendors reserve considerable timing margins against seldom worst-case conditions. To mitigate such pessimistic timing margins, we propose a temperature- and process-variation-aware timing detection and adaption DRAM (TPDA-DRAM) architecture. It equips in-situ cross-coupled detectors to monitor the voltage difference between bitline pairs, enabling estimation of timing margins caused by process and temperature variations. Moreover, TPDA-DRAM incorporates two collaborative timing adaption schemes: 1) a process-variation-aware timing adaption scheme (PVA) that selectively accelerates the access to weak cells, and 2) a temperature-variation-aware timing adaption scheme (TVA) that precisely adjusts timing parameters by adopting temperature information. Compared to prior art, the proposed detector reduces detection deviation by 54.8% and area overhead by 88.1%. The system-level evaluation in an eight-core system shows that TPDA-DRAM improves the average performance and energy efficiency by 20.5% and 15.0%, respectively. Yuxuan Qin, Chuxiong Lin, Zhang Luo, Weifeng He |
DAC | 6 |
| 2024 | Deciphering laser shock peening quality monitoring: Wavelet-driven network with interpretability
Zhifen Zhang, Zhengyao Du, Xizhang Chen, Guangrui Wen, Weifeng He, Xuefeng Chen 0002 |
Adv. Eng. Informatics | 8 |
| 2024 | Point cloud enhancement optimization and high-fidelity texture reconstruction methods for air material via fusion of 3D scanning and neural rendering
Qichun Hu, Yizhen Yin, Weifeng He, Senlin Zhu |
Expert Syst. Appl. | 6 |
| 2023 | TL-nvSRAM-CIM: Ultra-High-Density Three-Level ReRAM-Assisted Computing-in-nvSRAM with DC-Power Free Restore and Ternary MAC OperationsabstractAccommodating all the weights on-chip for large-scale NNs remains a great challenge for SRAM based computing-in-memory (SRAM-CIM) with limited on-chip capacity. Previous non-volatile SRAM-CIM (nvSRAM-CIM) addresses this issue by integrating high-density single-level ReRAMs on the top of high-efficiency SRAM-CIM for weight storage to eliminate the off-chip memory access. However, previous SL-nvSRAM-CIM suffers from poor scalability for an increased number of SL-ReRAMs and limited computing efficiency. To overcome these challenges, this work proposes an ultra-high-density three-level ReRAMs-assisted computing-in-nonvolatile-SRAM (TL-nvSRAM-CIM) scheme for large NN models. The clustered n-selector-n-ReRAM (cluster-nSnRs) is employed for reliable weight-restore with eliminated DC power. Furthermore, a ternary SRAM-CIM mechanism with differential computing scheme is proposed for energy-efficient ternary MAC operations while preserving high NN accuracy. The proposed TL-nvSRAM-CIM achieves 7.8x higher storage density, compared with the state-of-art works. Moreover, TL-nvSRAM-CIM shows up to 2.9x and 2.0x enhanced energy efficiency, respectively, compared to the baseline designs of SRAM-CIM and ReRAM-CIM, respectively. Dengfeng Wang, Liukai Xu, Songyuan Liu, Zhi Li 0058, Weifeng He, Xueqing Li 0002, Yanan Sun 0003 |
ICCAD | 6 |
| 2023 | A Metastability Inference and Avoidance Technique for Near-Threshold-Voltage Network-on-ChipabstractWith the application of low-power design technologies such as dynamic voltage and frequency scaling (DVFS) and globally asynchronous locally synchronization (GALS), a multi-voltage-/frequency-domain network-on-chip (NoC) suffers more and more serious metastability issue in inter-core data communication. To mitigate the metastability during the clock-domain crossing, a technique titled metastability inference and avoidance (MIAA) is presented. MIAA infers the potential metastability risk of a synchronizer's sampling clock through phase detection of a phase-related clock. MIAA avoids the occurrence of metastability by adaptively modulating the clock phase of the sampling clock once it infers the potential metastability risk. We designed a MIAA-based 40nm GALS$2\times 2$NoC that contains four independent voltage/frequency domains. The post-layout simulation results show that MIAA can well predict the metastability risks and reduce the probability of metastability to zero across a wide range of frequency ratios. The metastability mitigation allows us to use a single flip-flop instead of a multi-stage synchronizer for synchronization in the NoC, thereby improving the latency and throughput of the NoC by 40.2% and 79%, respectively. Chuxiong Lin, Weifeng He |
ISCAS | 5 |
| 2023 | An Area-Efficient Single-Phase-Clocked and Contention-Free Flip-Flop for Ultra-Low-Voltage OperationsabstractThis paper proposes an area-efficient single-phase-clocked and contention-free flip-flop (FF) targeting ultra-low-voltage (ULV) operations, named TSPC20. To ensure reliable operations in ULV regime, we eliminate all the contention paths and cut off the longest hold time path consisting of three stacking transistors in the conventional single-phase-clocked FF (TSPC18). Moreover, to further reduce the area and power consumption, we remove the redundant transistors through transistor merging and logical expression reorganization. TSPC20, with only 20 transistors, is the most area-efficient FF compared to prior FFs that can operate in ULV regime. Post-layout simulations with 28nm process shows that TSPC20 achieves 48% (54%) power reduction at 0.3V/6Mhz (0.9V/1.4GHz) considering 10% data activity ratio, compared to the conventional transmission-gate flip-flop (TGFF). The 1K Monte Carlo simulations verify that TSPC20 is functional down to 0.3V. Weifeng He |
ISCAS | 6 |
| 2023 | Surface stress monitoring of laser shock peening using AE time-scale texture image and multi-scale blueprint separable convolutional networks with attention mechanism
Zhifen Zhang, Zhengyao Du, Guangrui Wen, Weifeng He |
Expert Syst. Appl. | 7 |
| 2023 | CDAR-DRAM: Enabling Runtime DRAM Performance and Energy Optimization via In-Situ Charge Detection and Adaptive Data RestorationabstractWith the increasing of dynamic random access memory’s (DRAM) capacity, the refresh operation rapidly becomes a major concern to the performance of the current computational system. Moreover, conservative timing parameters adopted for access operations make an increasing amount of negative impact on system performance and energy efficiency. In this article, we propose an in-situ charge detection and adaptive data restoration DRAM (CDAR-DRAM) architecture, which can dynamically adjust the refresh rate and relax the constraints on access timing by removing pessimistic timing margins for PVT variations. CDAR-DRAM employs a low-cost skewed-inverter-based detector to monitor the bitline voltage in runtime and estimate real-time timing parameters of cells. Based on the detector, an adaptive refresh and restore scheme (CDAR-ref) is presented, which progressively reduces the refresh rate and partially restores cells’ voltage just enough for cells with sufficient charge, thereby optimizing both refresh and restoration operations. Moreover, a supplementary adaptive access scheme (CDAR-acc) is presented, which detects the runtime charge level of recently accessed rows and reduces access latency aggressively, benefitting workloads in a single-core system and memory nonintensive workloads in a multicore system. CDAR’s flexibility allows the two schemes to be combined. The evaluation shows that in an eight-core system, the combined scheme improves performance and energy efficiency by 15.2% and 22.6%, respectively. Yuxuan Qin, Chuxiong Lin, Weifeng He, Yanan Sun 0003, Zhigang Mao, Mingoo Seok |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2023 | A DFT-Compatible In-Situ Timing Error Detection and Correction Structure Featuring Low Area and Test OverheadabstractIn-situ timing error detection and correction (EDAC) structure is widely adopted in timing-error resilient circuits to reduce the conservative timing guardband induced by process, voltage, and temperature (PVT) variations. However, it introduces the latch-based datapath as well as extra detection and propagation logic, therefore challenges the design-for-testability (DFT) implementation. In this article, we propose a novel DFT-compatible EDAC structure with significant signal control simplification and test-pattern complexity reduction, featuring low area and test overhead. This structure leverages a new scannable EDAC cell (SEDC) which can be configured for timing EDAC in normal mode, or for shift operations as a flip-flop in scan mode. Specifically, the proposed detection logic can be controlled succinctly in scan shift operations and then observed via the global error propagation logic with simple control signal configurations during the test. Therefore, the sophisticated test pattern generation and critical path sensitization are removed. Based on the structure, a shift-based test method is presented to cover the EDAC structure with a low test pattern complexity and test time overheads. As compared with previous works, the proposed SEDC saves 30.5% area, 16.6% power, and 20.3% delay. Besides, our test method cooperating with the proposed EDAC structure reduces$149\times $and$23\times $static and at-speed test patterns, respectively, together with$232\times $static test cycles, and up to$25\times $at-speed test cycles on average, which proves the effectiveness for DFT. Hao Zhang 0151, Weifeng He, Yanan Sun 0003, Mingoo Seok |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2023 | CREAM: Computing in ReRAM-Assisted Energy- and Area-Efficient SRAM for Reliable Neural Network AccelerationabstractSRAM-based computing-in-memory (CIM) has been widely explored to accelerate neural networks (NNs). However, it is challenging to store all weights of many modern NNs due to limited on-chip SRAM capacity. This bottleneck induces a large amount of off-chip DRAM accesses and impedes the improvement of performance and energy efficiency. This paper proposes a new approach of computing in resistive random-access memory (ReRAM)-assisted energy- and area-efficient SRAM (CREAM) for accelerating large-scale NNs while eliminating the DRAM access. The NN weights are all stored in high-density on-chip ReRAMs and restored to the proposed non-volatile SRAM (nvSRAM) CIM cells with array-level parallelism. Furthermore, to deal with the influence of ReRAM and CMOS variations, a novel layer-wise and bit-wise weight-configuration search algorithm is proposed by leveraging different sensitivity of each layer in NN models. A data-aware weight-mapping method is also presented to efficiently map NN models to ReRAMs in CREAM for high computation parallelism. The experiment results show$10.3\times $weight storage density over the standard 6T SRAM array. Evaluations of ResNet-18 and VGG-9 on CIFAR-10/CIFAR-100 datasets show up to$3.47\times $and$1.70\times $energy efficiency over two baseline designs of SRAM-CIM and ReRAM-CIM, respectively, in addition to 15.6% higher accuracy than ReRAM-CIM under device variations. Yanan Sun 0003, Dengfeng Wang, Liukai Xu, Zhi Li 0058, Songyuan Liu, Weifeng He, Yongpan Liu, Huazhong Yang, Xueqing Li 0002 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2023 | BC-MVLiM: A Binary-Compatible Multi-Valued Logic-in-Memory Based on Memristive CrossbarsabstractLogic-in-memory with memristive crossbars is an attractive approach for realizing beyond von Neumann architectures. Multi-valued logic (MVL) containing more than two logic levels can enhance the computing speed with reduced number of logic operations. In this paper, a binary-compatible multi-valued logic-in-memory (BC-MVLiM) scheme is proposed with memristive dual-crossbars where both inputs and outputs are represented by the multi-level cells of memristors. Both of the binary and multiple-valued logic operations can be implemented in the proposed BC-MVLiM scheme depending on the radix of inputs. The proposed BC-MVLiM circuitry supports multiple row-wise and column-wise logic gates with multiple fan-ins and fan-outs for binary and ternary systems by leveraging both inter- and intra-crossbar operations. Experimental results show that the proposed BC-MVLiM-based multi-digit adder enhances the computation speed by up to 76.10% and 83.82%, for binary and ternary systems, respectively, as compared with the previously published memristive logic designs. By preventing the errors from propagating across multiple stages, the error rate of proposed BC-MVLiM is also reduced by up to 98.20% compared to the previous memristive logic designs in the presence of device variations. Yanan Sun 0003, Zhi Li 0058, Weifeng He, Qin Wang 0009, Zhigang Mao |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2023 | A Novel Approach for Surface Integrity Monitoring in High-Energy Nanosecond-Pulse Laser Shock Peening: Acoustic Emission and Hybrid-Attention CNNabstractThe high energy, transient, and nanosecond pulse of laser shock peening (LSP) renders real-time monitoring of material surface integrity challenging. Along these lines, this article proposes a novel method for real-time evaluation of the surface quality in the LSP process based on the multiple acoustic emission (AE) technique and hybrid attention-based convolutional neural networks (HACNN). First, a new quality index, surface hardness integrity (SHI) is constructed to fully characterize both the hardening rate and the impact depth of the material surface, while a good linear relationship is demonstrated between SHI and AE. Then, a feature extraction method, called wavelet packet energy cepstrum (WPEC), is proposed that requires minor signal processing expertise while being able to adaptively match the modal distribution of the broadband AE signals. Next, WPEC is fed into HACNN, where HA consists of channel attention and multiscale spatial attention (MSA). MSA combines the information flow at the global time, local time, and local frequency. Moreover, numerous comparisons were conducted with several traditional cepstrum and advanced attention mechanisms. The effectiveness of the proposed method is carefully verified by the LSP experiments for 7075 Al alloy achieving the highest average accuracy of 99.33% for the identification of four types of SHI. In addition, the sensitivity of HA in high-frequency event onset and offset is demonstrated by visualizing the gradient weights. Zhifen Zhang, Zhengyao Du, Guangrui Wen, Weifeng He |
IEEE Trans. Ind. Informatics | 6 |
| 2022 | CREAM: computing in ReRAM-assisted energy and area-efficient SRAM for neural network accelerationabstractComputing-in-memory has been widely explored to accelerate DNN. However, most existing CIM cannot store all NN weights due to limited SRAM capacity for edge AI devices, inducing a large amount off-chip DRAM access. In this paper, a new computing in ReRAM-assisted energy and area-efficient SRAM (CREAM) is proposed for implementing large-scale NNs while eliminating off-chip DRAM access. The weights of DNN are all stored in the high-dense on-chip ReRAM devices and restored to the proposed nvSRAM-CIM cells with array-level parallelism. A data-aware weight-mapping method is also proposed to enhance the CIM performance while fully exploiting the hardware utilization. Experiment results show that the proposed CREAM scheme enhances the storage density by up to 7.94x compared to the traditional SRAM arrays. The energy-efficiency of proposed CREAM is also enhanced by 2.14x and 1.99x, compared to the traditional SRAM-CIM with off-chip DRAM access and ReRAM-CIM circuits, respectively. Liukai Xu, Songyuan Liu, Zhi Li 0058, Dengfeng Wang, Yanan Sun 0003, Xueqing Li 0002, Weifeng He |
DAC | 8 |
| 2022 | MSLM-RF: A Spatial Feature Enhanced Random Forest for On-Board Hyperspectral Image ClassificationabstractHyperspectral imaging (HSI) greatly improves the capacity to identify and monitor ground objects due to the high spectral resolution. As the real-time remote sensing monitoring and warning tasks are getting more attention, new algorithms for low-power on-board classification are required to reduce the transmission time of satellite downlink. In this paper, we propose the Multi-Scale Local Maximum Random Forest (MSLM-RF) to significantly reduce the energy consumption while retaining high classification accuracy. The proposed MSLM-RF uses multi-scale maximum filters for spatial feature extraction and Random Forest for classification after spectral and spatial features fusion. The spatial features are efficiently extracted with low computational complexity by regarding the maximum light intensity values in different ranges of pixels as anchor points. MSLM-RF only consists of integer comparisons and a few additions, thereby eliminating the energy-hungry operations such as multiplication and exponentiation. According to experimental results on the HSI benchmark datasets, MSLM-RF delivers a better trade-off in accuracy and computational complexity than the state-of-the-art classification algorithms. Besides, MSLM-RF gets higher average classification accuracy and lower energy consumption than the previous on-board algorithms. The obtained results show the suitability of the proposed algorithm to accomplish practical real-time classification tasks on-board with low energy consumption. Shuai Yuan 0016, Yanan Sun 0003, Weifeng He, Qianrong Gu, Zhigang Mao, Shikui Tu |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2021 | CDAR-DRAM: An In-situ Charge Detection and Adaptive Data Restoration DRAM Architecture for Performance and Energy Efficiency ImprovementabstractAs the capacity of DRAM continues to grow, the refresh operation rapidly becomes the performance and power-efficiency bottleneck. Also, restore time, the time given for recharging cells post access, makes an increasingly large amount of negative impact on performance. To tackle these problems, in this paper, we propose an in-situ charge detection and adaptive data restoration DRAM (CDAR-DRAM) architecture, which can dynamically adjust the refresh rate and also relax the constraints on restore time. The proposed CDAR-DRAM employs a low-cost skewed-inverter-based detector, which can reduce the excessive timing margins that prior work added to guarantee the functionality of leaky DRAM cells under the worst-case temperature condition. Moreover, an adaptive DRAM refresh and restore scheme is proposed, which can switch automatically between two modes: (i) a refresh mode that supports adaptive refresh rate, and (ii) a restore mode that relaxes the constraints on restore time dynamically for cells having sufficient charge. With the transistor-and architecture-level simulations, we evaluate the CDAR-DRAM in an 8-core system across different workloads. Compared with the prior art, the proposed architecture achieves a 9.4% improvement in system performance and a 14.3% reduction in energy consumption, without requiring the time-consuming profiling process which many prior works employed. Chuxiong Lin, Weifeng He, Yanan Sun 0003, Zhigang Mao, Mingoo Seok |
DAC | 2 |
| 2021 | An Area-Efficient Scannable In Situ Timing Error Detection Technique Featuring Low Test Overhead for Resilient CircuitsabstractTiming error detection is a key technique for resilient circuits to explore the timing margins, yet it hinders the scan shift operations and increases the excessive test overhead. In this paper, we propose an area-efficient scannable in situ timing error detection technique consisting of a lightweight scannable error-detection cell and propagation logics, featuring low design-for-test effort and test overhead. The proposed error-detection cell fully reuses its main and shadow latches to construct the latch-based error-detection structure in normal mode, or the flip-flop-based datapath in scan mode. Therefore, it not only offers the time-borrowing ability to lower the correction overheads, but also supports the scan shift operations and detection logic tests. Besides, the dependency of error signal generation on the critical path sensitization is eliminated by configuring input and clock signals of error propagation logics, and thereby the detection and propagation logic can be tested easily. Benefiting from the technique, a set of test methods is presented with lower test pattern scales and test cycle overheads. As compared with previous works, the proposed cell saves at least 30.5% area overhead. Besides, experimental results across several benchmark circuits show that 116x of test patterns, 232x of static test cycles, and 26x of at-speed test cycles are saved on average, proving the effectiveness of the proposed technique for the design-for-test requirement. Hao Zhang 0151, Weifeng He, Yanan Sun 0003, Mingoo Seok |
ICCAD | 2 |
| 2021 | Investigation of Dynamic Leakage-Suppression Logic Techniques Crossing Different Technology Nodes from 180 nm Bulk CMOS to 7 nm FinFET Plus ProcessabstractLeakage power reduction techniques are crucial for energy-efficient circuits. This paper investigates the leakage suppression capability, performance, and reliability of dynamic leakage suppression logic (DLSL) and feedforward leakage self-suppression logic (FLSL) techniques, crossing different technology nodes from TSMC 180 nm bulk CMOS to 7 nm FinFET Plus process. Compared with CMOS benchmarks, experimental results show that DLSL-based benchmarks demonstrate a leakage power reduction for four orders of magnitude in 180 nm and 130 nm technologies, while only two orders of magnitude in other technologies. Moreover, FLSL offers a 4-28× performance improvement over DLSL at a cost of 2× leakage power. Zihan Lian, Hao Zhang 0151, Weifeng He, Yanan Sun 0003, Mingoo Seok |
ISCAS | 4 |
| 2021 | Design of Ternary Logic-in-Memory Based on Memristive Dual-CrossbarsabstractImplementing logic within memristive crossbar is an attractive approach to overcome the memory wall in conventional von Neumann architectures. Ternary logic with three logic levels can reduce the number of logic operations and enhance the computing speed compared to the binary logic. In this paper, a ternary logic-in-memory scheme is proposed based on the memristive dual-crossbar structure where the inputs and outputs are represented by the multi-level cells of memristors. Two inter-crossbar ternary logic gates and one intra-crossbar binary logic gate for both row and column-wise operations are supported in the proposed scheme to effectively reduce the operation latency. Experimental results show that the operation steps of the proposed multi-trit ternary adder are reduced by up to 83.82%, as compared with previously published binary memristive logic designs. The computation energy consumed by the proposed ternary adder is also reduced by up to 35.87% as compared to previously published binary IMPLY logic design. Yanan Sun 0003, Weifeng He, Qin Wang 0009 |
ISCAS | 3 |
| 2021 | An Energy-Efficient Logic Cell Library Design Methodology with Fine Granularity of Driving Strength for Near- and Sub-Threshold Digital CircuitsabstractCommercial multi-threshold standard logic cell libraries are designed for nominal super-threshold circuits. If blindly used at near- and sub-threshold voltages, such libraries exhibit excessively coarse granularity in driving strength, leading to sub-optimal logic synthesis and placement-and- routing results. To tackle this problem, a holistic methodology for designing a near- and sub-threshold standard cell library that has fine driving strength granularity is presented in this paper. Meanwhile, the proposed methodology leverages inverse narrow width effect, reverse short channel effect and forward body biasing to modulate the driving strength at low area overheads. Based on the proposed methodology, we develop a 65nm multi-threshold-voltage, multi-channel-length library and benchmark it against the commercial library across several common circuits. The results show a 26.6% reduction in power-delay product, a 28.1% reduction in energy-delay product, and a 27.0% reduction in leakage power at 5.8% area overhead on average, confirming the efficiency of the methodology in near- and sub-threshold digital circuits design. Hao Zhang 0151, Weifeng He, Yanan Sun 0003, Mingoo Seok |
ISCAS | 2 |
| 2021 | An Ultra-Low Leakage Bitcell Structure with the Feedforward Self-Suppression Scheme for Near-Threshold SRAMabstractLeakage power consumption has become a critical issue for low power Static Random-Access Memory (SRAM) design in the near-threshold regime. In this paper, an ultra-low leakage fourteen-transistor SRAM bitcell structure with the feedforward self-suppression scheme is presented. To reduce the leakage power significantly as well as maintain the data stability in hold state, a cross-coupled dynamic leakage- suppression inverter-based structure is adopted. Furthermore, the bypass scheme is employed to enable the speed modulation for bitcell read and write operations. As compared with state- of-the-art designs, a 65 nm 8kb SRAM array with the proposed bitcell structure achieves 75× leakage power, 45% write power as well as 65% read power reduction at 0.4V. Further comparisons with different processes verify up to 38k times leakage power reduction in 130nm planar process and 139× in 7nm plus FinFET process. Hao Zhang 0151, Weifeng He, Yanan Sun 0003, Mingoo Seok |
ISCAS | 3 |
| 2021 | High Energy-Efficient LDPC Decoder with AVFS System for NAND Flash MemoryabstractIn conventional Low-Density Parity-Check (LDPC) decoders, the real-time processing performance should meet its maximum decoding iterations for all packets and the work frequency or supply voltage is always fixed at a high level, which decreases its energy efficiency. In this paper, an energy- efficient LDPC decoding architecture with an adaptive voltage-frequency scaling (AVFS) scheme is presented. According to the usage of input packet FIFO related to variable decoding iterations, the architecture can dynamically adjust decoder's work frequency and supply voltage to reduce the processing energy while meeting its real-time processing requirement. Finally, the decoder is implemented with 28 nm CMOS process. Experimental results show that our decoder has a throughput of 1590 Mb/s when the raw bit error rate (RBER) of Flash memory is up to 10-2. The power consumption of the decoder can be reduced by 25%-62% and energy efficiency can be increased to 1.3-2.5 times under different AWGN noise. Jingtong Mo, Zihan Lian, Weifeng He |
ISCAS | 4 |
| 2021 | A 3.85-Gb/s 8 × 8 Soft-Output MIMO Detector With Lattice-Reduction-Aided Channel PreprocessingabstractThis article presents an 8 × 8 lattice-reduction-aided (LRA) soft-output multiple-input multiple-output (MIMO) detector for Chinese enhanced ultrahigh throughput (EUHT) wireless local area network (LAN) standard. The preprocessing algorithm combining simplified-sorting Cholesky decomposition and low-complexity decoupled lattice reduction (LDLR) is proposed to reduce computational complexity and latency with parallelism improvement. In addition, K-best detection adopts a sorting-reduced strategy utilizing approximate ordered sequence. Compared with other published LRA K-best detection algorithms, simulation results show that our proposed algorithm has performance improvement. In addition, in order to save hardware resources, a folded K-best architecture and an optimized intermediate storage strategy are introduced. Furthermore, a fully pipelined VLSI architecture is designed in Semiconductor Manufacturing International Corporation (SMIC) 40-nm 1P9M technology to support the 8 × 8.64 -QAM MIMO-OFDM system. The detector can achieve 3.85-Gb/s data throughput at 641-MHz clock frequency with 0.71-μs latency. The proposed detector is competitive in terms of latency, throughput, and area efficiency to state-of-the-art works and can meet the data-rate requirement of the EUHT standard. Zhuojun Liang, Dongxu Lv, Chao Cui, Haibao Chen, Weifeng He, Weiguang Sheng, Naifeng Jing, Zhigang Mao, Guanghui He 0002 |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2018 | Electronic-Photonic Integrated Circuit Design and Crosstalk Modeling for a High Density Multi-Lane MZM ArrayabstractTo meet the rapidly increasing bandwidth density demand in high-performance computers and data centers, the number of parallel lanes in a short-reach optical interconnect link is expected to increase by several folds. This paper presents a multi-lane traveling-wave Mach-Zehnder modulator (MZM) array design using silicon photonics technology. We treat the MZM as an electronic-photonic integrated circuit and adopt a newly developed multi-physics cross-layer design methodology. The design spans electrical/optical device modeling, electromagnetic simulation, and circuit analysis. The crosstalk issue within a closely-packed MZM array is investigated. The simulated 3-dB EO bandwidth of a single MZM is 17.5 GHz. The complete link considering crosstalk in the MZM array achieves BER-12at 25 Gbps. A prototype chip with 4 MZMs is designed using IMEC 50G silicon photonics technology, saving significant chip area compared to conventional design. Pengfei Ji, Yanan Sun 0003, Weifeng He |
ISCAS | 5 |
| 2017 | A static-placement, dynamic-issue framework for CGRA loop acceleratorabstractThis paper presents a static-placement, dynamic-issue (SPDI) framework for the coarse-grained reconfigurable architecture (CGRA) in order to tackle the inefficiencies of the static-issue, static-placement (SISP) CGRA. This framework includes the compiler that statically places the operations and hardware design, a SPDI CGRA, that automatically schedule the operations. We stress on introducing the SPDI CGRA in this paper. This newly designed hardware model adds the token buffer, which is capable of automatically scheduling the operations inside processing elements (PE), along with a router network that can effectively transform and control data flow among the PE array. This design lets the hardware share the responsibility for the compiler, making them cooperate to deal with the issuing, placement and routing problem. Evaluation of our study shows that our framework can reach on average 1.28, 1.30 and 1.33 higher than three state-of-the-art SISP CGRA using REGIMap, RS compile flow and the EPIMap approaches respectively. The area overhead is nearly 0.93% per token buffer entry for each PE relative to SISP CGRA. Zhongyuan Zhao 0004, Weiguang Sheng, Weifeng He, Zhigang Mao, Zhaoshi Li |
DATE | 3 |
| 2017 | A hardware-friendly hierarchical HEVC motion estimation algorithm for UHD applicationsabstractHigh Efficiency Video Coding (HEVC) standard has a superior video compression rate compared with previous H.264/AVC. At the same time, Ultra-high-definition (UHD) video applications are becoming a reality under the development of the display technology. In this paper, a hardware-friendly multi-layer HEVC motion estimation (ME) algorithm for UHD applications are proposed. To keep the computational regularity of the traditional full-search (FS) ME algorithm as well as reduce the computational complexity of ME in a large search range (SR), the basic layer of the proposed algorithm is to combine FS scheme in a core area with downsampling search scheme in a large peripheral area. Moreover, the finer layer of the algorithm employs a hexagon search scheme to perform further ME around the optimal match point generated by the basic layer. Integrating the proposed algorithm into the HM 15.0, experimental results show that our hardware-friendly algorithm can achieve 97.8% of computations reduction while only 0.77% of BD-rate loss on average. Consequently, the proposed algorithm is feasible for HEVC ME hardware design for UHD applications. Jiawei Gu, Guanghui He 0002, Weifeng He |
ISCAS | 4 |
| 2017 | A 0.2V 2.3pJ/Cycle 28dB output SNR hybrid Markov random field probabilistic-based circuit for noise immunity and energy efficiencyabstractIn this paper, two kinds of simplified cell structures for low voltage noise immunity and a hybrid Markov Random Field probabilistic-based circuit design technique are proposed to reduce the hardware overhead and improve the noise immunity. To demonstrate the proposed technique, four kinds of test chips with an 8-bit carry lookahead adder (CLA) are fabricated in a 130nm CMOS technology. Measurement results show the proposed hybrid MRF CLA improves 14% noise immunity, saves 53% energy consumption and reduces 11% circuit area than the other CLAs. Xuwei Jin, Wei Jin 0004, Hao Zhang 0151, Jianfei Jiang 0001, Weifeng He |
ISCAS | 5 |
| 2017 | A 12-bit 4928 × 3264 pixel CMOS image signal processor for digital still cameras
Wei Jin 0004, Guanghui He 0002, Weifeng He, Zhigang Mao |
Integr. | 3 |
| 2017 | A 0.33 V 2.5 μW cross-point data-aware write structure, read-half-select disturb-free sub-threshold SRAM in 130 nm CMOS
Wei Jin 0004, Weifeng He, Jianfei Jiang 0001, Haichao Huang, Xuejun Zhao, Yanan Sun 0003, Naifeng Jing |
Integr. | 2 |
| 2017 | In Situ Error Detection Techniques in Ultralow Voltage Pipelines: Analysis and OptimizationsabstractIn order to achieve high tolerance against process, voltage, and temperature variations in the ultralow voltage (ULV) circuits, in situ error detection and correction (EDAC) techniques were presented. However, circuits adding the capability of error detection incur large hardware overhead, especially in ULV due to larger delay variability. In this paper, we analyze the hardware overhead of error detection techniques in pipelines based on three different sequential elements: flip-flops, two-phase latches, and pulsed latches. By exploiting the cycle-borrowing ability, we propose a technique called sparse insertion of error detecting registers on the two-phase latch-based and pulsed-latch-based pipelines to reduce the sequential logic area. Furthermore, we propose a delay-padding methodology using a multi-Vtcell library in ULV circuits to reduce EDAC hardware overhead. The proposed techniques are applied on a benchmark six-stage pipeline operating at 0.35 V in a 65-nm CMOS. The analysis results show that our proposed techniques can reduce the total area by 26%-33% and the error detecting register count by 2.9-4.3× compared with conventional EDAC techniques. Wei Jin 0004, Seongjong Kim, Weifeng He, Zhigang Mao, Mingoo Seok |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2016 | Enabling in-situ logic-in-memory capability using resistive-RAM crossbar memoryabstractRecently, logic-in-memory (LIM) is gaining growing interest because it eliminates the unnecessary data movement between the memory and logic components that embarrasses both performance and power dissipation in modern microprocessors. However, most of the existing LIM just puts the logic and memory closer rather than a true integration due to the incompatibility of logic and memory circuit structures. In this paper, we propose a in-situ LIM design by leveraging the emerging ReRAM memory in a crossbar structure. It performs in-situ logic processing using the same memory cells based on the resistive states of ReRAMs with non-destructive operations, and therefore can exploit the large internal bandwidth available in the array without data readout. The logic exploration exposes that the proposed LIM can support different logical functions and get in-situ results without moving data in and out of the memory. We believe that the proposed design provides a promising solution for a true logic processing capability within memory. Naifeng Jing, Taozhong Li, Zhongyuan Zhao 0004, Wei Jin 0004, Yanan Sun 0003, Weifeng He, Zhigang Mao |
FPT | 6 |
| 2016 | High performance parallel turbo decoder with configurable interleaving network for LTE application
Zhiting Yan, Guanghui He 0002, Weifeng He, Shuaijie Wang, Zhigang Mao |
Integr. | 3 |
| 2015 | Redundancy based Interconnect Duplication to Mitigate Soft Errors in SRAM-based FPGAsabstractSoft error induced reliability problem has already become a major concern for modern SRAM-based FPGAs (Field Programmable Gate Arrays) even at the ground level. In this paper, we propose a duplication-with-recovery (DWR) technique to recover the configuration bit faults on interconnects, which contribute to the majority of soft errors in FPGAs. Based on a study on the detailed routing structure in real FPGAs, DWR leverages redundant resources for interconnect duplication and enables fault recovery with lightweight circuit-level support. Compared with traditional fault tolerant techniques, DWR retains the fault recovering capability but eliminates expensive copies. The experimental results show that a large portion of the interconnects can be protected, which in consequence significantly reduces the vulnerable configuration bits. In addition, DWR does not alter the placement and routing from standard design flow, and therefore does not affect the design closure but greatly improves the design reliability in a cost-effective way. Naifeng Jing, Jianfei Jiang 0001, Weifeng He, Zhigang Mao |
ICCAD | 5 |
| 2014 | Area and throughput efficient IDCT/IDST architecture for HEVC standardabstractHigh Efficiency Video Coding (HEVC) is new video coding standard beyond H.264/AVC. In this paper, an area and throughput efficient 2-D IDCT/IDST VLSI architecture for HEVC standard is presented. Adopting proposed data flow scheduling and shared constant multiplication structure, the architecture supports variable block size IDCT from 4×4 to 32×32 pixels as well as 4×4 pels IDST. Using 65nm technology, the synthesis results show that the maximum work frequency is 500MHz and the architecture hardware cost is about 145.4K gate count. Compared with previous work, our design achieves more than 50% reduction in hardware cost and 66% improvement in throughput efficiency. Experimental results show that the proposed architecture is able to deal with real-time HEVC IDCT/IDST of 4K×2K (4096×2048)@30 fps video sequence at 412MHz in average. In consequence, it offers a cost-effective solution for the future UHDTV applications. Ziyou Yao, Weifeng He, Guanghui He 0002, Zhigang Mao |
ISCAS | 2 |
| 2012 | A pre-emphasis circuit design for high speed on-chip global interconnectabstractOn-chip global interconnects are speed and power bottleneck in state-of-the-art chips. Pre-emphasis technique is an efficient way to improve the performance of the global communication. This paper first performs delay analysis of a global wire to work with a pre-emphasis circuit in time domain. Based on the analysis, a new pre-emphasis circuit design is proposed. Simulation results show that the pre-emphasis circuit can increase the link bandwidth by more than 40% and 20% in capacitive and capacitive-resistive coupled 10mm global link respectively. The new pre-emphasis circuit design can be applied in high speed global communication. Jianfei Jiang 0001, Weiguang Sheng, Zhigang Mao, Weifeng He |
ISCAS | 4 |
| 2012 | SEU fault evaluation and characteristics for SRAM-based FPGA architectures and synthesis algorithmsabstractReliability has become an increasingly important concern for SRAM-based field programmable gate arrays (FPGAs). Targeting SEU (single event upset) in SRAM-based FPGAs, this article first develops an SEU evaluation framework that can quantify the failure sensitivity for each configuration bit during design time. This framework considers detailed fault behavior and logic masking on a post-layout FPGA application and performs logic simulation on various circuit elements for fault evaluation. Applying this framework on MCNC benchmark circuits, we first characterize SEUs with respect to different FPGA circuits and architectures, for example, bidirectional routing and unidirectional routing. We show that in both routing architectures, interconnects not only contribute to the lion's share of the SEU-induced functional failures, but also present higher failure rates per configuration bits than LUTs. Particularly, local interconnect multiplexers in logic blocks have the highest failure rate per configuration bit. Then, we evaluate three recently proposed SEU mitigation algorithms, IPD, IPF, and IPV, which are all logic resynthesis-based with little or no overhead on placement and routing. Different fault mitigating capabilities at the chip level are revealed, and it demonstrates that algorithms with explicit consideration for interconnect significantly mitigate the SEU at the chip level, for example, IPV achieves 61% failure rate reduction on average against IPF with about 15%. In addition, the combination of the three algorithms delivers over 70% failure rate reduction on average at the chip level. The experiments also reveal that in order to improve fault tolerance at the chip level, it is necessary for future fault mitigation algorithms to concern not only LUT or interconnect faults, but also their interactions. We envision that our framework can be used to cast more useful insights for more robust FPGA circuits, architectures, and better synthesis algorithms. Naifeng Jing, Ju-Yueh Lee, Zhe Feng 0002, Weifeng He, Zhigang Mao, Lei He 0001 |
ACM Trans. Design Autom. Electr. Syst. | 4 |
| 2011 | Quantitative SEU Fault Evaluation for SRAM-Based FPGA Architectures and Synthesis AlgorithmsabstractThis paper studies the SEU (Single Event Upset) fault for SRAM-based FPGAs. Considering detailed fault behavior on various circuit elements in a post-layout FPGA application, we develop a simulation-based SEU evaluation tool that quantifies fault contribution for each configuration bit. Using this tool and MCNC benchmark circuits, we study the fault characteristics of FPGA circuits and architectures. We show that interconnects not only contribute to the lion share of functional failures, but also have higher failure rate per configuration bit than LUTs. Particularly, multiplexers in local interconnects have the highest failure rate per bit. We find that tuning LUT and cluster sizes helps to reduce the rate (up to 38% in our experiments). In addition, we evaluate two recent fault mitigation algorithms IPD and IPF, which reduce LUT faults by an average of 74% and 15% respectively. But when interconnects are taken into account, the reduction via IPD which considers only LUT faults is merely 6% on chip level. Yet the reduction via IPF which implicitly considers interconnect faults is still around 15%. Therefore, synthesis algorithm should be evaluated with interconnect faults and future algorithms should be developed with consideration of interconnect faults explicitly. Naifeng Jing, Ju-Yueh Lee, Zhe Feng 0002, Weifeng He, Zhigang Mao, Shi-Jie Wen, Richard Wong, Lei He 0001 |
FPL | 4 |
| 2011 | Mitigating FPGA interconnect soft errors by in-place LUT inversionabstractModern SRAM-based FPGAs (Field Programmable Gate Arrays) use multiplexer-based unidirectional routing, and SRAM configuration cells in these multiplexers contribute to the majority of soft errors in FPGAs. In this paper, we formulate an In-Placed inVersion (IPV) on LUT (Look-Up Table) logic polarities to reduce the Soft Error Rate (SER) at chip level, and reveal a locality and NP-Hardness of the IPV problem. We then develop an exact algorithm based on the binary integer linear programming (ILP) and also a heuristic based on the simulated annealing (SA), both enabled by the locality. We report results for the 10 largest MCNC combinational benchmarks synthesized by ABC and then placed and routed by VPR. The results show that IPV obtains close to 4× chip level SER reduction on average and SA is highly effective by obtaining the same SER reduction as ILP does. A recent work IPD has the largest LUT level SER reduction of 2.7× in literature, but its chip level SER reduction is merely 7% due to the dominance of interconnects. In contrast, SA-based IPV obtains nearly 4× chip level SER reduction and runs 30× faster. Furthermore, combining IPV and IPD leads to a chip level SER reduction of 5.3×. This does not change placement and routing, and does not affect design closure. To the best of our knowledge, our work is the first in-depth study on SER reduction for modern multiplexer-based FPGA routing by in-placed logic re-synthesis. Naifeng Jing, Ju-Yueh Lee, Weifeng He, Zhigang Mao, Lei He 0001 |
ICCAD | 3 |
| 2011 | Effective multi-standard macroblock prediction VLSI design for reconfigurable multimedia systemsabstractReconfigurable computing arrays facilitate the flexibility with high performance for regular and computation-intensive algorithms in multimedia processing. However, the efficiency of the irregular and control-intensive algorithms becomes the performance bottleneck of reconfigurable multimedia systems. In this paper, we propose the design and VLSI implementation of a novel memory efficient macroblock prediction and boundary strength (Bs) calculation engine. The control-intensive algorithms, including intra mode prediction, motion vector prediction, and Bs calculation, are implemented with 4x4 block level pipeline to achieve real-time decoding for H.264/AVC high profile and Chinese AVS Jizhun profile. Compared with existing designs, our design achieves 60% registers reduction for neighboring block load and update. Implementation results indicate that the proposed architecture can support 1920×1088@30fps of H.264 and AVS decoding at 86 MHz. Yuliang Tao, Guanghui He 0002, Weifeng He, Qin Wang 0009, Jun Ma 0012, Zhigang Mao |
ISCAS | 3 |
| 2011 | A thermal-aware task mapping flow for coarse-grain dynamic reconfigurable processorabstractThis paper presents a task level mapping flow for coarse-grained dynamic reconfigurable array processor based on static thermal-aware mapping techniques. The flow is composed of front-end SUIF tool, temporal partitioning algorithm, thermal aware sub-graph mapping algorithms and back-end RAM compiler to compile HLL task into binary code for the processor automatically. Using compact thermal model, the temperature distribution of each task sub-graph on reconfigurable RC array is pre-estimated statically. The runtime sequence of all task sub-graphs is generated ultimately with the random searching algorithm to balance the reconfigurable array's temperature. Experimental results show that the average maximum temperature and temperature distribution range can be reduced about 6.3°C and 12°C, respectively. Weifeng He, Naifeng Jing, Zhigang Mao |
ISCAS | 2 |
| 2011 | A clock-less transceiver for global interconnectabstractHigh speed and low power transceivers start to be used for global interconnection in state-of-the-art System-on-Chips (SoCs). In traditional transceivers, the bandwidth is largely dependent on the clock rate. This paper presents a clock-less transceiver for global interconnect. The asynchronous transceiver makes the data rate only depend on the link delay and can be conveniently used with low swing scheme to create a high speed and low power communication system. The transceiver is demonstrated and simulated. The simulation results indicate that the transceiver can be used in high speed and low power global communications. Jianfei Jiang 0001, Weiguang Sheng, Weifeng He, Zhigang Mao |
VLSI-SoC | 4 |
| 2011 | A 230mV 8-bit sub-threshold microprocessor for wireless sensor networkabstractA customized design flow for ultra-low power cell library and an 8-bit ultra-low power microprocessor for wireless sensor network application are presented in this paper. According to the logic pre-synthesis results of the 8-bit microprocessor HDL code, frequently used standard cells are collected to develop a customized sub-threshold cell library through size and structure modifications. The ultimate transistor-level netlist of the processor is generated by the RTL code re-synthesis results through cell substitution under the sub-threshold cell library. Using the SPICE simulator, experimental results show that our 8-bit microprocessor can work at a supply voltage as low as 230mV at full temperature range of all technology corners, which has a power only 79nW and the frequency 10 KHz at 230mV and room temperature. As a result, the proposed microprocessor provides a feasible solution for emerging energy-constrained applications. Wei Jin 0004, Weifeng He, Zhigang Mao |
VLSI-SoC | 3 |
| 2011 | Robust design of sub-threshold flip-flop cells for wireless sensor networkabstractAs a major sequential logic element, D flip-flop is an indispensible cell in logic cell library. In this paper, we proposed two improved sub-threshold D flip-flop circuits (mTGMS and emC2MOS D flip-flop) after conducting robustness analysis of several typical flip-flop circuits. Using SMIC 0.18um CMOS technology, the simulation results show that the minimum work voltage of our proposed mTGMS and emC2MOS is 0.19V and 0.18V, the minimum average power is 13.2pW and 14.1pW, while the minimum Power Delay Product (PDP) is 13aJ and 4.35aJ respectively. Wei Jin 0004, Weifeng He, Zhigang Mao |
VLSI-SoC | 3 |
| 2011 | A general statistical estimation for application mapping in Network-on-ChipabstractDesign space exploration is crucial to an optimal application mapping in Network-on-Chip. However, the optimality evaluation of the explored solution has been neglected in previous studies. In this paper, we propose an efficient and credible statistical estimation approach to evaluate the optimality of explored solutions with respect to the mapped communication, which is directly related to power dissipation in the network. Our approach is motivated by a basic statistical property on the solution space, and we consider the diversities in different complex on-chip network designs to make it more applicable. The statistical estimation and the optimality evaluation are validated in experiments by real and synthetic applications. It demonstrates an estimating error around 6% on average, which tends to be even smaller when problem scales up. We envision that the fidelity of our statistical estimation approach will promote its applicability in the promising Network-on-Chip designs. Naifeng Jing, Weifeng He, Zhigang Mao |
VLSI-SoC | 2 |
| 2010 | Statistical estimation and evaluation for communication mapping in Network-on-Chip
Naifeng Jing, Weifeng He, Yongxin Zhu 0001, Zhigang Mao |
Integr. | 2 |
| 2007 | An Improved Frame-Level Pipelined Architecture for High Resolution Video Motion EstimationabstractFrame-level pipelined motion estimation structure achieves high throughput by exploiting the explicit parallelism among motion estimation blocks. In this paper, an improved frame-level pipelined architecture for FSBM motion estimation is proposed. The design efforts are focused on reducing the size of internal data buffers and the hardware overheads for high resolution video motion estimation. Compared with previous high performance architectures, the proposed architecture employs the smallest number of data buffers as well as removes the data broadcasting operations and keeps nearly 100% fully pipelined computation. As a result, this architecture offers a feasible solution for SHDTV video pictures. Weifeng He, Zhigang Mao |
ISCAS | 1 |