EDBT 2026 Demo / reviewers in the wild / expert
Weiguang Sheng
dblp:86/4229 · also Wei-Guang Sheng
· DBLP profile ↗
33ranked-venue papers
3as first author
19since 2021 · last 2026
0000-0002-7831-526XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 27 · 1 first-author · 17 since 2021Software engineering, systems software and programming languages · 8 · 2 first-author · 3 since 2021Security and privacy · 3 · 2 first-authorDatabases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Viper: An ILP-Based Vectorization Framework for Fully Homomorphic Encryption
Weidong Yang 0007, Xinmo Li, Xiangmin Guo, Jianfei Jiang 0001, Naifeng Jing, Qin Wang 0009, Zhigang Mao, Weiguang Sheng |
ASP-DAC | 8 |
| 2026 | TPDA-DRAM: A Variation-Aware DRAM Improving System Performance via In-Situ Timing Margin Detection and Adaptive MitigationabstractDRAM latency remains a critical bottleneck in the performance of modern computing systems. However, the latency is excessively conservative due to the timing margins imposed by DRAM vendors to accommodate rare worst-case scenarios, such as weak cells and high temperatures. In this study, we introduce a temperature- and process-variation-aware timing detection and adaptation DRAM (TPDA-DRAM) architecture that dynamically mitigates timing margins at runtime. TPDA-DRAM leverages innovativein-situcross-coupled detectors to monitor voltage differences between bitline pairs inside DRAM arrays, ensuring precise detection of timing margins. Additionally, the proposed detector inherently accelerates the precharge operation of DRAM, thereby reducing the precharge latency by up to 62.5%. Building upon this architecture, we propose two variation-aware timing adaptation schemes: 1) a process-variation-aware adaptation (PVA) scheme that accelerates access to weak cells, mitigating process-induced timing margins, and 2) a temperature-variation-aware adaptation (TVA) scheme that leverages temperature information and the restoration truncation technique to reduce DRAM latency, mitigating temperature-induced timing margins. Evaluations on an eight-core computing system show that TPDA-DRAM improves average performance by 21.8% and energy efficiency by 18.2%. Yuxuan Qin, Chuxiong Lin, Guoming Rao, Weiguang Sheng, Weifeng He |
IEEE Trans. Computers | 5 |
| 2026 | HARMONY: A Hardware-Aware Mapping and Optimizing Framework for Computing-in-Memory AcceleratorsabstractThe increasing adoption of artificial intelligence has spurred the development of specialized deep neural network (DNN) accelerators. Among them, computing-in-memory (CIM) architectures are promising for their in-situ computation capability, which alleviates the computation and data movement bottlenecks of modern DNNs. However, the diversity of models and hardware designs makes it challenging to fully exploit CIM accelerators. Existing approaches often rely on manual mapping or provide limited automation, struggling to integrate general-purpose optimizations with CIM-specific features. In this work, we present HARMONY, a hardware-aware compilation framework for CIM accelerators. At its core is a hardware intermediate representation (IR) that unifies computational and memory abstractions. Based on this IR, HARMONY introduces an automatic mapping algorithm that identifies offloadable operators and constructs a hybrid software–hardware IR. This enables systematic integration of general-purpose and CIM-specific scheduling primitives within a unified search space, which is efficiently explored using reinforcement learning (RL). Extensive evaluations show that HARMONY supports a broader set of operators than existing CIM compilers and consistently delivers substantial performance and energy improvements across diverse DNN workloads. These results demonstrate that HARMONY provides both generality and efficiency, making it a practical compilation solution for CIM accelerators. Xinmo Li, Weidong Yang 0007, Naifeng Jing, Qin Wang 0009, Zhigang Mao, Weiguang Sheng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2026 | HiRe: A Hierarchical Reconfigurable Architecture for Large-Scale Multichiplet DNN AcceleratorsabstractMultichiplet deep neural network (DNN) accelerators have evolved as promising modular solutions, offering enhanced performance, scalability, and cost-effectiveness. These architectures, however, suffer from the escalating communication bottleneck with increasing scale, primarily stemming from rising hop count and worsening link underutilization. The bottleneck is exacerbated by conventional fixed interconnection networks’ nonadaptability to diverse DNN dataflows. Moreover, an efficient routing tailored for large-scale networks with deadlock-freedom is needed for performance. To address these scalability challenges, leveraging the low-latency links and abundant interconnection resources with reconfigurability in the active interposer, we propose HiRe, a hierarchical reconfigurable network-routing co-design architecture for large-scale multichiplet DNN accelerators. The architecture introduces reconfigurable nodes (RNs) across on-chip and interchiplet hierarchical networks, enabling dynamic bypassing and network reconfiguration. Based on the network, it incorporates an efficient deadlock-free routing that combines simulated annealing (SA)-based communication scheduling with greedy path selection. Through the network-routing co-design, HiRe reduces the hop count and enhances the link utilization. The HiRe architecture is implemented and synthesized in a 55-nm CMOS process. Experimental results demonstrate that HiRe achieves a 14.3%–45.2% EDP reduction and a 14.1%–35.0% latency reduction compared to state-of-the-art (SOTA) innovations, effectively mitigating the large-scale communication bottleneck. Dongxu Lyu, Jianfei Jiang 0001, Weiguang Sheng, Chen Zhang 0001, Guanghui He 0002 |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2025 | Principle-based Dataflow Optimization for Communication Lower Bound in Operator-Fused Tensor AcceleratorabstractAlthough design space exploration (DSE) is good at finding dataflow for optimal memory access in tensor accelerators, it is very timing-consuming and lacks architecture insight. In this study, we for the first time propose several principles for dataflow optimization that provides lower bound of memory communication for tensor operators such as matrix multiplication. Through these principles we can calculate the best tiling, scheduling and mapping for both intra- and inter-operator dataflow. In addition, we can identify all the tensor-wise opertor fusion that are profitable in memory communication, so we propose FuseCU, a new architecture that supports these profitable fusion which can be applied to existing spatial architectures for data movement saving. Experimental results show that FuseCU delivers 63.6%, 62.4% and 38.7% data movement saving and $1.33 \times, 1.25 \times$ and $1.14 \times$ speedup compared to the TPUv4i, Gemmini and Planaria designs without increasing buffer size or bandwidth. Additionally, FuseCU is open-sourced. Zelong Yuan, Weiguang Sheng, Jianfei Jiang 0001, Qin Wang 0009, Naifeng Jing |
DAC | 4 |
| 2025 | HEILP: An ILP-Based Scale Management Method for Homomorphic Encryption CompilerabstractRNS-CKKS, a fully homomorphic encryption (FHE) scheme, enabling secure computation on encrypted data, has widely be used in statistical analysis and data mining. However, developing RNS-CKKS programs requires substantial knowledge of cryptography, which is unfriendly to non-expert programmers. A critical obstacle is the scale management, which affects the complexity of programming and performance. Different FHE operations impose specific requirements on the scale and level, necessitating programmer intervention to ensure the recoverability of the results. Furthermore, operations at different levels have a significant impact on program performance. Existing methods rely on heuristic insights or iterative methods to manage the scales of ciphertexts. However, these methods lack a holistic understanding of the optimization space, leading to inefficient exploration and suboptimal performance. This work proposes HEILP, the first constrained-optimization-based approach for scale management in FHE. HEILP expresses node scale decision and scale management operation inserting as an integer linear programming model which can be solved with existing mathematical techniques in one shot. Our method creates a more comprehensive optimization space and enables a faster and more efficient exploration. Experimental results demonstrate that HEILP achieves an average performance improvement of 1.72 x over existing heuristic method, and outperforms a 1.19 x performance improvement with 48.65 x faster compilation time compared to the state-of-the-art iteration-based method. Weidong Yang 0007, Shuya Ji, Jianfei Jiang 0001, Naifeng Jing, Qin Wang 0009, Zhigang Mao, Weiguang Sheng |
DATE | 7 |
| 2025 | A Hierarchical 3-D Physical Design Method for Ultralarge-Scale Logic-on-Memory CGRA ChipabstractFace-to-face bonded 3-D (F2F 3D) technology, with the potential to significantly reduce chip area while enhancing performance, stands as one of the most promising ways to extend Moore’s Law. However, current 3-D physical design flows are often modifications of 2-D design flows and rely on technical personnel to manually modify technical files. Furthermore, existing research on 3-D design flow primarily focuses on module implementation, with very few studies addressing hierarchical design methods for large-scale chips. In this article, we first introduce a 3-D physical design flow which concurrently optimizes the timing of both the logic tier and the memory tier, achieving synchronized physical design for both tiers. Then, we develop a bottom-up hierarchical 3-D physical design flow to extend the 3-D design flow to large-scale chip design. Through coordinated power planning, clock tree design, and interconnect unit design, we enhance the power, performance, and area (PPA) metrics of the entire chip. Using our RTL-to-GDS physical design flow, we successfully implemented a 28-nm CMOS logic-on-memory (LoM) 3-D coarse-grained reconfigurable architecture (CGRA) chip with over 50 million gates. Experimental results demonstrate that our 3-D flow improves timing by 16.1% while reducing voltage drop by 38.6% compared to the 2-D design. In addition, the power-delay product (PDP) of the 3-D chip decreases by 10.2%, showcasing better performance. Zizheng Dong, Shuaipeng Li, Weijia Zhu, Ang Li 0045, Qin Wang 0009, Naifeng Jing, Weiguang Sheng, Jianfei Jiang 0001, Zhigang Mao |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2025 | IPDR: An Inter-Chiplet Priority-Driven Deadlock Resolution for 2-D/2.5-D Multichiplet Systems
Yaoyao Ye, Jianfei Jiang 0001, Weiguang Sheng, Ningyi Xu, Yong Lian 0001, Guanghui He 0002 |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2024 | VDA: A Simple but Efficient Virtual-Channel-Based Deadlock Avoidance Scheme for Scalable Chiplet NetworksabstractWith the escalating computation capability demands of AI and other applications, chiplet technology has emerged as a prominent force in the current market, offering scalability and cost-effectiveness. One of the most critical issues in chiplet-based systems lies in the implementation of deadlock-free routing in 2.5D architectures. However, existing routing algorithms for 2.5D chiplet-based networks typically impose turn restrictions or necessitate complex hardware modifications, posing significant obstacles to scalability and exponentially increasing design costs. To address existing issues, we propose VDA, a simple deadlock avoidance scheme with fully utilized virtual channels (VCs) and lightweight hardware overhead for scalable chiplet-based networks. By constructing a dedicated virtual network through VC assignment, we enable the existence of cyclic channel dependencies and reduce VC restrictions. Meanwhile, a loop topology at the interposer level is introduced to enhance transmission efficiency. Our evaluation demonstrates that VDA yields an average improvement of up to 32.96% in saturation throughput and reduces low-load latency by up to 13.62% under synthetic traffic patterns. Furthermore, our approach achieves an average runtime speedup of 1.7% ∼ 6.2% when executing realistic workload benchmarks compared to existing approaches, with only 0.2% area overhead. Duo Yu, Ang Li 0045, Naifeng Jing, Jianfei Jiang 0001, Weiguang Sheng, Qin Wang 0009 |
ACM Great Lakes Symposium on VLSI | 5 |
| 2024 | A Flexible and High-Precision Activation Function Unit Based on Equi-Error Partitioning AlgorithmabstractThe diversity of activation functions has gradually increased to accommodate different tasks in modern deep neural networks (DNNs). However, these novel activation functions involve more nonlinear operations relative to traditional activation functions, which increases the computational complexity. To address these issues, a piecewise linear (PWL) approximation algorithm called Equi-Error Partitioning Algorithm is proposed in this paper. The algorithm aims at balancing the errors between segments and solves the problem of excessive precision that exists in other PWL approximation methods and achieves on average 30.07× better mean squared error compared to the previous works. Based on this algorithm, we propose an activation function unit (AFU) which enables the addressing scheme of non-uniform segments and provides reconfigurability for all common activation functions by reloading parameters. End-to-end evaluation with several DNNs shows the accuracy loss is all less than 0.06% with 64 segments. Zelong Yuan, Siwei Yuan, Pengyu Liu 0004, Weiguang Sheng, Naifeng Jing |
ISCAS | 6 |
| 2024 | RecPIM: Efficient In-Memory Processing for Personalized Recommendation Inference Using Near-Bank ArchitectureabstractDeep learning (DL)-based personalized recommendation systems consume the major resources in modern AI data centers. The embedding layers with large memory capacity requirement and high bandwidth demand have been identified as the bottleneck of personalized recommendation inference. To mitigate the memory bandwidth bottleneck, near-memory processing (NMP) would be an effective solution which utilizes the through-silicon via (TSV) bandwidth within 3D-stacked DRAMs. However, existing NMP architectures suffer from the limited memory bandwidth caused by hard-to-scale TSVs. To overcome this obstacle, integrating the compute-logic near memory banks becomes a promising but challenging solution, since large memory capacity requirement limits the use of 3D-stacked DRAMs and irregular memory accesses lead to poor data locality, heavy TSV data traffic and low bank-level bandwidth utilization. To address this problem, we propose RecPIM, the first in-memory processing system for personalized recommendation inference using near-bank architecture based on 3D-stacked memory. From the hardware perspective, we introduce a heterogeneous memory system combined with 3D-stacked DRAM and DIMMs to accommodate large embedding tables and provide high bandwidth. By integrating processing logic units near memory banks on DRAM dies, our architecture can exploit the enormous bank-level bandwidth which is much higher than TSV bandwidth. Then, we integrate a small scratchpad memory to exploit the unique data reusability of DL-based personalized recommendation systems. Furthermore, we adopt a unidirectional data communication scheme to avoid additional cross-vault data transfer. From the software perspective, we present a customized programming model to facilitate memory management and task offloading. To reduce the data communication through TSVs and enhance the utilization of bank-level bandwidth, we develop an efficient data mapping scheme by partitioning the vector into smaller subvectors. Experimental results show that RecPIM achieves up to 2.58× speedup and 49.8% energy saving for data movement over the state-of-the-art NMP solution. Weidong Yang 0007, Shuya Ji, Jianfei Jiang 0001, Naifeng Jing, Qin Wang 0009, Zhigang Mao, Weiguang Sheng |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 8 |
| 2023 | An Efficient near-Bank Processing Architecture for Personalized Recommendation SystemabstractPersonalized recommendation systems consume the major resources in modern AI data centers. The memory-bound embedding layers with irregular memory access patterns have been identified as the bottleneck of recommendation systems. To overcome the memory challenges, near-memory processing (NMP) would be an effective solution which provides high bandwidth. Recent work proposes an NMP approach to accelerate the recommendation models by utilizing the through-silicon via (TSV) bandwidth in 3D-stacked DRAMs. However, the total bandwidth provided by TSVs is insufficient for a batch of embedding layers processed in parallel. In this paper, we propose a near-bank processing architecture to accelerate recommendation models. By integrating the compute-logic near memory banks on DRAM dies of the 3D-stacked DRAM, our architecture can exploit the enormous bank-level bandwidth which is much higher than TSV bandwidth. We also present a hardware/software interface for embedding layers offloading. Moreover, we propose an efficient mapping scheme to enhance the utilization of bank-level bandwidth. As a result, our architecture achieves up to 2.10X speedup and 31% energy saving for data movement over the state-of-the-art NMP solution for recommendation acceleration based on 3D-stacked memory. Weidong Yang 0007, Qin Wang 0009, Naifeng Jing, Jianfei Jiang 0001, Zhigang Mao, Weiguang Sheng |
ASP-DAC | 7 |
| 2023 | ACET: An Adaptive Clock Scheme Exploiting Comprehensive Timing Slack for Reconfigurable ProcessorsabstractTo ensure the correctness and reliability, digital circuits are designed with conservative timing margins to accommodate extreme variations in process, voltage, and temperature (PVT) and workload. However, worst-case scenarios rarely occur, leaving the reserved time margins unutilized, which leads to a waste of performance. This issue is particularly significant in reconfigurable processors, as they exhibit substantial workload timing slack in both spatial and temporal domains. Previous researches have mainly focused on either developing PVT slack or exploiting workload slack, but few have simultaneously considered both aspects. Additionally, directly applying existing timing enhancement techniques to reconfigurable processors is challenging due to their complex configurability and diminishing timing slack in array architectures.To address the above challenges, this paper introduces ACET, an Adaptive Clock scheme which Exploits Timing slack comprehensively through hardware-software co-optimization. On the hardware side, ACET incorporates an adaptive clock module that adjusts the clock period based on both workload and PVT conditions. The two conditions are obtained by employing a PVT delay monitor and encoding the workload-dependent delay into the configuration, respectively. Then timing information is transmitted to phase selection module for cycle-level adjustments, to leverage the temporal timing slack. On the software side, to further exploit the spatial timing slack, a scheduling algorithm is proposed, which heuristically rearranges the firing time of operations. Experiments demonstrate that ACET leads to an average performance increase of 70.1% or an equivalent energy saving of 35.6%, with the hardware overhead being only 0.56%. Shuya Ji, Weidong Yang 0007, Jianfei Jiang 0001, Naifeng Jing, Weiguang Sheng, Ang Li 0045, Qin Wang 0009 |
ICCD | 5 |
| 2023 | RTMDet-R2: An Improved Real-Time Rotated Object Detector
Haifeng Xiang, Naifeng Jing, Jianfei Jiang 0001, Weiguang Sheng, Zhigang Mao, Qin Wang 0009 |
PRCV (12) | 5 |
| 2022 | A Low Coupling and Lightweight Algorithm for Ship Detection in Optical Remote Sensing ImagesabstractIn recent years, many ship detection algorithms based on convolutional neural networks (CNNs) have been proposed to improve the performance of ship detection. However, with the increase in model complexity and size, it is challenging to deploy these models to resource-constrained edge platforms. In this letter, a low coupling algorithm that belongs to anchor-free methods is proposed for ship detection to reduce the model complexity and still obtain a competitive performance. The proposed low coupling network (LCNet) is easy to deploy and contributes to speeding up the inference and improving memory utilization. In addition, we propose a model compression process consisting of the quantization-aware training (QAT) method and a structural pruning method based on Taylor expansion, which can effectively reduce the model size according to hardware resource constraints. Comparative experimental results demonstrate that LCNet outperforms the state-of-the-art ship detection and natural object detection algorithms, with a 95.27% mAP and 88.91% F1 score on the HRSC2016 dataset. Our proposed model compression method also achieves a compression ratio of at least 80% with a negligible loss of performance. Guochao Deng, Qin Wang 0009, Jianfei Jiang 0001, Qirun Hong, Naifeng Jing, Weiguang Sheng, Zhigang Mao |
IEEE Geosci. Remote. Sens. Lett. | 6 |
| 2022 | An Efficient CNN Accelerator Using Inter-Frame Data Reuse of Videos on FPGAsabstractConvolutional neural networks (CNNs) have had great success when applied to computer vision technology, and many application-specific integrated circuit (ASIC) and field-programmable gate array (FPGA) CNN accelerators have been proposed. These accelerators primarily focus on the acceleration of a single input, and they are not particularly optimized for video applications. In this article, we focus on the similarities between continuous inputs in video, and we propose a YOLOv3-tiny CNN FPGA accelerator using incremental operation. The accelerator can skip the convolution operation of similar data between continuous inputs. We also use the Winograd algorithm to optimize the conv$3\times 3$operator in the YOLOv3-tiny network to further improve the accelerator’s efficiency. Experimental results show that our accelerator achieved 74.2 frames/s on ImageNet ILSVRC2015. Compared to the original network without Winograd algorithm and incremental operation, our design provides a$4.10\times $speedup. When compared with other YOLO network FPGA accelerators applied to video applications, our design provided a$3.13\times $–$18.34\times $normalized digital signal processor (DSP) efficiency and$1.10\times $–$14.2\times $energy efficiency. Shengzhao Li, Qin Wang 0009, Jianfei Jiang 0001, Weiguang Sheng, Naifeng Jing, Zhigang Mao |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2021 | Reducing Memory Access Conflicts with Loop Transformation and Data Reuse on Coarse-grained Reconfigurable ArchitectureabstractCoarse-Grained Reconfigurable Arrays (CGRAs) are promising to have low power consumption and high energy-efficiency characteristics as accelerators. Recent years, many research works focus on improving the programmability of the CGRAs by enabling the fast reconfiguration during execution. The performance of these CGRAs critically hinges upon the scheduling power of the compiler. One of the critical challenges is to reduce memory access conflicts using static compilation techniques. Memory accessing conflict brings the synchronization overhead which causes the pipelining stall and reduces CGRA performance. Existing compilers usually tackle this challenge by orchestrating the data placement of the on-chip global memory (OGM) in CGRA to let the parallel memory accesses avoid the bank conflict. However, we find bank conflict is not the only reason that causes the memory access conflicts. In some CGRAs, the bandwidth of the data network between OGM and processing element array (PEA) is also limited due to the low power design principle. The unbalanced network bandwidth loads is another reason that causes memory access conflicts. Furthermore, the redundant data access across iterations is one of the primary causes of memory access conflicts. Based on these observations, we provide a comprehensive and generalized compilation flow to reduce the memory conflicts. Firstly, we develop a loop transformation model to maximize the inter-iteration data reuse of the loops to reduce the memory accessing operations under the software pipelining scheme. Secondly, we enhance the bandwidth utilization of the network between OGM and PEA and avoid the bank conflict by providing a conflict-aware spatial mapping algorithm which can be easily integrated into existing CGRA modulo scheduling compilation flow. Experimental results show our method is capable of improving performance by an average of 44% comparing with state-of-the-art CGRA compiling flow. Yuge Chen, Zhongyuan Zhao 0004, Jianfei Jiang 0001, Guanghui He 0002, Zhigang Mao, Weiguang Sheng |
DATE | 6 |
| 2021 | Subgraph Decoupling and Rescheduling for Increased Utilization in CGRA ArchitectureabstractWhen coarse-grained reconfigurable array (CGRA) architecture is shifting towards general-purpose, some complex control flows, such as nested loop, conditional branch and data dependence, may embarrass it and reduce the processing element (PE) array utilization by breaking the intact dataflow graph (DFG) into multiple regions with inconsistent control regions. This paper proposes subgraph decoupling and rescheduling, which decouples the inconsistent regions into control-independent subgraphs. Each subgraph can be rescheduled with zero-cost domino context switching and parallelized to fully utilize the PE resources. Then, we propose lightweight hardware changes based on general CGRA architecture to enable our design. The experiment results show that our proposal can improve the performance and energy efficiency by 1.35× and 1.18× over a static-mapped CGRA (Plasticine), and by 1.27× and 1.45× over an instruction-driven CGRA (TIA). Qin Wang 0009, Jianfei Jiang 0001, Weiguang Sheng, Guanghui He 0002, Zhigang Mao, Naifeng Jing |
DATE | 4 |
| 2021 | A 3.85-Gb/s 8 × 8 Soft-Output MIMO Detector With Lattice-Reduction-Aided Channel PreprocessingabstractThis article presents an 8 × 8 lattice-reduction-aided (LRA) soft-output multiple-input multiple-output (MIMO) detector for Chinese enhanced ultrahigh throughput (EUHT) wireless local area network (LAN) standard. The preprocessing algorithm combining simplified-sorting Cholesky decomposition and low-complexity decoupled lattice reduction (LDLR) is proposed to reduce computational complexity and latency with parallelism improvement. In addition, K-best detection adopts a sorting-reduced strategy utilizing approximate ordered sequence. Compared with other published LRA K-best detection algorithms, simulation results show that our proposed algorithm has performance improvement. In addition, in order to save hardware resources, a folded K-best architecture and an optimized intermediate storage strategy are introduced. Furthermore, a fully pipelined VLSI architecture is designed in Semiconductor Manufacturing International Corporation (SMIC) 40-nm 1P9M technology to support the 8 × 8.64 -QAM MIMO-OFDM system. The detector can achieve 3.85-Gb/s data throughput at 641-MHz clock frequency with 0.71-μs latency. The proposed detector is competitive in terms of latency, throughput, and area efficiency to state-of-the-art works and can meet the data-rate requirement of the EUHT standard. Zhuojun Liang, Dongxu Lv, Chao Cui, Haibao Chen, Weifeng He, Weiguang Sheng, Naifeng Jing, Zhigang Mao, Guanghui He 0002 |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2020 | Enabling Resistive-RAM-based Activation Functions for Deep Neural Network AccelerationabstractThe Resistive-RAM (RRAM) based deep neural network (DNN) accelerators have shown great potential as they are good at solving matrix-vector multiplication (MVM). However, this computing paradigm does not benefit other NN operations like activation, which may be built upon various transcendental functions and require customized circuit as in current RRAM-based NN accelerators. In this paper, we propose the RRAM-CORDIC algorithm and crossbar design which enable various transcendental activation calculations on a RRAM crossbar just like MVM. By applying encoding and multi-iteration transformation, the RRAM-CORDIC can exploit higher MAC (multiply-and-accumulation) parallelism that is traditionally uneconomic in CMOS but now efficient in RRAM crossbar. In addition, it can work in a pipelined manner with high computing throughput. Experiment results show that the RRAM-CORDIC algorithm can sustain high accuracy on different transcendental functions, and deliver less than 0.5% NN accuracy loss on typical DNN inference. The elimination of CMOS circuit in turn can trade more computing resources for MVM in the same area budget that improves the performance up to 47% for different networks. Taozhong Li, Ning Guan, Qin Wang 0009, Guanghui He 0002, Weiguang Sheng, Zhigang Mao, Naifeng Jing |
ACM Great Lakes Symposium on VLSI | 6 |
| 2020 | Towards Higher Performance and Robust Compilation for CGRA Modulo SchedulingabstractCoarse-Grained Reconfigurable Architectures (CGRA) is a promising solution for accelerating computation intensive tasks due to its good trade-off in energy efficiency and flexibility. One of the challenging research topic is how to effectively deploy loops onto CGRAs within acceptable compilation time. Modulo scheduling (MS) has shown to be efficient on deploying loops onto CGRAs. Existing CGRA MS algorithms still suffer from the challenge of mapping loop with higher performance under acceptable compilation time, especially mapping large and irregular loops onto CGRAs with limited computational and routing resources. This is mainly due to the under utilization of the available buffer resources on CGRA, unawareness of critical mapping constraints and time consuming method of solving temporal and spatial mapping. This article focus on improving the performance and compilation robustness of the modulo scheduling mapping algorithm for CGRAs. We decomposes the CGRA MS problem into the temporal and spatial mapping problem and reorganize the processes inside these two problems. For the temporal mapping problem, we provide a comprehensive and systematic mapping flow that includes a powerful buffer allocation algorithm, and efficient interconnection & computational constraints solving algorithms. For the spatial mapping problem, we develop a fast and stable spatial mapping algorithm with backtracking and reordering mechanism. Our MS mapping algorithm is able to map loops onto CGRA with higher performance and faster compilation time. Experiment results show that given the same compilation time budget, our mapping algorithm generates higher compilation success rate. Among the successfully compiled loops, our approach can improve 5.4 to 14.2 percent performance and takes x24 to x1099 less compilation time in average comparing with state-of-the-art CGRA mapping algorithms. Zhongyuan Zhao 0004, Weiguang Sheng, Qin Wang 0009, Wenzhi Yin, Pengfei Ye, Jinchao Li, Zhigang Mao |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2019 | mRNA: Enabling Efficient Mapping Space Exploration for a Reconfiguration Neural AcceleratorabstractDeep learning accelerators have emerged to enable energy-efficient and high-throughput inference from edge devices such as self-driving cars and smartphones, to data centers for batch inference such as recommendation systems. However, the actual energy efficiency and throughput of a deep learning accelerator depends on the deep neural network (DNN) loop nest mapping on the processing element array of an accelerator. Moreover, the efficiency of a mapping dramatically changes by the target DNN layer dimensions and available hardware resources. Therefore, the optimal mapping search problem is a non-trivial high-dimensional optimization problem. Although several tools and frameworks exist for compiling to CPUs and GPUs, we lack similar tools for deep learning accelerators. To deal with the optimized mapping search problem in deep learning accelerators, we propose mRNA (mapper for reconfigurable neural accelerators), which automatically searches optimal mappings using heuristics based on domain knowledge about deep learning and an energy/runtime cost evaluation framework. mRNA targets MAERI, a recently proposed open-source deep learning accelerator that provides flexibility via reconfigurable interconnects, to run the unique mappings for each layer generated by mRNA. In realistic machine learning workloads from MLPerf, the optimal mappings identified by mRNA framework provides 15% to 26% lower runtime and 55% to 64% lower energy for convolutional layers and 24% to 67% lower runtime and maximum 67% lower energy for fully connected layers compared to simple reference mappings manually picked for each layer. Zhongyuan Zhao 0004, Hyoukjun Kwon, Sachit Kuhar, Weiguang Sheng, Zhigang Mao, Tushar Krishna |
ISPASS | 4 |
| 2019 | A New Cellular-Based Redundant TSV Structure for Clustered FaultsabstractDue to the winding level of the thinned wafers and the surface roughness of silicon dies, the quality of through-silicon vias (TSVs) varies during the fabrication and bonding process, which greatly reduces the yield of 3-D-ICs. The basic method to repair faulty TSVs (FTSVs) is to transfer the signals on FTSVs through regular TSVs. Many redundant TSV (RTSV) structures have been proposed to repair uniformly distributed FTSVs. For clustering FTSVs, a router-based RTSV structure appears to be a good scheme. But it is not an economical method, since the structure consumes many more hardware resources than normal structures. In this paper, we propose a cellular-based RTSV structure to utilize hardware resources more efficiently for a higher yield. We propose a corresponding algorithm for recovery-route searching. Simulation results show that for 1E6 TSVs and a TSV failure rate of 0.01%, our design consumes only 4.5% more area of all STSVs to achieve a yield above 99.9%. We compare our structure with several other designs and demonstrate the cost-effectiveness of the proposed technique. Qin Wang 0009, Zechen Liu, Jianfei Jiang 0001, Naifeng Jing, Weiguang Sheng |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2018 | Optimizing the data placement and transformation for multi-bank CGRA computing systemabstractThis paper provides a data placement optimization approach for Coarse-Grained Reconfigurable Architecture (CGRA) based computing platform in order to simultaneously optimize the performance of CGRA execution and data transformation between main memory and multi-bank memory. To achieve this goal, we have developed a performance model to evaluate the efficiency of data transformation and CGRA execution. This model is used for comparing the performances difference when using different data placement strategies. We search for the optimal data placement method by firstly choosing the method which generates the best CGRA execution efficiency from the candidates who can generate the optimal data transformation efficiency. Then we choose the best data placement strategy by comparing the performance of the selected strategy with the one generated through existing multi-bank optimization algorithm. Evaluation shows our approach is capable of optimizing the performance to 2.76x of state-of-the-art method when considering both data-transformation and CGRA execution efficiency. Zhongyuan Zhao 0004, Yantao Liu, Weiguang Sheng, Tushar Krishna, Qin Wang 0009, Zhigang Mao |
DATE | 3 |
| 2018 | Fail-Slow at Scale: Evidence of Hardware Performance Faults in Large Production Systems
Haryadi S. Gunawi, Riza O. Suminto, Russell Sears, Casey Golliher, Swaminathan Sundararaman, Tim Emami, Weiguang Sheng, Nematollah Bidokhti, Caitie McCaffrey, Gary Grider, Parks M. Fields, Kevin Harms, Robert B. Ross, Andree Jacobson, Robert Ricci, Kirk Webb, Peter Alvaro, H. Birali Runesha, Mingzhe Hao, Huaicheng Li |
FAST | 8 |
| 2018 | Fail-Slow at Scale: Evidence of Hardware Performance Faults in Large Production SystemsabstractFail-slow hardware is an under-studied failure mode. We present a study of 114 reports of fail-slow hardware incidents, collected from large-scale cluster deployments in 14 institutions. We show that all hardware types such as disk, SSD, CPU, memory, and network components can exhibit performance faults. We made several important observations such as faults convert from one form to another, the cascading root causes and impacts can be long, and fail-slow faults can have varying symptoms. From this study, we make suggestions to vendors, operators, and systems designers. Haryadi S. Gunawi, Riza O. Suminto, Russell Sears, Casey Golliher, Swaminathan Sundararaman, Tim Emami, Weiguang Sheng, Nematollah Bidokhti, Caitie McCaffrey, Deepthi Srinivasan, Biswaranjan Panda, Andrew Baptist, Gary Grider, Parks M. Fields, Kevin Harms, Robert B. Ross, Andree Jacobson, Robert Ricci, Kirk Webb, Peter Alvaro, H. Birali Runesha, Mingzhe Hao, Huaicheng Li |
ACM Trans. Storage | 8 |
| 2017 | A static-placement, dynamic-issue framework for CGRA loop acceleratorabstractThis paper presents a static-placement, dynamic-issue (SPDI) framework for the coarse-grained reconfigurable architecture (CGRA) in order to tackle the inefficiencies of the static-issue, static-placement (SISP) CGRA. This framework includes the compiler that statically places the operations and hardware design, a SPDI CGRA, that automatically schedule the operations. We stress on introducing the SPDI CGRA in this paper. This newly designed hardware model adds the token buffer, which is capable of automatically scheduling the operations inside processing elements (PE), along with a router network that can effectively transform and control data flow among the PE array. This design lets the hardware share the responsibility for the compiler, making them cooperate to deal with the issuing, placement and routing problem. Evaluation of our study shows that our framework can reach on average 1.28, 1.30 and 1.33 higher than three state-of-the-art SISP CGRA using REGIMap, RS compile flow and the EPIMap approaches respectively. The area overhead is nearly 0.93% per token buffer entry for each PE relative to SISP CGRA. Zhongyuan Zhao 0004, Weiguang Sheng, Weifeng He, Zhigang Mao, Zhaoshi Li |
DATE | 2 |
| 2015 | Designing ARINC653 Partition Constrained Scheduling for Secure Real Time Embedded AvionicsabstractBeing a high end embedded system, an avionic system calls for stringent real time constraints as well as secure guarantees. In terms of logical architecture, avionic systems have recently grown into the form of Integrated Modular Avionics (IMA) from the traditional federated avionics system whose redundancy level is overwhelming for modern large aircrafts. The key idea of IMA system lies in the rules of time and space partitioning, which guarantees system predictability and reliability. However, existing industrial practices of IMA partition and priority settings usually incur significant waste of resources, which would eventually lower the performance of IMA tasks in terms of latency or throughput. This issue was not properly addressed by previous researchers who assumed settings of priority variances and fixed partitions, which differ from practical applications. In this paper, a secure real time scheduling scheme with partition readjustment is proposed with inputs of features exhibited by tasks under partition. In our scheme, the resource costs are reduced by merging and restructuring partitions without compromising hard real time constraints. The simulation results of actual flight missions show that significant improvement by our method in terms of the average response time of tasks as well as number of partitions. Xiang Tao, Yongxin Zhu 0001, Yishu Mao 0001, Han Song, Weiguang Sheng, Weiwei Shi 0002 |
CSCloud | 7 |
| 2015 | Parasitic Parameters Impacts Investigation on Soft Error Rate by a Circuit Level FrameworkabstractIn highly reliable CMOS integrated circuits, parasitic parameters have dramatic impacts on SER(soft error rate) estimation and affect the design decision. We proposed a circuit level SER characterization framework(ASSET-SPI) to evaluate the impacts by conducting statistical fault injection experiments automatically on the circuit spice netlist containing parasitic parameters. Experiments on ISCAS benchmark circuits(implemented in 180nm process) demonstrate ASSET-SPI is feasible for circuit level SER evaluation. While experiments on inverter chains(implemented in 180, 130 and 65nm process) show parasitic parameters introduce -2.95% to 19.82% variation on SER. The results remind us that parasitic parameters should be considered in design time SER evaluation to avoid over pessimistic/optimistic SER estimation and inappropriate design decision. Weiguang Sheng, Zhongyuan Zhao 0004, Zhigang Mao |
PRDC | 1 |
| 2012 | A pre-emphasis circuit design for high speed on-chip global interconnectabstractOn-chip global interconnects are speed and power bottleneck in state-of-the-art chips. Pre-emphasis technique is an efficient way to improve the performance of the global communication. This paper first performs delay analysis of a global wire to work with a pre-emphasis circuit in time domain. Based on the analysis, a new pre-emphasis circuit design is proposed. Simulation results show that the pre-emphasis circuit can increase the link bandwidth by more than 40% and 20% in capacitive and capacitive-resistive coupled 10mm global link respectively. The new pre-emphasis circuit design can be applied in high speed global communication. Jianfei Jiang 0001, Weiguang Sheng, Zhigang Mao, Weifeng He |
ISCAS | 2 |
| 2011 | A clock-less transceiver for global interconnectabstractHigh speed and low power transceivers start to be used for global interconnection in state-of-the-art System-on-Chips (SoCs). In traditional transceivers, the bandwidth is largely dependent on the clock rate. This paper presents a clock-less transceiver for global interconnect. The asynchronous transceiver makes the data rate only depend on the link delay and can be conveniently used with low swing scheme to create a high speed and low power communication system. The transceiver is demonstrated and simulated. The simulation results indicate that the transceiver can be used in high speed and low power global communications. Jianfei Jiang 0001, Weiguang Sheng, Weifeng He, Zhigang Mao |
VLSI-SoC | 3 |
| 2009 | Soft error optimization of standard cell circuits based on gate sizing and multi-objective genetic algorithmabstractA radiation harden technique based on gate sizing and multi-objective genetic algorithm (MOGA) is developed to optimize the soft error tolerance of standard cell circuits. Soft error rate (SER), chip area and longest path delay are selected as the optimization goals and fast fitness evaluation algorithms for the three goals are developed and embedded into the MOGA. All the three goals are optimized simultaneously by optimally sizing the gates in the circuit, which is a complex NP-Complete problem and resolved by MOGA through exploring the global design space of the circuit. Syntax analysis technique is also employed to make the proposed framework can optimize not only pure combinational logic circuit but also the combinational parts of sequential logic circuit. Optimizing experiments carried out on ISCAS'85 and ISCAS'89 standard benchmark circuits show that the proposed optimization algorithm can decrease the SER 74.25% with very limited delay overhead (0.28%). Furthermore, the algorithm can also reduce the area for most of the circuit under test by average 5.23%. The proposed technique is proved to be better than other works in delay and area overhead and suitable to direct the design of soft error tolerance integrated circuits in high reliability realms. Weiguang Sheng, Liyi Xiao, Zhigang Mao |
DAC | 1 |
| 2008 | Versatile and Efficient Techniques for Speeding-Up Circuit Level Simulated Fault-Injection CampaignsabstractFault injection in circuit level has proved to be cumbersome and time-consuming when employed to characterize the soft error sensitivity of digital circuits, hence new generation of CAD tool is required to automate the faults insertion and the validation of soft error mitigation mechanisms of the circuits. This paper outlines the characteristics of a new fault-injection platform HSECT-SPI (HIT Soft Error Characterization Toolkit-Spice Based) and its evaluation in some benchmark circuits implemented with distinct processes and soft error hardening techniques. It also details some techniques devised and implemented within the platform to automate and speed-up the circuit level fault-injection experiments. Experimental results are provided, showing that the platform is efficient, accurate and can direct the design of soft error immune circuits with at least three orders of magnitudes speed gain. Weiguang Sheng, Liyi Xiao, Zhigang Mao |
PRDC | 1 |