EDBT 2026 Demo / reviewers in the wild / expert
Jianlei Yang 0001
dblp:99/9547-1
· DBLP profile ↗
74ranked-venue papers
15as first author
34since 2021 · last 2026
0000-0001-8424-7040ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 62 · 14 first-author · 27 since 2021Artificial intelligence and machine learning · 8 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SlimInfer: Accelerating Long-Context LLM Inference via Dynamic Token PruningabstractLong-context inference for Large Language Models (LLMs) is heavily limited by high computational demands. While several existing methods optimize attention computation, they still process the full set of hidden states at each layer, limiting overall efficiency. In this work, we propose SlimInfer, an innovative framework that aims to accelerate inference by directly pruning less critical prompt tokens during the forward pass. Our key insight is an information diffusion phenomenon: As information from critical tokens propagates through layers, it becomes distributed across the entire sequence. This diffusion process suggests that LLMs can maintain their semantic integrity when excessive tokens, even including these critical ones, are pruned in hidden states. Motivated by this, SlimInfer introduces a dynamic fine-grained pruning mechanism that accurately removes redundant tokens of hidden state at intermediate layers. This layer-wise pruning naturally enables an asynchronous KV cache manager that prefetches required token blocks without complex predictors, reducing both memory usage and I/O costs. Extensive experiments show that SlimInfer can achieve up to 2.53× time-to-first-token (TTFT) speedup and 1.88× end-to-end latency reduction for LLaMA3.1-8B-Instruct on a single RTX 4090, without sacrificing performance on LongBench. Lingkun Long, Rubing Yang, Yushi Huang, Desheng Hui, Jianlei Yang 0001 |
AAAI | 6 |
| 2026 | Focus-dLLM: Accelerating Long-Context Diffusion LLM Inference via Confidence-Guided Context FocusingabstractLingkun Long, Yushi Huang, Shihao Bai, Ruihao Gong, Jun Zhang, Ao Zhou, Jianlei Yang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Lingkun Long, Yushi Huang, Shihao Bai, Ruihao Gong, Jun Zhang 0004, Jianlei Yang 0001 |
ACL (1) | 7 |
| 2026 | EcoVLA: Energy-Efficient Device-Edge Co-inference for Vision-Language-Action Models Under Real-Time Constraints
Bo Dai 0015, Zeyu Hao, Lingkun Long, Chunming Hu, Jianlei Yang 0001 |
APPT | 8 |
| 2026 | MIREDO: MIP-Driven Resource-Efficient Dataflow Optimization for Computing-in-Memory AcceleratorabstractComputing-in-Memory (CIM) architectures have emerged as a promising solution for accelerating Deep Neural Networks (DNNs) by mitigating data movement bottlenecks. However, realizing the potential of CIM requires specialized dataflow optimizations, which are challenged by an expansive design space and strict architectural constraints. Existing optimization approaches often fail to fully exploit CIM accelerators, leading to noticeable gaps between theoretical and actual system-level efficiency. To address these limitations, we propose the MIREDO framework, which formulates dataflow optimization as a MixedInteger Programming (MIP) problem. MIREDO introduces a hierarchical hardware abstraction coupled with an analytical latency model designed to accurately reflect the complex data transfer behaviors within CIM systems. By jointly modeling workload characteristics, dataflow strategies, and CIM-specific constraints, MIREDO systematically navigates the vast design space to determine the optimal dataflow configurations. Evaluation results demonstrate that MIREDO significantly enhances performance, achieving up to $3.2 \times$ improvement across various DNN models and hardware setups. Xiaolin He, Cenlin Duan, Yingjie Qi, Xiao May, Jianlei Yang 0001 |
ASP-DAC | 5 |
| 2026 | CIMinus: Empowering Sparse DNN Workloads Modeling and Exploration on SRAM-Based CIM ArchitecturesabstractCompute-in-memory (CIM) has emerged as a pivotal direction for accelerating workloads in the field of machine learning, such as Deep Neural Networks (DNNs). However, the effectively exploitation of sparsity in CIM systems presents numerous challenges, due to the inherent limitations in their rigid array structures. Designing sparse DNN dataflows and developing efficient mapping strategies also become more complex when accounting for diverse sparsity patterns and the flexibility of a multi-macro CIM structure. Despite these complexities, there is still an absence of a unified systematic view and modeling approach for diverse sparse DNN workloads in CIM systems. In this paper, we propose CIMinus, a framework dedicated to cost modeling for sparse DNN workloads on CIM architectures. It provides an in-depth energy consumption analysis at the level of individual components and an assessment of the overall workload latency. We validate CIMinus against contemporary CIM architectures and demonstrate its applicability in two use-cases. These cases provide valuable insights into both the impact of sparsity patterns and the effectiveness of mapping strategies, bridging the gap between theoretical design and practical implementation. Yingjie Qi, Jianlei Yang 0001, Rubing Yang, Cenlin Duan, Xiaolin He, Ziyan He, Weitao Pan, Weisheng Zhao 0001 |
IEEE Trans. Computers | 2 |
| 2026 | GCoDE: Efficient Device-Edge Co-Inference for GNNs via Architecture-Mapping Co-SearchabstractGraph Neural Networks (GNNs) have emerged as the state-of-the-art graph learning method. However, achieving efficient GNN inference on edge devices poses significant challenges, limiting their application in real-world edge scenarios. This is due to the high computational cost of GNNs and limited hardware resources on edge devices, which prevent GNN inference from meeting real-time and energy requirements. As an emerging paradigm, device-edge co-inference shows potential for improving inference efficiency and reducing energy consumption on edge devices. Despite its potential, research on GNN device-edge co-inference remains scarce, and our findings show that traditional model partitioning methods are ineffective for GNNs. To address this, we propose GCoDE, the first automatic framework forGNN architecture-mappingCo-design and deployment onDevice-Edge hierarchies. By abstracting the device communication process into an explicit operation, GCoDE fuses the architecture and mapping scheme in a unified design space for joint optimization. Additionally, GCoDE’s system performance awareness enables effective evaluation of architecture efficiency across diverse heterogeneous systems. By analyzing the energy consumption of various GNN operations, GCoDE introduces an energy prediction method that improves energy assessment accuracy and identifies energy-efficient solutions. Using a constraint-based random search strategy, GCoDE identifies the optimal solution in 1.5 hours, balancing accuracy and efficiency. Moreover, the integrated co-inference engine in GCoDE enables efficient deployment and execution of GNN co-inference. Experimental results show that GCoDE can achieve up to 44.9× speedup and 98.2% energy reduction compared to existing approaches across diverse applications and system configurations. Jianlei Yang 0001, Yingjie Qi, Zhi Yang 0001, Weisheng Zhao 0001, Chunming Hu |
IEEE Trans. Computers | 2 |
| 2026 | Efficient SRAM-PIM Co-Design by Joint Exploration of Value-Level and Bit-Level SparsityabstractProcessing-in-memory (PIM) architectures mitigate the Von Neumann bottleneck by integrating computation units into memory arrays. Among PIM architectures, digital SRAMPIM has become a prominent approach, directly integrating digital logic within the SRAM array. However, the rigid crossbar architecture and full array activation pose challenges in efficiently utilizing value-level sparsity. Moreover, neural network models exhibit a high proportion of zero bits within non-zero values, which remain underutilized due to architectural constraints. To overcome these limitations, we present Dyadic Block PIM (DB-PIM), a groundbreaking algorithm-architecture co-design framework to harness both value-level and bit-level sparsity. At the algorithm level, our hybrid-grained pruning technique, combined with a novel sparsity pattern, enables effective sparsity management. Architecturally, DB-PIM incorporates a sparse network and customized digital SRAM-PIM macros, including input pre-processing unit (IPU), dyadic block multiply units (DBMUs), and Canonical Signed Digit (CSD)-based adder trees. It circumvents structured zero values in weights and bypasses unstructured zero bits within non-zero weights and block-wise all-zero bit columns in input features. As a result, the DBPIM framework skips a majority of unnecessary computations, thereby driving significant gains in computational efficiency. Experimental results demonstrate that our DB-PIM framework achieves up to 8.01× speedup and 85.28% energy savings, significantly boosting computational efficiency in digital SRAMPIM systems. Cenlin Duan, Jianlei Yang 0001, Yiou Wang, Yingjie Qi, Xiaolin He, Bonan Yan, Xiaotao Jia, Weisheng Zhao 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2026 | ACE-GNN: Adaptive GNN Co-Inference With System-Aware Scheduling in Dynamic Edge EnvironmentsabstractThe device-edge co-inference paradigm effectively bridges the gap between the high resource demands of Graph Neural Networks (GNNs) and limited device resources, making it a promising solution for advancing edge GNN applications. Existing research enhances GNN co-inference by leveraging offline model splitting and pipeline parallelism (PP), which enables more efficient computation and resource utilization during inference. However, the performance of these static deployment methods is significantly affected by environmental dynamics such as network fluctuations and multi-device access, which remain unaddressed. We present ACE-GNN, the first Adaptive GNN Co-inference framework tailored for dynamic Edge environments, to boost system performance and stability. ACE-GNN achieves performance awareness for complex multi-device access edge systems via system-level abstraction and two novel prediction methods, enabling rapid runtime scheme optimization. Moreover, we introduce a data parallelism (DP) mechanism in the runtime optimization space, enabling adaptive scheduling between PP and DP to leverage their distinct advantages and maintain stable system performance. Also, an efficient batch inference strategy and specialized communication middleware are implemented to further improve performance. Extensive experiments across diverse applications and edge settings demonstrate that ACE-GNN achieves a speedup of up to 12.7× and an energy savings of 82.3% compared to GCoDE, as well as 11.7× better energy efficiency than Fograph. Jianlei Yang 0001, Yingjie Qi, Xinming Wei, Cenlin Duan, Weisheng Zhao 0001, Chunming Hu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2026 | TinyFormer: Efficient Sparse Transformer Design and Deployment on Tiny DevicesabstractDeveloping deep learning models on tiny devices (e.g. Microcontroller units, MCUs) has attracted much attention in various embedded IoT applications. However, it is challenging to efficiently design and deploy recent advanced models (e.g. transformers) on tiny devices due to their severe hardware resource constraints. In this work, we proposeTinyFormer, a framework specifically designed to develop and deploy resource-efficient transformer models on MCUs. TinyFormer consists ofSuperNAS,SparseNAS, andSparseEngine. Separately, SuperNAS aims to search for an appropriate supernet from a vast search space. SparseNAS evaluates the best sparse single-path transformer model from the identified supernet. Finally, SparseEngine efficiently deploys the searched sparse models onto MCUs. To the best of our knowledge, SparseEngine is the first deployment framework capable of performing inference of sparse transformer models on MCUs. Evaluation results on the CIFAR-10 dataset demonstrate that TinyFormer can design efficient transformers with an accuracy of 96.1% while adhering to hardware constraints of 1MB storage and 320KB memory. Additionally, TinyFormer achieves significant speedups in sparse inference, up to$12.2\times $comparing to the CMSIS-NN library. TinyFormer is believed to bring powerful transformers into TinyML scenarios and to greatly expand the scope of deep learning applications. Jianlei Yang 0001, Jiacheng Liao, Fanding Lei, Meichen Liu, Lingkun Long, Han Wan, Bei Yu 0001, Weisheng Zhao 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 1 |
| 2025 | CIMFlow: An Integrated Framework for Systematic Design and Evaluation of Digital CIM ArchitecturesabstractDigital Compute-in-Memory (CIM) architectures have shown great promise in Deep Neural Network (DNN) acceleration by effectively addressing the “memory wall” bottleneck. However, the development and optimization of digital CIM accelerators are hindered by the lack of comprehensive tools that encompass both software and hardware design spaces. Moreover, existing design and evaluation frameworks often lack support for the capacity constraints inherent in digital CIM architectures. In this paper, we present CIMFlow, an integrated framework that provides an out-of-the-box workflow for implementing and evaluating DNN workloads on digital CIM architectures. CIMFlow bridges the compilation and simulation infrastructures with a flexible instruction set architecture (ISA) design, and addresses the constraints of digital CIM through advanced partitioning and parallelism strategies in the compilation flow. Our evaluation demonstrates that CIMFlow enables systematic prototyping and optimization of digital CIM architectures across diverse configurations, providing researchers and designers with an accessible platform for extensive design space exploration. Yingjie Qi, Jianlei Yang 0001, Yiou Wang, Dayu Wang, Cenlin Duan, Xiaolin He, Weisheng Zhao 0001 |
DAC | 2 |
| 2025 | Ultra Energy-Efficient Butterfly Counting in Bipartite Networks via Algorithm-Architecture Co-OptimizationabstractButterfly counting (BFC) problem, which counts the number of butterfly structure in a graph, is fundamental in bipartite network analysis. Recently, considerable efforts toward accelerations of BFC on both CPU and GPU platforms have been reported. However, the underlying BFC algorithms require repetitive vertex traversal and suffer from substantial latency and energy consumption concerns because data in large graphs has very limited reusability. In this paper, we introduce a hardware-software co-optimization approach to tackle these issues. A key innovation behind our approach is an algorithm that employs iterative lightweight arithmetical operations and facilitates highly parallel and pipelined processing. We further develop optimized data compression and pruning strategies to improve the efficiency of processing sparse data. These pivotal advancements are seamlessly integrated with a purpose-built hardware architecture to augment the overall implementation efficiency. Our proposed strategies are thoroughly evaluated on Zynq UltraScale+ FPGA platform. Compared with the state-of-the-art CPU (with 512 GB DRAM) and CPU+GPU (with 128 GB DRAM) implementations, our design achieve speedups of 15.84×and 1.35×, respectively, with only 4 GB DRAM. Meanwhile, our design’s energy efficiency is 50.14× over the CPU+GPU accelerator. Jianlei Yang 0001, Xiaotao Jia, Gang Qu 0001, Weisheng Zhao 0001 |
ICCAD | 3 |
| 2025 | Towards Affordable, Adaptive and Automatic GNN Training on CPU-GPU Heterogeneous PlatformsabstractGraph Neural Networks (GNNs) have been widely adopted due to their strong performance. However, GNN training often relies on expensive, high-performance computing platforms, limiting accessibility for many tasks. Profiling of representative GNN workloads indicates that substantial efficiency gains are possible on resource-constrained devices by fully exploiting available resources. This paper introduces$\mathrm{A}^{3} \text{GNN}$, a framework for Affordable, Adaptive, and Automatic GNN training on heterogeneous CPU-GPU platforms. It improves resource usage through locality-aware sampling and fine-grained parallelism scheduling. Moreover, it leverages reinforcement learning to explore the design space and achieve pareto-optimal trade-offs among throughput, memory footprint, and accuracy. Experiments show that$\mathrm{A}^{3}$GNN can bridge the performance gap, allowing seven Nvidia 2080Ti GPUs to outperform two A100 GPUs by up to$1.8 \times$in throughput with minimal accuracy loss. Yingjie Qi, Yiou Wang, Han Wan, Jianlei Yang 0001, Chunming Hu |
ICCD | 6 |
| 2025 | Finesse: An Agile Design Framework for Pairing-based Cryptography via Software/Hardware Co-DesignabstractPairing-based cryptography (PBC) is crucial in modern cryptographic applications.With the rapid advancement of adversarial research and the growing diversity of application requirements, PBC accelerators need regular updates in algorithms, parameter configurations, and hardware design.However, traditional design methodologies face significant challenges, including prolonged design cycles, difficulties in balancing performance and flexibility, and insufficient support for potential architectural exploration.To address these challenges, we introduce Finesse, an agile design framework based on co-design methodology.Finesse leverages a co-optimization cycle driven by a specialized compiler and a multi-granularity hardware simulator, enabling both optimized performance metrics and effective design space exploration.Furthermore, Finesse adopts a modular design flow to significantly shorten design cycles, while its versatile abstraction ensures flexibility across various curve families and hardware architectures.Finesse offers flexibility, efficiency, and rapid prototyping, comparing with previous frameworks.With compilation times reduced to minutes, Finesse enables faster iteration cycles and streamlined hardware-software co-design.Experiments on popular curves * Both authors contributed equally to this research. Tianwei Pan, Tianao Dai, Jianlei Yang 0001, Hongbin Jing, Zeyu Hao, Xiaotao Jia, Chunming Hu, Weisheng Zhao 0001 |
ISCA | 3 |
| 2024 | Towards Efficient SRAM-PIM Architecture Design by Exploiting Unstructured Bit-Level SparsityabstractBit-level sparsity in neural network models harbors immense untapped potential. Eliminating redundant calculations of randomly distributed zero-bits significantly boosts computational efficiency. Yet, traditional digital SRAM-PIM architecture, limited by rigid crossbar architecture, struggles to effectively exploit this unstructured sparsity. To address this challenge, we propose Dyadic Block PIM (DB-PIM), a groundbreaking algorithm-architecture co-design framework. First, we propose an algorithm coupled with a distinctive sparsity pattern, termed a dyadic block (DB), that preserves the random distribution of non-zero bits to maintain accuracy while restricting the number of these bits in each weight to improve regularity. Architecturally, we develop a custom PIM macro that includes dyadic block multiplication units (DBMUs) and Canonical Signed Digit (CSD)-based adder trees, specifically tailored for Multiply-Accumulate (MAC) operations. An input pre-processing unit (IPU) further refines performance and efficiency by capitalizing on block-wise input sparsity. Results show that our proposed co-design framework achieves a remarkable speedup of up to 7.69× and energy savings of 83.43%. Cenlin Duan, Jianlei Yang 0001, Yiou Wang, Yingjie Qi, Xiaolin He, Bonan Yan, Xiaotao Jia, Weisheng Zhao 0001 |
DAC | 2 |
| 2024 | WinoGen: A Highly Configurable Winograd Convolution IP Generator for Efficient CNN Acceleration on FPGAabstractThe convolution neural network (CNN) has been widely adopted in computer vision tasks. In the FPGA-based CNN accelerator design, Winograd convolution can effectively improve computation performance and save hardware resources. However, building efficient and highly compatible IP for arbitrary Winograd convolution on FPGA remains underexplored. To address this issue, we propose a novel and efficient reformulation of Winograd convolution, named Structured Direct Winograd Convolution (SDW). We further develop WinoGen, a Chisel-based highly configurable Winograd convolution IP generator. Given arbitrary input/output tile size and kernel size, it can generate optimized high-performance IP automatically. Meanwhile, our generated IP can be compatible with multiple kernel sizes and tile sizes. Experimental results show that the IP generated by WinoGen achieves DSP efficiency up to 3.80 GOPS/DSP and energy efficiency up to 652.77 GOPS/W while showing 2.45× and 3.10× improvements when processing a same CNN model compared with state-of-the-arts. Pengjia Li, Shixin Chen, Beichen Li 0003, Chong Tong, Jianlei Yang 0001, Tinghuan Chen, Bei Yu 0001 |
DAC | 7 |
| 2024 | GNNavigator: Towards Adaptive Training of Graph Neural Networks via Automatic Guideline ExplorationabstractGraph Neural Networks (GNNs) succeed significantly in many applications recently. However, balancing GNNs training runtime cost, memory consumption, and attainable accuracy for various applications is non-trivial. Previous training methodologies suffer from inferior adaptability and lack a unified training optimization solution. To address the problem, this work proposes GNNavigator, an adaptive GNN training configuration optimization framework. GN-Navigator meets diverse GNN application requirements due to our unified software-hardware co-abstraction, proposed GNNs training performance model, and practical design space exploration solution. Experimental results show that GNNavigator can achieve up to 3.1× speedup and 44.9% peak memory reduction with comparable accuracy to state-of-the-art approaches. Jianlei Yang 0001, Yingjie Qi, Bei Yu 0001, Weisheng Zhao 0001, Chunming Hu |
DAC | 2 |
| 2024 | Graph Neural Networks Automated Design and Deployment on Device-Edge Co-Inference SystemsabstractThe key to device-edge co-inference paradigm is to partition models into computation-friendly and computation-intensive parts across the device and the edge, respectively. However, for Graph Neural Networks (GNNs), we find that simply partitioning without altering their structures can hardly achieve the full potential of the co-inference paradigm due to various computational-communication overheads of GNN operations over heterogeneous devices. We present GCoDE, the first automatic framework for GNN that innovatively Co-designs the architecture search and the mapping of each operation on Device-Edge hierarchies. GCoDE abstracts the device communication process into an explicit operation and fuses the search of architecture and the operations mapping in a unified space for joint-optimization. Also, the performance-awareness approach, utilized in the constraint-based search process of GCoDE, enables effective evaluation of architecture efficiency in diverse heterogeneous systems. We implement the co-inference engine and runtime dispatcher in GCoDE to enhance the deployment efficiency. Experimental results show that GCoDE can achieve up to 44.9× speedup and 98.2% energy reduction compared to existing approaches across various applications and system configurations. Jianlei Yang 0001, Yingjie Qi, Zhi Yang 0001, Weisheng Zhao 0001, Chunming Hu |
DAC | 2 |
| 2024 | LLP-ECCA: A Low-Latency and Programmable Framework for Elliptic Curve Cryptography AcceleratorsabstractElliptic curve cryptography (ECC) plays a pivotal role in safeguarding data integrity and authentication in contemporary communication contexts, particularly within the domain of Intelligent Transport Systems (ITS). In the realm of ITS, vehicles communicate via the V2X (vehicle-to-everything) protocol, necessitating low-latency responses and minimal power consumption. Given the evolving nature of V2X protocol standards across the globe, programmability becomes a rigid requirement. However, existing strategies cannot meet all these vehicular equipment demands. This paper introduces a novel framework tailored for ECC acceleration to address the issues. Specifically, we propose the design of an Application Specific Instruction Set Processor (ASIP), augmented by pipeline and dual-issue techniques. Furthermore, the envisioned ASIP integrates a hybrid control framework founded on Finite State Machines (FSM), facilitating agile and effective management. Notably, a general GF(p256) Barrett modular multiplier is specially devised to optimize latency and area utilization. Experimental results on Xilinx Kintex Ultrscale+ FPGA demonstrate that the proposed ECC accelerator generates a signature within 131us and verifies a message within 181us, and the performance meets the requirements of today’s V2X standard. Tianao Dai, Jianlei Yang 0001, Zhaojun Lu, Xiaotao Jia, Gang Qu 0001, Weisheng Zhao 0001 |
ITC-Asia | 4 |
| 2024 | HGNAS: Hardware-Aware Graph Neural Architecture Search for Edge DevicesabstractGraph Neural Networks (GNNs) are becoming increasingly popular for graph-based learning tasks such as point cloud processing due to their state-of-the-art (SOTA) performance. Nevertheless, the research community has primarily focused on improving model expressiveness, lacking consideration of how to design efficient GNN models for edge scenarios with real-time requirements and limited resources. Examining existing GNN models reveals varied execution across platforms and frequent Out-Of-Memory (OOM) problems, highlighting the need for hardware-aware GNN design. To address this challenge, this work proposes a novel hardware-aware graph neural architecture search framework tailored for resource constraint edge devices, namely HGNAS. To achieve hardware awareness, HGNAS integrates an efficient GNN hardware performance predictor that evaluates the latency and peak memory usage of GNNs in milliseconds. Meanwhile, we study GNN memory usage during inference and offer a peak memory estimation method, enhancing the robustness of architecture evaluations when combined with predictor outcomes. Furthermore, HGNAS constructs a fine-grained design space to enable the exploration of extreme performance architectures by decoupling the GNN paradigm. In addition, the multi-stage hierarchical search strategy is leveraged to facilitate the navigation of huge candidates, which can reduce the single search time to a few GPU hours. To the best of our knowledge, HGNAS is the first automated GNN design framework for edge devices, and also the first work to achieve hardware awareness of GNNs across different platforms. Extensive experiments across various applications and edge devices have proven the superiority of HGNAS. It can achieve up to a$10.6\boldsymbol{\times}$speedup and an$82.5\%$peak memory reduction with negligible accuracy loss compared to DGCNN on ModelNet40. Jianlei Yang 0001, Yingjie Qi, Yumeng Shi, Cenlin Duan, Weisheng Zhao 0001, Chunming Hu |
IEEE Trans. Computers | 2 |
| 2024 | DDC-PIM: Efficient Algorithm/Architecture Co-Design for Doubling Data Capacity of SRAM-Based Processing-in-MemoryabstractProcessing-in-memory (PIM), as a novel computing paradigm, provides significant performance benefits from the aspect of effective data movement reduction. SRAM-based PIM has been demonstrated as one of the most promising candidates due to its endurance and compatibility. However, the integration density of SRAM-based PIM is much lower than other nonvolatile memory-based ones, due to its inherent 6T structure for storing a single bit. Within comparable area constraints, SRAM-based PIM exhibits notably lower capacity. Thus, aiming to unleash its capacity potential, we propose DDC-PIM, an efficient algorithm/architecture co-design methodology that effectively doubles the equivalent data capacity. At the algorithmic level, we propose a filter-wise complementary correlation (FCC) algorithm to obtain a bitwise complementary pair. At the architecture level, we exploit the intrinsic cross-coupled structure of 6T SRAM to store the bitwise complementary pair in their complementary states$(Q/\overline {Q})$, thereby maximizing the data capacity of each SRAM cell. The dual-broadcast input structure and reconfigurable unit support both depthwise and pointwise convolution, adhering to the requirements of various neural networks. Evaluation results show that DDC-PIM yields about$2.84\times $speedup on MobileNetV2 and$2.69\times $on EfficientNet-B0 with negligible accuracy loss compared with PIM baseline implementation. Compared with state-of-the-art SRAM-based PIM macros, DDC-PIM achieves up to$8.41\times $and$2.75\times $improvement in weight density and area efficiency, respectively. Cenlin Duan, Jianlei Yang 0001, Xiaolin He, Yingjie Qi, Yiou Wang, Ziyan He, Bonan Yan, Xiaotao Jia, Weitao Pan, Weisheng Zhao 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2024 | An Energy-Efficient Bayesian Neural Network Implementation Using Stochastic Computing MethodabstractThe robustness of Bayesian neural networks (BNNs) to real-world uncertainties and incompleteness has led to their application in some safety-critical fields. However, evaluating uncertainty during BNN inference requires repeated sampling and feed-forward computing, making them challenging to deploy in low-power or embedded devices. This article proposes the use of stochastic computing (SC) to optimize the hardware performance of BNN inference in terms of energy consumption and hardware utilization. The proposed approach adopts bitstream to represent Gaussian random number and applies it in the inference phase. This allows for the omission of complex transformation computations in the central limit theorem-based Gaussian random number generating (CLT-based GRNG) method and the simplification of multipliers as AND operations. Furthermore, an asynchronous parallel pipeline calculation technique is proposed in computing block to enhance operation speed. Compared with conventional binary radix-based BNN, SC-based BNN (StocBNN) realized by FPGA with 128-bit bitstream consumes much less energy consumption and hardware resources with less than 0.1% accuracy decrease when dealing with MNIST/Fashion-MNIST datasets. Xiaotao Jia, Huiyi Gu, Jianlei Yang 0001, Weitao Pan, Youguang Zhang, Sorin Cotofana, Weisheng Zhao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2023 | Hardware-Aware Graph Neural Network Automated Design for Edge Computing PlatformsabstractGraph neural networks (GNNs) have emerged as a popular strategy for handling non-Euclidean data due to their state-of-the-art performance. However, most of the current GNN model designs mainly focus on task accuracy, lacking in considering hardware resources limitation and real-time requirements of edge application scenarios. Comprehensive profiling of typical GNN models indicates that their execution characteristics are significantly affected across different computing platforms, which demands hardware awareness for efficient GNN designs. In this work, HGNAS is proposed as the first Hardware-aware Graph Neural Architecture Search framework targeting resource constraint edge devices. By decoupling the GNN paradigm, HGNAS constructs a fine-grained design space and leverages an efficient multi-stage search strategy to explore optimal architectures within a few GPU hours. Moreover, HGNAS achieves hardware awareness during the GNN architecture design by leveraging a hardware performance predictor, which could balance the GNN model accuracy and efficiency corresponding to the characteristics of targeted devices. Experimental results show that HGNAS can achieve about 10.6× speedup and 88.2% peak memory reduction with a negligible accuracy loss compared to DGCNN on various edge devices, including Nvidia RTX3080, Jetson TX2, Intel i7-8700K and Raspberry Pi 3B+. Jianlei Yang 0001, Yingjie Qi, Yumeng Shi, Weisheng Zhao 0001, Chunming Hu |
DAC | 2 |
| 2023 | Lossy and Lossless (L2) Post-training Model Size CompressionabstractDeep neural networks have delivered remarkable performance and have been widely used in various visual tasks. However, their huge sizes cause significant inconvenience for transmission and storage. Many previous studies have explored model size compression. However, these studies often approach various lossy and lossless compression methods in isolation, leading to challenges in achieving high compression ratios efficiently. This work proposes a post-training model size compression method that combines lossy and lossless compression in a unified way. We first propose a unified parametric weight transformation, which ensures different lossy compression methods can be performed jointly in a post-training manner. Then, a dedicated differentiable counter is introduced to guide the optimization of lossy compression to arrive at a more suitable point for later lossless compression. Additionally, our method can easily control a desired global compression ratio and allocate adaptive ratios for different layers. Finally, our method can achieve a stable 10× compression ratio without sacrificing accuracy and a 20× compression ratio with minor accuracy loss in a short time. Our code is available at https://github.com/ModelTC/L2_Compression. Yumeng Shi, Shihao Bai, Xiuying Wei, Ruihao Gong, Jianlei Yang 0001 |
ICCV | 5 |
| 2023 | NAND-SPIN-based processing-in-MRAM architecture for convolutional neural network acceleration
Yinglin Zhao, Jianlei Yang 0001, Bing Li 0017, Xingzhou Cheng, Xucheng Ye, Xiaotao Jia, Zhaohao Wang, Youguang Zhang, Weisheng Zhao 0001 |
Sci. China Inf. Sci. | 2 |
| 2023 | IMGA: Efficient In-Memory Graph Convolution Network Aggregation With Data Flow OptimizationsabstractAggregating features from neighbor vertices is a fundamental operation in graph convolution network (GCN). However, the sparsity in graph data creates poor spatial and temporal locality, causing dynamic and irregular memory access patterns and limiting the performance of aggregation on the Von Neumann architecture. The emerging processing-in-memory (PIM) architecture is based on emerging nonvolatile memory (NVM), like spin-orbit torque magnetic RAM (SOT-MRAM), and demonstrates promising prospects in alleviating the Von Neumann bottleneck. However, the limited memory capacity of PIM medium still incurs non-negligible data movements between PIM architecture and external memory. To solve this challenge, we propose an SOT-MRAM-based in-memory computing architecture, called IMGA, for efficient in-situ graph aggregation. Specifically, we design adaptive data flow management strategies that reuse vertex data in MRAM when processing graphs of different scales and adopt edge data as the control signal source to utilize the graph’s structural information. A reordering optimization strategy leveraging hardware–software co-design principle is proposed to further reduce the costly data movement. Experimental results demonstrate that IMGA achieves an average$2523\times $and$21\times $speedup, and 1.03E+6 and 1.04E+3 energy efficiency compared with CPU and GPU, respectively. Yuntao Wei, Shangtong Zhang, Jianlei Yang 0001, Xiaotao Jia, Zhaohao Wang, Gang Qu 0001, Weisheng Zhao 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2022 | Eventor: an efficient event-based monocular multi-view stereo accelerator on FPGA platformabstractEvent cameras are bio-inspired vision sensors that asynchronously represent pixel-level brightness changes as event streams. Event-based monocular multi-view stereo (EMVS) is a technique that exploits the event streams to estimate semi-dense 3D structure with known trajectory. It is a critical task for event-based monocular SLAM. However, the required intensive computation workloads make it challenging for real-time deployment on embedded platforms. In this paper, Eventor is proposed as a fast and efficient EMVS accelerator by realizing the most critical and time-consuming stages including event back-projection and volumetric ray-counting on FPGA. Highly paralleled and fully pipelined processing elements are specially designed via FPGA and integrated with the embedded ARM as a heterogeneous system to improve the throughput and reduce the memory footprint. Meanwhile, the EMVS algorithm is reformulated to a more hardware-friendly manner by rescheduling, approximate computing and hybrid data quantization. Evaluation results on DAVIS dataset show that Eventor achieves up to 24X improvement in energy efficiency compared with Intel i5 CPU platform. Jianlei Yang 0001, Yingjie Qi, Meng Dong, Yuhao Yang 0008, Runze Liu 0001, Weitao Pan, Bei Yu 0001, Weisheng Zhao 0001 |
DAC | 2 |
| 2022 | Triangle Counting Accelerations: From Algorithm to In-Memory Computing ArchitectureabstractTriangles are the basic substructure of networks and triangle counting (TC) has been a fundamental graph computing problem in numerous fields such as social network analysis. Nevertheless, like other graph computing problems, due to the high memory-computation ratio and random memory access pattern, TC involves a large amount of data transfers thus suffers from the bandwidth bottleneck in the traditional Von-Neumann architecture. To overcome this challenge, in this paper, we propose to accelerate TC with the emerging processing-in-memory (PIM) architecture through an algorithm-architecture co-optimization manner. To enable the efficient in-memory implementations, we come up to reformulate TC with bitwise logic operations (such as AND), and develop customized graph compression and mapping techniques for efficient data flow management. With the emerging computational Spin-Transfer Torque Magnetic RAM (STT-MRAM) array, which is one of the most promising PIM enabling techniques, the device-to-architecture co-simulation results demonstrate that the proposed TC in-memory accelerator outperforms the state-of-the-art GPU and FPGA accelerations by 12.2x and 31.8x, respectively, and achieves a 34x energy efficiency improvement over the FPGA accelerator. Jianlei Yang 0001, Yinglin Zhao, Xiaotao Jia, Rong Yin 0001, Xuhang Chen 0001, Gang Qu 0001, Weisheng Zhao 0001 |
IEEE Trans. Computers | 2 |
| 2022 | S2 Engine: A Novel Systolic Architecture for Sparse Convolutional Neural NetworksabstractConvolutional neural networks (CNNs) have achieved great success in performing cognitive tasks. However, execution of CNNs requires a large amount of computing resources and generates heavy memory traffic, which impose a severe challenge on computing system design. Through optimizing parallel executions and data reuse in convolution, systolic architecture demonstrates great advantages in accelerating CNN computations. However, regular internal data transmission path in traditional systolic architecture prevents the systolic architecture from completely leveraging the benefits introduced by neural network sparsity.Deployment of fine-grained sparsity on the existing systolic architectures is greatly hindered by the incurred computational overheads.In this work, we propose S2Engine a novel systolic architecture that can fully exploit the sparsity in CNNs with maximized data reuse. S2Engine transmits compressed data internally and allows each processing element to dynamically select an aligned data from the compressed dataflow in convolution. Compared to the naive systolic array, S2Engine achieves about 3.2 and about 3.0 improvements on speed and power efficiency, respectively. Jianlei Yang 0001, Wenzhi Fu, Xingzhou Cheng, Xucheng Ye, Pengcheng Dai, Weisheng Zhao 0001 |
IEEE Trans. Computers | 1 |
| 2022 | Accelerating Graph-Connected Component Computation With Emerging Processing-In-Memory ArchitectureabstractComputing the connected component (CC) of a graph is a basic graph computing problem, which has numerous applications like graph partitioning and pattern recognition. Existing methods for computing CC suffer from memory wall problems because of the frequent data transmission between CPU and memory. To overcome this challenge, in this article, we propose to accelerate CC computation with the emerging processing-in-memory (PIM) architecture through an algorithm–architecture co-design manner. The innovation lies in computing CC with bitwise logical operations (such as AND and OR), and the customized data flow management methods to accelerate computation and reduce energy consumption. As a proof of concept, experimental results with computational spin-transfer torque magnetic RAM (STT-MRAM) arrays demonstrate on average$19.8\times $and$12.4\times $speedups compared with the CPU and GPU implementations, and a$35.4 \times $energy efficiency improvement over the CPU implementation. Moreover, we investigate the potential associations between graph computing and bitwise Boolean logic, which could help design more general in-memory graph computing accelerators in the future. Xuhang Chen 0001, Xiaotao Jia, Jianlei Yang 0001, Gang Qu 0001, Weisheng Zhao 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2022 | Reconfigurable and Dynamically Transformable In-Cache-MPUF System With True Randomness Based on the SOT-MRAMabstractIn this paper, we present a reconfigurable Physically Unclonable Functions (PUF) based on the Spin-Orbit-Torque Magnetic Random-Access Memory (SOT-MRAM), which exploits thermal noise as the true dynamic entropy source. Therefore, the MRAM cells could be configured to random final states with stochastic switching mechanism. The proposed PUF is constructed and reconfigured by combining the small-capacity true random number generator (TRNG) and high-reliability secure hash algorithm (SHA-512), realizing the dynamic transformation between SOT-MRAM based last level cache and PUF (In-Cache-MPUF). Thanks to the full reconfigurability and the high endurance of SOT-MRAM, the proposed In-Cache-MPUF can achieve$10^{\textbf {14}}$maximum PUF bits per cell, which has greatly motivated the implementations compared with the traditional weak PUFs utilizing the static entropy source of process variations. The Monte-Carlo simulation results using 40 nm technology and a compact MTJ model show that the proposed PUF has desirable randomness as the digitized bit streams passing all the NIST tests, achieving 50.0428% uniqueness as well as 49.9236% uniformity. It also shows comparable reliability to the state-of-the-art works: a maximum bit error rate of 0.14% and 0.12% at 100 °C and 0.9 V, respectively. In addition, the system level performance is tested and validated by gem5. Zhengyi Hou, Zhaohao Wang, Chao Wang 0094, Min Wang 0033, You Wang 0002, Cenlin Duan, Jianlei Yang 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 8 |
| 2021 | FedSkel: Efficient Federated Learning on Heterogeneous Systems with Skeleton Gradients UpdateabstractFederated learning aims to protect users' privacy while performing data analysis from different participants. However, it is challenging to guarantee the training efficiency on heterogeneous systems due to the various computational capabilities and communication bottlenecks. In this work, we propose FedSkel to enable computation-efficient and communication-efficient federated learning on edge devices by only updating the model's essential parts, named skeleton networks. FedSkel is evaluated on real edge devices with imbalanced datasets. Experimental results show that it could achieve up to 5.52x speedups for CONV layers' back-propagation, 1.82x speedups for the whole training process, and reduce 64.8% communication cost, with negligible accuracy loss. Junyu Luo 0002, Jianlei Yang 0001, Xucheng Ye, Xin Guo 0008, Weisheng Zhao 0001 |
CIKM | 2 |
| 2021 | Brief Industry Paper: optimizing Memory Efficiency of Graph Neural Networks on Edge Computing PlatformsabstractGraph neural networks (GNN) have achieved state-of-the-art performance on various industrial tasks. However, the poor efficiency of GNN inference and frequent Out-of-Memory (OOM) problem limit the successful application of GNN on edge computing platforms. To tackle these problems, a feature decomposition approach is proposed for memory efficiency optimization of GNN inference. The proposed approach could achieve outstanding optimization on various GNN models, covering a wide range of datasets, which speeds up the inference by up to 3×. Furthermore, the proposed feature decomposition could significantly reduce the peak memory usage (up to 5× in memory efficiency improvement) and mitigate OOM problems during GNN inference. Jianlei Yang 0001, Yeqi Gao, Yingjie Qi, Yunli Chen, Pengcheng Dai, Weisheng Zhao 0001, Chunming Hu |
RTAS | 2 |
| 2021 | Fast Physics-Based Electromigration Analysis for Full-Chip Networks by Efficient Eigenfunction-Based SolutionabstractElectromigration (EM) becomes one of the most challenging reliability issues for current and future ICs in 10-nm technology and below. In this article, a novel method is proposed for the EM hydrostatic stress analysis on 2-D multibranch interconnect trees, which is the foundation of the EM reliability assessment for large-scale on-chip interconnect networks, such as on-chip power grid networks. The proposed method, which is based on an eigenfunction technique, could efficiently calculate the hydrostatic stress evolution for multibranch interconnect trees stressed with different current densities and nonuniformly distributed thermal effects. The proposed method solves the partial differential equations of transient EM stress more efficiently since it does not require any discretization either spatially or temporally, which is in contrast to numerical methods, such as the finite difference method and finite element method. The accuracy of the proposed transient analysis approach is validated against the analytical solution and commercial tools. The convergence of the proposed method is demonstrated by numerical experiments on practical power/ground networks, showing that only a small number of eigenfunction terms are necessary for the accurate solution. Thanks to its analytical nature, the proposed method is also utilized in efficient EM analysis techniques, such as searching for the void nucleation time by a modified bisection algorithm. The numerical results show that the proposed method is 10X-100X faster than the finite difference method and scales better for larger interconnect trees. Shaobin Ma, Sheldon X.-D. Tan, Chase Cook, Liang Chen 0025, Jianlei Yang 0001, Wenjian Yu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2021 | Efficient Computation Reduction in Bayesian Neural Networks Through Feature Decomposition and MemorizationabstractThe Bayesian method is capable of capturing real-world uncertainties/incompleteness and properly addressing the overfitting issue faced by deep neural networks. In recent years, Bayesian neural networks (BNNs) have drawn tremendous attention to artificial intelligence (AI) researchers and proved to be successful in many applications. However, the required high computation complexity makes BNNs difficult to be deployed in computing systems with a limited power budget. In this article, an efficient BNN inference flow is proposed to reduce the computation cost and then is evaluated using both software and hardware implementations. A feature decomposition and memorization (DM) strategy is utilized to reform the BNN inference flow in a reduced manner. About half of the computations could be eliminated compared with the traditional approach that has been proved by theoretical analysis and software validations. Subsequently, in order to resolve the hardware resource limitations, a memory-friendly computing framework is further deployed to reduce the memory overhead introduced by the DM strategy. Finally, we implement our approach in Verilog and synthesize it with a 45-nm FreePDK technology. Hardware simulation results on multilayer BNNs demonstrate that, when compared with the traditional BNN inference method, it provides an energy consumption reduction of 73% and a 4× speedup at the expense of 14% area overhead. Xiaotao Jia, Jianlei Yang 0001, Runze Liu 0001, Sorin Cotofana, Weisheng Zhao 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2020 | SparseTrain: Exploiting Dataflow Sparsity for Efficient Convolutional Neural Networks TrainingabstractTraining Convolutional Neural Networks (CNNs) usually requires a large number of computational resources. In this paper, SparseTrain is proposed to accelerate CNN training by fully exploiting the sparsity. It mainly involves three levels of innovations: activation gradients pruning algorithm, sparse training dataflow, and accelerator architecture. By applying a stochastic pruning algorithm on each layer, the sparsity of back-propagation gradients can be increased dramatically without degrading training accuracy and convergence rate. Moreover, to utilize both natural sparsity (resulted from ReLU or Pooling layers) and artificial sparsity (brought by pruning algorithm), a sparse-aware architecture is proposed for training acceleration. This architecture supports forward and back-propagation of CNN by adopting 1-Dimensional convolution dataflow. We have built a cycle-accurate architecture simulator to evaluate the performance and efficiency based on the synthesized design with 14nm FinFET technologies. Evaluation results on AlexNet/ResNet show that SparseTrain could achieve about 2.7× speedup and 2.2× energy efficiency improvement on average compared with the original training process. Pengcheng Dai, Jianlei Yang 0001, Xucheng Ye, Xingzhou Cheng, Junyu Luo 0002, Linghao Song, Yiran Chen 0001, Weisheng Zhao 0001 |
DAC | 2 |
| 2020 | TCIM: Triangle Counting Acceleration With Processing-In-MRAM ArchitectureabstractTriangle counting (TC) is a fundamental problem in graph analysis and has found numerous applications, which motivates many TC acceleration solutions in the traditional computing platforms like GPU and FPGA. However, these approaches suffer from the bandwidth bottleneck because TC calculation involves a large amount of data transfers. In this paper, we propose to overcome this challenge by designing a TC accelerator utilizing the emerging processing-in-MRAM (PIM) architecture. The true innovation behind our approach is a novel method to perform TC with bitwise logic operations (such as AND), instead of the traditional approaches such as matrix computations. This enables the efficient in-memory implementations of TC computation, which we demonstrate in this paper with computational Spin-Transfer Torque Magnetic RAM (STT-MRAM) arrays. Furthermore, we develop customized graph slicing and mapping techniques to speed up the computation and reduce the energy consumption. We use a device-to-architecture co-simulation framework to validate our proposed TC accelerator. The results show that our data mapping strategy could reduce 99.99% of the computation and 72% of the memory WRITE operations. Compared with the existing GPU or FPGA accelerators, our in-memory accelerator achieves speedups of 9× and 23.4×, respectively, and a 20.6× energy efficiency improvement over the FPGA accelerator. Jianlei Yang 0001, Yinglin Zhao, Yingjie Qi, Meichen Liu, Xingzhou Cheng, Xiaotao Jia, Gang Qu 0001, Weisheng Zhao 0001 |
DAC | 2 |
| 2020 | Accelerating CNN Training by Pruning Activation Gradients
Xucheng Ye, Pengcheng Dai, Junyu Luo 0002, Xin Guo 0008, Yingjie Qi, Jianlei Yang 0001, Yiran Chen 0001 |
ECCV (25) | 6 |
| 2020 | Towards Systems Education for Artificial Intelligence: A Course Practice in Intelligent Computing ArchitecturesabstractWith the rapid development of artificial intelligence (AI) community, education in AI is receiving more and more attentions. There have been many AI related courses in the respects of algorithms and applications, while not many courses in system level are seriously taken into considerations. In order to bridge the gap between AI and computing systems, we are trying to explore how to conduct AI education from the perspective of computing systems. In this paper, a course practice in intelligent computing architectures are provided to demonstrate the system education in AI era. The motivation for this course practice is first introduced as well as the learning orientations. The main goal of this course aims to teach students for designing AI accelerators on FPGA platforms. The elaborated course contents include lecture notes and related technical materials. Especially several practical labs and projects are detailed illustrated. Finally, some teaching experiences and effects are discussed as well as some potential improvements in the future. Jianlei Yang 0001, Xiaopeng Gao, Weisheng Zhao 0001 |
ACM Great Lakes Symposium on VLSI | 1 |
| 2020 | Dual-Plane Switch Architecture for Time-Triggered EthernetabstractTime-triggered Ethernet (TTE) technology introduces the concept of time-triggered on the basis of traditional Ethernet, so that it can achieve conflict-free and deterministic service forwarding without sacrificing compatibility. However, storage resources in industrial, aviation, aerospace and other equipment are limited. Therefore, it is important for TTEthernet to develop switching technologies with high storage efficiency and scalability. This paper proposes a dual plane switching (DPS) architecture for TTEthernet, which divides time-triggered services and event-triggered services into two planes for data forwarding. Experimental results show that using the TTE switch of this architecture has the advantages of high clock synchronization accuracy, high throughout, low transmission delay and small jitter of TTE service. Meng Dong, Zhiliang Qiu, Weitao Pan, Chenglei Kong, Jianlei Yang 0001 |
ACM Great Lakes Symposium on VLSI | 7 |
| 2020 | Computing-in-Memory Architecture Based on Field-Free SOT-MRAM with Self-Reference MethodabstractOn the current computing platforms, the memory wall between processor and memory has become the toughest challenge for the traditional Von-Neumann computer architecture. Computing-in-Memory (CIM) is taken as a promising approach to solving the above bottleneck in computing systems. In this paper, we propose a CIM platform with field-free spinorbit torque magnetic random access memory (SOT-MRAM). The self-reference (SelfRef) method is designed to enhance the read reliability and directly obtain logic results through memory-like read operations without adding logic cells. Memory read/write and logic operations, including NOT, AND/NAND and OR/NOR, can be implemented in the same SOT-MRAM chip. The speed and power penalties caused by SelfRef scheme are acceptable thanks to the ultrafast switching of the SOT. The read reliability and logic correctness of the proposed CIM are demonstrated by hybrid simulation on a 40 nm technology node. Chao Wang 0094, Zhaohao Wang, Jianlei Yang 0001, Youguang Zhang, Weisheng Zhao 0001 |
ISCAS | 4 |
| 2020 | TIPRDC: Task-Independent Privacy-Respecting Data Crowdsourcing Framework for Deep Learning with Anonymized Intermediate RepresentationsabstractThe success of deep learning partially benefits from the availability of various large-scale datasets. These datasets are often crowdsourced from individual users and contain private information like gender, age, etc. The emerging privacy concerns from users on data sharing hinder the generation or use of crowdsourcing datasets and lead to hunger of training data for new deep learning applications. One naive solution is to pre-process the raw data to extract features at the user-side, and then only the extracted features will be sent to the data collector. Unfortunately, attackers can still exploit these extracted features to train an adversary classifier to infer private attributes. Some prior arts leveraged game theory to protect private attributes. However, these defenses are designed for known primary learning tasks, the extracted features work poorly for unknown learning tasks. To tackle the case where the learning task may be unknown or changing, we present TIPRDC, a task-independent privacy-respecting data crowdsourcing framework with anonymized intermediate representation. The goal of this framework is to learn a feature extractor that can hide the privacy information from the intermediate representations; while maximally retaining the original information embedded in the raw data for the data collector to accomplish unknown learning tasks. We design a hybrid training method to learn the anonymized intermediate representation: (1) an adversarial training process for hiding private information from features; (2) maximally retain original information using a neural-network-based mutual information estimator. We extensively evaluate TIPRDC and compare it with existing methods using two image datasets and one text dataset. Our results show that TIPRDC substantially outperforms other existing methods. Our work is the first task-independent privacy-respecting data crowdsourcing framework. Ang Li 0005, Yixiao Duan, Huanrui Yang, Yiran Chen 0001, Jianlei Yang 0001 |
KDD | 5 |
| 2020 | An STT-MRAM based reconfigurable computing-in-memory architecture for general purpose computing
Xiaotao Jia, Jianlei Yang 0001, Weisheng Zhao 0001 |
CCF Trans. High Perform. Comput. | 6 |
| 2020 | Prototyping federated learning on edge computing systems
Jianlei Yang 0001, Yixiao Duan, Huanyu Zhou, Jingyuan Wang 0001, Weisheng Zhao 0001 |
Frontiers Comput. Sci. | 1 |
| 2020 | Hardware Security in Spin-based Computing-in-memory: Analysis, Exploits, and Mitigation TechniquesabstractComputing-in-memory (CIM) is proposed to alleviate the processor-memory data transfer bottleneck in traditional von Neumann architectures, and spintronics-based magnetic memory has demonstrated many facilitation in implementing CIM paradigm. Since hardware security has become one of the major concerns in circuit designs, this article, for the first time, investigates spin-based computing-in-memory (SpinCIM) from a security perspective. We focus on two fundamental questions: (1) How can the new SpinCIM computing paradigm be exploited to enhance hardware security?; (2) What security concerns has this new SpinCIM computing paradigm incurred? Jianlei Yang 0001, Yinglin Zhao, Xiaotao Jia, Gang Qu 0001, Weisheng Zhao 0001 |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2020 | SPINBIS: Spintronics-Based Bayesian Inference System With Stochastic ComputingabstractBayesian inference is an effective approach for solving statistical learning problems, especially with uncertainty and incompleteness. However, Bayesian inference is a computing-intensive task whose efficiency is physically limited by the bottlenecks of conventional computing platforms. In this paper, a spintronics-based stochastic computing (SC) approach is proposed for efficient Bayesian inference. The inherent stochastic switching behaviors of spintronic devices are exploited to build a stochastic bitstream generator (SBG) for SC with hybrid CMOS/magnetic tunnel junction (MTJ) circuits design. Aiming to improve the inference efficiency, an SBG sharing strategy is leveraged to reduce the required SBG array scale by integrating a switch network between SBG array and SC logic. A device-to-architecture level framework is proposed to evaluate the performance of spintronics-based Bayesian inference system (SPINBIS). Experimental results on data fusion applications have shown that SPINBIS could improve the energy efficiency about 12× than MTJ-based approach with 45% design area overhead and about 26× than FPGA-based approach. Xiaotao Jia, Jianlei Yang 0001, Pengcheng Dai, Runze Liu 0001, Yiran Chen 0001, Weisheng Zhao 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2020 | A Novel High Performance and Energy Efficient NUCA Architecture for STT-MRAM LLCs With Thermal ConsiderationabstractAs the speed gap of the modern processor and the off-chip main memory enlarges, on-chip cache capacity increases to sustain the performance scaling. As a result, the cache power occupies a large portion of the total power budget. Spin transfer torque magnetic memory (STT-MRAM) is proposed as a promising solution for the low power cache design due to its high integration density and ultralow leakage power. Nevertheless, the high write power and latency of STT-MRAM become new barriers for the commercialization of this emerging technology. In this paper, we investigate the thermal effect on the access performance of STT-MRAM, and observe that the temperature can affect the write delay and energy significantly. Then, we explore the nonuniform cache access (NUCA) design of the chip-multiprocessors with STT-MRAM-based last level cache (LLC). A thermal aware data migration policy, called “Thermosiphon,” which takes advantage of the thermal property of STT-MRAM, is proposed to reduce the LLC write energy. This policy splits the LLC into different regions dynamically based on the thermal distribution monitored by thermal sensors available on-chip, and adaptively migrates write intensive data among different thermal regions considering the thermal gradient. Compared to the conventional NUCA design, our proposed design can save 41.2% write energy at most and 13.01% on average with negligible hardware overhead. Bi Wu 0002, Pengcheng Dai, Yuanqing Cheng, Ying Wang 0001, Jianlei Yang 0001, Zhaohao Wang, Dijun Liu, Weisheng Zhao 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2019 | eSLAM: An Energy-Efficient Accelerator for Real-Time ORB-SLAM on FPGA PlatformabstractSimultaneous Localization and Mapping (SLAM) is a critical task for autonomous navigation. However, due to the computational complexity of SLAM algorithms, it is very difficult to achieve real-time implementation on low-power platforms. We propose an energy-efficient architecture for real-time ORB (Oriented-FAST and Rotated-BRIEF) based visual SLAM system by accelerating the most time-consuming stages of feature extraction and matching on FPGA platform. Moreover, the original ORB descriptor pattern is reformed as a rotational symmetric manner which is much more hardware friendly. Optimizations including rescheduling and parallelizing are further utilized to improve the throughput and reduce the memory footprint. Compared with Intel i7 and ARM Cortex-A9 CPUs on TUM dataset, our FPGA realization achieves up to 3× and 31× frame rate improvement, as well as up to 71× and 25× energy efficiency improvement, respectively. Runze Liu 0001, Jianlei Yang 0001, Yiran Chen 0001, Weisheng Zhao 0001 |
DAC | 2 |
| 2019 | Magnetic Skyrmion-Based Neural Recording System Design for Brain Machine InterfaceabstractNext-generation brain machine interface demand a high-channel-count neural recording system to wirelessly monitor activities of thousands of neurons. In order to achieve high-density neural recording, further development of single recording channel comprised of a neural amplifier front-end (AFE) and an analog-to-digit converter (ADC) is critical. Despite the great progress made in CMOS implementation of custom-designed neural recording system, hybrid limitations of increasing area and power consumption in line with Moore's law drove great demand for post-CMOS substitutes. Magnetic skyrmion with nano particle-like and non-volatile properties are of both fundamental and applied interests for future bio-inspired electronics. In this work, we propose a compact model including both AFE and ADC based on current-induced skyrmion motion. The proposed system achieved a power consumption of 0.63 pJ/channel with an area overhead of 0.14 μm2. The purpose of this work is to explore the feasibility of magnetic skyrmion for building large-scale, dense neuronal recording system which could pave a new way for future brain machine interface application. Biao Pan, Wang Kang 0001, Xing Chen 0012, Jinyu Bai, Jianlei Yang 0001, Youguang Zhang, Weisheng Zhao 0001 |
ISCAS | 5 |
| 2019 | SR-WTA: Skyrmion Racing Winner-Takes-All Module for Spiking Neural ComputingabstractSpiking neural network (SNN) has emerged as one of the popular architectures in complex pattern recognition and classification tasks. However, hardware implementation of such algorithms using conventional CMOS based neuron consume resources and power that are orders of magnitude higher than that in human brain. This can be attributed to the mismatch of the computational architecture between biological brain and the current Boolean logic computing platform. Magnetic skyrmions have been intensively studied as a prospective information carrier in neuromorphic computing hardware design. In this work, a compact time-domain skyrmion-racing winner-takes-all (SR-WTA) leaky-integrate-fire (LIF) spiking neuron network is presented for the first time. The skyrmion motion dynamics in the LIF neuron and the behaviors of the neuron network was investigated comprehensively. Both SPICE and micromagnetic simulations are performed to evaluate the functionality and performance of the proposed SR-WTA based SNN. Biao Pan, Wang Kang 0001, Xing Chen 0012, Jinyu Bai, Jianlei Yang 0001, Youguang Zhang, Weisheng Zhao 0001 |
ISCAS | 5 |
| 2019 | Exploiting Spin-Orbit Torque Devices As Reconfigurable Logic for Circuit ObfuscationabstractCircuit obfuscation is a frequently used approach to conceal logic functionalities in order to prevent reverse engineering attacks on fabricated chips. Efficient obfuscation implementations are expected with lower design complexity and overhead but higher attack difficulties. In this paper, an emerging obfuscation approach is proposed by leveraging spin-orbit torque (SOT) devices-based look-up-tables as reconfigurable logic to replace the carefully selected gates. It is essentially impossible to identify the obfuscated gate with SOTs inside according to the physical geometry characteristics because the configured functionalities are represented by magnetization states. Such an obfuscation approach makes the circuit security further improved with high exponential attack complexities. Experiments on MCNC and ISCAS 85/89 benchmark suits show that the proposed approach could reduce the area overheads due to obfuscation by 10% averagely. Jianlei Yang 0001, Qiang Zhou 0001, Zhaohao Wang, Hai Li 0001, Yiran Chen 0001, Weisheng Zhao 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2018 | Spintronics based stochastic computing for efficient Bayesian inference systemabstractBayesian inference is an effective approach for solving statistical learning problems especially with uncertainty and incompleteness. However, inference efficiencies are physically limited by the bottlenecks of conventional computing platforms. In this paper, an emerging Bayesian inference system is proposed by exploiting spintronics based stochastic computing. A stochastic bitstream generator is realized as the kernel components by leveraging the inherent randomness of spintronics devices. The proposed system is evaluated by typical applications of data fusion and Bayesian belief networks. Simulation results indicate that the proposed approach could achieve significant improvement on inference efficiencies in terms of power consumption and inference speed. Xiaotao Jia, Jianlei Yang 0001, Zhaohao Wang, Yiran Chen 0001, Hai Li 0001, Weisheng Zhao 0001 |
ASP-DAC | 2 |
| 2018 | A Scalable Pipelined Dataflow Accelerator for Object Region Proposals on FPGA PlatformabstractRegion proposal is critical for object detection while it usually poses a bottleneck in improving the computation efficiency on traditional control-flow architectures. We have observed region proposal tasks are potentially suitable for performing pipelined parallelism by exploiting dataflow driven acceleration. In this paper, a scalable pipelined dataflow accelerator is proposed for efficient region proposals on FPGA platform. The accelerator processes image data by a streaming manner with three sequential stages: resizing, kernel computing and sorting. First, Ping-Pong cache strategy is adopted for rotation loading in resize module to guarantee continuous output streaming. Then, a multiple pipelines architecture with tiered memory is utilized in kernel computing module to complete the main computation tasks. Finally, a bubble-pushing heap sort method is exploited in sorting module to find the top-k largest candidates efficiently. Our design is implemented with high level synthesis on FPGA platforms, and experimental results on VOC2007 datasets show that it could achieve about 3.67X speedups than traditional desktop CPU platform and >250X energy efficiency improvement than embedded ARM platform. Wenzhi Fu, Jianlei Yang 0001, Pengcheng Dai, Yiran Chen 0001, Weisheng Zhao 0001 |
FPT | 2 |
| 2018 | Power Supply Noise Aware Task Scheduling on Homogeneous 3D MPSoCs Considering the Thermal Constraint
Yinglin Zhao, Jianlei Yang 0001, Weisheng Zhao 0001, Aida Todri, Yuanqing Cheng |
J. Comput. Sci. Technol. | 2 |
| 2017 | Thermosiphon: A thermal aware NUCA architecture for write energy reduction of the STT-MRAM based LLCsabstractAs the speed gap of the modern processor and the off-chip main memory enlarges, on-chip cache capacity increases to sustain the performance scaling. As a result, the cache power occupies a large portion of the total power budget. STT-MRAM (Spin Transfer Torque Magnetic Memory) is proposed as a promising solution for the low power cache design due to its high integration density and ultra-low leakage. Nevertheless, the high write power and latency of STT-MRAM become new barriers for the commercialization of this emerging technology. In this paper, we investigate the thermal effect on the access performance of STT-MRAM and observe that the temperature can affect the write delay and energy significantly. Then, we explore the NUCA (Non-Uniform Cache Access) design of the CMPs (Chip-Multi-Processors)with STT-MRAM based LLC (Last Level Cache). A thermal aware data migration policy, called “Thermosiphon”, which takes advantage of the thermal property of STT-MRAM, is proposed to reduce the LLC write energy. This policy splits the LLC into different regions based on the thermal distribution and adaptively migrate write intensive data considering the temperature gradient among different thermal regions. Compared to the conventional NUCA design, our proposed design can save 22.5% write energy with negligible hardware overhead. Bi Wu 0002, Yuanqing Cheng, Pengcheng Dai, Jianlei Yang 0001, Youguang Zhang, Dijun Liu, Ying Wang 0001, Weisheng Zhao 0001 |
ICCAD | 4 |
| 2017 | Power Profile Equalizer: A Lightweight Countermeasure against Side-Channel AttackabstractPower attack is an important side-channel attack (SCA) method based on the correlation between measured power profile and internal switching activities. Various techniques have been proposed to prevent power attack. It has been noted that the on-chip power grid (PG) has a vital effect on the effectiveness of power attack by inducing a noise in the power profile. However, there is a lack of study on this intrinsic effect of PG. In this paper, we explore the methods of exploiting the PG-induced noise to counter with power attack. We note that the PG-induced noise strongly depends on the PG impedance and it can be regulated by adjusting the PG capacitor to control the power profile to fixed values, which contributes to reducing the power leakage. Further, we propose a novel adjustment technique for PG capacitor, i.e. power profile equalizer (PPE), as a lightweight (low-overhead) countermeasure against power attack. PPE exploits the regulated noise to equalize the power profile without violating the layout and supply noise constraints. To reduce the overheads, random walk is adopted to utilize the utmost on-chip resources. Moreover, PPE is implemented by optimizing PG which is an essential IC component rather than producing new circuits. As a result, PPE incurs low overheads. Experimental results show that PPE is able to improve the measurements to disclose (MTD) by 1800x while the area and power increase respectively by 0.12% and 0.91%. Chenguang Wang 0003, Yici Cai, Qiang Zhou 0001, Jianlei Yang 0001 |
ICCD | 5 |
| 2016 | Secure and Low-Overhead Circuit Obfuscation Technique with MultiplexersabstractCircuit obfuscation techniques have been proposed to conceal circuit's functionality in order to thwart reverse engineering (RE) attacks to integrated circuits (IC). We believe that a good obfuscation method should have low design complexity and low performance overhead, yet, causing high RE attack complexity. However, existing obfuscation techniques do not meet all these requirements. In this paper, we propose a polynomial obfuscation scheme which leverages special designed multiplexers (MUXs) to replace judiciously selected logic gates. Candidate to-be-obfuscated logic gates are selected based on a novel gate classification method which utilizes IC topological structure information. We show that this scheme is resilient to all the known attacks, hence it is secure. Experiments are conducted on ISCAS 85/89 and MCNC benchmark suites to evaluate the performance overhead due to obfuscation. Xiaotao Jia, Qiang Zhou 0001, Yici Cai, Jianlei Yang 0001, Gang Qu 0001 |
ACM Great Lakes Symposium on VLSI | 5 |
| 2016 | ODESY: a novel 3T-3MTJ cell design with optimized area DEnsity, scalability and latencYabstractThe STT-RAM (Spin-Transfer Torque Magnetic RAM) technology is a promising candidate for cache memory because of its high density, low standy-power, and non-volatility. As technology scales, especially under 40nm technology node, the read disturbance becomes severe since the read current approaches closely to the switching current. In addition, the read latency and access performance degrade significantly as well. The conventional 1T-1MTJ and 2T-2MTJ cell designs cannot address these challenges efficiently. In this paper, we propose a novel 3T-3MTJ cell structure using the advanced perpendicular MTJ technology. This memory cell has higher storage density and better performance, and is particularly suitable for the deeply scaled technology node. A two-stage sensing scheme is also proposed to facilitate the read operation of the 3T-3MTJ cell design. Circuit-level and architecture-level simulations show that the proposed 3T-3MTJ cell structure can achieve a better tradeoff between storage density, access performance, energy consumption, and reliability compared to the prior 1T-1MTJ and 2T-2MTJ cell structures. Linuo Xue, Yuanqing Cheng, Jianlei Yang 0001, Peiyuan Wang, Yuan Xie 0001 |
ICCAD | 3 |
| 2016 | Radiation-Induced Soft Error Analysis of STT-MRAM: A Device to Circuit ApproachabstractSpin-transfer torque magnetic random access memory (STT-MRAM) is a promising emerging memory technology due to its various advantageous features such as scalability, nonvolatility, density, endurance, and fast speed. However, the reliability of STT-MRAM is severely impacted by environmental disturbances because radiation strike on the access transistor could introduce potential write and read failures for 1T1MTJ cells. In this paper, a comprehensive approach is proposed to evaluate the radiation-induced soft errors spanning from device modeling to circuit level analysis. The simulation based on 3-D metal-oxide-semiconductor transistor modeling is first performed to capture the radiation-induced transient current pulse. Then a compact switching model of magnetic tunneling junction (MTJ) is developed to analyze the various mechanisms of STT-MRAM write failures. The probability of failure of 1T1MTJ is characterized and built as look-up-tables. This approach enables designers to consider the effect of different factors such as radiation strength, write current magnitude and duration time on soft error rate of STT-MRAM memory arrays. Meanwhile, comprehensive write and sense circuits are evaluated for bit error rate analysis under random radiation effects and transistors process variation, which is critical for performance optimization of practical STT-MRAM read and sense circuits. Jianlei Yang 0001, Peiyuan Wang, Yaojun Zhang, Yuanqing Cheng, Weisheng Zhao 0001, Yiran Chen 0001, Hai Li 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2016 | Temperature Impact Analysis and Access Reliability Enhancement for 1T1MTJ STT-RAMabstractSpin-transfer torque magnetic random access memory (STT-RAM) is a promising and emerging technology due to its many advantageous features such as scalability, nonvolatility, density, endurance, and fast access speed. However, the operation of STT-RAM is severely affected by environmental factors such as process variations and temperature. As the temperature rockets up in modern computing systems, it is highly desirable to understand thermal impact on STT-RAM operations and reliability. In this paper, a thermal-aware MTJ model, calibrated and validated by experimental measurements, is proposed as the basis for thoroughly thermal aware analysis of a 1T1MTJ STT-RAM cell structure. Using this model, we investigate temperature effect on memory cell access behavior in terms of access latency, energy, and reliability on a 45-nm technology node. Thermal impact on a more advanced 11-nm technology node is also evaluated in the paper. Additionally, we propose a thermal-aware design for STT-RAM sensing circuit using a body-biasing technique, which can enlarge read margin dramatically to enhance read reliability under temperature variations. Moreover, our proposed technique can suppress read disturbance effectively as well. Experimental results show that our proposed sensing circuit can enlarge read margin by 2.47× when reading “0” and 3.15× when reading “1,” and reduce read disturbance error rate by 55.6% on average. Bi Wu 0002, Yuanqing Cheng, Jianlei Yang 0001, Aida Todri, Weisheng Zhao 0001 |
IEEE Trans. Reliab. | 3 |
| 2016 | Alleviating Through-Silicon-Via Electromigration for 3-D Integrated Circuits Taking Advantage of Self-Healing EffectabstractThree-dimensional integration is considered to be a promising technology to tackle the global interconnect scaling problem for terascale integrated circuits (ICs). Three-dimensional ICs typically employ through-silicon-vias (TSVs) to vertically connect planar circuits. Due to its immature fabrication process, several defects, such as void, misalignment, and dust contamination, may be introduced. These defects can significantly increase current densities within TSVs and cause severe electromigration (EM) effects, which can degrade the reliability of 3-D ICs considerably. In this paper, we propose an effective framework to mitigate EM effect of the defective TSV. At first, we analyze various possible TSV defects and their impacts on EM reliability. Based on the observation that EM can be significantly alleviated by self-healing effect, we design an EM mitigation module to protect defective TSVs from EM. To guarantee EM mitigation efficiency, we propose two defective TSV protection schemes, i.e., neighbor sharing and global sharing. Experimental results show that the global-sharing scheme performs the best and can improve the EM mean time to failure by more than 70× on average with only 0.7% area overhead and less than 0.5% performance degradation compared with naked design without any EM protection. Yuanqing Cheng, Aida Todri, Jianlei Yang 0001, Weisheng Zhao 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2015 | Early stage real-time SoC power estimation using RTL instrumentationabstractEarly stage power estimation is critical for SoC architecture exploration and validation in modern VLSI design, but real-time, long time interval and accurate estimation is still challenging for system-level estimation and software/hardware tuning. This work proposes a model abstraction approach for real-time power estimation in the manner of machine learning. The singular value decomposition (SVD) technique is exploited to abstract the principle components of relationship between register toggling profile and accurate power waveform. The abstracted power model is automatically instrumented to RTL implementation and synthesized into FPGA platform for real-time power estimation by instrumenting the register toggling profile. The prototype implementation on three IP cores predicts the cycle-by-cycle power dissipation within 5% accuracy loss compared with a commercial power estimation tool. Jianlei Yang 0001, Liwei Ma, Yici Cai, Tin-Fook Ngai |
ASP-DAC | 1 |
| 2015 | A High-Speed Robust NVM-TCAM Design Using Body Bias FeedbackabstractAs manufacture process scales down rapidly, the design of ternary content-addressable memory (TCAM) requiring high storage density, fast access speed and low power consumption becomes very challenging. In recent years, many novel TCAM designs have been inspired by the research on emerging nonvolatile memory technologies, such as magnetic tunneling junction (MTJ), phase change memory (PCM), and memristor. These designs store a data as the resistive variable of a nonvolatile device, which usually results in limited sensing margin and therefore constrains the searching speed of TCAM architecture severely. To further enhance the performance and robustness of TCAMs, we proposed two novel cell designs that utilize MTJs as data storage units - the symmetrical dual-N structure and the asymmetrical P-N scheme. In both designs, a body bias feedback circuit is integrated to enlarge the sensing margins. Compared with an existing MTJ-based TCAM structure, the tolerance in gate voltage variation of the symmetrical dua-N (asymmetrical P-N) scheme can significantly improve 59.5% (21.2%). The latency and the dynamic energy consumption in one searching operation at the word length of 256 bits are merely 590.35ps (97.89ps) and 65.05fJ/bit (36.85fJ/bit), not even mentioning that the use of nonvolatile MTJ devices avoids unnecessary leakage power consumption. Bonan Yan, Yaojun Zhang, Jianlei Yang 0001, Hai Li 0001, Weisheng Zhao 0001, Pierre Chor-Fung Chia |
ACM Great Lakes Symposium on VLSI | 4 |
| 2015 | An overview on memristor crossabr based neuromorphic circuit and architectureabstractAs technology advances, artificial intelligence becomes pervasive in society and ubiquitous in our lives, which stimulates the desire for embedded-everywhere and human-centric intelligent computation paradigm. However, conventional instruction-based computer architecture was designed for algorithmic and exact calculations. It is not suitable for handling the applications of machine learning and neural networks that usually involve a large sets of noisy and incomplete natural data. Instead, neuromorphic systems inspired by the working mechanism of human brains create promising potential. Neuromorphic systems possess a massively parallel architecture with closely coupled memory and computing. Moreover, through the sparse utilizations of hardware resources in time and space, extremely high power efficiency can be achieved. In recent years, the use of memristor technology in neuromorphic systems has attracted growing attention for its distinctive properties, such as nonvolatility, reconfigurability, and analog processing capability. In this paper, we summarize the research efforts in the development of memristor crossbar based neuromorphic design from the perspectives of device modeling, circuit, architecture, and design automation. Yandan Wang, Bonan Yan, Chaofei Yang, Jianlei Yang 0001, Hai Li 0001 |
VLSI-SoC | 6 |
| 2015 | A Selected Inversion Approach for Locality Driven Vectorless Power Grid VerificationabstractVectorless power grid verification is a practical approach for early stage safety check without input current patterns. The power grid is usually formulated as a linear system and requires intensive matrix inversion and numerous linear programming (LP), which is extremely time-consuming for large-scale power grid verification. In this paper, the power grid is represented in the manner of domain-decomposition approach, and we propose a selected inversion technique to reduce the computation cost of matrix inversion for vectorless verification. The locality existence among power grids is exploited to decide which blocks of matrix inversion should be computed while remaining blocks are not necessary. The vectorless verification could be purposefully performed by this manner of selected inversion, while previous direct approaches are required to perform full matrix inversion and then discard small entries to reduce the complexity of LP. Meanwhile, constraint locality is proposed to improve the verification accuracy. In addition, a concept of quasi-Poisson block is introduced to exploit grid locality among realistic power grids and a scheme of pad-aware partitioning is proposed to enable the selected inversion approach available for practical use. Experimental results show that the proposed approach could achieve significant speedups compared with previous approaches while still guaranteeing the quality of solution accuracy. Jianlei Yang 0001, Yici Cai, Qiang Zhou 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2014 | Fast vectorless power grid verification using maximum voltage drop location estimationabstractPower grid integrity verification is critical for reliable chip design. Vectorless power grid verification provides a promising approach to evaluate the worst-case voltage fluctuations without the detailed information of circuit activities. Vectorless verification is usually required to solve numerous linear programming problems to obtain the worst-case voltage fluctuation throughout the grid, which is extremely time-consuming for large-scale verification. In this paper, a maximum voltage drop location estimation approach is proposed for efficient vectorless verification. The power grid nodes are grouped into disjoint subsets, and an estimation strategy is utilized to roughly locate the nodes which have the worst-case voltage drop in each group. Consequently, the verification problem size can be significantly reduced compared with accurate verification. Experimental results show that the proposed approach can achieve remarkable speedups with acceptable accuracy loss. Yici Cai, Jianlei Yang 0001 |
ASP-DAC | 3 |
| 2014 | Power supply noise aware evaluation framework for side channel attacks and countermeasuresabstractSide Channel Attack (SCA) aims to extract the secret information from cryptography chips by analyzing the leakage of physical parameters. Power analysis based SCA is a popular approach to obtain secret keys by monitoring the power consumption of cryptography chips. However, most SCA evaluation methods are performed on FPGA platforms while many parasitic physical effects cannot be revealed before the cryptography chips are taped out. Roughly ignoring these effects will significantly increase the attack difficulties due to the corresponding measurement noise. Power supply noise has been observed to be critical for power analysis based SCA. This paper demonstrates a power supply noise aware evaluation framework for practical side channel attack from cryptography system design to physical design. On-chip power delivery network is implemented among physical design stage. Consequently the supply noise of power network can be explored according to the post-layout implementation. Additionally, the countermeasures of cryptography chips could be enhanced by on-chip decapacitors placement due to its influences on the characteristics of power delivery network. Jianlei Yang 0001, Chenguang Wang 0003, Yici Cai, Qiang Zhou 0001 |
FPT | 1 |
| 2014 | Friendly Fast Poisson Solver Preconditioning Technique for Power Grid AnalysisabstractRobust and efficient algorithms for power grid analysis are crucial for both VLSI design and optimization. Due to the increasing size of power grids, IR drop analysis has become more computationally challenging both in runtime and memory consumption. This paper presents a Fast Poisson Solver (FPS) preconditioned method for unstructured power grids with unideal boundary conditions. Unstructured power grids are transformed to structured grids, which can be modeled as Poisson blocks by analytic formulation. The analytic formulation of transformed structured grids is adopted as an analytic preconditioner for original unstructured grids, in which the analytic preconditioner can be considered as a sparse approximate inverse technique. By combining this analytic preconditioner with robust conjugate gradient method, we demonstrate that this approach is totally robust for extremely large scale power grid simulations. Theoretical proof and experimental results show that iterations of our proposed method will hardly increase with the increasing of grid size as long as the pads density and the distribution range of metal conductance value have been decided. We demonstrate that the run efficiency of our approach is much higher than classical incomplete Cholesky factorization preconditioned conjugate gradient solver and random walk-based hybrid solver. Jianlei Yang 0001, Yici Cai, Qiang Zhou 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2014 | PowerRush: An Efficient Simulator for Static Power Grid AnalysisabstractEfficient power grid analysis is critical for modern very large scale integration design but is computationally challenging in runtime and memory consumption because of the increasing size of power grids. PowerRush is proposed as an efficient IR-drop simulator, which includes an efficient SPICE parser, a robust circuit builder, and a linear solver Algebraic MultiGrid Preconditioned Conjugate Gradient. The proposed AMG-PCG solver is a pure algebraic method, which can provide stable convergence without geometric information. Aggregation-based AMG with K-cycle acceleration is adopted as a preconditioner to improve the scalability of iterative method. In multigrid scheme, double pairwise aggregation technique is applied to matrix graph in coarsening to ensure low setup cost and memory requirement. Furthermore, K-cycle multigrid scheme is adopted to provide Krylov subspace acceleration at each level to guarantee enhanced robustness and scalability. The experimental results for large-scale power grids have shown that PowerRush has remarkable scalability both in runtime and memory consumption. DC analysis of power grid with 60-million nodes can be solved by PowerRush for 0.01 $mV$ accuracy within 150 s and 21.99 GB total memory used. Moreover, the proposed AMG-PCG solver can perform much better than widely used direct solver Cholmod and well-developed Hybrid solver both on runtime and memory consumption. Jianlei Yang 0001, Zuowei Li, Yici Cai, Qiang Zhou 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2013 | A multilevel ℌ-matrix-based approximate matrix inversion algorithm for vectorless power grid verificationabstractVectorless power grid verification technique makes it possible to estimate the worst-case voltage fluctuations of the on-chip power delivery network at the early design stage. For most of the existing vectorless verification algorithms, the sub-problem of linear system solution which computes the inverse of the power grid matrix takes up a large part of the computation time and has become a critical bottleneck of the whole algorithm. In this paper, we propose a new algorithm that combines the ℌ-matrix-based technique and the multilevel method to construct a data-sparse approximate inverse of the power grid matrix. Experimental results have shown that the proposed algorithm can obtain an almost linear complexity both in runtime and memory consumption for efficient vectorless power grid verification. Yici Cai, Jianlei Yang 0001 |
ASP-DAC | 3 |
| 2013 | Selected inversion for vectorless power grid verification by exploiting localityabstractVectorless power grid verification is a practical approach for early stage safety check without input current patterns. The power grid is usually formulated as a linear system and requires intensive matrix inversion and numerous linear programming, which is extremely time-consuming for large scale power grid verification. In this paper, the power grid is represented in the manner of domain-decomposition approach, and we propose a selected inversion technique to reduce the computation cost of matrix inversion for vectorless verification. The locality existence among power grids is exploited to decide which blocks of matrix inversion should be computed while remaining blocks are not necessary. The vectorless verification could be purposefully performed by this manner of selected inversion while previous direct approaches are required to perform full matrix inversion and then discard small entries to reduce the complexity of linear programming. Meanwhile, constraint locality is proposed to improve the verification accuracy. Experimental results show that the proposed approach could achieve significant speedups compared to previous approaches while still guaranteeing the quality of solution accuracy. Jianlei Yang 0001, Yici Cai, Qiang Zhou 0001 |
ICCD | 1 |
| 2012 | PowerRush : Efficient transient simulation for power grid analysisabstractTransient analysis is the most practical and effective approach for power grid validation, but which is very challengeable for large scale VLSI chips because it is really time consuming and requires large memory resources. In this paper we proposed a parallel transient simulation approach for efficient power grid analysis. Firstly we adopt symmetric formulation for NA equation of RLC power grid to reduce memory usage. Meanwhile, fast Cholesky factorization solver can be used to improve simulation efficiency. Secondly, we perform partition-based parallel transient simulation for naturally independent subnets without accuracy lost. Thirdly, we propose a composite simulation flow for efficient and practical transient analysis for industrial power grid. Finally, several industrial power grid benchmarks are evaluated on our approaches for high accurate transient simulation with extremely low memory consumption. Jianlei Yang 0001, Zuowei Li, Yici Cai, Qiang Zhou 0001 |
ICCAD | 1 |
| 2011 | Obstacle-avoiding and slew-constrained buffered clock tree synthesis for skew optimizationabstractBuered clock tree synthesis (CTS) is increasingly critical as VLSI technology continually scales down. Many researches have been done on this topic due to its key role in CTS, but current approaches either lack the obstacle-avoiding functionality or lead to large clock latency and/or skew. This paper presents a new obstacle-avoiding CTS approach with separate clock tree construction and buer insertion stages based on an integral view to explore the global optimization space. Aiming at skew optimization under constraints of slew and obstacles, our CTS approach features the clock tree construction stage with the obstacle-aware topology generation algorithm called OBB, balanced insertion of candidate buer positions, and a fast heuristic buer insertion algorithm. Experimental results show the eectiveness of our CTS approach with significantly improved skew and latency than [6] by 46% and 63% on average, and 15.3% reduction in skew than [5]. Our OBB heuristic obtains 36% improvement in skew than the classic balanced bipartition algorithm (BB) in [10]. Feifei Niu, Qiang Zhou 0001, Hailong Yao 0002, Yici Cai, Jianlei Yang 0001, Cliff C. N. Sze |
ACM Great Lakes Symposium on VLSI | 5 |
| 2011 | Fast poisson solver preconditioned method for robust power grid analysisabstractRobust and efficient algorithms for power grid analysis are crucial for both VLSI design and optimization. Due to the increasing size of power grids IR drop analysis has become more computationally challenging both in runtime and memory consumption. This work presents a fast Poisson solver preconditioned method for unstructured power grid with unideal boundary conditions. In fact, by taking the advantage of analytical formulation of power grids this analytical preconditioner can be considered as sparse approximate inverse technique. By combining this analytical preconditioner with robust conjugate gradient method, we demonstrate that this approach is totally robust for extremely large scale power grid simulations. Experimental results have shown that iterations of our proposed method will hardly increase with grid size increasing once the pads density and the range of metal resistances value distribution have been decided. We demonstrated that this approach solves an unstructured power grid with 2.56M nodes in only 1/3 iterations of classical ICCG solver, and achieves almost 20X speedups over the classical ICCG solver on runtime. Jianlei Yang 0001, Yici Cai, Qiang Zhou 0001 |
ICCAD | 1 |
| 2011 | PowerRush: A linear simulator for power gridabstractAs the increasing size of power grids, IR drop analysis has become more computationally challenging both in runtime and memory consumption. In this paper, we propose a linear complexity simulator named PowerRush, which consists of an efficient SPICE Parser, a robust circuit Builder and a linear solver. The proposed solver is a pure algebraic method which can provide an optimal convergence without geometric information. It is implemented by Algebraic Multigrid Preconditioned Conjugate Gradient method, in which an aggregation based algebraic multigrid with K-Cycle acceleration is adopted as a preconditioner to improve the robustness of conjugate gradient iterative method. In multigrid scheme, double pairwise aggregation technique is applied to the matrix graph in coarsening procedure to ensure low setup cost and memory requirement. Further, a K-Cycle multigrid scheme is adopted to provide Krylov subspace acceleration at each level to guarantee optimal or near optimal convergence. Experimental results on real power grids have shown that PowerRush has a linear complexity in runtime cost and memory consumption. The DC analysis of a 60 Million nodes power grid can be solved by PowerRush for 0.01mV accuracy in 170 seconds with 21.89GB memory used. Jianlei Yang 0001, Zuowei Li, Yici Cai, Qiang Zhou 0001 |
ICCAD | 1 |