Dashan Shang

dblp:186/8460 · DBLP profile ↗
← Back
11ranked-venue papers
0as first author
11since 2021 · last 2026
0000-0003-3573-8390ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 8 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021
YearPublicationVenuePosition
2026 An area/energy-efficient RRAM computing-in-memory macro with fully-charge-domain multi-bit computation
Shengzhe Yan, Zhuoyu Dai, Zeyu Guo 0002, Zhaori Cong, Zhihang Qian, Xiangqu Fu, Chunmeng Dou, Dashan Shang, Jinshan Yue
Sci. China Inf. Sci.11
2026 AgeBalance: Low-Cost Lifetime Extension for SRAM-Based PIM Accelerators
abstract
Although processing-in-memory (PIM) techniques have widely been used for deep neural networks (DNNs) acceleration, the inference performance of aged PIM-based accelerators remains to be investigated. This paper makes the first attempt to study Hot Carrier Injection (HCI) and Negative Bias Temperature Instability (NBTI) aging impacts on SRAM-based DNN accelerators, which provides a novel and unified framework, termedAgeBalancefor aging detection, analysis and mitigation. First, we discuss a convenient aging detection scheme. Then, we benchmark the inference accuracy drops of DNNs running on aged SRAM-based PIM accelerators. Finally, we propose a low-cost anti-aging training method without incurring additional hardware overhead on SRAM-based DNN accelerators. Extensive experimental results on MNIST, CIFAR10 and AG News datasets show that aging can cause the inference accuracy of shallow or deep DNNs to drop to about 10%, close to random guessing. The aging mitigation scheme proposed in this paper can largely restore the accuracy to the original. Moreover, the SRAM write overhead of our method is much reduced thanks to a score-based training approach, leading to a reduction of 5× to 10× writing energy compared to the traditional training method.
Ning Lin, Shaocong Wang 0001, Yangu He, Songqi Wang, Kwunhang Wong, Rongliang Fu, Wenxing Li, Tsung-Yi Ho, Dashan Shang, Xiaojuan Qi 0001, Xiaoming Chen 0003
IEEE Trans. Computers10
2025 Efficient Edge Vision Transformer Accelerator with Decoupled Chunk Attention and Hybrid Computing-In-Memory
abstract
Vision Transformers (ViTs) are new foundation models for vision applications. Edge-deploying ViTs to realize energy-saving, low-latency, and high-performance dense predictions have wide applications, such as autonomous driving and surveillance image analysis. However, the quadratic complexity of the self-attention mechanism renders ViTs slow and resource-intensive, particularly for pixel-level dense predictions that involve long contexts. Additionally, the pyramid-like architecture of modern ViT variants leads to an unbalanced workload, further reducing hardware utilization and decreasing the throughput of conventional edge devices. To this end, we propose an algorithm-hardware co-optimized edge ViT accelerator tailored for efficient dense predictions. At the algorithm level, we propose a decoupled chunk attention (DCA) mechanism implemented in a pipelined manner to reduce off-chip memory access, thereby enabling efficient dense predictions within limited on-chip memory. At the architecture level, we introduce a hybrid architecture that combines SRAM-based computing-in-memory (CIM) and nonvolatile RRAM storage to eliminate extensive off-chip memory access, with a fusion scheduling to balance workloads and minimize intermediate on-chip memory access. At the circuit level, a bit/element two-way-reconfigurable CIM macro is proposed to improve hardware utilization across pyramidal ViT blocks with varied matrix sizes. The experimental results on object detection, semantic segmentation, and depth estimation tasks demonstrate that our design can efficiently process patch lengths up to 16384 with a speedup of 18.5×-217.1×, a reduction in memory accesses of 1.7×-7.4×, and an improvement in energy efficiency of 1.8×, under less than 1% performance degradation.
Yi Li 0049, Zijian Ye, Xiangqu Fu, Songqi Wang, Shucheng Du, Ning Lin, Dashan Shang, Jinshan Yue, Xiaojuan Qi 0001, Feng Zhang 0014
DAC7
2025 Re4PUF: A Reliable, Reconfigurable ReRAM-based PUF Resilient to DNN and Side Channel Attacks
abstract
Resistive random-access memory (ReRAM) based Physical Unclonable Functions (PUFs) have emerged as an attractive hardware security primitive due to their low energy consumption and compact footprint. However, the reliability of existing ReRAM-based PUFs is challenged by read noise and temperature variations, as well as their resistance to Deep Neural Network (DNN) modeling attacks and Side Channel Attacks (SCAs). In this paper, we propose a novel 3T2R ReRAM-based reconfigurable PUF to address these challenges. By adopting the digital 3T2R voltage division cell design, we improve its reliability against ReRAM read noise and temperature variations, while the adjustable analog supply voltage of inverters enables quick, low-cost reconfigurability without reprogramming ReRAMs, effectively mitigating DNN modeling and SCA vulnerabilities. Our $\mathrm{Re}^{4}$ PUF chip has been experimentally validated, achieving a low Bit Error Rate (BER) of $1 \%$ at $85^{\circ} \mathrm{C}$, a 7.59 -fold reduction compared to existing ReRAM-based PUFs. It also demonstrates robust resistance to both DNN modeling attacks (MLP and Transformer) and SCAs, with success rates of approximately 50% and less than $70 \%$, respectively.
Ning Lin, Yangu He, Songqi Wang, Hegan Chen, Kwunhang Wong, Chuxin Li, Jichang Yang, Yongkang Han, Xiaoxin Xu, Dashan Shang
DAC16
2025 Guarder: A Stable and Lightweight Reconfigurable RRAM-based PIM Accelerator for DNN IP Protection
abstract
Deploying deep neural networks (DNNs) on conventional digital edge devices faces significant challenges due to high energy consumption. A promising solution is the processing-inmemory (PIM) architecture with resistive random-access memory (RRAM), but RRAM-based systems suffer from imprecise weights due to programming stochasticity and cannot effectively utilize conventional weight encryption/decryption intellectual property (IP) protection schemes. To address these issues, we propose a novel software-hardware co-design Guarder. On the hardware side, we introduce 3T2R cells to achieve reliable multiply-accumulate (MAC) operations and use reconfigurable inverter operating voltages to encode keys for encrypting DNNs on RRAM. On the software side, we implement a contrastive training method that ensures high model accuracy on authorized chips while degrading performance on unauthorized ones. This approach protects DNN IP with minimal hardware overhead while significantly mitigating the effects of RRAM programming stochasticity. Extensive experiments on tasks such as image classification (using MLP, ResNet, and ViT), segmentation (using SegFormer), and image generation (using DiT) validate the effectiveness of our method. The proposed contrastive training ensures negligible performance degradation on authorized chips, while performance on unauthorized chips drops to random guessing or generation. Compared to traditional RRAM accelerators, the 3T2R-based accelerator achieves a $1.41 \times$ reduction in area overhead and a $2.28 \times$ reduction in energy consumption.
Ning Lin, Yi Li 0049, Jiankun Li, Jichang Yang, Yangu He, Yukui Luo, Dashan Shang, Xiaoming Chen 0003, Xiaojuan Qi 0001
DAC7
2025 When Pipelined In-Memory Accelerators Meet Spiking Direct Feedback Alignment: A Co-Design for Neuromorphic Edge Computing
abstract
Spiking Neural Networks (SNNs) are increasingly favored for deployment on resource-constrained edge devices due to their energy-efficient and event-driven processing capabilities. However, training SNNs remains challenging because of the computational intensity of traditional backpropagation algorithms adapted for spike-based systems. In this paper, we propose a novel software-hardware co-design that introduces a hardware-friendly training algorithm, Spiking Direct Feedback Alignment (SDFA) and implement it on a Resistive Random Access Memory (RRAM)-based In-Memory Computing (IMC) architecture, referred to as PipeSDFA, to accelerate SNN training. Software-wise, the computational complexity of SNN training is reduced by the SDFA through the elimination of sequential error propagation. Hardware-wise, a three-level pipelined dataflow is designed based on IMC architecture to parallelize the training process. Experimental results demonstrate that the PipeSDFA training accelerator incurs less than 2% accuracy loss on five datasets compared to baselines, while achieving 1.1× ~10.5×and 1.37×~2.1× reductions in training time and energy consumption, respectively compared to PipeLayer.
Haoxiong Ren, Yangu He, Kwunhang Wong, Rui Bao, Ning Lin, Dashan Shang
ICCAD7
2025 A High-Density Energy-Efficient CNM Macro Using Hybrid RRAM and SRAM for Memory-Bound Applications
abstract
The big data era has facilitated various memory-centric algorithms, such as the Transformer decoder, neural network, stochastic computing (SC), and genetic sequence matching, which impose high demands on memory capacity, bandwidth, and access power consumption. The emerging nonvolatile memory devices and compute-near-memory (CNM) architecture offer a promising solution for memory-bound tasks. This work proposes a hybrid resistive random access memory (RRAM) and static random access memory (SRAM) CNM architecture. The main contributions include: 1) proposing an energy-efficient and high-density CNM architecture based on the hybrid integration of RRAM and SRAM arrays; 2) designing low-power CNM circuits using the logic gates and dynamic-logic adder with configurable datapath; and 3) proposing a broadcast mechanism with output-stationary workflow to reduce memory access. The proposed RRAM-SRAM CNM architecture and dataflow tailored for four distinct applications are evaluated at a 28-nm technology, achieving 4.62-TOPS$/$W energy efficiency and 1.20-Mb$/$mm2memory density, which shows$11.35\times $–$25.81\times $and$1.44\times $–$4.92\times $improvement compared to previous works, respectively.
Shengzhe Yan, Xiangqu Fu, Zhihang Qian, Zhi Li 0062, Zeyu Guo 0002, Zhuoyu Dai, Zhaori Cong, Chunmeng Dou, Feng Zhang 0014, Jinshan Yue, Dashan Shang
IEEE Trans. Very Large Scale Integr. Syst.12
2024 Older and Wiser: The Marriage of Device Aging and Intellectual Property Protection of DNNs
abstract
Deep neural networks (DNNs), such as the widely-used GPT-3 with billions of parameters, are often kept secret due to high training costs and privacy concerns surrounding the data used to train them. Previous approaches to securing DNNs typically require expensive circuit redesign, resulting in additional overheads such as increased area, energy consumption, and latency. To address these issues, we propose a novel hardware-software co-design approach for DNN intellectual property (IP) protection that capitalizes on the inherent aging characteristics of circuits and a novel differential orientation fine-tuning (DOFT) to ensure effective protection.
Ning Lin, Shaocong Wang 0001, Yue Zhang 0011, Yangu He, Kwunhang Wong, Arindam Basu, Dashan Shang, Xiaoming Chen 0003
DAC7
2024 LSMR: Synergy Randomness in Liquid State Machine and RRAM-based Analog-digital Accelerator
abstract
Bio-inspired event sensors are gaining popularity at the edge, such as in robots and wearable electronics. This trend necessitates learning vast amounts of sensory data on the edge, often in few-shot or even zero-shot scenarios, posing challenges in both software and hardware. This paper presents a novel software-hardware co-design to address these issues. Software-wise, we develop an SNN-ANN model, where the SNN encoder is a liquid state machine (LSM) that naturally processes events and significantly reduces learning complexity at the edge due to fixed random weights. The lightweight trainable ANN projection heads are optimized through contrastive learning, enabling zero-shot learning of multimodal events. Hardware-wise, we propose a hybrid analog (RRAM)-digital (CMOS) accelerator - LSMR. The analog in-memory computing core physically implements the LSM by leveraging RRAM stochasticity to generate fixed random weights. The digital core utilizes innovative reconfigurable systolic arrays to accelerate the contrastive learning of ANN projection heads. Extensive experimental outcomes from six neuromorphic datasets, encompassing visual, tactile, and auditory modalities, demonstrate that LSMR considerably improves energy efficiency by a range of 1.65× to 23.70×, in comparison to state-of-the-art edge devices. Simultaneously, it reduces training complexity by a range of 152.83× to 20,587.77× across various edge learning tasks.
Ning Lin, Songqi Wang, Xinyuan Zhang 0008, Shaocong Wang 0001, Yangu He, Woyu Zhang, Bo Wang 0153, Jiankun Li, Mingzi Li, Binbin Cui, Yi Li 0049, Jia Chen 0032, Chunwei Xia, Xiaoming Chen 0003, Dashan Shang
ICCAD16
2024 CMN: a co-designed neural architecture search for efficient computing-in-memory-based mixture-of-experts
abstract
Abstract Artificial intelligence (AI) has experienced substantial advancements recently, notably with the advent of large-scale language models (LLMs) employing mixture-of-experts (MoE) techniques, exhibiting human-like cognitive skills. As a promising hardware solution for edge MoE implementations, the computing-in-memory (CIM) architecture collocates memory and computing within a single device, significantly reducing the data movement and the associated energy consumption. However, due to diverse edge application scenarios and constraints, determining the optimal network structures for MoE, such as the expert’s location, quantity, and dimension on CIM systems remains elusive. To this end, we introduce a software-hardware co-designed neural architecture search (NAS) framework, C IM-based M oE N AS (CMN), focusing on identifying a high-performing MoE structure under specific hardware constraints. The results of the NYUD-v2 dataset segmentation on the RRAM (SRAM) CIM system reveal that CMN can discover optimized MoE configurations under energy, latency, and performance constraints, achieving 29.67 × ( 43.10 ×) energy savings, 175.44 ×( 109.89 ×) speedup, and 12.24 × smaller model size compared to the baseline MoE-enabled Visual Transformer, respectively. This co-design opens up an avenue toward high-performance MoE deployments in edge CIM systems.
Shihao Han, Sishuo Liu, Shucheng Du, Mingzi Li, Zijian Ye, Xiaoxin Xu, Dashan Shang
Sci. China Inf. Sci.9
2024 Erratum to: CMN: a co-designed neural architecture search for efficient computing-in-memory-based mixture-of-experts
Shihao Han, Sishuo Liu, Shucheng Du, Mingzi Li, Zijian Ye, Xiaoxin Xu, Dashan Shang
Sci. China Inf. Sci.9