EDBT 2026 Demo / reviewers in the wild / expert
Dunshan Yu
dblp:92/6680
· DBLP profile ↗
21ranked-venue papers
0as first author
8since 2021 · last 2026
—ORCID · unresolved
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 12 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MVP: A Mobile 3D-Stacked VLM Accelerator for Efficient Video Understanding by Leveraging Dynamic Sparse Attention PatternsabstractThe emergence of Vision-Language Models (VLMs) has enabled multimodal reasoning, e.g. video understanding, yet their extension to long-context inference remains bottlenecked by the “token explosion”. This surge in sequence length leads to prohibitive attention computation overhead and memory-bound KV cache access. While 3D-stacked logic-to-DRAM architectures offer high-bandwidth Processing Near Memory (PNM) capability, their distributed memory banks face severe workload imbalance due to the unique spatiotemporal sparsity patterns in video understanding tasks. In this paper, we introduce MVP, a 3D-stacked VLM accelerator featuring context-aware sparse attention (CASA) and online workload-aware hybrid parallelism scheduling. By leveraging dynamic sparse attention patterns, our design prunes redundant attention computation FLOPS and adaptively balances computation across hybrid bonding (HB) based many-core NoC architecture. Experimental results on VQA tasks demonstrate that our architecture achieves a 5.97× speedup and a 5.42× energy-efficiency improvement over the RTX 4080 GPU deployment, effectively mitigating the bottlenecks of VLM inference for video-understanding on mobile devices. Qianxu Wang, Dunshan Yu |
ISLPED | 3 |
| 2025 | Towards Effective and Sparse Adversarial Attack on Spiking Neural Networks via Breaking Invisible Surrogate GradientsabstractSpiking neural networks (SNNs) have shown their competence in handling spatial-temporal event-based data with low energy consumption. Similar to conventional artificial neural networks (ANNs), SNNs are also vulnerable to gradient-based adversarial attacks, wherein gradients are calculated by spatial-temporal back-propagation (STBP) and surrogate gradients (SGs). However, the SGs may be invisible for an inference-only model as they do not influence the inference results, and current gradient-based attacks are ineffective for binary dynamic images captured by the dynamic vision sensor (DVS). While some approaches addressed the issue of invisible SGs through universal SGs, their SGs lack a correlation with the victim model, resulting in sub-optimal performance. Moreover, the imperceptibility of existing SNN-based binary attacks is still insufficient. In this paper, we introduce an innovative potential-dependent surrogate gradient (PDSG) method to establish a robust connection between the SG and the model, thereby enhancing the adaptability of adversarial attacks across various models with invisible SGs. Additionally, we propose the sparse dynamic attack (SDA) to effectively attack binary dynamic images. Utilizing a generation-reduction paradigm, SDA can fully optimize the sparsity of adversarial perturbations. Experimental results demonstrate that our PDSG and SDA outperform state-of-the-art SNN-based attacks across various models and datasets. Specifically, our PDSG achieves 100% attack success rate on ImageNet, and our SDA obtains 82% attack success rate by modifying only 0.24% of the pixels on CIFAR10DVS. The code is available at https://github.com/ryime/PDSG-SDA. Li Lun, Kunyu Feng, Qinglong Ni, Ling Liang 0003, Yuan Wang 0001, Ying Li 0056, Dunshan Yu, Xiaoxin Cui |
CVPR | 7 |
| 2025 | A Compact Full Pipeline Architecture of SM3 Algorithm with High throughput and High EfficiencyabstractThis paper presents a compact full pipeline architecture of SM3 algorithm with high throughput and high efficiency, which is applied in high-performance security computing and large-scale data stream processing. The algorithm is reconstructed and a full pipeline computing architecture is proposed to achieve high throughput. In order to optimize the timing of algorithm, a compact architecture of compression function is further proposed, in which the critical path is reorganized and the adder delay of the path is optimized using CSA. Registers multiplexing is employed on the architecture of message word expansion for reducing area cost. Additionally, block RAM is employed to replace logic of polling and shifting, further optimizing the timing and reducing the LUT resources cost. The proposed architecture is implemented on a Xilinx Virtex-7 FPGA, and experimental results show that the designed SM3 achieves a peak frequency of 243 MHz, a high throughput of 124.42 Gbps, and a high efficiency of 15.22 Mbps/slice. Ruiting Wang, Xiaofeng Zou, Dunshan Yu |
ISCAS | 6 |
| 2024 | An Energy-Efficient Differential Frame Convolutional Accelerator with on-Chip Fusion Storage Architecture and Pixel-Level Pipeline Data FlowabstractConvolutional neural networks require a huge amount of computation in video applications. For some specific tasks, such as surveillance, differential frame convolution reuses inter-frame data and significantly reduces multiplication and accumulation. However, there are still some challenges in improving energy efficiency of differential frame convolution on chips. Firstly, differential frame convolution brings additional on-chip storage for reusing inter-frame data. Secondly, in post-processing of differential frame convolution, there are more memory accessing and arithmetic logic operations. Therefore, sparse working mode is of vital importance for the post-processing. In response to these challenges, this work proposes an on-chip fusion storage architecture for energy-efficient differential frame convolution and a pixel-level pipeline data flow that supports the sparsity of features. The simulation of our accelerator implemented in 28nm CMOS can achieve energy efficiency by 3.09× compared with other state-of-the-art works Zhenhui Dai, Yi Zhong 0002, Kunyu Feng, Yuan Wang 0001, Dunshan Yu, Xiaoxin Cui |
ISCAS | 9 |
| 2023 | Multimodal Affective States Recognition Based on Multiscale CNNs and Biologically Inspired Decision Fusion ModelabstractThere has been an encouraging progress in the affective states recognition models based on the single-modality signals as electroencephalogram (EEG) signals or peripheral physiological signals in recent years. However, multimodal physiological signals-based affective states recognition methods have not been thoroughly exploited yet. Here we propose Multiscale Convolutional Neural Networks (Multiscale CNNs) and a biologically inspired decision fusion model for multimodal affective states recognition. First, the raw signals are pre-processed with baseline signals. Then, the High Scale CNN and Low Scale CNN in Multiscale CNNs are utilized to predict the probability of affective states output for EEG and each peripheral physiological signal respectively. Finally, the fusion model calculates the reliability of each single-modality signals by the euclidean distance between various class labels and the classification probability from Multiscale CNNs, and the decision is made by the more reliable modality information while other modalities information is retained. We use this model to classify four affective states from the arousal valence plane in the DEAP and AMIGOS dataset. The results show that the fusion model improves the accuracy of affective states recognition significantly compared with the result on single-modality signals, and the recognition accuracy of the fusion result achieve 98.52 and 99.89 percent in the DEAP and AMIGOS dataset respectively. Yuxuan Zhao 0002, Xinyan Cao, Jinlong Lin, Dunshan Yu, Xixin Cao |
IEEE Trans. Affect. Comput. | 4 |
| 2022 | An Event-driven Spiking Neural Network Accelerator with On-chip Sparse WeightabstractSpiking neural networks (SNNs) have widely drew attention of recent research. With brain-spired dynamics and spike-based communication, SNN is supposed to be a more energy-efficient neural network than existing artificial neural network (ANN). To make better use of the temporal sparsity of spikes and spatial sparsity of weights in SNN, this paper presents a sparse SNN accelerator. It adopts a novel self-adaptive spike compressing and decompressing (SASCD) mechanism for different input spike sparsity, as well as on-chip compressed weight storage and processing. We implement the octa-core design on field programmable gate array (FPGA). The results demonstrate a peak performance of 35.84 GSOPs/s, which is equivalent to 358.4 GSOPs/s in dense SNN accelerators for 90% weight sparsity. For the single-layer perceptron model in rate coding implemented on the hardware, SASCD reduces the time step intervals from 2.15 $\mu$ s to 0.55 $\mu$ s. Yisong Kuang, Xiaoxin Cui, Chenglong Zou, Yi Zhong 0002, Zhenhui Dai, Zilin Wang 0001, Kefei Liu 0002, Dunshan Yu, Yuan Wang 0001 |
ISCAS | 8 |
| 2022 | ESSA: Design of a Programmable Efficient Sparse Spiking Neural Network AcceleratorabstractSpiking neural networks (SNNs) have been witnessing the developing trends to reduce the model size and improve the hardware efficiency for area- and energy-based applications, which are processed by model pruning and data compressions. However, it is challenging to exploit the unstructured sparsity of SNNs for the dense neuromorphic processors. In this article, we present an efficient sparse SNN accelerator (ESSA), which leverages both the temporal sparsity of spike events and the spatial sparsity of weights in SNN inference. It provides both the compressed weights for sparse SNNs and the uncompressed weights for compact SNNs. The self-adaptive spike compression is proposed for sparse spike scenarios, leading to the improvement of throughput by$3.2\times $. ESSA executes a flexible fan-in–fan-out tradeoff by using combinable dendrites, which overcomes the fan-in limitation in neuromorphic systems. Furthermore, a low-latency intrachip spike multicast method is adopted to reduce the resource overhead. Implemented on the Xilinx Kintex Ultrascale field-programmable gate array (FPGA), ESSA achieves an equivalent performance of 253.1 GSOP/s and an energy efficiency of 32.1 GSOP/W for 75% weight sparsity at 140 MHz. The implementation of a four-layer fully connected SNN is expected to perform$2.6~\mu \text{s}$per time step and the energy consumption is$14.6~\mu \text{J}$. Our results demonstrate that ESSA outperforms several state-of-the-art application-specific integrated circuit (ASIC) or FPGA neuromorphic processors. Yisong Kuang, Xiaoxin Cui, Zilin Wang 0001, Chenglong Zou, Yi Zhong 0002, Kefei Liu 0002, Zhenhui Dai, Dunshan Yu, Yuan Wang 0001, Ru Huang 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 8 |
| 2021 | A 28-nm 0.34-pJ/SOP Spike-Based Neuromorphic Processor for Efficient Artificial Neural Network ImplementationsabstractNeuromorphic hardware platforms inspired by human brain have emerged as novel non von Neumann computing architectures. They were proved excellent platforms for spiking neural network (SNN) implementations. However, implementing artificial neural networks (ANNs) on existing neuromorphic hardware platforms is still a daunting task because of critical limitations on coding scheme, maximum of fan-in, and highest weight precision in them. In this paper, we introduce a neuromorphic processor developed for various neural networks implementations including ANNs and SNNs. We employ spatio-temporal coding scheme based on spike events. By combining low-precision dendrites, the chip can implement weight precision between 1 bit and 8 bits and scalable fan-in. The 3.66-mm2chip fabricated in 28-nm CMOS with a maximum fan-in of 72 K per neuron demonstrates unprecedented compatibility with ANN applications compared to previously-proposed neuromorphic chips. Yisong Kuang, Xiaoxin Cui, Yi Zhong 0002, Kefei Liu 0002, Chenglong Zou, Zhenhui Dai, Dunshan Yu, Yuan Wang 0001, Ru Huang 0001 |
ISCAS | 7 |
| 2020 | A 3D Convolutional Neural Network for Emotion Recognition based on EEG SignalsabstractAs an important field of research in Human-Machine Interactions, emotion recognition based on the electroencephalography (EEG) signals has become common research. The traditional machine learning approaches use well-designed classifiers with hand-crafted features which may be limited to domain knowledge. Motivated by the outstanding performance of deep learning approaches in recognition tasks, we proposed a 3D convolutional neural network model to extract the spatial-temporal features automatically in the EEG signals. By the pre-processing method with baseline signals and the electrode topological structure relocated, the proposed model achieves a high accuracy rate of 96.61%, 96.43% in the Two class classification task (low/high arousal, low/high valence) and 93.53% in the Four class classification task (low arousal and low valence/high arousal and low valence/low arousal and high valence/high arousal and high valence) in the DEAP dataset, and 97.52%, 96.96% in the Two class classification task and 95.86% in the Four class classification task in the AMIGOS dataset. Yuxuan Zhao 0002, Jinlong Lin, Dunshan Yu, Xixin Cao |
IJCNN | 4 |
| 2018 | Polymorphic gate based IC watermarking techniquesabstractPolymorphic gates are reconfigurable devices whose functionality may vary in response to the change of execution environment such as temperature, supply voltage or external control signals. This feature makes them a perfect candidate for circuit watermarking. However, polymorphic gates are hard to find because they do not exhibit the traditional structure. In this paper, we report four dual-function polymorphic gates that we have discovered using an evolutionary approach. With these gates, we propose a circuit watermarking scheme that selectively replaces certain standard logic gates with the polymorphic gates. Experimental results on ISCAS and MCNC benchmark circuits demonstrate that this scheme introduces low overhead. More specifically, the average overhead in area, speed and power are 4.10%, 2.08% and 1.17% respectively when we embed 30-bit watermark sequences. These overheads increase to 6.36%, 4.75% and 2.08% respectively when 10% of the gates in the original circuits are replaced to embed watermark up to more than 300 bits. Xiaoxin Cui, Dunshan Yu, Omid Aramoon, Timothy Dunlap, Gang Qu 0001, Xiaole Cui |
ASP-DAC | 3 |
| 2018 | A Novel Polymorphic Gate Based Circuit Fingerprinting TechniqueabstractPolymorphic gates are reconfigurable devices that deliver multiple functionalities at different temperature, supply voltage or external inputs. Capable of working in different modes, polymorphic gate is a promising candidate for embedding secret information such as fingerprints. In this paper we report five polymorphic gates whose functionality varies in response to specific control input and propose a circuit fingerprinting scheme based on these gates. The scheme selectively replaces standard logic cells by polymorphic gates whose functionality differs with the standard cells only on Satisfiability Don't Care conditions. Additional dummy fingerprint bits are also introduced to enhance the fingerprint's robustness against attacks such as fingerprint removal and modification. Experimental results on ISCAS and MCNC benchmark circuits demonstrate that our scheme introduces low overhead. More specifically, the average overhead in area, speed and power are 4.04%, 6.97% and 4.15% respectively when we embed 64-bit fingerprint that consists of 32 real fingerprint bits and 32 dummy bits. This is only half of the overhead of the other known approach when they create 32-bit fingerprints. Xiaoxin Cui, Dunshan Yu, Omid Aramoon, Timothy Dunlap, Gang Qu 0001, Xiaole Cui |
ACM Great Lakes Symposium on VLSI | 3 |
| 2018 | A 108fsrms 0.45mW 100MS/s 1.25MHz bandwidth multi-bit ΔΣ time-to-digital converter with dynamic element matchingabstractA novel ΔΣ time-to-digital converter (TDC) with a time mode accumulator and a multi-bit quantizer is proposed in this work. Measurement time is reduced when compared with single-bit ΔΣ TDCs. A time difference adder consisting of gated delay-line based time-registers is used to serve as the time accumulator. A dynamic element matching algorithm is implemented to mitigate the performance loss degraded by the non-linearity of the multi-bit quantizer. The TDC is designed and simulated using a 65nm CMOS process and operates at a 100MHz sampling rate. For a 1.25MHz bandwidth, 108fsrmsintegrated noise or 2.4ps equivalent resolution is achieved. The power consumption is only 0.45 mW and the figure of merit (FoM) is calculated to be 154fJ/step. Yinxuan Lyu, Jianhua Feng, Chenfeng Tu, Linqi Shi, Hongfei Ye, Weixin Gai, Dunshan Yu |
ISCAS | 7 |
| 2018 | Evaluation of Dynamic-Adjusting Threshold-Voltage Scheme for Low-Power FinFET Circuits
Xiaoxin Cui, Yewen Ni, Dunshan Yu, Xiaole Cui |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2017 | Improving DFA attacks on AES with unknown and random faults
Nan Liao, Xiaoxin Cui, Dunshan Yu, Xiaole Cui |
Sci. China Inf. Sci. | 5 |
| 2016 | Ultralow-power high-speed flip-flop based on multimode FinFETs
Xiaoxin Cui, Nan Liao, Dunshan Yu, Xiaole Cui |
Sci. China Inf. Sci. | 5 |
| 2015 | Key characterization factors of accurate power modeling for FinFET circuits
Kaisheng Ma, Xiaoxin Cui, Nan Liao, Dunshan Yu |
Sci. China Inf. Sci. | 6 |
| 2014 | High-speed constant-time division module for Elliptic Curve Cryptography based on GF(2m)abstractTo achieve high performance scalar multiplication arithmetic in Elliptic Curve Cryptography (ECC) based on GF(2m), a high-speed constant-time division module with optimized architecture is proposed in this paper. Modified from the traditional extended Euclidean Great Common Divisor (GCD) division algorithm, the presented algorithm computes a single multiplicative inverse or division in constant m iterations, i.e. m clock cycles, in GF(2m), which obtains a tremendous reduction (specifically more than 50%) on computing time compared with previous works. Combined with the meticulously optimized architecture, this novel division module achieves lower area-time complexity, which makes it an excellent option for high performance ECC design. Xiaoxin Cui, Nan Liao, Dunshan Yu |
ISCAS | 7 |
| 2014 | Low power adiabatic logic based on FinFETs
Nan Liao, Xiaoxin Cui, Kaisheng Ma, Dunshan Yu |
Sci. China Inf. Sci. | 8 |
| 2014 | Ultra-low power dissipation of improved complementary pass-transistor adiabatic logic circuits based on FinFETs
Xiaoxin Cui, Nan Liao, Kaisheng Ma, Dunshan Yu |
Sci. China Inf. Sci. | 8 |
| 2013 | A Dynamic-Adjusting Threshold-Voltage Scheme for FinFETs low power designsabstractIn this paper, a novel device/circuit co-design scheme, namely Dynamic-Adjusting Threshold-Voltage Scheme (DATS) for independent-gate mode FinFET circuits has been proposed. The main idea of this scheme is that a pair of back-gate bias of FinFETs is adjusted dynamically to change threshold voltage according to the system operating frequency and operating mode, which could optimize circuit power, especially leakage power. The experimental and simulation result shows that the leakage power dissipation reduced greatly when circuits operate at the lower frequency, and the energy-delay product of FinFET circuits is reduced by 30% approximately. Xiaoxin Cui, Kaisheng Ma, Nan Liao, Dunshan Yu |
ISCAS | 8 |
| 2006 | An Efficient VLSI Implementation of Distributed Architecture for DWTabstractThis paper proposes an efficient and simple architecture for 9/7 discrete wavelet transform based on distributed arithmetic. To derive new proposed architecture, we consider the periodicity and symmetry of DWT to optimize the performance and reduce the computational redundancy. The inner product of coefficient matrix of DWT is distributed over the input by careful analysis of input, output and coefficient word lengths. In the coefficient matrix, linear maps are used to assign the necessary computation to processing elements in space domain. Moreover, the proposed architecture has regular data flow, and low control complexity. The result is a low hardware complexity DWT processor for 9/7 transforms, which allows two times faster clock than the direct implementation. This design is very suitable for image compression systems, e.g., JPEG2000 and MPEG4 Xixin Cao, Qingqing Xie, Chungan Peng, Qingchun Wang, Dunshan Yu |
MMSP | 5 |