VLDB 2026 Research / reviewers in the wild / expert
Shushan Qiao
dblp:21/10039
· DBLP profile ↗
33ranked-venue papers
0as first author
33since 2021 · last 2026
0000-0002-9102-2111ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 23 · 23 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A 0.277pW/bit Sleep 10T SRAM with Improved Read/Write SNM for Low Leakage Applications
Xiang Li 0164, Pengyuan Zhao, Minglong Jia, Linnan Li, Shushan Qiao |
ISCAS | 7 |
| 2026 | MPE: A Power-Efficient Edge-Device Mamba Processor with Multi-Dimensional Calculation-Compression Scheme
Zhou Wang 0005, Haochen Du, Jiuren Zhou, Xiguang Wu, Qiankun Li 0004, Yanqing Xu 0003, Hanqi Feng, Xiaonan Tang, Shushan Qiao, Yongke Wang, Anil A. Bharath, Emm Mic Drakakis |
ISCAS | 10 |
| 2026 | GTPE: A 28nm 33.12 TFLOPS/W GNN Training Processor with Unstructured Multi Threshold Pruning, Hybrid Multi-mode Approximate Computing and QUIRE Number System Support
Zhou Wang 0005, Haochen Du, Jiuren Zhou, Xiguang Wu, Qiankun Li 0004, Yanqing Xu 0003, Hanqi Feng, Xiaonan Tang, Shushan Qiao, Tian-Chun Ye 0001, Anil A. Bharath, Emm Mic Drakakis |
ISCAS | 10 |
| 2026 | GATPE: A High-Performance Edge-Device GAT Processor with Multi-Layer Data-Variation Mechanism
Zhou Wang 0005, Haochen Du, Jiuren Zhou, Xiguang Wu, Qiankun Li 0004, Yanqing Xu 0003, Hanqi Feng, Xiaonan Tang, Shushan Qiao, Anil A. Bharath, Emm Mic Drakakis |
ISCAS | 11 |
| 2026 | A 2-18 GHz LNA Based on Dual-mode Hybrid Coupled Frequency Response Reshaping Technique
Kuisong Wang, Henan Zhang, Xuming Sun, Shushan Qiao, Xiaoxin Liang |
ISCAS | 6 |
| 2026 | An Approximate Digital Compute-In-Memory Macro with Reconfigurable Computational Precision for Neural Network Acceleration
Heng You, Zixiao Zhan, Guanghua Zhao, Shushan Qiao, Jia Yuan |
ISCAS | 6 |
| 2026 | A Low-complex and Dual-SPSA Digital Predistortion for MIMO Transmitters with Crosstalk
Jiayin Song, Wulve Yang, Shushan Qiao |
ISCAS | 7 |
| 2026 | C-DPSS: Channel dual-phase sparsity pruning framework for spiking neural networksabstractSpiking Neural Networks (SNNs) have emerged as an essential paradigm for brain-inspired computing, achieving superior energy efficiency on neuromorphic hardware. However, as network scale increases, SNNs encounter growing challenges in deployment efficiency. While structured pruning provides a practical approach, existing methods typically rely on a dense-to-sparse training pattern, which incurs high computational costs during training and fails to exploit the efficiency of sparse computation from the outset. There is a lack of effective frameworks that enforce structured sparsity within a sparse-to-sparse training regime for SNNs. To bridge this gap, we propose the Channel Dual-phase Sparsity (C-DPSS) framework. This approach leverages Dynamic Sparse Training (DST) to enable efficient topology exploration while progressively enforcing hardware-friendly structured sparsity. Our methodology is grounded in the empirical observation that channel-level neuronal activity patterns stabilize earlier than synaptic weights under certain training regimes. Motivated by this temporal discrepancy, C-DPSS employs a dual-phase saliency metric during training. It defines an early exploration phase dominated by neuronal dynamics, and smoothly transitions to a late refinement phase driven by synaptic efficacy. Extensive experiments on both static and neuromorphic benchmarks demonstrate that C-DPSS provides a robust accuracy-efficiency trade-off. Notably, for VGG-16 on CIFAR-10, C-DPSS achieves 50% channel sparsity and reduces SOPs by 42.1% under S W = 0.90 and S C = 0.50 , with no observable accuracy degradation, showing a favorable accuracy–efficiency trade-off under structured channel pruning. We further examine the large-scale behavior of our approach through an ImageNet-1K feasibility study of sparse-to-sparse structured pruning. Our code is available at: https://github.com/JunLi0514/C-DPSS . Yiying Jiang, Jionghao Zhang, Hailong Zou, Mengdie Tao, Hang Ran, Bingchen Zhang, Shushan Qiao |
Neurocomputing | 13 |
| 2026 | MAFNet: Mamba-based asymmetric fusion network for event-based motion deblurring
Qihang Jiang, Weiliang Meng, Yumeng Ren, Hailong Zou, Shushan Qiao |
Inf. Sci. | 8 |
| 2026 | A bio-inspired spiking neural network with adaptive spatiotemporal filtering and depth-modulated synaptic plasticity for robust collision detection
Yumeng Ren, Qihang Jiang, Hang Ran, Shushan Qiao |
Neural Networks | 11 |
| 2026 | A Booth-Based Digital Compute-in-Memory Macro With Bitwise-Efficient Multiphase AccumulationabstractThis brief presents a Booth-based all-digital SRAM compute-in-memory (CIM) macro designed for high-efficiency multiply-and-accumulate (MAC) operations in artificial intelligence applications. The architecture incorporates three key innovations: 1) a pre-encoding (PENC) scheme utilizing a Booth-based consolidated truth table, designed to reduce encoder complexity and computing unit overhead; 2) a minimalist bitwise processing technique, employing only two transistors for bit retention and shifting, aimed to minimize computational complexity within the subcomputing array; and 3) a multiphase product summation approach, integrating unified negative compensation with global sign bit expansion, intended to lower power consumption during partial product accumulation and final addition. The proposed architecture is readily adaptable to multiplications across diverse bit widths and accommodates flexible CIM array sizes, ensuring high scalability. Measurement results from a 55-nm-CMOS process demonstrate that the proposed CIM macro achieves an energy efficiency of 16.8 TOPS/W under 8b/8b MAC operations. Yiying Jiang, Heng You, Shushan Qiao |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2026 | CTC-MVPT: A Cross-Temperature Continuous Tracking System for Minimum Voltage Point Based on a Universal Delay Chain With Multiple Monitoring PointsabstractThis brief presents a cross-temperature continuous tracking system for minimum voltage point (CTC-MVPT) based on a universal delay chain (UDC) with multiple monitoring points. By leveraging the timing information provided by the timing margin indicator and timing error indicator, the CTC-MVPT system enables continuous cross-temperature tracking of the minimum voltage point (MVP), significantly reducing power. Meanwhile, the two-step adjustment technique greatly shortens the time required for voltage regulation. In addition, this work further optimizes the 0-to-1 and 1-to-0 transition deviations of the UDC. Measurement results in a 22-nm technology show that the proposed CTC-MVPT system achieves up to 17.7% power reduction over the traditional adaptive voltage scaling (AVS) system as the temperature increases from$0~^{\circ }$C to$80~^{\circ }$C. Therefore, it takes only 1.17ms to adjust the voltage from 0.55 to 0.41 V, representing a 47.8% reduction compared to the traditional scheme. Under the TT corner, the maximum deviation rate of the UDC is only 1.5%, and the maximum deviation between 0-to-1 and 1-to-0 transitions is merely 1.2%. Meanwhile, the system achieves a power gain of up to 55.3% with only a 0.26% area overhead for timing monitoring. Gaoteng Zhang, Kangning Wang 0004, Jiliang Liu, Linnan Li, Shushan Qiao |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2026 | A 0.31-V 16-Kb 9T SRAM With Enhanced Sensing Margin and Read Performance for Low-Power ApplicationsabstractThis brief presents a low-power 9T static random access memory (SRAM) with enhanced read sensing margin and read performance. The read decoupled port of the proposed 9T SRAM cell achieves the enhanced sensing margin by mitigating the read bitline (RBL) leakage and improves the read performance through using one-transistor read path. The multithreshold voltage devices are used in SRAM cell for improving the leakage power and performance of SRAM. Additionally, an interleaved write wordline (WWL) structure is implemented to address the write half-select issue. The measurement results of the test chip fabricated in the 22-nm FDSOI technology demonstrate that the designed 9T SRAM achieves a minimum operation voltage of 0.31 V at 1.05 MHz and can operate at 60.5 MHz when the supply voltage is 0.5 V. The minimum active energy of 18.56 fJ/access-bit is obtained at 0.33 V. Furthermore, the designed SRAM exhibits a minimum leakage power of 0.11 pW/bitcell in the retention mode. Pengyuan Zhao, Linnan Li, Zhi Li 0090, Minglong Jia, Xiang Li 0164, Shushan Qiao |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2025 | STPE: An Energy-Efficient Edge-Device Transformer Inference Processor with Multi-Mode Data-Compression SchemeabstractTransformer-Based models have turned out to be very successful in many artificial intelligence (AI) tasks, outperforming traditional convolutional neural networks (CNNs), especially in the field of Natural Language Processing (NLP). Their success relies upon a self-attention mechanism which, when compared to CNNs, has a global rather than a local receptive domain. This article proposes an energy-efficient edge-device Transformer inference processor termed Smart Transformer Processing Element (STPE). Firstly, STPE sets up a Multi-Mode Indexing and Sparsity Scheme (MISS) for token association, and further reduces the computational load through in-situ computation; secondly, STPE exploits the Local Properties of Attention Mechanism (LPAM) to further reduce redundant and repetitive calculations in Transformer operations by means of a search band calculation and error correction mechanism; thirdly, STPE has designed a Quantization and Compression Parallel Method (QCPM) to improve the computing speed and hardware utilization under weak related (WR) token. Employing 28nm CMOS synthesis tools, the area of the proposed STPE processor is 7.33 mm2. Its peak energy efficiency is 84.15TOPS/W, which is 14.7 times higher than that of the H100 graphics processing unit (GPU) and 3.06 times higher than that of the most advanced Transformer processor. Zhou Wang 0005, Haochen Du, Vivek Mohan, Jiuren Zhou, Yanqing Xu 0003, Baoyi Han, Xiaonan Tang, Shushan Qiao, Shouyi Yin, Anil A. Bharath, Emmanuel M. Drakakis |
ISCAS | 9 |
| 2025 | A 230-nA Quiescent Current and Enhanced Transient Performance DC-DC Converter with Suppressed Under/Overshoot for IoT SoCsabstractThis paper presents a low quiescent current, transient-enhanced DC-DC buck converter with an overshoot and undershoot suppression scheme, designed to support low-power SoCs. To decrease undershoot or overshoot voltage during rapid load current transitions, an overshoot and undershoot voltage detection and response scheme is proposed. The buck converter employs an adaptive constant on time control mode, and to achieve automatic transition between PWM and PFM modes, a sub-threshold zero current detector(ZCD) is introduced. To reduce quiescent current and improve efficiency under light loads, a sampling-based low-power voltage reference circuit is proposed. Implemented with a 55 nm CMOS process, the buck converter operates with a load current range of 10 μA-10 mA. It achieves a quiescent current of 230 nA and a peak efficiency of 93.2%. During a 10 mA load variation, the converter exhibits an undershoot of 36 mV and an overshoot of 34 mV. Zhi Li 0090, Linnan Li, Shushan Qiao |
ISCAS | 5 |
| 2025 | AO-EDC: An Accuracy-Oriented Error Detection and Correction Scheme for DVFS System Based on Propagation Detection at Half-Path PointsabstractError detection and correction (EDC) techniques are conventionally employed in dynamic voltage and frequency scaling (DVFS) systems to eliminate timing margins preserved in ICs. However, existing error detection methods generate warning signals even when no violations occur, resulting in frequent corrections that degrade system throughput. This issue of inaccurate detection worsens with further scaling. To solve this issue, this paper presents an accuracy-oriented EDC scheme, AO-EDC. An accurate detection method, consisting of a propagation detector (PD) and an improved half-path-point insertion method, is proposed. PDs that filter out signals from other paths are inserted at the half-path points of selected paths, reducing inaccurate detection by up to 77.5%. In addition, a half-cycle clock gating circuit is proposed to reduce the redundancy of corrections. Implemented on a 22-nm process FIR circuit, post-layout simulation results demonstrate a 42.4% - 50.0% power reduction and a 106.7% - 1089.9% frequency gain at 0.9 - 0.55 V. The proposed scheme triggers fewer corrections and reduces correction timing waste, addressing the throughput degradation issue during over-scaling in DVFS systems. Zitao Liang, Jiliang Liu, Kangning Wang 0004, Linnan Li, Zhi Li 0090, Shushan Qiao |
ISCAS | 7 |
| 2025 | GPE: A High-Performance Edge GNN Inference Processor with Multi-Parallelism Format-Variation MechanismabstractRecently, Graph Neural Networks (GNNs) have shown great potential in terms of accuracy for problems that are well-described by graph representations, such as problems of path planning. However, implementing GNNs on mobile platforms is challenging as it requires a significant amount of computation and large memory. This article proposes a High-Performance Edge GNN Inference Processor termed GPE (GNN Processing Element). Firstly, GPE sets up Multi-Dimensional Indexing and Dynamic Pruning Schemes (MIDPS) for GNN networks, and achieves cross layer interconnection of multiple neighboring nodes via NOC (Network on Chip); secondly, GPE utilizes Graph Structure Adjacency Table Information (GSATI) of a GNN to further reduce redundant and repetitive calculations by means of repeated matching and difference transfer mechanisms; thirdly, GPE has a graph-based Multi Parallelism Simplification and Operation Method (MPSOM) to improve computing speed and hardware utilization under small data volumes. Using 28nm CMOS synthesis tools, the area of the proposed GPE processor is 5.37 square millimeters. Its peak energy efficiency is 21.5TOPS/W, which is 3.76 times higher than that of the H100 GPU (Graphics Processing Unit), while the energy consumption of GNN is 80.9% lower than the previous SOTA (State of Art) work. Zhou Wang 0005, Haochen Du, Jiuren Zhou, Yanqing Xu 0003, Vivek Mohan, Baoyi Han, Xiaonan Tang, Shushan Qiao, Shouyi Yin, Anil A. Bharath, Emmanuel M. Drakakis |
ISCAS | 9 |
| 2025 | Hardware Friendly Transformer Optimization with Dynamic Attention Matrix FusionabstractThe multi-head self-attention (MHSA) is the core component of the transformer, where dynamic matrix multiplications (DMM), particularly Q×KTand A′×V, pose significant challenges for hardware acceleration. To reduce DMM MACs, this paper proposes a Dynamic Attention Matrix Fusion (DAMF) method, which optimizes DMM from the attention algorithm. For Q × KT, a quadratic form fusion of WQWKweight matrices and an SVD approximation is introduced, transforming DMM into fewer scalar operations and eliminating the linear transformations for QK generation. For A′× V, this paper proposes approximating softmax using a Maclaurin series and power-of-2 as a shift factor, replacing A′× V with a hardware-friendly shift operation. Experimental results show that the proposed DAMF method does not cause significant accuracy loss in the BERT-base. Additionally, compared to MHSA with the same configuration, DAMF reduces parameters by 1.99 times, DMM MACs by 284 times, total MACs by 2.21 times and memory access by 2.71 times. Qingyao Yang, Shushan Qiao |
ISCAS | 5 |
| 2025 | DIME-Net: A Dual-Illumination Adaptive Enhancement Network Based on Retinex and Mixture-of-ExpertsabstractImage degradation caused by complex lighting conditions such as low-light and backlit scenarios is commonly encountered in real-world environments, significantly affecting image quality and downstream vision tasks. Most existing methods focus on a single type of illumination degradation and lack the ability to handle diverse lighting conditions in a unified manner. To address this issue, we propose a dual-illumination enhancement framework called DIME-Net. The core of our method is a Mixture-of-Experts illumination estimator module, where a sparse gating mechanism adaptively selects suitable S-curve expert networks based on the illumination characteristics of the input image. By integrating Retinex theory, this module effectively performs enhancement tailored to both low-light and backlit images. To further correct illumination-induced artifacts and color distortions, we design a damage restoration module equipped with Illumination-Aware Cross Attention and Sequential-State Global Attention mechanisms. In addition, we construct a hybrid illumination dataset, MixBL, by integrating existing datasets, allowing our model to achieve robust illumination adaptability through a single training process. Experimental results show that DIME-Net achieves competitive performance on both synthetic and real-world low-light and backlit datasets without any retraining. These results demonstrate its generalization ability and potential for practical multimedia applications under diverse and complex illumination conditions. Dingyi Wang, Shushan Qiao |
ACM Multimedia | 5 |
| 2025 | A Charge Domain SRAM Computing-in-Memory Macro With Quantized Interval-Optimized ADC and Input Bit-Level Sparsity-Optimized P2O-DAC for 8-b MAC OperationabstractComputing-in-memory (CIM) has recently gained significant attention as it achieves high energy efficiency and throughput for deep convolutional neural networks (DCNNs). In this brief, we present a static random access memory (SRAM) CIM macro aimed at improving the energy efficiency of edge devices when performing 8-b multiply-and-accumulate (MAC) operations. The proposed architecture implements the following: 1) a successive approximation register analog-to-digital converter (SAR ADC) readout circuit based on a weight-flip-store (WFS) coding scheme, where energy efficiency is improved by optimizing the quantized interval; 2) an input-relevant partial power-off digital-to-analog converter (P2O-DAC) using input bit-level sparsity to reduce power consumption; and 3) a pipeline structure for interleaving MAC computation and readout operation to minimize the redundancy when loading input data into the CIM array. Our proposed CIM macro is implemented in TSMC 40-nm CMOS technology. Postlayout simulation results show an average macro energy efficiency of 16.8 TOPS/W without input and weight value sparsity. Shukao Dou, Zupei Gu, Heng You, Shushan Qiao |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2025 | A General-Purpose Computing Core With Cooperative Motion Detection and Feature Extraction for Always-On PWM Image SensorsabstractThis article presents a mixed-signal general-purpose computing core designed to support real-time inference applications in always-on low-power pulsewidth modulation (PWM) CMOS image sensors (CISs), functioning as a processing-in-sensor (PIS) circuit. This core can be integrated into the columns of CISs to perform low-power edge processing on images, without affecting the pixel fill factor and imaging quality. It employs a coordinated mechanism in which motion detection (MD) triggers the activation of feature extraction (FE), thereby achieving an organic integration of MD and FE functionalities, and maximizing the system’s power efficiency. MD is implemented via in-column frame difference (FD), while FE is performed using a programmable-weight$3\times 3$convolution, a rectified linear unit activation function, and a$2\times 2$max-pooling (MP) operation. Both functionalities are computed based on real-time PWM signals from the CIS and the principle of current integration, with partial circuit reuse achieved through different switching operations. A 0.8-V computing core prototype, with an area of 720 × 272$\mu$m, was fabricated and verified using 0.18-$\mu $m standard CMOS technology. The experimental results at an image frame rate of 250 fps demonstrated an average power consumption of$5.23~\mu $W for MD and$17.53~\mu $W for FE. The prototype core computes the first two layers of an ultra lightweight convolutional neural network (CNN) for the task of MNIST digit classification, achieving an accuracy loss of only 0.86% compared to the ideal scenario. This analog computing core can be used in multimode, low-power, edge-intelligent vision sensors. Jinyu Gao, Aoming Zhan, Qihang Jiang, Feng Zhang 0014, Yong Chen 0005, Shushan Qiao |
IEEE Trans. Very Large Scale Integr. Syst. | 9 |
| 2025 | A Self-Calibrated Unified Voltage-and-Frequency Regulator System Design Based on Universal Logic Line CircuitabstractIn this brief, a unified voltage frequency regulator (UVFR) system is designed to eliminate the voltage margin induced by process, voltage, and temperature (PVT) variations. The frequency is regulated with voltage by a universal logic line oscillator (ULLO), which can protect the system from timing violations. The length of the ULLO is self-calibrated by a ULL-based time-digital converter (ULL-TDC) and an in situ half-critical path timing detector, where the ULL is designed to track the critical path delay. A fully synthesizable digital low dropout (DLDO) is designed with the ULL-TDC and a proportional differential (PD) circuit for voltage regulation. The proposed system is implemented in an ARM Cortex-M0 microcontroller in 22 nm technology. Simulation results show that the ULL can accurately track the critical path delay with a maximum variation of 3% at 0.6 V and 11.5% at 0.45 V. The UVFR system consumes 13.2–112 uW of power overhead, and eliminates the voltage margin by 22.3%–28% while reducing the power consumption by 35%–42.3%. Jiliang Liu, Zhi Li 0090, Kangning Wang 0004, Shushan Qiao |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2025 | Accelerating Unstructured Sparse DNNs via Multilevel Partial Sum Reduction and PE Array-Level Load BalancingabstractUnstructured pruning introduces significant sparsity in deep neural networks (DNNs), enhancing accelerator hardware efficiency. However, three critical challenges constrain performance gains: 1) complex fetching logic for nonzero (NZ) data pairs; 2) load imbalance across processing elements (PEs); and 3) PE stalls from write-back contention. This brief proposes an energy-efficient accelerator addressing these inefficiencies through three innovations. First, we propose a Cartesian-product output-row-stationary (CPORS) dataflow that inherently matches NZ data pairs by sequentially fetching compressed data. Second, a multilevel partial sum reduction (MLPR) strategy minimizes write-back traffic and converts random PE stalls into manageable load imbalance. Third, a kernel sorting and load scheduling (KSLS) mechanism resolves PE idle/stall and achieves PE array-level load balancing, attaining 76.6% average PE utilization across all sparsity levels. Implemented in 22-nmCMOS, the accelerator delivers$1.85\times $speedup and$1.4\times $energy efficiency over baseline and achieves 25.8 TOPS/W peak energy efficiency at 90% sparsity. Chendong Xia, Zhi Li 0090, Bing Li 0017, Shushan Qiao |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2024 | A 409mV, Sub-10nW Power-on Reset Circuit Using Adaptive Accuracy Adjustment for Low Voltage ApplicationsabstractA low power power-on reset (POR) circuit with low temperature coefficient, low quiescent current and small area is proposed in this paper. The POR circuit samples the power supply voltage through a high threshold transistor and then converts the voltage to a current, which will be compared with a native NMOS based current reference to obtain the reset signal. In order to reduce the power consumption in steady state, an adaptive accuracy adjustment mechanism is employed in the proposed POR circuit. The POR circuit uses a low accuracy but energy efficient structure to monitor the supply voltage in steady state, and when a voltage drop is observed, the POR circuit quickly switches to a high accuracy mode to get an accurate brown-out detection trip voltage. The POR circuit is implemented in 55nm CMOS process and the active area is just 107μm2. Post-layout simulation results show that the POR circuit has a POR trip voltage of 409mV, a static power of 9.11nW at a supply voltage of 0.45V. Besides, the temperature coefficient of the proposed POR circuit is only 31.76μV/°C over a temperature range of -40°C to 125°C. Heng You, Dashan Shi, Delong Shang, Shushan Qiao |
ISCAS | 5 |
| 2024 | Differentiable architecture search with multi-dimensional attention for spiking neural networks
Yilei Man, Linhai Xie, Shushan Qiao, Delong Shang |
Neurocomputing | 3 |
| 2024 | A 1-8b Reconfigurable Digital SRAM Compute-in-Memory Macro for Processing Neural NetworksabstractThis work presents a 1-8b reconfigurable digital SRAM compute-in-memory (CIM) macro, which significantly improves array utilization and energy efficiency under different input and weight configurations compared to previous works. To ensure the array utilization under different configurations, a row-based bitwise-summation-first digital CIM architecture is proposed. In addition, to realize flexible switching between signed and unsigned operations, a complete 2’s complement encoding method is adopted, which makes the computation of the sign bits consistent with that of the magnitude bits when performing signed operations, thus ensuring that each row of the CIM array can store the sign of the weight. Due to the support of reconfigurable bit width, the proposed CIM macro can be widely used in various neural networks for optimal efficiency. In order to better apply the CIM macro to binarized neural networks, a configurable bitwise multiplier is presented, which supports both AND and XNOR operations. Moreover, since the power consumption of the adder tree occupies a major part of the digital CIM macro, a 4–2 compressor based adder tree is presented to further improve the energy efficiency. Measurement results based on 55nm CMOS process show that the proposed CIM macro achieves an energy efficiency of up to 2238TOPS/W at 1b/1b and 44.82TOPS/W at 4b/4b MAC operations. Heng You, Delong Shang, Shushan Qiao |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2024 | CS-TTD: Triplet Transformer for Compressive Hyperspectral Target DetectionabstractDetecting targets in hyperspectral image under compressive sensing (CS) is crucial for real-time applications in front ends. However, applying the existing deep learning (DL) methods directly to compressed hyperspectral image has been challenging, including the uncertainty introduced by the random sensing matrix in the CS acquisition model, and the ineffectuality of traditional similarity metrics. In this article, we present a solution to these challenges by introducing a CS-based triplet transformer detector (CS-TTD). Our approach involves data augmentation using the restricted distribution property (RDP) to generate samples for a balanced number of positive samples, followed by the Siamese network with a triplet transformer to transition spectral vectors from the compressed domain to an embedding space. In addition, we introduce a combined convolution network (CCN) classification module to replace traditional metrics for obtaining classification results. To improve the training and make full use of dataset labels, we also present a two-stage training approach, known as intercategory separation and intracategory aggregation (ISIA), combined with hard-negative mining (HNM) and semi-HNM. Besides, to estimate the minimal number of compressed bands for sufficient detection performance, we introduce a kernel distribution estimation-based sequential iteration (KDE-SI). Our method achieves hyperspectral target detection (HTD) without reconstruction, and experimental results demonstrate that its performance at 0.16 compression ratio (CR) is comparative to the existing methods in the original domain. The code for this work is available athttps://github.com/nicyyyy/CS-TTD.git. Qingyao Yang, Shushan Qiao |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Compressive Hyperspectral Target Detection With Restricted Distribution PropertyabstractCompressive sensing is a widely utilized technique in hyperspectral imaging front-end, and hyperspectral target detection (HTD) in the compressive sensing domain is crucial for conserving computational and storage resources on front-end platforms and achieving real-time detection. However, existing HTD methods can only handle reconstructed hyperspectral images. To address this limitation, this study investigates a novel property under hyperspectral compressive sensing and proposes an algorithm framework for HTD without reconstruction. The newly proposed property, named Restricted Distribution Property (RDP), is based on assumptions about spectral random vectors and derivations from compressive sensing model. It indicates that the spectral vectors in the compressed hyperspectral follow a deterministic Gaussian distribution under specific conditions, and provides a generalized expression for their covariance matrix. Building upon this property, a KL divergence subspace distribution estimator is developed, enabling HTD without reconstruction. Experimental result demonstrates that the performance of proposed detector in compressed hyperspectral data is comparable to or even superior to conventional methods in original hyperspectral images. The result is inspiring for future research. Qingyao Yang, Dingyi Wang, Baoying Yu, Shushan Qiao |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | A Binary Keyword Spotting System with Error-Diffusion Based Feature Binarization
Dingyi Wang, Mengjie Luo, Shushan Qiao |
INTERSPEECH | 5 |
| 2022 | Low-complex and Highly-performed Binary Residual Neural Network for Small-footprint Keyword Spotting
Xiao Wang 0022, Shushan Qiao |
INTERSPEECH | 4 |
| 2022 | Error-Diffusion Based Speech Feature Quantization for Small-Footprint Keyword SpottingabstractNeural network based keyword spotting (KWS) system is a critical component for user interaction in current smart devices. Although small-footprint networks have been widely explored to reduce deployment overhead, low-precision input feature representation still lacks in-depth research. In this letter, an error-diffusion based speech feature quantization method is proposed. Specifically, our algorithm adapts image processing to quantize the input speech feature maps in arbitrary bits. Experiments show that in the 10-keyword KWS task, our 3-bit representation only brings a 0.45% average accuracy drop compared to the full-precision log-Mel spectrograms while others drop over 3%. In the 2 keywords task, our 3-bit representation produces no significant differences, while 1-bit quantization only leads to an average of 1.7% accuracy drop and is even capable of handling similar keywords and imbalanced data distribution. The result proves our method, to the best of our knowledge, is the first practical method that supports as low as 1-bit quantization for single-channel speech features in small-footprint KWS. In addition, we analyze the impact of error-diffusion directions and conclude that time-direction diffusion is more suitable for temporal convolutional networks. Mengjie Luo, Dingyi Wang, Shushan Qiao |
IEEE Signal Process. Lett. | 4 |
| 2021 | A 0.5V 36nW 10-Transistor Power-on-Reset Circuit with High AccuracyabstractIn this paper, a low voltage high accuracy 10- transistor power-on-reset circuit with brown-out-reset function is proposed. A native NMOS current reference based architecture is proposed to get high accuracy trip-voltage with a small area and power consumption. By adjusting the number of native NMOS transistors, a stable hysteresis window is obtained. Post-layout simulation results based on SMIC 55nm CMOS process show that the trip-voltage deviation of the proposed power-on-reset circuit is only 34mV under different temperature and process corners. Also, the trip-voltage of the proposed power-on-reset circuit shows great robustness to supply ramp time. The power consumption of the proposed circuit is as low as 36nW at 0.5V. Since the proposed power-on-reset circuit consists of only 10 transistors, the area is as low as 67.5μm2. Heng You, Jia Yuan, Zenghui Yu, Shushan Qiao |
ISCAS | 4 |
| 2021 | Low-Power Retentive True Single-Phase-Clocked Flip-Flop With Redundant-Precharge-Free OperationabstractAs basic components, optimizing power consumption of flip-flops (FFs) can significantly reduce the power of digital systems. In this article, an energy-efficient retentive true-single-phase-clocked (TSPC) FF is proposed. With the employment of input-aware precharge scheme, the proposed TSPC FF precharges only when necessary. In addition, floating node analysis and transistor level optimization are employed to further ensure the high energy efficiency of the FF without significantly increasing the area. Postlayout simulations based on SMIC 55-nm CMOS technology show that at a supply voltage of 1.2 V, the power consumption of the proposed FF is 84.37% lower than that of conventional transmission-gate flip-flop (TGFF) at 10% data activity. The reduction rate is increased to 98.53% as the data activity goes down to 0%. When the supply voltage decreases to 0.6 V, the proposed FF consumes only 0.411 fJ/cycle at 10% data activity, which is 84.23% lower than TGFF. Measurement results of ten test chips demonstrate the great energy efficiency of the proposed FF. Furthermore, the CK-to-Q delay of the proposed FF is 26.18% lower than that of TGFF at a supply voltage of 1.2 V. Heng You, Jia Yuan, Zenghui Yu, Shushan Qiao |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |