EDBT 2026 Demo / reviewers in the wild / expert
Junyi Qian
dblp:310/0068
· DBLP profile ↗
5ranked-venue papers
1as first author
5since 2021 · last 2026
0000-0001-9390-4567ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A 2.2-$μ$W Keyword Spotting System With Filterbank Learning and Attention-Aware Soft-Threshold Denoising in 28-nm CMOSabstractKeyword spotting (KWS) is crucial for speech interactions in mobile devices, yet its widespread adoption is constrained by the trade-off between low power and high accuracy in diverse noisy environments. Here, we present an ultra-low-power and robust KWS chip to address this challenge through architecture and circuit co-optimization. First, a learnable Mel filter bank is proposed and co-designed with a group-based time-multiplexed architecture to perform adaptive spectral reshaping in a log-Mel feature extractor, reducing the downstream network computations by 25.8% while preserving feature representation capability. Second, a broadcasted residual network (BC-ResNet) is designed to support both KWS and speaker verification (SV) classification tasks, which reduces additions by 43.5% and parameters by 5% compared to a lightweight depthwise separable CNN (DSCNN). Third, an attention-aware soft-threshold denoising strategy is developed to dynamically filter noisy features, yielding up to a 5.8% increase in accuracy over the baseline without denoising at 0dB signal-to-noise ratio (SNR). Finally, a multiplier-less processing element (PE) array combined with stacked-transistor-based ultra-low-leakage custom memory is designed, reducing hardware area by 31.3% and leakage power by 20.4%. Fabricated in a 28nm CMOS process, the proposed 12-class KWS chip consumes only$2.2\mu $W at 0.42V supply. It achieves 87.6%-93.5% accuracy on GSCD across 0dB to 20dB SNR levels, demonstrating high robustness for diverse scenarios. Zhihao Yan, Lishuo Deng, Junyi Qian, Weiwei Shan |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2025 | A 40nm Early-Warning AVFS Design with Path Activation-Based Monitoring Point OptimizationabstractThis work presents an early-warning adaptive voltage-frequency scaling (AVFS) system, offering a practical and innovative low-power solution for commercial IoT products. First, a novel monitoring point optimization strategy based on path activation analysis is proposed, which mitigates the risk of illegal voltage scaling and reduces monitoring costs by 18.7%. Second, a dual shadow register-based monitor for memory endpoints is designed to overcome the limitation of monitoring only D-type flip-flop (DFF) endpoints. Third, a single-chip solution that integrates both frequency and voltage regulation strategies is presented to alleviate the issues of off-chip voltage regulation, such as large noise ripple and long response times. Fabricated in a 40-nm CMOS process, a Cortex M0+ CPU with AVFS works within a frequency range of 50-100MHz and a voltage range of 0.75-1.1V. Furthermore, the proposed AVFS achieves significant power savings of up to 49.3% at FF, 25°C, and 100MHz. Kaize Zhou, Lishuo Deng, Junyi Qian, Xiaojie Yin, Kana Meng, Weiwei Shan |
ISCAS | 5 |
| 2025 | A Compound Timing Detection of Both Data Transition and Path Activation for Reliable In Situ Error Detection and CorrectionabstractTiming error detection and correction (EDAC) in resilient circuits helps eliminate excess timing margins. However, it faces misdetection risks when critical paths (CPs) remain inactive. We propose a 15-transistor timing error and path-activation detector (TEPD) capable of detecting both timing violations in late-arriving signals and CP activation, with robust operation down to 0.33 V. Regarding circuit-level error correction, we introduce an error-correcting flip-flop (ECFF), leveraging time-borrowing for zero-cycle response latency without requiring pipeline refresh. The custom-optimized ECFF adds only six transistors, increasing delay, dynamic power, and static power by 15%, 11%, and 14%, respectively, compared with a standard flip-flop, ensuring efficient error correction with minimal cost. A system-level voltage tuning strategy is further developed to handle continuous timing errors, ensuring robust adaptive voltage scaling (AVS) operation. Implemented on a neural network (NN) accelerator in the 28-nm CMOS, the system operates across a wide voltage range from 0.54 to 0.9 V. It achieves up to 52% power gain or 123% frequency gain at the near-threshold region, with negligible area and power overhead compared with the margined baseline. Lishuo Deng, Junyi Qian, Zhengguo Shen, Jingchen Wang, Zhangrui Qian, Longning Qi, Weiwei Shan |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2025 | An On-Chip-Training Keyword-Spotting Chip Using Interleaved Pipeline and Computation-in-Memory Cluster in 28-nm CMOSabstractTo improve the precision of keyword spotting (KWS) for individual users on edge devices, we propose an on-chip-training KWS (OCT-KWS) chip for private data protection while also achieving ultralow -power inference. Our main contributions are: 1) identity interchange and interleaved pipeline methods during backpropagation (BP), enabling the pipelined execution of operations that traditionally had to be performed sequentially, reducing cache requirements for loss values by 95.8%; 2) all-digital isolated-bitline (BL)-based computation-in-memory (CIM) macro, eliminating ineffective computations caused by glitches, achieving 2.03$\times$higher energy efficiency; and 3) multisize CIM cluster-based BP data flow, designing each CIM macro collaboratively to achieve all-time full utilization, reducing 47.2% of output feature map (Ofmap) access. Fabricated in 28-nm CMOS and enhanced with a refined library characterization methodology, this chip achieves both the highest training energy efficiency of 101.5 TOPS/W and the lowest inference energy of 9.9nJ/decision among current KWS chips. By retraining a three-class depthwise-separable convolutional neural network (DSCNN), detection accuracy on the private dataset increases from 80.8% to 98.9%. Junyi Qian, Peng Cao 0002, Xin Si, Weiwei Shan |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2023 | An All-Digital, 1.92-7.32 mV/LSB, 0.5-2 GS/s Sample Rate, and 0-Latency Prediction Voltage Sensor With Dynamic PVT Calibration for Droop Detection and AVS SystemabstractThe on-chip droop in processor may cause a severe voltage reduction resulting in a need for high-speed and high-resolution on-chip voltage sensors. However, traditional voltage sensors hardly achieve high resolution at GHz-level sampling rate and require multiple cycles to obtain quantized results. Therefore, we propose an on-chip all-digital voltage sensor with dynamic PVT calibration for droop detection and a unified voltage monitor and scaling (UVMS) systems for efficiency improvement. We propose a balanced ring oscillator for high resolution and Nyquist counter-based encoder for GHz-level sampling. A light-weight dynamic PVT calibration using temperature sensor is proposed to resist the PVT variations and nonlinearity. Then, an adaptive prediction mechanism is proposed for 0-latency voltage detection. Fabricated in a 28nm CMOS technology, our voltage sensor achieves a high resolution of 7.32 mV/LSB at a 2 GS/s or 1.92 mV/LSB at a 0.5 GS/s sample rate with 0-latency and a calibration error of only 1 LSB. Implemented in a BNN accelerator, proposed UVMS compresses the margin and achieves a power gain of 37% - 57% at frequencies ranging from 31 MHz to 337 MHz. Junyi Qian, Zhuo Chen 0039, Weiwei Shan |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |