EDBT 2026 Demo / reviewers in the wild / expert
Jun Zhou 0017
dblp:99/3847-17
· DBLP profile ↗
48ranked-venue papers
2as first author
34since 2021 · last 2026
0000-0003-2098-9621ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 40 · 2 first-author · 27 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An Energy-Efficient and End-to-End Keyword Spotting Processor Using 3D Broadcast Sparse Computing Array and Hybrid Sparsity Strategy PruningabstractKeyword spotting (KWS) has been widely applied in various low-power offline voice wake-up applications, with core requirements for low energy consumption, low cost, and low latency. Neural network (NN)-based KWS processors can achieve state-of-the-art (SOTA) accuracy, but NNs also introduce considerable computation and parameter overhead. NN accelerators based on model pruning and sparse computing can leverage the sparsity of activations and weights to reduce storage, energy consumption, and latency, albeit at the cost of increased hardware area and dense computation energy. However, compared to large NN models, KWS NN models typically have far fewer parameters (often less than a million), resulting in limited redundancy for pruning and thus low sparsity. Consequently, traditional model pruning and sparse computing methods are not suitable for KWS processor design. To address this, we propose: 1) a hybrid sparse pruning architecture (HSPA) that maps joint weight-activation pruning (JWAP) to weight-only pruning (WP), avoiding high hardware overhead JWAP-specific circuits while maintaining high sparsity; 2) a 3D broadcast sparse computing array (3D-BSCA), a novel 3D computing array supporting unstructured WP that uses a broadcast strategy to replace complex data routing and reduce hardware costs; and 3) a joint activation-index computation engine (JAICE) that designs and reuses a reconfigurable computing engine to generate zero-value indices, eliminating additional index computation hardware overhead. This work was implemented with layout and simulation in a 65nm CMOS process, achieving 91.9% accuracy and 1.37 μJ energy consumption on KWS tasks. Additionally, HSPA enabled a high sparsity of up to 55%, reducing energy consumption by 43.9%, while 3D-BSCA and JAICE collectively reduced hardware overhead by 40% compared to conventional sparse computing. Jianbiao Xiao, Hengxin Wang, Chiyu Zou, Yangning Hu, Jun Zhou 0017 |
ACM Great Lakes Symposium on VLSI | 5 |
| 2026 | TSA-P: A Triple-Sparsity-Aware Reconfigurable Few-Spikes-Neuron-Based SNN Processor
Aoyu Shen, Ruixin Mao, Jun Zhou 0017 |
ISCAS | 5 |
| 2026 | ERSRP: A 55nm 46.28 MPixels/(s·mm2) 104 FPS Efficient Real-Time Super-Resolution Processor with Layer-Fused Lightweight EngineabstractThis paper presents ERSRP, a 55nm Edge Real-Time Super-Resolution Processor that achieves 104 FPS at FHD resolution with a peak area efficiency of 46.28 MPixels/(s·mm2), 8.86× higher than the state-of-the-art. The processor is designed through a software-hardware co-optimization approach, addressing the low utilization, high memory demand, and workload imbalance challenges inherent in lightweight SR networks. At the algorithmic level, an Ultra-Lightweight Super-Resolution (ULSR) model is proposed that integrates depth-wise and point-wise separable blocks with a pixel-shuffle mechanism to achieve high-quality reconstruction (37.18 dB PSNR and 0.9581 SSIM on Set5) with only 5.62K parameters. At the hardware level, the ERSRP introduces a Lightweight Accelerated Engine (LAE) sup-porting a Layer Parallel Computing Scheme (LPCS) to improve lightweight operator throughput by 48.9%. A Point-wise Layer Fused Scheme (PLFS) further enhances utilization by 4.95× through inter-core and intra-core fusion without intermediate memory. Fabricated in a 55nm UMC CMOS process, the ERSRP achieves a throughput of 215.7 MPixels/s and supports 104 FPS real-time SR at FHD. Gaoxiang Wu, Liang Chang 0002, Jingke Wang, Zhicheng Hu, Xin Zhao 0044, Fengbin Tu, Jun Zhou 0017 |
ISCAS | 9 |
| 2026 | IDEA: A Real-Time Unified Accelerator for Dual-Task Image Restoration: Dehazing and Illumination Enhancement
Gaoxiang Wu, Jingke Wang, Liang Chang 0002, Jun Zhou 0017 |
ISCAS | 7 |
| 2025 | CREST: An Efficient Conjointly-trained Spike-driven Framework for Event-based Object Detection Exploiting Spatiotemporal DynamicsabstractEvent-based cameras feature high temporal resolution, wide dynamic range, and low power consumption, which are ideal for high-speed and low-light object detection. Spiking neural networks (SNNs) are promising for event-based object recognition and detection due to their spiking nature but lack efficient training methods, leading to gradient vanishing and high computational complexity, especially in deep SNNs. Additionally, existing SNN frameworks often fail to effectively handle multi-scale spatiotemporal features, leading to increased data redundancy and reduced accuracy. To address these issues, we propose CREST, a novel conjointly trained spike-driven framework to exploit spatiotemporal dynamics in event-based object detection. We introduce the conjoint learning rule to accelerate SNN learning and alleviate gradient vanishing. It also supports dual operation modes for efficient and flexible implementation on different hardware types. Additionally, CREST features a fully spike driven framework with a multi-scale spatiotemporal event integrator (MESTOR) and a spatiotemporal-IoU (ST-IoU) loss. Our approach achieves superior object recognition & detection performance and energy efficiency compared with state of-the-art SNN algorithms on three datasets, providing an efficient solution for event-based object detection algorithms suitable for SNN hardware implementation. Ruixin Mao, Aoyu Shen, Jun Zhou 0017 |
AAAI | 4 |
| 2025 | IRPE: Instance-level reconstruction-based 6D pose estimator
Le Jin, Guoshun Zhou, Zherong Liu, Yuanchao Yu, Jun Zhou 0017 |
Image Vis. Comput. | 7 |
| 2025 | VSLAM-BA: Algorithm and Hardware Co-Design for High Performance and Energy-Efficient Visual SLAM Backend Hardware AcceleratorabstractVisual Simultaneous Localization and Mapping (VSLAM) is a key localization technology for emerging applications such as autonomous driving and uncrewed aerial vehicles (UAVs). Compared with VSLAM frontend, VSLAM backend plays a more important role as it is employed to improve the localization accuracy. However, the VSLAM backend usually uses Bundle Adjustment (BA) as its core optimization method which is well-known for its large scale of problem construction, high computational complexity and high serialization of data processing, making it difficult to achieve high performance and energy efficiency on platforms such as CPUs or GPUs. Although there are some VSLAM backend accelerators proposed recently for addressing the above issues, they did not well exploit the data regularity and computational characteristics, resulting in limited performance/energy efficiency improvements or degraded accuracy. In this work, we propose VSLAM-BA which is a high performance and energy-efficient VSLAM backend accelerator with algorithm-hardware co-design. On the algorithm level, a keyframe-split-based Schur elimination scheme is proposed to reduce latency, power consumption and memory storage while maintaining accuracy. On the hardware level, a column-folding-based computing architecture is proposed to boost performance and energy efficiency. A loading-sensitive matrix-computing technique with an adaptive task scheduler is proposed to reduce the latency and energy consumption. Further, a recyclable computing technique with point-aware solver is proposed to reduce the memory and energy consumption. The experimental results show that the proposed VSLAM-BA achieves the highest performance (380 fps) and the highest energy efficiency (0.51 mJ per frame) with high accuracy and low memory storage, compared with the SOTA designs. The proposed accelerator can work with different VSLAM frontend for backend optimization of localization accuracy. Ye Liu 0011, Xiuyuan Qi, Shuang Hao 0005, Zili Huang, Neng Zhao, Ruixin Mao, Sixu Li, Ang Hu, Yu Long 0005, Shanshan Liu 0001, Jun Zhou 0017 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 16 |
| 2025 | An Ultra-High Performance and Scalable Optical Flow Hardware Accelerator Based on FPGA for Autonomous DrivingabstractOptical flow plays an extremely important role in the field of computer vision and extremely high real-time performance is required especially in autonomous driving. Traditional optical flow methods generally improve the accuracy of optical flow through the image pyramid technique. Nevertheless, the incorporation of the image pyramid elevates the computational complexity. Moreover, the data relationships between pyramid layers result in strong data dependencies, which renders it difficult to accelerate via parallel processing and makes it challenging to fulfill real-time demands in practical scenarios. To address this issue, in this paper, we propose an ultra-high performance and scalable optical flow hardware accelerator based on FPGA with several techniques, including an adaptive optical flow computation technique based on dynamic direction prediction to reduce computation without accuracy degradation, a highly scalable computing architecture with configurable numbers of PEs to improve the flexibility and hardware utilization under different hardware resource constraints, and a reconfigurable pyramid-layer pipeline technique to improve performance and reduce memory size. The proposed hardware accelerator was implemented and evaluated on a Xilinx FPGA ZCU104 achieving ultra-high performance (405 FPS) while maintaining high accuracy (AEE 1.02) compared with SOTA hardware accelerators. Ye Liu 0011, Shuang Hao 0005, Xiuyuan Qi, Zili Huang, Liang Chang 0002, Yu Long 0005, Jun Zhou 0017 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 11 |
| 2025 | PRADO: A Low-Latency and Energy-Efficient 6DoF Pose Refinement Accelerator With Domain-Specific ExplorationsabstractSix degrees of freedom (6DoF) pose estimation is a critical technique for applications involving humanoid robotics, autonomous driving, and virtual and augmented reality (VR/AR). Pose refinement plays a pivotal role in the 6DoF pose estimation, significantly enhancing accuracy by iteratively refining the initial pose derived from a single-shot pose estimation (SSPE) process. However, this iterative approach is time-and energy-consuming which is not suitable for resource-and power-constrained edge-devices. In this paper, we propose PRADO, an energy-efficient 6DoF pose refinement accelerator for edge applications. Several domain-specific explorations aimed at reducing processing latency and energy-consumption while maintaining high accuracy have been proposed, including an adaptive sparse correspondences sampling (ASCS) architecture to reduce redundant correspondences process, a hybrid static & dynamic pruning (HSDP) technique with co-designed hardware architecture to reduce computational complexity, and a cosine similarity-based adaptive early exit (CSAE2) technique to eliminate unnecessary iterations. Implemented and evaluated on the ZCU102 FPGA board, PRADO achieves the lowest processing latency of 61.2ms and energy consumption of 198.6mJ, and highest accuracy of 0.7112, outperforming state-of-the-art designs. Implemented with 28nm CMOS technology, PRADO achieves an even lower processing latency of 18.1ms and energy consumption of 12.6mJ. Haojie Wei, Le Jin, Junhao Zeng, Liwei Zhuang, Yu Long 0005, Jun Zhou 0017 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 10 |
| 2025 | An Energy-Efficient Block-Based Nonmaximum Suppression Engine for High-Parallel Postprocessing of Visual Object DetectionabstractNowadays, visual object detection (VOD) is widely used in many AI applications, such as autonomous driving, intelligent robotics, and smart surveillance. As an essential postprocessing step in VOD, nonmaximum suppression (NMS) is employed to generate bounding boxes as detection results. However, NMS is difficult to parallelize and computationally intensive, resulting in high processing latency and energy consumption. To address this issue, this brief proposes an energy-efficient block-based NMS engine that incorporates both algorithm- and hardware-level design techniques to improve processing speed and energy efficiency. These techniques include a block-NMS scheme, an adaptive hybrid sorting architecture (AHSA), and a reconfigurable pipeline-based block-NMS computation architecture. The proposed engine is implemented in 28-nm CMOS technology. Compared with the state-of-the-art designs, it achieves the highest performance (237.97 GOPS) and energy efficiency (8.71 TOPS/W), while delivering results fully equivalent to those of the original NMS. Yuchuan Gong, Haojie Wei, Hongtao Guo, Jiahao Zheng 0004, Qingyuan Hou, Zherong Liu, Jingxiao Zheng, Ye Liu 0011, Zhengning Wang, Jun Zhou 0017 |
IEEE Trans. Very Large Scale Integr. Syst. | 12 |
| 2025 | PIPECIM: Energy-Efficient Pipelined Computing-in-Memory Computation Engine With Sparsity-Aware TechniqueabstractComputing-in-memory (CIM) architecture has become a promising solution to improve the parallelism of the multiply-and-accumulation (MAC) operation for artificial intelligence (AI) processors. Recently, revived CIM engine partly relieves the memory wall issue by integrating computation in/with the memory. However, current CIM solutions still require large data movements with the increase of the practical neural network model and massive input data. Previous CIM works only considered computation without concern for the memory attribute, leading to a low memory computing ratio. This article presents a static-random access-memory (SRAM)-based digital CIM macro supporting pipeline mode and computation-memory-aware technique to improve the memory computing ratio. We develop a novel weight driver with fine-grained ping-pong operation, avoiding the computation stall caused by weight update. Based on our evaluation, the peak energy efficiency is 19.78 TOPS/W at the 22-nm technology node, 8-bit width, and 50% sparsity of the input feature map. Liang Chang 0002, Jingke Wang, Xin Zhao 0044, Wuyang Hao, Haining Tan, Yinhe Han 0001, Jun Zhou 0017 |
IEEE Trans. Very Large Scale Integr. Syst. | 11 |
| 2024 | Stellar: Energy-Efficient and Low-Latency SNN Algorithm and Hardware Co-Design with Spatiotemporal ComputationabstractThe brain-inspired Spiking Neural Network (SNN) has great potential to reduce energy consumption in AI applications. However, the state-of-the-art SNN algorithms focus on high accuracy and large sparsity by constructing complex neuron models with sparse spike generation, leading to low energy efficiency and high latency. The state-of-the-art SNN hardware designs are hard to exploit high data reuse and parallel processing dataflows due to the irregularity and time-dependency of the spikes. To address the above issues, in this work we propose STELLAR, an algorithm-hardware co-design framework exploiting rich spatiotemporal dynamics of the SNN for high energy efficiency and low latency while maintaining high accuracy. Firstly, based on the Few Spikes (FS) neuron, we propose few spikes backpropagation (FSBP) and its training flow with strong hardware awareness to adaptively train the deep SNN for a short time window and few spikes. The resulting sparse SNN enjoys rapid inference with few synaptic operations and competitive accuracy. Secondly, we propose a dedicated SNN architecture and spatiotemporal Row Stationary (stRS) dataflow to exploit large sparsity brought by the proposed algorithm for highly parallel and energy -efficient computation. Several techniques have been proposed to boost energy efficiency and speedup while maintaining accuracy, including the window-based parallel processing technique and the spatiotemporal encoding-based computation architecture. The experimental results show that 1) on the algorithm level, STELLAR outperforms the state-of-the-art SNN models with significantly fewer spikes and shorter time window on both static and neuromorphic datasets with higher or comparable accuracy; 2) on the architecture level, compared with several SOTA SNN hardware designs, STELLAR achieves up to 8.1 × energy efficiency and 7.1 × speedup. Ruixin Mao, Ye Liu 0011, Jun Zhou 0017 |
HPCA | 5 |
| 2024 | An Ultra-Low Power Time-Domain based SNN Processor for ECG ClassificationabstractWearable devices for ECG arrhythmia detection based on artificial neural networks (ANN) are very popular. However, the energy consumption of electrocardiogram (ECG) processing in ANN has become one of the most critical factors. One solution is using a spiking neural network (SNN), effectively reducing power consumption and improving energy efficiency. Nevertheless, the inevitable membrane potential storage and accumulation of SNN result in significant energy and area overheads. This paper proposes a time domain (TD) based SNN processor for ECG classification. We propose a novel memory delay unit (MDU), part of the memory delay line (MDL), to store and accumulate membrane potential. With this method, power consumption can be significantly reduced. Also, we propose a wave generator that works with MDL to maximize computing efficiency. Compared with digital neurons, our proposed TD neurons reduce power consumption by 32.5% and achieve a classification accuracy of 96.8%. It is very suitable for arrhythmia detection wearable devices. Haodong Fan, Liang Chang 0002, Junlu Zhou, Shuisheng Lin, Jun Zhou 0017 |
ISCAS | 6 |
| 2024 | An FPGA-based Ultra-High Performance and Scalable Optical Flow Hardware Accelerator for Autonomous DrivingabstractOptical flow plays an extremely important role in the field of computer vision and extremely high real-time performance is required especially in autonomous driving. Traditional optical flow methods generally improve the accuracy of optical flow through the image pyramid technique. However, the introduction of the image pyramid increases the computational complexity. Additionally, the data relationships between pyramid layers result in strong data dependencies, making it difficult to accelerate using parallel processing, and it is challenging to meet real-time requirements in practical scenarios. To address this issue, in this paper, we propose an FPGA-based ultra-high performance and scalable optical flow hardware accelerator with several techniques, including an adaptive optical flow computation technique based on dynamic direction prediction to reduce computation without accuracy degradation, a highly scalable computing architecture with configurable numbers of PEs to improve the flexibility and hardware utilization under different hardware resource constraints, and a reconfigurable pyramid-layer pipeline technique to improve performance and reduce memory size. The proposed hardware accelerator was implemented and evaluated on a Xilinx FPGA ZCU104 achieving ultra-high performance (405 FPS) while maintaining high accuracy (AEE 0.64). Ye Liu 0011, Shuang Hao 0005, Zili Huang, Xiuyuan Qi, Yu Long 0005, Jun Zhou 0017 |
ISCAS | 10 |
| 2024 | General Purpose Deep Learning Accelerator Based on Bit InterleavingabstractAlong with the rapid evolution of deep neural networks, the ever-increasing complexity imposes formidable computation intensity on the hardware accelerator. In this paper, we propose a novel computing philosophy called “bit interleaving” and the associate accelerator couple called “Bitlet” and Bitlet-X to maximally exploit the bit-level sparsity. Apart from the existing bit-serial/parallel accelerators, Bitlet leverages the abundant “sparsity parallelism” in the parameters to enforce the inference acceleration. Bitlet is versatile by supporting diverse precisions on a single platform, including floating-point 32 and fixed-point from 1b to 24b. The versatility enables Bitlet feasible for both efficient inference and training. Besides, by updating the key compute engine in the accelerator, Bitlet-X could furthermore improve the peak power consumption and efficiency for the inference-only scenario, with competitive accuracy. Empirical studies on 12 domain-specific deep learning applications highlight the following results: (1) up to 81×/21× energy efficiency improvement for training/inference over recent high-performance GPUs; (2) up to 15×/8× higher speedup/efficiency over state-of-the-art fixed-point accelerators; (3) 1.5mm2 area and scalable power consumption from 570mW (fp32) to 432mW (16b) and 365mW (8b) @28nm TSMC; (4) 1.3× improvement of the peak power efficiency for the Bitlet-X over Bitlet; (5) highly configurable justified by the ablation and sensitivity studies. Liang Chang 0002, Xin Zhao 0044, Zhicheng Hu, Jun Zhou 0017, Xiaowei Li 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2024 | HDSuper: High-Quality and High Computational Utilization Edge Super-Resolution Accelerator With Hardware-Algorithm Co-Design TechniquesabstractSuper-resolution (SR) techniques have been employed to construct high-definition images from low-quality images. Various neural networks have demonstrated excellent image-reconstruction quality in SR accelerators. However, deploying SR networks on edge devices is limited by resources and power consumption induced by significant algorithm parameters, computation complexity, and external memory accesses. This work explores the hardware algorithm co-design techniques to provide an end-to-end platform with a lightweight super-resolution network (LSR) and an efficient, high-quality SR accelerator HDSuper. For algorithm design, the improved depth-wise separable convolution and pixelshuffle layers are developed to reduce network size and computation complexity by considering the hardware constraints. Also, the improved channel attention (CA) blocks enhance the image reconstruction quality. For hardware accelerator design, we design a unified computing core (UCC) combined with an efficient flattening-and-allocation (F-A) mapping strategy to support various operators with high computational utilization. In addition, we design the patch computing scheme to reduce the external memory access of the hardware architecture. Based on the evaluation, the proposed algorithm achieves high-quality image reconstruction with$37.44dB$PSNR. Finally, the FPGA demonstration and ASIC layout under UMC 55nm are achieved with low power consumption ($2.08 W$and$152 mW$) under the lowest hardware resources compared to the state-of-the-art works. Xin Zhao 0044, Liang Chang 0002, Dongqi Fan, Zhicheng Hu, Ting Yue, Fengbin Tu, Jun Zhou 0017 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2024 | IPOCIM: Artificial Intelligent Architecture Design Space Exploration With Scalable Ping-Pong Computing-in-Memory MacroabstractComputing-in-memory (CIM) architecture has become a possible solution to designing an energy-efficient artificial intelligent processor. Various CIM demonstrators indicated the computing efficiency of CIM macro and CIM-based processors. However, previous studies mainly focus on macro optimization and low CIM capacity without considering the weight update strategy of CIM architecture. The artificial intelligence (AI) processor with a CIM engine practically induces issues, including updating memory data and supporting different operators. For instance, AI-oriented applications usually contain various weight parameters. The weight stored in the CIM architecture should be reloaded for the considerable gap between the capacity of CIM and growing weight parameters. The computation efficiency of the CIM architecture is reduced by the weight updating and waiting. In addition, the natural parallelism of CIM leads to the mismatch of various convolution kernel sizes in different networks and layers, which reduces hardware utilization efficiency. In this work, we develop a CIM engine with a ping-pong computing strategy as an alternative to typical CIM macro and weight buffer, hiding the data update latency and improving the data reuse ratio. Based on the ping-pong engine, we propose a flexible CIM architecture adapting to different sizes of neural networks, namely, intelligent pong computing-in memory (IPOCIM), with a fine-grained data flow mapping strategy. Based on the evaluation, IPOCIM can achieve a 1.27–$6.27\times $performance and 2.34–$5.30\times $energy efficiency improvement compared to the state-of-the-art works. Liang Chang 0002, Xin Zhao 0044, Ting Yue, Shuisheng Lin, Jun Zhou 0017 |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2023 | HDSuper: Algorithm-Hardware Co-design for Light-weight High-quality Super-Resolution AcceleratorabstractSuper-resolution (SR) networks have been gradually applied to embedded devices with good-quality image reconstruction. However, the hardware performance and power efficiency are limited by a large number of algorithm parameters, computation complexity, and hardware resources, obstructing the development of a high-quality SR accelerator. This paper proposes an end-to-end platform with a lightweight super-resolution network (LSR) and an efficient, high-quality super-resolution architecture HDSuper, to perform algorithm-hardware co-design for the SR accelerator. For algorithm design, we employ depth-wise separable convolution and pixelshuffle to reduce network size and computation complexity by considering the hardware constraints. For hardware design, we provide a unified computing core (UCC) combined with an efficient flattening-and-allocation (F-A) mapping strategy to support various operators with high computational utilization. We adopt the patch training method to reduce the external memory access of the hardware architecture. Based on the evaluation, the proposed algorithm achieves high-quality image reconstruction with 37.44dB PSNR. Finally, we implement the image reconstruction in FPGA demonstration, achieving high-quality image reconstruction with 2.08W power consumption under the lowest hardware resources compared to the state-of-the-art works. Liang Chang 0002, Xin Zhao 0044, Dongqi Fan, Zhicheng Hu, Jun Zhou 0017 |
DAC | 5 |
| 2023 | A FPGA-Based Iterative 6DoF Pose Refinement Processing Unit for Fast and Energy-Efficient Pose Estimation in Picking RobotsabstractFast and energy-efficient 6D pose estimation is essential for robotic applications, especially for picking robots in industrial scene. Introducing iterative pose refinement in the final stage of pose estimation pipeline can effectively improve the precision. However, this procedure can be very time consuming due to iterative convolutional neural network (CNN) inference on resource and power constrained platforms. In this paper, we propose a FPGA-based iterative pose refinement processing unit that achieves fast and energy-efficient pose estimation for picking robots. The design and implementation are based on a Xilinx Zynq UltraScale+ MPSoC. Our experimental results demonstrate that the proposed FPGA-Based processing unit is 21.78 times faster and 23.89 times more energy-efficient compared with the baseline, significantly enhances the speed and energy efficiency of pose refinement. The evaluation result on the datasets shows little accuracy drop compared to the baseline implementation. Le Jin, Guoshun Zhou, Liang Chang 0002, Jun Zhou 0017 |
IECON | 10 |
| 2023 | TDPRO: Time-Domain-Based Computing-in Memory Engine for Ultra-Low Power ECG ProcessorabstractFor the wearable biomedical signal detection, both high accuracy and low-power consumption are critical requirements. Various works have employed the neural network to improve the detecting accuracy and develop the biomedical processor. However, the biomedical processor with neural network engine contains massive data movements and large data buffers. One solution is the computing-in memory (CIM) architecture, which locates more data near the computing engine to reduce data movements. In traditional CIM-based solution, the detecting accuracy and power consumption is difficult to be optimized simultaneously, where the accuracy should be satisfied for the detection. To date, the time-domain computing engine have been developed to employ both digital and time domain computation. In this work, we present a high-precision time-domain engine to perform 8-bit multiplication and addition operation for the biomedical signal detection. With the high precision time-domain engine, we develop a CIM-based neural-network processor, namely TDPRO, to perform the detection of arrhythmia. In addition, we develop TD-zero-jumping (TDJ) and idle-shutdown (ISD) techniques according to signal features and data mapping strategy, further optimizing the power consumption. Based on our evaluation, the TD-based 8-bit mulitply-accumulation operation is robust, without declining the accuracy of biomedical signal detection. We design a ECG processor with the proposed TDPRO architecture, which obtains 98.60% high accuracy and 75.7% power saving compared to the recent the state-of-the-art study. Liang Chang 0002, Siqi Yang 0002, Zhiyuan Chang, Haodong Fan, Junlu Zhou, Jun Zhou 0017 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2023 | A High Accuracy and Low Power CNN-Based Environmental Sound Classification ProcessorabstractThe environmental sound classification (ESC) has attracted increasing attention as the environmental sound contains a wealth of information that can be used to detect particular events. However, so far, most of the existing work in ESC still remains in the stage of algorithm design and the design of ESC processor has not been thoroughly investigated. The existing ESC processor designs have issues in meeting low power consumption and high accuracy simultaneously due to the lack of joint-optimization between algorithm and hardware, and very few work has demonstrated a complete ESC system containing all the necessary modules. In this work, a high accuracy and low power CNN-based ESC processor has been proposed, featuring: 1) a big-small CNN-based reconfigurable ESC processing hardware architecture to reduce the power consumption and hardware overhead while maintaining high classification accuracy. 2) a Mel feature adaptation engine reusing the neural network processing unit to further reduce the power consumption. 3) an event-driven ESC processing technique to reduce the inference time and the power consumption. The design has been implemented on a Kintex-7 FPGA and achieves low power consumption of 0.313W with high accuracy of 84.5% for the ESC-50 dataset, outperforming other state-of-the-art ESC processors. Lujie Peng, Junyu Yang, Longke Yan, Xiben Jiao, Jianbiao Xiao, Liang Chang 0002, Yu Long 0005, Jun Zhou 0017 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 10 |
| 2023 | ADAS: A High Computational Utilization Dynamic Reconfigurable Hardware Accelerator for Super ResolutionabstractSuper-resolution (SR) based on deep learning has obtained superior performance in image reconstruction. Recently, various algorithm efforts have been committed to improving image reconstruction quality and speed. However, the inference of SR contains huge amounts of computation and data access, leading to low hardware implementation efficiency. For instance, the up-sampling with the deconvolution process requires considerable computation resources. In addition, the sizes of output feature maps of several middle layers are extraordinarily large, which is challenging to optimize, causing serious data access issues. In this work, we present an all-on-chip hardware architecture based on the deconvolution scheme and feature map segmentation strategy, namely ADAS, where all the generated data by the middle layers are buffered on-chip to avoid large data movements between on- and off-chip. In ADAS, we develop a hardware-friendly and efficient deconvolution scheme to accelerate the computation. Also, the dynamic reconfigurable process element (PE) combined with efficient mapping is proposed to enhance PE utilization up to nearly 100% and support multiple scaling factors. Based on our experimental results, ADAS demonstrates real-time image SR and better image reconstruction quality with PSNR (37.15 dB ) and SSIM (0.9587). Compared to baseline and validated with the FPGA platform, ADAS can support scaling factors of 2, 3, and 4, achieving 2.68 ×, 5.02 ×, and 8.28 × speedup. Liang Chang 0002, Xin Zhao 0044, Jun Zhou 0017 |
ACM Trans. Reconfigurable Technol. Syst. | 3 |
| 2022 | An energy-efficient seizure detection processor using event-driven multi-stage CNN classification and segmented data processing with adaptive channel selectionabstractRecently wearable EEG monitoring devices with seizure detection processor using convolutional neural network (CNN) have been proposed to detect the seizure onset of patients in real time for alert or stimulation purpose. High energy efficiency and accuracy are required for the seizure detection processor due to the tight energy constraint of wearable devices. However, the use of CNN and multi-channel processing nature of seizure detection result in significant energy consumption. In this work, an energy-efficient seizure detection processor is proposed, featuring multi-stage CNN classification, segmented data processing and adaptive channel selection to reduce the energy consumption while achieving high accuracy. The design has been fabricated and tested using a 55nm process technology. Compared with several state-of-the-art designs, the proposed design achieves the lowest energy per classification (0.32 μJ) with high sensitivity (97.78%) and low false positive rate per hour (0.5). Jiahao Liu 0006, Zirui Zhong, Hui Qiu, Jianbiao Xiao, Jiajing Fan, Zhaomin Zhang, Sixu Li, Siqi Yang 0002, Weiwei Shan, Shuisheng Lin, Liang Chang 0002, Jun Zhou 0017 |
DAC | 14 |
| 2022 | TDPRO: Ultra-low Power ECG Processor with High-Precision Time-Domain Computing EngineabstractIn wearable biomedical signal detection, the low-power consumption is a critical requirement. However, the process of biomedical signal detection with traditional neural-network processor is uneconomical for large data movements. A typical solution is the near memory computing (NMC) method, locating more data near the computing engine to save energy, where the detecting accuracy and power consumption is difficult to be optimized simultaneously. In addition, a suitable computing engine is needed to match both power and computation budget. In this work, we combine the NMC-based ECG processor equipped with a high-precision time-domain engine to perform the detection of arrhythmia, namely TDPRO. The proposed TDPRO supports high precision multiplication and addition operation with 8-bit input and weight parameters. Also, we propose TD-zero-jumping and idle-shutdown technique to further reduce 63%$\sim$ 91% power consumption of the time-domain engine. The error rate of 8-bit MAC operation in the TDPRO is 1.18%, which is suitable for the ECG detection. Liang Chang 0002, Siqi Yang 0002, Huinan Wang, Jianbo Xiao, Xin Zhao 0044, Shuisheng Lin, Jun Zhou 0017 |
ISCAS | 7 |
| 2022 | ReverSearch: Search-based energy-efficient Processing-in-Memory ArchitectureabstractRecent development of the processing-in-memory (PIM) architecture has demonstrated high efficiency by reducing data movements. However, the performance of the conventional PIM architecture is limited by several issues, including frequent bit-line operations, complicated control of data flow, and massive inter-macro data movements. In addition, both analog- and digital-PIM solutions have obstacles to meet requirement of high-precision computation. In this work, we explore the tradeoff between data movement and energy efficiency of PIM architecture. We develop a PIM architecture, namely ReverSearch, to accelerate multiple-and-accumulate operation, equipped with reverse searching engine and look up table operations. Also, the corresponding data mapping and data flow methods are provided to improve the performance of the ReverSearch architecture. Based on our evaluation, ReverSearch improves the energy efficiency by 17.26 × and 3.68 ×, compared to the baseline of LUT-Cache [1] and LAcc [2]. Weihang Li, Liang Chang 0002, Jiajing Fan, Xin Zhao 0044, Hengtan Zhang, Shuisheng Lin, Jun Zhou 0017 |
ISCAS | 7 |
| 2022 | LCSED: A low complexity CNN based SED model for IoT devices
Mingxue Yang, Lujie Peng, Yujiang Wang 0003, Zhenyuan Zhang 0004, Zhengxi Yuan, Jun Zhou 0017 |
Neurocomputing | 7 |
| 2022 | ULSED: An ultra-lightweight SED model for IoT devices
Lujie Peng, Junyu Yang, Jianbiao Xiao, Mingxue Yang, Yujiang Wang 0003, Haojie Qin, Xiaorong Li, Jun Zhou 0017 |
J. Parallel Distributed Comput. | 8 |
| 2022 | MobileSP: An FPGA-Based Real-Time Keypoint Extraction Hardware Accelerator for Mobile VSLAMabstractKeypoint extraction is a key technique for Visual Simultaneous Localization and Mapping (VSLAM). Recently, Convolutional Neural Network (CNN) has been used in the keypoint extraction for improving the accuracy. As one of the state-of-the-art CNN based keypoint extraction techniques, the SuperPoint ranked top in the CVPR2020 image matching challenge. However, the use of complex CNN makes it difficult to meet the real-time performance on a mobile platform with limited resource such as mobile robots and wearable Augmented Reality (AR) devices. In this work, based on the SuperPoint, we proposed an FPGA-based real-time keypoint extraction hardware accelerator through algorithm-hardware co-design for mobile VSLAM applications, which is named as MobileSP. Several algorithm and hardware level design techniques have been proposed to reduce the computation and improve the processing speed while maintaining high accuracy, including a partially shared detection & description encoding architecture, a pre-sorting based Non-Maximum Suppression (NMS) engine and a software-hardware hybrid pipeline computing technique. The design has been implemented and evaluated on a ZCU104 FPGA board. It achieves real-time performance of 42 fps with low Absolute Trajectory Error (ATE) of 1.82 cm simultaneously, outperforming several state-of-the-art designs. Ye Liu 0011, Xiuyuan Qi, Liang Chang 0002, Yu Long 0005, Jun Zhou 0017 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 8 |
| 2022 | ULECGNet: An Ultra-Lightweight End-to-End ECG Classification Neural NetworkabstractECG classification is a key technology in intelligent electrocardiogram (ECG) monitoring. In the past, traditional machine learning methods such as support vector machine (SVM) and K-nearest neighbor (KNN) have been used for ECG classification, but with limited classification accuracy. Recently, the end-to-end neural network has been used for ECG classification and shows high classification accuracy. However, the end-to-end neural network has large computational complexity including a large number of parameters and operations. Although dedicated hardware such as field-programmable gate array (FPGA) and application-specific integrated circuit (ASIC) can be developed to accelerate the neural network, they result in large power consumption, large design cost, or limited flexibility. In this work, we have proposed an ultra-lightweight end-to-end ECG classification neural network that has extremely low computational complexity (∼8.2k parameters & ∼227k multiplication/addition operations) and can be squeezed into a low-cost microcontroller (MCU) such as MSP432 while achieving 99.1% overall classification accuracy. This outperforms the state-of-the-art ECG classification neural network. Implemented on MSP432, the proposed design consumes only 0.4 mJ and 3.1 mJ per heartbeat classification for normal and abnormal heartbeats respectively for real-time ECG classification. Jianbiao Xiao, Jiahao Liu 0006, Huanqi Yang, Ning Wang 0070, Zhen Zhu 0005, Yu Long 0005, Liang Chang 0002, Jun Zhou 0017 |
IEEE J. Biomed. Health Informatics | 11 |
| 2022 | Multi-Dimensional Feature Combination Method for Continuous Blood Pressure Measurement Based on Wrist PPG SensorabstractThe cuff-less blood pressure (BP) monitoring method based on photoplethysmo- gram (PPG) makes it possible for long-term BP monitoring to prevent and treat cardiovascular and cerebrovascular events. In this paper, a portable BP prediction system based on feature combination and artificial neural network (ANN) is implemented. The robustness of the model is improved from three aspects. Firstly, an adaptive peak extraction algorithm was used to improve the accuracy of peaks and troughs detection. Secondly, multi-dimensional features were extracted and fused, including three groups of PPG-based features and one group of demographics-based features. Finally, a two-layer feedforward artificial neural networks algorithm was used for regression. Thirty-three subjects distributed in the three BP groups were recruited. The proposed method passed the European Society of Hypertension International Protocol revision 2010 (ESP-IP2). Experimental results show that the proposed method exhibits good accuracy for a diverse population with an estimation error of -0.07 ± 4.47 mmHg for SBP and 0.00 ± 3.61 mmHg for DBP. Moreover, the model tracked the BP of two subjects for half a month, laying the foundation work for daily BP monitoring. This work will contribute to the long-term wellness management and rehabilitation process, enabling timely detection and improvement of the user's physical health. Pan Yao, Ning Xue, Siyuan Yin, Changhua You, Yusen Guo, Yi Shi 0013, Tiezhu Liu, Jun Zhou 0017, Jianhai Sun, Chunxiu Liu |
IEEE J. Biomed. Health Informatics | 9 |
| 2021 | Energy-Efficient Spin-Orbit Torque MRAM Operations for Neural Network ProcessorabstractEmerging energy-efficient neural network processor is a promising hardware design to accelerate neural network algorithms with high performance and low power consumption. Typically, static random-access memory (SRAM) is employed to develop large buffers using in the processor. The bit cell of SRAM contains six transistors, leading to low density and large leakage current. In particular, several AI processors need multiple port and transfer-based SRAMs, which decrease the density and increase the power consumption. Recently, emerging spin-orbit torque magnetic random-access memory (SOT-MRAM) becomes a possible solution to replace the SRAM as working memory. However, more operations should be supported by the SOT- MRAM to provide sufficient functions, such as multiple-port memory, transpose memory, data-streaming operations. In this paper, we develop the working memory of neural network processor with SOT-MRAM to build the design library including the transpose operations, multiple-port memory, and data-streaming based buffer arrays. Equiped with those operations provided by SOT-MRAM, we can build high performance and energy-efficient neural network processors. Liang Chang 0002, Zixuan Zhu 0001, Zhen Zhu 0005, Siqi Yang 0002, Weihang Li, Jun Zhou 0017 |
ISCAS | 6 |
| 2021 | Energy-efficient computing-in-memory architecture for AI processor: device, circuit, architecture perspective
Liang Chang 0002, Zhaomin Zhang, Jianbiao Xiao, Zhen Zhu 0005, Weihang Li, Zixuan Zhu 0001, Siqi Yang 0002, Jun Zhou 0017 |
Sci. China Inf. Sci. | 10 |
| 2021 | Keyword spotting techniques to improve the recognition accuracy of user-defined keywords
Mingxue Yang, Zhengxi Yuan, Jun Zhou 0017 |
Neural Networks | 6 |
| 2021 | A Fast and Energy-Efficient SNN Processor With Adaptive Clock/Event-Driven Computation Scheme and Online LearningabstractIn the recent years, the spiking neural network (SNN) has attracted increasing attention due to its low energy consumption and online learning potential. However, the design of SNN processor has not been thoroughly investigated in the past, resulting in limited performance and energy consumption. In this work, a fast and energy-efficient SNN processor with adaptive clock/event-driven computation scheme and online learning capability has been proposed. Several techniques have been proposed to reduce the computation time and energy consumption, including Adaptive Clock- and Event-Driven Computing Scheme, Neighboring PE Borrowing Technique, Compressed Spike Routing Technique and Reconfigurable PE for Inference and Learning. Implemented on a Virtex-7 FPGA, the proposed design achieves computation time of 3.15 ms/image, inference energy consumption of$0.028~\mu $J/synapse/image and online learning energy consumption of$0.297~\mu $J/synapse/image for the MNIST 10-class dataset, which outperform several state-of-the-art SNN processors. The proposed SNN processor is suitable for real-time and energy-constrained applications. Sixu Li, Zhaomin Zhang, Ruixin Mao, Jianbiao Xiao, Liang Chang 0002, Jun Zhou 0017 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2020 | PRISM: Energy-Efficient Polymorphic Operation Based on Spin-Orbit Torque Memory for Reconfigurable ComputingabstractEmerging Non-Volatile Memories (NVMs) including resistive RAM (ReRAM), phase-change memory (PCM), and magnetic RAM (MRAM), have opened up new pathways for the NVM-based reconfigurable computing. Those NVMs technologies can achieve significant energy-efficient computational operations with only minor modification of the peripheral circuits. However, the supported operations are limited by the array structure and low energy-efficiency of implementing the computation using the memory array. In this paper, the Spin Orbit torque-MRAM based polymorphic circuits are proposed to support the reconfigurable computation for reducing the power consumption and improving the functionalities of the single memory array. With the high speed and energy-efficiency write operation, the proposed memory array support both read-out and write-in reconfigurable operations. Liang Chang 0002, Zhaohao Wang, Youguang Zhang, Weisheng Zhao 0001, Jun Zhou 0017 |
ISCAS | 6 |
| 2020 | Design and Characterization of Radiation-Hardened MCU for Space Application using Error Correction SRAM and Glitch Removal Clock Buffer CellabstractHigh-energy environmental radiation particles may affect the System on Chip (SoC) which causes errors in the combinational logic, sequential logic, and even in the clock network. In the latter, they appear as clock glitches that propagate, and eventually, incorrectly latch all the sequential circuits, such as flip-flops and latches attached to the clock node concerned. The impact on-chip functionality is usually fatal. In this work, we propose a low-cost adaptive clock glitch removal circuitry for radiation-resilient clock networks. The proposed technique can be adjusted based on the actual clock glitch profile, to ensure that the error in the clock network is removed, while the impact on the clock signal itself is minimized. It is fully synthesizable, and can thus be incorporated in a conventional digital flow. Using the proposed clock buffer cell, the implemented ARM Cortex M0 shows 85% error reduction when exposed to radiation with minimum area overhead. Anh-Tuan Do, Tony Tae-Hyoung Kim, Xin Liu 0015, Jun Zhou 0017 |
ISCAS | 4 |
| 2020 | Corrigendum to"Approximate error detection-correction for efficient adaptive voltage Over-Scaling"[Integration 63 (2018) 220-231]
Roberto Giorgio Rizzo, Andrea Calimera, Jun Zhou 0017 |
Integr. | 3 |
| 2020 | TSE-CNN: A Two-Stage End-to-End CNN for Human Activity RecognitionabstractHuman activity recognition has been widely used in healthcare applications such as elderly monitoring, exercise supervision, and rehabilitation monitoring. Compared with other approaches, sensor-based wearable human activity recognition is less affected by environmental noise and therefore is promising in providing higher recognition accuracy. However, one of the major issues of existing wearable human activity recognition methods is that although the average recognition accuracy is acceptable, the recognition accuracy for some activities (e.g., ascending stairs and descending stairs) is low, mainly due to relatively less training data and complex behavior pattern for these activities. Another issue is that the recognition accuracy is low when the training data from the test subject are limited, which is a common case in real practice. In addition, the use of neural network leads to large computational complexity and thus high power consumption. To address these issues, we proposed a new human activity recognition method with two-stage end-to-end convolutional neural network and a data augmentation method. Compared with the state-of-the-art methods (including neural network based methods and other methods), the proposed methods achieve significantly improved recognition accuracy and reduced computational complexity. Shuisheng Lin, Ning Wang 0070, Guanghai Dai, Yuxiang Xie, Jun Zhou 0017 |
IEEE J. Biomed. Health Informatics | 6 |
| 2019 | A High Throughput and Energy-Efficient Retina-Inspired Tone Mapping ProcessorabstractThis paper presents a high throughput and energy-efficient retina inspired tone mapping processor. Several hardware design techniques have been proposed to achieve high throughput and high energy efficiency, including data partition based parallel processing with S-shape sliding, adjacent frame feature sharing, multi-layer convolution pipelining and convolution filter compression with zero skipping convolution. The proposed processor has been implemented on a Xilinx's Virtex7 FPGA for demonstration. It is able to achieve a throughput of 189 frames per second for 1024*768 RGB images with 819 mW. Compared with several state-of-the-art tone mapping processors, the proposed processor achieves higher throughput and energy efficiency. It is suitable for high-speed and energy-constrained video enhancement applications such as autonomous vehicle and drone monitoring. Xiaoqiang Xiang, Yuxiang Xie, Jun Zhou 0017 |
FCCM | 6 |
| 2019 | Editorial TVLSI Positioning - Continuing and Accelerating an Upward TrajectoryabstractI. VLSI Systems: A Glance Into The Last Decades Since their inception in 1970s, VLSI systems have enabled several new technological capabilities and made them accessible to an unceasingly wider range of users, reaching a scale that has been exponentially increasing over the decades[1](seeFig. 1). Relentless integration of more complex systems has driven such remarkable evolution, as made possible by the inexorable miniaturization. As shown inFig. 1, more functionality has been crammed in a consistently smaller form factor, as exemplified by the physical volume shrinking of computers by 100 X/decade[2],[3]. At the same time, the energy per task has been decreasing at 10–100 X/decade, as shown inFig. 2, for several systems and system-on-chip subsystems[4]. This allowed packing more capabilities into the same power envelope, as generally observed in the electronic systems, even before the advent of the integrated circuit[5]. Massimo Alioto, Magdy S. Abadir, Tughrul Arslan, Chirn Chye Boon, Andreas Peter Burg, Chip-Hong Chang, Meng-Fan Chang, Yao-Wen Chang, Poki Chen, Pasquale Corsonello, Paolo Crovetti, Shiro Dosho, Rolf Drechsler, Ibrahim M. Elfadel, Ruonan Han 0001, Masanori Hashimoto, Chun-Huat Heng, Deuk Hyoun Heo, Tsung-Yi Ho, Houman Homayoun, Yuh-Shyan Hwang, Ajay Joshi, Rajiv V. Joshi, Tanay Karnik, Chulwoo Kim, Tony Tae-Hyoung Kim, Jaydeep P. Kulkarni, Volkan Kursun, Yoonmyung Lee, Hai Li 0001, Huawei Li 0001, Prabhat Mishra 0001, Baker Mohammad, Mehran Mozaffari Kermani, Makoto Nagata, Koji Nii, Partha Pratim Pande, Bipul Chandra Paul, Vasilis F. Pavlidis, José Pineda de Gyvez, Ioannis Savidis, Patrick Schaumont, Fabio Sebastiano, Anirban Sengupta 0003, Mingoo Seok, Mircea R. Stan, Mark Tehranipoor, Aida Todri, Marian Verhelst, Valerio Vignoli, Xiaoqing Wen, Jiang Xu 0001, Wei Zhang 0012, Zhengya Zhang, Jun Zhou 0017, Mark Zwolinski, Stacey Weber |
IEEE Trans. Very Large Scale Integr. Syst. | 55 |
| 2018 | A FPGA-based RO PUF with LUT-Based Self-Compare Structure and Adaptive Counter Time Period TuningabstractPUF can be used for IoT device authentication. This paper proposes a novel FPGA-based RO PUF with improved uniqueness and reliability. Firstly, the proposed PUF improves the uniqueness by LUT-based self-compare structure, which reduces the delay bias from systematic variations without special constraint on place & route and selection of challenge response pairs. Secondly, the proposed PUF improves the reliability by adaptive counter time period tuning based on real-time measured response stability. Implemented on Xilinx Spartan-6 FPGA, the proposed PUF shows better uniqueness and reliability than several state-of-the-arts FPGA-based RO PUFs. Jiayan Gan, Jun Zhou 0017, Ning Wang 0070 |
ISCAS | 2 |
| 2018 | Approximate Error Detection-Correction for efficient Adaptive Voltage Over-Scaling
Roberto Giorgio Rizzo, Andrea Calimera, Jun Zhou 0017 |
Integr. | 3 |
| 2017 | Early bird sampling: A short-paths free error detection-correction strategy for data-driven VOSabstractRazor is a milestone in the field of Error Detection&Correction strategies for low-power operation. Despite the impressive level of maturity, its application on circuits other than pipelined processors still remains an open issue. Firstly, the error detection mechanism relies on special flip-flops (FFs), the Razor-FFs, whose use imposes heavy hold-time fixing and large circuit area/power overheads; secondly, the error correction is performed through instruction replay, a practice that is not available (or very expensive to implement) in generic circuits. This work introduces Early Bird Sampling (EBS), a Razor variant that applies to low-power sequential circuits. The EBS allows to (i) solve the problem of short-path races bypassing tedious holdtime fixing design stages, (ii) reduce design overhead exploiting a local logic-masking mechanism for error correction. As a key feature, EBS enables Data-Driven Voltage Over-Scaling (DD-VOS), an aggressive dynamic voltage scaling strategy particularly suited for ultra-low power error-resilient applications. Simulation runs on a representative set of circuits provide a fair comparison with a standard Razor strategy. The collected results show EBS reduces area overheads (3.6% against 71.6% for Razor) and improves the voltage scaling profile achieving lower energy-per-operation (savings w.r.t. Razor range from 19.1% to 53.1%). Roberto Giorgio Rizzo, Valentino Peluso, Andrea Calimera, Jun Zhou 0017, Xin Liu 0015 |
VLSI-SoC | 4 |
| 2015 | An Area- and Energy-Efficient FIFO Design Using Error-Reduced Data Compression and Near-Threshold Operation for Image/Video ApplicationsabstractMany image/video processing algorithms require FIFO for filtering. The FIFO size is proportional to the length of the filters and input data width, causing large area and power consumption. We have proposed an energy- and area-efficient FIFO design for image/video applications through FIFO with error-reduced data compression (FERDC) and near-threshold operation. On architecture level, FERDC technique is proposed to reduce the size and power consumption of the FIFO by utilizing the spatial correlation between neighboring pixels and performing error-reduced data compression together with quantization to minimize the mean square error (MSE). On circuit level, near-threshold operation is adopted to achieve further power reduction while maintaining the required performance. To demonstrate the proposed FIFO, it has been implemented using a 0.18-μm CMOS process technology. The implementation covers different FIFO length, including 128, 256, 512, and 1024. The experimental results show that the proposed FIFO operating at 0.5 V and 28.57 MHz achieves up to 99%, 65%, and 34.91% reduction in dynamic power, leakage power, and area, respectively, with a small MSE of 2.76, compared with the conventional FIFO design. The proposed FIFO can be applied to a wide range of image/video signal processing applications to achieve high area and energy efficiency. Seyed Mohammad Ali Zeinolabedin, Jun Zhou 0017, Xin Liu 0015, Tony Tae-Hyoung Kim |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2014 | An area- and power-efficient FIFO with error-reduced data compression for image/video processingabstractFiltering is a key component of many digital image/video processing algorithms. It often requires FIFO to temporarily buffer the pixels data for later usage. The FIFO size is proportional to the length of the filters and input data width, causing large area and power consumption. This paper presents a technique named FIFO with error-reduced data compression (FERDC) to reduce the FIFO size for various filters. The proposed FERDC significantly reduces the area and power consumption while keeping the error metrics such as mean square error (MSE) and peak signal to noise ratio (PSNR) in the acceptable range. Simulation results of a two dimensional wavelet filter shows that the proposed FERDC technique achieves the FIFO size reduction of up to 44.44% with PSNR values larger than 39 dB, which leads to the reduction of at least 31.6% in the dynamic power and 44.44% in the leakage power. Seyed Mohammad Ali Zeinolabedin, Jun Zhou 0017, Xin Liu 0015, Tony Tae-Hyoung Kim |
ISCAS | 2 |
| 2011 | The impact of inverse narrow width effect on sub-threshold device sizingabstractSub-threshold operation has been proved to be successful to achieve minimum energy consumption. It is well known that the sub-threshold device sizing is different from super-threshold due to different current behavior. The previously reported sub-threshold sizing methods assume that the current is proportional to the transistor width. However, we have found that the inverse narrow width effect has a significant influence on the threshold voltage in the sub-threshold region, causing non-proportional current-width relationship. Sizing without considering this effect may result in significant imbalance in the rise and fall delay which degrades the performance, power consumption and the functional yield of the design. We have proposed a new sub-threshold sizing method to balance the rise and fall delay by taking into account the influence of inverse narrow width effect while minimizing the transistor size. Compared with the previous sub-threshold sizing method the delay and power-delay-product (PDP) are reduced by up to 35.4% and 73.4% with up to 57% saving in the area. Further, due to symmetric rise and fall delay the minimum operating voltage can be lowered by 8% which leads to another 16% of energy reduction. Jun Zhou 0017, Senthil Jayapal, Jan Stuyt, Jos Huisken, Harmke de Groot |
ASP-DAC | 1 |
| 2011 | A 40 nm inverse-narrow-width-effect-aware sub-threshold standard cell libraryabstractWe have investigated the impact of inverse narrow width effect on the threshold voltage and drain current in the near/sub-threshold region at three technology nodes (90 nm, 65 nm and 40 nm) and proposed a new sub-threshold device sizing method which is inverse-narrow-width-effect-aware to reduce the gate area, power consumption and delay. We applied the proposed sizing method in designing a 40 nm sub-threshold standard cell library. Compared with the sub-threshold standard cell library designed using the conventional sizing method, the proposed library has up to 20% less delay, up to 34% less power consumption and up to 47% less area. We used the proposed library for designing a digital base-band processor and achieved a total power consumption of around 5 μw with 6 MHz at 0.5 V, which is 17% better than the counterpart design. Jun Zhou 0017, Senthil Jayapal, Ben Busze, Jan Stuyt |
DAC | 1 |
| 2011 | A 36μW heartbeat-detection processor for a wireless sensor nodeabstractIn order to provide better services to elderly people, home healthcare monitoring systems have been increasingly deployed. Typically, these systems are based on wireless sensor nodes, and should utilize very low energy during their lifetimes, as they are powered by scavengers. In this article, we present an ultra-low power processing system for a wireless sensor node for very low duty cycle applications. In the CoolBio system-on-chip, we utilized several power reduction techniques at both the architecture level and the circuit level. These techniques include feature extraction, voltage and frequency scaling, clock and power gating and a redesign of key standard cells. In the design of the ultra-low power processing system, we paid special attention to the memory subsystem, as it is one of the most power-consuming modules in a design. We also designed a clock manager in order to reduce the power consumed by clocking, and a power manager that is able to power-off unutilized modules. The proposed wireless sensor node processing system consumes 36.4μW at 100MHz and 1.2V supply voltage, for a heartbeat-detection algorithm with a 0.01% duty cycle. Filipa Duarte, Jos Hulzink, Jun Zhou 0017, Jan Stuijt, Jos Huisken, Harmke de Groot |
ACM Trans. Design Autom. Electr. Syst. | 3 |