Zhuo Zou

dblp:39/3602 · DBLP profile ↗
← Back
64ranked-venue papers
0as first author
51since 2021 · last 2026
0000-0002-8546-1329ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 41 · 32 since 2021Computer networks · 7 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 6 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021
YearPublicationVenuePosition
2026 A Self-Supervised Neuromorphic Processor Using High-Dimensional Representations for Cognitive Map Navigation
abstract
This work proposes a self-supervised neuromorphic processor using high-dimensional representations for cognitive map navigation. By employing the Cognitive Map Learner (CML), it enables agents to explore and understand diverse environments online through random walks. To enhance path planning, the agent’s actions and observations are embedded into high-dimensional state spaces. This embedding creates a sense of direction, simplifying navigation into a retrieval process within an Associative Memory (AM). We design an energy-efficient processor that features a scalable multi-core hardware architecture with precision flexibility, combined with an on-chip random walk training engine. To balance the precision of the model with hardware overhead, two hardware-software co-design strategies are proposed. The first is a Content-Addressable Memory (CAM)-based approach for AM access, which reduces the number of memory access by up to 25%. The second involves high-dimensional matrix sparsity optimizations, reducing computation operations to less than 8%. We simulate this processor by a 40-nm CMOS technology, which has 2.88 mm2core area with 15.8 mW power at a frequency of 140 MHz. Compared to previous processors, our experiments show that the proposed processor achieves outstanding success rates of 99.9%, 96%, and 98.7% on 100 2D nodes, 125 3D nodes, and 25 abstract map nodes with obstacles, respectively. In terms of energy efficiency, it delivers a path planning result of 28 nJ/node and 35 nJ/node in 2D and 3D maps, offering a 1.2x to 2.9x improvement over the state-of-the-art.
Anqin Xiao, Luyu Yang, Yuhan He, Hengtan Zhang, Ziyi Yang 0014, Lirong Zheng 0001, Zhuo Zou
DATE7
2026 CIMS: A CAM-based In-Memory Sorting Architecture for Efficient Top-K Ranking
Yuhan He, Siheng Lei, Tianxi Hu, Lirong Zheng 0001, Zhuo Zou
ISCAS6
2026 Performance-Aware Design Space Exploration for Chiplet-Based Cortical Simulation Processors Considering Area Constraint
Fanxi Yang, Lufei Fan, Lirong Zheng 0001, Zhuo Zou
ISCAS8
2026 A Neuromorphic ASIC Design for Dexterous Hand Control
Hengtan Zhang, Yifu Liang, Zhongxue Gan 0001, Lirong Zheng 0001, Zhuo Zou
ISCAS8
2026 SIMBRAIN: A nonidealities-aware simulation framework for spiking neural networks based on memristor crossbars
Jiawei Xu 0002, Ruisi Shen, Dimitrios Stathis 0001, Lirong Zheng 0001, Zhuo Zou, Ahmed Hemani
Neurocomputing9
2026 FeDMus: Federated Dynamic Example Mining for Unlabeled Sensor Data in Industrial Automation
abstract
Machine learning-based time series analysis has numerous applications in industrial automation, particularly in identifying sensor signals collected from distributed machines. However, large-scale repetitive operations generate massive amounts of redundant data, posing challenges to centralized learning, which requires collecting all data. Federated active learning addresses these challenges by selecting the most informative samples from multiple local data distributions for training the local model, aggregating and broadcasting model weights without transferring raw data, thus preserving data privacy and reducing labeling costs. We propose a dynamic federated active learning strategy, FeDMus, Federated Dynamic Example Mining for Unlabeled Sensor Data. By monitoring the training states of local and global models, FeDMus optimizes the timing of sample selection and dynamically allocates the query budget. FeDMus includes a dynamic selector and a dynamic filter: the proposed selector dynamically allocates the query budget based on training loss and weight dispersion, selecting informative data based on uncertainty and diversity across both models; the filter dynamically adjusts the training set by reducing simple samples to prevent the model from getting stuck in redundant data. On the typical industrial dataset CNC machine, FeDMus achieves an accuracy of 91.96% using only 10% of the data. We validated our method on four datasets, and the experimental results show that FeDMus outperforms state-of-the-art methods in tasks such as anomaly detection, fault diagnosis, action and trajectory recognition.
Wanlin Yang, Tobias Schlagenhauf, Zhuo Zou
IEEE Trans Autom. Sci. Eng.5
2026 eBrainISA: Edge-Oriented Instruction Set Architecture for Hybrid Brain-Inspired Computing
Yujie Ying, Ziyi Yang 0014, Ling Liang 0003, Zegang Peng, Yifan Hu 0013, Zhuo Zou, Lei Deng 0003
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.9
2026 CAMPRO: A CAM-Based Processing-in-Memory Processor for Hyperdimensional Computing
abstract
This work introduces CAMPRO, a Content Addressable Memory (CAM)-based Processing-In-Memory (PIM) processor customized for Hyperdimensional Computing (HDC). CAMPRO leverages a 6T Split Word Lines (SWL) cell structure for its CAM, enabling efficient column-wise search for ultra-wide Hypervector (HV) storage and an optimized associative PIM architecture tailored to HDC operations, significantly enhancing energy efficiency. The four key operators of HDC, binding, bundling, permutation, and similarity, are mapped to the proposed architecture. CAMPRO enhances operational parallelism via approximate bundling and employs a hierarchical permutation method to mitigate the gap in flexible shift support within CAM-based PIM architecture. The fine grained pipelined operations boost processing efficiency and dynamically reclaim memory space to support larger models. The Two-Phase Bit Pruning (TPBP) strategy prunes redundant bits in class HVs across two computing stages to eliminate unnecessary computations, reducing operation counts by 73.6% and energy consumption by 65.2% while maintaining query precision. Simulated in a 22 nm CMOS process, CAMPRO occupies 1.13 mm2and consumes 0.99 mW at 200 MHz. CAMPRO demonstrates robust versatility and scalability across five datasets, including MUTAG, CIFAR10, MNIST, language classification, and EMG gesture recognition, using diverse encoding schemes. It achieves excellent energy efficiency and low latency from small to large-scale datasets. In language classification, it reduces inference energy by 99.2% compared to the similiar work. For EMG gesture recognition, it improves training and inference energy efficiency by 2.6x and 6.7x, and reduce inference latency by 73x compared to related works. On MNIST, it enhances energy efficiency by 11.3x and latency by 1.6x to the prior work, making it an efficient Artificial Intelligence of Things (AIoT) solution.
Yuhan He, Tianxi Hu, Anqin Xiao, Fanxi Yang, Hengtan Zhang, Lirong Zheng 0001, Zhuo Zou
IEEE Trans. Circuits Syst. I Regul. Pap.7
2026 PHENICS: A Scalable Neuromorphic FPGA Architecture for Million-Neuron Cortical Simulation With 4.6× Real-Time Acceleration
abstract
Traditional brain simulations using CPU/GPU architectures face critical limitations in both speed and scalability, particularly when modeling large-scale biological neural networks. Cortical models exhibit distinctive features, recurrent, random and sparse connectivity, and ultra-high synaptic density that fundamentally differ from the structured and dense architectures of deep neural networks (DNNs). As a result, conventional accelerators suffer from inefficiencies: communication overhead scales linearly with neuron count, and memory architectures exhibit poor utilization efficiency under complex connectivity, severely constraining scalability. To address these challenges, we presentPHENICS(Pyramidal Hierarchical Event-driven Neuromorphic Infrastructure for Cortical Simulation), a scalable FPGA-based architecture tailored for large-scale spiking neural network simulations. At the communication level, PHENICS introduces a pyramidal multi-tier on-chip network combined with aBusy-Aware Threshold Adaptation (BATA)routing strategy and a lightweight router design to mitigate network congestion and improve spike transmission efficiency. At the storage level, we employ a multi-level addressing scheme adapted to sparse, irregular synaptic connectivity and leverage high-bandwidth memory (HBM) for efficient large-scale synaptic access. PHENICS is implemented on the Xilinx Alveo U50 system, successfully simulating a one-million-neuron Leaky Integrate-and-Fire (LIF) cortical model with 4 billion synapses, achieving a$4.6\times $real-time acceleration. It achieves this with only 25% of the memory bandwidth available on GPU platforms, yet delivers a$5\times $speedup compared to GPU-based simulators. Furthermore, our architecture reduces synaptic storage overhead by$3\times $through compressed encoding and HBM-backed sparse addressing. Moreover, it exhibits sub-linear communication time growth with increasing neuron counts. These advancements pave the way for real-time simulation of billion-neuron brain-scale models, opening new frontiers in neuroscience and neuromorphic computing.
Fanxi Yang, Lufei Fan, Yuhan He, Hanwen Ou, Lirong Zheng 0001, Zhuo Zou
IEEE Trans. Circuits Syst. I Regul. Pap.8
2026 An Always-On Event-Triggered DVS Fall Detection Processor With Precision-Adaptive Inference in 40-nm CMOS
Ziyi Yang 0014, Jinqiao Yang, Quanshu Yan, Anqin Xiao, Lirong Zheng 0001, Zhuo Zou
IEEE Trans. Circuits Syst. I Regul. Pap.7
2026 Dynamic Grouping With a Self-Aware Computational Resource Allocation for Large-Scale Multi-Objective Optimization
Yuning Chen, Ziqing Zhou, Yi Liu 0027, Linqiang Hu, Zhuo Zou, Zhongxue Gan 0001, Chun Ouyang 0002
IEEE Trans. Evol. Comput.5
2026 Mobility-Aware Federated Learning: Optimizing Performance With Interpretable Models and Wireless Channel Resource Allocation
abstract
The integration of the Internet of Things (IoT) with Federated Learning (FL) offers a transformative approach to addressing the challenges of massive data processing and privacy preservation in distributed systems. As a decentralized machine learning paradigm, FL enables model training on distributed datasets while safeguarding data privacy, making it well-suited for IoT applications. However, the performance of wireless FL systems is often constrained by limited communication resources and the mobility of participating clients, which can disrupt efficient model training and convergence. In this paper, we propose a novel mobility-aware FL scheduling strategy that leverages interpretable machine learning to enhance resource allocation in wireless networks. A more effective and fair resource allocation strategy can be achieved by dynamically adjusting the weight of the model quality and the communication quality of the training participants. We evaluate the proposed strategy against traditional scheduling methods in both single and multi-base station scenarios. Simulation results reveal that our approach significantly enhances overall learning efficiency by prioritizing high-value local models. Furthermore, for mobile clients, we identify an optimal range of average speed and participant numbers that maximizes the performance of wireless FL systems, offering practical insights for real-world deployments.
Jichao Leng, Zihuai Lin, Ming Ding 0001, Zhuo Zou, Branka Vucetic
IEEE Trans. Mob. Comput.4
2025 A Computation and Energy Efficient Hardware Architecture for SSL Acceleration
abstract
In Computer Vision (CV), the deployment of Convolutional Neural Networks (CNNs) is often hindered by their substantial computational requirements and large labeled datasets. Self-supervised learning (SSL) serves as an effective approach to reducing the reliance on labeled data with the option of augmentation methods to infer and train CNNs. Excluding irrelevant features accelerates learning and improves optimization. We propose a Field-Programmable Gate Array (FPGA)-based hardware accelerator architecture tailored for SSL framework, leveraging its parallelism and reconfigurability to expedite block matching, optimize sparse convolutions, and manage data reuse, significantly improving resource and energy efficiency. The implementation and evaluation of our work on Xilinx ZCU102 FPGA working at 200 MHz confirm that the similarity finding part's FPGA accelerations with a low hardware overhead generates a latency of 0.0106 seconds, surpassing GPU and CPU, and in the sparse CNN's FPGA acceleration part, with the processing of VGG16 and ResNet50, compared with the related FPGA-based works, our design claims a maximum of 3.08× throughput improvement and 1.5× in energy efficiency.
Huidong Ji, Sheng Li 0019, Chen Ding 0010, Jiawei Xu 0001, Qitao Tan, Jun Liu 0075, Ao Li 0004, Xulong Tang, Lirong Zheng 0001, Geng Yuan, Zhuo Zou
ASP-DAC12
2025 DuoQ: A DSP Utilization-aware and Outlier-free Quantization for FPGA-based LLMs Acceleration
abstract
Quantization enables efficient deployment of large language models (LLMs) on FPGAs, but its presence of outliers affects the accuracy of the quantized model. Existing methods mainly deal with outliers through channel-wise or tokenwise isolation and encoding, which leads to expensive dynamic quantization. To address this problem, we introduce DuoQ, an FPGA-oriented algorithm-hardware co-design framework. DuoQ effectively eliminates outliers through learnable equivalent transformations and low-semantic token awareness in the quantization scheme part, facilitating per-tensor quantization with 4 -bits. We co-design the quantization algorithm and hardware architecture. Specifically, DuoQ accelerates end-to-end LLM through a novel DSP-aware PE unit design and encoder design. In addition, two types of post-processing units assist in the realization of nonlinear functions and dynamic token awareness. Experimental results show that compared with platforms with different architectures, DuoQ’s computational efficiency and energy efficiency are improved by up to $8.8 \times$ and $23.45 \times$. In addition, DuoQ has achieved accuracy improvements compared to other outlieraware software and hardware works.
Zhuoquan Yu, Huidong Ji, Junfu Wu, Xiaoze Yan, Lirong Zheng 0001, Zhuo Zou
DAC7
2025 NLBP: Efficient Training Spiking Neural Networks with Neuron-Level Back Propagation
abstract
Training brain-inspired spiking neural networks (SNNs) with multi-timestep backpropagation imposes substantial memory and computational overhead, hindering their scalability and deployment. To address these challenges, we propose a Neuron-Level Back Propagation (NLBP) method, which eliminates the need for temporal unfolding while maintaining competitive performance. We fit the relationship between a neuron’s average input current and its firing rate using an adaptive sigmoid function. Leveraging this mapping, we derive a neuron-level, single-step backpropagation rule that avoids explicit temporal unfolding. This design significantly reduces computational and memory costs when training deep, large-scale SNNs. Furthermore, NLBP unifies the training of integrate-and-fire (IF) and leaky integrate-and-fire (LIF) neuron models within a framework, and supports both soft and hard reset mechanisms, enhancing generality and practicality. Extensive experiments on pattern classification and object detection demonstrate that NLBP achieves competitive accuracy while reducing memory usage by up to 73.07% and training time by 66.04%.
Zikai Zhu, Yulong Yan, Longrun Xu, Jinqiao Yang, Lirong Zheng 0001, Zhuo Zou
ECAI6
2025 Perturbation-efficient Zeroth-order Optimization for Hardware-friendly On-device Training
abstract
Zeroth-order (ZO) optimization is an emerging deep neural network (DNN) training paradigm that offers computational simplicity and memory savings. However, this seemingly promising approach faces a significant and long-ignored challenge. ZO requires generating a substantial number of Gaussian random numbers, which poses significant difficulties and even makes it infeasible for hardware platforms, such as FPGAs and ASICs. In this paper, we identify this critical issue, which arises from the mismatch between algorithm and hardware designers. To address this issue, we proposed PeZO, a perturbation-efficient ZO framework. Specifically, we design random number reuse strategies to significantly reduce the demand for random number generation and introduce a hardware-friendly adaptive scaling method to replace the costly Gaussian distribution with a uniform distribution. Our experiments show that PeZO reduces the required LUTs and FFs for random number generation by 48.6% and 12.7%, and saves at maximum 86% power consumption, all without compromising training performance, making ZO optimization feasible for on-device training. To the best of our knowledge, we are the first to explore the potential of on-device ZO optimization, providing valuable insights for future research.
Qitao Tan, Sung-En Chang, Huidong Ji, Chence Yang, Ci Zhang, Jun Liu 0075, Zheng Zhan 0001, Zhenman Fang, Zhuo Zou, Yanzhi Wang 0001, Jin Lu 0001, Geng Yuan
ICCAD10
2025 Vision-Based Leader-Follower Formation Control with Distance-Angle Feedback Regulation
abstract
This paper presents a resource-efficient monocular vision framework for leader-follower formation control in GPS-denied environments. To address the challenges of markerless navigation and dynamic interference, our approach integrates geometry-constrained perception with a dual-loop PID control architecture. The main contributions are: (1) A dynamic inverse projection method that reduces scale drift by 62% through optical flow-verified bounding box normalization; (2) A cascaded PID architecture that decouples distance and angle control, achieving 15 FPS on embedded hardware; (3) An implicit communication paradigm enabling 92% occlusion recovery without explicit data exchange. Experimental results demonstrate a 5.2% mean distance error in the 10-25cm range, sub-centimeter adjustments under varying illumination, and robustness in textured environments. Comparative analysis shows a 31% reduction in tracking error compared to marker-based baselines. These findings highlight the potential of our framework for robust, scalable, and infrastructure-free multi-robot formation control.
Sunyao Zhou, Zhuo Zou, Yonghao Li, Muzhen He, Lizheng Liu
INDIN2
2025 CAP-HDC: A CAM-Based Processor for Hyperdimensional Computing
abstract
This paper presents CAP-HDC, a Content Addressable Memory (CAM)-based processor designed for Hyperdimensional Computing (HDC). CAP-HDC integrates the binding, bundling, permutation, and similarity operators of HDC into the in-memory associative processing framework that features high parallelism, thereby achieving low power consumption and latency. The CAM utilized in CAP-HDC is designed using the Split Word Lines (SWL) 6T bit cell, enabling column-wise searching across all rows simultaneously. The approximate bundling method is proposed to implement bundling in CAM with multiple Hypervectors (HVs) without sacrificing accuracy on the 5-Class Gesture dataset. The hierarchical permutation method is proposed to implement permutation with lower power consumption, achieving a reduction of 90.66% in power consumption compared to the direct circular shift method. CAP-HDC is simulated using the 22 nm CMOS process, occupying an area of 1.06 mm2 and consuming 1.08 mW at a clock frequency of 200 MHz with a 0.9 V power supply. Compared to previous works, CAP-HDC improves energy efficiency by 2.9x and latency by 2.4x on the MNIST dataset. For hand gesture prediction based on EMG signals, CAP-HDC achieves improvements of 3.1x in inference energy efficiency and 2.6x in encoding energy efficiency.
Yuhan He, Anqin Xiao, Tianxi Hu, Fanxi Yang, Hengtan Zhang, Lirong Zheng 0001, Zhuo Zou
ISCAS7
2025 A Neuromorphic Controller with On-Chip Learning for Robot Motion Control
abstract
Motion control is one of the most fundamental issues in robotics, with kinematics and dynamics serving as its core components. While most existing control systems rely on general-purpose processors with large areas and high power consumption. This paper proposes a neuromorphic controller with on-chip learning, satisfying the requirements of high control performance and low cost for robot motion control. The proposed controller consists of an Operational Space Control (OSC) unit and a Spiking Neural Networks (SNNs) processing unit, offering kinematic and dynamic motion control across different (4, 6, 7, and 9) Degrees of Freedom (DoF). Under external disturbances, its control precision and the convergence speed are enhanced by 2.83× and 1.78×, respectively, compared to standard proportional integrated-error derivative (PID) OSC controller. The controller is simulated under 40 nm CMOS technology, occupying a core area of 0.755 mm2and consuming 2.4 mW of power at a frequency of 100 MHz. Compared with other chips used for robot motion control, the proposed controller achieves 2.15× and 65× enhancements in core area and power consumption.
Hengtan Zhang, Jinqiao Yang, Yuhan He, Fanxi Yang, Lirong Zheng 0001, Zhuo Zou
ISCAS7
2025 Decentralized Learning in Space: A Framework for Efficient Model Training in LEO Constellations
abstract
Relying on ground-based infrastructure for model training in Low Earth Orbit (LEO) constellations introduces significant challenges, including intermittent and costly spaceground communication. To address these challenges, this paper proposes a decentralized learning framework that enables model training directly within satellite constellations, utilizing intra-and inter-plane inter-satellite links (ISLs) for information exchange. The paper develops a geo-spatial filtering mechanism that selects the most relevant satellites to form a smaller, more efficient constellation. It then introduces a novel Adaptive Halving-Doubling (AHD) algorithm that enables collective communication within the emerging partial ring topology. Experimental results validate the framework's effectiveness, efficiency, and scalability. Notably, the transmission energy was observed to constitute less than 10% of traditional full constellation (FC) setups, even with increasing model sizes. Moreover, communication overhead relative to FC setups decreased with increasing constellation size, highlighting the operational efficiency of the framework.
Christine Mwase, Kosta Dakic, Albert Kahira, Bassel Al Homssi, Zhuo Zou
WCNC5
2025 Self-aware collaborative edge inference with embedded devices for IIoT
Zhuoquan Yu, Yi Jin 0007, Christine Mwase, Zhuo Zou, Lirong Zheng 0001
Future Gener. Comput. Syst.7
2025 CDL-H: Cluster-Based Decentralized Learning With Heterogeneity-Aware Strategies for Industrial Internet of Things
abstract
In the Industrial Internet of Things (IIoT), leveraging edge clients collaboratively for decentralized learning is increasingly promising for emerging intelligent applications. However, the resources of edge clients are typically heterogeneous, such as computing power and training data, leading to a significant decrease in both learning efficiency and accuracy. Furthermore, resource constraints on the central server restrict scaling to more clients and data. To address these challenges, this article introduces a novel framework called clustered-based decentralized learning with heterogeneity-aware strategies (CDL-H). The framework incorporates two key strategies: 1) client clustering strategy and 2) differentiated iteration strategy. Different from the traditional approaches, CDL-H divides the clients into multiple clusters based on the similarity of their data distribution. Within each cluster, edge clients initiate multiple local iterations based on local computation and exchange models with other clients without the need for a central server. Additionally, the theoretical analysis on the convergence rate of the proposed framework is provided. To validate its effectiveness, extensive experiments are implemented based on popular tasks, such as image classification and fault identification in time-series data. The results demonstrate the robustness of the CDL-H framework across a wide range of computing power heterogeneity, with differences as large as$19 \times $and dropout rate as high as 50%. Moreover, CDL-H outperforms existing baselines, with up to 9.93% accuracy improvement on Fashion-MNIST and up to 6.10% accuracy improvement on CWRU.
Zhuoquan Yu, Jichao Leng, Zhuo Zou, Lirong Zheng 0001
IEEE Internet Things J.5
2025 SAIndust: A Self-Aware Heterogeneous Computing Framework for Industrial Internet of Things
abstract
Distributed collaborative automation and resource scheduling are important for improving the productivity of intelligent manufacturing in the Industrial Internet of Things (IIoT). However, current efforts at the edge layer, where a large number of operations converge and device interactions are concentrated, are inadequate in dealing with the resulting computational heterogeneity and dynamic changes in the operating environment. To address these issues, we propose a self-aware heterogeneous computing framework (SAIndust). First, we design and implement a fine-grained heterogeneous resource virtualization technology based on Kubernetes, which pools computing resources and implements circulation to improve resource utilization. Then, we design a self-aware method that drives distributed system state update and scheduling, which is an autonomic optimization framework for real-time scheduling. Finally, we build a physical prototype platform and develop a practical plug-and-play deployment and evaluation tools. Experiments with deep learning applications with different resource intensities show that its 1.54% and 1.85% GPU virtualization overheads and standard deviation of resource allocation can achieve good virtualization performance and high fidelity. On the other hand, while achieving a 56.9% reduction in the average age of information and only a 25.6% increase in the average CPU cost, SAIndust can reduce the resource saturation by an average of 8.71% and achieve a maximum throughput increase of 5.12× compared to related methods in medium-scale to ultra-large-scale edge clusters.
Zhuoquan Yu, Jichao Leng, Huidong Ji, Lirong Zheng 0001, Zhuo Zou
IEEE Internet Things J.6
2025 MemMIMO: A Simulation Framework for Memristor-Based Massive MIMO Acceleration
abstract
Memristor-based crossbar architectures have proven highly effective for matrix vector multiplication (MVM) operations, making them a promising solution for accelerating the MVMs widely used in precoding algorithms for multiple input multiple output (MIMO) wireless communication systems. However, real-world implementation of memristor-based computing systems face challenges due to commonly observed non-idealities in both the devices themselves and the circuits they’re built into. To facilitate a rapid design flow and investigate the impact of non-idealities, an integrated open-source simulation framework MemMIMO is developed. The simulation framework estimates the accuracy and hardware performance of the computing system, offering a variety of flexible design options. MemMIMO integrates a behavioral model of the mix-signal architecture with a digital front-end. There are three major building blocks in MemMIMO: the device fitting block, the mapping block, and the performance estimation block. These blocks work together to map the complex MVMs in precoding algorithms for MIMO systems to crossbar-based architectures that incorporate memristor models characterized by physical device behavior. Using two typical use cases targeting six-generation (6G) massive MIMO communication as case studies, MemMIMO is used to model different memristor devices, explore the impact of non-idealities on system accuracy, and benchmark circuit-level performance metrics including area, speed, and power.
Jiawei Xu 0002, Dimitrios Stathis 0001, Ruisi Shen, Lirong Zheng 0001, Zhuo Zou, Ahmed Hemani
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.7
2025 Toward Efficient Eye Tracking in AR/VR Devices: A Near-Eye DVS-Based Processor for Real-Time Gaze Estimation
abstract
This paper presents an efficient near-eye dynamic vision sensor (DVS)-based processor for real-time eye tracking in augmented reality/virtual reality (AR/VR) devices. The processor takes advantage of the sparse event data with fine time resolution from the DVS, addressing the need for high frame-rate, low-power, and accurate eye tracking on wearable devices with extended battery life. Exploiting the inherent sparsity of event data, we propose an event-density-based region of interest (ROI) determination method that operates directly on event stream, which requires$47\times $fewer operations than the traditional methods, effectively overcoming the latency problem caused by the heavy computational loads. To eliminate the issue of decreasing accuracy at the edges of the field of view (FoV), we customized and fine-tuned a neural network for gaze estimation, ensuring uniformly distributed sub-degree accuracy. An estimator with a streamlined output mapping strategy and an adaptive window-sliding convolution scheme is implemented for gaze estimation acceleration. The processor is designed and fabricated in UMC 40-nm LP technology with a core area of 2.52 mm2 and performs end-to-end eye tracking exclusively with the raw event stream from DVS, achieving an average accuracy of 0.91° within a$96^{\circ } \times 64^{\circ }$FoV. Operating at 200 MHz, it achieves a dynamic frame rate of up to 1.2 kHz and requires only$12.7~\mu $J of energy per gaze estimation. By integrating the DVS, the processor enables real-time, low-power, and accurate eye tracking, enhancing the immersive experience on AR/VR devices and offering intuitive and seamless interactions.
Shihang Tan, Jinqiao Yang, Ziyi Yang 0014, Qinyu Chen, Lirong Zheng 0001, Zhuo Zou
IEEE Trans. Circuits Syst. I Regul. Pap.7
2025 Real-Time Scheduling Framework for Multiagent Cooperative Logistics With Dynamic Supply Demands
abstract
In logistics systems with multiagent collaboration, one of the prevailing focus lies on modeling as the dynamic multiperiod vehicle routing problem (DMPVRP). This work introduces modifications to DMPVRP to align with the requirements of real factory operations, particularly with dynamic supply demands. A self-established multiagent dynamic scheduling framework has been proposed to adapt to dynamic environmental changes and make timely adjustments, which consists of two modules: dynamic path planning and machine assignment. The first module utilizes a self-designed multioperator two-stage evolutionary algorithm to dynamically update the routes for vehicles. The second module maintains the workload balance among vehicles in real time. Experimental results demonstrate that the proposed algorithm achieves optimal outcomes compared to three state-of-the-art algorithms, surpassing others by 20% in machine output and exhibiting 5% lower transportation costs. In addition, a case study from a steel cord manufacturing factory is conducted, demonstrating its capability to promptly enhance efficiency.
Yuning Chen, Yi Liu 0027, Hongda Zhang, Ziqing Zhou, Wenchao Ding 0001, Zhuo Zou, Chun Ouyang 0002, Zhongxue Gan 0001
IEEE Trans. Ind. Informatics7
2024 An Empirical Study of Distributed Deep Learning Training on Edge (Student Abstract)
abstract
Deep learning (DL), despite its success in various fields, remains expensive and inaccessible to many due to its need for powerful supercomputing and high-end GPUs. This study explores alternative computing infrastructure and methods for distributed DL on low-energy, low-cost devices. We experiment on Raspberry Pi 4 devices with ARM Cortex-A72 processors and train a ResNet-18 model on the CIFAR-10 dataset. Our findings reveal limitations and opportunities for future optimizations, paving the way for a DL toolset for low-energy edge devices.
Christine Mwase, Albert Kahira, Zhuo Zou
AAAI3
2024 FPGA-Based HPC for Associative Memory System
abstract
Associative memory plays a crucial role in the cognitive capabilities of the human brain. The Bayesian Confidence Propagation Neural Network (BCPNN) is a cortex model capable of emulating brain-like cognitive capabilities, particularly associative memory. However, the existing GPU-based approach for BCPNN simulations faces challenges in terms of time overhead and power efficiency. In this paper, we propose a novel FPGA-based high performance computing (HPC) design for the BCPNN-based associative memory system. Our design endeavors to maximize the spatial and timing utilization of FPGA while adhering to the constraints of the available hardware resources. By incorporating optimization techniques including shared parallel computing units, hybrid-precision computing for a hybrid update mechanism, and the globally asynchronous and locally synchronous (GALS) strategy, we achieve a maximum network size of $150 \times 10$ and a peak working frequency of 100 MHz for the BCPNN-based associative memory system on the Xilinx Alveo U200 Card. The tradeoff between performance and hardware overhead of the design is explored and evaluated. Compared with the GPU counterpart, the FPGA-based implementation demonstrates significant improvements in both performance and energy efficiency, achieving a maximum latency reduction of $33.25 \times$, and a power reduction of over $6.9 \times$, all while maintaining the same network configuration.
Yu Yang 0020, Dimitrios Stathis 0001, Ahmed Hemani, Anders Lansner, Jiawei Xu 0002, Lirong Zheng 0001, Zhuo Zou
ASPDAC9
2024 A Plug-and-Play Sensor System for Machine Fault Diagnosis in Lights-Out Manufacturing
abstract
Early awareness of machine faults in lights-out manufacturing allows manufacturers to prevent failures. However, contemporary anomaly detection systems demand significant investments in infrastructure and development. This paper introduces a plug-and-play sensor system to monitor machines in heavy-duty industries. We present a self-adaptive smart sensor system integrating data acquisition, model training, and inference into one system. Featuring a pre-trained feature extractor and a lightweight detector trained on-device after deployment, it offers autonomous machine data analysis to achieve machine fault diagnosis in a plug-and-play manner. We conduct a 25-day experiment with vibration signals and an adaptive anomaly detection algorithm in an actual manufacturing scenario. The evaluation shows that the proposed model occupied only 74KB of disk space and achieved 93.8% accuracy without expert knowledge.
Wanlin Yang, Zhuo Zou
IECON3
2024 A Fully Synthesizable Capacitorless Digital LDO for Distributed Power Delivery Network
abstract
This paper presents a fully synthesizable capacitorless digital low-dropout regulator (DLDO) for distributed power delivery networks in large-scale digital systems. A coarse-fine dual loop architecture is adopted for better transient response and higher output voltage accuracy. The coarse loop uses a CMP-triggered oscillator to achieve faster recovery under load voltage droop. An inverter-based droop detector is connected directly to the output voltage, which provides rapid detection of undershoot voltage and dispenses with bulky external capacitors. Moreover, the DLDO is implemented using only digital standard cells and auto place-and-route (P&R) tool. The fully synthesizability enables seamless integration into existing digital systems, providing flexibility and scalability. Therefore, the proposed DLDO offers a scalable and portable architecture that has a low design time cost for a distributed power delivery network. The proposed DLDO is implemented and simulated in the 40-nm CMOS technology, with a core area of 0.04 mm2. The range of input voltage is from 0.6 to 1.1 V with a 50-mV dropout voltage. When the current load is increased by 120 mA with a 2 ns edge time, the DLDO exhibits a voltage droop of 126 mV, a response time of 2.1 ns and a settle time of 7.5 ns, respectively. The maximum load current and peak current efficiency are 200 mA and 99.98%, respectively.
Chengwei Cao, Xiongchuan Huang, Zhuo Zou, Lirong Zheng 0001
ISCAS4
2024 A Near-Eye DVS-Based End-to-End Eye Tracking Processor for AR/VR Applications
abstract
This paper presents a near-eye DVS-based end-to-end eye tracking processor for augmented reality/virtual reality (AR/VR) devices, addressing the need for low-power, high frame rate and accurate eye tracking deployment on wearable devices with extended battery life. The processor features a dedicated hardware implementation for energy-efficient event-driven pre-processing and acceleration, and delivers accurate gaze estimation utilizing a customized neural network (NN) with constant-event-count sample partitioning method. It is capable of performing end-to-end eye tracking exclusively with the Dynamic Vision Sensor (DVS), achieving an average accuracy of 0.91° in a 96°×64° FoV. The processor is implemented and simulated using UMC 40-nm LP CMOS technology, with a core area of 1.88 mm2. Operating at a clock frequency of 200 MHz, it achieves a dynamic frame rate of up to 1250 Hz with a power dissipation of 8.09 mW, corresponding to a 20.82 µJ energy per gaze estimation. With the integration of DVS, this processor enables real-time, low-power and accurate eye tracking for an immersive interaction experience on AR/VR devices.
Shihang Tan, Quanshu Yan, Lirong Zheng 0001, Zhuo Zou
ISCAS5
2024 Spiking-HDC: A Spiking Neural Network Processor with HDC Classifier Enabling Transfer Learning
abstract
This work proposes Spiking-HDC, a spiking neural network (SNN) processing system with hyperdimensional computing (HDC) and its hardware design for domain transfer scenarios. The input data is firstly fed into a two-layer SNN, serving as a feature extractor. It is followed by a HDC classifier to process feature vectors using hypervectors in binary representation. Such a system leverages HDC’s capability of single-pass learning, which can be adopted to rapidly updating but highly similar tasks by fine-tuning the HDC classifier with limited labeled data. By our experiments, the proposed system demonstrates transfer learning accuracy of 94.76%, 87.12% and 94.37% with few-shot samples on N-MNIST, DVS-Gesture and MNIST datasets, respectively. To apply Spiking-HDC model to extreme edge inference tasks, a dedicated processor is designed and implemented. The simulated results in 40 nm CMOS process illustrate that it has 0.88 mm2core area and 1.8 mW power at 100 MHz frequency. In comparison to similar works, it achieves 3.8×-36× inference energy efficiency enhancement.
Anqin Xiao, Jinqiao Yang, Lirong Zheng 0001, Zhuo Zou
ISCAS5
2024 TSCM: A TCAM-Based Sparse Connection Memory Architecture in Neuromorphic Computing System for Cortical Simulation
abstract
The connection matrix requires significant memory capacity in large-scale Spiking Neural Networks (SNNs). However, the sparsity of connections in cortical models leads to memory capacity and energy inefficiencies. This paper proposes Ternary Content Addressable Memory (TCAM)-based Sparse Connection Memory (TSCM) architecture in neuromorphic computing systems for cortical simulation. The architecture consists of a TCAM for searching existing synaptic connections, a Static Random Access Memory (SRAM) for storing synapse addresses, and a three-stage circuit for memory access control. By leveraging the sparsity, the proposed memory architecture demonstrates improved area and energy efficiency compared to the conventional Direct Mapped Full-Address Memory (DMFAM) architecture. A TSCM macro is designed, simulated using UMC 40-nm CMOS technology, and evaluated across various scales of classical cortical models. Experimental results demonstrate that the TSCM architecture reduces area by 27.6% to 75.6% and energy consumption by 15.8% to 96.0% in cortical simulation with neurons ranging from 100k to 10M compared to DMFAM architecture.
Fanxi Yang, Yuhan He, Lirong Zheng 0001, Zhuo Zou
ISCAS5
2024 Robust Robot Formation Control Based on Streaming Communication and Leader-Follower Approach
abstract
In this paper, a non-visual robotic formation control method based on a streaming communication architecture and a leader-follower control model is studied. The proposed stream-based communication architecture is inspired by the flocking behavior of fish. We analogize it into a form resembling an N-ary tree for communication purposes. Communication proceeds to the next layer only when all nodes in the upper layer have completed the follower selection. We also introduce a fault-tolerance mechanism and a termination filtering mechanism to prevent multiple leaders from choosing the same follower, avoiding a scenario where the robots in the last layer enter an endless loop of follower selection. The proposed stream-based communication architecture, built upon serial and parallel tracking, can achieve more complex formations, such as rectangular formations, closely resembling real-world scenarios, significantly enhancing formation efficiency. Simulation experiments on the e-puck platform validate the effectiveness and robustness of this architecture.
Zhuo Zou, Xiaoming Hu 0001, Zhongxue Gan 0001, Lizheng Liu
IWCMC2
2024 DAI-NET: Toward communication-aware collaborative training for the industrial edge
abstract
The industrial edge generates an abundance of spatially distributed and dynamic data that needs to remain on-site for privacy and security reasons. Collaborative training at the edge can leverage this data to refine pre-trained models locally for specific industrial tasks and environments and have them adapt to local changes for enhanced performance, agility, and resilience. However, communication between the devices during training is a key bottleneck and is not modelled by existing frameworks such as MxNet, PyTorch and TensorFlow. This paper introduces DAI-NET, a co-simulation framework for examining communication and its associated costs, and provides results from an implementation using Python, OMNET++ and INET. To validate it and showcase its utility, the developed platform is applied in the analysis of (i) the performance and cost of collaboratively training a Multilayer Perceptron model, and (ii) the influence of computational heterogeneity. Communication costs generated during the training are captured at the device and system levels. In computationally heterogeneous clusters , the root cause of stragglers is exposed. In addition, the key performance contributors are identified to be a cluster’s computation capability and the variation in the relative computation capabilities of its devices. This study is particularly useful for Artificial Intelligence of Things (AIoT) systems, whose bandwidth and energy resources are limited. It lends the way for more practical research on communication-efficient algorithms, network protocols and architectures for the AIoT edge.
Christine Mwase, Yi Jin 0007, Tomi Westerlund, Hannu Tenhunen, Zhuo Zou
Future Gener. Comput. Syst.5
2024 MCU-Enabled Epileptic Seizure Detection System With Compressed Learning
abstract
Epilepsy is one of the most common neurological disorder diseases all over the world, which gives patients a huge burden in seizure-related disabilities. For epileptic seizure detection, encephalography (EEG) is a commonly used clinical approach. Recently, several Internet of Things (IoT)-based wearable monitoring systems using machine learning (ML) approaches have been proposed to assist real-time detection of epileptic seizure attack scenarios. Among these approaches, convolutional neural networks (CNNs) provide superior accuracy, at the expense of high computational complexity that is not friendly to resource-constrained wearable devices. In this work, we propose a compressed learning (CL)-based epileptic seizure detection system which enables implementing CNN on microcontroller units (MCUs). The proposed CL approach combines the measurement matrix of compressed sensing (CS) and a 1-D CNN. As a result, the input data size and parameter number of CNN can be significantly reduced while eliminating the complex reconstruction process of the traditional CS approach. Evaluated on the Bonn dataset, our proposed system ensures a 96.44% accuracy under a 0.1 compressed ratio (CR) corresponding to a$33.5\times $multiply-accumulate operations (MACs) reduction and a$21.9\times $decrease in model size compared with the baseline. Mapping such a model on an MSP432 MCU, the memory requirement is 50.13 kB and the power consumption is 13.4727 mW at 40 MHz frequency, corresponding to an energy consumption of$269.4~\mu \text{J}$/Classification, with a classification latency of 21.07 ms.
Liyu Qian, Yuxiang Huan, Yaojie Sun, Lirong Zheng 0001, Zhuo Zou
IEEE Internet Things J.7
2024 DTL-IDS: An optimized Intrusion Detection Framework using Deep Transfer Learning and Genetic Algorithm
abstract
In the dynamic field of the Industrial Internet of Things (IIoT), the networks are increasingly vulnerable to a diverse range of cyberattacks. This vulnerability necessitates the development of advanced intrusion detection systems (IDSs). Addressing this need, our research contributes to the existing cybersecurity literature by introducing an optimized Intrusion Detection System based on Deep Transfer Learning (DTL), specifically tailored for heterogeneous IIoT networks. Our framework employs a tri-layer architectural approach that synergistically integrates Convolutional Neural Networks (CNNs), Genetic Algorithms (GA), and bootstrap aggregation ensemble techniques. The methodology is executed in three critical stages: First, we convert a state-of-the-art cybersecurity dataset, Edge_IIoTset, into image data, thereby facilitating CNN-based analytics. Second, GA is utilized to fine-tune the hyperparameters of each base learning model, enhancing the model’s adaptability and performance. Finally, the outputs of the top-performing models are amalgamated using ensemble techniques, bolstering the robustness of the IDS. Through rigorous evaluation protocols, our framework demonstrated exceptional performance, reliably achieving a 100% attack detection accuracy rate. This result establishes our framework as highly effective against 14 distinct types of cyberattacks. The findings bear significant implications for the ongoing development of secure, efficient, and adaptive IDS solutions in the complex landscape of IIoT networks.
Shahid Latif, Wadii Boulila, Anis Koubaa, Zhuo Zou, Jawad Ahmad 0001
J. Netw. Comput. Appl.4
2024 CorTile: A Scalable Neuromorphic Processing Core for Cortical Simulation With Hybrid-Mode Router and TCAM
abstract
In neuromorphic processors, simulating large-scale Spiking Neural Networks (SNNs) for cortical models necessitates a significant increase in communication traffic and memory capacity, due to the lack of exploiting the sparsity of connections. Therefore, this paper proposes CorTile, a scalable neuromorphic processing core designed for cortical simulation. We propose a hybrid-mode router that supports Remote Unicast and Local Broadcast (RULB) routing method, leveraging the high local connectivity and low distal connectivity observed in cortical models. This approach achieves reductions of 36.7% in average router load, 40.7% in peak load, 51.2% in average link traffic, 41.7% in peak traffic, respectively, compared to conventional routing methods. Additionally, the proposed Ternary Content Addressable Memory (TCAM)-based Sparse Connection Memory (TSCM) architecture leads to 87.1% reduction in area and a 62.7% reduction in power consumption. These approaches effectively decrease communication traffic and mitigate the quadratic increase in memory requirements, achieving linear growth instead, thus achieving scalability. The proposed CorTile is simulated using UMC 40-nm CMOS process, occupying an area of 5.15 mm2, supporting a maximum of 8k neurons and 64M synapses. Evaluated using a typical macaque cortex model, it consumes 8.25 mW, with the router operating at 200 MHz and the other modules at 100 MHz. This design achieves an average router load of 12.33 Mpackets/s and peak link traffic of 21.16 MB/s. Thanks to the scalability of the proposed processing core that can be tiled into many-core processors, it paves the way for chiplets and multiple chip integration towards a brain-scale neuromorphic computing system.
Fanxi Yang, Yuhan He, Jinqiao Yang, Anqin Xiao, Lufei Fan, Lirong Zheng 0001, Zhuo Zou
IEEE Trans. Circuits Syst. I Regul. Pap.8
2023 Self-aware Collaborative Edge Inference with Embedded Devices for Task-oriented IIoT
abstract
The computing and communication resources of embedded devices are constrained and heterogeneous, resulting in a low quality-of-experience for compute-intensive applications in task-oriented industrial Internet of Things (IIoT), such as edge inference. To address these challenges, we first propose a model partitioning-based self-aware collaborative edge inference framework. Furthermore, the throughput-aware collaborative inference algorithm is designed for typical IIoT scenario, stacking tasks. Via jointly optimizing the partition layer and collaborative device selection, the optimal inference efficiency, maximum inference throughput, can be obtained. Finally, the performance of our proposal is demonstrated by extensive simulations and tests based on 10 Raspberry Pi 4Bs and popular models. Specifically, with the proposed algorithm, our platform reaches up to 14.77× throughput speed up for stacking tasks, which indicates that the the proposed design can improve the inference efficiency.
Zhuoquan Yu, Christine Mwase, Yi Jin 0007, Lirong Zheng 0001, Zhuo Zou
VTC Fall7
2023 A Low-Power Hybrid-Precision Neuromorphic Processor With INT8 Inference and INT16 Online Learning in 40-nm CMOS
abstract
In this work, we present a neuromorphic processor for artificial intelligence of things (AIoT) applications featuring low-power consumption, a small footprint, STDP-based online learning, and the ability to adapt to multiple applications. Hybrid precision, i.e., INT8 for inference and INT16 for training, is suggested to achieve balanced accuracy and energy efficiency. A precision-configurable leaky integrate-and-fire(LIF) neuron unit and a unified memory architecture are designed to maximize datapath reuse. A dynamic pruning technique is proposed to exploit the temporal sparsity, yielding synaptic operations reduction by 3.68x in training and 1.63x in inference, respectively. The design is implemented and fabricated in a 40-nm CMOS process, with a core area of 0.87 mm2. It is measured to consume a minimal power of$680~\mu \text{W}$at 70 MHz under a 0.75 V power supply, corresponding to 9.9 pJ per synaptic operation. Evaluated with typical spatial, temporal, and spatiotemporal datasets (MNIST, MIT-BIH, and N-MNIST), the proposed design achieve energy efficiency comparable to the best-in-class solutions with handcrafted training and customized ASICs, while demonstrating improved versatility across multiple applications with balanced accuracy, power consumption, and model adaptability.
Congyang Liu, Ziyi Yang 0014, Zikai Zhu, Haoming Chu, Yuxiang Huan, Lirong Zheng 0001, Zhuo Zou
IEEE Trans. Circuits Syst. I Regul. Pap.8
2023 ASLog: An Area-Efficient CNN Accelerator for Per-Channel Logarithmic Post-Training Quantization
abstract
Post-training quantization (PTQ) has been proven an efficient model compression technique for Convolution Neural Networks (CNNs), without re-training or access to labeled datasets. However, it remains challenging for a CNN accelerator to fulfill the efficiency potential of PTQ methods. A large number of PTQ techniques blindly pursue high theoretic compression effect and accuracy, ignoring their impact on the actual hardware implementation, which causes more hardware overhead than benefit. This paper introduces ASLog, a PTQ-friendly CNN accelerator that explores four key designs in an algorithm-hardware co-optimizing manner: the first practical 4-bit logarithmic PTQ pipeline SLogII, the multiplier-free arithmetic element (AE) design, the energy-efficient bias correction element (BCE) design, and the per-channel quantization friendly (PCF) architecture and dataflow. The proposed SLogII PTQ pipeline can push the limit of logarithmic PTQ to 4-bit with40% lower in power and area consumption compared with a common 8-bit multiplier. The BCE and PCF design proposed in this paper are the first to consider the hardware impact of the widely-used per-channel quantization and bias correction technique, enabling an efficient PTQ-friendly implementation with a small hardware overhead. The ASLog is validated in a UMC 40-nm process, with 12.2 TOPS/W energy efficiency and 0.80 mm2 core area. The ASLog can achieve 336.3 GOPS/mm2 area efficiency and >500 OPs/Byte operational intensity, which map to over$1.85\times $and$1.12\times $improvement compared with the previous related works.
Jiawei Xu 0002, Jiangshan Fan, Baolin Nan, Chen Ding 0010, Lirong Zheng 0001, Zhuo Zou, Yuxiang Huan
IEEE Trans. Circuits Syst. I Regul. Pap.6
2023 A Domain-Specific Accelerator for Ultralow Latency Market Data Distribution System
abstract
Ultralow latency parsing of financial data is gaining significance in the high-frequency trading of the security exchange market. Hardware-aided systems exhibit superior improvement of latency over traditional software solutions, but flexibility may suffer when processing the financial protocol. This article presents a domain-specific accelerator for the market data distribution system, which integrates a financial information exchange adapted for streaming (FAST) decoder, a 10-Gbps network interface, and a high-speed PCIe host interface into a single field programmable gate array (FPGA) acceleration card. The proposed FAST decoder adopts the finite state machines-coordinated sequence mapping table to achieve run-time reconfigurability over fine-grained FPGA programming, and 16 fields can be decoded simultaneously in a pipelined manner, resulting in a decoding latency of only 33 ns. Evaluated on the Xilinx Alveo U200 acceleration card, this work improves the latency of decoding a FAST message by 26%–72% than state-of-the-art FPGA designs, and outperforms the software baseline by$>27\times$in terms of latency covering both decoding and communication.
Yuxiang Huan, Chen Ding 0010, Yulong Yan, Jianjun Cui, Jiachen Wang 0008, Chuhuang Cai, Zhuo Zou, Lirong Zheng 0001
IEEE Trans. Ind. Informatics9
2023 An IoT-Based Wearable Labor Progress Monitoring System for Remote Evaluation of Admission Time to Hospital
abstract
Because contractions signal the approach of labor, pregnant women-especially primigravidas (i.e., women pregnant for the first time)-usually go to the hospital to seek medical intervention when they begin experiencing contractions, which is not conductive to good perinatal outcomes. Conventionally, uterine contraction monitoring requires specialized medical devices and relies on the doctor's clinical experience. Therefore, exploring an objective method to detect labor onset at home and avoid early hospital admission has essential importance. In this article, a labor progress monitoring system based on a sensing device, edge service, and Internet of things (IoT) platform is proposed, aiming to suggest suitable hospital admission times for low-risk primigravidas. The pregnant woman places the sensing device on her abdomen with the help of a belt to detect contraction activities. An intelligent edge service for contraction classification is deployed on a mobile phone. The system's artificial intelligence (AI)-assisted algorithm is lightweight, with 670 kB and 194 kB of memory dedicated to a convolutional neural network and long short-term memory, respectively. It classifies the pregnant woman as deferred admission, optional admission, or recommended admission according to different contraction states. An IoT platform connected to the hospital is implemented, providing professional suggestions from doctors. The test set collected in an emergency clinic shows that the proposed system can reach a classification accuracy of more than 96%. In conclusion, the proposed system enables remote labor progress monitoring at home and avoids early hospital admission.
Zhiqing Xiao, Lihua Xu, Zhuo Zou, Lirong Zheng 0001
IEEE J. Biomed. Health Informatics5
2022 Communication-efficient distributed AI strategies for the IoT edge
Christine Mwase, Yi Jin 0007, Tomi Westerlund, Hannu Tenhunen, Zhuo Zou
Future Gener. Comput. Syst.5
2022 A Hybrid-Mode On-Chip Router for the Large-Scale FPGA-Based Neuromorphic Platform
abstract
Large-scale neuromorphic computing requires the multi-chip network to provide high computing power. Efficient routing schemes and on-chip router design are necessary for handling various inter-chip transmission patterns. In this paper, we propose a hybrid-mode on-chip router that supports both multicast and unicast routing for the large-scale neuromorphic simulation. Two routing schemes, namely Cache-like Spike Weight Indexing and General Unicast Flow Control, are proposed to accommodate the chip-to-chip transmission of spike and non-spike data. This work is evaluated on a neuromorphic platform built with an$8\times 8$FPGA chips array. Running a simulation of 1M neurons at 200MHz, the proposed router achieves a processing latency of 25ns and a chip-to-chip latency of 287ns. Working in the unicast mode, the router can synchronize status flags of all chips within$5 ~\mu \text{s}$. Moreover, it reduces the peak spike traffic by 25.65% with the help of Load-aware Multicast Routing, compared with other multicast routing strategies.
Chen Ding 0010, Yuxiang Huan, Yulong Yan, Fanxi Yang, Lizheng Liu, Meigen Shen, Zhuo Zou, Lirong Zheng 0001
IEEE Trans. Circuits Syst. I Regul. Pap.8
2022 Edge-Based Collaborative Training System for Artificial Intelligence-of-Things
abstract
The descending of intelligence from the cloud to the heterogeneous and low-power edge in the Artificial Intelligence-of-Things prevents uploading user-sensitive information to the cloud. It brings an urgent demand for deploying training tasks collaboratively in industrial scenarios to manage data locally. This article proposes an edge-based collaborative training system for the smart factory which harnesses the intelligence of edge devices by balancing the computational and communicational resources and improving system dependability. Two typical scenarios of parts recognition and defect inspection are evaluated as a case study with our system. The feasibility and dependability of the presented system are verified with a platform composed of eight high-performance (Nvidia Jetson Nano) and eight low-performance edge devices (Raspberry Pi 4B). The efficiency under tradeoff between computational resource and network condition constraints in a cluster is tested to simulate real-case performance in smart factory scenarios. Our platform reaches the peak performance of 1167 images/s training efficiency on ResNet32 under a 125 MB/s bandwidth. Experimental results demonstrate that the proposed design can collaboratively perform training tasks with optimized efficiency and provide dependable collaborations for system fault detection and cluster extension.
Yi Jin 0007, Yulong Yan, Yuxiang Huan, Jiawei Xu 0002, Shancang Li, Prosanta Gope, Zhuo Zou, Lirong Zheng 0001
IEEE Trans. Ind. Informatics9
2021 Design Framework for SRAM-Based Computing-In-Memory Edge CNN Accelerators
abstract
This paper presents an architectural framework and an evaluation model for Static Random Access Memory (SRAM)-based Computing-in-Memory (CIM) edge Convolutional Neural Network (CNN) accelerators. To provide a baseline for system-level design perspectives, an architectural framework for SRAM-CIM design concerning the key design points in state-of-the-art works is proposed. Furthermore, a configurable evaluation model featuring top-down design flow based on the proposed framework is established to investigate design space explorations. Case studies validated the framework and evaluation model using LeNet-5, AlexNet and VGG-16 to achieve energy-aware optimizations. The optimized memory scale for LeNet-5 is "16 PEs and 120 tiles" with the minimal estimated inference energy of 0.0018J, while for AlexNet and VGG-16, "16 PEs and 120 tiles" is better achieving minimal energy consumption of 0.1733mJ and 0.6825mJ respectively. Estimation results highlight tradeoffs among data represent- tation parameters and memory partitioning parameters. This work provides specific SRAM-CIM design guidelines from a system-level perspective.
Yimin Wang 0001, Zhuo Zou, Lirong Zheng 0001
ISCAS2
2021 Self-aware distributed deep learning framework for heterogeneous IoT edge devices
Yi Jin 0007, Jiawei Cai, Jiawei Xu 0002, Yuxiang Huan, Yulong Yan, Yongliang Guo, Lirong Zheng 0001, Zhuo Zou
Future Gener. Comput. Syst.9
2021 An IoT-Based Anti-Counterfeiting System Using Visual Features on QR Code
abstract
This article presents an Internet-of-Things (IoT) anti-counterfeiting system that uses visual features combined with the quick response (QR) code. The visual features guarantee the authenticity of a product with the QR code for tracking and tracing. Two visual features, i.e., natural texture features and printed micro features are exploited in the proposed system. The natural texture features use the texture of fiber paper to achieve physical unclonable function (PUF), while the micro features are artificially generated for improved industrial manufacturability and reliability. Features are generated and registered in the production phase when the QR code is printed. In the anti-counterfeiting verification phase, the feature obtained through the feature extraction algorithm is compared with the record to calculate similarity, which indicates the verification result. Such an approach is fully compatible with the QR code-based logistic process without any additional manufacturing cost. A user-friendly application has been developed on a mobile platform that facilitates easy-to-use and affordable devices for verification, such as a mobile phone or a handheld code reader. The experimental results show 99.6% and 99.9% accuracy of anti-counterfeiting verification for texture features and micro features, respectively. The system with corresponding algorithms and software has been demonstrated in real-life products.
Yulong Yan, Zhuo Zou, Yu Gao 0042, Lirong Zheng 0001
IEEE Internet Things J.2
2021 IECA: An In-Execution Configuration CNN Accelerator With 30.55 GOPS/mm² Area Efficiency
abstract
It remains challenging for a Convolutional Neural Network (CNN) accelerator to maintain high hardware utilization and low processing latency with restricted on-chip memory. This paper presents an In-Execution Configuration Accelerator (IECA) that realizes an efficient control scheme, exploring architectural data reuse, unified in-execution controlling, and pipelined latency hiding to minimize configuration overhead out of the computation scope. The proposed IECA achieves row-wise convolution with tiny distributed buffers and reduces the size of total on-chip memory by removing 40% of redundant memory storage with shared delay chains. By exploiting a reconfigurable Sequence Mapping Table (SMT) and Finite State Machine (FSM) control, the chip realizes cycle-accurate Processing Element (PE) control, automatic loop tiling and latency hiding without extra time slots for pre-configuration. Evaluated on AlexNet and VGG-16, the IECA retains over 97.3% PE utilization and over 95.6% memory access time hiding on average. The chip is designed and fabricated in a UMC 55-nm process running at a frequency of 250 MHz and achieves an area efficiency of 30.55 GOPS/mm2and 0.244 GOPS/KGE (kilo-gate-equivalent), which makes an over$2.0\times $and$2.1\times $improvement, respectively, compared with that of previous related works. Implementation of the IEC control scheme uses only a 0.55% area of the 2.75 mm2core.
Boming Huang, Yuxiang Huan, Haoming Chu, Jiawei Xu 0002, Lizheng Liu, Lirong Zheng 0001, Zhuo Zou
IEEE Trans. Circuits Syst. I Regul. Pap.7
2021 A Wearable Hand Rehabilitation System With Soft Gloves
abstract
Hand paralysis is one of the most common complications in stroke patients, which severely impacts their daily lives. This article presents a wearable hand rehabilitation system that supports both mirror therapy and task-oriented therapy. A pair of gloves, i.e., a sensory glove and a motor glove, was designed and fabricated with a soft, flexible material, providing greater comfort and safety than conventional rigid rehabilitation devices. The sensory glove worn on the nonaffected hand, which contains the force and flex sensors, is used to measure the gripping force and bending angle of each finger joint for motion detection. The motor glove, driven by micromotors, provides the affected hand with assisted driving force to perform training tasks. Machine learning is employed to recognize the gestures from the sensory glove and to facilitate the rehabilitation tasks for the affected hand. The proposed system offers 16 kinds of finger gestures with an accuracy of 93.32%, allowing patients to conduct mirror therapy using fine-grained gestures for training a single finger and multiple fingers in coordination. A more sophisticated task-oriented rehabilitation with mirror therapy is also presented, which offers six types of training tasks with an average accuracy of 89.4% in real time.
Xiaoshi Chen, Shih-Ching Yeh, Lirong Zheng 0001, Zhuo Zou
IEEE Trans. Ind. Informatics7
2020 Thermal-Cycling-aware Dynamic Reliability Management in Many-Core System-on-Chip
abstract
Dynamic Reliability Management (DRM) is a common approach to mitigate aging and wear-out effects in multi- /many-core systems. State-of-the-art DRM approaches apply finegrained control on resource management to increase/balance the chip reliability while considering other system constraints, e.g., performance, and power budget. Such approaches, acting on various knobs such as workload mapping and scheduling, Dynamic Voltage/Frequency Scaling (DVFS) and Per-Core Power Gating (PCPG), demonstrated to work properly with the various aging mechanisms, such as electromigration, and Negative-Bias Temperature Instability (NBTI). However, we claim that they do not suffice for thermal cycling. Thus, we here propose a novel thermal-cycling-aware DRM approach for shared-memory many-core systems running multi-threaded applications. The approach applies a fine-grained control capable at reducing both temperature levels and variations. The experimental evaluations demonstrated that the proposed approach is able to achieve 39% longer lifetime than past approaches.
M. H. Haghbayan, Antonio Miele, Zhuo Zou, Hannu Tenhunen, Juha Plosila
DATE3
2020 An Autonomous Error-Tolerant Architecture Featuring Self-reparation for Convolutional Neural Networks
abstract
Convolutional neural networks are widely used in artificial intelligence and Internet of Things area. As the scale of convolutional neural network expands, more and more processing units are provided for it. The systems are easy prone to error, and any computing problems in any layer of the network will lead to wrong output results. Traditional multimode redundancy methods make the systems more complex, and increase power consumption. This paper proposes an autonomous error-tolerant architecture for convolutional neural networks. Taking the LeNet-5 as an example, the network layers of CNN are mapped on the AET architecture, an error-tolerant synapse is designed to discover the errors, an active evolution scheme is designed to handle unrecoverable errors and implement network reconfiguration. This design is implemented on FPGA, and the experimental results show that this architecture can realize effective error tolerance for convolutional neural network and has fast error recovery ability under the premise of ensuring the same recognition accuracy.
Lizheng Liu, Yuxiang Huan, Zhuo Zou, Xiaoming Hu 0001, Lirong Zheng 0001
VTC Spring3
2020 A Smart Dental Health-IoT Platform Based on Intelligent Hardware, Deep Learning, and Mobile Terminal
abstract
The dental disease is a common disease for a human. Screening and visual diagnosis that are currently performed in clinics possibly cost a lot in various manners. Along with the progress of the Internet of Things (IoT) and artificial intelligence, the internet-based intelligent system have shown great potential in applying home-based healthcare. Therefore, a smart dental health-IoT system based on intelligent hardware, deep learning, and mobile terminal is proposed in this paper, aiming at exploring the feasibility of its application on in-home dental healthcare. Moreover, a smart dental device is designed and developed in this study to perform the image acquisition of teeth. Based on the data set of 12 600 clinical images collected by the proposed device from 10 private dental clinics, an automatic diagnosis model trained by MASK R-CNN is developed for the detection and classification of 7 different dental diseases including decayed tooth, dental plaque, uorosis, and periodontal disease, with the diagnosis accuracy of them reaching up to 90%, along with high sensitivity and high specificity. Following the one-month test in ten clinics, compared with that last month when the platform was not used, the mean diagnosis time reduces by 37.5% for each patient, helping explain the increase in the number of treated patients by 18.4%. Furthermore, application software (APPs) on mobile terminal for client side and for dentist side are implemented to provide service of pre-examination, consultation, appointment, and evaluation.
Lizheng Liu, Jiawei Xu 0002, Yuxiang Huan, Zhuo Zou, Shih-Ching Yeh, Lirong Zheng 0001
IEEE J. Biomed. Health Informatics4
2018 A 3D Tiled Low Power Accelerator for Convolutional Neural Network
abstract
It remains a challenge to run Deep Learning in devices with stringent power budget in the Internet-of-Things. This paper presents a low-power accelerator for processing Convolutional Neural Networks on the embedded devices. The power reduction is realized by exploring data reuse in three different aspects, with regards to convolution, filter and input features. A systolic-like data flow is proposed and applied to rows of Processing Elements (PEs), which facilitate reusing the data during convolution. Reuse of input features and filters is achieved by arranging the PE array in a 3D tiled architecture, whose dimension is 3 × 14 × 4. Local storage within PEs is therefore reduced and only cost 17.75 kB, which is 20% of the state-of-the-art. With dedicated delay chains in each PE, this accelerator is reconfigurable to suit various parameter settings of convolutional layers. Evaluated in UMC 65 nm low leakage process, the accelerator can reach a peak performance of 84 GOPS and consume only 136 mW at 250 Mhz.
Yuxiang Huan, Jiawei Xu 0002, Lirong Zheng 0001, Hannu Tenhunen, Zhuo Zou
ISCAS5
2018 TMR Group Coding Method for Optimized SEU and MBU Tolerant Memory Design
abstract
This work proposes a fault tolerant memory design using the method of Triple Module Redundancy (TMR) group coding to tolerant the Single-Event Upset (SEU) and Multi-Bit Upset (MBU) influence on memory devices in space environment. The group coding method uses different models to partition and code each word line in memory with Hamming code to achieve best performance. TMR group coding method further increases the capability of self-correction for the errors occurred in parity bits. The evaluation results show that the suggested approach can obtain improved correctness for the memory output with optimized tradeoff between reliability and cost. At 5% error rate, the probability of correct output reaches 70.78% with small cost increment. To achieve 90% reliability, the accuracy improvement is 31.9% compared to TMR with 9% increased area. This solution proposed is evaluated on the memory rich micro-coded processor, but can be further extended to other memory-based processors that need high reliability for the SEU and MBU influence in aerospace applications.
Yi Jin 0007, Yuxiang Huan, Haoming Chu, Zhuo Zou, Lirong Zheng 0001
ISCAS4
2018 A Design of Autonomous Error-Tolerant Architectures for Massively Parallel Computing
Lizheng Liu, Yi Jin 0007, Yi Liu 0027, Yuxiang Huan, Zhuo Zou, Lirong Zheng 0001
IEEE Trans. Very Large Scale Integr. Syst.6
2017 Smart energy efficient gateway for Internet of mobile things
abstract
Internet of Things (IoT) is a fast developing vision in which physical quantities are digitized, processed and analyzed. Internet of Mobile Things (IoMT) as one of new domains of IoT, due to mobility, requires a more demanding and rigorous solution in many aspects, especially in terms of energy efficiency. We propose a solution consisting of energy efficient and fast hardware platform for building IoMT Fog layer facilities. Experimental results are presented to prove superiority of the proposed hardware in several aspects to popular general purpose platforms.
Igor Tcarenko, Yuxiang Huan, David Juhasz, Amir-Mohammad Rahmani, Zhuo Zou, Tomi Westerlund, Pasi Liljeberg, Lirong Zheng 0001, Hannu Tenhunen
CCNC5
2016 A 101.4 GOPS/W reconfigurable and scalable control-centric embedded processor for domain-specific applications
abstract
Increasing the energy efficiency and performance while providing the customizability and scalability is vital for embedded processors adapting to domain-specific applications such as Internet of Things. In this paper, we proposed a reconfigurable and scalable control-centric architecture, and implemented the design consisting of two cores and an on-chip multi-mode router in 65 nm technology. The reconfigurability is enabled by the restructurable sequence mapping table (SMT) thus the reorganizable functional units. Owing to the integration of the multi-mode router, on-chip or inter-chip network for multi-/many-core computing can be composed for performance extension on demand even in the post-fabrication stage. Control-centric design simplifies the control logic, shrinks the non-functional units and orchestrates the operations to increase the hard are utilization and reduce the excessive data movement for high energy efficiency. As a result, the processor can both conduct general-purpose processing with 29% smaller code size and application-specific processing with over 10 times performance improvement when implementing AES by SMT. The dual-core processor consumes 19.7 μW/MHz with die size of 3.5 mm2. The achieved energy efficiency is 101.4GOPS/W.
Zhuo Zou, Zhonghai Lu, Lirong Zheng 0001, Yuxiang Huan, Stefan Blixt
ISCAS2
2016 Design and implementation of multi-mode routers for large-scale inter-core networks
Zhuo Zou, Zhonghai Lu, Lirong Zheng 0001
Integr.2
2015 Implementing MVC Decoding on Homogeneous NoCs: Circuit Switching or Wormhole Switching
abstract
To implement multiview video decoding on network on-chip (NoC) based homogeneous multicore architectures, the selection of switching techniques for routers is one of the most important aspects for design space exploration. Circuit switching and wormhole switching are two most feasible switching techniques for on-chip networks. To choose the suitable switching technique, we perform the comparison on decoding speed of the whole system, link utilization and delay between circuit switching and wormhole switching for implementing eight-view QVGA video decoding on 4 × 4 NoCs at 30 fps. The required link bandwidths are both around 800 Mbps with the similar network utilization and delay. We conclude that, to implement multiview video decoding on homogeneous NoCs, circuit switching is more suitable considering the similar performance and lower cost compared with wormhole switching.
Zhuo Zou, Zhonghai Lu, Lirong Zheng 0001
PDP2
2014 A wirelessly-powered UWB sensor tag with time-domain sensor interface
abstract
This paper presents a wirelessly-powered sensor tag with a time-domain sensor interface for wireless sensing applications. The tag is remotely powered by RF wave. Instead of traditional approaches employing conventional ADCs for quantization and transmitter for data communication, in this work, a Pulse Position Modulator incorporating simple impulse radio UWB (IR-UWB) transmitter is proposed to convert and transmit the analog sensing information in time domain. The analog signal is compared with an adjustable triangular wave for analog to time conversion in signal-varying environments. Then a UWB transmitter converts the PPM signal to very short pulses and sends it back to the reader. The time interval of UWB pulses represents the original input signal in time domain which can be measured on the reader side by a time-to-digital conversion. This approach not only simplifies the ADC design but also relaxes the number of bits transmitted on the tag side. The sensor tag is designed in 180nm CMOS process. Simulation results demonstrate that the proposed approach reduce transmission power consumption by nearly 3 orders of magnitude over traditional approaches, while consuming only 85 µW for 1.5 MS/s sampling rate.
Dongxuan Bao, Zhuo Zou, Majid Baghaei Nejad, Lirong Zheng 0001
ISCAS2
2011 Analog front-end RX design for UWB impulse radio in 90nm CMOS
abstract
In this paper a reconfigurable differential Ultra Wideband-Impulse Radio (UWB-IR) energy receiver architecture has been simulated and implemented in UMC 90nm. The signal is amplified, rectified and integrated. By using an integration windowed scheme the SNR requirements are relaxed increasing the sensitivity. The design has been optimized for large bandwidths, low implementation area and configurability. The RX can be adapted to work at different data rates, processing gains, and channel environments. It works between the 3.1 – 4.8 GHz bands with OOK or PPM modulation with a tunable data rate up to 33Mb/s. In order to relax the ADC sampling time an interleave mode of operation has been implemented. It has a maximum power consumption of 22m W with a power supply of 1V. The complete RX occupies an area of 1.11mm2.
David Sarmiento M., Zhuo Zou, Jia Mao, Peng Wang 0092, Fredrik Jonsson, Lirong Zheng 0001
ISCAS2
2007 A Novel Passive Tag with Asymmetric Wireless Link for RFID and WSN Applications
abstract
In this paper, we present a radio-powered module with asymmetric wireless link utilizing ultra wideband radio system for RFID and wireless sensor applications. Our contribution includes using two different standards in uplink and downlink. Such as conventional RFIDs, incoming RF signal transmitted by reader is used to power the internal circuitry and receive the data. However, in upstream link, an IR-UWB transmitter is utilized. Unlike traditional RFID systems, due to great advantages of UWB communication, this tag is very robust to multi-path fading and collision problem and it is more secure against eavesdropping or jamming. The module consists of a power scavenging unit, a RF receiver, an IR-UWB transmitter, digital baseband controller, and an embedded UWB antenna are designed for integration on liquid-crystal polymer (LCP) substrate, using 0.18mum CMOS process technology.
Majid Baghaei Nejad, Zhuo Zou, Hannu Tenhunen, Lirong Zheng 0001
ISCAS2