VLDB 2026 Research / reviewers in the wild / expert
Lirong Zheng 0001
dblp:15/7206 · also Li-Rong Zheng 0001
· DBLP profile ↗
78ranked-venue papers
1as first author
38since 2021 · last 2026
0000-0001-9588-0239ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 52 · 1 first-author · 27 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 4 since 2021Computer networks · 7 · 4 since 2021Software engineering, systems software and programming languages · 3 · 2 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Security and privacy · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Self-Supervised Neuromorphic Processor Using High-Dimensional Representations for Cognitive Map NavigationabstractThis work proposes a self-supervised neuromorphic processor using high-dimensional representations for cognitive map navigation. By employing the Cognitive Map Learner (CML), it enables agents to explore and understand diverse environments online through random walks. To enhance path planning, the agent’s actions and observations are embedded into high-dimensional state spaces. This embedding creates a sense of direction, simplifying navigation into a retrieval process within an Associative Memory (AM). We design an energy-efficient processor that features a scalable multi-core hardware architecture with precision flexibility, combined with an on-chip random walk training engine. To balance the precision of the model with hardware overhead, two hardware-software co-design strategies are proposed. The first is a Content-Addressable Memory (CAM)-based approach for AM access, which reduces the number of memory access by up to 25%. The second involves high-dimensional matrix sparsity optimizations, reducing computation operations to less than 8%. We simulate this processor by a 40-nm CMOS technology, which has 2.88 mm2core area with 15.8 mW power at a frequency of 140 MHz. Compared to previous processors, our experiments show that the proposed processor achieves outstanding success rates of 99.9%, 96%, and 98.7% on 100 2D nodes, 125 3D nodes, and 25 abstract map nodes with obstacles, respectively. In terms of energy efficiency, it delivers a path planning result of 28 nJ/node and 35 nJ/node in 2D and 3D maps, offering a 1.2x to 2.9x improvement over the state-of-the-art. Anqin Xiao, Luyu Yang, Yuhan He, Hengtan Zhang, Ziyi Yang 0014, Lirong Zheng 0001, Zhuo Zou |
DATE | 6 |
| 2026 | WaferBRAIN: Whole-Brain Scale Neuromorphic Architecture Based on Wafer-Scale Integration
Yukun Feng, Liangyu Gan, Haoming Chu, Yufan He, Jiaxin Yin, Lirong Zheng 0001, Yuxiang Huan |
ISCA | 7 |
| 2026 | CIMS: A CAM-based In-Memory Sorting Architecture for Efficient Top-K Ranking
Yuhan He, Siheng Lei, Tianxi Hu, Lirong Zheng 0001, Zhuo Zou |
ISCAS | 5 |
| 2026 | Performance-Aware Design Space Exploration for Chiplet-Based Cortical Simulation Processors Considering Area Constraint
Fanxi Yang, Lufei Fan, Lirong Zheng 0001, Zhuo Zou |
ISCAS | 7 |
| 2026 | A Neuromorphic ASIC Design for Dexterous Hand Control
Hengtan Zhang, Yifu Liang, Zhongxue Gan 0001, Lirong Zheng 0001, Zhuo Zou |
ISCAS | 7 |
| 2026 | SIMBRAIN: A nonidealities-aware simulation framework for spiking neural networks based on memristor crossbars
Jiawei Xu 0002, Ruisi Shen, Dimitrios Stathis 0001, Lirong Zheng 0001, Zhuo Zou, Ahmed Hemani |
Neurocomputing | 8 |
| 2026 | CAMPRO: A CAM-Based Processing-in-Memory Processor for Hyperdimensional ComputingabstractThis work introduces CAMPRO, a Content Addressable Memory (CAM)-based Processing-In-Memory (PIM) processor customized for Hyperdimensional Computing (HDC). CAMPRO leverages a 6T Split Word Lines (SWL) cell structure for its CAM, enabling efficient column-wise search for ultra-wide Hypervector (HV) storage and an optimized associative PIM architecture tailored to HDC operations, significantly enhancing energy efficiency. The four key operators of HDC, binding, bundling, permutation, and similarity, are mapped to the proposed architecture. CAMPRO enhances operational parallelism via approximate bundling and employs a hierarchical permutation method to mitigate the gap in flexible shift support within CAM-based PIM architecture. The fine grained pipelined operations boost processing efficiency and dynamically reclaim memory space to support larger models. The Two-Phase Bit Pruning (TPBP) strategy prunes redundant bits in class HVs across two computing stages to eliminate unnecessary computations, reducing operation counts by 73.6% and energy consumption by 65.2% while maintaining query precision. Simulated in a 22 nm CMOS process, CAMPRO occupies 1.13 mm2and consumes 0.99 mW at 200 MHz. CAMPRO demonstrates robust versatility and scalability across five datasets, including MUTAG, CIFAR10, MNIST, language classification, and EMG gesture recognition, using diverse encoding schemes. It achieves excellent energy efficiency and low latency from small to large-scale datasets. In language classification, it reduces inference energy by 99.2% compared to the similiar work. For EMG gesture recognition, it improves training and inference energy efficiency by 2.6x and 6.7x, and reduce inference latency by 73x compared to related works. On MNIST, it enhances energy efficiency by 11.3x and latency by 1.6x to the prior work, making it an efficient Artificial Intelligence of Things (AIoT) solution. Yuhan He, Tianxi Hu, Anqin Xiao, Fanxi Yang, Hengtan Zhang, Lirong Zheng 0001, Zhuo Zou |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2026 | PHENICS: A Scalable Neuromorphic FPGA Architecture for Million-Neuron Cortical Simulation With 4.6× Real-Time AccelerationabstractTraditional brain simulations using CPU/GPU architectures face critical limitations in both speed and scalability, particularly when modeling large-scale biological neural networks. Cortical models exhibit distinctive features, recurrent, random and sparse connectivity, and ultra-high synaptic density that fundamentally differ from the structured and dense architectures of deep neural networks (DNNs). As a result, conventional accelerators suffer from inefficiencies: communication overhead scales linearly with neuron count, and memory architectures exhibit poor utilization efficiency under complex connectivity, severely constraining scalability. To address these challenges, we presentPHENICS(Pyramidal Hierarchical Event-driven Neuromorphic Infrastructure for Cortical Simulation), a scalable FPGA-based architecture tailored for large-scale spiking neural network simulations. At the communication level, PHENICS introduces a pyramidal multi-tier on-chip network combined with aBusy-Aware Threshold Adaptation (BATA)routing strategy and a lightweight router design to mitigate network congestion and improve spike transmission efficiency. At the storage level, we employ a multi-level addressing scheme adapted to sparse, irregular synaptic connectivity and leverage high-bandwidth memory (HBM) for efficient large-scale synaptic access. PHENICS is implemented on the Xilinx Alveo U50 system, successfully simulating a one-million-neuron Leaky Integrate-and-Fire (LIF) cortical model with 4 billion synapses, achieving a$4.6\times $real-time acceleration. It achieves this with only 25% of the memory bandwidth available on GPU platforms, yet delivers a$5\times $speedup compared to GPU-based simulators. Furthermore, our architecture reduces synaptic storage overhead by$3\times $through compressed encoding and HBM-backed sparse addressing. Moreover, it exhibits sub-linear communication time growth with increasing neuron counts. These advancements pave the way for real-time simulation of billion-neuron brain-scale models, opening new frontiers in neuroscience and neuromorphic computing. Fanxi Yang, Lufei Fan, Yuhan He, Hanwen Ou, Lirong Zheng 0001, Zhuo Zou |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2026 | An Always-On Event-Triggered DVS Fall Detection Processor With Precision-Adaptive Inference in 40-nm CMOS
Ziyi Yang 0014, Jinqiao Yang, Quanshu Yan, Anqin Xiao, Lirong Zheng 0001, Zhuo Zou |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2025 | A Computation and Energy Efficient Hardware Architecture for SSL AccelerationabstractIn Computer Vision (CV), the deployment of Convolutional Neural Networks (CNNs) is often hindered by their substantial computational requirements and large labeled datasets. Self-supervised learning (SSL) serves as an effective approach to reducing the reliance on labeled data with the option of augmentation methods to infer and train CNNs. Excluding irrelevant features accelerates learning and improves optimization. We propose a Field-Programmable Gate Array (FPGA)-based hardware accelerator architecture tailored for SSL framework, leveraging its parallelism and reconfigurability to expedite block matching, optimize sparse convolutions, and manage data reuse, significantly improving resource and energy efficiency. The implementation and evaluation of our work on Xilinx ZCU102 FPGA working at 200 MHz confirm that the similarity finding part's FPGA accelerations with a low hardware overhead generates a latency of 0.0106 seconds, surpassing GPU and CPU, and in the sparse CNN's FPGA acceleration part, with the processing of VGG16 and ResNet50, compared with the related FPGA-based works, our design claims a maximum of 3.08× throughput improvement and 1.5× in energy efficiency. Huidong Ji, Sheng Li 0019, Chen Ding 0010, Jiawei Xu 0001, Qitao Tan, Jun Liu 0075, Ao Li 0004, Xulong Tang, Lirong Zheng 0001, Geng Yuan, Zhuo Zou |
ASP-DAC | 10 |
| 2025 | DuoQ: A DSP Utilization-aware and Outlier-free Quantization for FPGA-based LLMs AccelerationabstractQuantization enables efficient deployment of large language models (LLMs) on FPGAs, but its presence of outliers affects the accuracy of the quantized model. Existing methods mainly deal with outliers through channel-wise or tokenwise isolation and encoding, which leads to expensive dynamic quantization. To address this problem, we introduce DuoQ, an FPGA-oriented algorithm-hardware co-design framework. DuoQ effectively eliminates outliers through learnable equivalent transformations and low-semantic token awareness in the quantization scheme part, facilitating per-tensor quantization with 4 -bits. We co-design the quantization algorithm and hardware architecture. Specifically, DuoQ accelerates end-to-end LLM through a novel DSP-aware PE unit design and encoder design. In addition, two types of post-processing units assist in the realization of nonlinear functions and dynamic token awareness. Experimental results show that compared with platforms with different architectures, DuoQ’s computational efficiency and energy efficiency are improved by up to $8.8 \times$ and $23.45 \times$. In addition, DuoQ has achieved accuracy improvements compared to other outlieraware software and hardware works. Zhuoquan Yu, Huidong Ji, Junfu Wu, Xiaoze Yan, Lirong Zheng 0001, Zhuo Zou |
DAC | 6 |
| 2025 | NLBP: Efficient Training Spiking Neural Networks with Neuron-Level Back PropagationabstractTraining brain-inspired spiking neural networks (SNNs) with multi-timestep backpropagation imposes substantial memory and computational overhead, hindering their scalability and deployment. To address these challenges, we propose a Neuron-Level Back Propagation (NLBP) method, which eliminates the need for temporal unfolding while maintaining competitive performance. We fit the relationship between a neuron’s average input current and its firing rate using an adaptive sigmoid function. Leveraging this mapping, we derive a neuron-level, single-step backpropagation rule that avoids explicit temporal unfolding. This design significantly reduces computational and memory costs when training deep, large-scale SNNs. Furthermore, NLBP unifies the training of integrate-and-fire (IF) and leaky integrate-and-fire (LIF) neuron models within a framework, and supports both soft and hard reset mechanisms, enhancing generality and practicality. Extensive experiments on pattern classification and object detection demonstrate that NLBP achieves competitive accuracy while reducing memory usage by up to 73.07% and training time by 66.04%. Zikai Zhu, Yulong Yan, Longrun Xu, Jinqiao Yang, Lirong Zheng 0001, Zhuo Zou |
ECAI | 5 |
| 2025 | CAP-HDC: A CAM-Based Processor for Hyperdimensional ComputingabstractThis paper presents CAP-HDC, a Content Addressable Memory (CAM)-based processor designed for Hyperdimensional Computing (HDC). CAP-HDC integrates the binding, bundling, permutation, and similarity operators of HDC into the in-memory associative processing framework that features high parallelism, thereby achieving low power consumption and latency. The CAM utilized in CAP-HDC is designed using the Split Word Lines (SWL) 6T bit cell, enabling column-wise searching across all rows simultaneously. The approximate bundling method is proposed to implement bundling in CAM with multiple Hypervectors (HVs) without sacrificing accuracy on the 5-Class Gesture dataset. The hierarchical permutation method is proposed to implement permutation with lower power consumption, achieving a reduction of 90.66% in power consumption compared to the direct circular shift method. CAP-HDC is simulated using the 22 nm CMOS process, occupying an area of 1.06 mm2 and consuming 1.08 mW at a clock frequency of 200 MHz with a 0.9 V power supply. Compared to previous works, CAP-HDC improves energy efficiency by 2.9x and latency by 2.4x on the MNIST dataset. For hand gesture prediction based on EMG signals, CAP-HDC achieves improvements of 3.1x in inference energy efficiency and 2.6x in encoding energy efficiency. Yuhan He, Anqin Xiao, Tianxi Hu, Fanxi Yang, Hengtan Zhang, Lirong Zheng 0001, Zhuo Zou |
ISCAS | 6 |
| 2025 | A Neuromorphic Controller with On-Chip Learning for Robot Motion ControlabstractMotion control is one of the most fundamental issues in robotics, with kinematics and dynamics serving as its core components. While most existing control systems rely on general-purpose processors with large areas and high power consumption. This paper proposes a neuromorphic controller with on-chip learning, satisfying the requirements of high control performance and low cost for robot motion control. The proposed controller consists of an Operational Space Control (OSC) unit and a Spiking Neural Networks (SNNs) processing unit, offering kinematic and dynamic motion control across different (4, 6, 7, and 9) Degrees of Freedom (DoF). Under external disturbances, its control precision and the convergence speed are enhanced by 2.83× and 1.78×, respectively, compared to standard proportional integrated-error derivative (PID) OSC controller. The controller is simulated under 40 nm CMOS technology, occupying a core area of 0.755 mm2and consuming 2.4 mW of power at a frequency of 100 MHz. Compared with other chips used for robot motion control, the proposed controller achieves 2.15× and 65× enhancements in core area and power consumption. Hengtan Zhang, Jinqiao Yang, Yuhan He, Fanxi Yang, Lirong Zheng 0001, Zhuo Zou |
ISCAS | 6 |
| 2025 | Self-aware collaborative edge inference with embedded devices for IIoT
Zhuoquan Yu, Yi Jin 0007, Christine Mwase, Zhuo Zou, Lirong Zheng 0001 |
Future Gener. Comput. Syst. | 8 |
| 2025 | CDL-H: Cluster-Based Decentralized Learning With Heterogeneity-Aware Strategies for Industrial Internet of ThingsabstractIn the Industrial Internet of Things (IIoT), leveraging edge clients collaboratively for decentralized learning is increasingly promising for emerging intelligent applications. However, the resources of edge clients are typically heterogeneous, such as computing power and training data, leading to a significant decrease in both learning efficiency and accuracy. Furthermore, resource constraints on the central server restrict scaling to more clients and data. To address these challenges, this article introduces a novel framework called clustered-based decentralized learning with heterogeneity-aware strategies (CDL-H). The framework incorporates two key strategies: 1) client clustering strategy and 2) differentiated iteration strategy. Different from the traditional approaches, CDL-H divides the clients into multiple clusters based on the similarity of their data distribution. Within each cluster, edge clients initiate multiple local iterations based on local computation and exchange models with other clients without the need for a central server. Additionally, the theoretical analysis on the convergence rate of the proposed framework is provided. To validate its effectiveness, extensive experiments are implemented based on popular tasks, such as image classification and fault identification in time-series data. The results demonstrate the robustness of the CDL-H framework across a wide range of computing power heterogeneity, with differences as large as$19 \times $and dropout rate as high as 50%. Moreover, CDL-H outperforms existing baselines, with up to 9.93% accuracy improvement on Fashion-MNIST and up to 6.10% accuracy improvement on CWRU. Zhuoquan Yu, Jichao Leng, Zhuo Zou, Lirong Zheng 0001 |
IEEE Internet Things J. | 6 |
| 2025 | SAIndust: A Self-Aware Heterogeneous Computing Framework for Industrial Internet of ThingsabstractDistributed collaborative automation and resource scheduling are important for improving the productivity of intelligent manufacturing in the Industrial Internet of Things (IIoT). However, current efforts at the edge layer, where a large number of operations converge and device interactions are concentrated, are inadequate in dealing with the resulting computational heterogeneity and dynamic changes in the operating environment. To address these issues, we propose a self-aware heterogeneous computing framework (SAIndust). First, we design and implement a fine-grained heterogeneous resource virtualization technology based on Kubernetes, which pools computing resources and implements circulation to improve resource utilization. Then, we design a self-aware method that drives distributed system state update and scheduling, which is an autonomic optimization framework for real-time scheduling. Finally, we build a physical prototype platform and develop a practical plug-and-play deployment and evaluation tools. Experiments with deep learning applications with different resource intensities show that its 1.54% and 1.85% GPU virtualization overheads and standard deviation of resource allocation can achieve good virtualization performance and high fidelity. On the other hand, while achieving a 56.9% reduction in the average age of information and only a 25.6% increase in the average CPU cost, SAIndust can reduce the resource saturation by an average of 8.71% and achieve a maximum throughput increase of 5.12× compared to related methods in medium-scale to ultra-large-scale edge clusters. Zhuoquan Yu, Jichao Leng, Huidong Ji, Lirong Zheng 0001, Zhuo Zou |
IEEE Internet Things J. | 5 |
| 2025 | MemMIMO: A Simulation Framework for Memristor-Based Massive MIMO AccelerationabstractMemristor-based crossbar architectures have proven highly effective for matrix vector multiplication (MVM) operations, making them a promising solution for accelerating the MVMs widely used in precoding algorithms for multiple input multiple output (MIMO) wireless communication systems. However, real-world implementation of memristor-based computing systems face challenges due to commonly observed non-idealities in both the devices themselves and the circuits they’re built into. To facilitate a rapid design flow and investigate the impact of non-idealities, an integrated open-source simulation framework MemMIMO is developed. The simulation framework estimates the accuracy and hardware performance of the computing system, offering a variety of flexible design options. MemMIMO integrates a behavioral model of the mix-signal architecture with a digital front-end. There are three major building blocks in MemMIMO: the device fitting block, the mapping block, and the performance estimation block. These blocks work together to map the complex MVMs in precoding algorithms for MIMO systems to crossbar-based architectures that incorporate memristor models characterized by physical device behavior. Using two typical use cases targeting six-generation (6G) massive MIMO communication as case studies, MemMIMO is used to model different memristor devices, explore the impact of non-idealities on system accuracy, and benchmark circuit-level performance metrics including area, speed, and power. Jiawei Xu 0002, Dimitrios Stathis 0001, Ruisi Shen, Lirong Zheng 0001, Zhuo Zou, Ahmed Hemani |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2025 | Toward Efficient Eye Tracking in AR/VR Devices: A Near-Eye DVS-Based Processor for Real-Time Gaze EstimationabstractThis paper presents an efficient near-eye dynamic vision sensor (DVS)-based processor for real-time eye tracking in augmented reality/virtual reality (AR/VR) devices. The processor takes advantage of the sparse event data with fine time resolution from the DVS, addressing the need for high frame-rate, low-power, and accurate eye tracking on wearable devices with extended battery life. Exploiting the inherent sparsity of event data, we propose an event-density-based region of interest (ROI) determination method that operates directly on event stream, which requires$47\times $fewer operations than the traditional methods, effectively overcoming the latency problem caused by the heavy computational loads. To eliminate the issue of decreasing accuracy at the edges of the field of view (FoV), we customized and fine-tuned a neural network for gaze estimation, ensuring uniformly distributed sub-degree accuracy. An estimator with a streamlined output mapping strategy and an adaptive window-sliding convolution scheme is implemented for gaze estimation acceleration. The processor is designed and fabricated in UMC 40-nm LP technology with a core area of 2.52 mm2 and performs end-to-end eye tracking exclusively with the raw event stream from DVS, achieving an average accuracy of 0.91° within a$96^{\circ } \times 64^{\circ }$FoV. Operating at 200 MHz, it achieves a dynamic frame rate of up to 1.2 kHz and requires only$12.7~\mu $J of energy per gaze estimation. By integrating the DVS, the processor enables real-time, low-power, and accurate eye tracking, enhancing the immersive experience on AR/VR devices and offering intuitive and seamless interactions. Shihang Tan, Jinqiao Yang, Ziyi Yang 0014, Qinyu Chen, Lirong Zheng 0001, Zhuo Zou |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2024 | FPGA-Based HPC for Associative Memory SystemabstractAssociative memory plays a crucial role in the cognitive capabilities of the human brain. The Bayesian Confidence Propagation Neural Network (BCPNN) is a cortex model capable of emulating brain-like cognitive capabilities, particularly associative memory. However, the existing GPU-based approach for BCPNN simulations faces challenges in terms of time overhead and power efficiency. In this paper, we propose a novel FPGA-based high performance computing (HPC) design for the BCPNN-based associative memory system. Our design endeavors to maximize the spatial and timing utilization of FPGA while adhering to the constraints of the available hardware resources. By incorporating optimization techniques including shared parallel computing units, hybrid-precision computing for a hybrid update mechanism, and the globally asynchronous and locally synchronous (GALS) strategy, we achieve a maximum network size of $150 \times 10$ and a peak working frequency of 100 MHz for the BCPNN-based associative memory system on the Xilinx Alveo U200 Card. The tradeoff between performance and hardware overhead of the design is explored and evaluated. Compared with the GPU counterpart, the FPGA-based implementation demonstrates significant improvements in both performance and energy efficiency, achieving a maximum latency reduction of $33.25 \times$, and a power reduction of over $6.9 \times$, all while maintaining the same network configuration. Yu Yang 0020, Dimitrios Stathis 0001, Ahmed Hemani, Anders Lansner, Jiawei Xu 0002, Lirong Zheng 0001, Zhuo Zou |
ASPDAC | 8 |
| 2024 | A Fully Synthesizable Capacitorless Digital LDO for Distributed Power Delivery NetworkabstractThis paper presents a fully synthesizable capacitorless digital low-dropout regulator (DLDO) for distributed power delivery networks in large-scale digital systems. A coarse-fine dual loop architecture is adopted for better transient response and higher output voltage accuracy. The coarse loop uses a CMP-triggered oscillator to achieve faster recovery under load voltage droop. An inverter-based droop detector is connected directly to the output voltage, which provides rapid detection of undershoot voltage and dispenses with bulky external capacitors. Moreover, the DLDO is implemented using only digital standard cells and auto place-and-route (P&R) tool. The fully synthesizability enables seamless integration into existing digital systems, providing flexibility and scalability. Therefore, the proposed DLDO offers a scalable and portable architecture that has a low design time cost for a distributed power delivery network. The proposed DLDO is implemented and simulated in the 40-nm CMOS technology, with a core area of 0.04 mm2. The range of input voltage is from 0.6 to 1.1 V with a 50-mV dropout voltage. When the current load is increased by 120 mA with a 2 ns edge time, the DLDO exhibits a voltage droop of 126 mV, a response time of 2.1 ns and a settle time of 7.5 ns, respectively. The maximum load current and peak current efficiency are 200 mA and 99.98%, respectively. Chengwei Cao, Xiongchuan Huang, Zhuo Zou, Lirong Zheng 0001 |
ISCAS | 5 |
| 2024 | A Near-Eye DVS-Based End-to-End Eye Tracking Processor for AR/VR ApplicationsabstractThis paper presents a near-eye DVS-based end-to-end eye tracking processor for augmented reality/virtual reality (AR/VR) devices, addressing the need for low-power, high frame rate and accurate eye tracking deployment on wearable devices with extended battery life. The processor features a dedicated hardware implementation for energy-efficient event-driven pre-processing and acceleration, and delivers accurate gaze estimation utilizing a customized neural network (NN) with constant-event-count sample partitioning method. It is capable of performing end-to-end eye tracking exclusively with the Dynamic Vision Sensor (DVS), achieving an average accuracy of 0.91° in a 96°×64° FoV. The processor is implemented and simulated using UMC 40-nm LP CMOS technology, with a core area of 1.88 mm2. Operating at a clock frequency of 200 MHz, it achieves a dynamic frame rate of up to 1250 Hz with a power dissipation of 8.09 mW, corresponding to a 20.82 µJ energy per gaze estimation. With the integration of DVS, this processor enables real-time, low-power and accurate eye tracking for an immersive interaction experience on AR/VR devices. Shihang Tan, Quanshu Yan, Lirong Zheng 0001, Zhuo Zou |
ISCAS | 4 |
| 2024 | Spiking-HDC: A Spiking Neural Network Processor with HDC Classifier Enabling Transfer LearningabstractThis work proposes Spiking-HDC, a spiking neural network (SNN) processing system with hyperdimensional computing (HDC) and its hardware design for domain transfer scenarios. The input data is firstly fed into a two-layer SNN, serving as a feature extractor. It is followed by a HDC classifier to process feature vectors using hypervectors in binary representation. Such a system leverages HDC’s capability of single-pass learning, which can be adopted to rapidly updating but highly similar tasks by fine-tuning the HDC classifier with limited labeled data. By our experiments, the proposed system demonstrates transfer learning accuracy of 94.76%, 87.12% and 94.37% with few-shot samples on N-MNIST, DVS-Gesture and MNIST datasets, respectively. To apply Spiking-HDC model to extreme edge inference tasks, a dedicated processor is designed and implemented. The simulated results in 40 nm CMOS process illustrate that it has 0.88 mm2core area and 1.8 mW power at 100 MHz frequency. In comparison to similar works, it achieves 3.8×-36× inference energy efficiency enhancement. Anqin Xiao, Jinqiao Yang, Lirong Zheng 0001, Zhuo Zou |
ISCAS | 4 |
| 2024 | TSCM: A TCAM-Based Sparse Connection Memory Architecture in Neuromorphic Computing System for Cortical SimulationabstractThe connection matrix requires significant memory capacity in large-scale Spiking Neural Networks (SNNs). However, the sparsity of connections in cortical models leads to memory capacity and energy inefficiencies. This paper proposes Ternary Content Addressable Memory (TCAM)-based Sparse Connection Memory (TSCM) architecture in neuromorphic computing systems for cortical simulation. The architecture consists of a TCAM for searching existing synaptic connections, a Static Random Access Memory (SRAM) for storing synapse addresses, and a three-stage circuit for memory access control. By leveraging the sparsity, the proposed memory architecture demonstrates improved area and energy efficiency compared to the conventional Direct Mapped Full-Address Memory (DMFAM) architecture. A TSCM macro is designed, simulated using UMC 40-nm CMOS technology, and evaluated across various scales of classical cortical models. Experimental results demonstrate that the TSCM architecture reduces area by 27.6% to 75.6% and energy consumption by 15.8% to 96.0% in cortical simulation with neurons ranging from 100k to 10M compared to DMFAM architecture. Fanxi Yang, Yuhan He, Lirong Zheng 0001, Zhuo Zou |
ISCAS | 4 |
| 2024 | MCU-Enabled Epileptic Seizure Detection System With Compressed LearningabstractEpilepsy is one of the most common neurological disorder diseases all over the world, which gives patients a huge burden in seizure-related disabilities. For epileptic seizure detection, encephalography (EEG) is a commonly used clinical approach. Recently, several Internet of Things (IoT)-based wearable monitoring systems using machine learning (ML) approaches have been proposed to assist real-time detection of epileptic seizure attack scenarios. Among these approaches, convolutional neural networks (CNNs) provide superior accuracy, at the expense of high computational complexity that is not friendly to resource-constrained wearable devices. In this work, we propose a compressed learning (CL)-based epileptic seizure detection system which enables implementing CNN on microcontroller units (MCUs). The proposed CL approach combines the measurement matrix of compressed sensing (CS) and a 1-D CNN. As a result, the input data size and parameter number of CNN can be significantly reduced while eliminating the complex reconstruction process of the traditional CS approach. Evaluated on the Bonn dataset, our proposed system ensures a 96.44% accuracy under a 0.1 compressed ratio (CR) corresponding to a$33.5\times $multiply-accumulate operations (MACs) reduction and a$21.9\times $decrease in model size compared with the baseline. Mapping such a model on an MSP432 MCU, the memory requirement is 50.13 kB and the power consumption is 13.4727 mW at 40 MHz frequency, corresponding to an energy consumption of$269.4~\mu \text{J}$/Classification, with a classification latency of 21.07 ms. Liyu Qian, Yuxiang Huan, Yaojie Sun, Lirong Zheng 0001, Zhuo Zou |
IEEE Internet Things J. | 6 |
| 2024 | CorTile: A Scalable Neuromorphic Processing Core for Cortical Simulation With Hybrid-Mode Router and TCAMabstractIn neuromorphic processors, simulating large-scale Spiking Neural Networks (SNNs) for cortical models necessitates a significant increase in communication traffic and memory capacity, due to the lack of exploiting the sparsity of connections. Therefore, this paper proposes CorTile, a scalable neuromorphic processing core designed for cortical simulation. We propose a hybrid-mode router that supports Remote Unicast and Local Broadcast (RULB) routing method, leveraging the high local connectivity and low distal connectivity observed in cortical models. This approach achieves reductions of 36.7% in average router load, 40.7% in peak load, 51.2% in average link traffic, 41.7% in peak traffic, respectively, compared to conventional routing methods. Additionally, the proposed Ternary Content Addressable Memory (TCAM)-based Sparse Connection Memory (TSCM) architecture leads to 87.1% reduction in area and a 62.7% reduction in power consumption. These approaches effectively decrease communication traffic and mitigate the quadratic increase in memory requirements, achieving linear growth instead, thus achieving scalability. The proposed CorTile is simulated using UMC 40-nm CMOS process, occupying an area of 5.15 mm2, supporting a maximum of 8k neurons and 64M synapses. Evaluated using a typical macaque cortex model, it consumes 8.25 mW, with the router operating at 200 MHz and the other modules at 100 MHz. This design achieves an average router load of 12.33 Mpackets/s and peak link traffic of 21.16 MB/s. Thanks to the scalability of the proposed processing core that can be tiled into many-core processors, it paves the way for chiplets and multiple chip integration towards a brain-scale neuromorphic computing system. Fanxi Yang, Yuhan He, Jinqiao Yang, Anqin Xiao, Lufei Fan, Lirong Zheng 0001, Zhuo Zou |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2023 | Self-aware Collaborative Edge Inference with Embedded Devices for Task-oriented IIoTabstractThe computing and communication resources of embedded devices are constrained and heterogeneous, resulting in a low quality-of-experience for compute-intensive applications in task-oriented industrial Internet of Things (IIoT), such as edge inference. To address these challenges, we first propose a model partitioning-based self-aware collaborative edge inference framework. Furthermore, the throughput-aware collaborative inference algorithm is designed for typical IIoT scenario, stacking tasks. Via jointly optimizing the partition layer and collaborative device selection, the optimal inference efficiency, maximum inference throughput, can be obtained. Finally, the performance of our proposal is demonstrated by extensive simulations and tests based on 10 Raspberry Pi 4Bs and popular models. Specifically, with the proposed algorithm, our platform reaches up to 14.77× throughput speed up for stacking tasks, which indicates that the the proposed design can improve the inference efficiency. Zhuoquan Yu, Christine Mwase, Yi Jin 0007, Lirong Zheng 0001, Zhuo Zou |
VTC Fall | 6 |
| 2023 | A Low-Power Hybrid-Precision Neuromorphic Processor With INT8 Inference and INT16 Online Learning in 40-nm CMOSabstractIn this work, we present a neuromorphic processor for artificial intelligence of things (AIoT) applications featuring low-power consumption, a small footprint, STDP-based online learning, and the ability to adapt to multiple applications. Hybrid precision, i.e., INT8 for inference and INT16 for training, is suggested to achieve balanced accuracy and energy efficiency. A precision-configurable leaky integrate-and-fire(LIF) neuron unit and a unified memory architecture are designed to maximize datapath reuse. A dynamic pruning technique is proposed to exploit the temporal sparsity, yielding synaptic operations reduction by 3.68x in training and 1.63x in inference, respectively. The design is implemented and fabricated in a 40-nm CMOS process, with a core area of 0.87 mm2. It is measured to consume a minimal power of$680~\mu \text{W}$at 70 MHz under a 0.75 V power supply, corresponding to 9.9 pJ per synaptic operation. Evaluated with typical spatial, temporal, and spatiotemporal datasets (MNIST, MIT-BIH, and N-MNIST), the proposed design achieve energy efficiency comparable to the best-in-class solutions with handcrafted training and customized ASICs, while demonstrating improved versatility across multiple applications with balanced accuracy, power consumption, and model adaptability. Congyang Liu, Ziyi Yang 0014, Zikai Zhu, Haoming Chu, Yuxiang Huan, Lirong Zheng 0001, Zhuo Zou |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2023 | ASLog: An Area-Efficient CNN Accelerator for Per-Channel Logarithmic Post-Training QuantizationabstractPost-training quantization (PTQ) has been proven an efficient model compression technique for Convolution Neural Networks (CNNs), without re-training or access to labeled datasets. However, it remains challenging for a CNN accelerator to fulfill the efficiency potential of PTQ methods. A large number of PTQ techniques blindly pursue high theoretic compression effect and accuracy, ignoring their impact on the actual hardware implementation, which causes more hardware overhead than benefit. This paper introduces ASLog, a PTQ-friendly CNN accelerator that explores four key designs in an algorithm-hardware co-optimizing manner: the first practical 4-bit logarithmic PTQ pipeline SLogII, the multiplier-free arithmetic element (AE) design, the energy-efficient bias correction element (BCE) design, and the per-channel quantization friendly (PCF) architecture and dataflow. The proposed SLogII PTQ pipeline can push the limit of logarithmic PTQ to 4-bit with40% lower in power and area consumption compared with a common 8-bit multiplier. The BCE and PCF design proposed in this paper are the first to consider the hardware impact of the widely-used per-channel quantization and bias correction technique, enabling an efficient PTQ-friendly implementation with a small hardware overhead. The ASLog is validated in a UMC 40-nm process, with 12.2 TOPS/W energy efficiency and 0.80 mm2 core area. The ASLog can achieve 336.3 GOPS/mm2 area efficiency and >500 OPs/Byte operational intensity, which map to over$1.85\times $and$1.12\times $improvement compared with the previous related works. Jiawei Xu 0002, Jiangshan Fan, Baolin Nan, Chen Ding 0010, Lirong Zheng 0001, Zhuo Zou, Yuxiang Huan |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2023 | A Domain-Specific Accelerator for Ultralow Latency Market Data Distribution SystemabstractUltralow latency parsing of financial data is gaining significance in the high-frequency trading of the security exchange market. Hardware-aided systems exhibit superior improvement of latency over traditional software solutions, but flexibility may suffer when processing the financial protocol. This article presents a domain-specific accelerator for the market data distribution system, which integrates a financial information exchange adapted for streaming (FAST) decoder, a 10-Gbps network interface, and a high-speed PCIe host interface into a single field programmable gate array (FPGA) acceleration card. The proposed FAST decoder adopts the finite state machines-coordinated sequence mapping table to achieve run-time reconfigurability over fine-grained FPGA programming, and 16 fields can be decoded simultaneously in a pipelined manner, resulting in a decoding latency of only 33 ns. Evaluated on the Xilinx Alveo U200 acceleration card, this work improves the latency of decoding a FAST message by 26%–72% than state-of-the-art FPGA designs, and outperforms the software baseline by$>27\times$in terms of latency covering both decoding and communication. Yuxiang Huan, Chen Ding 0010, Yulong Yan, Jianjun Cui, Jiachen Wang 0008, Chuhuang Cai, Zhuo Zou, Lirong Zheng 0001 |
IEEE Trans. Ind. Informatics | 10 |
| 2023 | An IoT-Based Wearable Labor Progress Monitoring System for Remote Evaluation of Admission Time to HospitalabstractBecause contractions signal the approach of labor, pregnant women-especially primigravidas (i.e., women pregnant for the first time)-usually go to the hospital to seek medical intervention when they begin experiencing contractions, which is not conductive to good perinatal outcomes. Conventionally, uterine contraction monitoring requires specialized medical devices and relies on the doctor's clinical experience. Therefore, exploring an objective method to detect labor onset at home and avoid early hospital admission has essential importance. In this article, a labor progress monitoring system based on a sensing device, edge service, and Internet of things (IoT) platform is proposed, aiming to suggest suitable hospital admission times for low-risk primigravidas. The pregnant woman places the sensing device on her abdomen with the help of a belt to detect contraction activities. An intelligent edge service for contraction classification is deployed on a mobile phone. The system's artificial intelligence (AI)-assisted algorithm is lightweight, with 670 kB and 194 kB of memory dedicated to a convolutional neural network and long short-term memory, respectively. It classifies the pregnant woman as deferred admission, optional admission, or recommended admission according to different contraction states. An IoT platform connected to the hospital is implemented, providing professional suggestions from doctors. The test set collected in an emergency clinic shows that the proposed system can reach a classification accuracy of more than 96%. In conclusion, the proposed system enables remote labor progress monitoring at home and avoids early hospital admission. Zhiqing Xiao, Lihua Xu, Zhuo Zou, Lirong Zheng 0001 |
IEEE J. Biomed. Health Informatics | 6 |
| 2022 | A Hybrid-Mode On-Chip Router for the Large-Scale FPGA-Based Neuromorphic PlatformabstractLarge-scale neuromorphic computing requires the multi-chip network to provide high computing power. Efficient routing schemes and on-chip router design are necessary for handling various inter-chip transmission patterns. In this paper, we propose a hybrid-mode on-chip router that supports both multicast and unicast routing for the large-scale neuromorphic simulation. Two routing schemes, namely Cache-like Spike Weight Indexing and General Unicast Flow Control, are proposed to accommodate the chip-to-chip transmission of spike and non-spike data. This work is evaluated on a neuromorphic platform built with an$8\times 8$FPGA chips array. Running a simulation of 1M neurons at 200MHz, the proposed router achieves a processing latency of 25ns and a chip-to-chip latency of 287ns. Working in the unicast mode, the router can synchronize status flags of all chips within$5 ~\mu \text{s}$. Moreover, it reduces the peak spike traffic by 25.65% with the help of Load-aware Multicast Routing, compared with other multicast routing strategies. Chen Ding 0010, Yuxiang Huan, Yulong Yan, Fanxi Yang, Lizheng Liu, Meigen Shen, Zhuo Zou, Lirong Zheng 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 9 |
| 2022 | Edge-Based Collaborative Training System for Artificial Intelligence-of-ThingsabstractThe descending of intelligence from the cloud to the heterogeneous and low-power edge in the Artificial Intelligence-of-Things prevents uploading user-sensitive information to the cloud. It brings an urgent demand for deploying training tasks collaboratively in industrial scenarios to manage data locally. This article proposes an edge-based collaborative training system for the smart factory which harnesses the intelligence of edge devices by balancing the computational and communicational resources and improving system dependability. Two typical scenarios of parts recognition and defect inspection are evaluated as a case study with our system. The feasibility and dependability of the presented system are verified with a platform composed of eight high-performance (Nvidia Jetson Nano) and eight low-performance edge devices (Raspberry Pi 4B). The efficiency under tradeoff between computational resource and network condition constraints in a cluster is tested to simulate real-case performance in smart factory scenarios. Our platform reaches the peak performance of 1167 images/s training efficiency on ResNet32 under a 125 MB/s bandwidth. Experimental results demonstrate that the proposed design can collaboratively perform training tasks with optimized efficiency and provide dependable collaborations for system fault detection and cluster extension. Yi Jin 0007, Yulong Yan, Yuxiang Huan, Jiawei Xu 0002, Shancang Li, Prosanta Gope, Zhuo Zou, Lirong Zheng 0001 |
IEEE Trans. Ind. Informatics | 10 |
| 2021 | Design Framework for SRAM-Based Computing-In-Memory Edge CNN AcceleratorsabstractThis paper presents an architectural framework and an evaluation model for Static Random Access Memory (SRAM)-based Computing-in-Memory (CIM) edge Convolutional Neural Network (CNN) accelerators. To provide a baseline for system-level design perspectives, an architectural framework for SRAM-CIM design concerning the key design points in state-of-the-art works is proposed. Furthermore, a configurable evaluation model featuring top-down design flow based on the proposed framework is established to investigate design space explorations. Case studies validated the framework and evaluation model using LeNet-5, AlexNet and VGG-16 to achieve energy-aware optimizations. The optimized memory scale for LeNet-5 is "16 PEs and 120 tiles" with the minimal estimated inference energy of 0.0018J, while for AlexNet and VGG-16, "16 PEs and 120 tiles" is better achieving minimal energy consumption of 0.1733mJ and 0.6825mJ respectively. Estimation results highlight tradeoffs among data represent- tation parameters and memory partitioning parameters. This work provides specific SRAM-CIM design guidelines from a system-level perspective. Yimin Wang 0001, Zhuo Zou, Lirong Zheng 0001 |
ISCAS | 3 |
| 2021 | Self-aware distributed deep learning framework for heterogeneous IoT edge devices
Yi Jin 0007, Jiawei Cai, Jiawei Xu 0002, Yuxiang Huan, Yulong Yan, Yongliang Guo, Lirong Zheng 0001, Zhuo Zou |
Future Gener. Comput. Syst. | 8 |
| 2021 | An IoT-Based Anti-Counterfeiting System Using Visual Features on QR CodeabstractThis article presents an Internet-of-Things (IoT) anti-counterfeiting system that uses visual features combined with the quick response (QR) code. The visual features guarantee the authenticity of a product with the QR code for tracking and tracing. Two visual features, i.e., natural texture features and printed micro features are exploited in the proposed system. The natural texture features use the texture of fiber paper to achieve physical unclonable function (PUF), while the micro features are artificially generated for improved industrial manufacturability and reliability. Features are generated and registered in the production phase when the QR code is printed. In the anti-counterfeiting verification phase, the feature obtained through the feature extraction algorithm is compared with the record to calculate similarity, which indicates the verification result. Such an approach is fully compatible with the QR code-based logistic process without any additional manufacturing cost. A user-friendly application has been developed on a mobile platform that facilitates easy-to-use and affordable devices for verification, such as a mobile phone or a handheld code reader. The experimental results show 99.6% and 99.9% accuracy of anti-counterfeiting verification for texture features and micro features, respectively. The system with corresponding algorithms and software has been demonstrated in real-life products. Yulong Yan, Zhuo Zou, Yu Gao 0042, Lirong Zheng 0001 |
IEEE Internet Things J. | 5 |
| 2021 | IECA: An In-Execution Configuration CNN Accelerator With 30.55 GOPS/mm² Area EfficiencyabstractIt remains challenging for a Convolutional Neural Network (CNN) accelerator to maintain high hardware utilization and low processing latency with restricted on-chip memory. This paper presents an In-Execution Configuration Accelerator (IECA) that realizes an efficient control scheme, exploring architectural data reuse, unified in-execution controlling, and pipelined latency hiding to minimize configuration overhead out of the computation scope. The proposed IECA achieves row-wise convolution with tiny distributed buffers and reduces the size of total on-chip memory by removing 40% of redundant memory storage with shared delay chains. By exploiting a reconfigurable Sequence Mapping Table (SMT) and Finite State Machine (FSM) control, the chip realizes cycle-accurate Processing Element (PE) control, automatic loop tiling and latency hiding without extra time slots for pre-configuration. Evaluated on AlexNet and VGG-16, the IECA retains over 97.3% PE utilization and over 95.6% memory access time hiding on average. The chip is designed and fabricated in a UMC 55-nm process running at a frequency of 250 MHz and achieves an area efficiency of 30.55 GOPS/mm2and 0.244 GOPS/KGE (kilo-gate-equivalent), which makes an over$2.0\times $and$2.1\times $improvement, respectively, compared with that of previous related works. Implementation of the IEC control scheme uses only a 0.55% area of the 2.75 mm2core. Boming Huang, Yuxiang Huan, Haoming Chu, Jiawei Xu 0002, Lizheng Liu, Lirong Zheng 0001, Zhuo Zou |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2021 | A Wearable Hand Rehabilitation System With Soft GlovesabstractHand paralysis is one of the most common complications in stroke patients, which severely impacts their daily lives. This article presents a wearable hand rehabilitation system that supports both mirror therapy and task-oriented therapy. A pair of gloves, i.e., a sensory glove and a motor glove, was designed and fabricated with a soft, flexible material, providing greater comfort and safety than conventional rigid rehabilitation devices. The sensory glove worn on the nonaffected hand, which contains the force and flex sensors, is used to measure the gripping force and bending angle of each finger joint for motion detection. The motor glove, driven by micromotors, provides the affected hand with assisted driving force to perform training tasks. Machine learning is employed to recognize the gestures from the sensory glove and to facilitate the rehabilitation tasks for the affected hand. The proposed system offers 16 kinds of finger gestures with an accuracy of 93.32%, allowing patients to conduct mirror therapy using fine-grained gestures for training a single finger and multiple fingers in coordination. A more sophisticated task-oriented rehabilitation with mirror therapy is also presented, which offers six types of training tasks with an average accuracy of 89.4% in real time. Xiaoshi Chen, Shih-Ching Yeh, Lirong Zheng 0001, Zhuo Zou |
IEEE Trans. Ind. Informatics | 6 |
| 2020 | An Autonomous Error-Tolerant Architecture Featuring Self-reparation for Convolutional Neural NetworksabstractConvolutional neural networks are widely used in artificial intelligence and Internet of Things area. As the scale of convolutional neural network expands, more and more processing units are provided for it. The systems are easy prone to error, and any computing problems in any layer of the network will lead to wrong output results. Traditional multimode redundancy methods make the systems more complex, and increase power consumption. This paper proposes an autonomous error-tolerant architecture for convolutional neural networks. Taking the LeNet-5 as an example, the network layers of CNN are mapped on the AET architecture, an error-tolerant synapse is designed to discover the errors, an active evolution scheme is designed to handle unrecoverable errors and implement network reconfiguration. This design is implemented on FPGA, and the experimental results show that this architecture can realize effective error tolerance for convolutional neural network and has fast error recovery ability under the premise of ensuring the same recognition accuracy. Lizheng Liu, Yuxiang Huan, Zhuo Zou, Xiaoming Hu 0001, Lirong Zheng 0001 |
VTC Spring | 5 |
| 2020 | A Smart Dental Health-IoT Platform Based on Intelligent Hardware, Deep Learning, and Mobile TerminalabstractThe dental disease is a common disease for a human. Screening and visual diagnosis that are currently performed in clinics possibly cost a lot in various manners. Along with the progress of the Internet of Things (IoT) and artificial intelligence, the internet-based intelligent system have shown great potential in applying home-based healthcare. Therefore, a smart dental health-IoT system based on intelligent hardware, deep learning, and mobile terminal is proposed in this paper, aiming at exploring the feasibility of its application on in-home dental healthcare. Moreover, a smart dental device is designed and developed in this study to perform the image acquisition of teeth. Based on the data set of 12 600 clinical images collected by the proposed device from 10 private dental clinics, an automatic diagnosis model trained by MASK R-CNN is developed for the detection and classification of 7 different dental diseases including decayed tooth, dental plaque, uorosis, and periodontal disease, with the diagnosis accuracy of them reaching up to 90%, along with high sensitivity and high specificity. Following the one-month test in ten clinics, compared with that last month when the platform was not used, the mean diagnosis time reduces by 37.5% for each patient, helping explain the increase in the number of treated patients by 18.4%. Furthermore, application software (APPs) on mobile terminal for client side and for dentist side are implemented to provide service of pre-examination, consultation, appointment, and evaluation. Lizheng Liu, Jiawei Xu 0002, Yuxiang Huan, Zhuo Zou, Shih-Ching Yeh, Lirong Zheng 0001 |
IEEE J. Biomed. Health Informatics | 6 |
| 2018 | A 3D Tiled Low Power Accelerator for Convolutional Neural NetworkabstractIt remains a challenge to run Deep Learning in devices with stringent power budget in the Internet-of-Things. This paper presents a low-power accelerator for processing Convolutional Neural Networks on the embedded devices. The power reduction is realized by exploring data reuse in three different aspects, with regards to convolution, filter and input features. A systolic-like data flow is proposed and applied to rows of Processing Elements (PEs), which facilitate reusing the data during convolution. Reuse of input features and filters is achieved by arranging the PE array in a 3D tiled architecture, whose dimension is 3 × 14 × 4. Local storage within PEs is therefore reduced and only cost 17.75 kB, which is 20% of the state-of-the-art. With dedicated delay chains in each PE, this accelerator is reconfigurable to suit various parameter settings of convolutional layers. Evaluated in UMC 65 nm low leakage process, the accelerator can reach a peak performance of 84 GOPS and consume only 136 mW at 250 Mhz. Yuxiang Huan, Jiawei Xu 0002, Lirong Zheng 0001, Hannu Tenhunen, Zhuo Zou |
ISCAS | 3 |
| 2018 | TMR Group Coding Method for Optimized SEU and MBU Tolerant Memory DesignabstractThis work proposes a fault tolerant memory design using the method of Triple Module Redundancy (TMR) group coding to tolerant the Single-Event Upset (SEU) and Multi-Bit Upset (MBU) influence on memory devices in space environment. The group coding method uses different models to partition and code each word line in memory with Hamming code to achieve best performance. TMR group coding method further increases the capability of self-correction for the errors occurred in parity bits. The evaluation results show that the suggested approach can obtain improved correctness for the memory output with optimized tradeoff between reliability and cost. At 5% error rate, the probability of correct output reaches 70.78% with small cost increment. To achieve 90% reliability, the accuracy improvement is 31.9% compared to TMR with 9% increased area. This solution proposed is evaluated on the memory rich micro-coded processor, but can be further extended to other memory-based processors that need high reliability for the SEU and MBU influence in aerospace applications. Yi Jin 0007, Yuxiang Huan, Haoming Chu, Zhuo Zou, Lirong Zheng 0001 |
ISCAS | 5 |
| 2018 | A Design of Autonomous Error-Tolerant Architectures for Massively Parallel Computing
Lizheng Liu, Yi Jin 0007, Yi Liu 0027, Yuxiang Huan, Zhuo Zou, Lirong Zheng 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2017 | Smart energy efficient gateway for Internet of mobile thingsabstractInternet of Things (IoT) is a fast developing vision in which physical quantities are digitized, processed and analyzed. Internet of Mobile Things (IoMT) as one of new domains of IoT, due to mobility, requires a more demanding and rigorous solution in many aspects, especially in terms of energy efficiency. We propose a solution consisting of energy efficient and fast hardware platform for building IoMT Fog layer facilities. Experimental results are presented to prove superiority of the proposed hardware in several aspects to popular general purpose platforms. Igor Tcarenko, Yuxiang Huan, David Juhasz, Amir-Mohammad Rahmani, Zhuo Zou, Tomi Westerlund, Pasi Liljeberg, Lirong Zheng 0001, Hannu Tenhunen |
CCNC | 8 |
| 2016 | A 101.4 GOPS/W reconfigurable and scalable control-centric embedded processor for domain-specific applicationsabstractIncreasing the energy efficiency and performance while providing the customizability and scalability is vital for embedded processors adapting to domain-specific applications such as Internet of Things. In this paper, we proposed a reconfigurable and scalable control-centric architecture, and implemented the design consisting of two cores and an on-chip multi-mode router in 65 nm technology. The reconfigurability is enabled by the restructurable sequence mapping table (SMT) thus the reorganizable functional units. Owing to the integration of the multi-mode router, on-chip or inter-chip network for multi-/many-core computing can be composed for performance extension on demand even in the post-fabrication stage. Control-centric design simplifies the control logic, shrinks the non-functional units and orchestrates the operations to increase the hard are utilization and reduce the excessive data movement for high energy efficiency. As a result, the processor can both conduct general-purpose processing with 29% smaller code size and application-specific processing with over 10 times performance improvement when implementing AES by SMT. The dual-core processor consumes 19.7 μW/MHz with die size of 3.5 mm2. The achieved energy efficiency is 101.4GOPS/W. Zhuo Zou, Zhonghai Lu, Lirong Zheng 0001, Yuxiang Huan, Stefan Blixt |
ISCAS | 4 |
| 2016 | Design and implementation of multi-mode routers for large-scale inter-core networks
Zhuo Zou, Zhonghai Lu, Lirong Zheng 0001 |
Integr. | 4 |
| 2016 | Design and simulation of a standing wave oscillator based PLLabstractA standing wave oscillator (SWO) is a perfect clock source which can be used to produce a high frequency clock signal with a low skew and high reliability. However, it is difficult to tune the SWO in a wide range of frequencies. We introduce a frequency tunable SWO which uses an inversion mode metal-oxide-semiconductor (IMOS) field-effect transistor as a varactor, and give the simulation results of the frequency tuning range and power dissipation. Based on the frequency tunable SWO, a new phase locked loop (PLL) architecture is presented. This PLL can be used not only as a clock source, but also as a clock distribution network to provide high quality clock signals. The PLL achieves an approximately 50% frequency tuning range when designed in Global Foundry 65 nm 1P9M complementary metal-oxide-semiconductor (CMOS) technology, and can be used directly in a high performance multi-core microprocessor. Youde Hu, Lirong Zheng 0001 |
Frontiers Inf. Technol. Electron. Eng. | 3 |
| 2015 | Latency-optimized stochastic LDPC decoder for high-throughput applicationsabstractStochastic decoding can be applied to Low-Density Parity-Check codes in order to achieve high throughput with less area. However, most architectures suffer from large decoding latencies, due to the mechanism of stochastic computation. In this paper, three novel strategies, including the LUT-based initialization, the posterior-information-based hard decision and the Bit-Flipping-based post processing, are proposed in order to reduce decoding latency and hence improve throughput. For the standard IEEE 802.3an (2048, 1723) code, simulation indicates 75.7% reduction in average decoding cycles at 4.5 dB with satisfied bit error rate. Moreover, hardware implementation shows that the area of variable node units is reduced significantly in SMIC 65 nm technology. Di Wu 0016, Yun Chen 0001, Qichen Zhang, Lirong Zheng 0001, Xiaoyang Zeng, Yeong-Luh Ueng |
ISCAS | 4 |
| 2015 | Implementing MVC Decoding on Homogeneous NoCs: Circuit Switching or Wormhole SwitchingabstractTo implement multiview video decoding on network on-chip (NoC) based homogeneous multicore architectures, the selection of switching techniques for routers is one of the most important aspects for design space exploration. Circuit switching and wormhole switching are two most feasible switching techniques for on-chip networks. To choose the suitable switching technique, we perform the comparison on decoding speed of the whole system, link utilization and delay between circuit switching and wormhole switching for implementing eight-view QVGA video decoding on 4 × 4 NoCs at 30 fps. The required link bandwidths are both around 800 Mbps with the similar network utilization and delay. We conclude that, to implement multiview video decoding on homogeneous NoCs, circuit switching is more suitable considering the similar performance and lower cost compared with wormhole switching. Zhuo Zou, Zhonghai Lu, Lirong Zheng 0001 |
PDP | 4 |
| 2015 | An Implantable RFID Sensor Tag toward Continuous Glucose MonitoringabstractThis paper presents a wirelessly powered implantable electrochemical sensor tag for continuous blood glucose monitoring. The system is remotely powered by a 13.56-MHz inductive link and utilizes an ISO 15693 radio frequency identification (RFID) standard for communication. This paper provides reliable and accurate measurement for changing glucose level. The sensor tag employs a long-term glucose sensor, a winding ferrite antenna, an RFID front-end, a potentiostat, a 10-bit sigma-delta analog to digital converter, an on-chip temperature sensor, and a digital baseband for protocol processing and control. A high-frequency external reader is used to power, command, and configure the sensor tag. The only off-chip support circuitry required is a tuned antenna and a glucose microsensor. The integrated chip fabricated in SMIC 0.13-μm CMOS process occupies an area of 1.2 mm ×2 mm and consumes 50 μW. The power sensitivity of the whole system is -4 dBm. The sensor tag achieves a measured glucose range of 0-30 mM with a sensitivity of 0.75 nA/mM. Xi Tan 0001, Xianliang Chen, Sizheng Chen, Lirong Zheng 0001, Hao Min |
IEEE J. Biomed. Health Informatics | 10 |
| 2014 | A wirelessly-powered UWB sensor tag with time-domain sensor interfaceabstractThis paper presents a wirelessly-powered sensor tag with a time-domain sensor interface for wireless sensing applications. The tag is remotely powered by RF wave. Instead of traditional approaches employing conventional ADCs for quantization and transmitter for data communication, in this work, a Pulse Position Modulator incorporating simple impulse radio UWB (IR-UWB) transmitter is proposed to convert and transmit the analog sensing information in time domain. The analog signal is compared with an adjustable triangular wave for analog to time conversion in signal-varying environments. Then a UWB transmitter converts the PPM signal to very short pulses and sends it back to the reader. The time interval of UWB pulses represents the original input signal in time domain which can be measured on the reader side by a time-to-digital conversion. This approach not only simplifies the ADC design but also relaxes the number of bits transmitted on the tag side. The sensor tag is designed in 180nm CMOS process. Simulation results demonstrate that the proposed approach reduce transmission power consumption by nearly 3 orders of magnitude over traditional approaches, while consuming only 85 µW for 1.5 MS/s sampling rate. Dongxuan Bao, Zhuo Zou, Majid Baghaei Nejad, Lirong Zheng 0001 |
ISCAS | 5 |
| 2014 | A Health-IoT Platform Based on the Integration of Intelligent Packaging, Unobtrusive Bio-Sensor, and Intelligent Medicine BoxabstractIn-home healthcare services based on the Internet-of-Things (IoT) have great business potential; however, a comprehensive platform is still missing. In this paper, an intelligent home-based platform, the iHome Health-IoT, is proposed and implemented. In particular, the platform involves an open-platform-based intelligent medicine box (iMedBox) with enhanced connectivity and interchangeability for the integration of devices and services; intelligent pharmaceutical packaging (iMedPack) with communication capability enabled by passive radio-frequency identification (RFID) and actuation capability enabled by functional materials; and a flexible and wearable bio-medical sensor device (Bio-Patch) enabled by the state-of-the-art inkjet printing technology and system-on-chip. The proposed platform seamlessly fuses IoT devices (e.g., wearable sensors and intelligent medicine packages) with in-home healthcare services (e.g., telemedicine) for an improved user experience and service efficiency. The feasibility of the implemented iHome Health-IoT platform has been proven in field trials. Geng Yang 0003, Matti Mäntysalo, Zhibo Pang, Sharon Kao-Walter, Qiang Chen 0014, Lirong Zheng 0001 |
IEEE Trans. Ind. Informatics | 9 |
| 2013 | A Hybrid Low Power Biopatch for Body Surface Potential MeasurementabstractThis paper presents a wearable biopatch prototype for body surface potential measurement. It combines three key technologies, including mixed-signal system on chip (SoC) technology, inkjet printing technology, and anisotropic conductive adhesive (ACA) bonding technology. An integral part of the biopatch is a low-power low-noise SoC. The SoC contains a tunable analog front end, a successive approximation register analog-to-digital converter, and a reconfigurable digital controller. The electrodes, interconnections, and interposer are implemented by inkjet-printing the silver ink precisely on a flexible substrate. The reliability of printed traces is evaluated by static bending tests. ACA is used to attach the SoC to the printed structures and form the flexible hybrid system. The biopatch prototype is light and thin with a physical size of 16 cm × 16 cm. Measurement results show that low-noise concurrent electrocardiogram signals from eight chest points have been successfully recorded using the implemented biopatch. Geng Yang 0003, Jian Chen 0001, Jia Mao, Hannu Tenhunen, Lirong Zheng 0001 |
IEEE J. Biomed. Health Informatics | 6 |
| 2012 | A multi-parameter bio-electric ASIC sensor with integrated 2-wire data transmission protocol for wearable healthcare systemabstractThis paper presents a fully integrated application specific integrated circuit (ASIC) sensor for the recording of multiple bio-electric signals. It consists of an analog front-end circuit with tunable bandwidth and programmable gain, a 6-input 8-bit successive approximation register analog to digital converter (SAR ADC), and a reconfigurable digital core. The ASIC is fabricated in a 0.18-µm 1P6M CMOS technology, occupies an area of 1.5 × 3.0 mm2, and totally consumes a current of 16.7 µA from a 1.2 V supply. Incorporated with the ASIC, an Intelligent Electrode can be dynamically configured for on-site measurement of different bio-signals. A 2-wire data transmission protocol is also integrated on chip. It enables the serial connection over a group of Intelligent Electrodes, thus minimizes the number of connecting cables. A wearable healthcare system is built upon a printed Active Cable and a scalable number of Intelligent Electrodes. The system allows synchronous processing of maximum 14-channel bio-signals. The ASIC performance has been successfully verified in in-vivo bio-electric recording experiments. Geng Yang 0003, Jian Chen 0001, Fredrik Jonsson, Hannu Tenhunen, Lirong Zheng 0001 |
DATE | 5 |
| 2012 | A high-resolution Time-to-Digital Converter based on parallel delay elementsabstractThis paper presents a flash-type Time to Digital Converter (TDC) based on parallel delay elements in 65-nm CMOS process technology. By using parallel delay elements the conversion resolution of the TDC becomes equal to the difference of delay elements rather than the delay time of each element. A Sensed Amplifier Flip Flop (SAFF) ensures narrow sampling window. Operating at 1.2-V supply, this TDC shows 3ps resolution with 0.5LSB of INL and 0.33LSB of DNL respectively and consumes average power 442μW. Fredrik Jonsson, Jian Chen 0001, Lirong Zheng 0001 |
ISCAS | 4 |
| 2012 | Bio-Patch Design and Implementation Based on a Low-Power System-on-Chip and Paper-Based Inkjet Printing TechnologyabstractThis paper presents the prototype implementation of a Bio-Patch using fully integrated low-power System-on-Chip (SoC) sensor and paper-based inkjet printing technology. The SoC sensor is featured with programmable gain and bandwidth to accommodate a variety of bio-signals. It is fabricated in a 0.18-ìm standard CMOS technology, with a total power consumption of 20 ìW from a 1.2 V supply. Both the electrodes and interconnections are implemented by printing conductive nano-particle inks on a flexible photo paper substrate using inkjet printing technology. A Bio-Patch prototype is developed by integrating the SoC sensor, a soft battery, printed electrodes and interconnections on a photo paper substrate. The Bio-Patch can work alone or operate along with other patches to establish a wired network for synchronous multiple-channel bio-signals recording. The measurement results show that electrocardiogram and electromyogram are successfully measured in in-vivo tests using the implemented Bio-Patch prototype. Geng Yang 0003, Matti Mäntysalo, Jian Chen 0001, Hannu Tenhunen, Lirong Zheng 0001 |
IEEE Trans. Inf. Technol. Biomed. | 6 |
| 2011 | Evaluating Sustainability, Environmental Assessment and Toxic Emissions during Manufacturing Process of RFID Based SystemsabstractThe present state of the art research in the direction of embedded systems demonstrate that analysis of life-cycle, sustainability and environmental assessment have not been a core focus for researchers. To maximize a researcher's contribution in formulating environmentally friendly products, devising green manufacturing processes and services, there is a strong need to enhance life-cycle awareness and sustainability understandings among embedded systems researchers, so that the next generation of engineers will be able to realize the goal of a sustainable life-cycle. In this work an attempt has been made to investigate and evaluate the life-cycle management and environmental assessment in fabricating processes of the RFID based systems. We have chosen a general life cycle assessment approach which involves the collection and evaluation of quantitative data on the inputs and outputs of materials and energy associated with the RFID based systems. Based on the developed generic models, we have obtained the results in terms of environmental emissions for a production of paper substrate printed RFID antennas. We also make an attempt to raise several sustainability issues and quantify the toxic emissions during the manufacturing process. Rajeev Kumar Kanth, Pasi Liljeberg, Hannu Tenhunen, Qiansu Wan, Yasar Amin, Botao Shao, Qiang Chen 0014, Lirong Zheng 0001 |
DASC | 8 |
| 2011 | Analog front-end RX design for UWB impulse radio in 90nm CMOSabstractIn this paper a reconfigurable differential Ultra Wideband-Impulse Radio (UWB-IR) energy receiver architecture has been simulated and implemented in UMC 90nm. The signal is amplified, rectified and integrated. By using an integration windowed scheme the SNR requirements are relaxed increasing the sensitivity. The design has been optimized for large bandwidths, low implementation area and configurability. The RX can be adapted to work at different data rates, processing gains, and channel environments. It works between the 3.1 – 4.8 GHz bands with OOK or PPM modulation with a tunable data rate up to 33Mb/s. In order to relax the ADC sampling time an interleave mode of operation has been implemented. It has a maximum power consumption of 22m W with a power supply of 1V. The complete RX occupies an area of 1.11mm2. David Sarmiento M., Zhuo Zou, Jia Mao, Peng Wang 0092, Fredrik Jonsson, Lirong Zheng 0001 |
ISCAS | 7 |
| 2011 | Stochastic coverage in event-driven sensor networksabstractOne of the primary tasks of sensor networks is to detect events in a field of interest (FoI). To quantify how well events are detected in such networks, coverage of events is a fundamental problem to be studied. However, traditional studies mostly focus on analyzing the coverage of the FoI, which is usually called area coverage. In this paper, we propose an analytic method to evaluate the performance of event coverage in sensor networks with randomly deployed sensor nodes and stochastic event occurrences. We provide formulas to calculate the probabilities of event coverage and event missing. The numerical results show how these two probabilities change with the sensor and event densities. Moreover, simulations are conducted to validate the analytic method. This method can provide guidelines for determining the amount of sensor nodes to achieve a certain level of coverage in event-driven sensor networks. Huimin She, Zhonghai Lu, Axel Jantsch, Dian Zhou, Lirong Zheng 0001 |
PIMRC | 5 |
| 2011 | Modeling and analysis of Rayleigh fading channels using stochastic network calculusabstractDeterministic network calculus (DNC) is not suitable for deriving performance guarantees for wireless networks due to their inherently random behaviors. In this paper, we develop a method for Quality of Service (QoS) analysis of wireless channels subject to Rayleigh fading based on stochastic network calculus. We provide closed-form stochastic service curve for the Rayleigh fading channel. With this service curve, we derive stochastic delay and backlog bounds. Simulation results verify that the bounds are reasonably tight. Moreover, through numerical experiments, we show the method is not only capable of deriving stochastic performance bounds, but also can provide guidelines for designing transmission strategies in wireless networks. Huimin She, Zhonghai Lu, Axel Jantsch, Dian Zhou, Lirong Zheng 0001 |
WCNC | 5 |
| 2010 | COSMO: CO-Simulation with MATLAB and OMNeT++ for Indoor Wireless NetworksabstractSimulations are widely used to design and evaluate new protocols and applications of indoor wireless networks. However, the available network simulation tools face the challenges of providing accurate indoor channel models, three-dimensional (3-D) models, model portability, and effective validation. In order to overcome these challenges, this paper presents a new CO-Simulation framework based on MATLAB and OMNeT++ (COSMO) to rapidly build credible simulations for indoor wireless networks. A hierarchical ad hoc passive RFID network for indoor tag locating is described as a case study, demonstrating the significance and efficiency of COSMO compared with other network simulators. COSMO surpasses other network simulators in terms of workload and validity. Zhonghai Lu, Qiang Chen 0014, Xiaolang Yan, Lirong Zheng 0001 |
GLOBECOM | 5 |
| 2010 | A Low Delay Multiple Reader Passive RFID System Using Orthogonal TH-PPM IR-UWBabstractNA Zhonghai Lu, Zhibo Pang, Xiaolang Yan, Qiang Chen 0014, Lirong Zheng 0001 |
ICCCN | 6 |
| 2010 | A switch mode resonating H-Bridge polar transmitter using RF ΣΔ modulationabstractUsing saturated power amplifier (PA) as the last stage, polar transmitter has the potential to be the most power efficient architecture to transmit large Peak-to-Average Ratio (PAR) signals. In this work, a polar transmitter using H-Bridge configured Class-D amplifiers is proposed. To fully exploit low voltage resource, maintain linearity and meet the spectrum mask requirements, RF Sigma-Delta Modulation (SDM) is used. An on-chip transformer based filter network is designed to filter out SDM noise and provide load matching. The system verification is carried out by using Matlab passband simulation on a 13dB PAR mobile WiMAX signal. Evaluation of noise shaping and spectral regrowth shows the proposed architecture can achieve -45dBc/10kHz ACPR in a 140MHz bandwidth range. This provides a solid ground for the circuit design work. Liang Rong, Fredrik Jonsson, Lirong Zheng 0001 |
ISCAS | 3 |
| 2009 | Analytical Evaluation of Retransmission Schemes in Wireless Sensor NetworksabstractRetransmission has been adopted as one of the most popular schemes for improving transmission reliability in wireless sensor networks. Many previous works have been done on reliable transmission issues in experimental ways, however, there still lack of analytical techniques to evaluate these solutions. Based on the traffic model, service model and energy model, we propose an analytical method to analyze the delay and energy metrics of two categories of retransmission schemes: hop-by- hop retransmission (HBH) and end-to-end retransmission (ETE). With the experiment results, the maximum packet transfer delay and energy efficiency of these two scheme are compared in several scenarios. Moreover, the analytical results of transfer delay are validated through simulations. Our experiments demonstrate that HBH has less energy consumption at the cost of larger transfer delay compared with ETE. With the same target success probability, ETE is superior on the delay metric for low bit-error- rate (BER) cases, while HBH is superior for high BER cases. Huimin She, Zhonghai Lu, Axel Jantsch, Dian Zhou, Lirong Zheng 0001 |
VTC Spring | 5 |
| 2009 | Two-Dimensional and Three-Dimensional Integration of Heterogeneous Electronic Systems Under Cost, Performance, and Technological ConstraintsabstractPresent day market demand for high-performance high-density portable hand-held applications has shifted the focus from 2-D planar system-on-a-chip-type single-chip solutions to alternatives such as tiled silicon and single-level embedded modules as well as 3-D die stacks. Among the various choices, finding an optimal solution for system implementation deals usually with cost, performance, power, thermal, and technological tradeoff analyses at the system conceptual level. It has been estimated that decisions made in the first 20% of the design cycle influence up to 80% of the final product cost. In this paper, we discuss realistic metrics appropriate for performance and cost tradeoff analyses both at the system conceptual level in the early stages of the design cycle and in the implementation phase, for verification. In order to validate the proposed metrics and methodology, two ubiquitous electronic systems are analyzed under various implementation schemes and the performance tradeoffs discussed. This case study is used to highlight the importance of a cost and performance tradeoff analysis early in the design flow. Roshan Weerasekera, Dinesh Pamunuwa, Lirong Zheng 0001, Hannu Tenhunen |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2008 | Digital calibration of gain and linearity in a CMOS RF mixerabstractThis paper presents a new digital calibration technique that allows CMOS Gilbert cell down-conversion mixers to meet their block specifications under large process, temperature and power supply variations. The gain and IIP3 are calibrated by regulating the current of the input differential pair and by switching the loads. IIP2 calibration is achieved by using a novel technique that consists of offset voltages cancellation in the switching pairs. The technique is tested by calibrating a 0.18um CMOS mixer in several corner conditions. It is found that by using this calibration technique, the Gilbert cell mixer can achieve yields comparable to digital circuits, hence making it amenable to AMS SoC integration. Saul Rodriguez 0001, Ana Rusu, Lirong Zheng 0001, Mohammed Ismail 0001 |
ISCAS | 3 |
| 2008 | Design and implementation of a fully reconfigurable chipless RFID tag using Inkjet printing technologyabstractIn this paper, a novel fully reconfigurable chipless RFID tag has been presented. 8-bit data are encoded by impedance mismatches along transmission line. Inkjet printing is used to reconfigure for special tag IDs. By integrating inkjet printing technology, printable tags are not only economically feasible but also technically efficient. The remarkable idea has been successfully validated both by simulation and experimental measurements. Linlin Zheng, Saul Rodriguez 0001, Lu Zhang 0011, Botao Shao, Lirong Zheng 0001 |
ISCAS | 5 |
| 2008 | Modeling of On-Chip Bus Switching Current and Its Impact on Noise in Power Supply GridabstractIn this paper, an analytical model for the current draw of an on-chip bus is presented. The model is combined with an on-chip power supply grid model in order to analyze noise caused by switching buses in a power supply grid. The bus is modeled as distributed resistance-inductance-capacitance (RLC) lines that are capacitively and inductively coupled to each other. Different switching patterns and driver skewing times are also included in the model. The power supply grid is modeled as a network ofRLCsegments. The model is verified by comparing it to HSPICE. The error was below 8%. The model is applied to determine the influence of driver skewing times on maximum power supply noise. Sampo Tuuna, Lirong Zheng 0001, Jouni Isoaho, Hannu Tenhunen |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2008 | Minimal-Power, Delay-Balanced Smart Repeaters for Global Interconnects in the Nanometer RegimeabstractA smart repeater is proposed for driving capacitively-coupled, global-length on-chip interconnects that alters its drive strength dynamically to match the relative bit pattern on the wires and thus the effective capacitive load. This is achieved by partitioning the driver into main and assistant drivers; for a higher effective load capacitance both drivers switch, while for a lower effective capacitance the assistant driver is quiet. In a UMC 0.18-mum technology the potential energy saving is around 10% and the reduction in jitter 20%, in comparison to a traditional repeater for typical global wire lengths. It is also shown that the average energy saving for nanometer technologies is in the range of 20% to 25%. The driver architecture exploits the fact that as feature sizes decrease, the capacitive load per transistor shrinks, whereas global wire loads remain relatively unchanged. Hence, the smaller the technology, the greater the potential saving. Roshan Weerasekera, Dinesh Pamunuwa, Lirong Zheng 0001, Hannu Tenhunen |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2007 | Extending systems-on-chip to the third dimension: performance, cost and technological tradeoffsabstractBecause of the today’s market demand for high- performance, high-density portable hand-held applications, elec- tronic system design technology has shifted the focus from 2-D planar SoC single-chip solutions to different alternative options as tiled silicon and single-level embedded modules as well as 3- D integration. Among the various choices, finding an optimal solution for system implementation dealt usually with cost, performance and other technological trade-off analysis at the system conceptual level. It has been identified that the decisions made within the first 20% of the total design cycle time will ultimately result upto 80% of the final product cost. In this paper, we discuss appropriate and realistic metric for performance and cost trade-off analysis both at system conceptual level (up-front in the design phase) and at implementation phase for verification in the three-dimensional integration. In order to validate the methodology, two ubiquitous electronic systems are analyzed under various implementation schemes and discuss the pros and cons of each of them. Roshan Weerasekera, Lirong Zheng 0001, Dinesh Pamunuwa, Hannu Tenhunen |
ICCAD | 2 |
| 2007 | A Novel Passive Tag with Asymmetric Wireless Link for RFID and WSN ApplicationsabstractIn this paper, we present a radio-powered module with asymmetric wireless link utilizing ultra wideband radio system for RFID and wireless sensor applications. Our contribution includes using two different standards in uplink and downlink. Such as conventional RFIDs, incoming RF signal transmitted by reader is used to power the internal circuitry and receive the data. However, in upstream link, an IR-UWB transmitter is utilized. Unlike traditional RFID systems, due to great advantages of UWB communication, this tag is very robust to multi-path fading and collision problem and it is more secure against eavesdropping or jamming. The module consists of a power scavenging unit, a RF receiver, an IR-UWB transmitter, digital baseband controller, and an embedded UWB antenna are designed for integration on liquid-crystal polymer (LCP) substrate, using 0.18mum CMOS process technology. Majid Baghaei Nejad, Zhuo Zou, Hannu Tenhunen, Lirong Zheng 0001 |
ISCAS | 4 |
| 2006 | An innovative receiver architecture for autonomous detection of ultra-wideband signalsabstractUltra wideband radio is an emerging wireless standard that uses sub-nanosecond pulses to transmit data, resulting in several GHz bandwidths. The problem of generating a synchronized template respect to the received signal grows in complexity as the signal bandwidth increases. In this paper, innovative, low cost, non-coherent receiver architecture is proposed for autonomous detection of ultra wideband signals. The receiver self-generates a synchronous template and hence, no transmitter-reference synchronizer is required. We validate its performance via simulations compared with coherent receivers and conventional non-coherent receivers, the new architecture is found much more robust to timing noise and hence greatly facilitates the synchronise problem in UWB receiver. Majid Baghaei Nejad, Lirong Zheng 0001 |
ISCAS | 2 |
| 2004 | A study on the implementation of 2-D mesh-based networks-on-chip in the nanometre regime
Dinesh Pamunuwa, Johnny Öberg, Lirong Zheng 0001, Mikael Millberg, Axel Jantsch, Hannu Tenhunen |
Integr. | 3 |
| 2004 | Interconnect intellectual property for Network-on-Chip (NoC)
Lirong Zheng 0001, Hannu Tenhunen |
J. Syst. Archit. | 2 |
| 2003 | Layout, Performance and Power Trade-Offs in Mesh-Based Network-on-Chip Architectures
Dinesh Pamunuwa, Johnny Öberg, Lirong Zheng 0001, Mikael Millberg, Axel Jantsch |
VLSI-SOC | 3 |
| 2003 | Maximizing throughput over parallel wire structures in the deep submicrometer regimeabstractIn a parallel multiwire structure, the exact spacing and size of the wires determine both the resistance and the distribution of the capacitance between the ground plane and the adjacent signal carrying conductors, and have a direct effect on the delay. Using closed-form equations that map the geometry to the wire parasitics and empirical switch factor based delay models that show how repeaters can be optimized to compensate for dynamic effects, we devise a method of analysis for optimizing throughput over a given metal area. This analysis is used to show that there is a clear optimum configuration for the wires which maximizes the total bandwidth. Additionally, closed form equations are derived, the roots of which give close to optimal solutions. It is shown that for wide buses, the optimal wire width and spacing are independent of the total width of the bus, allowing easy optimization of on-chip buses. Our analysis and results are valid for lossy interconnects as are typical of wires in submicron technologies. Dinesh Pamunuwa, Lirong Zheng 0001, Hannu Tenhunen |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2000 | Combating digital noise in high speed ULSI circuits using binary BCH encodingabstractIncreased integration in deep submicron (DSM) technologies has caused very high increases in the RLC parasitics which affect the coupling of noise to signals propagating over interconnect. Error free transmission on-chip will no longer be guaranteed, This paper examines the issue of high speed signaling in DSM and proposes the use of particular BCH codes to improve the bit error rate in the face of noise. We conclude from our results that it is possible to achieve a considerable coding gain by choosing the code properly. Dinesh Pamunuwa, Lirong Zheng 0001, Hannu Tenhunen |
ISCAS | 2 |
| 2000 | Efficient and accurate modeling of power supply noise on distributed on-chip power networksabstractIn this paper, we propose an efficient and accurate modeling technique for power supply noise estimation over on-chip power lines which are modeled as distributed RCL networks. With this model, peak noise on the power lines that includes on-chip resistive and inductive voltage drops, switching noise on packages, and on-chip decoupling effects, can be computed very efficiently and accurately. The model is verified by SPICE simulations. Lirong Zheng 0001, Hannu Tenhunen |
ISCAS | 1 |