EDBT 2026 Demo / reviewers in the wild / expert
Terry Tao Ye
dblp:12/6657-1 · also Terry Ye 0001
· DBLP profile ↗
28ranked-venue papers
5as first author
20since 2021 · last 2026
0000-0002-4359-3550ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 22 · 5 first-author · 15 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Computer networks · 2 · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Late Breaking Results: Algorithm-Hardware Co-Design of a Sparsity-Aware Dense-Sparse Scheme for DNN AcceleratorsabstractDeep neural networks (DNNs) in modern applications have increased the demand for energy-efficient DNN inference solutions, especially on resource-constrained platforms. However, the growing model capacity of DNNs incurs significant memory traffic and energy consumption. To address these challenges, we propose a novel solution that presents an algorithm-hardware co-design for reconfigurable DNN acceleration. This design exploits value- and bit-level sparsity to minimize memory footprint and enhance computational efficiency. To achieve this, the proposed algorithm leverages a static dense-sparse storage format, along with a dynamic bit-processing scheme that removes non-contributing bits. Building on this algorithm, a flexible processing element array is designed to perform LUT-based shift-accumulate operations, with fine-grained per-layer configurability. Experimental results show that this design yields 13–24% storage savings across the evaluated DNN models, while delivering up to 8.4× effective sparsity. Based on post-implementation FPGA results (from our RTL design), the proposed accelerator delivers 1.41× lower LUT usage than state-of-the-art design at similar throughput. Yueting Li 0001, Terry Tao Ye, Weisheng Zhao 0001 |
DATE | 3 |
| 2026 | ReNN-RV: Run-Time PE Reconfiguration for DNN Inference Acceleration With Custom RISC-V ISAabstractDeep neural network (DNN) accelerators integrated with RISC-V Instruction Set Architecture (ISA) extensions have enabled efficient computing on resource-constrained platforms. However, their specialization in regular compute patterns limits effectiveness on irregular workloads, making it challenging to achieve high throughput and energy efficiency. To tackle these challenges, we present ReNN-RV, which integrates a computation-aware RISC-V ISA extension with an instructiondriven processing pipeline to efficiently accelerate run-time reconfigurable processing elements (RePEs). The computation-aware ISA employs configurable opcodes and custom encodings to support fine-grained task scheduling, while an instructiondriven pipeline implements it with minimal control complexity. Moreover, theRePEaccelerator provides seamless switching between multiply-accumulate (MAC) and non-MAC operations by configuring a path multiplexer to realize multiple operators at run time. Experimental results demonstrate that ReNN-RV achieves average reductions of 14.6× in cycle count and 15.3× in execution time across representative DNN workloads compared with the baseline RISC-V design. On average, ReNN-RV outperforms state-of-the-art designs by 10.1× for energy efficiency and 10.3× for computational throughput. Yueting Li 0001, Terry Tao Ye, Ngai Wong 0001, Zhenhua Zhu 0002, Yongfu Li 0002, Weisheng Zhao 0001 |
IEEE Trans. Computers | 2 |
| 2026 | RV-WINO: A RISC-V Neural Network Accelerator Based on Winograd Algorithm Fabricated in 55-nm CMOS ProcessabstractThe rapid evolution of artificial intelligence (AI) in IoT applications necessitates the execution of inference tasks on edge devices. However, the deployment of computation-intensive neural networks on resource-constrained edge systems presents a significant challenge. This brief presents the RV-WINO processor, the first silicon implementation of a RISC-V processor based on the Winograd algorithm for convolution and general matrix multiplication (GEMM) acceleration. The processor incorporates a Winograd module, which significantly reduces multiplication operations during convolutions, leading to a substantial decrease in energy consumption. In addition, the processor includes a matrix multiplication module that reuses the multipliers of the Winograd module, accelerating fully connected and dot product operations in neural networks. The RV-WINO processor fabricated in a 55-nm CMOS process achieves the peak computational performance of 0.95 and 2.39 GOPS in INT32 and INT8 modes, with its peak energy efficiency reaching 112 and 237 GOPS/W. In convolutional neural network (CNN) inference tasks, the execution time is reduced by over 80% compared with the baseline processor. Yucong Huang, Qu Lu, Xinyu Kang, Yuru Li, Qi Wang 0051, Terry Tao Ye |
IEEE Trans. Very Large Scale Integr. Syst. | 8 |
| 2025 | Logic Gate Network Inference Acceleration with RISC-V Custom Instruction SetabstractLogic Gate Networks (LGNs) exploit the similarity between neural networks and logic circuit networks and replace the neurons with logic gates.Consequently, the computation inside the neurons can be replaced by Boolean operations (16 operations for two-input logic).LGNs can be implemented by logic-based instructions in processors and significantly reduce the computation overhead during inference.However, the encoding and decoding processes at the input and output stages of LGNs face efficiency challenges when using traditional RISC-V instruction sets.This limitation arises because these processes rely on one-bit operations, which cannot fully utilize the 32-bit bandwidth of standard instructions.In this work, we proposed four custom RISC-V-based instructions to accelerate the encoding and decoding processes of LGNs.An applicationspecific RISC-V processor, called RV-LGN, has been implemented on FPGA and synthesized using Synopsys® Design Compiler with the CMOS 55nm process.The custom instructions can be called via in-line assembly in C code, making RV-LGN highly promising for implementation in edge devices.Benchmark tests on MIT-BIH, MNIST, and CIFAR-10 classification tasks demonstrate that RV-LGN achieves a runtime reduction of over 87% compared to a generic RISC-V RV32IM ISA processor.Additionally, power consumption during LGN inference is significantly reduced.For the MIT-BIH dataset, the energy consumption is 0.098 µJ/Beat, while MNIST and CIFAR-10 tasks require 0.18 µJ/Image and 0.51 µJ/Image, respectively.These results highlight the superior efficiency of RV-LGN compared to other processors. Chenxi Feng, Xinyu Kang, Yuru Li, Yucong Huang, Terry Tao Ye |
CF | 6 |
| 2025 | Lightweight MCU Implementation (Under 512KB Flash and 256KB RAM) of Neural Networks for Bird Call ClassificationabstractNeural network implementation on resourceconstrained MCUs for edge devices had always been a challenging task, where both the computation logics as well as parameter storage have to be counted and utilized efficiently. In this paper, we propose techniques including lightweight architectural modifications, mixed-precision quantization, and knowledge distillation to implement a complete neural network under tightly budgeted hardware resources. An optimized SqueezeNet network, called BirdCallNet is constructed for bird call classification and used as a case study to demonstrate these techniques. The bird call Mel spectrograms are resized to$112 \times 112$single-channel inputs, and the model was quantized (with symmetric quantization for weights and asymmetric quantization for activations) through ONNX Runtime from FP32 to INT8. The architecture is refined by replacing standard convolutions with time-frequency separable convolutions and integrating efficient channel attention (ECA) blocks, further improving performance under tight constraints. The BirdCallNet can be successfully executed on a 32-bit microcontroller (STM32H743IIT6) with only$\mathbf{5 1 2 K B}$(around$\mathbf{4 3 8 K B}$) flash and 256 KB (around 228 KB) of RAM, with the performance comparable to other implementations with much more hardware overhead. A Python-based GUI is built that allows audio files to be transformed into Mel spectrograms for real-time processing. The techniques proposed in this paper could be used as guidelines for neural network implementation on edge devices. Terry Tao Ye |
HPCC | 2 |
| 2025 | A Low-Noise, High-Input-Impedance Pre-Amplifier for Piezoelectric MEMS MicrophoneabstractThis paper proposes a low-noise, high-input-impedance pre-amplifier for piezoelectric MEMS microphones. It connects directly to the piezoelectric audio sensor without the need for auxiliary biases and operates under a low supply voltage, with a low power consumption, and a small silicon footprint. To recognize a wide range of sound pressure levels with low THD, the pre-amplifier can switch between two modes, i.e., a 20dB high-gain mode and a 10dB low-gain mode. It also incorporates an impedance-enhanced ESD design for the IO pads and a source follower for impedance conversion and input noise suppression. The design is implemented in 0.18-µm CMOS process, and the performance estimation is based on post-layout simulation. The input referred noise of the proposed pre-amplifier is 8.77µVrms in low-gain mode and 3.65µVrms in high-gain mode. it also features a DC input impedance of 190 GΩ, 74.56 dB PSRR, 10.4/20.9 dB system gain, 0.136%/0.044% THD with 94 dBSPL input. It consumes 86.5 µA of current under a 1.6-3.6 V power supply and occupies an active silicon area of 0.23 mm2. Weiye Song, Yucong Huang, Terry Tao Ye |
ISCAS | 5 |
| 2025 | NNia-8: An 8-Core RISC-V Neural Network Inference Accelerator with Efficient Processing Elements and Memory Utilization
Yucong Huang, Xinyu Kang, Yuru Li, Qi Wang 0051, Terry Tao Ye |
NPC (2) | 6 |
| 2025 | qLIF: Mitigating the memory and computation overhead to implement spiking convolutional neural networks
Silong Li, Xinyu Kang, Chunlin Yu, Terry Tao Ye |
Neural Comput. Appl. | 5 |
| 2025 | Spike-Count Reduction Techniques for Low Power Spiking Neural NetworksabstractSpiking neural network (SNN) has demonstrated its great potential in low-power neuromorphic applications. In SNN, computation activities are associated with the arrival and firing of spikes, its power consumption is directly correlated with the number of spikes propagated in the network. In this paper, we explore two methods to reduce the spike-count in the network, aiming to reduce the power consumption of SNN. We use Poisson distribution function in the input layer and through adjusting the correlation (called the gain in the paper) between the probability of spike generation and input values, the number of spikes in the input layer can be reduced to only 20% of the baseline model with the accuracy degradation of less than 1%. We also exploit the leaky-integrate-and-fire (LIF) mechanism and use the refractory period to reduce the generation of spikes from the neurons in the hidden layers. Through this method, the spike-count is reduced by 20% $$\sim $$ 50% while the performance degradation is still less than 1%. These two spike-count reduction techniques are implemented in Verilog RTL; the power simulation results demonstrate significant power reduction in performing SNN computations. We further discovered that for different network architectures, these two techniques have different trade-offs to achieve optimal spike-count reduction while maintaining satisfactory results. Compared with other spike-count reduction techniques, the proposed scheme is efficient and straightforward for hardware implementation, making it well-suited for edge computing scenarios. Xinyu Kang, Zhitao Yang, Terry Tao Ye |
Neural Process. Lett. | 4 |
| 2025 | RV-SCNN: A RISC-V Processor With Customized Instruction Set for SNN and CNN Inference Acceleration on Edge PlatformsabstractThe rapid advancement of artificial intelligence (AI) applications has driven an increasing demand for conducting inference tasks on edge devices. However, implementing computation-intensive neural networks on resource-constrained edge systems remains a significant challenge. In this article, we propose a novel processor architecture called RV-SCNN to address this challenge. The architecture is based on the RISC-V generic instruction set and incorporates various single instruction multiple data (SIMD) custom instruction extensions to accelerate the computation of spike neural networks (SNNs) and convolutional neural networks (CNNs), enabling efficient execution of complex neural network models. The core operators of the processor are shared by both SNN and CNN operations, thus supporting both computation modes. Other acceleration implementations include an internal hardware loop control unit that reduces the instruction overhead, an address calculation unit and an interlayer fusion unit that minimize the memory access overhead, as well as an image to column (IM2COL) unit that improves the computational efficiency of the$3 \times 3$convolutions in SNNs and CNNs. The custom instructions are called through inline assembly in the C program, providing higher flexibility compared to traditional ASICs and supporting custom complex SNN/CNN network structures. Compared to traditional instruction sets, the RV-SCNN processor reduces the execution time of CNNs and SNNs by over 90%. We validate the processor on FPGA platform and evaluate its performance under CMOS 55-nm process. The processor achieves an operational efficiency of 9.88 pJ/SOP in SNN network inference tasks, while the peak energy efficiency reaches 679 GOPS/W in CNN network inference. Chenxi Feng, Xinyu Kang, Qi Wang 0051, Yucong Huang, Terry Tao Ye |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2024 | RV-GEMM: Neural Network Inference Acceleration with Near-Memory GEMM Instructions on RISC-VabstractGeneral Matrix Multiply (GEMM), as a fundamental operation in neural network, plays an important role in artificial intelligence and signal processing applications. In this paper, we proposed three SMID RISC-V custom instructions to accelerate GEMM computations, supporting multiple precisions including 32-bit, 16-bit and 8-bit fixed. Furthermore, we implemented address calculation and loop control units along with the GEMM acceleration module to reduce the memory access overhead. These three GEMM custom instructions, along with the near-memory optimization units, were incorporated in the RV-GEMM processor and implemented on the FPGA platform for speedup evaluation. It was also compiled in Synopsys Design Compiler with CMOS 55nm process for hardware overhead estimation. Compared to the baseline RISC-V processor, for GEMM computations under precisions of 32-bit, 16-bit and 8-bit fixed, the RV-GEMM processor achieved speedup ratios of 15.8×, 28.7× and 42.5×. The peak energy efficiency also reached 260 GOPS/W, 420 GOPS/W and 609 GOPS/W, respectively. Chenxi Feng, Bingzhen Chen, Qi Wang 0051, Yucong Huang, Terry Tao Ye |
CF | 6 |
| 2024 | RWriC: A Dynamic Writing Scheme for Variation Compensation for RRAM-based In-Memory ComputingabstractRRAM-based compute-in-memory (CIM) suffers from programming variation issues, specifically device-to-device variation (DDV) and cycle-to-cycle variation (CCV), which can have a detrimental impact on inference accuracy. To address these variation issues, we propose RWriC, a dynamic Writing scheme for variation Compensation for RRAM-based CIM. RWriC sequentially programs the weights, implemented by multiple RRAM cells, starting from the high significance cell (HSC) and moving towards the low significance cell (LSC). This approach leverages the knowledge of current cumulative errors and the programming targets (PTs) of other RRAM cells to dynamically adjust the PT of the RRAM currently under programming. By shifting the PT of HSC, RWriC enables the LSC to compensate for the programming errors of the HSC. Moreover, when the variation is substantial, RWriC allows the magnitude of LSC to be scaled up, providing an even wider compensation range. Through the combined application of the shifting and scaling techniques, experimental results show that the inference accuracy for ResNet50 on the CIFAR-10 dataset only drops by 0.9% under 18% device variation. In comparison to the conventional writing scheme, our RWriC approach achieves a 5-11x improvement in variation robustness for ResNet50 and Yolov8 across different tasks. Yucong Huang, Jingyu He, Kwang-Ting Cheng, Chi-Ying Tsui, Terry Tao Ye |
DAC | 5 |
| 2024 | Exploring RFID Technology for Wireless Control of Smart AntennasabstractIn recent years, Radio Frequency Identification (RFID) technology has been applied in various fields, including warehouse management, smart transportation, passive sensors, and position tracking. This paper presents a novel application of RFID technology in the field of smart antennas, specifically for the wireless control of the pattern and working frequency of the Yagi-Uda antenna through the Ultra High Frequency (UHF)-RFID protocol. By leveraging RFID technology, our innovative approach controls the passive resonator’s length, enabling it to function as a director or reflector. This solution effectively addresses the challenges of pattern instability caused by metal control cables, which are commonly associated with conventional reconfigurable antennas. The proposed antenna operates at 2.49 GHz in directional mode and at 2.45 GHz and 2.51 GHz in bidirectional mode. In the directional mode, the measured gain is approximately 7.5 dBi, exhibiting a remarkable front-back ratio of 11 dB. Furthermore, the entire system demonstrates excellent energy efficiency, with a mere power consumption of 12 W, and the wireless control range reaches up to 33 m. Chuankui Shen, Zhengxing Wang, Junwei Wu 0002, Qiang Cheng 0002, Terry Tao Ye |
IEEE Internet Things J. | 5 |
| 2024 | Optimizing CNN Computation Using RISC-V Custom Instruction Sets for Edge PlatformsabstractBenefit from the custom instruction extension capabilities, RISC-V architecture can be optimized for many domain-specific applications. In this paper, we propose seven RISC-V SIMD (single instruction multiple data) custom instructions that can significantly optimize the convolution, activation and pool operations in CNN inference computation. More specifically, instruction CONV23 can greatly speed up the operation ofF(2 × 2, 3 × 3). With the adoption of Winograd algorithm, the number of multiplications can be reduced from 36 to 16, and the execution time is also reduced from 140 to 21 clock cycles. These custom instructions can be executed in batch mode within the acceleration module where the immediate data can be reused, so the latency and energy overhead associated with excess memory accesses can be eliminated. Using inline assembler in C language, the custom instructions can be called and compiled together with C source code. A revised RISC-V processor, RI5CY-Accel is constructed on FPGA to accommodate these custom instructions. Revised LeNet-5, VGG16 and ResNet18 model; called LeNet-Accel, VGG16-Accel and ResNet18-Accel are also optimized based on RI5CY-Accel architecture. Benchmark experiments demonstrated that the inference of LeNet-Accel, VGG16-Accel and ResNet18-Accel based on RI5CY-Accel can greatly reduce the execution latency by over 76.6%, 88.8% and 87.1%, with the total energy consumption saving of 74.8%, 87.8% and 85.1% respectively. Bingzhen Chen, Chenxi Feng, Qi Wang 0051, Terry Tao Ye |
IEEE Trans. Computers | 7 |
| 2023 | RVComp: Analog Variation Compensation for RRAM-Based in-Memory ComputingabstractResistive Random Access Memory (RRAM) has shown great potential in accelerating memory-intensive computation in neural network applications. However, RRAM-based computing suffers from significant accuracy degradation due to the inevitable device variations. In this paper, we propose RVComp, a fine-grained analog Compensation approach to mitigate the accuracy loss of in-memory computing incurred by the Variations of the RRAM devices. Specifically, weights in the RRAM crossbar are accompanied by dedicated compensation RRAM cells to offset their programming errors with a scaling factor. A programming target shifting mechanism is further designed with the objectives of reducing the hardware overhead and minimizing the compensation errors under large device variations. Based on these two key concepts, we propose double and dynamic compensation schemes and the corresponding support architecture. Since the RRAM cells only account for a small fraction of the overall area of the computing macro due to the dominance of the peripheral circuitry, the overall area overhead of RVComp is low and manageable. Simulation results show RVComp achieves a negligible 1.80% inference accuracy drop for ResNet18 on the CIFAR-10 dataset under 30% device variation with only 7.12% area and 5.02% power overhead and no extra latency. Jingyu He, Yucong Huang, Miguel Angel Lastras-Montaño, Terry Tao Ye, Chi-Ying Tsui, Kwang-Ting Cheng |
ASP-DAC | 4 |
| 2022 | Multiplication Through a Single Look-Up-Table (LUT) in CNN Inference ComputationabstractParameter quantization with lower bit width is the common approach to reduce the computation loads in CNN inference. With the parameters being replaced by fixed-width binaries, multiplication operations can be replaced by the look-up-table (LUT), where the multiplier-multiplicand operands serve as the table index, and the precalculated products serve as table elements. Because the histogram profiles of the parameters in different layers/channels differ significantly in CNN, previous LUT-based computation methods have to use different LUTs for each layer/channel, and consequently demand larger memory space along with extra access time and power consumption. In this work, we first normalize the parameters Gaussian profiles of different layers/channels to have similar means and variances, and further quantize the normalized parameters into fixed width through nonlinear quantization. Because of the normalized parameters profile, we can use one single compact LUT ($16\times 16$entries) to replace all multiplication operations in the whole network. Furthermore, the normalization procedure also reduces the errors induced from quantization. Experiments demonstrate that with a compact 256-entry LUT, we can achieve the accuracy comparable to the results from 32-bit floating-point calculation; while significantly reducing the computation loads and memory spaces, along with power consumption and hardware resources. Qi Wang 0051, Terry Tao Ye |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2021 | Customized Instruction on RISC-V for Winograd-Based Convolution AccelerationabstractConvolution operation accounts for the major work-load in convolutional neural networks (CNN). However, standard instruction set for RISC-V processor cannot efficiently perform the matrix convolution between kernel and input matrices. In this paper, we construct a custom instruction under the RISC-V ISA that can perform the F(2×2,3×3) convolution within one single execution. Particularly, optimized by the Winograd algorithm, the operation only needs 16 multiplications instead of 36 multiplications as needed by standard ISA. Benefit from this cycles, as compared to 140 cycles using standard instructions. Thenew instruction, F(2×2,3×3) can be calculated within 19 clock power consumed during convolution operation is also reduced significantly. Jianghan Zhu, Qi Wang 0051, Can He, Terry Tao Ye |
ASAP | 5 |
| 2021 | 55nm CMOS Analog Circuit Implementation of LIF and STDP Functions for Low-Power SNNsabstractSpiking neural networks (SNNs) demonstrate great potentials to achieve low-power computation for AI applications. SNN uses spike trains, instead of binary bit-steams to encode input and output information, therefore, analog implementation of SNN will have more advantages than digital implementation in terms of power consumption and hardware overheads. Leaky Integrate-and-Fire (LIF) and Spike Timing Dependent Plasticity (STDP) models are the two fundamental mechanisms of SNN operation. In this paper, we propose a 55nm analog CMOS implementation of the LIF and STDP functions. Testing results demonstrate that the circuit can closely imitate the behavior of the LIF and STDP mechanisms, while demanding a much lower power consumption (around 1nJ per spike with the pulse width of 0.5ms). The proposed LIF and STDP circuits can be used as building blocks to construct a complete SNN architecture. Zhitao Yang, Zhujiang Han, Yucong Huang, Terry Tao Ye |
ISLPED | 4 |
| 2021 | A Method of Fast M-type Approximation on Probabilistic Shaping for Optical CommunicationabstractProbabilistic shaping is widely used in optical communication to increase the data rate. Because the number of symbols is an integer, when modulating the input signals into the predefined symbol sequence, the modulated probabilistic distribution is a discrete M-type distribution. In this paper, a fast method of M-type approximation based on information divergence is proposed. Compared with traditional algorithm, the proposed fast M-type approximation can greatly reduce the computation complexity while achieving the same results. This paper also proves mathematically that the information divergence increment has boundaries for M-type approximation near the ideal distribution. Based on the bounds of information divergence increment, the M-type approximation distribution boundary related to ideal distribution can be obtained. Jianghan Zhu, Xu Li 0001, Yibo Lv, Terry Tao Ye |
VTC Fall | 4 |
| 2021 | E-Textile Battery-Less Displacement and Strain Sensor for Human Activities TrackingabstractWearable sensors demand flexible and concealable devices that can be seamlessly integrated with apparels and clothes. E-textile sensors and devices, built from specially functionalized fibers and yarns and constructed with embroidery and woven processes, provide an ideal solution for different wearable applications. In this article, we propose a textile-based, embroidered passive strain and displacement sensor functioned as a UHF RFID antenna that can be aesthetically integrated with garment fabrics. The antenna structure consists of a dipole arm and a coupling loop mounted with an RFID chip. The displacement between the loop and dipole arm will alter the impedance of the antenna and consequently modulate the backscattered signals from the RFID chip to the reader. The sensor is passive, i.e., no battery is needed on the sensor and the sensing information can be wirelessly retrieved by the reader from a few meters away. The structure is sensitive to strains and displacement caused by body movement. Our experiments demonstrate different applications for this sensor, i.e., when the sensor is embroidered on different locations on human body, such as breast, knees, and elbows, it can be used as a respiration monitor, gesture indicator, etc. to conveniently track and log various forms of body movements. Mengxia Yu, Bingyi Xia, Silong Wang, Mingxun Chen, Siqi Dai, Tingzhe Wang, Terry Tao Ye |
IEEE Internet Things J. | 9 |
| 2020 | Fabrics-Based Embroidered Passive Displacement Sensors for On-Body Applications
Bingyi Xia, Terry Tao Ye |
EWSN | 4 |
| 2020 | Analog Circuit Implementation of Neurons with Multiply-Accumulate and ReLU FunctionsabstractAlthough Artificial Neural Networks (ANNs) are inspired by biological neural systems, most of ANNs today are implemented with digital circuitry and use binary values in computation. In recent years, analog-based neuromorphic system has gained lots of attention as it provides a natural interface for brain-machine interaction. In this paper, we present analog designs of a complete neuron system, where the Multiply-Accumulate (MAC) and Rectified Linear Unit (ReLU) functions are all implemented in analog circuits. The design uses SMIC 55nm standard LP CMOS process node and operates at low supply voltage (1.2 V). The simulation results in SPECTRE demonstrate that the MAC's linear error is no more than 0.5% and total harmonic distortion (THD) is less than 1.6% when the inputs vary from peak (-10 µA) to peak (10 µA) at 10 MHz, the -3dB bandwidth is 288 MHz, the maximum power consumption is 540 µW and the static power consumption is 493 µW under 100MHz input signal frequency. More specifically, our design is resilient to the fluctuation of power supply, which helps to achieve high precision of computation. Yucong Huang, Zhitao Yang, Jianghan Zhu, Terry Tao Ye |
ACM Great Lakes Symposium on VLSI | 4 |
| 2020 | Analog Circuit Implementation of LIF and STDP Models for Spiking Neural NetworksabstractSpiking Neural Networks (SNN) is one special implementation of Artificial Neural Networks (ANN), where the input signals are encoded in the temporal relationship between consecutive spikes (spike trains) instead of real-numbered values. Nevertheless, SNN is believed to be a closer representation of the biological neural system, because it imitates the current spikes that are transmitted between neurons in real biological systems. Practical and simplified SNN models include the Leaky Integrate-and-Fire (LIF) function of the neurons and the Spiking Timing Dependent Plasticity (STDP) function of the synapses. While most ANN architectures can be implemented with digital logic gates, SNN is more suitable to be implemented in analog circuits. In this paper, we propose revised analog circuit implementations of SNN neurons with LIF and STDP functions. Compared with previous works by other researchers, our proposed analog designs use fewer components and can be cascaded to form a complete neural system. The circuits are designed and simulated with SMIC 55nm CMOS LP process. The simulated results demonstrate that the analog neural system can work under a very small current (less than 10 µA) and voltage supply (1.0 Volts), and consumes less power consumption than digital implementations. Zhitao Yang, Yucong Huang, Jianghan Zhu, Terry Tao Ye |
ACM Great Lakes Symposium on VLSI | 4 |
| 2004 | Packetization and routing analysis of on-chip multiprocessor networks
Terry Tao Ye, Luca Benini, Giovanni De Micheli |
J. Syst. Archit. | 1 |
| 2003 | Physical Planning for On-Chip Multiprocessor Networks and Switch FabricsabstractOn-chip implementation of multiprocessor systems requires the planarization of the interconnect network onto the silicon floorplan. Manual floorplanning approaches will become increasingly more difficult and ineffective as multiprocessor complexity increases. Compared with traditional ASIC architectures, multiprocessors have homogeneous processing elements and regular network topologies. Therefore, traditional ASIC floorplanning methodologies based on macro placement are not effective in this domain. We propose an automated physical planning tool, called REGULAY, which can generate floorplans for different topologies under different design constraints. Compared with traditional floorplanning approaches, REGULAY shows significant advantages in reducing the total interconnect wire-length while preserving the regularity and hierarchy of the network topology. Terry Tao Ye, Giovanni De Micheli |
ASAP | 1 |
| 2003 | Packetized On-Chip Interconnect Communication Analysis for MPSoCabstractInterconnect networks play a critical role in shared memory multi-processor systems-on-chip (MPSoC) designs. MPSoC performance and power consumption are greatly affected by the packet dataflows that are transported on the network. In this paper, by introducing a packetized on-chip communication power model, we discuss the packetization impact on MPSoC performance and power consumption. Particularly, we propose a quantitative analysis method to evaluate the relationship between different design options (cache, memory, packetization scheme, etc.) at the architectural level. From the benchmark experiments, we show that optimal performance and power tradeoff can be achieved by the selection of appropriate packet sizes. Terry Tao Ye, Luca Benini, Giovanni De Micheli |
DATE | 1 |
| 2002 | Analysis of power consumption on switch fabrics in network routersabstractIn this paper, we introduce a framework to estimate the power consumption on switch fabrics in network routers. We propose different modeling methodologies for node switches, internal buffers and interconnect wires inside switch fabric architectures. A simulation platform is also implemented to trace the dynamic power consumption with bit-level accuracy. Using this framework, four switch fabric architectures are analyzed under different traffic throughput and different numbers of ingress/egress ports. This framework and analysis can be applied to the architectural exploration for low power high performance network router designs. Terry Tao Ye, Giovanni De Micheli, Luca Benini |
DAC | 1 |
| 2000 | Data Path Placement with RegularityabstractAs more data processing functions are integrated into systems-on-chip, data path is becoming a critical part of the whole VLSI design. However, traditional physical design methodology can not satisfy the data path performance requirement because it has no knowledge of the data path bit-sliced structure. In this paper, an Abstract Physical Model (APM) is proposed to extract bit-slice regularity information from Data Flow Graph (DFG) and it is used for interconnect and congestion planning. A two step heuristic algorithm is introduced to optimize the linear placement of APM to satisfy both the wire length and routing track budget. Terry Tao Ye, Giovanni De Micheli |
ICCAD | 1 |