EDBT 2026 Demo / reviewers in the wild / expert
Zhiyi Yu
dblp:09/4195
· DBLP profile ↗
67ranked-venue papers
7as first author
38since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 56 · 6 first-author · 30 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 3 since 2021Software engineering, systems software and programming languages · 3 · 3 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Spiking-NeRF: Neural Graphics Acceleration With Spiking Feature Encoding for Edge 3D RenderingabstractNeural Radiance Fields (NeRF) have demonstrated remarkable potential for high-fidelity 3D scene reconstruction and rendering. However, achieving real-time performance remains a major challenge due to two critical bottlenecks: the high memory demand of multi-resolution hash encoding and the considerable computational cost of floating-point interpolation. To address these limitations, we propose Spiking-NeRF, a braininspired algorithm-hardware co-design framework. On the algorithm side, we introduce a spiking feature encoding scheme based on Integrate-and-Fire (IF) neurons, which transforms continuous voxel features into sparse spikes, reducing hash storage overhead by $\mathbf{7 5 \%}$. We further propose a global importance-based pruning strategy that compresses hash tables by $\mathbf{7 1. 3 \%}$ by removing lowaccessed entries. To reduce interpolation complexity, we design a hard-threshold weight discretization method that eliminates floating-point operations in favor of bitwise logic. On the hardware side, we accelerate critical stages of the NeRF pipeline by integrating a spike-skipping mechanism that dynamically bypasses hash entries, reducing memory traffic by 32.46%. We also co-optimize on-chip storage by leveraging access locality patterns across different resolution levels of the hash structure. Experimental results demonstrate that Spiking-NeRF achieves real-time rendering performance while maintaining high visual fidelity. Compared to edge GPUs, our design improves throughput by $111.2 \times$ and reduces power consumption by $41.67 \times$. Against SOTA NeRF accelerators, Spiking-NeRF achieves up to $2.48 \times$ higher throughput and $5.45 \times$ lower energy usage, underscoring the potential of spike-based computing for next-generation lowpower neural graphics systems. Jianzhen Gao, Wei Liu 0118, Hengyi Zhou, Zhiyi Yu, Shanlin Xiao |
ASP-DAC | 5 |
| 2026 | An Algorithm-Hardware Co-Design for Efficient and Robust Spiking Neural Networks via SparsityabstractSNN deployment faces a dilemma: rate codes are power-hungry while temporal codes lack noise resilience. This paper proposes an SNN algorithm-hardware co-design, which uses sparse coding and a zero-skipping accelerator to alleviate the rate-temporal trade-off. The design reduces network spike count by 88% compared to rate coding while enhancing fault tolerance. Benchmarked against state-of-the-art rate-coding and temporal-coding accelerators, the prototype saves 88% and 89% energy, achieves $4.5 \times$ and $26.8 \times$ higher throughput, and uses 82% fewer LUTs, enabling efficient and robust edge inference. Wei Liu 0118, Yinsheng Chen, Jilong Luo, Yusa Wang, Zhiyi Yu, Shanlin Xiao |
ASP-DAC | 5 |
| 2026 | NBCache: An Efficient and Scalable Non-Blocking Cache for Coherent Multi-Chiplet Systems
Zhirong Ye, Tao Lu 0012, Zhaolin Li, Zhiyi Yu, Mingyu Wang 0003 |
ASP-DAC | 6 |
| 2026 | LRM-GPU: Alleviating Synchronization Overhead for Multi-Chiplet GPU Architecture
Baiqing Zhong, Zhirong Ye, Haiqiu Huang, Zhaolin Li, Zhiyi Yu, Mingyu Wang 0003 |
HPCA | 7 |
| 2026 | SynapseHD: A unified training framework for bridging spiking neural networks and hyperdimensional computing
Lingfeng Zhou, Huiyao Wang, Jinghai Wang, Zhiyi Yu, Shanlin Xiao |
Neurocomputing | 5 |
| 2026 | HARD: A Heterogeneous Last-Level Cache Architecture With Readless Hierarchical Tag and Dynamic-LRU PolicyabstractThis paper proposes a Heterogeneous Last Level Cache Architecture with Readless Hierarchical Tag and Dynamic-LRU Policy (HARD), designed to enhance system performance and reliability by leveraging the complementary advantages of Static Random-Access Memory (SRAM) and Spin-Orbit Torque Magnetic Random-Access Memory (SOT-MRAM). In the Data RAM, by integrating SOT-MRAM, HARD increases cache capacity under the same area constraints, effectively alleviating performance bottlenecks caused by memory access latency. To better manage the designed Data RAM, we introduce a Dynamic Least Recently Used (Dynamic-LRU) policy, which dynamically adjusts data placement based on access patterns, storing access-intensive data in the SRAM region while migrating access-sparse data to the SOT-MRAM region. This approach not only ensures the performance of the last level cache in high-speed access scenarios but also optimizes storage resource utilization and significantly reduces write pressure on SOT-MRAM. Furthermore, in the Tag RAM, to address the read disturbance issue of SOT-MRAM, we introduce a readless hierarchical tag structure. This structure also employs a heterogeneous design of SRAM and SOT-MRAM, reducing unnecessary matching operations through hierarchical architecture and readless comparison schemes, thereby lowering energy consumption and improving the reliability of the Tag RAM. Experimental results demonstrate that HARD excels in complex data access scenarios, achieving a 40.6× improvement in Mean Time To Failure (MTTF) for the Tag RAM and a 6.3% reduction in miss rate attributed to the increased cache capacity under the same area budget for the Data RAM, offering an efficient and scalable cache solution for high-performance computer architectures. Nan Li 0069, Weichong Chen, Zhiyi Yu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2026 | MAX-SM: High-Utilization Dynamic SM Partitioning for Heterogeneous Workloads on Multitasking Chiplet-Based GPUs
Mingyu Wang 0003, Tao Lu 0012, Baiqing Zhong, Zhaolin Li, Zhiyi Yu |
IEEE Trans. Parallel Distributed Syst. | 7 |
| 2025 | SCSC: Leveraging Sparsity and Fault-Tolerance for Energy-Efficient Spiking Neural NetworksabstractSpiking neural networks (SNNs) are more energy-efficient for processing sparse spike signals and demonstrate better fault tolerance compared to artificial neural networks (ANNs). In neuromorphic chips, synaptic weight access and neuron computation operations constitute 75%-95% of the chip's energy consumption. Therefore, our primary strategy to achieve highly energy-efficient SNNs is to enhance network sparsity while leveraging SNNs' high fault tolerance to reduce both weight access and neuron computation energy. The coding module is an essential component of SNNs, responsible for encoding non-spiking inputs into spike trains. However, previous coding schemes often exhibit poor sparsity or fault tolerance performance. Thus, we propose a novel coding scheme for SNNs: spiking convolutional sparse coding (SCSC). SCSC utilizes convolutional kernels as dictionaries and achieves sparsity through neural layers. Additionally, dynamic firing thresholds in neural layers balance sparsity with network performance and fault tolerance. The experimental results indicate that SCSC can increase network sparsity by 10%-20% and achieves higher accuracy than baseline networks when dealing with disturbances. Furthermore, we utilize approximate DRAM to store synaptic weights and selectively deactivate specific neuronal computing modules. With only a 1% decrease in accuracy, SCSC can reduce synaptic weight access energy by 29% and neuronal computing energy by 49%. Wei Liu 0118, Jinghai Wang, Zhiyi Yu, Shanlin Xiao |
ASP-DAC | 6 |
| 2025 | Towards In-Situ Neuromorphic Computing Architecture for Event Stream Super-ResolutionabstractEvent-based cameras, with their unique event stream representation, effectively mitigate motion blur in highspeed, high-exposure scenarios but suffer from low spatial resolution. To address this, we propose a super-resolution hardware accelerator for event streams based on Spiking Neural Networks (SNNs). In terms of network architecture, we incorporate hardware-friendly algorithmic designs by simplifying neuron models and optimizing convolution operations. On the hardware side, the design adopts a hierarchical structure featuring a highly parallel computational array. Additionally, by proposing a Kernel-Channel-Timestamp-Row (KCTR) dataflow and dual-pipeline structure, the design achieves in-situ computing, eliminating intermediate storage within layers and significantly reducing inter-layer spike storage. Evaluations on the N-MNIST and ASL-DVS datasets demonstrate root mean square errors (RMSE) of $\mathbf{1. 2 9 6}$ and $\mathbf{0. 1 2 1}$ for reconstructed super-resolution event streams. In downstream applications, the classification accuracies reach 98.84% and 99.73%, respectively. The proposed accelerator, designed using a 28 nm CMOS process, improves reconstruction speed by 95.6% compared to a GPU, operates at 500 MHz, and consumes only 0.546 pJ per synaptic operation. Yihe Yu, Wei Liu 0118, Jinghai Wang, Zhiyi Yu, Shanlin Xiao |
DAC | 7 |
| 2025 | FDAIMC: A Fully-Differential Analog In-Memory-Computing for MAC in MRAM with Accuracy Calibration Under Process and Voltage VariationabstractAnalog in-memory-computing (AIMC) is adopted extensively in non-volatile memory for multibit multiply-and-accumulate (MAC) operation. However, the low-on/off-ratio feature of magnetic tunnel junction (MTJ) impedes a high-performance AIMC macro based on spin transfer torque magnetic random access memory (STT-MRAM). Secondly, because of the uncertainty feature of a mixed-signal system under process and voltage variation, a calibration support is indispensable. Moreover, the incompatibility between a nonlinear analog signal and a linear digital signal hinders accurate computation and calibration support. To overcome these challenges, this work proposes a STT-MRAM-AIMC macro featuring: 1) a 2-level-differential cell array and a linear computing scheme with a calibration support in analog domain; 2) an analog-digital-conversion (ADC) system, including a slew-rate-independent voltage-to-time converter (SRIVTC) scheme and a self-triggered time-to-MAC value converter (STTMC) scheme; 3) a compact layout design for high area efficiency. Finally, an average accuracy of 95.44% is obtained under the TT&0.9V corner. By using the calibration strategy, the average accuracy of 97.8% and 88.6% are obtained under FF&0.945V and SS&0.855V separately, with over 30% enhancement. Furthermore, a 1.64~21.18 times area FoM than state of the art is obtained. An energy efficiency of 87.2~312.4 TOPS/W is obtained. Weichong Chen, Ruida Hong, Jinghai Wang, Ningyuan Yin, Zhiyi Yu |
DATE | 6 |
| 2025 | A Low-Complexity True Random Number Generation Scheme Using 3D-NAND Flash MemoryabstractUnpredictable true random numbers are essential in cryptographic applications and secure communications. However, implementing True Random Number Generators (TRNGs) typically requires specialized hardware devices. In this paper, we propose a low-complexity true random number extraction scheme that can be implemented in endpoint systems containing 3D-NAND flash memory chips, addressing the need for random numbers without requiring additional complex hardware. We successfully utilized the randomness of the rapid charging and discharging of shallow charge traps in 3D-NAND memory as an entropy source. The proposed approach only requires conventional user-mode erase, program, and read operations, without any special timing control. We successfully extracted random bitstream using this scheme without a post-debiasing process. We evaluated the randomness of the generated bitstream using the NIST SP 800–22 statistical test suite, and it passed all 15 tests. Ruibin Zhou, Xianping Liu, Yungen Peng, Zhiyi Yu |
DATE | 7 |
| 2025 | Faster-SNN: Towards Faster and Better Spiking Neural Networks with Hybrid Neural CodingabstractInspired by the heterogeneity of the brain, hybrid neural coding in SNN models has garnered increasing attention from researchers. However, most prior research relies on the ANN2SNN conversion method, which results in large time steps and decreased energy efficiency. To overcome these limitations, we propose a faster SNN model (Faster-SNN) based on hybrid neural coding and direct training. Faster-SNN assigns different coding schemes to the input layer, hidden layer, and output layer to achieve hybrid coding. The input layer uses temporal & spatial attention coding (TSAC), which incorporates a spatio-temporal attention mechanism to enhance spatio-temporal information processing. In addition, the hidden layer employs an optimized Burst-LIF neuron to implement burst coding, effectively leveraging the residual information in the membrane potential to improve information transfer efficiency. Finally, the output layer uses TTFS coding to ensure accurate and rapid decision-making. Experimental results demonstrate that our model achieves high-accuracy inference with extremely low latency through the use of hybrid neural coding and direct training methods. Yinsheng Chen, Jilong Luo, Zhiyi Yu, Shanlin Xiao |
ICME | 3 |
| 2025 | Stair-LIF: Boosting the Representation of Spiking Neural Networks with Learnable Incremental Multi-Threshold NeuronsabstractSpiking neural networks (SNNs) have shown remarkable potential in processing spatio-temporal data by mimicking biological neuronal mechanisms and achieving low computational costs. However, previous SNNs often rely on neuron models with fixed and single threshold voltages and binary spikes across layers during training, which limits their capacity for accurate information representation and reduces their biological plausibility. Inspired by the diversity of neuronal behaviors in different brain regions, we propose a novel neuron called Stair-LIF, which introduces learnable incremental multi-threshold mechanisms to enhance neuronal representational capacity and utilizes multi-spike firing to improve the precision of information transmission. Furthermore, we propose a channel-wise parameterization method to expand representational capacity among Stair-LIF. Experimental results on static datasets (CIFAR-10, CIFAR-100) and neuromorphic dynamic datasets (CIFAR10-DVS and DVS128 Gesture) demonstrate that the Starir-LIF neuron achieves state-of-the-art performance. Jilong Luo, Yinsheng Chen, Jinghai Wang, Zhiyi Yu, Shanlin Xiao |
ICME | 5 |
| 2025 | C3ache: Towards Hierarchical Cache-Centric Computing for Sparse Matrix Multiplication on GPGPUs
Mingyu Wang 0003, Baiqing Zhong, Haiqiu Huang, Guangjie Cao, Zhiyi Yu |
MICRO | 6 |
| 2025 | CINOC: Computing in Network-On-Chip With Tiled Many-Core Architectures for Large-Scale General Matrix MultiplicationsabstractLarge-scale general matrix multiplications (LMMs) are the key bottlenecks in various computation domains such as Transformer applications. However, it is a challenge to perform LMMs efficiently on traditional multi/many-core processor systems due to the large amount of memory access and the tight dependence of data transmission. By analyzing the aforementioned problems, we propose a computing in network-on-chip paradigm to perform LMMs by mitigating the performance losses caused by limited on-chip cache resources and memory bandwidth. Specifically, we propose a co-design of computable network-on-chip and the last-level cache method in tiled many-core architectures, which can reconstruct the redundant cache capacity as computable input buffer to balance the demands of computing, storage, and communication for the running LMM applications. Furthermore, a data-aware thread execution mechanism is also proposed to maximize the computational efficiency of thread streams in computable network. At the software level, memory-friendly matrix partitioning strategy, hybrid routing method and programming model are designed to bridge the gap between application demands and mismatched hardware/software interfaces. Experimental evaluations demonstrate that this proposed work achieves a computational latency reduction of 45% compared to the state-of-the-art GPU architecture, and the inference performance is improved by$2\times $of the GPT network. Mingyu Wang 0003, Jiahua Yan, Tao Lu 0012, Zhiyi Yu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2025 | A Low-Cost Secure Branch Predictor to Mitigate the Speculative Attacks by Disrupting Setup PhaseabstractMany types of speculative attacks that exploit branch prediction bear a malicious training process on the branch predictor (BP) in the setup phase. Currently, defense mechanisms on the BP have been rarely studied, and the existing works are restricted to typical scenarios. In this article, we propose a low-cost secure BP to mitigate speculative attacks by monitoring suspicious branch prediction behaviors in all circumstances. We propose a secure mechanism to evaluate the risk level for every branch in the pattern history table. Additionally, we utilize a random number generator to randomly invert the prediction result so as to disrupt the training process, according to the risk level and the generated random number. The maximum inversion probability can be real-time configured during operation. Typically, we implement it on the BI-MODE BP in XuanTie-C910 RISC-V core with a minimal system on chip on field-programmable gate array (FPGA). The realistic hardware evaluation under SPEC2017 with Linux shows that, under the optimal tradeoff between security and performance with the maximum inversion probability of 20%, the hardware and performance overhead are less than 1%, and the speculative attacks including Spectre v1.x and Meltdown can be mitigated with the rates of 44%–88%. Runye Ding, Yao Liu 0006, Zhiyi Yu |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2025 | A Hybrid Stochastic-Binary Computing Batch Normalization Engine for Low-Power On-Chip Learning Spiking Neural NetworksabstractBatch normalization (BN) has proven to be a critical component in speeding up the training of deep spiking neural networks in deep learning. However, conventional BN implementations face significant challenges in terms of excessive off-chip memory bandwidth requirements and complex circuit designs, hindering their applicability for on-chip training in spiking neural networks (SNNs). This article introduces a novel hybrid stochastic-binary computing BN engine (HBN) that strikes an optimal balance between computational efficiency and hardware resource utilization, enabling efficient on-chip learning for SNNs. While conventional binary-mode BN engines offer temporal efficiency, they demand substantial hardware resources. In contrast, stochastic computing (SC)-based BN approaches reduce hardware overhead but introduce latency penalties and necessitate additional random number generation (RNG) circuits. To overcome these limitations, we propose a hybrid architecture that seamlessly integrates binary and stochastic computing (SC) paradigms. Our co-designed methodology effectively balances computational latency and hardware footprint. This is achieved by a rounding-free SC multiplier unified with binary-circuit map ping, which eliminates latency and RNG overheads. Extensive validation across both static image datasets and neuromorphic datasets demonstrates that HBN maintains algorithmic fidelity while achieving unprecedented computational efficiency. Simulation results reveal 98.7% reduction in floating-point operations (FLOPs), 98.5% latency improvement, and 98.2% energy consumption reduction compared with conventional BN implementations. FPGA implementation on the ZCU102 platform demonstrates practical hardware advantages, including 74.9% reduction in look-up table (LUT) utilization, 83.6% decrease in flip-flop (FF) count, and 13.7% reduction in block RAM (BRAM) allocation. Notably, the design achieves 63.7% power reduction compared with state-of-the-art implementations while maintaining complete DSP-free operation. Wei Liu 0118, Zhiyi Yu, Shanlin Xiao |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2025 | ASNA-Flow: An Efficient Asynchronous Neuromorphic Accelerator for Real-Time Event-Based Optical FlowabstractOptical flow estimation constitutes a fundamental computational challenge in computer vision, with critical applications object trajectory prediction, depth reconstruction, and autonomous navigation systems. The emergence of neuromorphic vision systems, integrating event-driven cameras with spiking neural networks (SNNs), has recently gained attention as a promising paradigm for edge deployment of optical flow estimation due to their advantages in ultralow power and resource efficiency. However, current neuromorphic computing platforms lack specialized architectures optimized for this problem domain. Existing implementations either prioritize configurable architectures at the expense of energy efficiency or employ intricate hardware control mechanisms to manage the asynchronous and sparse computing patterns inherent in SNNs. To address these limitations, we present ASNA-Flow, an event-driven asynchronous neuromorphic accelerator featuring a pioneering algorithm–hardware co-design framework specifically tailored for event-based optical flow estimation. Our methodology encompasses three key innovations: 1) a hardware-aware algorithm optimization that maintains computational fidelity while enhancing implementation efficiency; 2) systematic data pattern analysis to inform architectural decisions; and 3) novel exploitation of optical flow’s spatial locality characteristics to enable efficient sparse computing. Implemented in TSMC 28-nm CMOS technology, ASNA-Flow achieves real-time performance of 104 frames per second (FPS) with ultralow power consumption of 7.9 mW, demonstrating superior energy efficiency of 0.3 pJ per synaptic operation (SOP). This work establishes the first dedicated neuromorphic computing solution that simultaneously addresses the temporal sparsity, event-driven processing, and energy constraints inherent in optical flow estimation tasks. Jinghai Wang, Jilong Luo, Lingfeng Zhou, Zhiyi Yu, Shanlin Xiao |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2025 | An End-to-End Bundled-Data Asynchronous Circuits Design Flow: From RTL to GDSabstractAsynchronous circuits with low power and robustness are revived in emerging applications such as the Internet of Things (IoT) and neuromorphic chips, thanks to clock-less and event-driven mechanisms. However, the lack of mature computer-aided design (CAD) tools for designing large-scale asynchronous circuits results in low design efficiency and high cost. This article proposes an end-to-end bundled-data (BD) asynchronous circuit design flow, which can facilitate building asynchronous circuits, even if the designer has little or no asynchronous circuit foundation. Three features that enable this are: 1) a lightweight circuit converter developed in Python can convert circuits from synchronous descriptions to corresponding asynchronous ones at register transfer level (RTL). Desynchronization flow helps designers maintain a “synchronization mentality” to construct asynchronous circuits; 2) a synchronization-like verification method is proposed for asynchronous circuits so that it can be functionally verified before synthesis. Avoids the risk of rework after logic defects are discovered during the synthesis and implementation, as asynchronous circuits often cannot be simulated until gate-level (GL) netlist generation; and 3) the whole implementation flow from RTL to graphic data system (GDS) is based on commercial electronic design automation (EDA) tools. Similar to the design flow of synchronous circuits, it helps designers implement asynchronous circuits with “synchronization habits.” Furthermore, to validate this methodology, two asynchronous processors were, respectively, implemented and evaluated in the TSMC 28-nm CMOS process. Compared to their synchronous counterparts, the general-purpose asynchronous RISC-V processor achieves 20.5% power savings. And the domain-specific asynchronous spiking neural network (SNN) accelerator achieves 58.46% power savings and$2.41\times $energy efficiency improvement at 70% input spike sparsity. Jinghai Wang, Shanlin Xiao, Jilong Luo, Lingfeng Zhou, Zhiyi Yu |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2024 | An Efficient Asynchronous Circuits Design Flow with Backward Delay Propagation ConstraintabstractAsynchronous circuits have recently become more popular in Internet of Things (IoT) and neural network chips because of their potential low power consumption. However, due to the lack of Electronic Design Automation (EDA) tools, the asynchronous circuits design efficiency remains low and faces challenges in large-scale applications. This paper proposes a new asynchronous circuits design flow using traditional EDA tools, and applies a new backward delay propagation constraint (BDPC) method. In this method, control paths and data paths are tightly coupled and analyzed together to improve the accuracy of static timing analysis. Compared to previous works, the proposed design flow and constraint method offer significant advantages in terms of accuracy and efficiency. To verify this flow, an asynchronous RISC-V processor was implemented on TSMC 65nm process. Compared to synchronous version, asynchronous processor achieves a power optimization of 17.4 % while main-taining the same speed and area. Lingfeng Zhou, Shanlin Xiao, Huiyao Wang, Jinghai Wang, Zeyang Xu, Zhiyi Yu |
DATE | 7 |
| 2024 | Sparsespikformer: A Co-Design Framework for Token and Weight Pruning in Spiking TransformerabstractAs the third-generation neural network, the Spiking Neural Network (SNN) has the advantages of low power consumption and high energy efficiency, making it suitable for implementation on edge devices. However, despite these advantages, SNN still faces accuracy limitations when compared to Artificial Neural Network (ANN). More recently, the most advanced SNN, Spikformer, combines the self-attention module from Transformer with SNN to achieve accuracy comparable to that of ANN. Additionally, to improve the final accuracy, the researchers adopt larger channel dimensions in MLP layers, leading to an increased number of redundant model parameters. To effectively decrease the computational complexity and weight parameters of the model, we explore the Lottery Ticket Hypothesis (LTH) and discover a very sparse (>90%) subnetwork that achieves comparable performance to the original network. Furthermore, we also design a lightweight token selector module, which can remove unimportant background information from images based on the average spike firing rate of neurons, selecting only essential foreground image tokens to participate in attention calculation. Experimental results demonstrate that our co-design framework can significantly reduce 90% model parameters and cut down Giga Floating-Point Operations (GFLOPs) by 20% while maintaining the accuracy of the original model. Shanlin Xiao, Zhiyi Yu |
ICASSP | 4 |
| 2024 | CCacheSim: A Circuit-Architecture Cross-Level Simulation Framework for SRAM-Based in-Cache Computing System EvaluationabstractSRAM-based Compute-In-Memory (CIM) circuits have demonstrated significant performance and energy efficiency advantages. Although numerous frameworks or tools have emerged for simulating CIM-based systems, most frameworks are tailored for specific DNN accelerators and rarely consider SRAM-CIM solutions in general processor systems because it is difficult to establish an effective mechanism to build the CIM data path to integrate the SRAM-CIM module that is tightly coupled with the cache hierarchy into the system. To address this problem, we propose a circuit-architecture cross-level simulation framework named CCacheSim for in-cache computing system. CCacheSim integrates the simulation of SRAM-CIM circuit timing and energy consumption characteristics, providing circuit-level accuracy evaluation support for in-cache computing system simulations. For the circuit level, the SRAM-CIM model can automatically generate a circuit-level netlist and conduct accurate simulation through corresponding configurations, thus balancing accuracy and agility for early design exploration. For the architectural level, to efficiently support the portable integration of SRAM-CIM module to in-cache computing system, a configurable hardware programming interface is implemented within the cache model to manage the interaction of the control stream between processor and cache for CIM tasks. Moreover, a request queue based access mechanism is proposed to ensure the completeness of the operands required by CIM tasks. To validate the proposed framework, CCacheSim is implemented to simulate varying configurations of in-cache computing systems and SRAM-CIM modules. The results prove that CCacheSim can conduct accurate performance and energy consumption evaluation for given processor architecture with given CIM module. CCacheSim can support flexible and effective design space exploration for in-cache computing system. Baiqing Zhong, Mingyu Wang 0003, Yicong Zhang, Zhiyi Yu |
ICCD | 5 |
| 2024 | Atomic Cache: Enabling Efficient Fine-Grained Synchronization with Relaxed Memory Consistency on GPGPUs Through In-Cache Atomic OperationsabstractGeneral-purpose graphics processing unit (GPGPU), widely recognized as an exceptional computing platform for de-ploying emerging parallel applications, requires strict adherence to atomicity and memory consistency models for shared variable synchronization. This is crucial to ensure deterministic execution and leverage the performance advantages of the GPGPU single-instruction -multiple-threads architecture. However, the escalating demand for shared variable updates across thread blocks, notably in applications like deep neural networks and graph analysis, significantly exacerbates the serialization overhead of atomic operations due to the von Neumann bottleneck. Additionally, the overhead introduced by memory fences supporting the memory consistency model further complicates this fine-grained synchronization requirement. To address these challenges, this paper proposes Atomic Cache, facilitating an In-Cache computing hardware-software co-design for GPGPUs. At the software level, we propose relaxed memory consistency based on non-ordering commutativity to alleviate the execution of in-cache atomic operations, thereby mitigating the performance overhead of memory fences. At the hardware level, we present the In-Situ Store Atomic Cache Macro, which empowers the Atomic Cache to efficiently execute atomic logic and arithmetic operations within the cache array. This innovation alleviates the von Neumann bottleneck associated with serialized execution of atomic operations. The experimental evaluation results demonstrate that the Atomic Cache can save more than 60% of memory access energy while incurring only 9.42% chip area overhead. Furthermore, it not only delivers an average speedup ratio of 2.59 × and an IPC performance improvement of 1.48× for RISC-V GPGPUs, but also achieves an average speedup ratio of 1.31 × and an IPC performance improvement of 39.92% when compared to state-of-the-art designs employing local atomic buffers. Yicong Zhang, Mingyu Wang 0003, Wangguang Wang, Yangzhan Mai, Haiqiu Huang, Zhiyi Yu |
MICRO | 6 |
| 2024 | Enhancing text classification with attention matrices based on BERTabstractSummary Text classification is a critical task in the field of natural language processing. While pre‐trained language models like BERT have made significant strides in improving performance in this area, the distinctive dependency information that is present in text has not been fully exploited. Besides, BERT mostly captures phrase‐level information in lower layers, which becomes progressively weaker with the increasing depth of layers. To address these limitations, our work focuses on enhancing text classification through the incorporation of Attention Matrices, particularly in the fine‐tuning process of pre‐trained models like BERT. Our approach, named AM‐BERT, leverages learned dependency relationships as external knowledge to enhance the pre‐trained model by generating attention matrices. In addition, we introduce a new learning strategy that enables the model to retain learned phrase‐level structure information. Extensive experiments and detailed analysis on multiple benchmark datasets demonstrate the effectiveness of our approach in text classification tasks. Furthermore, we show that AM‐BERT achieves stable performance improvements also in named entity recognition tasks. Zhiyi Yu, Jialin Feng |
Expert Syst. J. Knowl. Eng. | 1 |
| 2024 | Enhancing aspect-based sentiment analysis with dependency-attention GCN and mutual assistance mechanism
Jialin Feng, Zhiyi Yu |
J. Intell. Inf. Syst. | 3 |
| 2024 | Contrastive learning for unsupervised sentence embeddings using negative samples with diminished semantics
Zhiyi Yu, Jialin Feng |
J. Supercomput. | 1 |
| 2024 | Better-Than-Worst-Case: A Frequency Adaptation Asynchronous RISC-V Core With Vector ExtensionabstractIn recent years, asynchronous circuits have become more popular in neural network chips and the Internet of Things (IoT) due to their potential advantages of low-power consumption and high performance. However, the existing design methods for asynchronous circuits are still constrained by critical paths, increasing power consumption and hindering the further improvement of performance. In this article, a fully digital design method for frequency adaptation asynchronous bundled-data (BD) circuits is proposed. The proposed method is straightforward, effective, widely applicable, and independent of asynchronous controllers. It allows to automatically work on different frequencies as required, which can improve performance and reduce power consumption, achieving better-than-worst-case. To verify the proposed method, an asynchronous RISC-V processor with vector acceleration extension is designed on both TSMC 65-nm process and field-programmable gate array (FPGA) platform. According to the postlayout simulation results, compared with its synchronous version, the asynchronous processor achieves a 10% speed improvement (from 227.3 to 250 MHz) with a 37% power reduction (from 135 to 85$\mu$W/MHz) under ideal conditions. Even under the worst conditions, the asynchronous processor achieves equivalent performance to the synchronous processor, while still reducing power consumption by 29% (from 133 to 95$\mu$W/MHz). On the FPGA platform, asynchronous processor also achieves higher speed while lower power consumption. Lingfeng Zhou, Shanlin Xiao, Huiyao Wang, Jinghai Wang, Zeyang Xu, Zhiyi Yu |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2024 | Toward Efficient Asynchronous Circuits Design Flow Using Backward Delay Propagation ConstraintabstractIn recent years, asynchronous circuits have gained attention in neural network chips and Internet of Things (IoT) due to their potential advantages of low power and high performance. However, design efficiency of asynchronous circuits remains low and faces challenges in large-scale applications because of the lack of electronic design automation (EDA) support. This article presents a new bundled-data (BD) asynchronous circuits’ design flow using traditional EDA tools, including a new backward delay propagation constraint (BDPC) method. In this method, control paths and data paths are analyzed together in a tightly coupled approach to improve the accuracy of static timing analysis (STA). Compared with other design flows, the proposed design flow and constraint method show significant advantages in aspects of STA accuracy, design efficiency, and design applicability, and solving the congestion issues of field-programmable gate array (FPGA) in a previous work. An asynchronous RISC-V processor was implemented to verify the method, with selective handshake technology to further reduce power. Compared with the synchronous processor, the asynchronous processor achieves a 17.4% power optimization on the TSMC 65-nm process and a 48.3% dynamic power savings on the FPGA while maintaining the same frequency and resource utilization. Lingfeng Zhou, Shanlin Xiao, Huiyao Wang, Jinghai Wang, Zeyang Xu, Zhiyi Yu |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2023 | CRAFT: Common Router Architecture for Throughput Optimization
Jiahua Yan, Mingyu Wang 0003, Zhiyi Yu |
ICA3PP (3) | 4 |
| 2023 | LWSDP: Locality-Aware Warp Scheduling and Dynamic Data Prefetching Co-design in the Per-SM Private Cache of GPGPUsabstractGeneral Purpose Graphics Processing Units (GPG-PUs) employ frequent context switching to mask the long-latency of memory operations. However, GPGPUs still suffer from stagnation due to the incomplete overlapping of memory operations. To alleviate this stagnation and enhance Memory-Level Parallelism (MLP), it is crucial to overlap and minimize memory operations. This paper conducts a comprehensive analysis of data locality in GPGPUs and proposes an approach called Locality-Aware Warp Scheduling and Dynamic Data Prefetching (LWSDP) Co-design in the Per-SM Private Cache of GPGPUs, which effectively utilizes data locality to improve MLP. In addition to employing a coordinated scheduler and dynamic data prefetching, we incorporate Prefetching Requests Admitted Cache Access Re-execution (PRA-CAR) to mitigate the adverse impact of excessive prefetching memory requests on memory saturation. Experimental results demonstrate that LWSDP achieves an average 33.02% performance improvement and an average 28.16% miss rate reduction compared to the previous schedulers on data locality-sensitive kernels. Wangguang Wang, Mingyu Wang 0003, Yicong Zhang, Yukun Wei, Zhiyi Yu |
ICPADS | 5 |
| 2023 | A Scalable Deadlock-Free Static Routing Algorithm for Chiplet-Based SystemsabstractThe utilization of the Chiplet methodology can accelerate VLSI system development and provide better flexibility. Building interconnection networks across multiple Chiplets and ensuring high-performance deadlock-free routing in systems with diverse irregular topologies is a challenging task.To avoid the reordering introduced by adaptive routing algorithms, a scalable static deadlock-free routing algorithm specifically designed for Chiplets is proposed. This approach capitalizes on static routing, a feature that ensures consistent message order due to fixed paths, fundamentally averting reordering issues. By adaptively configuring the state of the local router, it is possible to proactively initiate turns or exit detour loops, thereby effectively preventing deadlocks. Furthermore, by employing state configuration, the routing remains deadlock-free even when there are variations in the number of Chiplets, scale, or internal topology. This showcases the system’s scalability.Due to limited wiring resources, we chose classical routing algorithms up*/down* to compare, and the results showed a significant latency advantages, with a saturation injection rate approximately 1.5 to 2 times higher. Mingyu Wang 0003, Yicong Zhang, Tao Lu 0012, Zhiyi Yu |
ICPADS | 5 |
| 2023 | A 1.97 TFLOPS/W Configurable SRAM-Based Floating-Point Computation-in-Memory Macro for Energy-Efficient AI ChipsabstractFloating-point (FP) computation-in-memory (CIM) technology is increasingly demanded by low-power neural network training. In this work, we propose an energy-efficient configurable SRAM-based FP CIM macro. A mantissa parallel alignment method is proposed to improve calculation speed and accuracy in FP multiply-accumulation (MAC) operations. The separated mantissa CIM and exponent CIM are designed to enable pipelining of exponent and mantissa operations to increase computation throughput. Furthermore, the macro can be flexibly set to BF16 or FP32 precision by configuring accumulators. The proposed FP CIM macro is analyzed in 40 nm CMOS technology, and the estimated area is 0.48 mm2, The simulation results show that the macro achieves a frequency of 294 MHz in 1.1 V. In BF16 mode, the macro can achieve a peak throughput of 56.5 GFLOPS and an energy efficiency of 1.97 TFLOPS/W while the peak throughput and energy efficiency are 16 GFLOPS and 0.62 TFLOPS/W in FP32 mode. Yangzhan Mai, Mingyu Wang 0003, Chuanghao Zhang, Baiqing Zhong, Zhiyi Yu |
ISCAS | 5 |
| 2023 | Computing Resistance-Style Image Sensors for Artificial Neural NetworksabstractToday, machine vision experiences large latency due to big data processing, which is a barrier to time-critical applications. To address this issue, in-sensor computing was presented in the past. Here, we present a scheme of computing in a magnetic tunneling junction (MTJ) sensor array for proof-of-principle. Using the MTJ sensor array, the functions of artificial neural network (ANN) classifiers and autoencoders were verified. The time for correct classification of one picture was less than$9~\mu \text{s}$. The power consumed in the sensor array can be decreased according to the square law without affecting the results. Our work shows universal circuits and algorithms to compute in resistance-style ANN image sensors with promising energy efficiency. Guihua Zhao, Yating Peng, Yizhan Wang, Caihua Wan, Xianping Liu, Yu Zhang 0248, Xiufeng Han, Weichong Chen, Zhiyi Yu |
IEEE Internet Things J. | 9 |
| 2023 | TensorCache: Reconstructing Memory Architecture With SRAM-Based In-Cache Computing for Efficient Tensor Computations in GPGPUsabstractGeneral purpose graphics processing units (GPGPUs) have emerged as a convincing and pivotal computing platform for deep learning applications. However, the fundamental tensor computations for neural networks on GPGPUs are still restricted by the von Neumann bottleneck. The memory bandwidth and energy consumption of moving a large amount of neural network data between the memory hierarchy and computational units of GPGPUs dominate the overall computational cost. To address these challenges, this article proposes TensorCache to reconstruct memory architecture with static random-access memory (SRAM)-based In-Cache Computing for efficient tensor computations in GPGPUs. It provides an innovative digital SRAM processing-in-memory (PIM) solution by transforming the cache array into large-scale PIM units, effectively mitigating the significant performance and energy consumption losses caused by data movement. To enable efficient hardware-software co-design for TensorCache, a decoupled architecture-based SRAM-PIM macro (SPM) is introduced at the hardware level, supporting in-memory bit-parallel comparison (IMBC) and near-memory radix-4 booth encoder (NRBE) for efficient mixed-precision floating-point (FP) tensor computations. At the software level, a programming model leveraging the GPGPU’s flexible programmability is proposed to bridge the gap between application demands and mismatched hardware/software interfaces. Experimental evaluations demonstrate that TensorCache achieves up to$38.59\times $speedup and$16.26\times $throughput enhancement compared to GPU CUDA Cores. Furthermore, it attains an acceleration of up to$1.78\times $and$3.87\times $throughput improvement compared to GPU Tensor Cores, while saving power consumption in tensor computations by over 90% with a mere 21% chip area overhead. Yicong Zhang, Mingyu Wang 0003, Yangzhan Mai, Zhiyi Yu |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2021 | 3D-VNPU: A Flexible Accelerator for 2D/3D CNNs on FPGAabstractThree-dimensional convolutional neural networks (3D CNNs) have proven to be outstanding in applications such as video analysis, 3-dimension geometric data, and 3-dimension medical image diagnosis. Compared to 2D CNNs, 3D CNNs require high computational complexity to get spatio-temporal features while Winograd algorithm can significantly reduce the amount of computation. Prior works based on 3D Winograd accelerators are only applied to stride-1 convolution, however, most of the popular 3D CNNs contain stride-2 convolution layers. In this paper, we propose a novel flexible Winograd-based decomposition method (FWDM) to apply the 3D Winograd to different strides convolution. Evaluation results show that FWDM reduces computational complexity by a factor of 3.2 for C3D, 2.9 for 3D ConvNet, and 2.6 for 3D ResNet-18. Furthermore, we design a flexible computing engine to stretch the use range of the decomposition method. Coupling FWDM and computing engine, a Winograd-based, 2D/3D CNNs compatible, highly efficient, and flexible accelerator (3D-VNPU) is proposed. Finally, we demonstrate the effectiveness of 3D-VNPU on FPGA platform (Xilinx ZCU102) and achieve 1.35TOPS for C3D, 1.2TOPS for 3D ResNet-18, and 1.1TOPS for VGG-16. DSP efficiency outperforms other CNNs accelerators 2.57~15.3x compared with prior works in FPGA of C3D. Compared to GPU and CPU, our accelerator achieves improvement up to 37.9x in performance relative to CPU and 11.8x in energy efficiency relative to GPU. Huipeng Deng, Jian Wang 0080, Huafeng Ye, Shanlin Xiao, Xiangyu Meng 0003, Zhiyi Yu |
FCCM | 6 |
| 2021 | High-parallelism Inception-like Spiking Neural Networks for Unsupervised Feature Learning
Mingyuan Meng, Lei Bi 0001, Jinman Kim, Shanlin Xiao, Zhiyi Yu |
Neurocomputing | 6 |
| 2021 | A Data-Driven Asynchronous Neural Network AcceleratorabstractDeep neural networks (DNNs) are revolutionizing machine learning, with unprecedented accuracy on many AI tasks. Energy-efficient neural acceleration is crucial in broadening DNN applications in cloud and mobile end devices. However, power-hungry clock networks limit the energy-efficiency of DNN accelerators. In this work, we propose a novel DNN hardware accelerator, called the asynchronous neural network processor (AsNNP). At the heart of AsNNP is a scalable hierarchy matrix multiply unit, with bit-serial processing elements working in parallel. It replaces the global clock networks with asynchronous handshake protocols to realize the synchronization and communication between each part, minimizing the dynamic power. Meanwhile, a fine-grain asynchronous pipeline based on weak-conditioned half-buffer (WCHB) is introduced to pipe successive computations in a data-driven manner, i.e., once data arrives computation begins, maximizing the throughput. These techniques enable AsNNP to work in a fully data-driven asynchronous communication fashion with optimized energy-efficiency. The proposed accelerator is implemented with quasi-delay-insensitive (QDI) clockless logic family and evaluated in a 65 nm process. Compared with the synchronous baseline, simulation results show that AsNNP offers 2.2× higher equivalent frequency and 1.59× lower power. Compared with state-of-the-art DNN accelerators, AsNNP shows 1.17×-4.97× energy-efficiency improvement. Shanlin Xiao, Weikun Liu, Junshu Lin, Zhiyi Yu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2021 | An MTJ-Based Asynchronous System With Extremely Fine-Grained Voltage ScalingabstractIn this work, we present an asynchronous MTJ-CMOS hybrid system with extremely fine-grained voltage scaling (EFGVS) technique. The supply voltage of the system is turned on/off by asynchronous bundled-data handshake signals. The MTJ write circuit is variation robust, self-terminated and redundant-write preventing. Besides, EFGVS and asynchronous data driven handshake enhance the timing robustness by removing the matching delay elements and reducing timing assumptions. The completion detection circuitry is also simplified. A RISC processor with proposed techniques is designed and fabricated by 55nm technology. The voltage supplies are turned on/off within tens of picoseconds. The energy for writing MTJs is 0.45pJ. The tolerance of the minimum TMR due to process variation is 75%. Sleep mode leakage power can be reduced over 10× by powering off modules with the Break-even Time of 23.6 ns. Because of the extremely fine-grained voltage scaling, no more than 20% of the modules are powered on during the execution. Ningyuan Yin, Baofa Huang, Xiaobai Chen, Zhiyi Yu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2020 | SPA: Stochastic Probability Adjustment for System Balance of Unsupervised SNNsabstractSpiking neural networks (SNNs) receive widespread attention because of their low-power hardware characteristic and brain-like signal response mechanism, but currently, the performance of SNNs is still behind Artificial Neural Networks (ANNs). We build an information theory-inspired system called Stochastic Probability Adjustment (SPA) system to reduce this gap. The SPA maps the synapses and neurons of SNNs into a probability space where a neuron and all connected pre-synapses are represented by a cluster. The movement of synaptic transmitter between different clusters is modeled as a Brownian-like stochastic process in which the transmitter distribution is adaptive at different firing phases. We experimented with a wide range of existing unsupervised SNN architectures and achieved consistent performance improvements. The improvements in classification accuracy have reached 1.99% and 6.29% on the MNIST and EMNIST datasets respectively. Mingyuan Meng, Shanlin Xiao, Zhiyi Yu |
ICPR | 4 |
| 2020 | Spiking Inception Module for Multi-layer Unsupervised Spiking Neural NetworksabstractSpiking Neural Network (SNN), as a brain-inspired approach, is attracting attention due to its potential to produce ultra-high-energy-efficient hardware. Competitive learning based on Spike-Timing-Dependent Plasticity (STDP) is a popular method to train an unsupervised SNN. However, previous unsupervised SNNs trained through this method are limited to a shallow network with only one learnable layer and cannot achieve satisfactory results when compared with multi-layer SNNs. In this paper, we eased this limitation by: 1) We proposed a Spiking Inception (Sp-Inception) module, inspired by the Inception module in the Artificial Neural Network (ANN) literature. This module is trained through STDP-based competitive learning and outperforms the baseline modules on learning capability, learning efficiency, and robustness. 2)We proposed a Pooling-Reshape-Activate (PRA) layer to make the Sp-Inception module stackable. 3)We stacked multiple Sp-Inception modules to construct multilayer SNNs. Our algorithm outperforms the baseline algorithms on the hand-written digit classification task, and reaches state-of-the-art results on the MNIST dataset among the existing unsupervised SNNs. Mingyuan Meng, Shanlin Xiao, Zhiyi Yu |
IJCNN | 4 |
| 2020 | A Low-Cost and High-Throughput NoC-Aware Chip-to-Chip InterconnectionabstractTo offer sufficient computing capability, current hardware systems tend to be equipped with multiple chips, while each chip integrates tens to hundreds of cores with Network-on-Chip (NoC) architecture. Thus, chip-to-chip interconnection for NoCs becomes indispensable. However, state-of-the-art inter-chip interconnection, like PCIe or SRIO, suffers from latency and bandwidth bottlenecks especially for communication-intensive tasks. In this paper, we propose a NoC-aware chip-to-chip interconnection scheme. In addition to a lightweight interconnection architecture and protocol, we utilize NoC routers to improve inter-chip connection. A virtual-channel router with transmission priority in the NoC is proposed to eliminate the congestion in the chip-to-chip interconnection. The interconnection system is implemented in Verilog RTL and verified on the Xilinx ZCU102 evaluation kit, achieving up to 10Gb/s/lane line rate with low resource utilization. Wenkang Liao, Yuhao Guo, Shanlin Xiao, Zhiyi Yu |
ISCAS | 4 |
| 2020 | NeuronLink: An Efficient Chip-to-Chip Interconnect for Large-Scale Neural Network AcceleratorsabstractLarge-scale neural network (NN) accelerators typically consist of several processing nodes, which could be implemented as a multi- or many-core chip and organized via a network-on-chip (NoC) to handle the heavy neuron-to-neuron traffic. Multiple NoC-based NN chips are connected through chip-to-chip interconnection networks to further boost the overall neural acceleration capability. Huge amounts of multicast-based traffic travel on-chip or cross chips, making the interconnection network design more challenging and become the bottleneck of the NN system performance and energy. In this article, we propose coupling intrachip and interchip communication techniques, called NeuronLink, for NN accelerators. Regarding the intrachip communication, we propose scoring crossbar arbitration, arbitration interception, and route computation parallelization techniques for virtual-channel routing, leading to a high-throughput NoC with a lower hardware cost for multicast-based traffic. Regarding the interchip communication, we propose a lightweight and NoC-aware chip-to-chip interconnection scheme, enabling efficient interconnection for NoC-based NN chips. In addition, we evaluate the proposed techniques on a four connected NoC-based deep neural network (DNN) chips with four field-programmable gate arrays (FPGAs). The experimental results show that the proposed interconnection network can efficiently manage the data traffic inside DNNs with high-throughput and low-overhead against state-of-the-art interconnects. Shanlin Xiao, Yuhao Guo, Wenkang Liao, Huipeng Deng, Huanliang Zheng, Jian Wang 0080, Gezi Li, Zhiyi Yu |
IEEE Trans. Very Large Scale Integr. Syst. | 10 |
| 2019 | A 68-mw 2.2 Tops/w Low Bit Width and Multiplierless DCNN Object Detection Processor for Visually Impaired PeopleabstractDeep convolutional neural network (DCNN) object detection is a powerful solution in visual perception, but it requires huge computation and communication costs. We proposed a fast and low-power always-on object detection processor that allows visually impaired people to understand their surroundings. We designed an automatic DCNN quantization algorithm that successfully quantizes the data to 8-bit fix-points with 32 values and uses 5-bit indexes to represent them, reducing hardware cost by over 68% compared to the 16-bit DCNN, with negligible accuracy loss. A specific hardware accelerator is designed, which uses reconfigurable process engines to realize multi-layer pipelines to significantly reduce or eliminate the off-chip temporary data transfer. A lookup table is used to implement all multiplications in convolutions to reduce the power significantly. The design is fabricated in SMIC 55-nm technology, and the post-layout simulation shows only 68-mw power at 1.1-v voltage with 155 Go/s performance, achieving 2.2 Top/w energy efficiency. Xiaobai Chen, Jinglong Xu, Zhiyi Yu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2018 | An Automatic Task Partition Method for Multi-core SystemabstractIn this paper, an automated task partition method for multi-core system is proposed. To explore the full parallelism of an application written in sequential languages such as C/C++, we first present a coarse-grain intermediate representation called Function-ANd-Statement (FANS) which takes function call structure as well as statement structure into account. Based on the FANS intermediate representation, we propose a node fusion technique called Stratify And Grain-Controlled Fusion (SAGCF) to partition the whole application into many subtasks with the goal of maximizing parallelism in space and time dimensions as well as minimizing communication. All of these proposed techniques are implemented in an open source Automatic Task Partition Framework (ATPF). Finally, the feasibility of the proposed method is demonstrated by several cases. Ming-e Jing, Yibo Fan, Xiaoyong Xue, Xiaoyang Zeng, Zhiyi Yu |
ISCAS | 6 |
| 2018 | A Flexible and Energy-Efficient Convolutional Neural Network Acceleration With Dedicated ISA and Accelerator
Xiaobai Chen, Zhiyi Yu |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2017 | An FPGA-Based Hardware Accelerator for Traffic Sign DetectionabstractTraffic sign detection plays an important role in a number of practical applications, such as intelligent driver assistance and roadway inventory management. In order to process the large amount of data from either real-time videos or large off-line databases, a high-throughput traffic sign detection system is required. In this paper, we propose an FPGA-based hardware accelerator for traffic sign detection based on cascade classifiers. To maximize the throughput and power efficiency, we propose several novel ideas, including: 1) rearranged numerical operations; 2) shared image storage; 3) adaptive workload distribution; and 4) fast image block integration. The proposed design is evaluated on a Xilinx ZC706 board. When processing high-definition (1080p) video, it achieves the throughput of 126 frames/s and the energy efficiency of 0.041 J/frame. Weijing Shi, Xin Li 0001, Zhiyi Yu, Gary Overett |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2015 | Special Issue on Emerging Many-Core Systems for Exascale ComputingabstractNo abstract available. Masoud Daneshtalab, Farhad Mehdipour, Zhiyi Yu, Hannu Tenhunen |
ACM J. Emerg. Technol. Comput. Syst. | 3 |
| 2015 | A 65 nm Cryptographic Processor for High Speed Pairing ComputationabstractPairings are attractive and competitive cryptographic primitives for establishing various novel and powerful information security schemes. This paper presents a flexible and high-performance processor for cryptographic pairings over pairing-friendly curves at high security levels. In this design, hardware for Fp2arithmetic is optimized to accelerate the pairing computation, and especially a combined modular multiplier, which implements (AB + CD) based on Montgomery method, is proposed. This combined multiplier has the data path delay close to that of a single multiplier implementing (AB) but saves 20% area cost compared with two single multipliers. The Design I of the proposed processor is the first fabricated chip for pairing cryptography. An improved version, Design II, is implemented using TSMC 65-nm CMOS technology and achieves the working frequency of 633 MHz after placing and routing. As demonstrated, the optimal ate pairings of 126- and 128-bit security can be computed by Design II in 0.521 and 0.554 ms, respectively. These results outperform the hardware implementations reported by previous works. Jun Han 0003, Zhiyi Yu, Xiaoyang Zeng |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2015 | Many-Core Processors Granularity Evaluation by Considering Performance, Yield, and Lifetime ReliabilityabstractNetwork-on-chip based many-core processors break the limitations faced with single-core processors, and can bring high performance, improved yield, and high reliability. They are widely considered as the most promising platform. One of the most important questions is: what kind of granularity of cores (in other words, the number of cores and the area of each core under specific area constraint of the chip) is the best for many-core processors? Performance is widely used as the most important metric to determine the choice, but recently, reliability and yield are also becoming the first-class constraints in many application domains. This paper presents a novel performance model by considering practical program style, including pipelined and parallel program styles, as well as intercore communication, and proposes a new metric combing performance, reliability, and yield of many-core processors to choose the core granularity. According to our paper, with a given area (300 mm2) and certain applications, the optimal result is obtained using 8 × 8 mesh when all three factors are considered. Jianming Yu, Yueming Yang, Xiaodong Zhang 0024, Zhiyi Yu |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2015 | Design and Analysis of Highly Energy/Area-Efficient Multiported Register Files With Read Word-Line Sharing Strategy in 65-nm CMOS ProcessabstractThis brief proposes an ultralow-voltage four-read-port and two-write-port multiported register file with a novel architecture of read word-line sharing strategy for energy/area efficiency. Static read circuits and memory cells with nonminimum channel length are introduced to improve the ultralow-voltage performance. The chip of this register file is fabricated in 65-nm LP CMOS process and occupies the area of 0.019 mm$^{2}$ . Test results show that the minimum operation voltage is 320 mV with its corresponding max frequency 110 KHz. The minimum energy consumption is 0.94 pJ/cycle at the point of 400 mV, 850 KHz, corresponding to 0.15 fJ/port/bit/cycle after normalization. Compared with the state-of-the-art designs, it improves energy efficiency by 25% and saves the area by 58.7%. Xiaoyang Zeng, Yuejun Zhang, Shujie Tan, Jun Han 0003, Zhang Zhang 0004, Xu Cheng 0002, Zhiyi Yu |
IEEE Trans. Very Large Scale Integr. Syst. | 9 |
| 2014 | An Efficient Implementation of Montgomery Multiplication on Multicore Platform With Optimized Algorithm, Task Partitioning, and Network ArchitectureabstractThe modular multiplication (MM) is a key operation in cryptographic algorithms, such as RSA and elliptic-curve cryptography. Multicore processor is a suitable platform to implement MM because of its flexibility, high performance, and energy-efficiency. In this paper, we propose a block-level parallel algorithm for MM with quotient pipelining and optimally map it on a network-on-chip-based multicore platform equipped with broadcasting mechanism. Aiming at highest performance, a theoretical speedup model for parallel MM is also developed for parameter exploration that optimizes task partitioning. Experimental results based on a multicore prototype show that compared with the sequential MM on single core, the parallel implementation proposed in this paper maximizes the speedup ratio with regard to given intercore communication latency. Renfeng Dou, Jun Han 0003, Yifan Bo, Zhiyi Yu, Xiaoyang Zeng |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2013 | Implementation and optimization of 3780-point FFT on multi-core systemabstractThe 3780-point FFT is a main component of the time domain synchronous OFDM (TDS-OFDM) system in the Chinese Digital Terrestrial Multimedia Broadcasting (DTMB) national standard. In this paper, we proposed a pure software solution for the 3780-point FFT on a multi-core processor to achieve high performance and high flexibility. A new 12-point FFT implementation is used to improve system performance significantly since it is one of the key modules in 3780-point FFT. Together with some other techniques such as optimized assembly code, this 3780-point FFT improves the throughput by 39.26% and reduces the number of instructions by 19.9% compared with the non-optimized method. The throughput achieves 13.595 Msamples/s and meets the requirement of DMBT standard. Ming-e Jing, Zhiyi Yu, Xiaoyang Zeng, Jiayi Sheng, Haofan Yang 0001 |
ISCAS | 2 |
| 2013 | Time-Division-Multiplexer based routing algorithm for NoC systemabstractIn this paper, we present a routing algorithm based on the Time-Division-Multiplexer technique for routing table based Network-on-Chip (NoC) routers to decrease the demand of the system bandwidth while ensuring deadlock free. To fully use the communication resources of NoC — channels, banker algorithm is adopted to allocate and recycle the resources, and a weighted maze algorithm is utilized to determine if there is an available path for the current communication process. Experimental results show that the bandwidth requirement with the proposed algorithm decreases by 71.4% compared with the odd-even algorithm. Ming-e Jing, Zhiyi Yu, Xiaoyang Zeng, Liyang Zhou |
ISCAS | 2 |
| 2013 | A low power register file with asynchronously controlled read-isolation and software-directed write-discardingabstractThe register file (RF) consumes a large portion of power and is often a hotspot in microprocessor. In this paper, we propose and exploit several approaches to reducing both read and write access frequency to RF to reduce its power consumption and power density. Asynchronously controlled read-isolation is inserted in D Stage to prevent unused RF read access, triggered by a custom designed local asynchronous clock network, without changing the pipeline architecture and critical path. Software-directed write-discarding adopts static speculation algorithms to exploit short-lived values and determine their lifetime, with architectural supports to discard unnecessary writeback. Our approaches reduce RF access frequency by 27% for read and 50% for write, respectively. Moreover, 37% of RF power is eliminated with negligible overhead in area and almost no impact on the performance. Zheng Yu 0001, Xueqiu Yu, Xiaoyang Zeng, Zhiyi Yu |
ISCAS | 5 |
| 2013 | Parallelization of Radix-2 Montgomery Multiplication on Multicore PlatformabstractMontgomery multiplication is the kernel operation in public key ciphers. Aiming at parallel implementation of Montgomery multiplication, this brief presents an improved task partitioning of the Montgomery multiplication algorithm for the multicore platform with area-efficient processors. Several multicore platforms are designed to verify the efficiency of parallelization. The fastest platform takes 3460 cycles to finish a 1024-b Montgomery multiplication, which is six times faster than a single MIPS processor and three times faster than the pSHS parallelization based on a platform with eight MicroBlaze cores. Jun Han 0003, Zhiyi Yu, Xiaoyang Zeng |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2012 | Evaluating performance of manycore processors with various granularities considering yield and lifetime reliabilityabstractPerformance is one of the most important targets in MPSoC design, and it is affected by a lot of factors, such as area, yield, application, and lifetime. Based on the performance, yield and lifetime reliability modeling and analysis of MPSoC, this work proposes a metric directing how to choose the granularity of MPSoC at high level design in order to obtain high performance when yield and lifetime reliability are considered. The results show that, with the given area (300mm2) and certain applications, the optimal performance is obtained at 3×3 mesh, and optimal design becomes 4×4 mesh when yield is considered, and it will prefer 5×5 mesh or 6×6 mesh when lifetime reliability is further considered. Yueming Yang, Zewen Shi, Jianming Yu, Liulin Zhong, Xiaoyang Zeng, Zhiyi Yu |
ISCAS | 6 |
| 2012 | A pure software ldpc decoder on a multi-core processor platform with reduced inter-processor communication costabstractAs an error correction code, Low Density Parity Check (LDPC) code has been widely used in various communication standards such as WiMAX and DVB-S2. But these continuously-evolving communication standards and the high development cost and low-flexibility of hardwired ASIC solutions have pushed LDPC researchers to turn to more cost-efficient and flexible implementation, and thus the multi-core processor based implementation of LDPC decoder is gaining increasing attention in the last few years. However, the performance of the multi-core processor based implementation is far below the hardwired ASICs, with one of the key reasons that the cost of communication between processors is very high. Three approaches are proposed in this paper to reduce the communication cost, including: optimized algorithm partitioning to reduce communication traffic, utilizing imbalanced communication between tasks to optimize mapping and reduce overall communication distance, and simplified data sending-receiving mechanism to reduce the cost of identifying received data. By using these approaches, the communication time of the proposed implementation of LDPC decoder only accounts for 12.2% of total decoding time, which generally occupies 50% decoding time in the previously reported LDPC decoders on multi-core processors. And our work can achieve better throughput performance under the same hardware condition compared with other state-of-the-art works. Yan Ying, Kaidi You, Liyang Zhou, Heng Quan, Ming-e Jing, Zhiyi Yu, Xiaoyang Zeng |
ISCAS | 6 |
| 2012 | Task-binding based branch-and-bound algorithm for NoC mappingabstractNetwork-on-Chip (NoC) architecture is drawing intensive attention since it promises to maintain high performance in handling complex communication issues as the number of on-chip components increases. Mapping a given application onto the multi-core processors on NoC to obtain a high performance is a significant challenge. In this paper, we propose an optimized branch-and-bound (B&B) mapping algorithm to reduce the communication energy or improve the mapping efficiency by binding the tasks together when they have a large communication volume. Experimental results show that the proposed algorithm can achieve high performance in a short time compared with the traditional algorithm. For example, when mapping 64 tasks onto an 8×8 NoC system, with the approximate run time, 14.72% and 64.11% average energy consumption is saved compared with the original B&B and simulated annealing (SA) algorithms, respectively. Liyang Zhou, Ming-e Jing, Liulin Zhong, Zhiyi Yu, Xiaoyang Zeng |
ISCAS | 4 |
| 2011 | A reconfigurable and deadlock-free routing algorithm for 2D Mesh Network-on-ChipabstractThis paper presents a reconfigurable and deadlock- free routing (RDR) algorithm. It can be reconfigured to adapt to the modification of the topology due to faulty routers. It is evaluated from the point of view of performance penalty under various fault patterns. Meanwhile deadlock-freedom and reconfigure mechanism issues are addressed. Fault-tolerance capability, re-configurability and scalability are further evaluated and compared to several other routing algorithms. Zewen Shi, Yueming Yang, Xiaoyang Zeng, Zhiyi Yu |
ISCAS | 4 |
| 2011 | Fault tolerant computing for stream DSP applications using GALS multi-core processorsabstractThis paper presents a multi-core processor with globally asynchronous locally synchronous (GALS) clocking style designed to achieve soft error tolerance for stream DSP applications, and to maintain system throughput energy efficiently. Each processor in the chip can be combined with one of its neighbor processors to run the same programs and their results are equivalence checked to detect the soft error occurrence. When error occurs in some processor, the program in that processor (not the whole chip) is re-executed from the saved state to recover from the error. Due to the programming model of stream DSP applications, each processor can be isolated by FIFOs in the proposed multi-core processors, and fault detection and recovery can be done with low overhead. Furthermore, the GALS clocking style allows adjusting the frequency of the processors hit by a soft error-not the frequency of the whole chip-to maintain the system throughput, which results high energy efficiency. Zhiyi Yu, Zewen Shi, Xiaoyang Zeng |
ISCAS | 1 |
| 2010 | A scalable and fault-tolerant routing algorithm for NoCsabstractComputing design has been moving to multi-core or many-core domain and Network-on-chip (NoC) is upcoming. However, manufacturing defects and hard malfunction are inevitable, and fault-tolerant routing algorithm is important to provide the required communication in spite of failures. The proposed algorithm, referred to as scalable and fault-tolerant distributed routing (SFDR), partitions the system into nine regions using the concept of divide-and-conquer. Each region guarantees fault-tolerance of one's own area and the whole system still works no matter where the fault node locates. The novel routing algorithm has excellent scalability with hardware cost keeping constant independent of system size. The router has been synthesized using SMIC 0.13um CMOS process and there is almost no hardware overhead compared to Logic-Based Distributed Routing (LBDR) which is only partially fault-tolerant and hardware cost reduces up to 42% compared to table-based routing. Zewen Shi, Kaidi You, Yan Ying, Bei Huang, Xiaoyang Zeng, Zhiyi Yu |
ISCAS | 6 |
| 2010 | A Low-Area Multi-Link Interconnect Architecture for GALS Chip MultiprocessorsabstractA new inter-processor communication architecture for chip multiprocessors is proposed which has a low area cost, flexible routing capability, and supports globally asynchronous locally synchronous (GALS) clocking styles. To achieve a low area cost, the proposed statically-configurable asymmetric architecture assigns large buffer resources to only the nearest neighbor interconnect and much smaller buffer resources for long distance interconnect. To maintain flexible routing capability, each neighboring processor pair has multiple connecting links. The architecture supports long distance communication in GALS systems by transferring the source clock with the data signals along the entire path for write synchronization. Compared to a traditional dynamically-configurable interconnect architecture with symmetric buffer allocation and single-links between neighboring processor pairs, this implementation has approximately two times smaller communication circuitry area with a similar routing capability. Area and speed estimates are obtained with the physical design of seven chips in 0.18-¿m CMOS. Zhiyi Yu, Bevan M. Baas |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2009 | High Performance, Energy Efficiency, and Scalability With GALS Chip MultiprocessorsabstractChip multiprocessors with globally asynchronous locally synchronous (GALS) clocking styles are promising candidates for processing computationally-intensive and energy-constrained workloads. The GALS methodology simplifies clock tree design, provides opportunities to use clock and voltage scaling jointly in system submodules to achieve high energy efficiencies, and can also result in easily scalable clocking systems. However, its use typically also introduces performance penalties due to additional communication latency between clock domains. We show that GALS chip multiprocessors (CMPs) with large inter-processor first-inputs-first-outputs (FIFOs) buffers can inherently hide much of the GALS performance penalty while executing applications that have been mapped with few communication loops. In fact, the penalty can be driven tozerowith sufficiently large FIFOs and the removal of multiple-loop communication links. We present an example mesh-connected GALS chip multiprocessor and show it has a less than 1% performance (throughput) reduction on average compared to the corresponding synchronous system for many DSP workloads. Furthermore, adaptive clock and voltage scaling for each processor provides an approximately 40% power savings without any performance reduction. These results compare favorably with the GALS uniprocessor, which compared to the corresponding synchronous uniprocessor, has a reported greater than 10% performance (throughput) reduction and an energy savings of approximately 25% using dynamic clock and voltage scaling for many general purpose applications. Zhiyi Yu, Bevan M. Baas |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2008 | A low-area interconnect architecture for chip multiprocessorsabstractA new inter-processor communication architecture for chip multiprocessors is proposed which has a low area cost and flexible routing capability. To achieve a low area cost, the proposed statically-configurable asymmetric architecture assigns large buffer resources only to the nearest neighbor interconnect and much smaller buffer resources for long distance interconnect. To maintain flexible routing capability, each neighboring processor pair has two connecting links. Compared to a traditional dynamically-configurable interconnect architecture with symmetric buffer allocation and single-links between neighboring processor pairs, this implementation has approximately 2 times smaller communication circuitry area with a similar routing capability. Area and speed estimates are obtained with the physical design of seven chips in 0.18 μm CMOS. Zhiyi Yu, Bevan M. Baas |
ISCAS | 1 |
| 2007 | A Scalable Dual-Clock FIFO for Data Transfers Between Arbitrary and Haltable Clock DomainsabstractA robust, scalable, and power efficient dual-clock first-input first-out (FIFO) architecture which is useful for transferring data between modules operating in different clock domains is presented. The architecture supports correct operation in applications where multiple clock cycles of latency exist between the data producer, FIFO, and the data consumer; and with arbitrary clock frequency changes, halting, and restarting in either or both clock domains. The architecture is demonstrated in both a 0.18- mum CMOS full-custom design and a 0.18-mum CMOS standard cell design used in a globally asynchronous locally synchronous array processor. It achieves 580-MHz operation and 10.3-mW power dissipation while performing simultaneous FIFO read and write operations at 1.8 V. Ryan W. Apperson, Zhiyi Yu, Michael J. Meeuwsen, Tinoosh Mohsenin, Bevan M. Baas |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2006 | Hardware and applications of AsAP: An asynchronous array of simple processors
Bevan M. Baas, Zhiyi Yu, Michael J. Meeuwsen, Omar Sattari, Ryan W. Apperson, Eric W. Work, Jeremy W. Webb, Michael A. Lai, Daniel Gurman, Jason Cheung, Dean Nguyen Truong, Tinoosh Mohsenin |
Hot Chips Symposium | 2 |
| 2006 | Implementing Tile-based Chip Multiprocessors with GALS Clocking StylesabstractThis paper investigates implementation techniques for tile-based chip multiprocessors with Globally Asynchronous Locally Synchronous (GALS) clocking styles. These architectures can simplify the physical design flow since they allow focusing on a single processor when designing an entire chip. However, they also introduce challenges to maintain system robustness and scalability. We propose a physical design flow for these architectures, investigate timing issues for robust implementations, and propose methods to take full advantage of their potential scalability. As a design example, we present data from a recently implemented single-chip 6 x 6 tile-based GALS processing array. Zhiyi Yu, Bevan M. Baas |
ICCD | 1 |