Yintao Liu 0001

dblp:49/426-1 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
5since 2021 · last 2026
0009-0002-3538-4676ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Late Breaking Results: A Power-Efficient RISC-V Baseband System-on-Chip for Multi-Standard Integrated Sensing and Communications
abstract
We present Ishtar, a power-efficient RISC-V baseband system-on-chip (SoC) tailored for multi-standard integrated sensing and communications (ISAC) in low-altitude wireless networks (LAWNs). Ishtar integrates a hierarchical scheduling scheme and a system-level power-gating architecture that dynamically controls power domains to balance performance and energy efficiency. It supports dynamic task scheduling across heterogeneous protocols using a domain-specific, graph-based representation. Implemented in 40 nm technology and running at 300 MHz, Ishtar achieves better normalized efficiency than state-of-the-art SDR SoCs, delivering real-time multi-standard sniffing under stringent power and area constraints.
Limin Jiang, Yi Shi 0004, Yihao Shen, Yintao Liu 0001, Siyi Xu, Qingyu Deng, Shan Cao 0001, Zhiyuan Jiang, Sheng Zhou 0001
DATE4
2026 Venusian: Rapid Wireless Baseband Validation via High-Level Programming and FPGA-Based RISC-V Accelerator Co-Design
Limin Jiang, Yi Shi 0004, Yihao Shen, Yintao Liu 0001, Shan Cao 0001, Zhiyuan Jiang, Sheng Zhou 0001
IEEE Trans. Very Large Scale Integr. Syst.4
2025 A Hierarchical Dataflow-Driven Heterogeneous Architecture for Wireless Baseband Processing
abstract
Wireless baseband processing (WBP) is a key element of wireless communications, with a series of signal processing modules to improve data throughput and counter channel fading. Conventional hardware solutions, such as digital signal processors (DSPs) and more recently, graphic processing units (GPUs), provide various degrees of parallelism, yet they both fail to take into account the cyclical and consecutive character of WBP. Furthermore, the large amount of data in WBPs cannot be processed quickly in symmetric multiprocessors (SMPs) due to the unpredictability of memory latency. To address this issue, we propose a hierarchical dataflow-driven architecture to accelerate WBP. A pack-and-ship approach is presented under a non-uniform memory access (NUMA) architecture to allow the subordinate tiles to operate in a bundled access and execute manner. We also propose a multi-level dataflow model and the related scheduling scheme to manage and allocate the heterogeneous hardware resources. Experiment results demonstrate that our prototype achieves 2× and 2.3× speedup in terms of normalized throughput and single-tile clock cycles compared with GPU and DSP counterparts in several critical WBP benchmarks. Additionally, a link-level throughput of 288 Mbps can be achieved with a 45-core configuration.
Limin Jiang, Yi Shi 0004, Yintao Liu 0001, Qingyu Deng, Siyi Xu, Yihao Shen, Fangfang Ye, Shan Cao 0001, Zhiyuan Jiang
ASP-DAC3
2025 Zoozve: A Strip-Mining-Free RISC-V Vector Extension with Arbitrary Register Grouping Compilation Support (WIP)
abstract
Vector processing is crucial for boosting processor performance and efficiency, particularly with data-parallel tasks. The RISC-V ”V” Vector Extension (RVV) enhances algorithm efficiency by supporting vector registers of dynamic sizes and their grouping. Nevertheless, for very long vectors, the static number of RVV vector registers and its power-of-two grouping can lead to performance restrictions. To counteract this limitation, this work introduces Zoozve, a RISC-V vector instruction extension that eliminates the need for strip-mining. Zoozve allows for flexible vector register length and count configurations to boost data computation parallelism. With a data-adaptive register allocation approach, Zoozve permits any register groupings and accurately aligns vector lengths, cutting down register overhead and alleviating performance declines from strip-mining. Additionally, the paper details Zoozve’s compiler and hardware implementations using LLVM and SystemVerilog. Initial results indicate Zoozve yields a minimum 10.10× reduction in dynamic instruction count for fast Fourier transform (FFT), with a mere 5.2% increase in overall silicon area.
Siyi Xu, Limin Jiang, Yintao Liu 0001, Yihao Shen, Yi Shi 0004, Shan Cao 0001, Zhiyuan Jiang
LCTES3
2025 A Heterogeneous CNN Compilation Framework for RISC-V CPU and NPU Integration Based on ONNX-MLIR
abstract
With the continuous advancement of convolutional neural networks (CNNs), many neural network processing units (NPU) have emerged in recent years. NPUs offer advantages such as improved energy efficiency and low latency compared to traditional processors. However, NPUs often struggle to keep pace with the growing complexity of algorithmic models, which limits the implementation of certain AI applications. Central processing units (CPU) are known for their versatility, but are computationally inefficient. Combining the strengths of both architectures could present a viable solution to these challenges. Despite this potential, there is currently insufficient academic focus on CPU/NPU co-computing, and few architectures or toolchains have been developed to support such integration. In this paper, we propose a RISC-V CPU/NPU co-computing-based compilation framework to solve the problem of mapping from algorithms to heterogeneous architectures. To bridge the computational gap between heterogeneous platforms, within ONNX-MLIR, we developed three phases to support operator packing, algorithm adaption, and data coordination. To integrate heterogeneous compilation platforms, we introduce four additional functions to handle task scheduling, weight quantization, weight rearrangement, and memory management. In particular: (1) to optimize the transition overhead, an operator packing technique is introduced by extending the ONNX-MLIR framework with custom operators, which enhances NPU efficiency by reducing the number of transitions between operations. (2) At the aspect of task scheduling, a three-mode task scheduler is developed to allow users to customize task distribution for correctness verification, performance optimization, or detailed analysis on heterogeneous platforms, which demonstrates the versatility of our framework. (3) In terms of memory management, an optimized scheme for CPU/NPU architectures is proposed to ensure accurate data access. This memory management employs a dual-end approach with self-checks and a release-after-use mechanism to further improve efficiency. Experimental results demonstrate that the framework effectively generates instructions for various CNNs, significantly enhancing the efficiency and performance of heterogeneous architectures. It achieves up to a 7.06× reduction in code density and up to a 5.58× improvement in memory usage, narrowing the computational gap between CPUs and NPUs.
Shan Cao 0001, Meiling Yang, Yintao Liu 0001, Yu Li 0051, Beining Zhao 0001, Xinyu Chen 0007, Zhiyuan Jiang
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3