EDBT 2026 Demo / reviewers in the wild / expert
Yihao Shen
dblp:370/2160
· DBLP profile ↗
7ranked-venue papers
1as first author
7since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 4 · 4 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Late Breaking Results: A Power-Efficient RISC-V Baseband System-on-Chip for Multi-Standard Integrated Sensing and CommunicationsabstractWe present Ishtar, a power-efficient RISC-V baseband system-on-chip (SoC) tailored for multi-standard integrated sensing and communications (ISAC) in low-altitude wireless networks (LAWNs). Ishtar integrates a hierarchical scheduling scheme and a system-level power-gating architecture that dynamically controls power domains to balance performance and energy efficiency. It supports dynamic task scheduling across heterogeneous protocols using a domain-specific, graph-based representation. Implemented in 40 nm technology and running at 300 MHz, Ishtar achieves better normalized efficiency than state-of-the-art SDR SoCs, delivering real-time multi-standard sniffing under stringent power and area constraints. Limin Jiang, Yi Shi 0004, Yihao Shen, Yintao Liu 0001, Siyi Xu, Qingyu Deng, Shan Cao 0001, Zhiyuan Jiang, Sheng Zhou 0001 |
DATE | 3 |
| 2026 | Venusian: Rapid Wireless Baseband Validation via High-Level Programming and FPGA-Based RISC-V Accelerator Co-Design
Limin Jiang, Yi Shi 0004, Yihao Shen, Yintao Liu 0001, Shan Cao 0001, Zhiyuan Jiang, Sheng Zhou 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2025 | A Hierarchical Dataflow-Driven Heterogeneous Architecture for Wireless Baseband ProcessingabstractWireless baseband processing (WBP) is a key element of wireless communications, with a series of signal processing modules to improve data throughput and counter channel fading. Conventional hardware solutions, such as digital signal processors (DSPs) and more recently, graphic processing units (GPUs), provide various degrees of parallelism, yet they both fail to take into account the cyclical and consecutive character of WBP. Furthermore, the large amount of data in WBPs cannot be processed quickly in symmetric multiprocessors (SMPs) due to the unpredictability of memory latency. To address this issue, we propose a hierarchical dataflow-driven architecture to accelerate WBP. A pack-and-ship approach is presented under a non-uniform memory access (NUMA) architecture to allow the subordinate tiles to operate in a bundled access and execute manner. We also propose a multi-level dataflow model and the related scheduling scheme to manage and allocate the heterogeneous hardware resources. Experiment results demonstrate that our prototype achieves 2× and 2.3× speedup in terms of normalized throughput and single-tile clock cycles compared with GPU and DSP counterparts in several critical WBP benchmarks. Additionally, a link-level throughput of 288 Mbps can be achieved with a 45-core configuration. Limin Jiang, Yi Shi 0004, Yintao Liu 0001, Qingyu Deng, Siyi Xu, Yihao Shen, Fangfang Ye, Shan Cao 0001, Zhiyuan Jiang |
ASP-DAC | 6 |
| 2025 | Zoozve: A Strip-Mining-Free RISC-V Vector Extension with Arbitrary Register Grouping Compilation Support (WIP)abstractVector processing is crucial for boosting processor performance and efficiency, particularly with data-parallel tasks. The RISC-V ”V” Vector Extension (RVV) enhances algorithm efficiency by supporting vector registers of dynamic sizes and their grouping. Nevertheless, for very long vectors, the static number of RVV vector registers and its power-of-two grouping can lead to performance restrictions. To counteract this limitation, this work introduces Zoozve, a RISC-V vector instruction extension that eliminates the need for strip-mining. Zoozve allows for flexible vector register length and count configurations to boost data computation parallelism. With a data-adaptive register allocation approach, Zoozve permits any register groupings and accurately aligns vector lengths, cutting down register overhead and alleviating performance declines from strip-mining. Additionally, the paper details Zoozve’s compiler and hardware implementations using LLVM and SystemVerilog. Initial results indicate Zoozve yields a minimum 10.10× reduction in dynamic instruction count for fast Fourier transform (FFT), with a mere 5.2% increase in overall silicon area. Siyi Xu, Limin Jiang, Yintao Liu 0001, Yihao Shen, Yi Shi 0004, Shan Cao 0001, Zhiyuan Jiang |
LCTES | 4 |
| 2025 | Skeleton-based action recognition through dual-granularity feature fusion with self-adapting graph convolution and multi-scale temporal convolution
Hao Chen 0049, Yihao Shen, Yuanxiang Zhang, Xiaoying Pan |
Neurocomputing | 2 |
| 2025 | Unlimited Vector Processing for Wireless Baseband Based on RISC-V ExtensionabstractWireless baseband processing (WBP) serves as an ideal scenario for utilizing vector processing, which excels in managing data-parallel operations due to its parallel structure. However, conventional vector architectures face certain constraints such as limited vector register sizes, reliance on power-of-two vector length (VL) multipliers, and vector permutation capabilities tied to specific architectures. To address these challenges, we have introduced an instruction set extension (ISE) based on RISC-V known as unlimited vector processing (UVP). This extension enhances both the flexibility and efficiency of vector computations. UVP employs a novel programming model that supports non-power-of-two register groupings (RGs) and hardware strip mining, thus enabling smooth handling of vectors of varying lengths while reducing the software strip-mining burden. Vector instructions are categorized into symmetric and asymmetric classes, complemented by specialized load/store strategies to optimize execution. Moreover, we present a hardware implementation of UVP featuring sophisticated hazard detection mechanisms, optimized pipelines for symmetric tasks such as fixed-point multiplication and division, and a robust permutation engine for effective asymmetric operations. Comprehensive evaluations demonstrate that UVP significantly enhances performance, achieving up to$3.0\times $and$2.1\times $speedups in matrix multiplication and fast Fourier transform (FFT) tasks, respectively, when measured against lane-based vector architectures. Our synthesized register transfer level (RTL) for a 16-lane configuration using SMIC 40-nm technology spans 0.94 mm2and achieves an area efficiency of 21.2 GOPS/mm2. Limin Jiang, Yi Shi 0004, Yihao Shen, Shan Cao 0001, Zhiyuan Jiang, Sheng Zhou 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2023 | Parallel Computing for Energy-Efficient Baseband Processing in O-RAN: Synchronization and OFDM Implementation Based on SPMDabstractOpen radio access network (O-RAN) is considered as a viable method for reducing the cost and enhancing the energy efficiency of cellular networks, due to its native incorporation of intelligence and open interfaces. However, the processing delay of software-based wireless protocol stacks has hindered its development. This paper presents the implementation of parallel computing acceleration for an LTE baseband system based on single program multiple data (SPMD) methodology, and proposes detailed optimization strategies for the time-consuming synchronization and OFDM modulation modules in the system. Experiment results based on the implicit SPMD program compiler (ISPC) show that continuous memory access has a significant impact on the final acceleration effect. Moreover, the processing speed of software-based physical layer can be increased up to 10–30 times through SPMD. Yihao Shen, Shan Cao 0001, Zhiyuan Jiang, Sheng Zhou 0001 |
GLOBECOM | 1 |