VLDB 2026 Research / reviewers in the wild / expert
Ruozhou Xiao
dblp:309/9320
· DBLP profile ↗
8ranked-venue papers
0as first author
8since 2021 · last 2026
0009-0001-9862-3278ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 7 · 7 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Boosting Vector Instruction Throughput in RISC-V via a Hybrid Decoupled Architecture with VLIW-Driven Execution
Weiying Wang, Qingchen Zhai, Ruozhou Xiao |
ISCAS | 3 |
| 2026 | RTCore: A RISC-V Processor Featuring Nested Hardware Loop Optimization for Real-time Control System
Qingchen Zhai, Weiying Wang, Yuanyang Xiang, Kunyu Zong, Ruozhou Xiao |
ISCAS | 7 |
| 2026 | Breaking the Local Optima Barrier in Branch Predictor Design Space Exploration: An LLM-Based Initialization Strategy for PSO
Qingchen Zhai, Yuanyang Xiang, Weiying Wang, Kunyu Zong, Jingshuai Jiang, Ruozhou Xiao |
ISCAS | 8 |
| 2026 | LoopHint: A Compiler-Assisted Loop Branch Predictor for Embedded DSPsabstractLoop-intensive computations dominate embedded DSP applications, making loop branch prediction critical for maintaining pipeline efficiency. However, a gap exists between resource-heavy dynamic predictors and inflexible Zero Overhead Looping (ZOL) hardware. Conventional predictors are often too costly for embedded silicon, while traditional ZOL is limited by strict constraints on loop size, nesting, and data-dependent exit conditions. To address these limitations, we introduce LoopHint, a hardware-software co-design approach that provides a flexible, learning-free loop prediction mechanism. LoopHint utilizes compiler-time static analysis to identify deterministic loop patterns and dependency chains, communicating this metadata to a simplified hardware loop table via specialized instructions. This co-design approach allows the processor to bypass the warm-up and aliasing issues of dynamic predictors while supporting a broader range of loop structures than ZOL. Experimental results on DSP kernels and the Embench-IoT benchmark suite show that LoopHint achieves an average 67% reduction in branch mispredictions. With minimal impact on the area and power, LoopHint provides a scalable solution for high-efficiency control flow in modern embedded DSPs. Yuanyang Xiang, Ruozhou Xiao, Zhiwei Zhang 0026 |
LCTES | 3 |
| 2025 | A Lightweight RISC-V Multi-Core Interaction Framework For Embedded Real-Time ApplicationsabstractMulti-core architectures have been adopted to enhance computation capability in embedded real-time systems, yet the interaction between cores affects performance, hardware complexity, and utilization of the systems. Current approaches suffer from high complexity or limited flexibility. We present a lightweight task-based multi-core interaction framework compatible with RISC-V specifications, balancing complexity and flexibility. The method proposed relies on RISC-V Core Local Interrupt Controller (CLIC) with customized extensions on the standard AMBA bus offering fully programmable capability for task offloading/trigger and shared resource protection. Compared with the interconnect between TI’s C2000 Series cores and its Control Law Accelerator(CLA) co-processors, experiments show the method reduces more than 57% of the response cycles cost by multi-core task management, but only introduces overhead of less than 1% area of a dual-issue RISC-V core with 7-stage in-order pipeline. Yuanyang Xiang, Weiying Wang, Qingchen Zhai, Kunyu Zong, Ruozhou Xiao |
ISCAS | 7 |
| 2025 | Using Greedy-Enhanced Heuristic to Optimize Instruction Scheduling for RISC-V DSabstractWith the rapid growth of the Internet of Things (IoT) and embedded systems, the demand for RISC-V architecture in DSPs (Digital Signal Processors) has increased significantly, making processor performance optimization a key research focus. As a core aspect of compiler optimization, instruction scheduling significantly affects execution efficiency and system performance. However, existing heuristic scheduling methods often fail to fully exploit hardware features. To address these limitations, this paper proposes a greedy-enhanced heuristic scheduling strategy that integrates the greedy algorithm’s fast decision-making with an adaptive weighted scoring model to prevent local optima. This approach effectively balances multiple factors such as instruction latency, register pressure, and resource conflicts, thereby enhancing scheduling performance. Experimental results on the RISC-V DSP platform, SummerCore, show that the proposed method achieves significant performance gains, with an average improvement of over 5% compared to LLVM -O3 optimization across multiple benchmarks. Yuanyang Xiang, Qingchen Zhai, Ruozhou Xiao |
ISCAS | 4 |
| 2025 | Instruction Level Parallelism Optimizations in a High-Performance Dual-Issue RISC-V Processor for Real-Time Control SystemsabstractThis paper presents optimization efforts on instruction-level parallelism in D-RTCore, a seven-stage dual-issue RISC-V processor designed for real-time control applications. To enhance instruction fetch and dispatch efficiency, we introduce four key techniques: branch prediction optimization, variable-length hardware loops, dynamic instruction fusion, and out-of-order write-back without a reorder buffer (ROB). Performance evaluations reveal that D-RTCore achieves cycle reductions of up to 79% compared to the industry-standard TI C28x. For example, it executes the complex fast fourier transform with a 44% reduction in cycles and the fast fourier transform magnitude with a remarkable 79% decrease in cycles. These results highlight D-RTCore’s advanced execution capabilities, making it a competitive solution for real-time control systems that require efficient processing. Qingchen Zhai, Weiying Wang, Yuanyang Xiang, Kunyu Zong, Ruozhou Xiao |
ISCAS | 7 |
| 2024 | LLM Based End-to-end Branch Predictor Optimization GeneratorabstractThe branch predictor is one of the crucial components influencing the performance of modern processors. This study aims to explore the optimization space for configurable parameters within the branch predictor and, based on the parameter results, generate RTL code end-to-end through a large language model, thereby enhancing the efficiency of processor microarchitecture design. In the research process, we take the TAGE branch predictor as an example, explore it towards the real-time control domain function library, and construct a DSP processor branch predictor for an eight-stage pipelined dual-issue architecture based on LLM. Qingchen Zhai, Ruozhou Xiao |
ASAP | 3 |