Yinan Xu 0001

dblp:76/10563-1 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
10since 2021 · last 2026
0000-0002-9702-2542ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 2 first-author · 9 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 FastDSE: Enabling Efficient CPU Microarchitecture Design Space Exploration with FPGA Acceleration
abstract
Design Space Exploration (DSE) is essential for tuning CPU microarchitectural parameters to achieve favorable trade-offs among performance, power, and area. Prior efforts have mainly focused on improving search algorithms to find better design points; however, the growing complexity of modern processors has made design evaluation increasingly time consuming, with simulation costs becoming the dominant bottleneck that limits the efficiency of DSE. To address this challenge, we present FastDSE, an open-source FPGA accelerated DSE framework. FastDSE offloads RTL simulation to the FPGA and executes the DSE search algorithm on the host. To support rapid evaluation of different design points, it decouples logical microarchitectural parameters from physical hardware resources with minor RTL modifications, enabling runtime CPU parameter reconfiguration without FPGA re-synthesis. To support diverse DSE search algorithms and processor designs, FastDSE provides lightweight APIs for host-FPGA interaction and simulation metrics collection. Evaluated on an out-of-order RISC-V processor and representative algorithms, FastDSE reduces simulation time by 99.87%, boosts exploration throughput by up to 68.7 ×, and delivers up to 30.7% PPA improvement over the baseline.
Kaifan Wang, Jiabin Wu, Yinan Xu 0001, Ninghui Sun, Yungang Bao
ACM Great Lakes Symposium on VLSI3
2026 TraceRTL: Agile Performance Evaluation for Microarchitecture Exploration
abstract
While agile chip development methodologies have accelerated RTL design and simulation, performance evaluation remains constrained by three challenges: (1) inefficient feature prototyping caused by the tight coupling between functional correctness and performance evaluation, particularly for large-scale, error-prone microarchitectures; (2) limited workloads due to incomplete peripheral/software environments or unavailable source code; and (3) time-consuming warm-up phases in sampling-based simulation, required to mitigate cold-start effects. To address these challenges, we propose TRACERTL, an agile, trace-driven performance evaluation methodology that decouples the functional and performance components of CPU RTL designs. It introduces three techniques: (1) a trace-driven performance exploration framework that bypasses full functional correctness while preserving performance accuracy; (2) a trace transformation technique, TraceBridge, that replays traces across different formats and instruction sets; and (3) a fast warm-up strategy, TraceDedup, that eliminates redundant traces and efficiently initializes microarchitectural states. Using TRACERTL, we develop the first trace-driven RTL CPU derived from XiangShan, a high-performance out-of-order RISC-V processor. TRACERTL achieves performance accuracies of 99.87% and 99.86% on SPECint2017 and SPECfp2017, respectively. With TraceBridge, we evaluate x86-based Google workload traces on a RISC-V RTL CPU and reveal distinct memory-bound behavior. TraceDedup further accelerates warm-up phases in sampling-based simulations by$\text{1. 5} \times$to$\text{1 1. 8} \times$.
Zifei Zhang 0001, Yinan Xu 0001, Sa Wang, Dan Tang 0002, Yungang Bao
HPCA2
2026 Democratizing and Accelerating Hardware Verification with Software-Native Optimization
Yunlong Xie, Zhicheng Yao, Fangyuan Song, Junyue Wang, Haojin Tang, Yinan Xu 0001, Ziyuan Gao, Duan Yu, Jiayi Rao, Junyu Yue, Yunqi Lu, Zechen Yang, Xu An, Qi Ge, Jiuyue Ma, Jian-Yi Meng, Kan Shi, Dan Tang 0002, Sa Wang, Yungang Bao
ISCA8
2026 CodeV: Empowering LLMs With HDL Generation Through Multilevel Summarization
abstract
The design flow of processors, particularly in hardware description languages (HDL) like Verilog and Chisel, is complex and costly. While recent advances in large language models (LLMs) have significantly improved coding tasks in software languages such as Python, their application in HDL generation remains limited due to the scarcity of high-quality HDL data. Traditional methods of adapting LLMs for hardware design rely on synthetic HDL datasets, which often suffer from low quality because even advanced LLMs like GPT perform poorly in the HDL domain. Moreover, these methods focus solely on chat tasks and the Verilog language, limiting their application scenarios. In this paper, we observe that: (1) HDL code collected from the real world is of higher quality than code generated by LLMs. (2) LLMs like GPT-3.5 excel in summarizing HDL code rather than generating it. (3) An explicit language tag can help LLMs better adapt to the target language when there is insufficient data. Based on these observations, we propose an efficient LLM fine-tuning pipeline for HDL generation that integrates a multi-level summarization data synthesis process with a novel Chat-FIM-Tag supervised fine-tuning method. The pipeline enhances the generation of HDL code from natural language descriptions and enables the handling of various tasks such as chat and infilling incomplete code. Utilizing this pipeline, we introduce CodeV, a series of HDL generation LLMs. Among them, CodeV-All not only possesses a more diverse range of language abilities (Verilog and Chisel) and a broader scope of tasks (Chat and FIM), but also achieves performance on VerilogEval that is comparable to that of CodeV-Verilog fine-tuned on Verilog only, making them the first series of open-source LLMs designed for multi-scenario HDL generation. Code, models, and dataset: https://github.com/IPRC-DIP/CodeV.
Yang Zhao 0013, Chongxiao Li, Pengwei Jin, Muxin Song, Yinan Xu 0001, Ziyuan Nan, Mingju Gao, Tianyun Ma, Yansong Pan, Rui Zhang 0040, Xishan Zhang, Zidong Du, Qi Guo 0001, Xing Hu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2025 DiffTest-H: Toward Semantic-Aware Communication in Hardware-Accelerated Processor Verification
Kunlin You, Yinan Xu 0001, Kehan Feng, Luoshan Cai, Yaoyang Zhou, Yungang Bao
MICRO2
2024 PathFuzz: Broadening Fuzzing Horizons with Footprint Memory for CPUs
abstract
Coverage metrics have been widely adopted to quantify the completeness of hardware verification. Recently, coverage-guided fuzzing has emerged as a popular method for automatically creating test inputs toward higher verification coverage reach. However, we observe that its effectiveness on CPUs is hindered by limited sources of seed corpus and efficiency of mutations. To broaden the fuzzing horizons, this paper proposes the PathFuzz framework incorporating an efficient input format for fuzzing CPUs, the footprint memory, with seed corpus from real-world large-scale programs. Experiments demonstrate that using PathFuzz reaches over 95% verification coverage with four long-standing bugs newly identified in two well-known open-source CPU designs.
Yinan Xu 0001, Sa Wang, Dan Tang 0002, Ninghui Sun, Yungang Bao
DAC1
2024 XiangShan: An Open-Source Project for High-Performance RISC-V Processors Meeting Industrial-Grade Standards
abstract
•Overview •Microarchitecture design •Agile development platform •Applications in industry & academia •Summary
Kaifan Wang, Yinan Xu 0001, Zifei Zhang 0001, Guokai Chen, Linjuan Zhang, Dan Tang 0002, Ninghui Sun, Yungang Bao
HCS3
2023 Functional Verification for Agile Processor Development: A Case for Workflow Integration
Yinan Xu 0001, Kaifan Wang, Huaqiang Wang, Linjuan Zhang, Zifei Zhang 0001, Dan Tang 0002, Sa Wang, Kan Shi, Ninghui Sun, Yungang Bao
J. Comput. Sci. Technol.1
2022 Towards Developing High Performance RISC-V Processors Using Agile Methodology
abstract
While research has shown that the agile chip design methodology is promising to sustain the scaling of computing performance in a more efficient way, it is still of limited usage in actual applications due to two major obstacles: 1) Lack of tool-chain and developing framework supporting agile chip design, especially for large-scale modern processors. 2) The conventional verification methods are less agile and become a major bottleneck of the entire process. To tackle both issues, we propose MINJIE, an open-source platform supporting agile processor development flow. MINJIE integrates a broad set of tools for logic design, functional verification, performance modelling, pre-silicon validation and debugging for better development efficiency of state-of-the-art processor designs. We demonstrate the usage and effectiveness of MINJIE by building two generations of an open-source superscalar out-of-order RISC-V processor code-named XIANGSHAN using agile methodologies. We quantify the performance of XIANGSHAN using SPEC CPU2006 benchmarks and demonstrate that XIANGSHAN achieves industry-competitive performance.
Yinan Xu 0001, Dan Tang 0002, Guokai Chen, Lingrui Gou, Qianruo Li, Zuojun Li, Jiazhan Tan, Huaqiang Wang, Huizhe Wang, Kaifan Wang, Chuanqi Zhang, Fawang Zhang, Linjuan Zhang, Zifei Zhang 0001, Yaoyang Zhou, Yike Zhou, Jiangrui Zou, Ye Cai 0001, Dandan Huan, Zusong Li, Jiye Zhao, Qiyuan Quan, Xingwu Liu, Sa Wang, Kan Shi, Ninghui Sun, Yungang Bao
MICRO1
2021 Omegaflow: a high-performance dependency-based architecture
abstract
This paper investigates how to better track and deliver dependency in dependency-based cores to exploit instruction-level parallelism (ILP) as much as possible. To this end, we first propose an analytical performance model for the state-of-art dependency-based core, Forwardflow, and figure out two vital factors affecting its upper bound of performance. Then we propose Omegaflow,a dependency-based architecture adopting three new techniques, which respond to the discovered factors. Experimental results show that Omegaflow improves IPC by 24.6% compared to the state-of-the-art design, approaching the performance of the OoO architecture with an ideal scheduler (94.4%) without increasing the clock cycle and consumes only 8.82% more energy than Forwardflow.
Yaoyang Zhou, Chuanqi Zhang, Yinan Xu 0001, Huizhe Wang, Sa Wang, Ninghui Sun, Yungang Bao
ICS4