Wenshuo Yue

dblp:316/3290 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
4since 2021 · last 2026
0000-0002-4339-0489ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Reconfigurable Computing Challenge: FPGA-Gym-v2: FPGA-Based RL Environment Acceleration with LLM-Assisted Onboarding
abstract
Reinforcement learning (RL) repeatedly executes environment rollouts computation, where environment stepping may dominate training time when many environments run in parallel. Building on our prior FPGA-Gym/PEARL acceleration backbone, this work introduces FPGA-Gym-v2, a upgraded framework that keeps RL inference, training, and replay-buffer management on the host CPU/GPU, but offloads massively parallel environment stepping to an FPGA. It reduces host–FPGA traffic with compact encoding, on-chip state residency, and compute–communication overlap. Across representative Gymnasium workloads, FPGA-Gym-v2 achieves 4.36×–972.6× throughput speedup over EnvPool and up to 4699.3× over VectorEnv, while reducing DQN/PPO training time by 7%–15% over EnvPool and 18%–61% over VectorEnv. Further, FPGA-Gym-v2 adds an large-language-model-assisted (LLM-assisted) onboarding flow that generates environment-specific specifications, wrappers, Verilog HDL skeletons, and golden tests under explicit validation gates. The code is available at https://github.com/Selinaee/FPGA_Gym.
Hongxiao Zhao, Wenshuo Yue, Yihan Fu, Daijing Shi, Anjunyi Fan, Bonan Yan
FCCM3
2025 PEARL: FPGA-Based Reinforcement Learning Acceleration with Pipelined Parallel Environments
abstract
Reinforcement learning (RL) is an effective machine learning approach that enables artificial intelligence agents to perform complex tasks and make decisions in dynamic situations. Training an RL agent demands its repetitive interaction with the environment to learn optimal policies. To efficiently collect training data, parallelizing environments is a widely used technique by enabling simultaneous interactions between multiple agents and environments. However, existing CPU-based RL software frameworks face a key challenge of slow multi-environmental update computation. To solve this problem, we present a novel FPGA-based RL accelerating framework-PEARL. PEARL instantiates multiple parallel environments and accelerates them with a carefully designed pipeline scheme to hide data transfer latency within the computation time. We evaluate PEARL on respective RL environments and achieve 4.36 x to 972.6 x speedup over the existing fastest software-based framework for parallel environment execution. When scaling the number of environments from 1024 to 43008 (42x) in CliffWalking benchmark, the power consumption increases marginally by 3%, while LUT and flip-flops utilization rise by 2.24 x and 3.08 x, respectively. This demonstrates efficient resource usage and power management in PEARL. Further, PEARL allows users to define and add their environments within the framework flexibly. We have established an open-source repository for users to utilize and expand. We also implement PEARL with the existing RL algorithm and save 7% -15% training time. All the source code is available online https://github.com/Selinaee/FPGA_Gym.
Hongxiao Zhao, Wenshuo Yue, Yihan Fu, Daijing Shi, Anjunyi Fan, Yuchao Yang 0001, Bonan Yan
DATE3
2025 PROCA: Programmable Probabilistic Processing Unit Architecture with Accept/Reject Prediction & Multicore Pipelining for Causal Inference
abstract
Causal inference is an important field in data science and cognitive artificial intelligence. It requires the construction of complex probabilistic models to describe the causal relationships between random variables. Probabilistic models rely on probabilistic programming as a flexible framework. However, the computing speed of probabilistic programming is often hindered by the extensive use of Markov chain Monte Carlo (MCMC) algorithms, even though they are powerful in Bayesian inference. To accelerate MCMC, this work presents PROCA, a programmable MCMC-based probabilistic processing unit architecture. PROCA exploits processing-in-memory function units to generate new samples of Markov chains. PROCA is programmable to execute the computation for arbitrary forms of posterior distribution formulas that software probabilistic programming frameworks support. We develop a novel accept/reject prediction methodology to accelerate the sequential MCMC computation, thereby introducing efficient multi-core pipelining methods. We implement and validate the PROCA architecture with commercial process development kits. The implementation is evaluated based on 9 representative benchmarks, covering PyMC official tutorial probabilistic problems, single-variable probabilistic problems, and real-world causal inference problems. Our comprehensive experiments demonstrate that PROCA achieves a speedup of 172~4871 $\times$ compared to Intel Xeon Gold CPU, $42 \sim 1058 \times$ compared to NVIDIA A100 GPU, and $1.765 \times$ over state-of-the-art MCMC accelerators, respectively. PROCA achieves comparable statistical robustness to the software probabilistic programming frameworks. Compared with state-of-the-art MCMC domain-specific accelerators, our design boosts the energy efficiency by $9.47 \times$.
Yihan Fu, Anjunyi Fan, Wenshuo Yue, Hongxiao Zhao, Daijing Shi, Qiuping Wu, Yaoyu Tao, Yuchao Yang 0001, Bonan Yan
HPCA3
2024 Probabilistic Compute-in-Memory Design for Efficient Markov Chain Monte Carlo Sampling
abstract
Markov chain Monte Carlo (MCMC) is a widely used sampling method in modern artificial intelligence and probabilistic computing systems. It involves repetitive random number generations and thus often dominates the latency of probabilistic model computing. Hence, we propose a compute-in-memory (CIM) based MCMC design as a hardware acceleration solution. This work investigates SRAM bitcell stochasticity and proposes a novel “pseudo-read” operation, based on which we offer a block-wise random number generation circuit scheme for fast random number generation. Moreover, this work proposes a novel multi-stage exclusive-OR gate (MSXOR) design method to generate strictly uniformly distributed random numbers. The probability error deviating from a uniform distribution is suppressed under$10^{-6}$. Also, this work presents a novel in-memory copy circuit scheme to realize data copy inside a CIM sub-array, significantly reducing the use of R/W circuits for power saving. Evaluated in a commercial 28-nm process development kit, this CIM-based MCMC design generates 4-bit$\sim$32-bit samples with an energy efficiency of 0.53 pJ/sample and high throughput of up to 1066.7M samples/s. Compared to conventional processors, the overall energy efficiency improves$2.12\times10^{9}$to$9.58\times10^{9}$times.
Yihan Fu, Daijing Shi, Anjunyi Fan, Wenshuo Yue, Yuchao Yang 0001, Ru Huang 0001, Bonan Yan
IEEE Trans. Circuits Syst. I Regul. Pap.4