Yuchao Yang 0001

dblp:212/7311-1 · DBLP profile ↗
← Back
19ranked-venue papers
0as first author
15since 2021 · last 2026
0000-0003-4674-4059ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 10 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 5 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021
YearPublicationVenuePosition
2026 MGPA: A Memristor-based Genome Processing Accelerator for Single-cell RNA Sequencing
Lianfeng Yu, Yihang Zhu, Yaoyu Tao, Yuchao Yang 0001
DATE8
2026 Phase-Change Memory Based Passive Timing Circuits for DRAM Refresh Management
Anjunyi Fan, Yingming Lu, Bonan Yan, Yuchao Yang 0001
ISCAS6
2026 A Logic-Assisted Hybrid-Mode Ising Machine Based on a 3D Topology for 3D Path Planning
Runhong Tang, Xiangrui Wang, Yuchao Yang 0001, Yuqi Su
ISCAS3
2025 PEARL: FPGA-Based Reinforcement Learning Acceleration with Pipelined Parallel Environments
abstract
Reinforcement learning (RL) is an effective machine learning approach that enables artificial intelligence agents to perform complex tasks and make decisions in dynamic situations. Training an RL agent demands its repetitive interaction with the environment to learn optimal policies. To efficiently collect training data, parallelizing environments is a widely used technique by enabling simultaneous interactions between multiple agents and environments. However, existing CPU-based RL software frameworks face a key challenge of slow multi-environmental update computation. To solve this problem, we present a novel FPGA-based RL accelerating framework-PEARL. PEARL instantiates multiple parallel environments and accelerates them with a carefully designed pipeline scheme to hide data transfer latency within the computation time. We evaluate PEARL on respective RL environments and achieve 4.36 x to 972.6 x speedup over the existing fastest software-based framework for parallel environment execution. When scaling the number of environments from 1024 to 43008 (42x) in CliffWalking benchmark, the power consumption increases marginally by 3%, while LUT and flip-flops utilization rise by 2.24 x and 3.08 x, respectively. This demonstrates efficient resource usage and power management in PEARL. Further, PEARL allows users to define and add their environments within the framework flexibly. We have established an open-source repository for users to utilize and expand. We also implement PEARL with the existing RL algorithm and save 7% -15% training time. All the source code is available online https://github.com/Selinaee/FPGA_Gym.
Hongxiao Zhao, Wenshuo Yue, Yihan Fu, Daijing Shi, Anjunyi Fan, Yuchao Yang 0001, Bonan Yan
DATE7
2025 PROCA: Programmable Probabilistic Processing Unit Architecture with Accept/Reject Prediction & Multicore Pipelining for Causal Inference
abstract
Causal inference is an important field in data science and cognitive artificial intelligence. It requires the construction of complex probabilistic models to describe the causal relationships between random variables. Probabilistic models rely on probabilistic programming as a flexible framework. However, the computing speed of probabilistic programming is often hindered by the extensive use of Markov chain Monte Carlo (MCMC) algorithms, even though they are powerful in Bayesian inference. To accelerate MCMC, this work presents PROCA, a programmable MCMC-based probabilistic processing unit architecture. PROCA exploits processing-in-memory function units to generate new samples of Markov chains. PROCA is programmable to execute the computation for arbitrary forms of posterior distribution formulas that software probabilistic programming frameworks support. We develop a novel accept/reject prediction methodology to accelerate the sequential MCMC computation, thereby introducing efficient multi-core pipelining methods. We implement and validate the PROCA architecture with commercial process development kits. The implementation is evaluated based on 9 representative benchmarks, covering PyMC official tutorial probabilistic problems, single-variable probabilistic problems, and real-world causal inference problems. Our comprehensive experiments demonstrate that PROCA achieves a speedup of 172~4871 $\times$ compared to Intel Xeon Gold CPU, $42 \sim 1058 \times$ compared to NVIDIA A100 GPU, and $1.765 \times$ over state-of-the-art MCMC accelerators, respectively. PROCA achieves comparable statistical robustness to the software probabilistic programming frameworks. Compared with state-of-the-art MCMC domain-specific accelerators, our design boosts the energy efficiency by $9.47 \times$.
Yihan Fu, Anjunyi Fan, Wenshuo Yue, Hongxiao Zhao, Daijing Shi, Qiuping Wu, Yaoyu Tao, Yuchao Yang 0001, Bonan Yan
HPCA10
2024 MeMCISA: Memristor-Enabled Memory-Centric Instruction-Set Architecture for Database Workloads
abstract
The exponential growth of data exerts great pressure on hardware design for database systems. Memory-centric computing (MCC) architecture, which enable compute capabilities near or inside memory storage, demonstrate great potential in enhancing the efficiency of database operations with higher compute parallelism and reduced data movements. However, existing MCC architecture mainly focus on artificial intelligence (AI) computations and those designed for database applications can only run a limited number of standalone queries such as SORT or JOIN, lacking efficient support for increasingly diverse and complex database workloads. For example, realizing a commercial recommendation engine on database requires supporting workloads including but not limited to vector aggregation, convolution or$N$-hop neighborhoods computing, etc. In this work, we develop a memristor-enabled memory-centric instruction-set architecture (MeMCISA) aiming to efficiently accelerate versatile workloads in modern database systems. MeMCISA features scalable multi-bank memristor-based storage organization with near-memory circuitries and caches in banks. An out-of-order (O0O) scheduling scheme is designed for MeMCISA based on a vector instruction set with four types of instructions (bit-level, element-level, vector-level, and control-level), combining memristor-enabled in-memory computing and near-memory computing to efficiently run workloads with varying computational kernels and data sizes. MeMCISA can support parallel instruction executions across different memristor banks as well as different hardware modules within a memristor bank. Furthermore, we develop data dependency handling mechanisms to support vector dependency scenarios in MeMCISA that do not exist in conventional scalar-based instruction sets. A prototype MeMCISA is implemented based on a 40nm CMOS technology with necessary peripheral hardware including instruction buffer and instruction scheduler. To accurately study MeMCISA performance in real-world database systems, a software-hardware co-designed framework integrating reconfigurable MeMCISA prototype is created that can support end-to-end simulations for database workloads starting from raw software codes. Based on this framework, we evaluate MeMCISA performance with standalone database queries as well as complex database workloads from representative benchmarks including UniBench, neural collaborative filtering (NCF), and ResNet-18. Simulation results demonstrate that MeMCISA achieves up to 41.84 × ~ 1767.70 × in speed compared to general-purpose processors (CPUs/GPUs).
Yihang Zhu, Lianfeng Yu, Anjunyi Fan, Longhao Yan, Zhaokun Jing, Bonan Yan, Pek Jun Tiw, Yaoyu Tao, Yuchao Yang 0001
MICRO11
2024 Tuning the ferroelectricity of Hf0.5Zr0.5O2 with alloy electrodes
Keqin Liu, Bingjie Dang, Jinxuan Bai, Zelun Pan, Ru Huang 0001, Yuchao Yang 0001
Sci. China Inf. Sci.9
2024 Probabilistic Compute-in-Memory Design for Efficient Markov Chain Monte Carlo Sampling
abstract
Markov chain Monte Carlo (MCMC) is a widely used sampling method in modern artificial intelligence and probabilistic computing systems. It involves repetitive random number generations and thus often dominates the latency of probabilistic model computing. Hence, we propose a compute-in-memory (CIM) based MCMC design as a hardware acceleration solution. This work investigates SRAM bitcell stochasticity and proposes a novel “pseudo-read” operation, based on which we offer a block-wise random number generation circuit scheme for fast random number generation. Moreover, this work proposes a novel multi-stage exclusive-OR gate (MSXOR) design method to generate strictly uniformly distributed random numbers. The probability error deviating from a uniform distribution is suppressed under$10^{-6}$. Also, this work presents a novel in-memory copy circuit scheme to realize data copy inside a CIM sub-array, significantly reducing the use of R/W circuits for power saving. Evaluated in a commercial 28-nm process development kit, this CIM-based MCMC design generates 4-bit$\sim$32-bit samples with an energy efficiency of 0.53 pJ/sample and high throughput of up to 1066.7M samples/s. Compared to conventional processors, the overall energy efficiency improves$2.12\times10^{9}$to$9.58\times10^{9}$times.
Yihan Fu, Daijing Shi, Anjunyi Fan, Wenshuo Yue, Yuchao Yang 0001, Ru Huang 0001, Bonan Yan
IEEE Trans. Circuits Syst. I Regul. Pap.5
2023 Accelerating Neural-ODE Inference on FPGAs with Two-Stage Structured Pruning and History-based Stepsize Search
abstract
Neural ordinary differential equation (Neural-ODE) outperforms conventional deep neural networks (DNNs) in modeling continuous-time or dynamical systems by adopting numerical ODE integration onto a shallow embedded NN. However, Neural-ODE suffers from slow inference due to the costly iterative stepsize search in numerical integration, especially when using higher-order Runge-Kutta (RK) methods and smaller error tolerance for improved integration accuracy. In this work, we first present algorithmic techniques to speedup RK-based Neural-ODE inference: a two-stage coarse-grained/fine-grained structured pruning method based on top-K sparsification that reduces the overall computations by more than 60% in the embedded NN and a history-based stepsize search method based on past integration steps that reduces the latency for reaching accepted stepsize by up to 77% in RK methods. A reconfigurable hardware architecture is co-designed based on proposed speedup techniques, featuring three processing loops to support programmable embedded NN and a variety of higher-order RK methods. Sparse activation processor with multi-dimensional sorters is designed to exploit structured sparsity in activations. Implemented on a Xilinx Virtex-7 XC7VX690T FPGA and experimented on a variety of datasets, the prototype accelerator using a more complex 3rd-order RK method achieves more than 2.6x speedup compared to the latest Neural-ODE FPGA accelerator using the simplest Euler method. Compared to a software execution on Nvidia A100 GPU, the inference speedup can be up to 18x.
Jing Wang 0172, Lianfeng Yu, Bonan Yan, Yaoyu Tao, Yuchao Yang 0001
FPGA6
2023 Memristive dynamics enabled neuromorphic computing systems
Bonan Yan, Yuchao Yang 0001, Ru Huang 0001
Sci. China Inf. Sci.2
2023 COPPER: a combinatorial optimization problem solver with processing-in-memory architecture
abstract
The combinatorial optimization problem (COP), which aims to find the optimal solution in discrete space, is fundamental in various fields. Unfortunately, many COPs are NP-complete, and require much more time to solve as the problem scale increases. Troubled by this, researchers may prefer fast methods even if they are not exact, so approximation algorithms, heuristic algorithms, and machine learning have been proposed. Some works proposed chaotic simulated annealing (CSA) based on the Hopfield neural network and did a good job. However, CSA is not something that current general-purpose processors can handle easily, and there is no special hardware for it. To efficiently perform CSA, we propose a software and hardware co-design. In software, we quantize the weight and output using appropriate bit widths, and then modify the calculations that are not suitable for hardware implementation. In hardware, we design a specialized processing-in-memory hardware architecture named COPPER based on the memristor. COPPER is capable of efficiently running the modified quantized CSA algorithm and supporting the pipeline further acceleration. The results show that COPPER can perform CSA remarkably well in both speed and energy.
Bingzhe Wu, Wei Hu 0003, Guangyu Sun 0003, Yuchao Yang 0001
Frontiers Inf. Technol. Electron. Eng.7
2022 Fast and Scalable Memristive In-Memory Sorting with Column-Skipping Algorithm
abstract
Memristive in-memory sorting has been proposed recently to improve hardware sorting efficiency. Using iterative in-memory min computations, data movements between memory and external processing units can be eliminated for improved latency and energy efficiency. However, the bit-traversal algorithm to search the min requires a large number of column reads on memristive memory. In this work, we propose a column-skipping algorithm with help of a near-memory circuit. Redundant column reads can be skipped based on recorded states for improved latency and hardware efficiency. To enhance the scalability, we develop a multi-bank management that enables column-skipping for dataset stored in different memristive memory banks. Prototype column-skipping sorters are implemented with a 1T1R memristive memory in 40nm CMOS technology. Experimented on a variety of sorting datasets, the length-1024 32-bit column-skipping sorter with state recording of 2 demonstrates up to 4.08× speedup, 3.14× area efficiency and 3.39× energy efficiency, respectively, over the latest memristive in-memory sorting.
Lianfeng Yu, Zhaokun Jing, Yuchao Yang 0001, Yaoyu Tao
ISCAS3
2022 VSDCA: A Voltage Sensing Differential Column Architecture Based on 1T2R RRAM Array for Computing-in-Memory Accelerators
abstract
Non-volatile memory (NVM) such as RRAM and PCM has become the key component in high energy efficiency computing-in-memory (CIM) architectures. However, the computing accuracy and energy efficiency improvement of conventional 1T1R RRAM array based current sensing CIM scheme is hindered by device variation and large output current. In this work, we propose a voltage sensing differential column architecture (VSDCA) based on 1T2R RRAM array for binary memory and CIM applications. The memory mode of VSDCA macro can improve$1.12\times $to$5.29\times $relative read margin compared to conventional 1T1R current sensing memory. The computing mode supports 8-bit input, 9-bit weight and 18-bit output high precision and rows fully parallel computing. The VSDCA macro design is evaluated under SMIC 40 nm technology node, the energy efficiency for the high precision CIM reaches 39.52 TOPS/W. The CIFAR10 inference accuracy of the simulated VGG16 and ResNet18 model is 85.91% and 89.32% respectively.
Zhaokun Jing, Bonan Yan, Yuchao Yang 0001, Ru Huang 0001
IEEE Trans. Circuits Syst. I Regul. Pap.3
2021 In-memory computing with emerging nonvolatile memory devices
Caidie Cheng, Pek Jun Tiw, Yimao Cai, Xiaoqin Yan, Yuchao Yang 0001, Ru Huang 0001
Sci. China Inf. Sci.5
2021 NAS4RRAM: neural network architecture search for inference on RRAM-based accelerators
Zhihang Yuan, Jingze Liu, Longhao Yan, Haoxiang Chen 0003, Bingzhe Wu, Yuchao Yang 0001, Guangyu Sun 0003
Sci. China Inf. Sci.7
2020 Efficient 16 Boolean logic and arithmetic based on bipolar oxide memristors
Mingyuan Ma, Liying Xu, Zhenhua Zhu 0002, Qingxi Duan, Yu Wang 0002, Ru Huang 0001, Yuchao Yang 0001
Sci. China Inf. Sci.10
2019 A General Logic Synthesis Framework for Memristor-based Logic Design
abstract
Memristor-based logic design gives an alternative solution to improve the energy efficiency of computing systems, benefiting from combining the memory with computing units. Inspired by this thought, previous work has demonstrated various memristor-based logic families with different attributes and computation patterns. Besides, some logic synthesis tools are designed for specific memristive logic implementations. However, the poor universality and the neglect of realistic constraints in memory largely restrict the utility of these logic synthesis tools. In this paper, we propose a general logic synthesis framework for memristor-based logic design, containing a universal abstract description method for memristive logic, a mapping rules generator, and a synthesis and mapping flow. The proposed logic synthesis framework is suitable for various types of existing memristor-based logic families and takes the memory status into consideration. It is also possible to handle future memristive devices and logic families by providing the universal abstraction interface. Furthermore, we also design a circuit-partitioning-based synthesis acceleration strategy to tackle with the long synthesis time problem. Experimental results show that, our framework can generate mapping results under the restriction of limited resource, while the existing synthesis tools may fail under the same restriction, and achieve comparable synthesis results with the same resource as the existing synthesis tools, which is enough for computation and storage. And the proposed acceleration scheme can achieve ~ 1000× speedup compared with the initial one.
Zhenhua Zhu 0002, Mingyuan Ma, Jialong Liu, Liying Xu, Xiaoming Chen 0003, Yuchao Yang 0001, Yu Wang 0002, Huazhong Yang
ICCAD6
2019 Investigation of NbOx-based volatile switching device with self-rectifying characteristics
Yichen Fang, Zongwei Wang 0001, Caidie Cheng, Zhizhen Yu, Yuchao Yang 0001, Yimao Cai, Ru Huang 0001
Sci. China Inf. Sci.6
2018 Integration of biocompatible organic resistive memory and photoresistor for wearable image sensing application
Yichen Fang, Zongwei Wang 0001, Yuchao Yang 0001, Jintong Xu, Yimao Cai, Ru Huang 0001
Sci. China Inf. Sci.5