VLDB 2026 Research / reviewers in the wild / expert
Miaoxiang Yu
dblp:339/8985
· DBLP profile ↗
11ranked-venue papers
1as first author
11since 2021 · last 2026
0000-0002-4382-9009ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 11 · 1 first-author · 11 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficient LLM Decoding on Ryzen AI NPUsabstractWe propose an efficient and scalable LLM decoding framework optimized for AMD Ryzen AI NPUs, leveraging two novel techniques: FusedDQP and FlowKV. FusedDQP fuses dequantization with projection to minimize memory operations and latency, while FlowKV introduces a pipelined, bandwidth-optimized approach for KV cache access across compute tiles (CT). Together, these methods deliver substantial improvements in both speed and energy efficiency without altering model accuracy. Our solution achieves up to 14.2× speedup and 2.66× power efficiency gains compared to existing state-of-the-art (SOTA) NPU baselines, demonstrating linear scalability with CT count and robustness across LLaMA-3.1/3.2 model variants (1B, 3B, and 8B parameters). We also benchmark against CPU and iGPU on the same platform, our performance surpasses CPU and iGPU (up to 1.8x and 16.2x speedup), while delivering substantially improved energy efficiency (up to 3.63x and 11.38x for CPU and iGPU, respectively). Zhenyu Xu 0007, Miaoxiang Yu, Jillian Cai, Qing Yang 0001, Tao Wei 0001 |
DATE | 2 |
| 2026 | An Efficient Dataflow Framework for DiT-Based Image GenerationabstractWe present a general and efficient dataflow framework for mapping Diffusion Transformer (DiT)-based models onto tiled mesh accelerators. The framework explicitly orchestrates tiling, streaming, and Direct Memory Access (DMA)-driven data movement to efficiently execute attention and large matrix multiplication (MM). We use the AMD Ryzen AI Neural Processing Unit (NPU) [1] as a representative edge-class tiled-mesh dataflow platform. In our implementation, the DiT denoising stage runs on the NPU, while the text encoder and Variational Autoencoder (VAE) decoder remain on the CPU.The iterative denoising loop dominates end-to-end latency. We therefore focus our acceleration and optimization efforts on the denoising module, while keeping the text-embedding and VAE-decoder stages on the CPU. Each denoising step reuses the same transformer backbone, and the dominant operations within a step are MM, self-attention, and MLP submodules that combine MM with elementwise nonlinear operations. The NPU is organized as a two-dimensional array of compute tiles (CTs), also referred to as AIEs. Our design partitions input matrices into tiles and maps independent tile computations across multiple AIE cores to exploit spatial parallelism.We evaluate the framework on two representative DiT-based text-to-image models, FLUX.1-schnell and Z-Image-Turbo [2], [3], and achieve end-to-end generation of a high-quality image in just over one minute. Compared with the integrated GPU (iGPU), the NPU delivers comparable generation latency while being approximately 4.4× more power efficient (package). For FLUX.1-schnell, the iGPU achieves 10.5 s/step, while the NPU achieves 13.3 s/step. For Z-Image-Turbo, the iGPU achieves 10.1 s/step, while the NPU achieves 9.8 s/step, showing that the NPU can match and slightly surpass iGPU denoising latency. These results translate into substantially improved energy efficiency for on-device image generation. Although demonstrated on a specific NPU and two representative models, the proposed dataflow framework is neither hardware-specific nor model-specific. It generalizes naturally to other diffusion-family models and mesh-based dataflow accelerators. Our results highlight the potential of spatial NPUs as energy-efficient platforms for next-generation on-device generative AI.Code Availability: The source code and setup instructions are available at: https://github.com/jia1217/vigenflow. Yazhe Zhang, Shouyu Du, Zhenyu Xu 0007, Miaoxiang Yu, Dingjiang Yan, Zhiheng Ni, Qing Yang 0001, Tao Wei 0001 |
FCCM | 4 |
| 2026 | Exploring Real-Time Power Electronics Simulation on AMD AIEsabstractHardware-in-the-Loop (HIL) simulation is a critical technique for validating embedded controllers in power-electronic systems such as electric vehicles and data centers, where switching frequencies can exceed 100kHz. Achieving real-time performance at these frequencies requires sub-microsecond simulation steps while maintaining sufficient numerical accuracy. Existing commercial HIL solutions predominantly rely on FPGA platforms due to their fine-grained timing control and deterministic execution. However, the emergence of spatial, dataflow-oriented accelerators raises the question of whether alternative architectures can meet these stringent requirements. Shouyu Du, Zhenyu Xu 0007, Miaoxiang Yu, Jillian Cai, Yeonho Jeong, Tao Wei 0001 |
FPGA | 3 |
| 2026 | FPGA-Based Adjoint Method Accelerator for Rapid Optical Inverse DesignabstractInverse design is an emerging theory in material, photonics, and other fields. The Adjoint Method (AM) is an efficient gradient-based optimization strategy commonly used in inverse design. The optimization of inverse design frameworks faces fundamental challenges in achieving computational efficiency and scalability. While previous approaches have explored both algorithmic optimization and hardware acceleration, the inherent trade-offs between memory consumption and processing throughput continue to limit performance at scale. In this paper, we propose a high-speed field-programmable gate array (FPGA)-based accelerator to enhance the performance of the AM for fast inverse design. The accelerator is designed to significantly reduce computation time and enhance scalability, enabling more efficient and rapid design iterations. The accelerator is implemented on an AMD Alveo U280 FPGA, and we compared the FPGA and graphic process unit (GPU) performance on a series of inverse design tasks. The experimental results demonstrate that the proposed FPGA accelerator offers superior performance in terms of speed, especially in smaller design sizes compared to the NVIDIA A100 GPU. This significant improvement is due, in part, to the fully pipelined architecture and the low memory bandwidth requirement of the accelerator, which efficiently handles the enormous data throughput required for these tasks. Lianyou Lai, Zhiheng Ni, Miaoxiang Yu, Zhenyu Xu 0007, Weijian Xu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2025 | Tile-Level Pipeline for Linear Scalable Stencil Computation on AMD AI EnginesabstractStencil computation is an essential method, particularly useful for numerical simulations in areas like acoustics, heat transfer, and electromagnetism. Recent studies have utilized AMD AI Engines (AIEs) for stencil computations by configuring multiple AIE tiles within a Compute Unit to exploit task-level parallelism, achieving notable speedup through concurrent task execution. However, this setup suffers from suboptimal performance due to high memory bandwidth demands, resulting in underutilization of the available AIE tiles. This work introduces a Tile-Level Pipeline architecture designed for stencil computation that operates with constant memory bandwidth. This approach achieves linear scalability, where performance scales linearly with the number of AIE tiles, and ensures full utilization of all AIE tiles on the chip. We empirically demonstrate these benefits using AIEs. Zhenyu Xu 0007, Miaoxiang Yu, Yazhe Zhang, Jillian Cai, Qing Yang 0001, Tao Wei 0001 |
FPGA | 2 |
| 2024 | TwinStep Network (TwNet): a Neuron-Centric Architecture Achieving Rapid TrainingabstractRecurrent Neural Networks (RNNs) face challenges with the Back Propagation Through Time (BPTT) algorithm, leading to substantial computational and memory demands in training, especially on GPUs. Inspired by biological neural systems, we introduce the TwinStep Network (TwNet) via algorithm/architecture co-design, achieving online training via a neuron-centric design. At its core, TwinStep signifies that both the forward pass (inference) and back propagation (training) steps for each neuron happen concurrently. This approach, which more closely resembles biological neural processes, eliminates the necessity of storing the intermediate state of each neuron at each time steps, as required in BPTT. Consequently, it overcomes the limitation on the number of time steps that can be included in the BPTT training process. Uniquely, TwNet's “pipeline parallelism” facilitates serial processing and concurrent handling of multiple time steps. We implemented TwNet on FPGAs with a fully pipelined architecture. It achieves up to 885x speedup in training several popular RNN testbenches in comparison with other state-of-the-art approaches while maintaining accuracy, marking an advancement in online RNN training and potential applications. Zhenyu Xu 0007, Miaoxiang Yu, Jillian Cai, Qing Yang 0001, Tao Wei 0001 |
ASAP | 2 |
| 2024 | An FPGA-Enabled Framework for Rapid Automated Design of Photonic Integrated CircuitsabstractThis paper introduces an FPGA-enabled framework to accelerate the automated design process for Photonic Integrated Circuit (PIC) devices. PICs are foreseen as a foundation for the next-generation semiconductors. However, the complexity of PIC design presents considerable challenges. Machine Learning (ML) techniques have shown promise in the realm of PIC design. The primary hurdle, however, is the extended training duration, solely constrained by the slow electromagnetic (EM) Finite-Difference Time-Domain (FDTD) solver. We propose a fast framework with a dedicated FPGA FDTD accelerator tailor-designed to speed up the PIC simulation. Benchmarking was carried out against commercial tools, with the single-FPGA accelerator outperforming both a multicore CPU and a GPU cluster. We taped out and evaluated the PIC devices designed through the proposed framework, and the experimental outcomes aligned. This demonstrates the full design circle, showcasing that the proposed framework enabled by FPGA breaks the current bottleneck in this domain. Zhenyu Xu 0007, Miaoxiang Yu, Jillian Cai, Saddam Gafsi, Judson Douglas Ryckman, Qing Yang 0001, Tao Wei 0001 |
FPGA | 2 |
| 2023 | A Novel FPGA-Based Circuit Simulator for Accelerating Reinforcement Learning-Based Design of Power ConvertersabstractHigh-efficiency energy conversion systems have become increasingly important due to their wide use in all electronic systems such as data centers, smart mobile devices, E-vehicles, medical instruments, and so forth. Complex and interdependent parameters make optimal designs of power converters challenging to get. Recent research has shown that machine learning (ML) algorithms, such as reinforcement learning (RL), show great promise in design of such converter circuits. A trained RL agent can search for optimal design parameters for power conversion circuit topologies under targeted application requirements. Training an RL agent requires numerous circuit simulations. It requires significantly more training iterations when the tolerance of circuit components due to manufacturing inconsistency, aging, and temperature variation is considered. As a result, they may take days to complete, primarily because of the slow time-domain circuit simulation. This paper proposes a new FPGA architecture that accelerates the circuit simulation and hence substantially speeds up the RL-based design method for power converters. Our new architecture supports all power electronic circuit converters and their variations. It substantially improves the training speed of RL-based design methods. High-level synthesis (HLS) was used to build the accelerator on Amazon Web Service (AWS) F1 instance. An AWS virtual PC hosts the training algorithm. The host interacts with the FPGA accelerator by updating the circuit parameters, initiating simulation, and collecting the simulation results during training iterations. A script was created on the host side to facilitate this design method to convert a netlist containing circuit topology and parameters into core matrices in the FPGA accelerator. Experimental results showed$\mathbf{60}\times$overall speedup of our RL-based design method in comparison with using a popular commercial simulator, PowerSim. Zhenyu Xu 0007, Miaoxiang Yu, Jillian Cai, Qing Yang 0001, Yeonho Jeong, Tao Wei 0001 |
ASAP | 2 |
| 2023 | A Heterogeneous Computer Architecture Accelerating Reinforcement Learning-based Design for Silicon Photonic DevicesabstractThis paper proposes a framework to substantially accelerate Reinforcement Learning (RL)-based design method for Photonic Integrated Circuit (PIC) devices. PICs are widely anticipated to underpin the forthcoming generation of semiconductor chips. However, the complexity of PIC design, which includes hundreds of degrees of freedom (DOF), presents considerable challenges. Machine Learning (ML) techniques, inclusive of RL, have demonstrated their effectiveness in the domain of PIC design. The primary hurdle, however, is the extended training duration, primarily constrained by the sluggish electromagnetic (EM) solver, specifically, the Finite-Difference Time-Domain (FDTD) solver. We have engineered a novel computational architecture that can be deployed on cloud-based systems using a cluster of Central Processing Units (CPUs), Field-Programmable Gate Arrays (FPGAs), and Graphics Processing Units (GPUs). An FPGA-FDTD accelerator, which capitalizes on the high memory bandwidth of on-chip memory (OCM), has been specifically designed to simulate planar PIC devices. Each FPGA-FDTD accelerator, also denoted as an FPGA kernel, functions as an autonomous RL environment. The host machine, in conjunction with the FPGA kernel, is designated as a worker node within the cluster. A functional prototype has been successfully implemented on the Amazon Web Services (AWS) cloud. The framework has effectively designed numerous PIC devices, and experimental results show the architecture outperforms existing methods significantly in design speed while maintaining or exceeding their design quality. Notably, the framework's versatility requires minimal adjustments for a broad range of devices, significantly reducing design time and promising to expedite PIC innovation. Miaoxiang Yu, Zhenyu Xu 0007, Jillian Cai, Qing Yang 0001, Tao Wei 0001 |
ASAP | 1 |
| 2023 | A Finite-Difference Time-Domain (FDTD) solver with linearly scalable performance in an FPGA clusterabstractThis paper presents an FPGA cluster-based Finite-Difference Time-Domain (FDTD) accelerator that offers a linear speedup with the number of FPGAs participating in computation within the cluster. FDTD is a numeric method for simulating electromagnetic wave propagation and interactions with diverse materials and structures. Recent advancements in machine learning-based design and optimization techniques for photonic integrated circuits and microwave circuits, known as inverse design, have demonstrated remarkable success. Inverse design necessitates numerous FDTD simulations, and the high-performance FDTD accelerator enables rapid design automation, which is crucial for accelerating innovation. Our proposed accelerator comprises deeply pipelined FDTD cell update kernels that can traverse multiple FPGAs via high-speed optical links, effectively utilizing available resources across all FPGAs in a cluster. The architecture includes a head node and a flexible number of cascaded server nodes, together with custom cross-FPGA data routing kernels integrated into the "Open Cloud Testbed" (OCT) FPGA infrastructure to facilitate seamless data transfer. The proposed accelerator is developed on an existing platform, OCT FPGA. Our experiments reveal that, for a 4096×4096 2.5D FDTD simulation, each server node (Xilinx Alveo U280) can achieve 86.4 Giga-cells updates per second (GCUPS), and the head node can achieve 38.4 GCUPS. The overall speed with 4 server nodes is 38.4 + 4×86.4 = 384 GCUPS. Zhenyu Xu 0007, Miaoxiang Yu, Jillian Cai, Qing Yang 0001, Tao Wei 0001 |
CLUSTER | 2 |
| 2023 | A Novel FPGA Simulator Accelerating Reinforcement Learning-Based Design of Power ConvertersabstractHigh-efficiency energy conversion systems have become increasingly important due to their wide use in all electronic systems such as data centers, smart mobile devices, E-vehicles, medical instruments, and so forth. Complex and interdependent parameters make optimal designs of power converters challenging to get. Recent research has shown that reinforcement learning (RL) shows great promise in the design of such converter circuits. A trained RL agent can search for optimal design parameters for power conversion circuit topologies under targeted application requirements. Training an RL agent requires numerous circuit simulations. As a result, they may take days to complete, primarily because of the slow time-domain circuit simulation. Zhenyu Xu 0007, Miaoxiang Yu, Qing Yang 0001, Yeonho Jeong, Tao Wei 0001 |
FPGA | 2 |