Zhenyu Xu 0007

dblp:76/934-7 · DBLP profile ↗
← Back
18ranked-venue papers
11as first author
17since 2021 · last 2026
0000-0002-3635-2409ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 18 · 11 first-author · 17 since 2021Software engineering, systems software and programming languages · 2 · 2 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Efficient LLM Decoding on Ryzen AI NPUs
abstract
We propose an efficient and scalable LLM decoding framework optimized for AMD Ryzen AI NPUs, leveraging two novel techniques: FusedDQP and FlowKV. FusedDQP fuses dequantization with projection to minimize memory operations and latency, while FlowKV introduces a pipelined, bandwidth-optimized approach for KV cache access across compute tiles (CT). Together, these methods deliver substantial improvements in both speed and energy efficiency without altering model accuracy. Our solution achieves up to 14.2× speedup and 2.66× power efficiency gains compared to existing state-of-the-art (SOTA) NPU baselines, demonstrating linear scalability with CT count and robustness across LLaMA-3.1/3.2 model variants (1B, 3B, and 8B parameters). We also benchmark against CPU and iGPU on the same platform, our performance surpasses CPU and iGPU (up to 1.8x and 16.2x speedup), while delivering substantially improved energy efficiency (up to 3.63x and 11.38x for CPU and iGPU, respectively).
Zhenyu Xu 0007, Miaoxiang Yu, Jillian Cai, Qing Yang 0001, Tao Wei 0001
DATE1
2026 An Efficient Dataflow Framework for DiT-Based Image Generation
abstract
We present a general and efficient dataflow framework for mapping Diffusion Transformer (DiT)-based models onto tiled mesh accelerators. The framework explicitly orchestrates tiling, streaming, and Direct Memory Access (DMA)-driven data movement to efficiently execute attention and large matrix multiplication (MM). We use the AMD Ryzen AI Neural Processing Unit (NPU) [1] as a representative edge-class tiled-mesh dataflow platform. In our implementation, the DiT denoising stage runs on the NPU, while the text encoder and Variational Autoencoder (VAE) decoder remain on the CPU.The iterative denoising loop dominates end-to-end latency. We therefore focus our acceleration and optimization efforts on the denoising module, while keeping the text-embedding and VAE-decoder stages on the CPU. Each denoising step reuses the same transformer backbone, and the dominant operations within a step are MM, self-attention, and MLP submodules that combine MM with elementwise nonlinear operations. The NPU is organized as a two-dimensional array of compute tiles (CTs), also referred to as AIEs. Our design partitions input matrices into tiles and maps independent tile computations across multiple AIE cores to exploit spatial parallelism.We evaluate the framework on two representative DiT-based text-to-image models, FLUX.1-schnell and Z-Image-Turbo [2], [3], and achieve end-to-end generation of a high-quality image in just over one minute. Compared with the integrated GPU (iGPU), the NPU delivers comparable generation latency while being approximately 4.4× more power efficient (package). For FLUX.1-schnell, the iGPU achieves 10.5 s/step, while the NPU achieves 13.3 s/step. For Z-Image-Turbo, the iGPU achieves 10.1 s/step, while the NPU achieves 9.8 s/step, showing that the NPU can match and slightly surpass iGPU denoising latency. These results translate into substantially improved energy efficiency for on-device image generation. Although demonstrated on a specific NPU and two representative models, the proposed dataflow framework is neither hardware-specific nor model-specific. It generalizes naturally to other diffusion-family models and mesh-based dataflow accelerators. Our results highlight the potential of spatial NPUs as energy-efficient platforms for next-generation on-device generative AI.Code Availability: The source code and setup instructions are available at: https://github.com/jia1217/vigenflow.
Yazhe Zhang, Shouyu Du, Zhenyu Xu 0007, Miaoxiang Yu, Dingjiang Yan, Zhiheng Ni, Qing Yang 0001, Tao Wei 0001
FCCM3
2026 Exploring Real-Time Power Electronics Simulation on AMD AIEs
abstract
Hardware-in-the-Loop (HIL) simulation is a critical technique for validating embedded controllers in power-electronic systems such as electric vehicles and data centers, where switching frequencies can exceed 100kHz. Achieving real-time performance at these frequencies requires sub-microsecond simulation steps while maintaining sufficient numerical accuracy. Existing commercial HIL solutions predominantly rely on FPGA platforms due to their fine-grained timing control and deterministic execution. However, the emergence of spatial, dataflow-oriented accelerators raises the question of whether alternative architectures can meet these stringent requirements.
Shouyu Du, Zhenyu Xu 0007, Miaoxiang Yu, Jillian Cai, Yeonho Jeong, Tao Wei 0001
FPGA2
2026 FPGA-Based Adjoint Method Accelerator for Rapid Optical Inverse Design
abstract
Inverse design is an emerging theory in material, photonics, and other fields. The Adjoint Method (AM) is an efficient gradient-based optimization strategy commonly used in inverse design. The optimization of inverse design frameworks faces fundamental challenges in achieving computational efficiency and scalability. While previous approaches have explored both algorithmic optimization and hardware acceleration, the inherent trade-offs between memory consumption and processing throughput continue to limit performance at scale. In this paper, we propose a high-speed field-programmable gate array (FPGA)-based accelerator to enhance the performance of the AM for fast inverse design. The accelerator is designed to significantly reduce computation time and enhance scalability, enabling more efficient and rapid design iterations. The accelerator is implemented on an AMD Alveo U280 FPGA, and we compared the FPGA and graphic process unit (GPU) performance on a series of inverse design tasks. The experimental results demonstrate that the proposed FPGA accelerator offers superior performance in terms of speed, especially in smaller design sizes compared to the NVIDIA A100 GPU. This significant improvement is due, in part, to the fully pipelined architecture and the low memory bandwidth requirement of the accelerator, which efficiently handles the enormous data throughput required for these tasks.
Lianyou Lai, Zhiheng Ni, Miaoxiang Yu, Zhenyu Xu 0007, Weijian Xu
IEEE Trans. Circuits Syst. I Regul. Pap.4
2025 HEDWIG: Homomorphic Encryption Accelerator Design Using BFV-HPS With HiGh-Speed Fixed-Point Approximation
Antian Wang, Weihang Tan, Zhenyu Xu 0007, Tao Wei 0001, Caiwen Ding, Keshab K. Parhi, Yingjie Lao
FPGA3
2025 Tile-Level Pipeline for Linear Scalable Stencil Computation on AMD AI Engines
abstract
Stencil computation is an essential method, particularly useful for numerical simulations in areas like acoustics, heat transfer, and electromagnetism. Recent studies have utilized AMD AI Engines (AIEs) for stencil computations by configuring multiple AIE tiles within a Compute Unit to exploit task-level parallelism, achieving notable speedup through concurrent task execution. However, this setup suffers from suboptimal performance due to high memory bandwidth demands, resulting in underutilization of the available AIE tiles. This work introduces a Tile-Level Pipeline architecture designed for stencil computation that operates with constant memory bandwidth. This approach achieves linear scalability, where performance scales linearly with the number of AIE tiles, and ensures full utilization of all AIE tiles on the chip. We empirically demonstrate these benefits using AIEs.
Zhenyu Xu 0007, Miaoxiang Yu, Yazhe Zhang, Jillian Cai, Qing Yang 0001, Tao Wei 0001
FPGA1
2024 TwinStep Network (TwNet): a Neuron-Centric Architecture Achieving Rapid Training
abstract
Recurrent Neural Networks (RNNs) face challenges with the Back Propagation Through Time (BPTT) algorithm, leading to substantial computational and memory demands in training, especially on GPUs. Inspired by biological neural systems, we introduce the TwinStep Network (TwNet) via algorithm/architecture co-design, achieving online training via a neuron-centric design. At its core, TwinStep signifies that both the forward pass (inference) and back propagation (training) steps for each neuron happen concurrently. This approach, which more closely resembles biological neural processes, eliminates the necessity of storing the intermediate state of each neuron at each time steps, as required in BPTT. Consequently, it overcomes the limitation on the number of time steps that can be included in the BPTT training process. Uniquely, TwNet's “pipeline parallelism” facilitates serial processing and concurrent handling of multiple time steps. We implemented TwNet on FPGAs with a fully pipelined architecture. It achieves up to 885x speedup in training several popular RNN testbenches in comparison with other state-of-the-art approaches while maintaining accuracy, marking an advancement in online RNN training and potential applications.
Zhenyu Xu 0007, Miaoxiang Yu, Jillian Cai, Qing Yang 0001, Tao Wei 0001
ASAP1
2024 An FPGA-Enabled Framework for Rapid Automated Design of Photonic Integrated Circuits
abstract
This paper introduces an FPGA-enabled framework to accelerate the automated design process for Photonic Integrated Circuit (PIC) devices. PICs are foreseen as a foundation for the next-generation semiconductors. However, the complexity of PIC design presents considerable challenges. Machine Learning (ML) techniques have shown promise in the realm of PIC design. The primary hurdle, however, is the extended training duration, solely constrained by the slow electromagnetic (EM) Finite-Difference Time-Domain (FDTD) solver. We propose a fast framework with a dedicated FPGA FDTD accelerator tailor-designed to speed up the PIC simulation. Benchmarking was carried out against commercial tools, with the single-FPGA accelerator outperforming both a multicore CPU and a GPU cluster. We taped out and evaluated the PIC devices designed through the proposed framework, and the experimental outcomes aligned. This demonstrates the full design circle, showcasing that the proposed framework enabled by FPGA breaks the current bottleneck in this domain.
Zhenyu Xu 0007, Miaoxiang Yu, Jillian Cai, Saddam Gafsi, Judson Douglas Ryckman, Qing Yang 0001, Tao Wei 0001
FPGA1
2023 A Novel FPGA-Based Circuit Simulator for Accelerating Reinforcement Learning-Based Design of Power Converters
abstract
High-efficiency energy conversion systems have become increasingly important due to their wide use in all electronic systems such as data centers, smart mobile devices, E-vehicles, medical instruments, and so forth. Complex and interdependent parameters make optimal designs of power converters challenging to get. Recent research has shown that machine learning (ML) algorithms, such as reinforcement learning (RL), show great promise in design of such converter circuits. A trained RL agent can search for optimal design parameters for power conversion circuit topologies under targeted application requirements. Training an RL agent requires numerous circuit simulations. It requires significantly more training iterations when the tolerance of circuit components due to manufacturing inconsistency, aging, and temperature variation is considered. As a result, they may take days to complete, primarily because of the slow time-domain circuit simulation. This paper proposes a new FPGA architecture that accelerates the circuit simulation and hence substantially speeds up the RL-based design method for power converters. Our new architecture supports all power electronic circuit converters and their variations. It substantially improves the training speed of RL-based design methods. High-level synthesis (HLS) was used to build the accelerator on Amazon Web Service (AWS) F1 instance. An AWS virtual PC hosts the training algorithm. The host interacts with the FPGA accelerator by updating the circuit parameters, initiating simulation, and collecting the simulation results during training iterations. A script was created on the host side to facilitate this design method to convert a netlist containing circuit topology and parameters into core matrices in the FPGA accelerator. Experimental results showed$\mathbf{60}\times$overall speedup of our RL-based design method in comparison with using a popular commercial simulator, PowerSim.
Zhenyu Xu 0007, Miaoxiang Yu, Jillian Cai, Qing Yang 0001, Yeonho Jeong, Tao Wei 0001
ASAP1
2023 A Heterogeneous Computer Architecture Accelerating Reinforcement Learning-based Design for Silicon Photonic Devices
abstract
This paper proposes a framework to substantially accelerate Reinforcement Learning (RL)-based design method for Photonic Integrated Circuit (PIC) devices. PICs are widely anticipated to underpin the forthcoming generation of semiconductor chips. However, the complexity of PIC design, which includes hundreds of degrees of freedom (DOF), presents considerable challenges. Machine Learning (ML) techniques, inclusive of RL, have demonstrated their effectiveness in the domain of PIC design. The primary hurdle, however, is the extended training duration, primarily constrained by the sluggish electromagnetic (EM) solver, specifically, the Finite-Difference Time-Domain (FDTD) solver. We have engineered a novel computational architecture that can be deployed on cloud-based systems using a cluster of Central Processing Units (CPUs), Field-Programmable Gate Arrays (FPGAs), and Graphics Processing Units (GPUs). An FPGA-FDTD accelerator, which capitalizes on the high memory bandwidth of on-chip memory (OCM), has been specifically designed to simulate planar PIC devices. Each FPGA-FDTD accelerator, also denoted as an FPGA kernel, functions as an autonomous RL environment. The host machine, in conjunction with the FPGA kernel, is designated as a worker node within the cluster. A functional prototype has been successfully implemented on the Amazon Web Services (AWS) cloud. The framework has effectively designed numerous PIC devices, and experimental results show the architecture outperforms existing methods significantly in design speed while maintaining or exceeding their design quality. Notably, the framework's versatility requires minimal adjustments for a broad range of devices, significantly reducing design time and promising to expedite PIC innovation.
Miaoxiang Yu, Zhenyu Xu 0007, Jillian Cai, Qing Yang 0001, Tao Wei 0001
ASAP2
2023 A Finite-Difference Time-Domain (FDTD) solver with linearly scalable performance in an FPGA cluster
abstract
This paper presents an FPGA cluster-based Finite-Difference Time-Domain (FDTD) accelerator that offers a linear speedup with the number of FPGAs participating in computation within the cluster. FDTD is a numeric method for simulating electromagnetic wave propagation and interactions with diverse materials and structures. Recent advancements in machine learning-based design and optimization techniques for photonic integrated circuits and microwave circuits, known as inverse design, have demonstrated remarkable success. Inverse design necessitates numerous FDTD simulations, and the high-performance FDTD accelerator enables rapid design automation, which is crucial for accelerating innovation. Our proposed accelerator comprises deeply pipelined FDTD cell update kernels that can traverse multiple FPGAs via high-speed optical links, effectively utilizing available resources across all FPGAs in a cluster. The architecture includes a head node and a flexible number of cascaded server nodes, together with custom cross-FPGA data routing kernels integrated into the "Open Cloud Testbed" (OCT) FPGA infrastructure to facilitate seamless data transfer. The proposed accelerator is developed on an existing platform, OCT FPGA. Our experiments reveal that, for a 4096×4096 2.5D FDTD simulation, each server node (Xilinx Alveo U280) can achieve 86.4 Giga-cells updates per second (GCUPS), and the head node can achieve 38.4 GCUPS. The overall speed with 4 server nodes is 38.4 + 4×86.4 = 384 GCUPS.
Zhenyu Xu 0007, Miaoxiang Yu, Jillian Cai, Qing Yang 0001, Tao Wei 0001
CLUSTER1
2023 A Novel FPGA Simulator Accelerating Reinforcement Learning-Based Design of Power Converters
abstract
High-efficiency energy conversion systems have become increasingly important due to their wide use in all electronic systems such as data centers, smart mobile devices, E-vehicles, medical instruments, and so forth. Complex and interdependent parameters make optimal designs of power converters challenging to get. Recent research has shown that reinforcement learning (RL) shows great promise in the design of such converter circuits. A trained RL agent can search for optimal design parameters for power conversion circuit topologies under targeted application requirements. Training an RL agent requires numerous circuit simulations. As a result, they may take days to complete, primarily because of the slow time-domain circuit simulation.
Zhenyu Xu 0007, Miaoxiang Yu, Qing Yang 0001, Yeonho Jeong, Tao Wei 0001
FPGA1
2022 Software defined optical time-domain reflectometer
abstract
The rapid growth of high-speed transceiver technology has met the demand for higher data rates in modern society. The field of optical communication has taken advantage of this growth by using these transceivers with small form-factor pluggable modules. This work demonstrates that these readily available, widely deployed, and commercialized modules can be converted into a state-of-the art software defined optical time-domain reflectometer (SD-OTDR). Enabled by the reconfigurable computing resource, the SD-OTDR can realize in-situ diagnostics of optical fiber without adding any overhead to existing systems. This software defined reflectometer can obtain sub-cm spatial resolutions with a measurement sensitivity on the order of -65dB. This is made possible by using the advantages offered by the reconfigurable fabric in modern System-on-Chip platforms often found in communication networks.
Thomas Mauldin, Zhenyu Xu 0007, Tao Wei 0001
FCCM2
2022 Highly Scalable Runtime Countermeasure Against Microprobing Attacks on Die-to-Die Interconnections in System-in-Package
abstract
The emerging System-in-Package (SiP) technology has enabled multiple dies fabricated on a single chip for high performance and energy efficiency. Die-to-die (D2D) communication in SiP is typically unencrypted, exposing sensitive data to possible microprobing attacks. In this paper, we propose an on-chip microprobe detection circuit together with a noise canceling technique to protect D2D buses for future SiP security. The proposed method utilizes the metastable state of a flip-flop to detect the small timing variation caused by the inevitable loading effect of a microprobe. This design requires minimum digital resources with high scalability. Uniquely, the proposed design protects D2D buses at runtime without interfering with normal data transfers. In addition, it introduces zero latency to the communication channel. We built the detection circuit in a Xilinx ZYNQ Ultrascale+ SoC to prove its feasibility. Dynamic partial reconfiguration function is employed to create the test platform and emulate D2D interconnections as well as microprobing attacks on them. To demonstrate its potential to be used in standard communication protocols, we integrated the detection circuit with a fully functional Advanced eXtensible Interface (AXI) bus. Experimental results show that the proposed runtime detection method is effective, resource-efficient, and reliable under temperature-varying environments.
Zhenyu Xu 0007, Thomas Mauldin, Qing Yang 0001, Tao Wei 0001
FPGA1
2022 A Novel Interconnection Architecture for Secured Die-to-Die Communication in System-in-Package
abstract
The emerging System-in-Package (SiP) technology has enabled multiple dies fabricated on a single chip for high performance and energy efficiency. Die-to-die (D2D) communication in SiP is typically unencrypted, exposing sensitive data to possible microprobing attacks. This paper presents a new architecture design that protects D2D interconnection network from such microprobing attacks. Our new design provides instant detection of a microprobe at run time with no interference with normal data transfers and easily scalable to hundreds of D2D interconnection buses. The uniqueness of the new architecture is extremely simple and readily applicable to any system on a chip architecture. The trick is exploiting metastable state of a flip-flop (FF) to detect the small timing variation caused by the inevitable loading effect of a microprobe. To demonstrate its effectiveness and performance, a working prototype has been built on an emulation testbench in a Xilinx ZYNQ Ultrascale+ SoC to prove its feasibility. Dynamic partial reconfiguration (DPR) function is an enabler to create the test platform. DPR emulates both unprobed and probed scenarios on a D2D bus, and allow us to switch in between as required. To show our new design can be easily integrated to standard communication protocols, we implemented our prototype on a fully functional Advanced eXtensible Interface (AXI) bus. Experimental results show that the proposed runtime detection method is effective, resource-efficient, and reliable under temperature-varying environments.
Zhenyu Xu 0007, Qing Yang 0001, Tao Wei 0001
NAS1
2021 Runtime Detection of Probing/Tampering on Interconnecting Buses
abstract
It has been reported that physical probing on an off-chip bus can reveal confidential information in an electronic system. An attacker can use non-invasive and inexpensive electric probes (or interposers) to measure signals from circuit traces, such as the memory bus between the memory controller and a memory module. This paper describes a method to detect any bus probing/tampering by tracking the phase shift of output digital waveforms, induced by input impedance change at the bus transmitter (Tx). A low-overhead digital circuit based on flip-flop's metastability is built around the Tx using a field-programmable logic gate array (FPGA) to precisely measure the phase shift of output signals. Uniquely, the output data launched by the Tx is used as a stimulus signal, thus, the proposed method holds the advantage of detecting probing attacks at run-time. That is, the detection action operates in parallel with the normal data transfer on a bus without any interference, imposing zero latency to the communication channel. In order to show its feasibility in a real-world communication protocol, we implemented the proposed method in the DDR memory controller on an FPGA board (Xilinx ZCU104). The working prototype is able to protect a memory bus between the FPGA board and a DDR4 DIMM with a data rate of 2400MT/s. Experimental results show that the proposed method can be used to countermeasure interposer attacks, probing attacks, and cold boot attacks. We believe that the proposed method can be implemented in a variety of communication channels.
Zhenyu Xu 0007, Thomas Mauldin, Qing Yang 0001, Tao Wei 0001
FCCM1
2021 Minimal Overhead Optical Time-Domain Reflectometer Via I/O Integrated Data Converter Enabled by Field Programmable Voltage Offset
Thomas Mauldin, Zhenyu Xu 0007, Tao Wei 0001
FPL2
2020 A Bus Authentication and Anti-Probing Architecture Extending Hardware Trusted Computing Base Off CPU Chips and Beyond
abstract
Tamper-proof hardware designs present a great challenge to computer architects. Most existing research limits hardware trusted computing base (TCB) to a CPU chip and anything off the CPU chip is vulnerable to probing and tampering. This paper introduces a new hardware design that provides strong defenses against physical attacks on interconnecting buses between chips in a computer system thereby extending the hardware TCB beyond CPU chips. The new approach is referred to as DIVOT: Detecting Impedance Variations Of Transmission-lines (Tx-lines). Every Tx-line in a computer system, such as a bus and interconnection wire has a unique, intrinsic, and fingerprint-like property: Impedance Inhomogeneity Pattern (IIP), i.e. the impedance distribution over distance. Such unpredictable, uncontrollable, and non-reproducible IIP fingerprints can be used to authenticate a Tx-line to ensure the confidentiality and integrity of data being transmitted. In addition, physical probes perturb the electromagnetic (EM) field around a Tx-line, leading to an altered IIP. As a result, runtime monitoring of IIPs can also be used to actively detect physical probing, snooping, and wire-tapping on buses. While the physics behind the IIP is known, the major technical breakthrough of DIVOT is the new integrated time domain reflectometer, iTDR, that is capable of carrying out in-situ and runtime monitoring of a Tx-line without interfering with normal data transfers. The iTDR is based on two innovations: analog-to-probability conversion (APC) and probability density modulation (PDM). The iTDR performs runtime IIP measurements noninvasively and is CMOS-compatible allowing it to be integrated with any interface logic connected to a bus. DIVOT is a generic, scalable, cost-effective, and low-overhead security solution for any computer system from servers to embedded computers in smart mobile devices and IoTs. To demonstrate the proposed architecture, a working prototype of DIVOT has been built on an FPGA as a proof of concept. Experimental results clearly showed the feasibility and performance of DIVOT for both hardware authentication and tamperproof applications. More specifically, the probability of correctly identifying a bus is close to 1 with an equal error rate (EER) of less than 0.06% at room temperature. We present an example design that incorporates DIVOT into an off-chip memory bus to protect against physical attacks including probing/snooping, tampering, and cold boot attacks.
Zhenyu Xu 0007, Thomas Mauldin, Zheyi Yao, Shuyi Pei, Tao Wei 0001, Qing Yang 0001
ISCA1