VLDB 2026 Research / reviewers in the wild / expert
Jillian Cai
dblp:358/7030
· DBLP profile ↗
8ranked-venue papers
0as first author
8since 2021 · last 2026
0009-0001-5006-223XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 8 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficient LLM Decoding on Ryzen AI NPUsabstractWe propose an efficient and scalable LLM decoding framework optimized for AMD Ryzen AI NPUs, leveraging two novel techniques: FusedDQP and FlowKV. FusedDQP fuses dequantization with projection to minimize memory operations and latency, while FlowKV introduces a pipelined, bandwidth-optimized approach for KV cache access across compute tiles (CT). Together, these methods deliver substantial improvements in both speed and energy efficiency without altering model accuracy. Our solution achieves up to 14.2× speedup and 2.66× power efficiency gains compared to existing state-of-the-art (SOTA) NPU baselines, demonstrating linear scalability with CT count and robustness across LLaMA-3.1/3.2 model variants (1B, 3B, and 8B parameters). We also benchmark against CPU and iGPU on the same platform, our performance surpasses CPU and iGPU (up to 1.8x and 16.2x speedup), while delivering substantially improved energy efficiency (up to 3.63x and 11.38x for CPU and iGPU, respectively). Zhenyu Xu 0007, Miaoxiang Yu, Jillian Cai, Qing Yang 0001, Tao Wei 0001 |
DATE | 3 |
| 2026 | Exploring Real-Time Power Electronics Simulation on AMD AIEsabstractHardware-in-the-Loop (HIL) simulation is a critical technique for validating embedded controllers in power-electronic systems such as electric vehicles and data centers, where switching frequencies can exceed 100kHz. Achieving real-time performance at these frequencies requires sub-microsecond simulation steps while maintaining sufficient numerical accuracy. Existing commercial HIL solutions predominantly rely on FPGA platforms due to their fine-grained timing control and deterministic execution. However, the emergence of spatial, dataflow-oriented accelerators raises the question of whether alternative architectures can meet these stringent requirements. Shouyu Du, Zhenyu Xu 0007, Miaoxiang Yu, Jillian Cai, Yeonho Jeong, Tao Wei 0001 |
FPGA | 4 |
| 2025 | Tile-Level Pipeline for Linear Scalable Stencil Computation on AMD AI EnginesabstractStencil computation is an essential method, particularly useful for numerical simulations in areas like acoustics, heat transfer, and electromagnetism. Recent studies have utilized AMD AI Engines (AIEs) for stencil computations by configuring multiple AIE tiles within a Compute Unit to exploit task-level parallelism, achieving notable speedup through concurrent task execution. However, this setup suffers from suboptimal performance due to high memory bandwidth demands, resulting in underutilization of the available AIE tiles. This work introduces a Tile-Level Pipeline architecture designed for stencil computation that operates with constant memory bandwidth. This approach achieves linear scalability, where performance scales linearly with the number of AIE tiles, and ensures full utilization of all AIE tiles on the chip. We empirically demonstrate these benefits using AIEs. Zhenyu Xu 0007, Miaoxiang Yu, Yazhe Zhang, Jillian Cai, Qing Yang 0001, Tao Wei 0001 |
FPGA | 4 |
| 2024 | TwinStep Network (TwNet): a Neuron-Centric Architecture Achieving Rapid TrainingabstractRecurrent Neural Networks (RNNs) face challenges with the Back Propagation Through Time (BPTT) algorithm, leading to substantial computational and memory demands in training, especially on GPUs. Inspired by biological neural systems, we introduce the TwinStep Network (TwNet) via algorithm/architecture co-design, achieving online training via a neuron-centric design. At its core, TwinStep signifies that both the forward pass (inference) and back propagation (training) steps for each neuron happen concurrently. This approach, which more closely resembles biological neural processes, eliminates the necessity of storing the intermediate state of each neuron at each time steps, as required in BPTT. Consequently, it overcomes the limitation on the number of time steps that can be included in the BPTT training process. Uniquely, TwNet's “pipeline parallelism” facilitates serial processing and concurrent handling of multiple time steps. We implemented TwNet on FPGAs with a fully pipelined architecture. It achieves up to 885x speedup in training several popular RNN testbenches in comparison with other state-of-the-art approaches while maintaining accuracy, marking an advancement in online RNN training and potential applications. Zhenyu Xu 0007, Miaoxiang Yu, Jillian Cai, Qing Yang 0001, Tao Wei 0001 |
ASAP | 3 |
| 2024 | An FPGA-Enabled Framework for Rapid Automated Design of Photonic Integrated CircuitsabstractThis paper introduces an FPGA-enabled framework to accelerate the automated design process for Photonic Integrated Circuit (PIC) devices. PICs are foreseen as a foundation for the next-generation semiconductors. However, the complexity of PIC design presents considerable challenges. Machine Learning (ML) techniques have shown promise in the realm of PIC design. The primary hurdle, however, is the extended training duration, solely constrained by the slow electromagnetic (EM) Finite-Difference Time-Domain (FDTD) solver. We propose a fast framework with a dedicated FPGA FDTD accelerator tailor-designed to speed up the PIC simulation. Benchmarking was carried out against commercial tools, with the single-FPGA accelerator outperforming both a multicore CPU and a GPU cluster. We taped out and evaluated the PIC devices designed through the proposed framework, and the experimental outcomes aligned. This demonstrates the full design circle, showcasing that the proposed framework enabled by FPGA breaks the current bottleneck in this domain. Zhenyu Xu 0007, Miaoxiang Yu, Jillian Cai, Saddam Gafsi, Judson Douglas Ryckman, Qing Yang 0001, Tao Wei 0001 |
FPGA | 3 |
| 2023 | A Novel FPGA-Based Circuit Simulator for Accelerating Reinforcement Learning-Based Design of Power ConvertersabstractHigh-efficiency energy conversion systems have become increasingly important due to their wide use in all electronic systems such as data centers, smart mobile devices, E-vehicles, medical instruments, and so forth. Complex and interdependent parameters make optimal designs of power converters challenging to get. Recent research has shown that machine learning (ML) algorithms, such as reinforcement learning (RL), show great promise in design of such converter circuits. A trained RL agent can search for optimal design parameters for power conversion circuit topologies under targeted application requirements. Training an RL agent requires numerous circuit simulations. It requires significantly more training iterations when the tolerance of circuit components due to manufacturing inconsistency, aging, and temperature variation is considered. As a result, they may take days to complete, primarily because of the slow time-domain circuit simulation. This paper proposes a new FPGA architecture that accelerates the circuit simulation and hence substantially speeds up the RL-based design method for power converters. Our new architecture supports all power electronic circuit converters and their variations. It substantially improves the training speed of RL-based design methods. High-level synthesis (HLS) was used to build the accelerator on Amazon Web Service (AWS) F1 instance. An AWS virtual PC hosts the training algorithm. The host interacts with the FPGA accelerator by updating the circuit parameters, initiating simulation, and collecting the simulation results during training iterations. A script was created on the host side to facilitate this design method to convert a netlist containing circuit topology and parameters into core matrices in the FPGA accelerator. Experimental results showed$\mathbf{60}\times$overall speedup of our RL-based design method in comparison with using a popular commercial simulator, PowerSim. Zhenyu Xu 0007, Miaoxiang Yu, Jillian Cai, Qing Yang 0001, Yeonho Jeong, Tao Wei 0001 |
ASAP | 3 |
| 2023 | A Heterogeneous Computer Architecture Accelerating Reinforcement Learning-based Design for Silicon Photonic DevicesabstractThis paper proposes a framework to substantially accelerate Reinforcement Learning (RL)-based design method for Photonic Integrated Circuit (PIC) devices. PICs are widely anticipated to underpin the forthcoming generation of semiconductor chips. However, the complexity of PIC design, which includes hundreds of degrees of freedom (DOF), presents considerable challenges. Machine Learning (ML) techniques, inclusive of RL, have demonstrated their effectiveness in the domain of PIC design. The primary hurdle, however, is the extended training duration, primarily constrained by the sluggish electromagnetic (EM) solver, specifically, the Finite-Difference Time-Domain (FDTD) solver. We have engineered a novel computational architecture that can be deployed on cloud-based systems using a cluster of Central Processing Units (CPUs), Field-Programmable Gate Arrays (FPGAs), and Graphics Processing Units (GPUs). An FPGA-FDTD accelerator, which capitalizes on the high memory bandwidth of on-chip memory (OCM), has been specifically designed to simulate planar PIC devices. Each FPGA-FDTD accelerator, also denoted as an FPGA kernel, functions as an autonomous RL environment. The host machine, in conjunction with the FPGA kernel, is designated as a worker node within the cluster. A functional prototype has been successfully implemented on the Amazon Web Services (AWS) cloud. The framework has effectively designed numerous PIC devices, and experimental results show the architecture outperforms existing methods significantly in design speed while maintaining or exceeding their design quality. Notably, the framework's versatility requires minimal adjustments for a broad range of devices, significantly reducing design time and promising to expedite PIC innovation. Miaoxiang Yu, Zhenyu Xu 0007, Jillian Cai, Qing Yang 0001, Tao Wei 0001 |
ASAP | 3 |
| 2023 | A Finite-Difference Time-Domain (FDTD) solver with linearly scalable performance in an FPGA clusterabstractThis paper presents an FPGA cluster-based Finite-Difference Time-Domain (FDTD) accelerator that offers a linear speedup with the number of FPGAs participating in computation within the cluster. FDTD is a numeric method for simulating electromagnetic wave propagation and interactions with diverse materials and structures. Recent advancements in machine learning-based design and optimization techniques for photonic integrated circuits and microwave circuits, known as inverse design, have demonstrated remarkable success. Inverse design necessitates numerous FDTD simulations, and the high-performance FDTD accelerator enables rapid design automation, which is crucial for accelerating innovation. Our proposed accelerator comprises deeply pipelined FDTD cell update kernels that can traverse multiple FPGAs via high-speed optical links, effectively utilizing available resources across all FPGAs in a cluster. The architecture includes a head node and a flexible number of cascaded server nodes, together with custom cross-FPGA data routing kernels integrated into the "Open Cloud Testbed" (OCT) FPGA infrastructure to facilitate seamless data transfer. The proposed accelerator is developed on an existing platform, OCT FPGA. Our experiments reveal that, for a 4096×4096 2.5D FDTD simulation, each server node (Xilinx Alveo U280) can achieve 86.4 Giga-cells updates per second (GCUPS), and the head node can achieve 38.4 GCUPS. The overall speed with 4 server nodes is 38.4 + 4×86.4 = 384 GCUPS. Zhenyu Xu 0007, Miaoxiang Yu, Jillian Cai, Qing Yang 0001, Tao Wei 0001 |
CLUSTER | 3 |