EDBT 2026 Demo / reviewers in the wild / expert
Xilang Zhou
dblp:380/5578
· DBLP profile ↗
5ranked-venue papers
3as first author
5since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 5 · 3 first-author · 5 since 2021Software engineering, systems software and programming languages · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | PipeViT: Accelerating Vision Transformers via Intra-Layer PipeliningabstractVision Transformers (ViTs) have achieved high performance across various computer vision tasks by leveraging the attention mechanism. However, the attention module in ViTs severely hindered inference performance due to its low operational intensity. Existing approaches improve ViTs efficiency through pruning, sparsity, and linearization, but at the cost of fine-tuning overhead and accuracy degradation. In this paper, we propose PipeViT, a memory-efficient and low-latency accelerator for ViTs inference. The key insight of PipeViT is to exploit intra-layer acceleration opportunities. Specifically, we first fuse the attention operations into a single operator to reduce memory access overhead. Then, we divide the input of attention into multiple tiles to reduce the on-chip memory requirement. Finally, we pipeline the tiled attention computation to improve overall throughput. Based on the optimized dataflow, we design a heterogeneous dual-core architecture for efficient pipeline execution. Furthermore, to maximize hardware utilization, the architecture can be reconfigured into a single core with higher parallelism during the execution of the feed-forward network. Experimental results show that PipeViT achieves up to $19.3 \times 1.5 \times, 2.1 \times$, and $2.0 \times$ improvements in Frames Per Second (FPS) compared to state-of-the-art accelerators, including ViTA, Auto-ViT, MEViT, and HeatViT. Additionally, PipeViT achieves up to $8.0 \times$ and $2.6 \times$ higher energy efficiency compared to CPU and GPU implementations, respectively. Xilang Zhou, Yiheng Xu, Haodong Lu 0001, Jun Yu 0010, Kun Wang 0005 |
ASP-DAC | 1 |
| 2025 | AttentionLib: A Scalable Optimization Framework for Automated Attention Acceleration on FPGAabstractThe self-attention mechanism is a fundamental component within transformer-based models. Nowadays, as the length of sequences processed by large language models (LLMs) continues to increase, the attention mechanism has gradually become a bottleneck in model inference. The LLM inference process can be separated into two phases: prefill and decode. The latter contains memory-intensive attention computation, making FPGA-based accelerators an attractive solution for acceleration. However, designing accelerators tailored for the attention module poses a challenge, requiring substantial manual work. To automate this process and achieve superior acceleration performance, we propose AttentionLib, an MLIR-based framework. AttentionLib automatically performs fusion dataflow optimization for attention computations and generates high-level synthesis code in compliance with hardware constraints. Given the large design space, we provide a design space exploration (DSE) engine to automatically identify optimal fusion dataflows within the specified constraints. Experimental results show that AttentionLib is effective in generating well-suited accelerators for diverse attention computations and achieving superior performance under hardware constraints. Notably, the accelerators generated by AttentionLib exhibit at least a 25.1 × improvement compared to the baselines solely automatically optimized by Vitis HLS. Furthermore, these designs outperform GPUs in decode workloads, showcasing over a 2× speedup for short sequences. Xilang Zhou, Faxian Sun, Jianli Chen, Jun Yu 0010, Kun Wang 0005 |
DATE | 2 |
| 2025 | TaiChi: Efficient Execution for Multi-DNNs Using Graph-Based SchedulingabstractApplications constructed with multiple Deep Neural Networks (multi-DNNs) are growing rapidly in edge and data center. However, executing multi-DNNs efficiently remains chal-lenging because multi-DNNs are inherently heterogeneous. The diverse operators, dependencies and performance requirements of multi-DNNs lead to high costs of encoding and generalization. We introduce Taichi, a graph-based framework for efficiently scheduling multi-DNNs on multi-core accelerators. Specifically, Taichi consists of two phases: (1) a graph neural network (GNN) is utilized to automatically capture the features from the graph structure of multi-DNNs and (2) reinforcement learning (RL) is employed to find an optimal online scheduling strategy. Evaluation results show that TaiChi reduces latency by 1.1-2.4 x and 1.1-1.6x compared to SJF and MAGMA, and improves throughput by 26.4-63.7% and 18.6-33.7%, respectively. Moreover, TaiChi achieves an average speedup of 779 x in scheduling runtime compared to MAGMA. Xilang Zhou, Zhuoheng Wan, Jianli Chen |
DATE | 1 |
| 2024 | PipeFuser: Building Flexible Pipeline Architecture for DNN Accelerators via Layer FusionabstractIn this paper, we propose a fused-pipeline architecture that leverages the layer fusion technique to harness the strengths of both non-pipeline and full-pipeline architectures while mitigating their disadvantages. In particular, we observe that the performance of the fused-pipeline accelerators is significantly influenced by the layer fusion strategies and intra-layer mapping schemes. To optimize and rapidly employ the fused-pipeline architecture, we present an end-to-end automation framework, named PipeFuser. At the core of PipeFuser is a genetic algorithm (GA)-based co-design engine, which is used to acquire near-optimal hardware configurations in the vast design space. Experimental results demonstrate that our fused-pipeline architecture achieves 2.3 × to 3.3 × higher performance over the non-pipeline design and 1.9 × to 2.5 × speedup compared to the full-pipeline architecture, with greater deployment flexibility. Xilang Zhou, Haodong Lu 0001, Kun Wang 0005 |
ASPDAC | 1 |
| 2024 | DNNMapper: An Elastic Framework for Mapping DNNs to Multi-die FPGAsabstractDeep Neural Networks (DNNs) have stimulated intensive FPGA-based acceleration solutions, and multi-die FPGAs offer abundant resources for implementing large-scale DNN workloads. However, current FPGA frameworks overlook the opportunities of optimization on multi-die FPGAs. In this paper, we propose an automated framework named DNNMapper, for mapping DNNs to multi-die FPGAs. With careful consideration of the unique architectural characteristics and resource constraints of multi-die FPGAs, DNNMapper involves model partitioning and resource allocation as two critical processes that map DNN layers onto respective FPGA dies and efficiently allocate hardware resources. DNNMapper employs a co-design engine based on a genetic algorithm, which co-optimizes model partitioning and resource allocation. Experimental results demonstrate that accelerators generated by DNNMapper offer superior performance and scalability, achieving up to 2× higher throughput and 1.3× to 1.9× higher DSP density. Moreover, our accelerator demonstrates a frequency improvement from 1.28× to 1.69×. Xilang Zhou, Haodong Lu 0001, Kun Wang 0005 |
ISCAS | 2 |