EDBT 2026 Demo / reviewers in the wild / expert
Xuechao Wei
dblp:08/10954
· DBLP profile ↗
22ranked-venue papers
5as first author
12since 2021 · last 2025
0000-0002-0996-2260ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 21 · 5 first-author · 11 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Klotski v2: Improved DNN Model Orchestration Framework for Dataflow Architecture AcceleratorsabstractDataflow architecture accelerators are a new kind of scalable DNN accelerators. For an instruction, the availability of input operands solely determines the beginning of executions. DNN model orchestration determines how to partition, schedule, and map the computation to the underlying hardware. In this article, we propose the Klotski v2 framework to solve DNN model orchestration for dataflow architecture accelerators. First, a Bayesian optimization-based entropy-directed partition algorithm is proposed to transform a DNN model into$\mu $ops. Second, a unified formal formulation for$\mu $ops scheduling and mapping is presented. Third, a two-stage methodology is proposed to decouple the scheduling and mapping. Fourth, a Hilbert curve-based mapping heuristic is proposed to enhance problem-solving efficiency, improving the tradeoff between solution quality and algorithm runtime. Extensive results show that Klotski v2 can achieve an average of 21.57% higher execution performance improvement than the previous methodologies. With the Hilbert curve-based mapping heuristic, we improve the algorithm efficiency by an average of 63.50% across different DNN workloads. Xuechao Wei, Youwei Zhuo, Yi Cai 0003, Hongzhong Zheng, Bei Yu 0001, Yuan Xie 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2024 | PT-Map: Efficient Program Transformation Optimization for CGRA MappingabstractCoarse-Grained Reconfigurable Array (CGRA) is a parallel architecture providing high energy efficiency and spatial-temporal re-configurability. Beyond loop scheduling for throughput optimization, program transformation is also crucial in CGRA mapping to optimize overall performance and efficiency. However, existing studies on program transformation optimization face challenges in exploring the transformation space systematically and evaluating candidates efficiently, leading to sub-optimal results. To tackle these challenges, this paper introduces PT-Map, an efficient program transformation optimization framework for CGRA mapping. PT-Map defines a comprehensive transformation space and employs a CGRA-specialized top-down exploration approach. It also incorporates a bottom-up evaluation scheme using architectural parameters and a graph neural network-based predictive model. Experiments demonstrate that PT-Map achieves up to 2.95X/1.80X speedups and 59.0%/23.2% energy-delay-product (EDP) reductions over the state-of-the-art approaches MapZero and PBP, respectively. Bizhao Shi, Tuo Dai, Jiaxi Zhang 0001, Xuechao Wei, Guojie Luo |
DAC | 4 |
| 2024 | RadiK: Scalable and Optimized GPU-Parallel Radix Top-K SelectionabstractTop-k selection, which identifies the largest or smallest k elements from a data set, is a fundamental operation in data-intensive domains such as databases and deep learning, so its scalability and efficiency are critical for these high-performance systems. However, previous studies on its efficient GPU implementation are mostly merge-based and rely heavily on fast but size-limited on-chip memory, thereby limiting scalability with a restricted upper bound on k. This work introduces RadiK, a scalable and optimized GPU-parallel radix top-k selection that supports significantly larger k values than existing methods without compromising efficiency, regardless of input length and batch size. RadiK incorporates a novel optimization framework tailored for high memory bandwidth and resource utilization, achieving up to 2.5 × speedup over the prior art for non-batch queries and up to 4.8 × speedup for batch queries. In addition, we propose an adaptive scaling technique that strengthens robustness, which further provides up to 2.7 × speedup on highly adversarial input distributions. Bole Zhou, Jiejing Zhang, Xuechao Wei, Yinghan Li 0002, Yingda Chen |
ICS | 4 |
| 2024 | ArkVale: Efficient Generative LLM Inference with Recallable Key-Value EvictionabstractLarge Language Models (LLMs) are widely used in today's tasks of natural language processing.
To support applications like multi-turn chats, document understanding, and content generation, models with long context lengths are growing in importance.
However, managing long contexts brings substantial challenges due to the expansion of key-value cache (KV cache). Longer KV cache requires larger memory, limiting the batch-size thus decreasing throughput. Also, computing attention over long KV cache incurs more memory access, hurting the end-to-end latency.
Prior works find that it is sufficient to use only the recent and high-impact tokens for attention computation, allowing the eviction of less vital tokens to shrink cache size.
Nonetheless, we observe a dynamic shift in token importance across different decoding steps. Tokens initially evicted might regain importance after certain decoding steps.
To address this, we propose ArkVale, a page-based KV cache manager that can recognize and recall currently important tokens evicted before. We asynchronously copy the filled page into external memory (e.g., CPU memory) as backup and summarize it into a much smaller digest by constructing the bounding-volume of its keys. Before attention computation, we measure all pages' importance based on their digests, recall the important ones, evict the unimportant ones, and select the top-ranked pages for attention computation.
Experiment results show that ArkVale performs well on various long context tasks with negligible accuracy loss under 2k$\sim$4k cache budget and can improve decoding latency to $2.2\times$ and batching throughput to $4.6\times$ because it applies attention on only a small subset of pages and reduce per-sample memory usage of KV cache. Renze Chen, Zhuofeng Wang, Beiquan Cao, Size Zheng 0001, Xuechao Wei, Shengen Yan, Meng Li 0004, Yun Liang 0001 |
NeurIPS | 7 |
| 2024 | POSTER: RadiK: Scalable Radix Top-K Selection on GPUsabstractBy identifying the k largest or smallest elements in a set of data, top-k selection is critical for modern high-performance databases and machine learning systems, especially with large data volumes. However, previous studies on its GPU implementation are mostly merge-based and rely heavily on the high-speed but size-limited on-chip memory, thereby resulting in a restricted upper bound on k. This paper introduces RadiK, a highly optimized GPU-parallel radix top-k selection that is scalable with k, input length, and batch size. With a carefully designed optimization framework targeting high memory bandwidth and resource utilization, RadiK supports far larger k than the prior art, achieving up to 2.5× speedup for non-batch queries and up to 4.8× speedup for batch queries. We also propose a lightweight refinement that strengthens the robustness of RadiK against skewed distributions by adaptively scaling the input elements. Bole Zhou, Jiejing Zhang, Xuechao Wei, Yinghan Li 0002, Yingda Chen |
PPoPP | 4 |
| 2023 | Klotski: DNN Model Orchestration Framework for Dataflow Architecture AcceleratorsabstractDataflow architecture accelerators are a new kind of scalable DNN accelerators. The availability of input operands of the instructions solely determines the execution of instructions. This paper proposes the Klotski framework to solve DNN model orchestration for dataflow architecture accelerators. First, a Bayesian optimization-based entropy-directed partition algorithm is proposed to transform a DNN model into$\mu \mathbf{ops}$. Second, a unified formal formulation for$\mu \mathbf{ops}$scheduling and mapping is presented. Third, a two-stage methodology is proposed to decouple the scheduling and mapping, making the solution feasible. Extensive results show that Klotski outperforms baselines in runtime by an average of 9.55% and 48.48%. Xuechao Wei, Youwei Zhuo, Yi Cai 0003, Hongzhong Zheng, Bei Yu 0001, Yuan Xie 0001 |
ICCAD | 2 |
| 2023 | ArchExplorer: Microarchitecture Exploration Via Bottleneck AnalysisabstractDesign space exploration (DSE) for microarchitecture parameters is an essential stage in microprocessor design to explore the trade-offs among performance, power, and area (PPA). Prior work either employs excessive expert efforts to guide microarchitecture parameter tuning or demands high computing resources to prepare datasets and train black-box prediction models for DSE. Jiayi Huang 0001, Xuechao Wei, Yuzhe Ma, Sicheng Li 0001, Hongzhong Zheng, Bei Yu 0001, Yuan Xie 0001 |
MICRO | 3 |
| 2023 | Efficient Super-Resolution System With Block-Wise Hybridization and Quantized Winograd on FPGAabstractSuper-resolution (SR) techniques aim to restore a high-resolution (HR) image from low-resolution (LR) images, which are often used to assist the enhancement of image/video quality under the rapid development of HR and high-frame-rate media. Recently, neural network (NN)-based methods perform much better image reconstruction quality than classical approaches. However, the unacceptable computation complexity as well as the huge memory footprints of NNs limit the throughputs and scalability of these SR systems. In this work, we analyze several key issues in the design of NN-based SR systems first. Then, we propose a three-level systematic optimization methodology for SR systems to reduce computation overhead and keep image quality. At the algorithm level, we introduce image blocking to SR tasks and develop a block-wise SR algorithm based on the hybrid of NN and interpolation with a consistent image block evaluation metric. The configurable hybrid parameters help the SR algorithm to achieve a flexible tradeoff between the computation overhead and image quality. At the operator level, we focus on the transpose convolution operators commonly used for upsampling in SR NNs. We propose an efficient Winograd-based transposed convolution acceleration method. Through the efficient subconvolutions conversion and the Winograd specialization, this methods enables unified Winograd transformations and simplified data access patterns. At the data level, we propose a novel quantization method for Winograd-aware SR NNs to get better-quantized accuracy. Comprehensive evaluations demonstrate the effectiveness of these optimizations. Our SR system reduces a large number of multiplications with great scalability and supports 4K@120 fps and 8K@30 fps outputs with acceptable image quality degradation. Bizhao Shi, Jiaxi Zhang 0001, Zhuolun He, Xuechao Wei, Sicheng Li 0001, Guojie Luo, Hongzhong Zheng, Yuan Xie 0001 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2023 | An Intermediate-Centric Dataflow for Transposed Convolution Acceleration on FPGAabstractTransposed convolution has been prevailing in convolutional neural networks (CNNs), playing an important role in multiple scenarios such as image segmentation and back-propagation process of training CNNs. This mainly benefits from the ability to up-sample the input feature maps by interpolating new information from the input feature pixels. However, the backward-stencil computation constrains its performance and hindered its wide application in diverse platforms. Moreover, in contrast to the efforts on accelerating the convolution, there is a rare investigation on the acceleration of transposed convolution that is identically compute-intensive as the former. For acceleration of transposed convolution, we propose an intermediate-centric dataflow scheme, in which we decouple the generation of the intermediate patch from its further process, aim at efficiently performing the backward-stencil computation . The intermediate-centric dataflow breaks the transposed convolution into several phases/stages, achieving feeding the input feature maps and performing the backward-stencil computation in a pipelining manner. It also provides four-degree computation parallelism and efficient data reuse of input feature maps/weights. Furthermore, we also theoretically analyze the irregular data dependence leveraging the polyhedral model, which constrains the parallel computing of transposed convolution. Additionally, we devise an optimization problem to explore the design space and automatically generate the optimal design configurations for different transposed convolutional layers and hardware platforms. By selecting the representative transposed convolutional layers from DCGAN, FSRCNN, and FCN, we generate the corresponding accelerator arrays of intermediate-centric dataflow on the Xilinx Alveo U200 platform and reach the performance of 3.92 TOPS, 2.72 TOPS, and 4.76 TOPS, respectively. Zhengzheng Ma, Tuo Dai, Xuechao Wei, Guojie Luo |
ACM Trans. Embed. Comput. Syst. | 3 |
| 2022 | GNNear: Accelerating Full-Batch Training of Graph Neural Networks with near-Memory ProcessingabstractRecently, Graph Neural Networks (GNNs) have become state-of-the-art algorithms for analyzing non-euclidean graph data. However, to realize efficient GNN training is challenging, especially on large graphs. The reasons are many-folded: 1) GNN training incurs a substantial memory footprint. Full-batch training on large graphs even requires hundreds to thousands of gigabytes of memory. 2) GNN training involves both memory-intensive and computation-intensive operations, challenging current CPU/GPU platforms. 3) The irregularity of graphs can result in severe resource under-utilization and load-imbalance problems. Zhe Zhou 0002, Cong Li 0008, Xuechao Wei, Xiaoyang Wang 0006, Guangyu Sun 0003 |
PACT | 3 |
| 2022 | 2022 ICCAD CAD Contest Problem C: Microarchitecture Design Space ExplorationabstractIt is vital to select microarchitectures to achieve good trade-offs between performance, power, and area in the chip development cycle. Combining high-level hardware description languages and optimization of electronic design automation tools empowers microarchitecture exploration at the circuit level. Due to the extremely large design space and high runtime cost to evaluate a microarchitecture, ICCAD 2022 CAD Contest Problem C calls for an effective design space exploration algorithm to solve the problem. We formulate the research topic as a contest problem and provide benchmark suites, contest benchmark platforms, etc., for all contestants to innovate and estimate their algorithms. Sicheng Li 0001, Xuechao Wei, Bizhao Shi, Yen-Kuang Chen, Yuan Xie 0001 |
ICCAD | 3 |
| 2022 | PetS: A Unified Framework for Parameter-Efficient Transformers Serving
Zhe Zhou 0002, Xuechao Wei, Jiejing Zhang, Guangyu Sun 0003 |
USENIX ATC | 2 |
| 2020 | FTDL: A Tailored FPGA-Overlay for Deep Learning with High ScalabilityabstractFast inference is of paramount value to a wide range of deep learning applications. This work presents FTDL, a highly-scalable FPGA overlay framework for deep learning applications, to address the architecture and hardware mismatch faced by traditional efforts. The FTDL overlay is specifically optimized for the tiled structure of FPGAs, thereby achieving post-place-and-route operating frequencies exceeding 88 % of the theoretical maximum across different devices and design scales. A flexible compilation framework efficiently schedules matrix multiply and convolution operations of large neural network inference on the overlay and achieved over 80 % hardware efficiency on average. Taking advantage of both high operating frequency and hardware efficiency, FTDL achieves 402.6 and 151.2 FPS with GoogLeNet and ResNet50 on ImageNet, respectively, while operating at a power efficiency of 27.6 GOPS/W, making it up to 7.7× higher performance and 1.9× more power-efficient than the state-of-the-art. Runbin Shi, Yuhao Ding, Xuechao Wei, He Li 0008, Hang Liu 0001, Hayden Kwok-Hay So, Caiwen Ding |
DAC | 3 |
| 2020 | FTDL: An FPGA-tailored Architecture for Deep Learning SystemsabstractHardware acceleration of deep learning (DL) systems has been increasingly studied to achieve desirable performance and energy efficiency. The FPGA strikes a balance between high energy efficiency and fast development cycle and therefore is widely used as a DNN accelerator. However, there exists an architecture-layout mismatch in the current designs, which introduces scalability and flexibility issues, leading to irregular routing and resource imbalance problems. To address these limitations, in this work, we propose FTDL, an FPGA-tailored architecture with a parameterized and hierarchical hardware that is adaptive to different FPGA devices. FTDL has the following novelties: (i) At the architecture level, FTDL consists of Tiled Processing Elements (TPE) and super blocks, to achieve a near-to-theoretical digital signal processing (DSP) operating-frequency of 650 MHz. More importantly, FTDL is configurable and delivers good scalability, i.e., the timing is stabilized even when the design is scaled-up to 100% resource utilization for different deep learning systems. (ii) In workload compilation, FTDL provides a compiler that manages to map the DL workloads to the architecture level in an optimal manner. Experimental results show that for most benchmark layers in MLPerf, FTDL achieves an over 80% hardware efficiency. Runbin Shi, Yuhao Ding, Xuechao Wei, Hang Liu 0001, Hayden Kwok-Hay So, Caiwen Ding |
FPGA | 3 |
| 2019 | Overcoming Data Transfer Bottlenecks in FPGA-based DNN Accelerators via Layer Conscious Memory ManagementabstractDeep Neural Networks (DNNs) are becoming more and more complex than before. Previous hardware accelerator designs neglect the layer diversity in terms of computation and communication behavior. On-chip memory resources are underutilized for the memory bounded layers, leading to suboptimal performance. In addition, the increasing complexity of DNN structures makes it difficult to do on-chip memory allocation. To address these issues, we propose a layer conscious memory management framework for FPGA-based DNN hardware accelerators. Our framework exploits the layer diversity and the disjoint lifespan information of memory buffers to efficiently utilize the on-chip memory to improve the performance of the layers bounded by memory and thus the entire performance of DNNs. It consists of four key techniques working coordinately with each other. We first devise a memory allocation algorithm to allocate on-chip buffers for the memory bound layers. In addition, buffer sharing between different layers is applied to improve on-chip memory utilization. Finally, buffer prefetching and splitting are used to further reduce latency. Experiments show that our techniques can achieve 1.36X performance improvement compared with previous designs. Xuechao Wei, Yun Liang 0001, Jason Cong |
DAC | 1 |
| 2019 | Overcoming Data Transfer Bottlenecks in DNN Accelerators via Layer-Conscious Memory ManagmentabstractDeep Neural Networks (DNNs) are rapidly evolving to satisfy the performance and accuracy requirements in many real world applications. The evolution renders DNNs more and more complex in terms of network topology, data sizes and layer types. Currently most state-of-the-art DNN accelerators adopt a uniform memory hierarchy (UMH) design methodology, which means that the data transferring of all convolutional and fully connected layers must go through the same memory levels. Unfortunately, for some layers, the performance is always bounded by off-chip memory transferring. It is caused by the saturating of data reuse happening in on-chip buffers, resulting in underutilization of on-chip memory. To address this issue, we propose a layer-conscious memory hierarchy (LCMH) methodology for DNN accelerators. LCMH could determine the memory levels of all the layers according to their requirements for off-chip memory bandwidth and on-chip buffer size for the data sources. As a result, the off-chip memory footprints of memory bounded layers could be avoided by keeping the data of them on chip. In addition, we provide architectural support for the accelerators equipped with LCMH. Experimental results show that designs with layer- conscious memory management could achieve up to 36% speedup compared with the designs wth UMH and 5% improvement over state-of-the-art designs. Xuechao Wei, Yun Liang 0001, Peng Zhang 0007, Cody Hao Yu, Jason Cong |
FPGA | 1 |
| 2019 | Frequency Improvement of Systolic Array-Based CNNs on FPGAsabstractFPGAs are commercially available off-the-shelf for implementing convolutional neural network (CNN) accelerators to trade off accuracy, performance, and power. Systolic array architecture for CNN accelerators on FPGAs has the potential to run at a high frequency due to its regular and simple interconnections. However, current FPGA CAD tools are unable to synthesize and layout systolic arrays in high quality. In this paper, we identify the reasons for the frequency degradation of systolic array designs for CNN accelerators. We also propose two methods to improve the frequency at the front-end and the back-end, respectively. The experimental results show that our methods are able to achieve 1.29 × higher frequency and attain 1.5TOPS for the VGG16 network on the Xilinx KCU1500 platform. Jiaxi Zhang 0001, Wentai Zhang 0001, Guojie Luo, Xuechao Wei, Yun Liang 0001, Jason Cong |
ISCAS | 4 |
| 2018 | TGPA: tile-grained pipeline architecture for low latency CNN inferenceabstractFPGAs are more and more widely used as reconfigurable hardware accelerators for applications leveraging convolutional neural networks (CNNs) in recent years. Previous designs normally adopt a uniform accelerator architecture that processes all layers of a given CNN model one after another. This homogeneous design methodology usually has dynamic resource underutilization issue due to the tensor shape diversity of different layers. As a result, designs equipped with heterogeneous accelerators specific for different layers were proposed to resolve this issue. However, existing heterogeneous designs sacrifice latency for throughput by concurrent execution of multiple input images on different accelerators. In this paper, we propose an architecture named Tile-Grained Pipeline Architecture (TGPA) for low latency CNN inference. TGPA adopts a heterogeneous design which supports pipelining execution of multiple tiles within a single input image on multiple heterogeneous accelerators. The accelerators are partitioned onto different FPGA dies to guarantee high frequency. A partition strategy is designd to maximize on-chip resource utilization. Experiment results show that TGPA designs for different CNN models achieve up to 40% performance improvement than homogeneous designs, and 3X latency reduction over state-of-the-art designs. Xuechao Wei, Yun Liang 0001, Cody Hao Yu, Peng Zhang 0007, Jason Cong |
ICCAD | 1 |
| 2017 | Throughput optimization for streaming applications on CPU-FPGA heterogeneous systemsabstractStreaming processing is an important technology that finds applications in networking, multimedia, signal processing, etc. However, it is very challenging to design and implement streaming applications as they impose complex constraints. First, the tasks involved in the streaming applications must complete the computation under a latency constraint. Second, streaming systems are built under more and more stringent power budget. Hence, power capping technique is employed to manage the power consumption for streaming systems. To accommodate these needs, heterogeneous systems that consist of CPUs and FPGAs are becoming increasingly popular due to their performance and power benefits. In this paper, we optimize the throughput for streaming applications on CPU-FPGA heterogeneous system under latency and power constraints. We develop two algorithms to map the tasks onto the heterogeneous system and order their execution by exploiting the heterogeneity in architectural capabilities and task characteristics. We also employ pipelining to improve the throughput by overlapping the execution of different frames and use frequency scaling to adjust the execution of tasks for power saving. Experiments using a variety of streaming applications show that our heterogeneous solution can successfully meet the latency and power constraints for the cases where the CPU implementation fails. Furthermore, our technique can improve the throughput by 37.32% on average. Xuechao Wei, Yun Liang 0001, Tao Wang 0004, Songwu Lu, Jason Cong |
ASP-DAC | 1 |
| 2017 | Automated Systolic Array Architecture Synthesis for High Throughput CNN Inference on FPGAsabstractConvolutional neural networks (CNNs) have been widely applied in many deep learning applications. In recent years, the FPGA implementation for CNNs has attracted much attention because of its high performance and energy efficiency. However, existing implementations have difficulty to fully leverage the computation power of the latest FPGAs. In this paper we implement CNN on an FPGA using a systolic array architecture, which can achieve high clock frequency under high resource utilization. We provide an analytical model for performance and resource utilization and develop an automatic design space exploration framework, as well as source-to-source code transformation from a C program to a CNN implementation using systolic array. The experimental results show that our framework is able to generate the accelerator for real-life CNN models, achieving up to 461 GFlops for floating point data type and 1.2 Tops for 8-16 bit fixed point. Xuechao Wei, Cody Hao Yu, Peng Zhang 0007, Youxiang Chen, Yun Liang 0001, Jason Cong |
DAC | 1 |
| 2012 | Distributed replay protocol for distributed uniprocessorsabstractData speculation technique has been heavily exploited in various scenarios of architecture design. It bridges the time or space gap between data producer and data consumer, which gives opportunities to processors to gain significant speedups. However, large instruction windows, deep pipeline and increasing latency of on-chip communication make data misspeculation very expensive in modern processors. Mengjie Mao, Hong An, Bobin Deng, Xuechao Wei, Wenting Han |
ICS | 5 |
| 2012 | FlexBFS: a parallelism-aware implementation of breadth-first search on GPUabstractIn this paper, we present FlexBFS, a parallelism-aware implementation for breadth-first search on GPU. Our implementation can adjust the computation resources according to the feedback of available parallelism dynamically. We also optimized our program in three ways: (1)a simplified two-level queue management,(2)a combined kernel strategy and (3)a high-degree vertices specialization approach. Our experimental results show that it can achieve 3~20 times speedup against the fastest serial version, and can outperform the TBB based multi-threading CPU version and the previous most effective GPU version on all types of input graphs. Gu Liu, Hong An, Wenting Han, Xuechao Wei, Xulong Tang |
PPoPP | 7 |