EDBT 2026 Demo / reviewers in the wild / expert
Wenqi Lou
dblp:211/7977
· DBLP profile ↗
30ranked-venue papers
7as first author
28since 2021 · last 2026
0000-0002-2240-6672ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 24 · 7 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CloserToMe: A Unified Framework for Accurate and Transferable Latency Prediction Across Heterogeneous DevicesabstractHardware accelerators such as GPUs, NPUs, and FPGAs are essential to meeting AI’s computational demands. With the proliferation of heterogeneous devices across cloud and edge, various model optimization techniques adapt to diverse hardware characteristics through operator transformations and structural modifications. Accurate, efficient latency prediction enables rapid selection of optimal strategies across hardware backends. Many existing methods treat hardware as a black-box executor, directly regressing latency without explicitly modeling the intricate interactions between neural network (NN) structures and device-specific execution behaviors. To address these challenges, we introduce a new modeling perspective that captures the interaction between neural architectures and hardware execution. To capture device-specific characteristics, we propose two complementary modeling strategies. The Device Behavior Signature Selector (DBSel) characterizes hardware execution behavior by selectively probing a small set of representative architectures, forming a compact, workload-driven profile. In parallel, we construct capability vectors that capture the hierarchical memory of each device and compute characteristics, providing a structured abstraction of its architectural capacity. To unify both behavioral and structural views, we introduce the Hardware–Operation Dialogue Module (HODM), which models fine-grained interactions between neural operators and hardware properties. Together, these components empower CloserToMe to deliver accurate and transferable latency predictions across unseen and diverse platforms. Cheng Tang 0004, Guochong Sui, Wenqi Lou, Jiayi Tuo, Wenqian Xie, Yinkang Gao, Yixuan Zhu, Lei Gong 0003, Chao Wang 0003, Xuehai Zhou |
AAAI | 3 |
| 2026 | Window-Diffusion: Accelerating Diffusion Language Model Inference with Windowed Token Pruning and Caching
Fengrui Zuo, Zhiwei Ke, Wenqi Lou, Lei Gong 0003, Chao Wang 0003, Xuehai Zhou |
APPT | 5 |
| 2026 | Two-Stage Hierarchy-Aware Learning with Gradient Conflict Mitigation for HLS Latency and Resource Prediction
Zhiwei Ke, Fengrui Zuo, Wenqi Lou, Chao Wang 0003, Xuehai Zhou |
Euro-Par (2) | 4 |
| 2026 | Realizable N:M Sparse Transformer Inference via Search-Kernel Co-design
Wenqi Lou, Zhiguang Wang, Zhiwei Ke, Fengrui Zuo, Chao Wang 0003, Xuehai Zhou |
Euro-Par (1) | 2 |
| 2026 | UniCoX: A Unified Cost Model for Tensorized Program Tuning Across Ubiquitous AcceleratorsabstractTensorized programs leverage hardware intrinsics on accelerators to boost tensor computation performance. With the rise of hardware customization, massive accelerators and intrinsics have emerged, posing engineering challenges for manually coding efficient tensorized programs. Tensorized program tuning in deep learning compilers (DLCs) is considered an effective approach. At the core of program tuning relies the design of the cost model to predict performance. However, there is currently a lack of cost models specifically for tensorized programs, which severely hinders the co-optimization of deep learning compilers and hardware accelerators.In this paper, we propose UniCoX, a unified cost model specifically for tensorized program tuning across ubiquitous accelerators. We systematically analyze the design challenges introduced by tensorized programs from the perspectives of feature representation and transfer prediction. For feature representation, we leverage attention mechanisms to mine key software schedule features and design corresponding aligned hardware features, resulting in a unified cross-accelerator feature representation. For transfer prediction, by integrating lifelong learning and transfer learning with data sampling strategies, we propose a unified transfer prediction strategy to keep pace with the rapid development of accelerators. To meet training and testing demands, we construct TensorizeSetX, a dataset dedicated to tensorized program tuning. Results show that UniCoX achieves the state-of-the-art accuracy while supporting low-cost and flexible transfer prediction. It can accelerate search time by 11.3’ and improve inference speed by 1.9’ within the state-of-the-art tensorized program tuning framework, TVM MetaSchedule. Lei Gong 0003, Wenqi Lou, Qianyu Cheng, Xianglan Chen, Chao Wang 0003, Xuehai Zhou |
IEEE Trans. Computers | 3 |
| 2026 | UniSparTa: A Unified Sparse Tensor Program Tuning FrameworkabstractSparse tensor computation is widely used in deep learning and scientific computing. However, diverse sparse data and algorithmic characteristics at the application level, combined with the diversity of hardware platforms, pose significant challenges for efficient sparse tensor program optimization. Manually crafted operator libraries are time-consuming to develop and lack portability. To address this, we propose UniSparTa, a unified sparse tensor program tuning framework that automatically generates high-performance programs. First, we extract unified optimization principles for high-performance sparse tensor programs and propose a domain-specific language (DSL) to automatically generate a high-quality design space without manual intervention. Second, by analyzing the general distribution of the design space, we introduce an adaptive search strategy combining Deep Q-Networks (DQN) and Simulated Annealing (SA). Finally, to avoid the unacceptable time cost of real measurement during tuning, we propose a unified cost model based on multimodal fusion to accurately predict program performance. Furthermore, by leveraging data augmentation and transfer learning, we enable low-cost transfer prediction across different sparse data patterns, algorithms, and hardware platforms. Results show that, compared to the state-of-the-art operator library MKL, the manually optimized scheme ASpT, the tensor compiler TVM, and the sparse tensor tuning framework WACO, UniSparTa achieves average speedups of 1.98×, 2.75×, 6.13×, and 1.75×, respectively. Moreover, UniSparTa significantly accelerates the tuning process. Lei Gong 0003, Xiangjun Qu, Cheng Tang 0004, Wenqi Lou, Qianyu Cheng, Xianglan Chen, Chao Wang 0003, Xuehai Zhou |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2026 | LORA: A Latency-Oriented Recurrent Architecture for Large Language Model on Multi-FPGA Platform With Communication OptimizationabstractThe remarkable performance of Large Language Models (LLMs) has driven their widespread deployment in data centers to support diverse user-facing applications. However, the rapidly growing computational and storage demands of these models have made single-device deployment increasingly impractical. Prior research on LLM inference has primarily addressed this challenge through algorithmic optimizations such as quantization or by integrating customized hardware acceleration frameworks. As model parameters continue to scale, multi-device deployment has become a necessary approach for enabling efficient LLM inference. Nevertheless, constructing low-latency multi-device platforms for LLMs inference using available FPGA or GPU accelerators remains constrained by inefficient synchronization schemes or limited compute intensity in current architectures. Furthermore, existing solutions often lack co-optimized designs that effectively integrate communication with computation. To address these limitations, this paper proposes LORA, a low-latency end-to-end LLMs acceleration platform utilizing multiple FPGAs. Firstly, we optimize the synchronization timing within the LLMs to minimize storage, computation, and BRAM overhead. Secondly, we tightly couple communication and computation through techniques such as pipeline overlapping and input data packing. Next, we deploy homogeneous accelerators on each FPGA device, leveraging a recurrent architecture to further reduce inference latency. Finally, we apply FPGA-specific optimizations and conduct performance modeling and analysis of the acceleration framework to select optimal deployment parameters for various computational tasks. Implemented on Xilinx Alveo U280 FPGAs, LORA-F and LORA-Q achieve average speedups of 14.4× and 32.6×, respectively, compared to NVIDIA V100 GPUs when running modern LLMs. Compared with existing multi-FPGA accelerator platforms, LORA-F and LORA-Q demonstrate average performance improvements of up to 2.6× and 4.3×, respectively. Zhendong Zheng, Qianyu Cheng, Wenqi Lou, Lei Gong 0003, Xianglan Chen, Chao Wang 0003, Xuehai Zhou |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2026 | TETRIS: A Novel FPGA Virtualization Framework for Fine-grained Sharing via Hierarchical ReconfigurationabstractField-Programmable Gate Arrays (FPGAs) are increasingly used in cloud platforms to accelerate diverse workloads, thanks to their reconfigurability and high performance. However, in multi-tenant cloud environments, existing FPGA virtualization mechanisms fail to align with dynamic application demands due to their static partitioning methods, leading to significant internal fragmentation. To address these issues, we present Tetris , an FPGA virtualization framework that supports flexible and dynamic reconfigurable resource allocation to improve cloud platform deployment efficiency. Specifically, Tetris deploys nested dynamic reconfigurable regions on FPGAs and adopts a tree structure to manage reconfigurable resources, enabling fine-grained resource reallocation at runtime. Enabled by the proposed fat-tree-based data transmission architecture between reconfigurable regions, Tetris can map dataflow-based applications onto these regions effectively. Additionally, Tetris offers dual-level resource optimization strategies to help system to balance the resource utilization and run-time compilation overhead. We evaluate Tetris with dataflow HLS benchmarks. Experimental results show that Tetris achieves a 1.16 \(\times\) improvement in resource utilization compared to advanced FPGA virtualization frameworks and delivers a 1.3 \(\times\) increase in run-time throughput, incurring less than 30% additional compilation latency. Tetris enables scalable, on-demand cloud FPGA acceleration, improving adaptability across different workloads. Wenbin Teng, Wenqi Lou, Lei Gong 0003, Chao Wang 0003, Xuehai Zhou |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2026 | MoE-Sched: Enabling Efficient FPGA Deployment of Mixture-of-Experts Vision Transformers via Coordinated SchedulingabstractCompared to traditional Vision Transformers (ViT), Mixture-of-Experts ViTs (MoE-ViTs) are introduced to scale model size without a proportional increase in computational complexity, making them a new research focus. Given the high performance and reconfigurability, field-programmable gate array (FPGA)-based accelerators for MoE-ViTs emerge, delivering substantial gains over general-purpose processors. However, existing accelerators often fall short of efficiently managing the highly dynamic and sparse computation patterns, resulting in suboptimal tradeoffs between resource utilization and performance. To address the inefficiencies in deploying MoE-ViTs on FPGAs, we present MoE-Sched, a novel end-to-end accelerator that embraces a scheduling-centric design philosophy. Rather than optimizing isolated kernels, MoE-Sched coordinates multilevel scheduling, from fine-grained intrakernel streaming to module reuse and multidie mapping, to holistically balance latency, bandwidth (BW), and resource usage. We further integrate a hardware-aware quantization scheme tailored for streaming attention and sparse expert execution, preserving accuracy while minimizing overhead. Experimental results demonstrate that our accelerator achieves nearly 100 frames/s on M3ViT-tiny, a$3.13\times $improvement in throughput, and over 75% energy reduction compared to state-of-the-art (SOTA) FPGA MoE accelerators, while maintaining less than 1% accuracy loss across vision benchmarks. Our implementation will be open-sourced. Jiale Dong, Wenqi Lou, Zhendong Zheng, Yunji Qin, Lei Gong 0003, Chao Wang 0003, Xuehai Zhou |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2025 | UniCoS: A Unified Neural and Accelerator Co-Search Framework for CNNs and ViTsabstractCurrent algorithm-hardware co-search works often suffer from lengthy training times and inadequate exploration of hardware design spaces, leading to suboptimal performance. This work introduces UniCoS, a unified framework for co-optimizing neural networks and accelerators for CNNs and Vision Transformers (ViTs). By introducing a novel training-free proxy that evaluates accuracy within seconds and a clustering-based algorithm for exploring heterogeneous dataflows, UniCoS efficiently navigates the design spaces of both architectures. Experimental results demonstrate that the solutions generated by UniCoS consistently surpass state-of-the-art (SOTA) methods (e.g., $3.54 \times$ energy-delay product (EDP) improvement with a $1.76 \%$ higher accuracy on ImageNet) while requiring notably reduced search time (up to $48 \times, \sim 3$ hours). The code is available at https://github.com/mine7777/Unicos.git. Wenqi Lou, Cheng Tang 0004, Hongbing Wen, Yunji Qin, Lei Gong 0003, Chao Wang 0003, Xuehai Zhou |
DAC | 2 |
| 2025 | CoQMoE: Co-Designed Quantization and Computation Orchestration for Mixture-of-Experts Vision Transformer on FPGA
Jiale Dong, Wenqi Lou, Zhendong Zheng, Lei Gong 0003, Chao Wang 0003, Xuehai Zhou |
Euro-Par (2) | 4 |
| 2025 | Spectral Enhanced Tuning: An Efficient Plug-and-Play Framework for Frequency-Aware DehazingabstractImage dehazing aims to recover clear images from hazy counterparts; however, current deep learning methods tend to rely on low-frequency information while neglecting multi-frequency integration, limiting their performance. To address this, we introduce Spectral Enhanced Tuning (SET), a modular framework composed of lightweight, plug-and-play components that can be easily integrated into existing dehazing networks to improve performance via efficient multi-frequency feature fusion. The framework features a Frequency Spectrum Decoder (FSD) that utilizes localized windows for high-frequency details and average pooling for low-frequency global features, improving efficiency over traditional global processing. Additionally, the Dual-Domain Guided Attention (DDGA) computes importance maps across frequency and spatial domains, surpassing simple concatenation or addition, while the Direction-Aware Convolution Module (DAConv) captures spatial details for better reconstruction. Experimental results demonstrate that SET achieves significant dehazing performance gains with negligible increases in parameters and computational cost, validating its efficiency and effectiveness. Cheng Tang 0004, Wenqi Lou, Qianyu Cheng, Jiayi Tuo, Tianhao Jiang, Chao Wang 0003, Xuehai Zhou |
ICME | 2 |
| 2025 | Automated FPGA Accelerator Generation Framework for Transformers with Dataflow OptimizationabstractTransformers have revolutionized natural language processing (NLP) and computer vision (CV) tasks, yet their deployment remains constrained by the quadratic complexity of self-attention. While FPGA-based accelerators offer promising solutions, existing designs struggle with fixed dataflow patterns and limited hardware specialization across diverse model architectures and sequence lengths. This paper presents AutoTrans, an automated framework for generating optimized FPGA accelerators tailored to Transformers. To improve dataflow flexibility, we propose configurable fusion granularities at the tensor, row, and block levels for self-attention layers, enabling fine-grained trade-offs between computational efficiency and resource utilization. For hardware specialization, we develop a parameterized heterogeneous multi-core architecture featuring dedicated compute engines for attention and linear layers, guided by a genetic algorithm-based design space exploration (DSE) strategy. Experimental results demonstrate the effectiveness of AutoTrans across varied application scenarios. On the ZCU102 board, AutoTrans achieves up to 621 GOPS when accelerating BERT-base with 4 K-token inputs, yielding a 1.27 × to 1.91 × improvement in energy efficiency over previous designs. For ViT-Base with 256-token inputs on the Alveo U50, AutoTrans attains up to 1548 GOPS, achieving up to 1.96 × higher compute density, highlighting its scalability and adaptability across NLP and vision domains. Wenqi Lou, Yunji Qin, Chao Wang 0003, Lei Gong 0003, Xuehai Zhou |
ICPP | 1 |
| 2025 | UbiMoE: A Ubiquitous Mixture-of-Experts Vision Transformer Accelerator With Hybrid Computation Pattern on FPGAabstractCompared to traditional Vision Transformers (ViT), Mixture-of-Experts Vision Transformers (MoE-ViT) are introduced to scale model size without a proportional increase in computational complexity, making them a new research focus. Given the high performance and reconfigurability, FPGA-based accelerators for MoE-ViT emerge, delivering substantial gains over general-purpose processors. However, existing accelerators often fall short of fully exploring the design space, leading to suboptimal trade-offs between resource utilization and performance. To overcome this problem, we introduce UbiMoE, a novel end-to-end FPGA accelerator tailored for MoE-ViT. Leveraging the unique computational and memory access patterns of MoE-ViTs, we develop a latency-optimized streaming attention kernel and a resource-efficient reusable linear kernel, effectively balancing performance and resource consumption. To further enhance design efficiency, we propose a two-stage heuristic search algorithm that optimally tunes hardware parameters for various FPGA resource constraints. Compared to state-of-the-art (SOTA) FPGA designs, UbiMoE achieves 1.34× and 3.35× throughput improvements for MoE-ViT on Xilinx ZCU102 and Alveo U280 platforms, respectively, while enhancing energy efficiency by 1.75× and 1.54×. Our implementation is available at https://github.com/DJ000011/UbiMoE. Jiale Dong, Wenqi Lou, Zhendong Zheng, Yunji Qin, Lei Gong 0003, Chao Wang 0003, Xuehai Zhou |
ISCAS | 2 |
| 2025 | TSI: A Time-Semantic Instruction Set for Deterministic Data-Flow Execution in Real-Time Embedded SystemsabstractReal-Time Embedded Systems (RTES) are widely used in safety-critical devices, where deterministic data flow is essential to system verification and reliable execution. It requires that each consumer task instance reads data from the deterministic producer task instance. In software based on general-purpose computing instruction sets, communication-related instruction execution order couples data flow among tasks, necessitating a deterministic execution order of these instructions to preserve data-flow determinism. However, enforcing this order complicates software, and suffers from priority inversion and variable execution overheads, which significantly increases task worst-case response times (WCRT) and response time variability. This paper identifies the cause of above issues as the semantics of general-purpose instruction sets, under which dataflow determinism relies on the deterministic execution order of communication-related instructions. To address this, we make the following contributions. First, we propose Time-Semantic Instruction set (TSI), which supports memory access using both addresses and timestamps. TSI enables data-flow determinism without strict instruction ordering. Second, we design a TSI-enabled implementation compatible with conventional memory systems. Third, we provide two TSI-based deterministic data-flow programming paradigms, along with correctness proofs. Finally, we evaluate TSI hardware cost and implement a cycle-accurate simulator based on a TSI-extended RISC-V. Experiments demonstrate that, under reasonable memory overhead, our approach reduces programming complexity and achieves up to$21.6 \times$reduction in WCRT and up to$89.6 \times$reduction in response time variability compared to existing methods. Yinkang Gao, Yixuan Zhu, Lei Gong 0003, Wenqi Lou, Chao Wang 0003, Xi Li 0003, Xuehai Zhou |
RTSS | 6 |
| 2025 | Optimizing utilization in logical execution time system with preserved externally-observable timed I/O semantics
Caixu Zhao, Yinkang Gao, Yixuan Zhu, Lei Gong 0003, Wenqi Lou, Xi Li 0003 |
J. Syst. Archit. | 7 |
| 2025 | Picasso: Analyzing Prompt Design for Text-to-Image Generative Diffusion Models from a Temporal-Spatial PerspectiveabstractVisual Concept Implantation (VCI) is essential in text-to-image fields. While VCI methods in diffusion models have matured, Visual Concept Disentanglement (VCD) remains an unexplored area. VCD involves analyzing Prompt Spaces trained in VCI to produce disentangled SubPrompts or SubCones for exploring interpretability. However, challenges arise due to Prompt Space design complexity and feature information extraction. We propose Picasso, a unified framework for VCD in diffusion models (DM). Our contributions include: for performance evaluation: Transforming VCD in DM into a regular clustering task by Visualization based on SubCones (VbSC); for unified framework design: Picasso processes diverse Prompt Spaces using spatially clusterable features; for Prompt Space exploration: Introducing a temporal-spatial SOTA Prompt Design subset based on temporal features. Our method provides a feasible mechanism for VCD in DM. Through functional and interpretability validation methods, we will comprehensively evaluate the effectiveness of our proposed method in visual concept implantation tasks and verify the correctness of the parameter space design principles. Picasso will be released at https://haoyu-cai.github.io/our_picasso . Haoyu Cai, Wenqi Lou, Chao Wang 0003, Xuehai Zhou |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2024 | Enhancing Long Sequence Input Processing in FPGA-Based Transformer Accelerators through Attention FusionabstractAttention-based transformers have achieved significant performance breakthroughs in natural language processing (NLP) and computer vision (CV) tasks. Meanwhile, the ever-increasing length of today’s input sequences puts much pressure on computing devices. FPGAs are widely used to accelerate Transformer inference due to their high energy efficiency and flexibility. However, most of the existing FPGA-based Transformer accelerators are oriented to small input lengths, making it hard to accelerate long input sequences. To this end, we design an efficient Transformer accelerator for FPGA and long-sequence input scenarios. We use the tiling softmax algorithm to fuse attention computation, eliminating the memory and bandwidth bottleneck in the attention layer and allowing our accelerator to support arbitrary input sequence lengths. We use BERT-Base on the Alveo U50 board for evaluation, and our implementation achieves computational efficiency improvements of 1.09 ∼ 2.48 × over prior FPGA accelerators. Besides, our accelerator can support up to 175K input sequence length when running BERT-like structures, far more than previous designs. Yunji Qin, Wenqi Lou, Chao Wang 0003, Lei Gong 0003, Xuehai Zhou |
ACM Great Lakes Symposium on VLSI | 2 |
| 2024 | AutoSparse: A Source-to-Source Format and Schedule Auto- Tuning Framework for Sparse Tensor ProgramabstractSparse tensor computation plays a crucial role in modern deep learning workloads, and its expensive computational cost leads to a strong demand for high-performance oper-ators. However, developing high-performance sparse operators is exceptionally challenging and tedious. Existing vendor operator libraries fail to keep pace with the evolving trends in new algorithms. Sparse tensor compilers simplify the development and optimization of operator, but existing work either requires significant engineering effort for tuning or suffers from limitations in search space and search strategies, which creates unavoidable cost and efficiency issues. In this paper, we propose AutoSparse, a source-to-source auto-tuning framework that targets sparse for-mat and schedule for sparse tensor program. Firstly, AutoSparse designs a sparse tensor DSL based on dynamic computational graph at the front-end, and proposes a sparse tensor program computational pattern extraction and automatic design space generation scheme based on it. Second, AutoSparse's back-end designs an adaptive exploration strategy based on reinforcement learning and heuristic algorithm to find the optimal format and schedule configuration in a large-scale design space. Compared to prior work, developers using AutoSparse do not need to specify tuning design space relied on any compilation or hardware knowledge. We use the SuiteS parse dataset to compare with four state-of-the-art baselines, namely, the high-performance operator library MKL, the manually-based optimisation scheme ASpT, the auto-tuning-based framework TVM-S and WACO. The results demonstrate that AutoSparse achieves average speedups of 1.92-$2.48 \times. 1.19-6.34 \times$. and$1.47-2.23\times$for the SpMV, SpMM, and SDDMM operators, respectively. We will open-source AutoSparse at https://github.com/Qu-Xiangjun/AutoSparse. Xiangjun Qu, Lei Gong 0003, Wenqi Lou, Qianyu Cheng, Xianglan Chen, Chao Wang 0003, Xuehai Zhou |
ICCD | 3 |
| 2024 | UniCoMo: A Unified Learning-Based Cost Model for Tensorized Program TuningabstractTensorized programs use hardware intrinsics on accelerators to significantly improve tensor computation performance. The trend of hardware customization has led to the emergence of massive hardware accelerators and intrinsics, which poses a significant engineering challenge for manually coding efficient tensorized programs. Tensorized program tuning in deep learning compilers (DLCs) is considered an effective approach to address the challenge. At the core of program tuning relies the design of the cost model, but currently there is still a lack of cost models specifically designed for tensorized programs, which hampers the co-optimization of DLCs and hardware accelerators. In this paper, we propose UniCoMo, a unified cost model for tensorized program tuning on various hardware platforms and hardware intrinsics. We first analyze the design challenges introduced by tensorized programs for cost models in terms of feature representation and transfer prediction. And then, for feature representation, we propose a unified feature representation for tensorized programs by using program behavior as a template, mining program features with the schedule attention matrix, and incorporating hardware intrinsic abstraction. For transfer prediction, we propose a unified transfer prediction strategy for tensorized program cost model based on lifelong learning and transfer learning. To meet training and testing requirements, we constructed a dataset dedicated to tensorized program tuning. Results show that UniCoMo maintains the state-of-the-art accuracy while significantly improving adaptability to diverse execution environments and enabling flexible transfer prediction. It can speed up search time by 9.8× and improve inference speed by 1.9× open-sourced at https://github.com/ZhW-loop/UniCoMo. Lei Gong 0003, Wenqi Lou, Qianyu Cheng, Xianglan Chen, Chao Wang 0003, Xuehai Zhou |
ICCD | 3 |
| 2024 | Fine-Grained Shared Cache Interference Analysis Using Basic Block's Execution TimeabstractShared last-level caches in multi-core architectures may lead to mutual interference in memory accesses between cores, resulting in additional access latency. To ensure the accuracy of programs' worst-case execution time (WCET) analysis, the inter-core interference of shared caches must be analyzed. This involves getting the time when memory access may occur between cores (i.e., inter-core context). But due to the uncertainty of inputs, describing inter-core context becomes extremely challenging. Existing works often consider the entire execution time of program as the time when all memory accesses may occur to enhance scalability, which ignores the timing information within programs and may lead to an overestimation of interference. This paper proposes a fine-grained shared cache interference analysis method based on the execution time of basic blocks. It employs a path-based strategy to estimate the execution time of basic blocks and use it as the lifecycles of memory accesses within the blocks, which can effectively eliminate impossible interferences. Experiments show that compared to the all interference method, we can reduce WCET by 65% in the best case and by 14% on average, with only a 130% increase in average analysis time. Yixuan Zhu, Wenqi Lou, Yinkang Gao, Binze Jiang, Xiaohang Gong, Xi Li 0003 |
ICCD | 2 |
| 2024 | MFNAS: Multi-fidelity Exploration in Neural Architecture Search with Stable Zero-Shot Proxy
Wenqi Lou, Yunji Qin, Lei Gong 0003, Chao Wang 0003, Xuehai Zhou |
PRICAI (1) | 2 |
| 2024 | Unleashing Network/Accelerator Co-Exploration Potential on FPGAs: A Deeper Joint SearchabstractRecently, algorithm-hardware co-exploration for neural networks (NNs) has become the key to obtaining high-quality solutions. However, previous efforts for FPGAs focus on neural architecture search (NAS) while lacking hardware architecture search (HAS), thus limiting the full potential of co-design. Although expanding the scope of HAS offers performance potential, the exponentially increased joint search space presents a formidable challenge. To address this, we propose a deep and efficient framework, which jointly searches for Networks and Accelerators for FPGAs in a balanced co-search space. First, we adjust the NAS space and then introduce a block-level bitwidth search on the software side. Meanwhile, we design a hardware-friendly quantization algorithm to facilitate hardware efficiency and accuracy. Second, we design a dataflow-configurable hardware unit with computation and memory access optimizations for quantized multiplication. Based on this, we incorporate critical heterogeneous multicore architecture exploration on the hardware side. Third, to enable rapid hardware feedback in the enlarged HAS space, we perform resource and performance modeling and design a fast hardware generation algorithm based on the genetic algorithm. Specifically, we apply optimization techniques, like mapping space pruning, greedy bandwidth allocation, and coarse-grained search, to speed up this process. We validate in edge and cloud scenarios. Experimental results show that efficiently explores a significantly larger joint space and provides high-quality solutions. Compared with previous state-of-the-art co-design works, the searched CNN-accelerator pairs improve the throughput by 2.07× ~ 7.10× and energy efficiency by 1.41× ~ 2.27× under similar accuracy on the ImageNet dataset. Wenqi Lou, Lei Gong 0003, Chao Wang 0003, Jiaming Qian, Xuan Wang 0020, Changlong Li 0006, Xuehai Zhou |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2024 | FlexBCM: Hybrid Block-Circulant Neural Network and Accelerator Co-Search on FPGAsabstractBlock-circulant matrix (BCM) compression has garnered much attention in the hardware acceleration of convolutional neural networks (CNNs) due to its regularity and efficiency. However, constrained by the difficulty of exploring the compression parameter space, existing BCM-based methods often apply a uniform compression parameter to all CNN models’ layers, losing the compression’s flexibility. Additionally, independently optimizing models or accelerators makes achieving the optimal tradeoff between model accuracy and hardware efficiency challenging. To this end, we propose FlexBCM, a joint exploration framework that efficiently explores both the parameter compression and hardware parameter space to generate customized hybrid BCM-compressed CNN and field-programmable gate array (FPGA) accelerator solutions. On the algorithmic side, leveraging the idea of neural architecture search (NAS), we design an efficient differentiable sampling method to rapidly evaluate the accuracy of candidate subnets. Additionally, we devise a hardware-friendly frequency domain quantization scheme for BCM computation. On the hardware side, we develop the efficient and parameter-configurable convolutional core (ConvPU) alongside the BCM computing core (BCMPU). The BCMPU can flexibly accommodate different compression parameters at runtime, incorporate complex-number DSP packing and conjugate symmetry optimizations. For model-to-hardware evaluation, we construct accurate latency and resource consumption models. Moreover, we design a fast hardware generation algorithm based on the coarse-grained search to provide prompt feedback on the hardware evaluation of the current subnet. Finally, we validate FlexBCM on the Xilinx ZCU102 FPGA and compare its compressed CNN-accelerator solutions with previous state-of-the-art works. Experimental results demonstrate that FlexBCM achieves 1.21–3.02 times higher-computational efficiency for ResNet18 and ResNet34 models while maintaining an acceptable accuracy loss on the ImageNet dataset. Wenqi Lou, Yunji Qin, Xuan Wang 0020, Lei Gong 0003, Chao Wang 0003, Xuehai Zhou |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2023 | NAF: Deeper Network/Accelerator Co-Exploration for Customizing CNNs on FPGAabstractRecently, algorithm and hardware co-design for neu-ral networks (NNs) has become the key to obtaining high-quality solutions. However, prior works lack consideration of the underlying hardware and thus suffer from a severely unbalanced neural architecture and hardware architecture search (NA-HAS) space on FPGAs, failing to unleash the performance potential. Nevertheless, a deeper joint search leads to a larger (multiplicative) search space, highly challenging the search. To this end, we propose an efficient differentiable search framework NAF, which jointly searches the networks (e.g., operations and bitwidths) and accelerators (e.g., heterogeneous multicores and mappings) under a balanced NA-HAS space. Concretely, we design a coarse-grained hardware-friendly quantization algorithm and integrate it at a block granularity into the co-search process. Meanwhile, we design a highly optimized block processing unit (BPU) with key dataflow configurable. Afterward, a dynamic hardware generation algorithm based on modeling and heuristic rules is designed to perform the critical HAS and fast generate hardware feedback. Experimental results show that compared with the previous state-of-the-art (SOTA) co-design works, NAF improves the throughput by$1.99\times\sim 6.84\times$on Xilinx ZCU102 and energy efficiency by 17%~88% under similar accuracy on the ImageNet dataset. Wenqi Lou, Jiaming Qian, Lei Gong 0003, Xuan Wang 0020, Chao Wang 0003, Xuehai Zhou |
DATE | 1 |
| 2023 | hAP: A Spatial-von Neumann Heterogeneous Automata Processor with Optimized Resource and IO Overhead on FPGAabstractRegular expression (REGEX) matching tasks drive much research on automata processors (AP). Among them, the von Neumann AP can efficiently utilize on-chip memory to process the Deterministic Finite Automata (DFA), but it is limited to small REGEX sets due to the DFA's state explosion problem. For large REGEX sets, the spatial AP based on Nondeterministic Finite Automaton (NFA) is the mainstream choice. However, there are two problems with previous FPGA-based spatial AP. First, it cannot obtain a balanced FPGA resource usage (LUT and BRAM), which easily leads to resource shortage. Second, to compress the report output data of large REGEX sets, it uses dynamic report compression, which not only consumes a lot of FPGA resources but also limits performance. Xuan Wang 0020, Lei Gong 0003, Wenqi Lou, Weiya Wang, Chao Wang 0003, Xuehai Zhou |
FPGA | 4 |
| 2022 | TCL-Net: A Lightweight and Efficient Dehazing Network with Frequency-Domain Fusion and Multi-Angle Attention
Cheng Tang 0004, Wenqi Lou |
ACCV (4) | 2 |
| 2022 | OctCNN: A High Throughput FPGA Accelerator for CNNs Using Octave Convolution AlgorithmabstractWith the rapid development of convolutional neural networks (CNNs), FPGAs have become one of the most attractive candidates for deploying CNNs. However, previous FPGA solutions based on the traditional convolution are still limited by computational power. In this article, we introduce the octave convolution (OctConv) into the CNN accelerator design for the first time to improve the hardware acceleration efficiency and design a dedicated OctPU for mapping OctConv to FPGAs, which employs a parallel dataflow pattern to exploit the parallelism of OctConv. Then, we present a novel and scalable architecture that dynamically combines the inter-layer pipelined structure and multi-layer reuse structure. Meanwhile, to obtain the optimized solution, we build a multidimensional performance and resource analysis model and a two-stage search algorithm based on greedy and heuristic algorithms. We evaluate our proposal by implementing VGG16 and ResNet50 on the Xilinx VU9P FPGA. Experimental results show that our prototypes can achieve an average of 3321 GOP/s for the convolutional layers for VGG16 and 2873 GOP/s for the overall ResNet50 using OctConv. Compared to previous works based on the traditional convolution, our prototypes own a 1.72 to 2.33 speedup in throughput and a 2.01 to 5.18 improvement in computational density. Our design also presents an excellent compromise performance and generalization Wenqi Lou, Lei Gong 0003, Chao Wang 0003, Zidong Du, Xuehai Zhou |
IEEE Trans. Computers | 1 |
| 2020 | OctCNN: An Energy-Efficient FPGA Accelerator for CNNs using Octave Convolution AlgorithmabstractRecently, embedded FPGAs have been explored as a potential platform for deploying machine learning on edge-devices due to their high energy efficiency and low cost. However, the lack of resources also makes the deployment of CNN on FPGAs more challenging. In this paper, we present OctCNN, which utilizes the octave convolution (OctConv) algorithm to optimize the FPGA-based CNN accelerator. We first propose a novel architecture for deploying OctConv on FPGAs and then present a resource and performance analysis model to guide a fast design space exploration. As a case study, we implement a classic CNN model, VGG16, on Xilinx ZC702. Results show, compared to the mobile-class CPU and GPU, OctCNN achieves$\mathbf{16.88}\times$and$\mathbf{2.43}\times$energy efficiency, respectively. Besides, it has a promising energy efficiency compared to previous FPGA accelerators. Wenqi Lou, Chao Wang 0003, Lei Gong 0003, Xuehai Zhou |
CLUSTER | 1 |
| 2019 | RV-CNN: Flexible and Efficient Instruction Set for CNNs Based on RISC-V Processors
Wenqi Lou, Chao Wang 0003, Lei Gong 0003, Xuehai Zhou |
APPT | 1 |