EDBT 2026 Demo / reviewers in the wild / expert
Wendong Mao
dblp:228/6705
· DBLP profile ↗
32ranked-venue papers
7as first author
25since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 21 · 2 first-author · 19 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 4 first-author · 2 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Precision-Scalable Accelerator for Compressive Hyperspectral Image Reconstruction with a Lightweight DUN
Shengzhi Qiang, Wendong Mao, Zhongfeng Wang 0001 |
ASP-DAC | 3 |
| 2026 | A Lightweight Algorithm-Hardware Co-design for Real-Time Video Frame Interpolation
Jisheng Zhang, Qiwei Dong, Wendong Mao, Zhongfeng Wang 0001 |
ISCAS | 4 |
| 2026 | Deep Lookup NetworkabstractConvolutional neural networks are constructed with massive operations with different types and are highly computationally intensive. Among these operations, multiplication operation is higher in computational complexity and usually requires more energy consumption with longer inference time than other operations, which hinders the deployment of convolutional neural networks on mobile devices. In many resource-limited edge devices, complicated operations can be calculated via lookup tables to reduce computational cost. Motivated by this, in this paper, we introduce a generic and efficient lookup operation which can be used as a basic operation for the construction of neural networks. Instead of calculating the multiplication of weights and activation values, simple yet efficient lookup operations are adopted to compute their responses. To enable end-to-end optimization of the lookup operation, we construct the lookup tables in a differentiable manner and propose several training strategies to promote their convergence. By replacing computationally expensive multiplication operations with our lookup operations, we develop lookup networks for the image classification, image super-resolution, and point cloud classification tasks. It is demonstrated that our lookup networks can benefit from the lookup operations to achieve higher efficiency in terms of energy consumption and inference speed while maintaining competitive performance to vanilla convolutional networks. Extensive experiments show that our lookup networks produce state-of-the-art performance on different tasks (both classification and regression tasks) and different data types (both images and point clouds). Yulan Guo, Longguang Wang, Wendong Mao, Yingqian Wang 0002, Li Liu 0002, Wei An 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | CDM-QTA: Quantized Training Acceleration for Efficient LoRA Fine-Tuning of Diffusion ModelabstractFine-tuning large diffusion models for custom applications demands substantial power and time, which poses significant challenges for efficient implementation on mobile devices. In this paper, we develop a novel training accelerator specifically for Low-Rank Adaptation (LoRA) of diffusion models, aiming to streamline the process and reduce computational complexity. By leveraging a fully quantized training scheme for LoRA fine-tuning, we achieve substantial reductions in memory usage and power consumption while maintaining high model fidelity. The proposed accelerator features flexible dataflow, enabling high utilization for irregular and variable tensor shapes during the LoRA process. Experimental results show up to 1.81× training speedup and 5.50× energy efficiency improvements compared to the baseline, with minimal impact on image generation quality. Jinming Lu, Minghao She, Wendong Mao, Zhongfeng Wang 0001 |
ISCAS | 3 |
| 2025 | Trio-ViT: Post-Training Quantization and Acceleration for Softmax-Free Efficient Vision TransformerabstractMotivated by the huge success of Transformers in the field of natural language processing (NLP), Vision Transformers (ViTs) have been rapidly developed and achieved remarkable performance in various computer vision tasks. However, their huge model sizes and intensive computations hinder ViTs’ deployment on embedded devices, calling for effective model compression methods, such as quantization. Unfortunately, due to the existence of hardware-unfriendly and quantization-sensitive non-linear operations, particularly Softmax, it is non-trivial to completely quantize all operations in ViTs, yielding either significant accuracy drops or non-negligible hardware costs. In response to challenges associated with standard ViTs, we focus our attention towards the quantization and acceleration for efficient ViTs, which not only eliminate the troublesome Softmax but also integrate linear attention with low computational complexity, and propose Trio-ViT accordingly. Specifically, at the algorithm level, we develop a tailored post-training quantization engine taking the unique activation distributions of Softmax-free efficient ViTs into full consideration, aiming to boost quantization accuracy. Furthermore, at the hardware level, we build an accelerator dedicated to the specific Convolution-Transformer hybrid architecture of efficient ViTs, thereby enhancing hardware efficiency. Extensive experimental results consistently prove the effectiveness of our Trio-ViT framework. Particularly, we can gain up to$\uparrow {3.6}\times $,$\uparrow {5.0}\times $, and$\uparrow {7.3}\times $FPS under comparable accuracy over state-of-the-art ViT accelerators, as well as$\uparrow {6.0}\times $,$\uparrow {1.5}\times $, and$\uparrow {2.1}\times $DSP efficiency. Codes are available athttps://github.com/shihuihong214/Trio-ViT. Huihong Shi, Haikuo Shao, Wendong Mao, Zhongfeng Wang 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2025 | A Unified Accelerator for All-in-One Image Restoration Based on Prompt Degradation LearningabstractAll-in-one image restoration (IR) recovers images from various unknown distortions by a single model, such as rain, haze, and blur. Transformer-based IR methods have significantly improved the visual effects of the restored images. However, deploying complex IR models on edge devices is challenging due to massive parameters and intensive computations. Moreover, existing accelerators are typically customized for a single task, resulting in severe resource underutilization when executing multiple tasks. Therefore, this paper develops an algorithm-hardware co-design framework to accelerate a novel CNN-Transformer cooperative model for multiple IR tasks. Firstly, on the algorithm level, an Efficient Restoration Foundational Model (ERFM) is proposed to recover corrupted images from various degradations with low model complexity. Secondly, to guide adaptive corruption removal, a novel prompt learning scheme is introduced to fuse context-related degradation cues and boost high-quality reconstruction. Thirdly, on the hardware level, an integer approximation method is proposed to avoid expensive hardware overhead caused by complex nonlinear operations, such as layer normalization and softmax while maintaining comparable IR quality. Moreover, a head stationary dataflow and softmax fusion mechanism are designed to reduce data movement and enhance on-chip resource utilization. Finally, an overall hardware architecture is developed and implemented in TSMC 28 nm CMOS technology. Experimental results show that our ERFM achieves better visual perception than other baselines on seven challenging IR tasks without task-specific fine-tuning. Moreover, compared to other accelerators for vision Transformers, our design can achieve 3.3$\times$and 3.7$\times$improvements in throughput and energy efficiency. Qiwei Dong, Wendong Mao, Zhongfeng Wang 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2025 | An Energy-Efficient Neuromorphic Accelerator Based on Deformable Spiking Transformer for Dynamic Vision SensorabstractNowadays, brain-inspired Spiking Neural Networks (SNNs) have been proven to effectively process Dynamic Vision Sensor (DVS) event data streams due to their event-driven computation and temporal characteristics. Among the various SNN models, the spiking-based Transformer has demonstrated superior performance. However, there is few effort focused on designing dedicated accelerators for spiking-based Transformers. In this paper, we propose an energy-efficient neuromorphic accelerator based on a novel spiking-based Transformer for DVS. At the algorithmic level, we propose a Deformable Spiking Transformer (DST), which incorporates novel Spike-Driven Deformable Attention modules to enhance feature extraction while reducing computational complexity. At the hardware level, we design an energy-efficient Spiking Convolution Core and Spiking Attention Core to efficiently support sparse spiking convolutions and deformable attention mechanisms in the DST. Moreover, to leverage dynamic sparsity and minimize processing latency, we introduce a sparse spiking computing flow that enables parallel processing of sparse computations in the spiking convolutions. Based on algorithm-hardware co-optimization, we develop an energy-efficient neuromorphic accelerator for DVS processing and implement it in TSMC 28nm CMOS technology. Experimental results show that the DST achieves promising accuracy on the DVS datasets while maintaining competitive inference latency. Compared to prior hardware designs for SNNs, the proposed accelerator has the highest peak throughput. In comparison to Transformer accelerators, it achieves at least$1.11 \times $improvement in energy efficiency. Wendong Mao, Zhongfeng Wang 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2025 | A CPU+FPGA OpenCL Heterogeneous Computing Platform for Multi-Kernel PipelineabstractOver the past decades, Field-Programmable Gate Arrays (FPGAs) have become a choice for heterogeneous computing due to their flexibility, energy efficiency, and processing speed. OpenCL is used in FPGA heterogeneous computing for its high-level abstraction and cross-platform compatibility. Previous works have introduced optimization techniques in OpenCL for FPGAs to leverage FPGA-specific advantages. However, the multi-kernel pipeline technique, which can raise throughput and resource utilization, has not performed well. This article presents a CPU+FPGA heterogeneous platform with a novel execution model to optimize multi-kernel pipeline. Firstly, we extend OpenCL by introducing new APIs and additional functions to represent the execution model. Secondly, a hardware-software co-scheduling scheme is employed to manage execution. Thirdly, we design a holistic development flow and toolkit to facilitate the deployment of algorithms on the platform or the integration of RTL IP cores to the OpenCL environment. We validate the platform using a Range Doppler algorithm. The proposed development flow and integrated toolchain enhance the efficiency of integrating traditional RTL IP cores into the OpenCL environment. Experimental results demonstrate that, with a comparable processing speed (averaging 95%) to traditional RTL implementations, the platform successfully establishes the multi-kernel pipelines. Leveraging the multi-kernel pipeline, the platform achieves a significant improvement in multi-frame processing speed compared to traditional OpenCL. Yuefei Wang, Wendong Mao, Lang Feng 0001, Jin Sha 0001, Zhongfeng Wang 0001 |
ACM Trans. Design Autom. Electr. Syst. | 2 |
| 2025 | An Energy-Efficient FPGA Accelerator for Swin TransformerabstractRecently, transformers have shown strong performance in tasks such as computer vision and natural language processing. Notably, Swin Transformer has gained significant attention for its low computational complexity and impressive performance in computer vision tasks, due to its window attention mechanism and hierarchical architecture. However, these features also make hardware deployment more complicated. In this brief, we present an energy-efficient field-programmable gate array (FPGA) accelerator for Swin Transformer to support the hierarchical architecture and execute the window attention. First, we introduce a systolic array with alterable datapath (SAAD) to conduct the window attention. Second, we split the patch merging operation and design a data rearrangement module, which reduces the computing latency induced by the data rearrangement in Swin Transformer. Third, we present a parallelized dual-array dataflow to support different computing operations in Swin Transformer. We implement the accelerator on the Xilinx XCZU19EG platform. The proposed architecture achieves a throughput per digital signal processing (DSP) of 0.630 giga operations per second (GOPS)/DSP, which is$1.94\times $higher than existing works. Yuefei Wang, Wendong Mao, Huihong Shi, Jin Sha 0001, Zhongfeng Wang 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2025 | VCNPU: An Algorithm-Hardware Co-Optimized Framework for Accelerating Neural Video CompressionabstractVideo compression is essential for storing and transmitting video content. Real-time decoding is indispensable for delivering a seamless user experience. Neural video compression (NVC) integrates traditional coding techniques with deep learning, resulting in impressive compression efficiency. However, the real-time deployment of advanced NVC models encounters challenges due to their high complexity and extensive off-chip memory access. This article presents a novel NVC accelerator, called video compression neural processing unit (VCNPU), via an algorithm-hardware co-design framework. First, at the algorithmic level, a reparameterizable video compression network (RepVCN) is proposed to aggregate multiscale features and boost video compression quality. RepVCN can be equivalently transformed into a streamlined structure without extra computations after training. Second, a mask-sharing pruning strategy is proposed to compress RepVCN in the fast transform domain. It effectively prevents the destruction of sparse patterns caused by model simplification, maintaining the model capacity. Third, at the hardware level, a reconfigurable sparse computing module is designed to flexibly support sparse fast convolutions and deconvolutions of the compact RepVCN. Besides, a hybrid layer fusion pipeline is advocated to reduce off-chip data communication caused by extensive motion and residual features. Finally, based on the joint optimization of computation and communication, our VCNPU is constructed to realize adaptive adjustments of various decoding qualities and is implemented under TSMC 28-nm CMOS technology. Extensive experiments demonstrate that our RepVCN provides superior coding quality over other video compression baselines. Meanwhile, our VCNPU achieves$6.7\times $improvements in throughput,$2.9\times $in area efficiency, and$4\times $in energy efficiency compared to prior video processors. Wendong Mao, Zhongfeng Wang 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2025 | DATA: A Memory-Efficient Deformable Transformer Accelerator via Neural Architecture SearchabstractTransformer is widely used in the field of artificial intelligence (AI) due to its excellent feature extraction capabilities. Its variant, the deformable transformer, is highly appreciated in autonomous driving and robotics since it employs a deformable attention mechanism to enhance feature extraction. However, due to its out-of-order memory access and data dependency, the deployment of the deformable transformer on mobile devices is much limited. To address these problems, a deformable attention transformer accelerator (DATA) is proposed in this work to speed up the processing by co-optimizing the algorithm and hardware. Specifically, we propose a memory-aware neural architecture search (NAS) method for deformable attention by constructing a continuous search space to automatically obtain a memory-efficient feature map slicing scheme. Based on the proposed slicing scheme, we design an efficient data flow to avoid the memory access conflict problem. In addition, a space-division multiplexing and time-division multiplexing hardware computing module is introduced to perform computations in the deformable attention layer, which greatly improves the utilization of hardware resources. Finally, the proposed accelerator is implemented on an Xilinx platform. In comparison, the proposed method achieves a maximum$2.42\times $improvement in computational efficiency over prior arts, and the memory access requirement is reduced to 12.5% of the baseline. Mingfan Zhao, Wendong Mao, Zhongfeng Wang 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2024 | A Computationally Efficient Neural Video Compression Accelerator Based on a Sparse CNN-Transformer Hybrid NetworkabstractVideo compression is widely used in digital television, surveillance systems, and virtual reality. Real-time video decoding is crucial in practical scenarios. Recently, neural video compression (NVC) combines traditional coding with deep learning, achieving impressive compression efficiency. Nevertheless, the NVC models involve high computational costs and complex memory access patterns, challenging real-time hardware implementations. To relieve this burden, we propose an algorithm and hardware co-design framework named NVCA for video decoding on resource-limited devices. Firstly, a CNN-Transformer hybrid network is developed to improve compression performance by capturing multi-scale non-local features. In addition, we propose a fast algorithm-based sparse strategy that leverages the dual advantages of pruning and fast algorithms, sufficiently reducing computational complexity while maintaining video compression efficiency. Secondly, a reconfigurable sparse computing core is designed to flexibly support sparse convolutions and deconvolutions based on the fast algorithm-based sparse strategy. Furthermore, a novel heterogeneous layer chaining dataflow is incorporated to reduce off-chip memory traffic stemming from extensive inter-frame motion and residual information. Thirdly, the overall architecture of NVCA is designed and synthesized in TSMC 28nm CMOS technology. Extensive experiments demonstrate that our design provides superior coding quality and up to 22.7x decoding speed improvements over other video compression designs. Meanwhile, our design achieves up to 2.2x improvements in energy efficiency compared to prior accelerators. Wendong Mao, Huihong Shi, Zhongfeng Wang 0001 |
DATE | 2 |
| 2024 | An FPGA-Based Reconfigurable Accelerator for Convolution-Transformer Hybrid EfficientViTabstractVision Transformers (ViTs) have achieved significant success in computer vision. However, their intensive computations and massive memory footprint challenge ViTs’ deployment on embedded devices, calling for efficient ViTs. Among them, EfficientViT, the state-of-the-art one, features a Convolution-Transformer hybrid architecture, enhancing both accuracy and hardware efficiency. Unfortunately, existing accelerators cannot fully exploit the hardware benefits of EfficientViT due to its unique architecture. In this paper, we propose an FPGA-based accelerator for EfficientViT to advance the hardware efficiency frontier of ViTs. Specifically, we design a reconfigurable architecture to efficiently support various operation types, including lightweight convolutions and attention, boosting hardware utilization. Additionally, we present a time-multiplexed and pipelined dataflow to facilitate both intra- and inter-layer fusions, reducing off-chip data access costs. Experimental results show that our accelerator achieves up to 780.2 GOPS in throughput and 105.1 GOPS/W in energy efficiency at 200MHz on the Xilinx ZCU102 FPGA, which significantly outperforms prior works. Haikuo Shao, Huihong Shi, Wendong Mao, Zhongfeng Wang 0001 |
ISCAS | 3 |
| 2024 | A Precision-Scalable Vision Accelerator for Robotic ApplicationsabstractRobot vision systems, by providing abundant and crucial environmental information, enable robots to intelligently perceive environment and make autonomous decisions. However, DNN-based models targeting the depth estimation task as well as other robotic applications tend to be computationally complex, bringing challenges to the efficient deployment on edge devices. In this paper, we propose a precision-scalable vision accelerator for robotic applications. Firstly, we develop a bit-level computing strategy to build the fundamental processing unit for precision scalability, reducing hardware complexity dramatically. Secondly, we present an efficient processing unit group with optimizations in parallelism scalability and overhead reduction. Thirdly, an energy-efficient architecture and the dataflow are proposed, enabling the accelerator to flexibly handle various operations in visual tasks targeting robotic applications like depth estimation. Our design is synthesized under TSMC 28nm CMOS technology. Experiments show that our design achieves a 6.05 TOPS/W energy efficiency as well as 2.23× and 2.13× area efficiency compared with the previous precision-scalable accelerators and the stereo vision accelerators respectively. Haoran Zeng, Wendong Mao, Zhongfeng Wang 0001 |
ISCAS | 2 |
| 2024 | NASA-F: FPGA-Oriented Search and Acceleration for Multiplication-Reduced Hybrid NetworksabstractThe costly multiplications challenge the deployment of modern deep neural networks (DNNs) on resource-constrained devices. To promote hardware efficiency, prior works have built multiplication-free models. However, they are generally inferior to their multiplication-based counterparts in accuracy, calling for multiplication-reduced hybrid models to marry the benefits of both approaches. To achieve this goal, recent works, i.e., NASA and NASA+, have developedNeuralArchitectureSearch (NAS) andAcceleration frameworks to search for and accelerate such hybrid models via a tailored differentiable NAS (DNAS) engine and dedicated ASIC-based accelerators. In this paper, we delve deeper into the inherent advantages of FPGAs and present an enhanced approach called NASA-F, which focuses on FPGA-oriented search and acceleration for hybrid models. Specifically,on the algorithm level, we develop a tailored one-shot supernet-based NAS engine to streamline the search for hybrid models, eliminating the need for executing NAS for each deployment as well as additional training/finetuning steps.On the hardware level, we develop a chunk-based accelerator to fully leverage the diverse hardware resources available on FPGAs for the acceleration of heterogeneous layers in hybrid models, aiming to enhance both hardware utilization and throughput. Extensive experimental results consistently validate the superiority of our NASA-F framework, e.g., we can gain$\uparrow 0.67\%$top-1 accuracy over the prior work NASA on CIFAR100 even without additional training steps for searched models. Additionally, we can achieve up to$\uparrow 1.86\times $throughout and$\uparrow 2.16\times $FPS with$\uparrow 0.39$% top-1 accuracy over the state-of-the-art multiplication-based system on Tiny-ImageNet. Codes are available athttps://github.com/shihuihong214/NASA-F. Huihong Shi, Yang Xu 0090, Yuefei Wang, Wendong Mao, Zhongfeng Wang 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2024 | P2-ViT: Power-of-Two Post-Training Quantization and Acceleration for Fully Quantized Vision TransformerabstractVision transformers (ViTs) have excelled in computer vision (CV) tasks but are memory-consuming and computation-intensive, challenging their deployment on resource-constrained devices. To tackle this limitation, prior works have explored ViT-tailored quantization algorithms but retained floating-point scaling factors, which yield nonnegligible requantization overhead, limiting ViTs’ hardware efficiency and motivating more hardware-friendly solutions. To this end, we propose P2-ViT, the first power-of-two (PoT) posttraining quantization (PTQ) and acceleration framework to accelerate fully quantized ViTs. Specifically, as for quantization, we explore a dedicated quantization scheme to effectively quantize ViTs with PoT scaling factors, thus minimizing the requantization overhead. Furthermore, we propose coarse-to-fine automatic mixed-precision quantization to enable better accuracy-efficiency tradeoffs. In terms of hardware, we develop a dedicated chunk-based accelerator featuring multiple tailored subprocessors to individually handle ViTs’ different types of operations, alleviating reconfigurable overhead. In addition, we design a tailored row-stationary dataflow to seize the pipeline processing opportunity introduced by our PoT scaling factors, thereby enhancing throughput. Extensive experiments consistently validate P2-ViT’s effectiveness. Particularly, we offer comparable or even superior quantization performance with PoT scaling factors when compared with the counterpart with floating-point scaling factors. Besides, we achieve up to$10.1\times $speedup and$36.8\times $energy saving over GPU’s Turing Tensor Cores, and up to$1.84\times $higher computation utilization efficiency against SOTA quantization-based ViT accelerators. Codes are available athttps://github.com/shihuihong214/P2-ViT. Huihong Shi, Xin Cheng 0015, Wendong Mao, Zhongfeng Wang 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2023 | S$$^2$$R: Exploring a Double-Win Transformer-Based Framework for Ideal and Blind Super-Resolution
Minghao She, Wendong Mao, Huihong Shi, Zhongfeng Wang 0001 |
ICANN (6) | 2 |
| 2023 | An Efficient Accelerator Based on Lightweight Deformable 3D-CNN for Video Super-ResolutionabstractDeformable convolutional networks (DCNs) have shown outstanding potential in video super-resolution with their powerful inter-frame feature alignment. However, deploying DCNs on resource-limited devices is challenging, due to their high computational complexity and irregular memory accesses. In this work, an algorithm-hardware co-optimization framework is proposed to accelerate the DCNs on field-programmable gate array (FPGA). Firstly, at the algorithm level, an anchor-based lightweight deformable network (ALDNet) is proposed to extract spatio-temporal information from the aligned features, boosting the visual effects with low model complexity. Secondly, to reduce intensive multiplications, an innovative shift-based deformable 3D convolution is developed using low-cost bit shifts and additions, maintaining comparable reconstruction quality. Thirdly, at the hardware level, a dedicated critical processing core, together with a block-level interleaving storage scheme, is presented to avoid dynamic and irregular memory accesses caused by the deformable convolutions. Finally, an overall architecture is designed to accelerate the ALDNet and implemented on an Intel Stratix 10GX platform. Experimental results demonstrate that the proposed design can provide significantly better visual perception than other FPGA-based super-resolution implementations. Meanwhile, compared with the prior hardware accelerators, our design can achieve$2.75\times $and$1.63\times $improvements in terms of throughput and energy efficiency, respectively. Wendong Mao, Zhongfeng Wang 0001 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 2 |
| 2023 | Intelligent Typography: Artistic Text Style Transfer for Complex Texture and StructureabstractText style transfer is an important task to render artistic texts from a reference image or style, and is widely desired in many visual creations. Previous works have brought some efficient methods for text style transfer, which facilitate users to design various artistic texts automatically. However, these works mainly focus on relatively simple text effects, and do not perform well on complex reference styles. In this paper, we propose a coarse-to-fine framework to generate exquisite texts with complex texture and structure in an unsupervised way, achieving real-time control of style scales (i.e., text stylistic degree or deformation degree). The key idea is to decouple the overall task into two steps, prototype generation and detail refinement, and explore delicate networks for each step to imitate the features at different levels. Based on this idea, in the first step, we present a novel pro-gen GAN to generate prototypes of artistic texts using the reference style, and develop a deformable module to empower the pro-gen GAN to continuously characterize the multi-scale shape features without network retraining. Furthermore, we propose a mix-attention training scheme for text style transfer, which can avoid artifacts and retain a clear text background. In the second step, we introduce two optimized networks for detail refinements. Experimental results show that the proposed method can synthesize exquisite stylized texts with complex reference styles, and surpass the state of the arts in texture reconstruction, contour imitation, and text image quality drastically. Wendong Mao, Shuai Yang 0001, Huihong Shi, Jiaying Liu 0001, Zhongfeng Wang 0001 |
IEEE Trans. Multim. | 1 |
| 2023 | FTA-GAN: A Computation-Efficient Accelerator for GANs With Fast Transformation AlgorithmabstractNowadays, generative adversarial network (GAN) is making continuous breakthroughs in many machine learning tasks. The popular GANs usually involve computation-intensive deconvolution operations, leading to limited real-time applications. Prior works have brought several accelerators for deconvolution, but all of them suffer from severe problems, such as computation imbalance and large memory requirements. In this article, we first introduce a novel fast transformation algorithm (FTA) for deconvolution computation, which well solves the computation imbalance problem and removes the extra memory requirement for overlapped partial sums. Besides, it can reduce the computation complexity for various types of deconvolutions significantly. Based on FTA, we develop a fast computing core (FCC) and the corresponding computing array so that the deconvolution can be efficiently computed. We next optimize the dataflow and storage scheme to further reuse on-chip memory and improve the computation efficiency. Finally, we present a computation-efficient hardware architecture for GANs and validate it on several GAN benchmarks, such as deep convolutional GAN (DCGAN), energy-based GAN (EBGAN), and Wasserstein GAN (WGAN). The experimental results show that our design can reach 2211 GOPS under 185-MHz working frequency on Intel Stratix 10SX field-programmable gate array (FPGA) board with satisfactory visual results. In brief, the proposed design can achieve more than 2× hardware efficiency improvement over previous designs, and it can reduce the storage requirement drastically. Wendong Mao, Peixiang Yang, Zhongfeng Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2022 | Accelerate Three-Dimensional Generative Adversarial Networks Using Fast AlgorithmabstractThree-dimensional generative adversarial networks (3D-GAN) have attracted widespread attention in three-dimension (3D) visual tasks. 3D deconvolution (DeConv), as an important computation of 3D-GAN, significantly increases computational complexity compared with 2D DeConv. 3D DeConv has become a bottleneck for the acceleration of 3D-GAN. Previous accelerators suffer from several problems, such as large memory requirements and resource underutilization. To handle the above issues, a fast algorithm for 3D DeConv (F3DC) is proposed in this paper. F3DC applies a fast algorithm to reduce the number of multiplications and achieves a significant algorithmic strength reduction. Besides, F3DC removes the extra memory requirement for overlapped partial sums and avoids computational imbalance to fully utilize resources. Moreover, we design an F3DC-based hardware architecture, which consists of four fast processing units (FPUs). Each FPU includes a pre-process module, a EWMM module and a post-process module for F3DC transformation. By implementing our design on the Xilinx VC709 platform for 3D-GAN, we achieve a throughput up to 1700 GOPS and 4× computational efficiency improvement compared with prior works. Ziqi Su, Wendong Mao, Zhongfeng Wang 0001, Jun Lin 0001 |
ISCAS | 2 |
| 2022 | A Reconfigurable Approach for Deconvolutional Network Acceleration with Fast AlgorithmabstractRecently, deconvolutional neural network (DeCNN) has attracted widespread attention in various applications. The deconvolution (DeConv), as the main operation in DeCNN, has become the bottleneck of acceleration, due to its high computational complexity. Previous works have introduced fast algorithms such as the cascaded fast FIR algorithm (CFFA) and the Winograd algorithm to reduce the computational complexity of DeConv for the applications on mobile devices. Since these fast algorithms need different computing parameters to accelerate various operations, directly applying these methods to process DeCNNs with different kernels usually causes limited flexibility. To address this problem, we propose a reconfigurable scheme based on the fast transformation algorithm (FTA) to accelerate multiple types of DeConvs, minimizing the hardware overhead for reconfigurability. Based on this scheme, a reconfigurable hardware architecture is developed to support several types of DeConvs. In addition, an adaptive dataflow is proposed to handle different convolutional layers. The presented design can support several types of operations and achieve up to 222.54 GOPS under 210 MHz on the Intel Arria 10SX FPGA platform, which shows our design can obtain better flexibility and computational efficiency compared with prior arts. Peixiang Yang, Wendong Mao, Zhongfeng Wang 0001, Jun Lin 0001 |
ISCAS | 2 |
| 2022 | An Efficient FPGA-based Accelerator for Deep ForestabstractDeep Forest is a prominent machine learning algorithm known for its high accuracy in forecasting. Compared with deep neural networks, Deep Forest has almost no multiplication operations and has better performance on small datasets. However, due to the deep structure and large forest quantity, it suffers from large amounts of calculation and memory consumption. In this paper, an efficient hardware accelerator is proposed for deep forest models, which is also the first work to implement Deep Forest on FPGA. Firstly, a delicate node computing unit (NCU) is designed to improve inference speed. Secondly, based on NCU, an efficient architecture and an adaptive dataflow are proposed, in order to alleviate the problem of node computing imbalance in the classification process. Moreover, an optimized storage scheme in this design also improves hardware utilization and power efficiency. The proposed design is implemented on an FPGA board, Intel Stratix V, and it is evaluated by two typical datasets, ADULT and Face Mask Detection. The experimental results show that the proposed design can achieve around $40 \times$ speedup compared to that on a 40 cores high performance x86 CPU. Jiapeng Luo, Wendong Mao, Zhongfeng Wang 0001 |
ISCAS | 3 |
| 2022 | A robust framework for multi-view stereopsis
Wendong Mao, Mingjie Wang 0002, Hui Huang 0004, Minglun Gong |
Vis. Comput. | 1 |
| 2021 | LITNet: A Light-weight Image Transform Net for Image Style TransferabstractRecently, style transfer networks have received widespread attention in computer vision field, which combine stylistic features from a style image and content information from a content image to generate an output. However, the high-resolution synthesized outputs come at the cost of intensive computation, making it difficult to employ style transfer networks on embedded devices with limited computational resources. To address this issue, we propose a compression algorithm for one of the influential CNN-based style transfer networks, which is named Image Transform Net (ITNet), and gain a Light-weight Image Transform Net (LITNet) accordingly. To improve the performance of ITNet, normalization layers and the structure of upsampling blocks are modified, and depthwise separable convolutions combined with width multiplier are employed to obtain a brand-new light-weight network. However, since the representation ability of light-weight networks is too weak for unsupervised learning tasks such as style transfer, directly using the above techniques to compress the model leads to unstable training processes and yields poor outputs. To solve this problem, a novel distillation loss is proposed to convert unsupervised learning into supervised learning. Besides, the weights between the original losses and the distillation loss are balanced for better visual results. Experimental results demonstrate the effectiveness of our LITNet. With minimal visual quality degradation, the light-weight network can achieve more than 67 × compression in model size and 63× reduction in FLOPs. Codes and pre-trained models are available at https://github.com/shihuihong214/LITNet. Huihong Shi, Wendong Mao, Zhongfeng Wang 0001 |
IJCNN | 2 |
| 2020 | No-reference image sharpness assessment based on discrepancy measures of structural degradation
Hao Cai 0004, Mingjie Wang 0002, Wendong Mao, Minglun Gong |
J. Vis. Commun. Image Represent. | 3 |
| 2020 | BSD-GAN: Branched Generative Adversarial Network for Scale-Disentangled Representation Learning and Image SynthesisabstractWe introduce BSD-GAN, a novel multi-branch and scale-disentangled training method which enables unconditional Generative Adversarial Networks (GANs) to learn image representations at multiple scales, benefiting a wide range of generation and editing tasks. The key feature of BSD-GAN is that it is trained in multiple branches, progressively covering both the breadth and depth of the network, as resolutions of the training images increase to reveal finer-scale features. Specifically, each noise vector, as input to the generator network of BSD-GAN, is deliberately split into several sub-vectors, each corresponding to, and is trained to learn, image representations at a particular scale. During training, we progressively "de-freeze" the sub-vectors, one at a time, as a new set of higher-resolution images is employed for training and more network layers are added. A consequence of such an explicit sub-vector designation is that we can directly manipulate and even combine latent (sub-vector) codes which model different feature scales. Extensive experiments demonstrate the effectiveness of our training method in scale-disentangled learning of image representations and synthesis of novel image contents, without any extra labels and without compromising quality of the synthesized high-resolution images. We further demonstrate several image generation and manipulation applications enabled or improved by BSD-GAN. Zili Yi, Hao Cai 0004, Wendong Mao, Minglun Gong, Hao (Richard) Zhang |
IEEE Trans. Image Process. | 4 |
| 2020 | F-DNA: Fast Convolution Architecture for Deconvolutional Network AccelerationabstractDeconvolutional neural network (DeCNN), such as fully convolutional network (FCN) and generative adversarial network (GAN), has shown great potential in various vision tasks. Convolution and deconvolution, the two major operations of DeCNN, both require real-time hardware acceleration. However, some previous designs for deconvolutions require large memory for overlapped results, while others incur computation imbalance and cause resource underutilization. In this article, we propose an efficient method to convert deconvolutions to convolutions, which enables balanced computations to make full use of processing elements. Based on the fast FIR algorithm, a reconfigurable conv-deconv unit (RCU) with low complexity is designed, which can support various types of convolutions and deconvolutions. By exploiting the computing characteristics of RCUs, a computation-balance scheme is developed to eliminate large memory requirements caused by overlapped results. In addition, a fast convolution architecture for deconvolutional network acceleration (F-DNA) is proposed. The dataflow of F-DNA improves the computation efficiency through input data reuse. The architecture is implemented on Xilinx Virtex-UltraScale, for two typical DeCNNs, DCGAN and FSRCNN. Implementation results show that the proposed design outperforms existing works significantly, particularly in terms of computation efficiency and memory requirements. Wendong Mao, Jun Lin 0001, Zhongfeng Wang 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2019 | A Global-Matching Framework for Multi-View Stereopsis
Wendong Mao, Minglun Gong, Xin Huang 0030, Hao Cai 0004, Zili Yi |
CAIP (1) | 1 |
| 2019 | Methodology for Efficient Reconfigurable Architecture of Generative Neural NetworkabstractGenerative neural networks have been developing rapidly in the field of deep learning nowadays. Generative models have obtained much popularity in various applications such as image generation, reading comprehension and style transfer. Convolutional (CONV) and deconvolutional (DeCONV) layers are typical components of generative neural networks. The use of traditional convolution accelerators will cause problems of overlapping and resource under-utilization while doing deconvolutions. There is little research on acceleration of deconvolution implementations. In this paper, we propose efficient reconfigurable architecture of generative neural networks. Firstly, the fast reconfigurable unit (FRU) based on cascaded fast FIR algorithm (CFFA) is proposed to support both convolutions and deconvolutions. The problems of overlapping and resource under-utilization are solved. Secondly, the reconfigurable architecture on the basis of FRUs for CONV and DeCONV layers is proposed accordingly. Thirdly, a novel shift scale quantization method is proposed to uniformly quantize CONV and DeCONV layers. Only integer computations are required with the quantization method. Finally, we choose a typical generative neural network and implement it on Xilinx Zynq ZC706. It is estimated that the performance reaches 62.85 GOPS under 330MHz working frequency on Xilinx ZC706. In brief, the proposed design outperforms existing works significantly, particularly surpasses related reconfigurable design by more than 20 times in terms of performance density. Wendong Mao, Jichen Wang, Jun Lin 0001, Zhongfeng Wang 0001 |
ISCAS | 1 |
| 2019 | Semi-Dense Stereo Matching Using Dual CNNsabstractA robust solution for semi-dense stereo matching is presented. It utilizes two CNN models for computing stereo matching cost and performing confidence-based filtering, respectively. Compared to existing CNNs-based matching cost generation approaches, our method feeds additional global information into the network so that the learned model can better handle challenging cases, such as lighting changes and lack of textures. Through utilizing non-parametric transforms, our method is also more self-reliant than most existing semi-dense stereo approaches, which rely highly on the adjustment of parameters. The experimental results based on Middlebury Stereo dataset demonstrate that the proposed approach outperforms the state-of-the-art semi-dense stereo approaches. Wendong Mao, Mingjie Wang 0002, Jun Zhou 0023, Minglun Gong |
WACV | 1 |
| 2019 | Multi-Scale Convolution Aggregation and Stochastic Feature Reuse for DenseNetsabstractRecently, Convolution Neural Networks (CNNs) obtained huge success in numerous vision tasks. In particular, DenseNets have demonstrated that feature reuse via dense skip connections can effectively alleviate the difficulty of training very deep networks and that reusing features generated by the initial layers in all subsequent layers has strong impact on performance. To feed even richer information into the network, a novel adaptive Multi-scale Convolution Aggregation module is presented in this paper. Composed of layers for multi-scale convolutions, trainable cross-scale aggregation, maxout, and concatenation, this module is highly non-linear and can boost the accuracy of DenseNet while using much fewer parameters. In addition, due to high model complexity, the network with extremely dense feature reuse is prone to overfitting. To address this problem, a regularization method named Stochastic Feature Reuse is also presented. Through randomly dropping a set of feature maps to be reused for each mini-batch during the training phase, this regularization method reduces training costs and prevents co-adaptation. Experimental results on CIFAR-10, CIFAR-100 and SVHN benchmarks demonstrated the effectiveness of the proposed methods. Mingjie Wang 0002, Jun Zhou 0023, Wendong Mao, Minglun Gong |
WACV | 3 |