Xuegong Zhou

dblp:15/645 · DBLP profile ↗
← Back
31ranked-venue papers
6as first author
10since 2021 · last 2026
0000-0003-4178-4094ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 28 · 5 first-author · 9 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 FHPSAC: FPGA-based High-Parallelism SAC Accelerator
abstract
Reinforcement learning (RL) enables autonomous decision-making in applications such as robotics and control, and Soft Actor-Critic (SAC) is a leading model-free algorithm for continuous tasks. However, SAC’s small-batch training updates and fine-grained computation lead to heavy scheduling and kernel-launch overheads on GPU, limiting efficiency. In this work, we present FHPSAC (FPGA-based High-Parallelism SAC Accelerator), the first FPGA-accelerated architecture dedicated to SAC training. First, we propose a hardware–software co-designed on-chip memory hierarchy to statically partition and allocate SAC’s training data for conflict-free parallel access. Second, we build a high-parallelism accelerator with a tensor core for GEMM (General Matrix Multiply) and a lightweight unit for irregular elementwise/reduction kernels. Finally, we implement the full system on a Xilinx XCVU9P FPGA and demonstrate significant speedup with low power. Experimental results show that FHPSAC obtains 5.33–14.85× speedup compared with the Intel Xeon Gold 6130 CPU, while outperforming an NVIDIA A100-SXM4 GPU by 2.90–10.43× in training latency with an average power of 40.17 W. FHPSAC substantially reduces SAC training latency, providing a computational foundation for large-scale SAC deployments.
Jiabin Xu, Wang Fan, Xuegong Zhou, Wei Cao 0002, Fengzhe Zhang, Fan Zhang 0044, Xinsheng Yu 0001
FCCM3
2026 An Agile Deployment System for Password Recovery on FPGA
abstract
Hardware-based acceleration of password recovery remains a pressing challenge, as CPUs and GPUs struggle to efficiently process modern cryptographic primitives. Although FPGAs offer superior performance-per-watt, their widespread adoption is limited by long development cycles, manual optimization, and the absence of an end-to-end deployment framework that jointly accelerates password generation and verification. To address this gap, we propose the first agile, end-to-end FPGA-based password recovery system that unifies deep-learning-driven password generation and cryptographic verification within a single deployment workflow. The framework consists of: (1) a customized Neural Processing Unit (NPU) that accelerates GAN-based password generation models such as PassGAN; (2) an automated, template-based accelerator generator for verification kernels, built on reusable Chisel hardware primitives; and (3) a multi-objective Design Space Exploration (DSE) engine that co-optimizes kernel-level parameters (e.g., loop unrolling) and system-level parallelism to determine globally optimal FPGA configurations. We deploy the system on a heterogeneous platform combining a Zynq MPSoC with dual Virtex UltraScale+ FPGAs. Experimental results show that the NPU outperforms an NVIDIA Tesla V100 by 82.16% in PassGAN inference throughput. The full system achieves 1.90× higher end-to-end throughput and 2.32× better energy efficiency than GPU-based implementations, and delivers an average 32.58% speedup over state-of-the-art FPGA-only verification designs. These results demonstrate the practicality and scalability of our architecture for real-world password recovery workflows.
Liming Deng, Guowei Zhu, Xitian Fan, Guangwei Xie, Mingqian Sun, Xuegong Zhou, Wei Cao 0002, Fan Zhang 0044, Xinsheng Yu 0001
IEEE Trans. Computers6
2025 Agile Design Flow for Cryptographic Hardware Accelerators
abstract
This paper presents an agile design flow for cryptographic hardware accelerators, which supports the automated generation of RTL code for efficient hardware accelerators intended for FPGA deployment. An automated design method for hardware accelerators based on domain-specific hardware design templates (DST) is proposed. Based on the characteristics of cryptographic algorithms, we designed a parameterized DST and constructed a corresponding specific hardware operator library (SHWOL). We adopted a novel operator matching strategy based on subgraph isomorphism to realize the mapping of user input algorithms in the operator library, thereby deriving the DST parameters for code generation thus finalizing the RTL code generation with design space exploration (DSE). When implemented on an FPGA, compared with the existing high-level synthesis (HLS) tools, the code generated by our proposed design flow has an LUT efficiency (throughput/number of LUTs) of up to 483× and energy efficiency (throughput/power) of up to 676×.
Liming Deng, Guowei Zhu, Wei Cao 0002, Xitian Fan, Xuegong Zhou
ICCD5
2025 AHCA: Agile Design Framework for Hashcat Acceleration Based on FPGA
abstract
This article presents AHCA, an agile design framework for Field Programmable Gate Array (FPGA)-based Hashcat acceleration that automates the generation of optimized register transfer level (RTL) code. Our approach is centered on a proposed automated design method using a parameterized domain-specific template (DST) and a specific hardware operator library. The framework analyzes an algorithm’s graph to extract key hardware operators and their interconnection network. To support diverse user inputs, we introduce an innovative operator matching strategy using subgraph isomorphism, which maps algorithms to our operator library. This matched information, combined with design space exploration (DSE), is used to configure the DST and generate the final RTL code, avoiding redundancy for previously implemented algorithms. Compared to state-of-the-art high-level synthesis (HLS) tools, AHCA demonstrates a maximum performance enhancement of 797×, a Look-Up table (LUT) efficiency improvement of up to 105×, and an energy efficiency gain of up to 676×. When deployed on an FPGA for password cracking, the AHCA-generated hardware achieves a 63.95× enhancement in energy efficiency over CPUs and a 4.71× improvement over GPUs.
Liming Deng, Guowei Zhu, Xitian Fan, Wei Cao 0002, Xuegong Zhou, Fan Zhang 0044, Shaobo Yang
ACM Trans. Reconfigurable Technol. Syst.5
2025 DVHetero: A Framework for Designing and Validating Heterogeneous SoC with RISC-V Processor and CGRA
abstract
CGRA, as a coprocessor in SoCs, has been widely studied. However, there is limited research on how to efficiently debug and verify SoCs composed of CGRAs and processors during the design process. To address this gap, we introduce DVHetero. DVHetero incorporates a simulation and validation framework, SoCDiff, which enables comprehensive SoC simulation, debugging, and rapid error localization. Using this verification framework, we successfully implemented and validated the entire SoC. The SoC includes a Chisel-based CGRA generator and provides a pipelined CGRA architecture template. The CGRA is tightly integrated with the RISC-V processor, allowing for efficient DMA-based data transfer and MMIO support within the SoC. The pipelined CGRA architecture generated by DVHetero shows a 1.27× improvement in area efficiency and a 10.54× increase in mapping speed compared to the state-of-the-art CGRA framework, HierCGRA. Additionally, compared to state-of-the-art CGRA-SoC systems FDRA, DVHetero demonstrates a 1.67× increase in execution speed and a 4.34× improvement in area efficiency.
Guowei Zhu, Liming Deng, Kaisen Zhang, Wang Fan, Boyin Jin, Wei Cao 0002, Fengzhe Zhang, Xuegong Zhou, Fan Zhang 0044, Xinsheng Yu 0001
ACM Trans. Reconfigurable Technol. Syst.8
2024 An Efficient Reinforcement Learning Based Framework for Exploring Logic Synthesis
abstract
Logic synthesis is a crucial step in electronic design automation tools. The rapid developments of reinforcement learning (RL) have enabled the automated exploration of logic synthesis. Existing RL based methods may lead to data inefficiency, and the exploration approaches for FPGA and ASIC technology mapping in recent works lack the flexibility of the learning process. This work proposes ESE, a reinforcement learning based framework to efficiently learn the logic synthesis process. The framework supports the modeling of logic optimization and technology mapping for FPGA and ASIC. The optimization for the execution time of the synthesis script is also considered. For the modeling of FPGA mapping, the logic optimization and technology mapping are combined to be learned in a flexible way. For the modeling of ASIC mapping, the standard cell based optimization and LUT optimization operations are incorporated into the ASIC synthesis flow. To improve the utilization of samples, the Proximal Policy Optimization model is adopted. Furthermore, the framework is enhanced by supporting MIG based synthesis exploration. Experiments show that for FPGA technology mapping on the VTR benchmark, the average LUT-Level-Product and script runtime are improved by more than 18.3% and 12.4% respectively than previous works. For ASIC mapping on the EPFL benchmark, the average Area-Delay-Product is improved by 14.5%.
Xuegong Zhou, Hao Zhou 0008, Lingli Wang
ACM Trans. Design Autom. Electr. Syst.2
2023 An Optimized GIB Routing Architecture with Bent Wires for FPGA
abstract
Field-programmable gate arrays (FGPAs) are widely used because of the superiority in flexibility and lower non-recurring engineering cost. How to optimize the routing architecture is a key problem for FPGA architects because it has a large impact on FPGA area, delay, and routability. In academia, the routing architecture is mainly based on the connection blocks (CBs) and switch blocks (SBs), whereas most research has focused on SB architectures, such as Wilton, Universal, and Disjoint SB patterns. In this article, we propose a novel unidirectional routing architecture—general interconnection block (GIB)—to improve FPGA performance. With the GIB architecture, logic block (LB) pins can directly connect with the adjacent GIBs without programmable switches. Inside a GIB, LB pins can connect to the routing channel tracks on the four sides of a GIB. In particular, the logic pins from different neighboring LBs that connect to the same GIB can connect with each other with only one programmable switch. In addition, we enhance VTR to support the GIB with bent wires and develop a searching framework based on the simulated annealing algorithm to search for a near-optimal distribution of wire types. We evaluate the GIB architecture on VTR 8 with the provided benchmark circuits. The experimental results show that the GIB architecture with length-4 wires can achieve 9.5% improvement on the critical path delay and 11.1% improvement on the area-delay product compared to the VTR CB-SB architecture with length-4 wires. After exploring mixed wire types, the optimized GIB architecture can further improve the delay by 16.4% and area-delay product by 17.1% compared to the CB-SB architecture with length-4 wires.
Kaichuang Shi, Xuegong Zhou, Hao Zhou 0008, Lingli Wang
ACM Trans. Reconfigurable Technol. Syst.2
2022 Efficient Reinforcement Learning Framework for Automated Logic Synthesis Exploration
abstract
Logic synthesis is a crucial step in electronic design automation tools for integrated circuit design. In recent years, the development of reinforcement learning (RL) has enabled the designers to automatically explore the logic synthesis process. Existing RL based methods typically use conventional on-policy models, which leads to data inefficiency. Moreover, the exploration approach for FPGA technology mapping in recent works lacks the flexibility of the learning process. In this work, we propose ESE, a reinforcement learning based framework to efficiently learn the logic synthesis process. The framework supports the modeling for both the logic optimization and the FPGA technology mapping. The reward functions and terminal conditions in the RL environment are designed to efficiently guide the optimization of the metrics and execution time. For the modeling of FPGA mapping, the logic optimization and technology mapping are combined to be learned in a flexible way. Moreover, the Proximal Policy Optimization model is adopted to improve the utilization of samples. The proposed framework is evaluated on several common benchmarks. For the logic optimization on the EPFL benchmark, compared with previous works, the proposed method obtains an 11.3% improvement in the average quality (node-level-product) and reduces the execution time by 13.7%. For the FPGA technology mapping on the VTR benchmark, our method improves the average quality (LUT-level-product) by 14.8%, and reduces the execution time by 14.4% compared with the recent work.
Xuegong Zhou, Hao Zhou 0008, Lingli Wang
FPT2
2021 FastCGRA: A Modeling, Evaluation, and Exploration Platform for Large-Scale Coarse-Grained Reconfigurable Arrays
abstract
Coarse-Grained Reconfigurable Arrays (CGRAs) provide sufficient flexibility in domain-specific applications with high hardware efficiency, which make CGRAs suitable for fast-evolving fields such as neural network acceleration and edge computing. To meet the requirement of the fast evolution, we propose FastCGRA, the modeling, mapping, and exploration platform for large-scale CGRAs. FastCGRA supports hierarchical architecture description and automatic switch module generation. Connectivity-aware packing and graph partition algorithms are designed to reduce the complexity of placement and routing. The graph homomorphism placement algorithm in FastCGRA enables efficient placement on large-scale CGRAs. The packing and placement algorithms cooperate with a negotiation-based routing algorithm to form an integral mapping procedure. FastCGRA can support the modeling and mapping of large-scale CGRAs with significantly higher placement and routing efficiency than existing platforms. The automatic switch module generation method can reduce the complexity of CGRA interconnection design. With these features, FastCGRA can boost the exploration of large-scale CGRAs.
Su Zheng, Kaisen Zhang, Yaoguang Tian, Wenbo Yin, Lingli Wang, Xuegong Zhou
FPT6
2021 MRI-based brain tumor segmentation using FPGA-accelerated neural network
abstract
BACKGROUND: Brain tumor segmentation is a challenging problem in medical image processing and analysis. It is a very time-consuming and error-prone task. In order to reduce the burden on physicians and improve the segmentation accuracy, the computer-aided detection (CAD) systems need to be developed. Due to the powerful feature learning ability of the deep learning technology, many deep learning-based methods have been applied to the brain tumor segmentation CAD systems and achieved satisfactory accuracy. However, deep learning neural networks have high computational complexity, and the brain tumor segmentation process consumes significant time. Therefore, in order to achieve the high segmentation accuracy of brain tumors and obtain the segmentation results efficiently, it is very demanding to speed up the segmentation process of brain tumors. RESULTS: Compared with traditional computing platforms, the proposed FPGA accelerator has greatly improved the speed and the power consumption. Based on the BraTS19 and BraTS20 dataset, our FPGA-based brain tumor segmentation accelerator is 5.21 and 44.47 times faster than the TITAN V GPU and the Xeon CPU. In addition, by comparing energy efficiency, our design can achieve 11.22 and 82.33 times energy efficiency than GPU and CPU, respectively. CONCLUSION: We quantize and retrain the neural network for brain tumor segmentation and merge batch normalization layers to reduce the parameter size and computational complexity. The FPGA-based brain tumor segmentation accelerator is designed to map the quantized neural network model. The accelerator can increase the segmentation speed and reduce the power consumption on the basis of ensuring high accuracy which provides a new direction for the automatic segmentation and remote diagnosis of brain tumors.
Siyu Xiong, Guoqing Wu 0003, Xitian Fan, Zhongcheng Huang, Wei Cao 0002, Xuegong Zhou, Shijin Ding, Jinhua Yu 0003, Lingli Wang, Zhifeng Shi
BMC Bioinform.7
2020 Fast Exact NPN Classification by Co-Designing Canonical Form and Its Computation Algorithm
abstract
NPN classification of Boolean functions is a powerful technique used in many practical applications, including logic synthesis, technology mapping, architecture exploration, circuit restructuring, and approximate logic synthesis. Computing the canonical form of a function is the most common approach to NPN classification. Exact classification of practical functions is an open problem because there are difficult functions beyond the capability of the state-of-the-art exact algorithms, which may take several months to compute a canonical form. This article proposes a new approach to exact NPN classification, in which a series of canonical forms and the algorithms to compute them are designed together. As a result, the runtime of the exact classification for difficult functions is effectively controlled by making both representation and computation cost-aware. Experimental results show that the proposed algorithm can perform exact classification of the worst-case 16-input functions in less than 3 minutes. This indicates that, for the first time, the problem of exact classification can be effectively solved for any Boolean functions with up to 16 inputs arising in practical applications.
Xuegong Zhou, Lingli Wang, Alan Mishchenko
IEEE Trans. Computers1
2019 ARBSA: Adaptive Range-Based Simulated Annealing for FPGA Placement
abstract
Placement has always been the most time-consuming part of the field programmable gate array (FPGA) compilation flow. Conventional simulated annealing has been unable to keep pace with ever increasing sizes of designs and FPGA chip resources. Without utilizing information of the circuit topology, it relies on large amounts of random swap operations, which are time-costly. This paper proposes an adaptive range-based algorithm to improve the behavior of swap operations and limit the swap distances by introducing the concept of range-limiting strategy for nets. It avoids unnecessary design space exploration, and thus can converge to near-optimal solutions much more quickly. The experimental results are based on the Titan benchmarks, which contain 4K to 30K blocks, including logic array blocks, inputs and outputs, digital signal processors, and random access memories. This approach achieves$2.82\boldsymbol \times $speed up, 4.8% reduction on wire length, 4.1% improvement on critical path compared with the SA from VTR with wire length-driven optimization, and$1.78\boldsymbol \times $speed up, 10% reduction on wire length, 2% reduction on critical path with path timing-driven optimization. It also manifests better scalability on larger benchmarks.
Junqi Yuan, Jialing Chen, Lingli Wang, Xuegong Zhou, Yinshui Xia
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2019 Fast Adjustable NPN Classification Using Generalized Symmetries
abstract
NPN classification of Boolean functions is a powerful technique used in many logic synthesis and technology mapping tools in both standard cell and FPGA design flows. Computing the canonical form is the most common approach of Boolean function classification. This article proposes two different hybrid NPN canonical forms and a new algorithm to compute them. By exploiting symmetries under different phase assignment as well as higher-order symmetries, the search space of NPN canonical form computation is pruned and the runtime is dramatically reduced. Nevertheless, the runtime for some difficult functions remains high. Fast heuristic method can be used for such functions to compute semi-canonical forms in a reasonable time. The proposed algorithm can be adjusted to be a slow exact algorithm or a fast heuristic algorithm with lower quality. For exact NPN classification, the proposed algorithm is 40× faster than state-of-the-art. For heuristic classification, the proposed algorithm has similar performance as state-of-the-art with a possibility to trade runtime for quality.
Xuegong Zhou, Lingli Wang, Alan Mishchenko
ACM Trans. Reconfigurable Technol. Syst.1
2018 Fast Adjustable NPN Classification using Generalized Symmetries
abstract
NPN classification of Boolean functions is a powerful technique used in many logic synthesis and technology mapping tools in FPGA design flows. Computing the canonical form of a function is the most common approach of Boolean function classification. In this paper, a novel algorithm for computing NPN canonical form is proposed. By exploiting symmetries under different phase assignments and higher-order symmetries of Boolean functions, the search space of NPN canonical form computation is pruned and the runtime is dramatically reduced. The algorithm can be adjusted to be a slow exact algorithm or a fast heuristic algorithm with lower quality. For exact classification, the proposed algorithm achieves a 30× speedup compared to a state-of-the-art algorithm. For heuristic classification, the proposed algorithm has similar performance as the state-of-the-art algorithm with a possibility to trade runtime for quality.
Xuegong Zhou, Lingli Wang, Peiyi Zhao, Alan Mishchenko
FPL1
2017 Accelerating low bit-width convolutional neural networks with embedded FPGA
abstract
Convolutional Neural Networks (CNNs) can achieve high classification accuracy while they require complex computation. Binarized Neural Networks (BNNs) with binarized weights and activations can simplify computation but suffer from obvious accuracy loss. In this paper, low bit-width CNNs, BNNs and standard CNNs are compared to show that low bit-width CNNs is better suited for embedded systems. An architecture based on the two-stage arithmetic unit (TSAU) as the basic processing element is proposed to process each layer iteratively for low bit-width CNN accelerators. Then the DoReFa-Net which is trained with weights and activations represented in 1 bit and 2 bits respectively is implemented on Zynq XC7Z020 FPGA with a 410.2 GOPS performance. The accelerator can meet the real-time requirement of embedded applications with a 106 FPS throughput and a 73.1% top-5 accuracy on the ImageNet dataset. The accelerator outperforms existing FPGA-based CNN accelerators in the tradeoff among accuracy, energy and resource efficiency.
Wei Cao 0002, Xuegong Zhou, Lingli Wang
FPL4
2017 A scalable hybrid architecture for high performance data-parallel applications
abstract
This paper presents a scalable hybrid architecture for high performance data-parallel applications on tightly coupled shared-memory CPU-FPGA systems such as the Xilinx Zynq SoC. The aims of the proposed architecture are: 1)to simplify the development of hardware acceleration for dataparallel applications; 2)to reach the performance limit caused by memory access and/or hardware resource available on an FPGA; 3)to reduce the overhead caused by task scheduling and device drivers. The proposed architecture can be used as a generic template to implement data-parallel applications. Each task in an application is mapped to one hardware accelerator, which is called “kernel”. Several identical instances of each hardware kernel execute concurrently to provide parallelism. By deploying the maximum number of instances of the hardware kernel, we make full use of the bandwidth of memory access and the resources available on the FPGA. In order to improve performance further, task scheduling and device drivers are implemented as a hardware scheduler called DmaScheduler on FPGA hardware. Experimental results show 2.93x-51.25x speedup on Zynq FPGA for applications of image processing, Black Scholes option pricing, matrix multiplication and clustering algorithm, compared with existing FPGA implementations.
Moucheng Yang, Jifang Jin, Xuegong Zhou, Lingli Wang
FPT4
2017 RBSA: Range-based simulated annealing for FPGA placement
abstract
Placement has always been the most time-consuming part in the FPGA compilation flow. Traditional simulated annealing has been unable to keep pace with ever increasing sizes of designs and FPGA chip resources. Without utilizing information of the circuit topology, it relies on large amounts of random swap operations, which are time-costly. This paper proposes a range-based algorithm to improve the behavior of swap operations and limit the swap distances by introducing the concept of range limiting for nets. It avoids unnecessary design space exploration, and thus can converge to near-optimal solutions much more quickly. The Titan benchmarks we have tested on contains 4K to 30K blocks, which include LABs, IOs, DSPs and RAMs. This approach achieves 2.05X speed up on average compared with the SA from VTR while preserving the placement quality of both the wire length and critical path. It also manifests better scalability towards larger benchmarks.
Junqi Yuan, Lingli Wang, Xuegong Zhou, Yinshui Xia
FPT3
2016 A high performance FPGA-based accelerator for large-scale convolutional neural networks
abstract
In recent years, convolutional neural networks (CNNs) based machine learning algorithms have been widely applied in computer vision applications. However, for large-scale CNNs, the computation-intensive, memory-intensive and resource-consuming features have brought many challenges to CNN implementations. This work proposes an end-to-end FPGA-based CNN accelerator with all the layers mapped on one chip so that different layers can work concurrently in a pipelined structure to increase the throughput. A methodology which can find the optimized parallelism strategy for each layer is proposed to achieve high throughput and high resource utilization. In addition, a batch-based computing method is implemented and applied on fully connected layers (FC layers) to increase the memory bandwidth utilization due to the memory-intensive feature. Further, by applying two different computing patterns on FC layers, the required on-chip buffers can be reduced significantly. As a case study, a state-of-the-art large-scale CNN, AlexNet, is implemented on Xilinx VC709. It can achieve a peak performance of 565.94 GOP/s and 391 FPS under 156MHz clock frequency which outperforms previous approaches.
Huimin Li 0005, Xitian Fan, Wei Cao 0002, Xuegong Zhou, Lingli Wang
FPL5
2016 High performance Deformable Part Model accelerator based on FPGA
abstract
Deformable Part Model (DPM) is one of the best algorithms for image-based object detection. However, the high computation intensity leads to relatively long detecting time. Even with the powerful CPU or GPU computing system, it is still too slow for practical applications. To solve this problem, this paper proposes a high performance DPM accelerator based on FPGA, where a dedicated JPEG decoder is integrated to process the images with the 1080p JPEG format. Pipelined architecture and data reuse strategies are developed to achieve the high throughput and energy efficiency. The proposed accelerator can process input images with 22 fps at the frequency of 156MHz on Xilinx VC709 board, which outperforms previous approaches.
Qi Zhan, Wei Cao 0002, Xuegong Zhou, Lingli Wang
FPT5
2015 UniStream: A unified stream architecture combining configuration and data processing
abstract
This paper proposes UniStream, a unified stream architecture based on point-to-point stream channels combining both bitstream configuration and data stream processing. In addition, unified APIs are provided to support bitstream configuration and data stream processing, as well as the stream interconnect. A cost model is also presented for the overhead on the stream interconnect, hardware task configuration and data stream processing at system level, which can be used during the early stage of development. The flexibility and high efficiency of UniStream are demonstrated on Xilinx Virtex-5 and Virtex-6 FPGAs. Experimental results on bitstream configuration/ read-back, data encryption/decryption and Discrete Cosine Transformation show that performance can be significantly improved with different stream modes.
Jian Yan 0002, Jifang Jin, Ying Wang 0032, Xuegong Zhou, Philip H. W. Leong, Lingli Wang
FPL4
2015 An adaptive cross-layer fault recovery solution for reconfigurable SoCs
abstract
Due to the technology scaling, the reconfigurable SoCs built on SRAM-based FPGAs become more susceptible to radiation and aging effects. This paper proposes an adaptive cross-layer fault recovery solution based on hardware/software co-design for reconfigurable SoCs. By pyramidal structure design and cross-layer adaptivity, our solution gives both consideration to hardware circuit integrity at the hardware level and application operating normality at the software level with reduced correction cost. The experiment result shows that the proposed solution can efficiently increase the system reliability and decrease the correction cost.
Jifang Jin, Jian Yan 0002, Xuegong Zhou, Lingli Wang
FPT3
2013 Implementation of high performance hardware architecture of OpenSURF algorithm on FPGA
abstract
This paper proposes a high performance hardware architecture of Speeded Up Robust Features (SURF) algorithm based on OpenSURF. In order to achieve high processing frame rate, the hardware architecture is designed with several characteristics. Firstly, a sliding window method is proposed to extract feature points in parallel at selected scale levels. As a result, the time cost in feature extraction can be greatly reduced. Secondly, data reuse strategy is proposed in orientation generation and descriptor generation to reduce the memory access times. In this way, 3.87x and 2.25X speedup are achieved respectively. Thirdly, the integral image is segmented to buffer in different memory blocks in order to support multiple data accessing in one clock cycle, which will further reduce the whole calculating time of our implementation. The hardware architecture is implemented on an XC6VSX475T FPGA with 156 MHz and its maximal frame rate for VGA format image can reach 356 frames per second (fps), which is 6.25 times frame rate of OpenSURF running on a server with a Xeon 5650 processor, and 6 times the reported frame rate of the recent implementation on three Vritex4 FPGAs [8].
Xitian Fan, Chenlu Wu, Wei Cao 0002, Xuegong Zhou, Shengye Wang, Lingli Wang
FPT4
2013 An FPGA-cluster-accelerated match engine for content-based image retrieval
abstract
In this paper, a high-performance match engine for content-based image retrieval is proposed. Highly customized floating-point(FP) units are designed, to provide the dynamic range and precision of standard FP units, but with considerably less area than standard FP units. Match calculation arrays with various architectures and scales are designed and evaluated. An CBIR system is built on a 12-FPGA cluster. Inter-FPGA connections are based on standard 10-Gigabyte Ethernet. The whole FPGA cluster can compare a query image against 150 million library images within 10 seconds, basing on detailed local features. Compared with the Intel Xeon 5650 server based solution, our implementation is 11.35 times faster and 34.81 times more power efficient.
Chenlu Wu, Xuegong Zhou, Wei Cao 0002, Shengye Wang, Lingli Wang
FPT3
2013 A hardware implementation of Bag of Words and Simhash for image recognition
abstract
Algorithms such as Bag of Words and Simhash have been widely used in image recognition. To achieve better performance as well as energy-efficiency, a hardware implementation of these two algorithms is proposed in this paper. To the best of our knowledge, it is the first time that these algorithms have been implemented on hardware for image recognition purpose. The proposed implementation is able to generate a fingerprint of an image and find the closest match in the database accurately. It is implemented on Xilinx's Virtex-6 SX475T FPGA. Tradeoffs between high performance and low hardware overhead are obtained through proper parallelization. The experimental result shows that the proposed implementation can process 1,018 images per second, approximately 17.8x faster than software on Intel's 12-thread Xeon X5650 processor. On the other hand, the power consumption is 0.35x compared to software-based implementation. Thus, the overall advantage in energy-efficiency is as much as 46x. The proposed architecture is scalable, and is able to meet various requirements of image recognition.
Shengye Wang, Xuegong Zhou, Wei Cao 0002, Chenlu Wu, Xitian Fan, Lingli Wang
FPT3
2013 SPREAD: A Streaming-Based Partially Reconfigurable Architecture and Programming Model
abstract
Partially reconfigurable systems are promising computing platforms for streaming applications, which demand both hardware efficiency and reconfigurable flexibility. To realize the full potential of these systems, a streaming-based partially reconfigurable architecture and unified software/hardware multithreaded programming model (SPREAD) is presented in this paper. SPREAD is a reconfigurable architecture with a unified software/hardware thread interface and high throughput point-to-point streaming structure. It supports dynamic computing resource allocation, runtime software/hardware switching, and streaming-based multithreaded management at the operating system level. SPREAD is designed to provide programmers of streaming applications with a unified view of threads, allowing them to exploit thread, data, and pipeline parallelism; it enhances hardware efficiency while simplifying the development of streaming applications for partially reconfigurable systems. Experimental results targeting cryptography applications demonstrate the feasibility and superior performance of SPREAD. Moreover, the parallelized Advanced Encryption Standard (AES), Data Encryption Standard (DES), and Triple DES (3DES) hardware threads on field-programmable gate arrays show 1.61-4.59 times higher power efficiency than their implementations on state-of-the-art graphics processing units.
Ying Wang 0032, Xuegong Zhou, Lingli Wang, Jian Yan 0002, Wayne Luk, Chenglian Peng, Jiarong Tong
IEEE Trans. Very Large Scale Integr. Syst.2
2012 A partially reconfigurable architecture supporting hardware threads
abstract
As a promising computing platform for stream processing, partially reconfigurable systems have shown their hardware efficiency and reconfiguration flexibility. This paper presents a partially reconfigurable architecture supporting hardware threads. It gives a unified software/hardware thread interface and high throughput point-to-point streaming structure. Dynamic computing resource allocation and streaming-based multi-threaded management are also provided at operating system level. It is easy for programmers to exploit the inherent thread, data and pipeline parallelism in a unified view of threads, enhancing hardware efficiency while improving productivity. The experimental results on a cryptography application demonstrate the feasibility and superior performance. Moreover, the parallelized AES, DES and 3DES hardware threads on field-programmable gate arrays show 1.61-4.59 times higher power efficiency than their implementations on state-of-the-art graphics processing units.
Ying Wang 0032, Jian Yan 0002, Xuegong Zhou, Lingli Wang, Wayne Luk, Chenglian Peng, Jiarong Tong
FPT3
2010 General switch box modeling and optimization for FPGA routing architectures
abstract
This paper explores the FPGA routing architecture based on a new concept of “general switch box (GSB)” to improve the performance of FPGA. Compared with the existing CB/SB routing architecture and CS-box architecture, the proposed GSB architecture has much larger exploration space. Experimental results with MCNC benchmark circuits show that the performance of FPGAs with GSB is about 24.3% better than the CB/SB architecture with the same segment distribution in terms of product of channel width and delay using 0.17% less routing switches for the single wire length. For the two types of wire segments, we propose an architecture with 13.3% performance improvement at the cost of about 0.8% increase in switch number compared to the single wire length GSB architecture.
Kejie Ma, Lingli Wang, Xuegong Zhou, Sheldon X.-D. Tan, Jiarong Tong
FPT3
2007 Online Hybrid Task Scheduling in Reconfigurable Systems
abstract
This paper mainly discusses online tasks scheduling problem on hybrid CPU-FPGA reconfigurable systems. In these systems, hybrid tasks may be binary codes executed on CPU as well as hardware logic circuits implemented on FPGA. Tasks scheduling algorithms of conventional operating systems are not suitable for scheduling hybrid tasks on CPU-FPGA architecture. Based on a real reconfigurable system prototype, we present a task scheduler model and correlative algorithm for scheduling software, hardware and hybrid tasks. This algorithm combines tasks allocation, tasks placement with tasks migration. Simulation results have demonstrated this algorithm provides preferable scheduling performance and reduces the scheduling rejection rate by making use of the great flexibility of hybrid tasks.
Xuegong Zhou, Ying Wang 0032, Chenglian Peng
CSCWD2
2007 Fast On-line Task Placement and Scheduling on Reconfigurable Devices
abstract
This paper focus on on-line placement and scheduling of tasks with known executing time on reconfigurable devices. The notion of recognition-earliest for scheduling algorithms is introduced, that is the algorithm can arrange the start time of a newly arrived task as early as possible. A new scheduling algorithm is proposed. By exploit the knowledge about temporal properties of each task, the algorithm attains recognition-earliest. A fast placement algorithm is also presented. The evaluation results show that the proposed placement algorithm is one of the fastest algorithm, and the proposed scheduling algorithm achieves the best performance compared with previous algorithms, while has a quite low runtime cost.
Xuegong Zhou, Ying Wang 0032, XunZhang Huang, Chenglian Peng
FPL1
2006 On-line scheduling of real-time tasks for reconfigurable computing system
abstract
Efficient task scheduling is very important for obtaining high performance in reconfigurable computing system. Previous researches mostly concentrate on the spatial placement of tasks, and did not pay enough attention to temporal factors. This paper focuses on the on-line scheduling of real-time tasks with known executing time, and introduces the notion of recognition-complete for scheduling algorithms, that is the algorithm can arrange the start time of a newly arrived task as early as possible. A new on-line scheduling algorithm is proposed, which achieves recognition-complete by using the technique of time window. The simulation results show that the proposed algorithm gains a prominent improvement in scheduling performance over previous algorithms
Xuegong Zhou, Ying Wang 0032, XunZhang Huang, Chenglian Peng
FPT1
2005 ESIDE - embedded system integrated design environment based on Eclipse platform
abstract
Traditional embedded system design procedure involves a mixture of design tools to cope with software and hardware design at several abstract levels. Data exchange between tools and cooperation between designers are difficult in such environment. This paper presents an embedded system integrated design environment (ESIDE), which applies a top-down design flow and integrates design tools at different abstract levels. It provides a consistent mechanism for the data exchange between tools and an efficient way to manage various kinds of intellectual property (IP) components with an IP database. ESIDE supports design exchange and management between various designers so that a large number of designers are able to collaborate efficiently in a complex embedded system design.
Xuegong Zhou, Chenglian Peng
CSCWD (1)1