Peng Zhang 0007

dblp:21/1048-7 · DBLP profile ↗
← Back
33ranked-venue papers
1as first author
11since 2021 · last 2025
0009-0006-6841-2257ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 26 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Software engineering, systems software and programming languages · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2025 Activation and Weight Distribution Balancing for Optimal Post-Training Quantization in Learned Image Compression
abstract
Recently, Learned Image Compression (LIC) models have garnered significant attention due to their superior performance in comparison to traditional image codecs. However, the growing complexity of these deep learning-based models results in high memory consumption and computational load, which limits their practical deployment. Quantization has emerged as a promising technique to reduce both the storage requirements and computational overhead. Despite its success in high-level vision tasks like image recognition and object detection, quantization techniques applied to LIC models remain underexplored. In this work, we identify the unique challenges of quantizing LIC models, specifically focusing on the impact of latent distribution ranges in high-bitrate. We observe that the activation layers of high-bitrate models exhibit a wider distribution range, which causes significant performance degradation after quantization. Furthermore, we explore the limitations of existing LIC quantization schemes, such as per-channel quantization for activation layers, which result in poor hardware acceleration performance and increased data storage overhead. To address these challenges, we propose the activation and weight distribution balancing post-training quantization (AWDB-PTQ) method for LIC models, which uses a coarse-to-fine strategy to optimize balancing coefficients. In addition, we employ per-tensor activation quantization and symmetric uniform quantization to better facilitate hardware acceleration. Experimental results demonstrate that our proposed method outperforms existing methods in terms of both compression performance and computational efficiency. Our code and data are available at: https://github.com/jie-yu16/AWDB-PTQ.
Songping Mai, Peng Zhang 0007, Yucheng Jiang
ACM Multimedia3
2025 Hardware-friendly rate estimation algorithm and architecture design for AVS3
Yunyao Yan, Guoqing Xiang, Jie Chen 0001, Xiaofeng Huang, Peng Zhang 0007, Huizhu Jia
Multim. Tools Appl.5
2024 iEDA: An Open-source infrastructure of EDA
abstract
By leveraging the power of open-source software, the EDA tool offers a cost-effective and flexible solution for designers, researchers, and hobbyists alike. Open-source EDA promotes collaboration, innovation, and knowledge sharing within the EDA community. It emphasizes the role of the toolchain in accelerating the development of electronic systems, reducing design costs, and improving design quality. This paper presents an open-source EDA project, iEDA, aiming to build a basic infrastructure for EDA technology evolution and closing the industrial-academic gap in the EDA area. As the foundation for developing EDA tools and researching EDA algorithms and technologies, iEDA is mainly composed of file system, database, manager, operator and interface. To demonstrate the effectiveness of iEDA, we implement and tape out four chips of different scales (from 700k to 500M gates) on different process nodes (110nm and 28nm) with iEDA. iEDA is publicly available on the project home page https://github.com/OSCC-Project/iEDA.
Zengrong Huang, Simin Tao, Zhipeng Huang 0009, Chunan Zhuang, Yihang Qiu, Guojie Luo, Huawei Li 0001, Haihua Shen, Mingyu Chen 0001, Dongbo Bu, Wenxing Zhu, Ye Cai 0001, Xiaoming Xiong, Yi Heng, Peng Zhang 0007, Bei Yu 0001, Biwei Xie, Yungang Bao
ASPDAC19
2024 Enhanced Screen Content Image Compression: A Synergistic Approach for Structural Fidelity and Text Integrity Preservation
abstract
With the rapid development of video conferencing and online education applications, screen content image (SCI) compression has become increasingly crucial. Recently, deep learning techniques have made significant progress in compressing natural images, surpassing the performance of traditional standards like versatile video coding. However, directly applying these methods to SCIs is challenging due to the unique characteristics of SCIs. In this paper, we propose a synergistic approach to preserve structural fidelity and text integrity for SCIs. Firstly, external prior guidance is proposed to enhance structural fidelity and text integrity by providing global spatial attention. Then, a structural enhancement module is proposed to improve the preservation of structural information by enhanced spatial feature transform. Finally, the loss function is optimized for better compression efficiency in text regions by weighted mean square error. Experimental results show that the proposed method achieves 13.3% BD-Rate saving compared to the baseline window attention convolutional neural networks (WACNN) on the JPEGAI, SIQAD, SCID, and MLSCID datasets on average. Our code is available at https://github.com/vpaHduGroup/SFTIP_SCC.
Fangtao Zhou, Xiaofeng Huang, Peng Zhang 0007, Meng Wang 0017, Zhao Wang 0004, Yang Zhou 0052, Haibing Yin
ACM Multimedia3
2024 A hardware-friendly algorithm for LCU-level pipe-lined integer motion estimation
Xizhong Zhu, Guoqing Xiang, Peng Zhang 0007
Multim. Tools Appl.3
2023 An Efficient Real-Time Hardware Architecture for Deblocking Filter in AVS3
abstract
To achieve higher video compression efficiency to cope with the demand for ultra high definition video applications, the AVS3 standard has been proposed recently. As a block partition-based coding standard, AVS3 suffers from the blocking artifact problem especially at low bitrates, which can be alleviated by deblocking filter. This paper presents an efficient hardware architecture for the deblocking filter for AVS3 with high-throughput. First, a fast and hardware-friendly algorithm is proposed to localize the blocking artifact boundaries. Then, a buffer organization method and a data caching strategy are proposed to solve the problem of data dependency among coding units. Based on the proposed optimized algorithm and caching strategy, a four-stage pipelined deblocking filter module architecture is designed. The experimental results show that the proposed architecture can achieve 4K@120fps video processing at 100MHz which is sufficient for real-time application.
Xiaofeng Huang, Guoqing Xiang, Xizhong Zhu, Jiaojiao Yang, Peng Zhang 0007, Huizhu Jia
ICME6
2023 A Hardware-efficient Unified Motion Estimation for Video Coding
abstract
Motion estimation (ME) is one of the most critical tools in video coding and consumes the majority of the encoding complexity. Three types of ME are utilized in the latest video coding standards, namely integer, fractional, and affine MEs. They are implemented as three searches for the integer motion vector (IMV), fractional motion vector (FMV), and control point motion vectors (CPMVs). Many algorithms were proposed to reduce the complexity for them individually, but the overall overhead of three searches is still challenging for hardware implementations. Therefore, we propose a hardware-efficient Unified Motion Estimation (UME) to derive three types of MVs with only one search. An IME with sub-block refinement is performed to collect extra motion information while searching for the IMV. The FMV and CPMVs are then derived from the collected information using a mixed error surface and an overdetermined system. Compared to the default ME algorithms in VVC, the time cost for ME is reduced by 41.63% with a coding loss of only 0.87% under LDB configuration. For hardware implementations, the minimum required resources and corresponding latency are significantly reduced by 75.35% and 69.17%, respectively.
Xizhong Zhu, Guoqing Xiang, Peng Zhang 0007, Huizhu Jia
ACM Multimedia3
2023 A Reconfigurable Multiple Transform Selection Architecture for VVC
abstract
Video coding plays an important role in the highly information-based world as videos contribute the largest part of network traffic. The latest video coding standard Versatile Video Coding (VVC) introduces a new transform scheme multiple transform selection (MTS), which brings considerable coding gains at the expense of high coding complexity. In this article, we propose a reconfigurable MTS architecture that supports all transform types in VVC with square and rectangular sizes ranging from$4\times $4 to 32$\times32$. Firstly, we explore the features of three types of transform matrices and extract the features that are beneficial to designing a unified architecture. Then, we present an improved calculation scheme for general transforms, where the transform matrix is decomposed into two simpler matrices to increase the similarity and decrease the complexity of matrices involved in three types of transform operations. Thanks to the improved calculated scheme, a unified shift-adder unit (SAU) is designed and highly reused by different types. Moreover, we provide a twirling two-point splicing (T2S) scheme to improve reusability and deal with issues of data mismatch when conducting discrete cosine transform (DCT)-II of different sizes. As a consequence, an architecture with constant throughput of 32 pixels/cycle is implemented and specified in Verilog HDL. The synthesis results indicate that the application specific integrated circuit (ASIC)-based and field-programmable gate array (FPGA)-based hardware architectures achieve significant advantages both in area reduction and power consumption compared to existing methods in the literature.
Zhijian Hao, Heming Sun, Guoqing Xiang, Peng Zhang 0007, Xiaoyang Zeng, Yibo Fan
IEEE Trans. Very Large Scale Integr. Syst.4
2022 Efficient Algorithm and Hardware Architecture for Rate Estimation in Mode Decision of AVS3
abstract
Towards enabling advanced video coding for emerging ap-plications, the AVS3 standard has been developed recently, achieving twice the coding efficiency of the AVS2 stan-dard through complex coding tools including advanced rate-distortion optimization (RDO) to select the best mode. The bit-rates are produced with the Advanced-Entropy-Coding (AEC) in the RDO process of AVS3. However, AEC dom-inates the time complexity of RDO and among the steps, con-text updating and interval subdivision are performed recur-sively, which is not conducive to real-time application, espe-cially for the hardware implementation. Thus this paper pro-poses an adaptive rate estimation algorithm with a piece-wise linear function that is very friendly to hardware implemen-tation to accelerate the rate estimation process in the RDO for AVS3 practical applications. The proposed architecture can meet the requirement of 4K@120fps ultra-high-definition videos at 200 MHz, whereas the BD-Rate increases only by 0.67% under the All-Intra (AI) configuration.
Yunyao Yan, Guoqing Xiang, Huizhu Jia, Yuan Li 0014, Peng Zhang 0007, Jie Chen 0001
ICME6
2022 A 3.1 Gbin/s advanced entropy coding hardware design for AVS3
abstract
AVS3 is a newly proposed video coding standard by the Audio Video coding Standard Workgroup, demonstrating higher compression efficiency than the High Efficiency Video Coding standard. Advanced entropy coding is one of the performance bottlenecks of the AVS3 standard video encoder due to the strong data dependency in its arithmetic coding process. A novel arithmetic encoding hardware structure is presented in this paper, and as we know, this is the first paper on AVS3 AEC hardware implementation. Firstly, we select and apply the typical optimization schemes adopted in HEVC context-based adaptive binary arithmetic coding designs. Secondly, Utilizing the unique characteristics of AVS3 AEC, we propose mathematical reordering, variable-clock-cycle range updating and variable-clock-cycle context modeling methods to optimize the critical path. Our design can encode 2.6457 bins per clock cycle, and the corresponding throughput is 3131 Mbin/s in Globalfoundries 28nm process. Compared with the basic anchor structure, it has obtained a performance improvement of 319% and can meet the 8k@l20fps ultra-high-definition video encoding requirements.
Yujie Cai, Xiaoyang Zeng, Yibo Fan, Peng Zhang 0007, Guoqing Xiang, Haibing Yin
ISCAS5
2022 An Area-efficient Unified Transform Architecture for VVC
abstract
The next-generation video coding standard Versatile Video Coding (VVC) adopts Multiple Transform Selection (MTS) to the transform module, improving coding efficiency at the expense of high computational complexity. Compared to High Efficiency Video Coding (HEVC), VVC supports larger sizes and extends the transform types to Discrete Cosine Transform (DCT)-II, Discrete Sine Transform (DST)-VII, and DCT-VIII. This paper presents an area-efficient unified architecture for VVC. To reduce the area consumption, we propose an optimized calculation scheme for general transformations where the transform matrix is decomposed into two simpler matrices named the Low-value matrix and the Error matrix. Based on the decomposition algorithm, Shift-Addition Units (SAUs)-based circuits are designed to conduct matrix multiplication and can be reused by three types. As a result, this unified architecture is capable of performing all types and sizes in VVC. The synthesis results indicate that this architecture achieves an area reduction of 37.9% $\sim$ 72.2% compared with related works for 32-point transforms.
Zhijian Hao, Qi Zheng 0004, Yibo Fan, Guoqing Xiang, Peng Zhang 0007, Heming Sun
ISCAS5
2019 Overcoming Data Transfer Bottlenecks in DNN Accelerators via Layer-Conscious Memory Managment
abstract
Deep Neural Networks (DNNs) are rapidly evolving to satisfy the performance and accuracy requirements in many real world applications. The evolution renders DNNs more and more complex in terms of network topology, data sizes and layer types. Currently most state-of-the-art DNN accelerators adopt a uniform memory hierarchy (UMH) design methodology, which means that the data transferring of all convolutional and fully connected layers must go through the same memory levels. Unfortunately, for some layers, the performance is always bounded by off-chip memory transferring. It is caused by the saturating of data reuse happening in on-chip buffers, resulting in underutilization of on-chip memory. To address this issue, we propose a layer-conscious memory hierarchy (LCMH) methodology for DNN accelerators. LCMH could determine the memory levels of all the layers according to their requirements for off-chip memory bandwidth and on-chip buffer size for the data sources. As a result, the off-chip memory footprints of memory bounded layers could be avoided by keeping the data of them on chip. In addition, we provide architectural support for the accelerators equipped with LCMH. Experimental results show that designs with layer- conscious memory management could achieve up to 36% speedup compared with the designs wth UMH and 5% improvement over state-of-the-art designs.
Xuechao Wei, Yun Liang 0001, Peng Zhang 0007, Cody Hao Yu, Jason Cong
FPGA3
2018 Automated accelerator generation and optimization with composable, parallel and pipeline architecture
abstract
CPU-FPGA heterogeneous architectures feature flexible acceleration of many workloads to advance computational capabilities and energy efficiency in today's datacenters. This advantage, however, is often overshadowed by the poor programmability of FPGAs. Although recent advances in high-level synthesis (HLS) significantly improve the FPGA programmability, it still leaves programmers facing the challenge of identifying the optimal design configuration in a tremendous design space. In this paper we propose the composable, parallel and pipeline (CPP) microarchitecture as an accelerator design template to substantially reduce the design space. Also, by introducing the CPP analytical model to capture the performance-resource trade-offs, we achieve efficient, analytical-based design space exploration. Furthermore, we develop the AutoAccel framework to automate the entire accelerator generation process. Our experiments show that the AutoAccel-generated accelerators outperform their corresponding software implementations by an average of 72x for a broad class of computation kernels.
Jason Cong, Peng Wei 0004, Cody Hao Yu, Peng Zhang 0007
DAC4
2018 S2FA: an accelerator automation framework for heterogeneous computing in datacenters
abstract
Big data analytics using the JVM-based MapReduce framework has become a popular approach to address the explosive growth of data sizes. Adopting FPGAs in datacenters as accelerators to improve performance and energy efficiency also attracts increasing attention. However, the integration of FPGAs into such JVM-based frameworks raises the challenge of poor programmability. Programmers must not only rewrite Java/Scala programs to C/C++ or OpenCL, but, to achieve high performance, they must also take into consideration the intricacies of FPGAs. To address this challenge, we present S2FA (Spark-to-FPGA-Accelerator), an automation framework that generates FPGA accelerator designs from Apache Spark programs written in Scala. S2FA bridges the semantic gap between object-oriented languages and HLS C while achieving high performance using learning-based design space exploration. Evaluation results show that our generated FPGA designs achieve up to 49.9× performance improvement for several machine learning applications compared to their corresponding implementations on the JVM.
Cody Hao Yu, Peng Wei 0004, Max Grossman, Peng Zhang 0007, Vivek Sarkar, Jason Cong
DAC4
2018 TGPA: tile-grained pipeline architecture for low latency CNN inference
abstract
FPGAs are more and more widely used as reconfigurable hardware accelerators for applications leveraging convolutional neural networks (CNNs) in recent years. Previous designs normally adopt a uniform accelerator architecture that processes all layers of a given CNN model one after another. This homogeneous design methodology usually has dynamic resource underutilization issue due to the tensor shape diversity of different layers. As a result, designs equipped with heterogeneous accelerators specific for different layers were proposed to resolve this issue. However, existing heterogeneous designs sacrifice latency for throughput by concurrent execution of multiple input images on different accelerators. In this paper, we propose an architecture named Tile-Grained Pipeline Architecture (TGPA) for low latency CNN inference. TGPA adopts a heterogeneous design which supports pipelining execution of multiple tiles within a single input image on multiple heterogeneous accelerators. The accelerators are partitioned onto different FPGA dies to guarantee high frequency. A partition strategy is designd to maximize on-chip resource utilization. Experiment results show that TGPA designs for different CNN models achieve up to 40% performance improvement than homogeneous designs, and 3X latency reduction over state-of-the-art designs.
Xuechao Wei, Yun Liang 0001, Cody Hao Yu, Peng Zhang 0007, Jason Cong
ICCAD5
2017 Automated Systolic Array Architecture Synthesis for High Throughput CNN Inference on FPGAs
abstract
Convolutional neural networks (CNNs) have been widely applied in many deep learning applications. In recent years, the FPGA implementation for CNNs has attracted much attention because of its high performance and energy efficiency. However, existing implementations have difficulty to fully leverage the computation power of the latest FPGAs. In this paper we implement CNN on an FPGA using a systolic array architecture, which can achieve high clock frequency under high resource utilization. We provide an analytical model for performance and resource utilization and develop an automatic design space exploration framework, as well as source-to-source code transformation from a C program to a CNN implementation using systolic array. The experimental results show that our framework is able to generate the accelerator for real-life CNN models, achieving up to 461 GFlops for floating point data type and 1.2 Tops for 8-16 bit fixed point.
Xuechao Wei, Cody Hao Yu, Peng Zhang 0007, Youxiang Chen, Yun Liang 0001, Jason Cong
DAC3
2017 HLScope+, : Fast and accurate performance estimation for FPGA HLS
abstract
High-level synthesis (HLS) tools have vastly increased the productivity of field-programmable gate array (FPGA) programmers with design automation and abstraction. However, the side effect is that many architectural details are hidden from the programmers. As a result, programmers who wish to improve the performance of their design often have difficulty identifying the performance bottleneck. It is true that current HLS tools provide some estimate of the performance with a fixed loop count, but they often fail to do so for programs with input-dependent execution behavior. Also, their external memory latency model does not accurately fit the actual bus-based shared memory architecture. This work describes a high-level cycle estimation methodology to solve these problems. To reduce the time overhead, we propose a cycle estimation process that is combined with the HLS software simulation. We also present an automatic code instrumentation technique that finds the reason for stall accurately in on-board execution. The experimental results show that our framework provides a cycle estimate with an average error rate of 1.1% and 5.0% for compute- and DRAM-bound modules, respectively, for ADM-PCIE-7V3 board. The proposed metmethodhod is about two orders of magnitude faster than the FPGA bitstream generation.
Peng Zhang 0007, Peng Li 0031, Jason Cong
ICCAD2
2016 Software Infrastructure for Enabling FPGA-Based Accelerations in Data Centers: Invited Paper
abstract
This paper focuses on the development of an infrastructure to enable FPGA-based acceleration in data centers. We present an initial version of an integrated solution that includes automated compilation for accelerator generation, runtime accelerator resource scheduling and management, and acceleration libraries for FPGA-based customized computing for big data applications. The solution can help overcome some of the main challenges with FPGA-based accelerated computing. It has the potential to bring significant performance and energy efficiency improvement for data center applications.
Jason Cong, Muhuan Huang, Peichen Pan, Di Wu 0010, Peng Zhang 0007
ISLPED5
2016 An Optimal Microarchitecture for Stencil Computation Acceleration Based on Nonuniform Partitioning of Data Reuse Buffers
abstract
High-level synthesis (HLS) tools have made significant progress in compiling high-level descriptions of computation into highly pipelined register-transfer level specifications. The high-throughput computation raises a high data demand. To prevent data accesses from being the bottleneck, on-chip memories are used as data reuse buffers to reduce off-chip accesses. Also memory partitioning is explored to increase the memory bandwidth by scheduling multiple simultaneous memory accesses to different memory banks. Prior work on memory partitioning of data reuse buffers is limited to uniform partitioning. In this paper, we perform an early-stage exploration of nonuniform memory partitioning. We use the stencil computation, a popular communication-intensive application domain, as a case study to show the potential benefits of nonuniform memory partitioning. Our novel method can always achieve the minimum memory size and the minimum number of memory banks, which cannot be guaranteed in any prior work. We develop a generalized microarchitecture to decouple stencil accesses from computation, and an automated design flow to integrate our microarchitecture with the HLS-generated computation kernel for a complete accelerator.
Jason Cong, Peng Li 0031, Bingjun Xiao, Peng Zhang 0007
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2015 CMOST: a system-level FPGA compilation framework
abstract
Programming difficulty is a key challenge to the adoption of FPGAs as a general high-performance computing platform. In this paper we present CMOST, an open-source automated compilation flow that maps C-code to FPGAs for acceleration. CMOST establishes a unified framework for the integration of various system-level optimizations and for different hardware platforms. We also present several novel techniques on integrating optimizations in CMOST, including task-level dependence analysis, block-based data streaming, and automated SDF generation. Experimental results show that automatically generated FPGA accelerators can achieve over 8x speedup and 120x energy gain on average compared to the multi-core CPU results from similar input C programs. CMOST results are comparable to those obtained after extensive manual source-code transformations followed by high-level synthesis.
Peng Zhang 0007, Muhuan Huang, Bingjun Xiao, Hui Huang 0001, Jason Cong
DAC1
2015 Resource-Aware Throughput Optimization for High-Level Synthesis
abstract
With the emergence of robust high-level synthesis tools to automatically transform codes written in high-level languages into RTL implementations, the programming productivity when synthesising accelerators improves significantly. However, although the state-of-the-art high-level synthesis tools can offer high-quality designs for simple nested loop kernels, there is still a significant performance gap between the synthesized and the optimal design for real world complex applications with multiple loops.
Peng Li 0031, Peng Zhang 0007, Louis-Noël Pouchet, Jason Cong
FPGA2
2014 An Optimal Microarchitecture for Stencil Computation Acceleration Based on Non-Uniform Partitioning of Data Reuse Buffers
abstract
High-level synthesis (HLS) tools have made significant progress in compiling high-level descriptions of computation into highly pipelined register-transfer level (RTL) specifications. The high-throughput computation raises a high data demand. To prevent data accesses from being the bottleneck, on-chip memories are used as data reuse buffers to reduce off-chip accesses. Also memory partitioning is explored to increase the memory bandwidth by scheduling multiple simultaneous memory accesses to different memory banks. Prior work on memory partitioning of data reuse buffers is limited to uniform partitioning. In this paper, we perform an early-stage exploration of non-uniform memory partitioning. We use the stencil computation, a popular communication-intensive application domain, as a case study to show the potential benefits of non-uniform memory partitioning. Our novel method can always achieve the minimum memory size and the minimum number of memory banks, which cannot be guaranteed in any prior work. We develop a generalized microarchitecture to decouple stencil accesses from computation, and an automated design flow to integrate our microarchitecture with the HLS-generated computation kernel for a complete accelerator.
Jason Cong, Peng Li 0031, Bingjun Xiao, Peng Zhang 0007
DAC4
2014 FPGA Acceleration for Simultaneous Medical Image Reconstruction and Segmentation
abstract
The conventional approach of computed tomography (CT) is to solve each image processing task individually in sequence. An obvious drawback is that the measured data is only used once at the first step, and the possible errors, from noises in the measured data, inappropriate modeling, or inappropriate parameters, are not easy to be corrected and will be propagated into the later steps. As a consequence, approaches that combine the reconstruction and the specific processing task have become popular. This work adopts an iterative algorithm with simultaneous reconstruction and segmentation using the Mumford-Shah model, which can be applied not only to regularize the ill-posedness of the tomographic reconstruction problem, but also to compute segmentation directly from the measured data. The Mumford-Shah model is both mathematically and computationally difficult. In this paper, we accelerated this computation and data intensive application by FPGA devices and achieved 9.24X speedup over the conventional CPU implementation.
Peng Li 0031, Thomas Page, Guojie Luo, Wentai Zhang 0001, Peng Zhang 0007, Peter Maass, Ming Jiang 0001, Jason Cong
FCCM6
2014 Combining computation and communication optimizations in system synthesis for streaming applications
abstract
Data streaming is a widely-used technique to exploit task-level parallelism in many application domains such as video processing, signal processing and wireless communication. In this paper we propose an efficient system-level synthesis flow to map streaming applications onto FPGAs with consideration of simultaneous computation and communication optimizations. The throughput of a streaming system is significantly impacted by not only the performance and number of replicas of the computation kernels, but also the buffer size allocated for the communications between kernels. In general, module selection/replication and buffer size optimization were addressed separately in previous work. Our approach combines these optimizations together in system scheduling which minimizes the area cost for both logic and memory under the required throughput constraint. We first propose an integer linear program (ILP) based solution to the combined problem which has the optimal quality of results. Then we propose an iterative algorithm which can achieve the near-optimal quality of results but has a significant improvement on the algorithm scalability for large and complex designs. The key contribution is that we have a polynomial-time algorithm for an exact schedulability checking problem and a polynomial-time algorithm to improve the system performance with better module implementation and buffer size optimization. Experimental results show that compared to the separate scheme of module select/replication and buffer size optimization, the combined optimization scheme can gain 62% area saving on average under the same performance requirements. Moreover, our heuristic can achieve 2 to 3 orders of magnitude of speed-up in runtime, with less than 10% area overhead compared to the optimal solution by ILP.
Jason Cong, Muhuan Huang, Peng Zhang 0007
FPGA3
2013 Memory partitioning for multidimensional arrays in high-level synthesis
abstract
Memory partitioning is widely adopted to efficiently increase the memory bandwidth by using multiple memory banks and reducing data access conflict. Previous methods for memory partitioning mainly focused on one-dimensional arrays. As a consequence, designers must flatten a multidimensional array to fit those methodologies. In this work we propose an automatic memory partitioning scheme for multidimensional arrays based on linear transformation to provide high data throughput of on-chip memories for the loop pipelining in high-level synthesis. An optimal solution based on Ehrhart points counting is presented, and a heuristic solution based on memory padding is proposed to achieve a near optimal solution with a small logic overhead. Compared to the previous one-dimensional partitioning work, the experimental results show that our approach saves up to 21% of block RAMs, 19% in slices, and 46% in DSPs.
Peng Li 0031, Peng Zhang 0007, Chen Zhang 0001, Jason Cong
DAC3
2013 Efficient system-level mapping from streaming applications to FPGAs (abstract only)
abstract
Streaming processing is an important computation model that represents many applications in various domains such as video processing, signal processing and wireless communication. FPGA is a natural platform for streaming applications because the task-level pipelined parallelism can be efficiently implemented on FPGA by its customizable communication and memory architecture. In this paper we propose an efficient design space exploration algorithm to map kernels of streaming applications onto FPGAs. We aim at finding the most area-efficient selections of hardware modules from the implementation library while satisfying the system performance requirement. In particular, we consider both module selection and replication techniques. Design metrics are formulated in our high-level model based on these two techniques. In addition, we extend the analytic formulations in previous work by supporting complex stream graph structures like feedback loops. The proposed iterative exploration algorithm is based on the system of difference constraint (SDC) and thus can be solved in polynomial time. Compared to previous mainstream ILP-based solutions, our proposed algorithm is scalable and practical in large systems. Both the ILP formulation and our proposed iterative exploration mechanism are applied to a set of streaming applications from StreamIt benchmarks and also to one real example MPEG-4 decoder. Experiments demonstrate that our design space exploration algorithm can efficiently find a feasible solution with an average 5.7% area overhead.
Jason Cong, Muhuan Huang, Peng Zhang 0007
FPGA3
2013 Polyhedral-based data reuse optimization for configurable computing
abstract
Many applications, such as medical imaging, generate intensive data traffic between the FPGA and off-chip memory. Significant improvements in the execution time can be achieved with effective utilization of on-chip (scratchpad) memories, associated with careful software-based data reuse and communication scheduling techniques. We present a fully automated C-to-FPGA framework to address this problem. Our framework effectively implements data reuse through aggressive loop transformation-based program restructuring. In addition, our proposed framework automatically implements critical optimizations for performance such as task-level parallelization, loop pipelining, and data prefetching.
Louis-Noël Pouchet, Peng Zhang 0007, P. Sadayappan, Jason Cong
FPGA2
2013 Automatic multidimensional memory partitioning for FPGA-based accelerators (abstract only)
abstract
With the increase of data processing throughput in reconfigurable computing, data parallelism is now crucial for the performance of FPGA-based accelerators. However, most of the data parallelism optimizations are still performed manually by experienced hardware designers. Memory partitioning is widely adopted to efficiently increase the memory bandwidth by using multiple memory banks and reducing data access conflict. Previous methods for memory partitioning mainly focused on one-dimensional arrays. As a consequence, designers must flatten a multidimensional array to fit those methodologies, but it makes the partition related to the dimensional width of the array. In this work we propose an automatic memory partitioning scheme for multidimensional arrays to provide high data throughput of on-chip memories for the loop pipelining in high-level synthesis. Linear transformation is applied to optimize the layout of the data elements in the memory banks, with the partition unrelated to the dimensional width. Two transformation vectors are used to map the original data element onto different banks and different inner bank offsets. The vector for the optimal bank mapping is decided by non-conflict access constraint. In addition, a memory padding technique is proposed to find a vector for inner bank offset with a trade-off between practicality and optimality. We use six benchmarks with different access patterns to prove our idea. Compared to the previous one-dimensional partitioning work, the experimental results show that our approach saves up to 21% of block RAMs, 19% in slices, and 46% in DSPs.
Peng Li 0031, Peng Zhang 0007, Chen Zhang 0001, Jason Cong
FPGA3
2012 An integrated and automated memory optimization flow for FPGA behavioral synthesis
abstract
Behavioral synthesis tools have made significant progress in compiling high-level programs into register-transfer level (RTL) specifications. But manually rewriting code is still necessary in order to obtain better quality of results in memory system optimization. In recent years different automated memory optimization techniques have been proposed and implemented, such as data reuse and memory partitioning, but the problem of integrating these techniques into an applicable flow to obtain a better performance has become a challenge. In this paper we integrate data reuse, loop pipelining, memory partitioning, and memory merging into an automated optimization flow (AMO) for FPGA behavioral synthesis. We develop memory padding to help in the memory partitioning of indices with modulo operations. Experimental results on Xilinx Virtex-6 FPGAs show that our integrated approach can gain an average 5.8× throughput and 4.55× latency improvement compared to the approach without memory partitioning. Moreover, memory merging saves up to 44.32% of block RAM (BRAM).
Peng Zhang 0007, Xu Cheng 0001, Jason Cong
ASP-DAC2
2012 Optimizing memory hierarchy allocation with loop transformations for high-level synthesis
abstract
For the majority of computation-intensive application systems, off-chip memory bandwidth is a critical bottleneck for both performance and power consumption. The efficient utilization of limited on-chip memory resources plays a vital role in reducing the off-chip memory accesses. This paper presents an efficient approach for optimizing the on-chip memory allocation by loop transformations in the imperfectly nested loops. We analytically model the on-chip buffer size and off-chip bandwidth after affine loop transformation, loop fusion/distribution and code motion. Branch-and-bound and knapsack reuse techniques are proposed to reduce the computation complexity in finding optimal solutions. Experimental results show that our scheme can save 40% of on-chip memory size with the same bandwidth consumption compared to the previous approaches.
Jason Cong, Peng Zhang 0007, Yi Zou 0001
DAC2
2012 Combining module selection and replication for throughput-driven streaming programs
abstract
Streaming processing is widely adopted in many data-intensive applications in various domains. FPGAs are commonly used to realize these applications since they can exploit inherent data parallelism and pipelining in the applications to achieve a better performance. In this paper we investigate the design space exploration problem (DSE) when mapping streaming applications onto FPGAs. Previous works narrowly focus on using techniques like replication or module selection to meet the throughput target. We propose to combine these two techniques together to guide the design space exploration. A formal formulation and solution to this combined problem is presented in this paper. Our objective is to optimize the total area cost subject to the throughput constraint. In particular, we are able to handle the feedback loops in the streaming programs, which, to the best of our knowledge, has never been discussed in previous work. Our methodology is evaluated with high-level synthesis tools, and we demonstrate our workflow on a set of benchmarks that vary from module kernel design such as FFT to large designs such as an MPEG-4 decoder.
Jason Cong, Muhuan Huang, Bin Liu 0006, Peng Zhang 0007, Yi Zou 0001
DATE4
2012 Memory partitioning and scheduling co-optimization in behavioral synthesis
abstract
Achieving optimal throughput by extracting parallelism in behavioral synthesis often exaggerates memory bottleneck issues. Data partitioning is an important technique for increasing memory bandwidth by scheduling multiple simultaneous memory accesses to different memory banks. In this paper we present a vertical memory partitioning and scheduling algorithm that can generate a valid partition scheme for arbitrary affine memory inputs. It does this by arranging non-conflicting memory accesses across the border of loop iterations. A mixed memory partitioning and scheduling algorithm is also proposed to combine the advantages of the vertical and other state-of-art algorithms. A set of theorems is provided as criteria for selecting a valid partitioning scheme. This is followed by an optimal and scalable memory scheduling algorithm. By utilizing the property of constant strides between memory addresses in successive loop iterations, an address translation optimization technique for an arbitrary partition factor is proposed to improve performance, area and energy efficiency. Experimental results show that on a set of real-world medical image processing kernels, the proposed mixed algorithm with address translation optimization can gain speed-up, area reduction and power savings of 15.8%, 36% and 32.4% respectively, compared to the state-of-art memory partitioning algorithm.
Peng Li 0031, Peng Zhang 0007, Guojie Luo, Tao Wang 0004, Jason Cong
ICCAD3
2011 Combined loop transformation and hierarchy allocation for data reuse optimization
abstract
External memory bandwidth is a crucial bottleneck in the majority of computation-intensive applications for both performance and power consumption. Data reuse is an important technique for reducing the external memory access by utilizing the memory hierarchy. Loop transformation for data locality and memory hierarchy allocation are two major steps in data reuse optimization flow. But they were carried out independently. This paper presents a combined approach which optimizes loop transformation and memory hierarchy allocation simultaneously to achieve global optimal results on external memory bandwidth and on-chip data reuse buffer size. We develop an efficient and optimal solution to the combined problem by decomposing the solution space into two subspaces with linear and nonlinear constraints respectively. We show that we can significantly prune the solution space without losing its optimality. Experimental results show that our scheme can save up to 31% of on-chip memory size compared to the separated two-step method when the memory hierarchy allocation problem is not trivial. Also, run-time complexity is acceptable for the practical cases.
Jason Cong, Peng Zhang 0007, Yi Zou 0001
ICCAD2