Jingfei Jiang

dblp:76/1876 · DBLP profile ↗
← Back
33ranked-venue papers
1as first author
19since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 22 · 1 first-author · 12 since 2021Artificial intelligence and machine learning · 6 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Dual cuts for efficiency: Streamlining NAS with train-time pruning of supernet and search space
Di Niu 0001, Hengyue Pan, Jingfei Jiang, Jinwei Xu
Knowl. Based Syst.3
2025 Multi-modal Parallelism Scheduling for Heterogeneous Multicore Computing Systems
Jingfei Jiang, Jinwei Xu
ICA3PP (2)2
2025 DiffRS: An Extensible Diffusion Model for Remote Sensing Image Generation
abstract
Remote sensing image generation is of great value for virtual environment creation and adversarial learning for fake news detection. It could also address the learning sample shortage in the region of interest. However, most current image generation methods are limited to producing images of fixed sizes, few studies on extensible natural image generation largely focus on the stitching of random contents, lacking effective exploration of contextual information, which weakens the coherence of the extended images. To address this problem, we propose an extensible generation method for remote sensing images with the model DiffRS. This approach allows for sequential extension of arbitrary sizes by exploring the generated neighboring regions. The method is particularly suitable for scenes like remote sensing images where a generation block could cover multiple independent targets, rather than natural image tasks which may stitch across regions to form a completely target. Compared to the state-of-the-art extensible generation methods, DiffRS could improve the large scale image generation with better structure consistency, richer details and higher realism. Experiments showed that DiffRS could improve the FID score by 4.6% and 3.2% respectively in comparison with the MultiDiffusion and Mixture of Diffuser models.
Xin Niu 0002, Jingfei Jiang, Hengyue Pan
ICASSP3
2025 FQuant: Fast Quantization with Adaptive Resolution via the Clustering Algorithm
Linagwei Li, Jingfei Jiang, Jinwei Xu, Shunan Zhou, Minghua Zhu
ICIC (21)2
2025 NP4Q: Nice Point for Post-training Quantization of Object Detection Models
abstract
Post-Training Quantization (PTQ) methods have been widely applied in neural network model compression due to their ability to compress models without the need for retraining the model weights. Recently, various quantization methods for object detection models have been proposed. Unfortunately, object detection models are highly sensitive to quantization, particularly in low-bit-width scenarios, due to the high dynamic range of values. Addressing the long-tailed distribution of activations in object detection models, we propose NP4Q, a quantization framework that employs a piecewise non-uniform clustering quantization method. NP4Q is inspired by our findings that long-tailed activations can be characterized using the mean and standard deviation in a similar manner. Consequently, NP4Q utilizes a piecewise quantization strategy. Specifically, NP4Q first introduces a Nice Boundary Point (NBP) using the mean and standard deviation to partition the activations into head and tail segments. For the tail segment, a simple MinMax uniform quantization is sufficient, while for the head segment, the Meanshift Clustering Quantization (MCQ) is proposed to better handle the majority of the data. By leveraging NBP and MCQ, NP4Q can rapidly generate high-precision quantized models. Experiments demonstrate that NP4Q effectively segments the long tail and enhances the detection accuracy of quantized models, particularly in low-bit configurations. For instance, NP4Q pushes the accuracy of YOLOv5s to 52.2% and RetinaNet to 35.3% in 4-bit INT.
Shunan Zhou, Jingfei Jiang, Liangwei Li, Minghua Zhu
IJCNN2
2025 SAPFIS: a parallel fuzzy inference system for air combat situation assessment
Jingfei Jiang, Jinwei Xu, Pengbo Wu
J. Supercomput.2
2025 SPDFA: A Novel Dataflow Fusion Sparse Deep Neural Network Accelerator
abstract
Unstructured sparse pruning significantly reduces the computational and parametric complexities of deep neural network models. Nevertheless, the highly irregular nature of sparse models limits their performance and efficiency on traditional computing platforms, thereby prompting the development of specialized hardware solutions. To improve computational efficiency, we introduce the Sparse Dataflow Fusion Accelerator (SPDFA), a specialized architecture meticulously designed for sparse deep neural networks. Firstly, we present a non-blocking data distribution-computing engine that integrates inner product and column product. This engine boosts computational efficiency by decomposing matrix multiplication and convolution into rectangular matrix-vector multiplications. Secondly, we implement a computation array to further exploit the parallelism, and design an on-chip buffer structure that supports multi-line memory access mode. Lastly, to bolster the adaptability of our accelerator, we propose an innovative macroinstruction set coupled with a micro-kernel scheme. Furthermore, we refine the macroinstruction issue strategy, thereby further enhancing computational efficiency. Our evaluation results demonstrate that SPDFA achieves an average 1.29 \(\times\) –2.38 \(\times\) improvement in computational efficiency compared to the state-of-the-art SpMM accelerators when applied to unstructured sparse deep neural network models. Furthermore, its performance outperforms existing sparse neural network accelerators by a factor of 1.03 \(\times\) –1.83 \(\times\) . Additionally, SPDFA exhibits excellent scalability with a scaling efficiency exceeding 80%.
Jinwei Xu, Jingfei Jiang, Xifu Qian, Yong Dou
ACM Trans. Reconfigurable Technol. Syst.2
2024 HPFIA: A High-Performance Fuzzy Inference Accelerator for Situation Assessment on Airborne Equipment
abstract
The ever-increasing multi-source fusion information perceived by situation assessment system pose a computational challenge to current airborne equipment. Fuzzy inference method introduced in situation assessment could effectively adapt to the incompleteness and uncertainty of situational information, but still struggling to meet the high-performance requirements under limited hardware resources on airborne equipment. Leveraging hardware accelerators (GPUs, FPGAs, etc.) to accelerate intensive computation like situation factor evaluation has become paramount. Since the lack of relevant acceleration methods in state-of-the-art researches, this paper presents the first-ever FPGA accelerator of the HPFIA to optimize calculation performance of situation assessment. Our designed accelerator delivers up to 237.48× performance improvement and 3230.15× better energy efficiency ratio over the software implementation on three universal computing platforms.
Jingfei Jiang, Jinwei Xu
HPCC2
2024 Funnel: An Efficient Sparse Attention Accelerator with Multi-Dataflow Fusion
abstract
The self-attention mechanism is the core component of Transformer, which provides a powerful ability to understand the sequence context. However, the self-attention mechanism also suffers from a large amount of redundant computation. Model sparsification can effectively reduce computational load, but the irregularity of non-zeros introduced by sparsification significantly decreases hardware efficiency. This paper proposes Funnel, an accelerator that dynamically predicts sparse attention patterns and efficiently processes unstructured sparse data. Firstly, we adopt a fast quantization method based on lookup table to minimize the cost of sparse patterns prediction. Secondly, we propose Funnel Computing Unit (FCU), a hardware architecture that efficiently handles sparse attention through multi-dataflow fusion. Sampled Dense-Dense Matrix Multiplication (SDDMM) and Sparse-Dense Matrix Multiplication (SpMM) are core components of sparse attention mechanism. FCU unifies the computation ways of matrix inner product and row-wise product to support SDDMM and SpMM at the same time, which greatly reduces the storage and movement overhead of intermediate results. Lastly, we devise a lightweight buffer and data tiling strategy tailored to the proposed accelerator, aimed at enhancing data reuse. Experiments demonstrate that our accelerator achieves 0.10-0.25 sparsity with small accuracy loss. When computing the self-attention layer, it attains hardware efficiency ranging from 60% to 85%. Compared to CPU and GPU, it achieves 5.60x and 8.20x speedup. Compared to the state-of-the-art attention accelerators A3, SpAtten, FTRANS, and Sanger, it achieves 7.37x, 4.52x, 9.58x, and 3.08x speedup.
Shenghong Ma, Jinwei Xu, Jingfei Jiang, Dongsheng Li 0001
ISPA3
2024 Multi-relation Neural Network Recommendation Model Based on Knowledge Graph Embedding Algorithm
Hongpu Liu, Jingfei Jiang, Lingshu Kong, Jingshu Wang
KSEM (1)2
2024 End-To-End High-Quality Transformer Object Detection Model Applied to Human Head Detection
Rongchun Li, Peng Qiao, Jingfei Jiang
PRCV (12)4
2024 Efficient SpMM Accelerator for Deep Learning: Sparkle and Its Automated Generator
abstract
Deep learning (DL) technology has made breakthroughs in a wide range of intelligent tasks, such as vision, language, recommendation systems, and so on. Sparse matrix multiplication (SpMM) is the key computation kernel of most sparse models. Conventional computing platforms, such as CPUs, GPUs, and AI chips with regular processing units, are unable to effectively support sparse computation due to their fixed structure and instruction sets. This work extends Sparkle, an accelerator architecture, which is developed specifically for processing SpMM in DL. During the balanced data loading process, some modifications are implemented to enhance the flexibility of the Sparkle architecture. Additionally, a Sparkle generator is proposed to accommodate diverse resource constraints and facilitate adaptable deployment. Leveraging Sparkle’s structural parameters and template-based design methods, the generator enables automatic Sparkle circuit generation under varying parameters. An instantiated Sparkle accelerator is implemented on the Xilinx xqvu11p FPGA platform with a specific configuration. Compared to the state-of-the-art SpMM accelerator SIGMA, the Sparkle accelerator instance improves the sparse computing efficiency by about 10 to 20 \(\%\) . Furthermore, the Sparkle instance achieved 7.76 \(\times\) higher performance over the Nvidia Orin NX GPU. More instances of accelerators with different parameters were evaluated, demonstrating that the Sparkle architecture can effectively accelerate SpMM.
Shiyao Xu, Jingfei Jiang, Jinwei Xu, Xifu Qian
ACM Trans. Reconfigurable Technol. Syst.2
2023 Auto-Divide GNN: Accelerating GNN Training with Subgraph Division
Zhejiang Ran, Ke-shi Ge, Zhiquan Lai, Jingfei Jiang, Dongsheng Li 0001
Euro-Par5
2022 Sparkle: A High Efficient Sparse Matrix Multiplication Accelerator for Deep Learning
abstract
Deep learning (DL) technology is applied to a wide range of intelligent tasks across vision, language, recommendation systems, etc. Large DL models with high sparsity become critical for various intelligent applications and require an energy-efficient hardware accelerator. Sparse-dense matrix multiplication (SpMM) is a key computation kernel widely used in most sparse and large DL workloads. However, traditional computing platforms such as CPU, GPU, and Al chips with regular processing units are limited to support sparsity by their fixed structures. In this work, a specific SpMM accelerator named Sparkle is proposed which achieves high performance and high computational efficiency. A block-wise arrangement approach is proposed in Sparkle to process matrix multiplications. A novel compressed sparsity format, the pointer-bitmap, is designed to simplify the decoding process and improve the efficiency of data loading. Grouped PEs and configurable hierarchical reduction network are deployed to leverage sparsity, further enhancing the utilization of the compute resources. Sparkle is implemented using the Xilinx xqvu11p FPGA. A diverse set of matrices in DL workloads are evaluated and Sparkle achieves 2.1× higher energy efficiency over the NVIDIA TITAN X GPU. Our experiments also show that Sparkle roughly promotes 26% compute efficiency better than state-of-the-art sparse accelerators SIGMA.
Shiyao Xu, Jingfei Jiang, Jinwei Xu, Chaorun Liu, Yuanhong He
ICCD2
2022 MLPs: Efficient Training of MiniGo on Large-scale Heterogeneous Computing System
abstract
Deep Reinforcement Learning has been successfully applied in various applications and achieved impressive performance compared with previous traditional methods but suffers from high computation cost and long training time. MLPerf takes deep reinforcement learning as one of the benchmark tracks and provides a single node training version of MiniGo as a reference. A key challenge is to achieve efficient MiniGo training on a large-scale computing system. According to the training computation pattern in MiniGo and the characteristics of our large-scale heterogeneous computing system, we propose a MultiLevel Parallel strategy, MLPs, including task-level parallelism between nodes, CPU-DSP heterogeneous parallelism, and DSP multi-core parallelism. The proposed method reduces the overall execution time from 43 hours to 16 hours while scaling the node size from 1067 to 4139. The scaling efficiency is 69.1%. According to our fitting method, the scaling efficiency is 46.5% when scaling to 8235 nodes. The experimental results show that the proposed method achieves the efficient training of MiniGo on the largescale heterogeneous computing system.
Peng Qiao, Zhouyu He, Rongchun Li, Jingfei Jiang, Yong Dou, Dongsheng Li 0001
ICPADS4
2022 Evaluating a New Attention Framework Based on Matrix Blocking for Attention Models on FPGAs
abstract
The attention mechanism has recently shown superior performance in natural language processing and computer vision tasks. But its complex dataflow and large-scale matrix calculation with huge computing and memory overhead pose a great challenge for the design of hardware accelerators. And previous solutions that benefited from matrix partitioning are bounded by the softmax function. In this paper, we propose a new attention framework that can dramatically improve the performance of attention model inference for long sequence tasks on FPGAs. We design a novel accelerator architecture that employs two systolic arrays and a ping-pong structure to accelerate attention calculation. Meanwhile, we propose an analytical model to predict resource usage and performance, which guides a fast design space exploration. Experiments using the state-of-the-art BERT demonstrate the design achieves 4.61 and 1.24× improvement in speed and energy efficiency compared to CPU and GPU on the Xilinx XCZU11EG platform.
Jingfei Jiang, Jinwei Xu
ICTAI2
2021 RFC-HyPGCN: A Runtime Sparse Feature Compress Accelerator for Skeleton-Based GCNs Action Recognition Model with Hybrid Pruning
abstract
Skeleton-based Graph Convolutional Networks (GCNs) models for action recognition have achieved excellent prediction accuracy in the field. However, limited by large model and computation complexity, GCNs for action recognition like 2s-AGCN have insufficient power-efficiency and throughput on GPU. Thus, the demand of model reduction and hardware acceleration for low-power GCNs action recognition application becomes continuously higher.To address challenges above, this paper proposes a runtime sparse feature compress accelerator with hybrid pruning method: RFC-HyPGCN. First, this method skips both graph and spatial convolution workloads by reorganizing the multiplication order. Following spatial convolutions channel-pruning dataflow, a coarse-grained pruning method on temporal filters is designed, together with sampling-like fine-grained pruning on time dimension. Later, we come up with an architecture where all convolutional layers are mapped on chip to pursue high throughput. To further reduce storage resource utilization, online sparse feature compress format is put forward. Features are divided and encoded into several banks according to presented format, then bank storage is split into depth-variable mini-banks. Furthermore, this work applies quantization, input-skipping and intra-PE dynamic data scheduling to accelerate the model. In experiments, proposed pruning method is conducted on 2s-AGCN, acquiring 3.0x-8.4x model compression ratio and 73.20% graph-skipping efficiency with balancing weight pruning. Implemented on Xilinx XCKU-115 FPGA, the proposed architecture has the peak performance of 1142 GOP/s and achieves up to 9.19x and 3.91x speedup over high-end GPU NVIDIA 2080Ti and NVIDIA V100, respectively. Compared with latest accelerator for action recognition GCNs models, our design reaches 22.9x speedup and 28.93% improvement on DSP efficiency.
Dong Wen 0004, Jingfei Jiang, Jinwei Xu, Yang Zhao 0003, Yong Dou
ASAP2
2021 A high-throughput scalable BNN accelerator with fully pipelined architecture
Jingfei Jiang, Jinwei Xu, Peng Zhang 0035, Dong Wen 0004, Yong Dou
CCF Trans. High Perform. Comput.2
2021 An energy-efficient convolutional neural network accelerator for speech classification based on FPGA and quantization
Dong Wen 0004, Jingfei Jiang, Yong Dou, Jinwei Xu
CCF Trans. High Perform. Comput.2
2020 A Dynamic Mapping Model for General CNN Accelerator Based on FPGA
Jingfei Jiang, Jinwei Xu
NPC2
2019 Enhancing 2D Representation via Adjacent Views for 3D Shape Retrieval
abstract
Multi-view shape descriptors obtained from various 2D images are commonly adopted in 3D shape retrieval. One major challenge is that significant shape information are discarded during 2D view rendering through projection. In this paper, we propose a convolutional neural network based method, CenterNet, to enhance each individual 2D view using its neighboring ones. By exploiting cross-view correlations, CenterNet learns how adjacent views can be maximally incorporated for an enhanced 2D representation to effectively describe shapes. We observe that a very small amount of, e.g., six, enhanced 2D views, are already sufficient for a panoramic shape description. Thus, by simply aggregating features from six enhanced 2D views, we arrive at a highly compact yet discriminative shape descriptor. The proposed shape descriptor significantly outperforms state-of-the-art 3D shape retrieval methods on the ModelNet and ShapeNetCore55 benchmarks, and also exhibits robustness against object occlusion.
Zhaoqun Li, Biao Leng, Jingfei Jiang
ICCV5
2017 An FPGA-based processor for training convolutional neural networks
abstract
Convolutional neural networks (CNNs) have gained great success in various computer vision applications. However, training a CNN model is computation-intensive and time-consuming. Hence training is mainly processed on large clusters of high-performance processors like server CPUs and GPUs. In this paper, we propose an FPGA-based processor design to accelerate the training process of CNNs. We first analyze the operations in all types of CNN layers in the training process. A uniform computation engine design is proposed to efficiently carry out all kinds of operations based on the analysis. Then a scalable accelerator framework is presented that exploits the parallelism further by unrolling the loops in two levels. The proposed accelerator design is demonstrated by implementing a processor on the Xilinx ZU19EG FPGA working at 200 MHz. The evaluation results on a group of CNN models show that our processor is 5.7 to 10.7-fold faster than the software implementations on the Intel Core i5-4440 CPU(@3.10GHz).
Yong Dou, Jingfei Jiang, Qiang Wang 0006, Paul Chow
FPT3
2017 Throughput-Optimized FPGA Accelerator for Deep Convolutional Neural Networks
abstract
Deep convolutional neural networks (CNNs) have gained great success in various computer vision applications. State-of-the-art CNN models for large-scale applications are computation intensive and memory expensive and, hence, are mainly processed on high-performance processors like server CPUs and GPUs. However, there is an increasing demand of high-accuracy or real-time object detection tasks in large-scale clusters or embedded systems, which requires energy-efficient accelerators because of the green computation requirement or the limited battery restriction. Due to the advantages of energy efficiency and reconfigurability, Field-Programmable Gate Arrays (FPGAs) have been widely explored as CNN accelerators. In this article, we present an in-depth analysis of computation complexity and the memory footprint of each CNN layer type. Then a scalable parallel framework is proposed that exploits four levels of parallelism in hardware acceleration. We further put forward a systematic design space exploration methodology to search for the optimal solution that maximizes accelerator throughput under the FPGA constraints such as on-chip memory, computational resources, external memory bandwidth, and clock frequency. Finally, we demonstrate the methodology by optimizing three representative CNNs (LeNet, AlexNet, and VGG-S) on a Xilinx VC709 board. The average performance of the three accelerators is 424.7, 445.6, and 473.4GOP/s under 100MHz working frequency, which outperforms the CPU and previous work significantly.
Yong Dou, Jingfei Jiang, Jinwei Xu, Shijie Li 0002, Yongmei Zhou, Yingnan Xu
ACM Trans. Reconfigurable Technol. Syst.3
2016 Automatic code generation of convolutional neural networks in FPGA implementation
abstract
Convolutional neural networks (CNNs) have gained great success in various computer vision applications. However, state-of-the-art CNN models are computation-intensive and hence are mainly processed on high performance processors like server CPUs and GPUs. Owing to the advantages of high performance, energy efficiency and reconfigurability, Field-Programmable Gate Arrays (FPGAs) have been widely explored as CNN accelerators. In this paper, we propose parallel structures to exploit the inherent parallelism and efficient computation units to perform operations in convolutional and fully-connected layers. Further, an automatic generator is proposed to generate Verilog HDL source code automatically according to high-level hardware description language. Execution time, DSP consumption and performance are analytically modeled based on some critical design variables. We demonstrate the automatic methodology by implementing two representative CNNs (LeNet and AlexNet) and evaluate the execution time models by comparing estimated and measured values. Our results show that the proposed automatic methodology yields hardware design with good performance and saves much developing round time.
Yong Dou, Jingfei Jiang, Jinwei Xu
FPT3
2016 Improved Survey Propagation on Graphics Processing Units
Yang Zhao 0003, Jingfei Jiang, Pengbo Wu
GPC2
2016 Performance modeling of hyper-scale custom machine for the principal steps in block Wiedemann algorithm
Jingfei Jiang
J. Supercomput.2
2016 Coarse-Grained Architecture for Fingerprint Matching
abstract
Fingerprint matching is a key procedure in fingerprint identification applications. The minutiae-based fingerprint matching algorithm is one of the most typical algorithms achieving a reasonably correct recognition rate. This study proposes a coarse-grained parallel architecture called fingerprint matching core (FMC) to accelerate fingerprint matching. The proposed architecture has a two-level parallel structure (i.e., parallel among groups (PAG) and parallel in group (PIG)). A multirequest controller is added to the PAG structure to obtain a concurrent operation of the multiple processing element group (PEG). The DDR3 controller is used in the PIG structure to read eight minutiae from eight different fingerprints and realize the simultaneous computation of the eight PEs. The whole system is implemented on a Xilinx FPGA board with a Virtex VII XC7VX485T chip. The 16-PEG FMC achieves a throughput of about 9.63 million fingerprint pairs per second, which is larger than that achieved on a Tesla K20c platform. The software execution times are also measured on the 2.93GHz Intel Xeon 5670, 2.3GHz AMD Opteron(tm) Processor 6376, and Tesla K20c platforms. The Intel Xeon 5670 has two processors with 12 cores, and the AMD Opteron(tm) Processor 6376 has two processors with 16 cores. Moreover, the throughput is about 31 times that achieved on a 2.93GHz Intel Xeon 5670 single core.
Jinwei Xu, Jingfei Jiang, Yong Dou, Xiaolong Shen
ACM Trans. Reconfigurable Technol. Syst.2
2015 Optimized deep belief networks on CUDA GPUs
abstract
A deep belief network (DBN) is an important branch of deep learning models and has been successfully applied in many machine learning and pattern recognition fields such as computer vision and speech recognition. However, the training of billions of parameters in DBN is computationally challenging for modern central processing units (CPUs). Many studies have reported the efficient implementations of the pre-training process of DBNs for graphics processing units (GPUs), but few studies have mentioned the fine-tuning process of DBNs. In this paper, we describe an efficient DBN implementation on the GPU, including the pre-training and fine-tuning processes. Experimental results show that our proposed method on the GPU (NVIDIA Tesla K40c) achieves up to 22 speedups on the pre-training process and 33 speedups on the fine-tuning processes compared with conventional CPU (Intel Core i7-4790K) implementations. Moreover, the performance of our algorithm is superior to that of the OpenBLAS library on the CPU and the CUBLAS library on the GPU.
Teng Li 0010, Yong Dou, Jingfei Jiang, Yueqing Wang
IJCNN3
2013 Effect of fixed-point arithmetic on deep belief networks (abstract only)
abstract
Deep Belief Networks (DBNs) are state-of-the-art learning algorithms building on a subset of neural networks, Restricted Boltzmann Machine (RBM). DBNs are computationally intensive posing the question of whether DBNs can be FPGA accelerated. Fixed-point arithmetic can have an important influence on the execution time and prediction accuracy of a DBN. Previous studies have focused only on customized RBM accelerators with a fixed data-width. Our results experiments demonstrate that variable data-widths can obtain similar performance levels. We can also observe that the most suitable data-widths for different types of DBN are not unique or fixed. From this we conclude that a DBN accelerator should support various data-widths rather than only fixed one as done in previous work. The processing performance of DBN accelerators in FPGA is almost always constrained not by the capacity of the processing units, but by their on-chip RAM capacity and speed. We propose an efficient memory sub-system combining junction and padding methods to reduce bandwidth usage for DBN accelerators, which shows that supporting various data-widths is not as difficult as it may sound. The cost is only little in hardware terms and does not affect the critical path. We design a generation tool to help users reconfiguring the memory sub-system with arbitrary data-width flexibly. Our tool can also be used as an advanced IP core generator above FPGA memory controller supporting parallel memory access in irregular data-width for other applications.
Jingfei Jiang, Rongdong Hu, Mikel Luján
FPGA1
2010 An Efficient Coding Scheme for Tolerating Double Disk Failures
abstract
A new MDS array erasure code, called DA-Code, which can tolerate double disk erasures for highly reliable data storage system is proposed in this paper. The DA-Code requires only XOR operations and achieves optimal encoding, updating and decoding complexity. The parity symbols are evenly distributed in the array, overcoming the bottleneck effects of repeated write operation. Detailed DA-Code's decoding algorithm for correcting double disk failures is provided. Analysis result shows that the new coding scheme has excellent performance. Thus, the DA-Code is practically very meaningful for storage systems which need high reliability.
Rongdong Hu, Jingfei Jiang
HPCC3
2010 A Unified Co-Processor Architecture for Matrix Decomposition
Yong Dou, Jie Zhou 0007, Guiming Wu, Jingfei Jiang, Yuanwu Lei, Shi-Ce Ni
J. Comput. Sci. Technol.4
2009 Fine-grained parallel application specific computing for RNA secondary structure prediction using SCFGS on FPGA
abstract
In the field of RNA secondary structure prediction, the CYK (Coche-Younger-Kasami) algorithm is a most popular methods using SCFG (stochastic context-free grammars) model. However, general purpose parallel computers including SMP multiprocessors or cluster systems exhibit low parallel efficiency and they are too expensive to be used easily for many research institutes. FPGA chips provide a new approach to accelerate the CYK algorithm by exploiting fine-grained custom design. The CYK algorithm shows complicated data dependence, in which the dependence distance is variable, and the dependence direction is also across two dimensions. We propose a systolic array structure including one master PE and multiple slave PEs for fine grain hardware implementation on FPGA. We partition tasks by columns and assign tasks to PEs for load balance. We exploit data reuse schemes to reduce the need to load matrix from external memory. To our knowledge, our implementation with 16 PEs is the only FPGA accelerator implementing the complete CYK/inside algorithm. The experimental results show a factor of more than 14 speedup over the Infernal-0.55 software running on a PC platform with Pentium 4 2.66GHz CPU. The computational power of our platform with FPGA accelerator is comparable to a PC cluster consisting of 20 Intel-Xeon CPUs for RNA secondary structure prediction using SCFGs, but the hardware cost and power consumption is only about 15% and 10% of the latter respectively.
Yong Dou, Fei Xia 0003, Jingfei Jiang
CASES3
2009 A Fine-grained Pipelined Implementation of the LINPACK Benchmark on FPGAs
abstract
Previous works have projected that the peak performance of FPGAs can outperform that of the general purpose processors. However, no work actually compares the performance between FPGAs and CPUs using the standard benchmarks such as the LINPACK benchmark. We propose and implement an FPGA-based hardware design of the LINPACK benchmark, the key step of which is LU decomposition with pivoting. We introduce a fine-grained pipelined LU decomposition algorithm that enables optimum performance by exploiting fine-grained pipeline parallelism. A scalable linear array of processing elements (PEs), which is the core component of our hardware design, is proposed to implement this algorithm. To the best of our knowledge, this is the first reported FPGA-based pipelined implementation of LU decomposition with pivoting. A total of 19 PEs can be integrated into an Altera Stratix II EP2S130F1020C5 on our self-designed development board. Experimental results show that the speedup up to 6.14 can be achieved relative to a Pentium 4 processor for the LINPACK benchmark.
Guiming Wu, Yong Dou, Yuanwu Lei, Jie Zhou 0007, Jingfei Jiang
FCCM6