VLDB 2026 Research / reviewers in the wild / expert
Fengbo Ren
dblp:23/11242
· DBLP profile ↗
28ranked-venue papers
1as first author
11since 2021 · last 2024
0000-0002-6509-8753ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 6 since 2021Systems, architecture and hardware · 10 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 7 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Tokenmotion: Motion-Guided Vision Transformer for Video Camouflaged Object Detection VIA Learnable Token SelectionabstractThe area of Video Camouflaged Object Detection (VCOD) presents unique challenges in the field of computer vision due to texture similarities between target objects and their surroundings, as well as irregular motion patterns caused by both objects and camera movement. In this paper, we introduce TokenMotion (TMNet), which employs a transformer-based model to enhance VCOD by extracting motion-guided features using a learnable token selection. Evaluated on the challenging MoCA-Mask dataset, TMNet achieves state-of-the-art performance in VCOD. It outperforms the existing state-of-the-art method by a 12.8% improvement in weighted F-measure, an 8.4% enhancement in S-measure, and a 10.7% boost in mean IoU. The results demonstrate the benefits of utilizing motion-guided features via learnable token selection within a transformer-based framework to tackle the intricate task of VCOD. The code of our work will be available when the paper is accepted. Zifan Yu, Erfan Bank Tavakoli, Meida Chen, Suya You, Raghuveer M. Rao, Sanjeev Agarwal, Fengbo Ren |
ICASSP | 7 |
| 2024 | FSpGEMM: A Framework for Accelerating Sparse General Matrix-Matrix Multiplication Using Gustavson's Algorithm on FPGAsabstractGeneral sparse matrix–matrix multiplication (SpGEMM) is integral to many high-performance computing (HPC) and machine learning applications. However, prior field-programmable gate array (FPGA)-based SpGEMM accelerators either use the inner product algorithm with wasted and costly operations or Gustavson’s algorithm with a cache-based hardware architecture suffering from long-latency cache miss penalties and limited to embedded devices. In this work, we propose framework for accelerating SpGEMM (FSpGEMM), an OpenCL-based SpGEMM framework for accelerating Gustvason’s algorithm that includes an FPGA kernel implementing a throughput-optimized and scalable hardware architecture compatible with high-bandwidth memory (HBM) or traditional DDR-based memory. In addition, to address the irregular memory access patterns incurred by Gustavson’s algorithm, we propose a new buffering scheme tailored to Gustavson’s algorithm enabled by a new compressed sparse vector (CSV) format for representing sparse matrices and a row reordering technique as a preprocessing step to improve data reuse, and consequently, resource utilization. The proposed framework includes a host program implementing preprocessing functions for reordering input matrices and storing them in the proposed CSV format for further use. We implemented FSpGEMM using Intel FPGA SDK for OpenCL and experimented with a benchmark of sparse matrices selected from the SuiteSparse Matrix Collection on a Bittware 520N-MX FPGA board. The results show that the reordering technique improves the performance on average by 20.3% compared with the baseline. Finally, FSpGEMM outperforms the state-of-the-art (SOTA) FPGA implementation by an average of$2.23\times $in terms of execution cycles with the same benchmark and memory system configuration for a fair comparison. Erfan Bank Tavakoli, Michael Riera, Masudul Hassan Quraishi, Fengbo Ren |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2023 | Automatic Error Detection in Integrated Circuits Image Segmentation: A Data-Driven ApproachabstractDue to the complicated nanoscale structures of current integrated circuits(IC) builds and low error tolerance of IC image segmentation tasks, most existing automated IC image segmentation approaches require human experts for visual inspection to ensure correctness, which is one of the major bottlenecks in large-scale industrial applications. In this paper, we present the first data-driven automatic error detection approach that targets two types of IC segmentation errors: wire and via errors. On an IC image dataset collected from real industry, we demonstrate that, by adapting existing CNN-based approaches of image classification and image translation with additional pre-processing and post-processing techniques, we are able to achieve recall/precision of 0.92/0.93 in wire error detection and 0.96/0.90 in via error detection, respectively. Bruno Machado Trindade, Zifan Yu, Chris Pawlowicz, Fengbo Ren |
ICASSP | 6 |
| 2023 | Enhanced Low-Resolution LiDAR-Camera Calibration via Depth Interpolation and Supervised Contrastive LearningabstractMotivated by the increasing application of low-resolution LiDAR, we target the problem of low-resolution LiDAR-camera calibration in this work. The main challenges are two-fold: sparsity and noise in point clouds. To address the problem, we propose to apply depth interpolation to increase the point density and supervised contrastive learning to learn noise-resistant features. The experiments on RELLIS-3D demonstrate that our approach achieves an average mean absolute rotation/translation errors of 0.15cm/0.33° on 32-channel LiDAR point cloud data, which significantly outperforms all reference methods. Zifan Yu, Suya You, Raghuveer M. Rao, Sanjeev Agarwal, Fengbo Ren |
ICASSP | 6 |
| 2023 | TransUPR: A Transformer-based Plug-and-Play Uncertain Point Refiner for LiDAR Point Cloud Semantic SegmentationabstractCommon image-based LiDAR point cloud semantic segmentation (LiDAR PCSS) approaches have bottlenecks resulting from the boundary-blurring problem of convolution neural networks (CNNs) and quantitation loss of spherical projection. In this work, we propose a transformer-based plug-and-play uncertain point refiner, i.e., TransUPR, to refine selected uncertain points in a learnable manner, which leads to an improved segmentation performance. Uncertain points are sampled from coarse semantic segmentation results of 2D image segmentation where uncertain points are located close to the object boundaries in the 2D range image representation and 3D spherical projection background points. Following that, the geometry and coarse semantic features of uncertain points are aggregated by neighbor points in 3D space without adding expensive computation and memory footprint. Finally, the transformer-based refiner, which contains four stacked self-attention layers, along with an MLP module, is utilized for uncertain point classification on the concatenated features of self-attention layers. As the proposed refiner is independent of 2D CNNs, our TransUPR can be easily integrated into any existing image-based LiDAR PCSS approaches, e.g., CENet. Our TransUPR with the CENet achieves state-of-the-art performance, i.e., 68.2% mean Intersection over Union (mIoU) on the Semantic KITTI benchmark, which provides a performance improvement of 0.6% on the mIoU compared to the original CENet. Zifan Yu, Meida Chen, Suya You, Raghuveer M. Rao, Sanjeev Agarwal, Fengbo Ren |
IROS | 7 |
| 2022 | STPLS3D: A Large-Scale Synthetic and Real Aerial Photogrammetry 3D Point Cloud Dataset
Meida Chen, Qingyong Hu, Zifan Yu, Hugues Thomas, Andrew Feng, Kyle McCullough, Fengbo Ren, Lucio Soibelman |
BMVC | 8 |
| 2022 | An Experimental Study on Transferring Data-Driven Image Compressive Sensing to Bioelectric SignalsabstractThe emerging area of bioelectric signal compressive sensing(CS) has shown great potential in health care applications. However, improving the reconstruction accuracy of compressively sensed bioelectric signals remains a challenging problem. In recent years, data-driven image CS methods have achieved significant improvements in reconstruction accuracy over conventional model-based image CS methods. In this paper, we conduct an experimental study on transferring existing data-driven image CS methods to bioelectric signals. Through our investigation of five critical factors affecting the reconstruction performance of bioelectric signals, we conclude that existing data-driven image CS methods can be transferred to ECG signals with high reconstruction accuracy. Our experimental results show that transferred data-driven image CS methods can achieve up to 8.08-2.73 SNR improvement over the reference method on ECG signal reconstruction across compression ratios of 2-8x. Jonathan Zhao, Fengbo Ren |
ICASSP | 3 |
| 2022 | A Data-Driven Approach for Automated Integrated Circuit Segmentation of Scan Electron Microscopy ImagesabstractThis paper proposes an automated data-driven integrated circuit segmentation approach of scan electron microscopy (SEM) images inspired by state-of-the-art CNN-based image perception methods. Based on the requirements derived from real industry applications, we take wire segmentation and via detection algorithms to generate integrated circuit segmentation maps from SEMs in our approach. On SEM images collected in the industrial applications, our method achieves an average of 50.71 on Electrically Significant Difference (ESD) in the wire segmentation task and 99.05% F1 score in the via detection task, which achieves about 85% and 8% improvements over the reference method, respectively. Zifan Yu, Bruno Machado Trindade, Pullela Sneha, Erfan Bank Tavakoli, Chris Pawlowicz, Fengbo Ren |
ICIP | 8 |
| 2021 | Characterizing Loop Acceleration in Heterogeneous ComputingabstractComputation intensive applications usually consist of multiple nested or flattened loops. These loops are the main building blocks of the applications and embody a specific type of execution pattern. In order to reduce the running time of the loops, developers need to analyze the loops in the code and try to parallelize them on hardware accelerators, such as GPUs, TPUs, and FPGAs, which are increasingly available in the cloud. Unfortunately, the lack of understanding of loop characteristics and the ability of hardware accelerators in handling these types of loops prevents developers from choosing the right platform to develop their applications in the cloud. Also, developing and optimizing code for a specific accelerator is a time-consuming effort. To address these issues, this paper studies the effectiveness of different processors in accelerating common patterns of loops. It identifies five important types of loops that commonly exist in real-world applications, and presents Loopy, the implementations of these loops optimized for different architectures. Using Loopy, the paper also evaluates different hardware in accelerating the loop patterns. The result reveals the architectural differences among different accelerators with regard to different loop patterns. It also provides insights for the developers to choose the right accelerators for their applications. The current version of Loopy supports both FPGAs and GPUs, which are the most versatile and available accelerators. Saman Biookaghazadeh, Fengbo Ren, Ming Zhao 0002 |
CLOUD | 2 |
| 2021 | FSCHOL: An OpenCL-based HPC Framework for Accelerating Sparse Cholesky Factorization on FPGAsabstractThe proposed FSCHOL framework consists of an FPGA kernel implementing a throughput-optimized hardware architecture for accelerating the supernodal multifrontal algorithm for sparse Cholesky factorization and a host program implementing a novel scheduling algorithm for finding the optimal execution order of supernodes computations for an elimination tree on the FPGA to eliminate the need for off-chip memory access for storing intermediate results. Moreover, the proposed scheduling algorithm minimizes on-chip memory requirements for buffering intermediate results by resolving the dependency of parent nodes in an elimination tree through temporal parallelism. Experiment results for factorizing a set of sparse matrices in various sizes from SuiteSparse Matrix Collection show that the proposed FSCHOL implemented on an Intel Stratix 10 GX FPGA development board achieves on average 5.5× and 9.7× higher performance and 10.3× and 24.7× lower energy consumption than implementations of CHOLMOD on an Intel Xeon E5-2637 CPU and an NVIDIA V100 GPU, respectively. Erfan Bank Tavakoli, Michael Riera, Masudul Hassan Quraishi, Fengbo Ren |
SBAC-PAD | 4 |
| 2021 | A Survey of System Architectures and Techniques for FPGA VirtualizationabstractFPGA accelerators are gaining increasing attention in both cloud and edge computing because of their hardware flexibility, high computational throughput, and low power consumption. However, the design flow of FPGAs often requires specific knowledge of the underlying hardware, which hinders the wide adoption of FPGAs by application developers. Therefore, the virtualization of FPGAs becomes extremely important to create a useful abstraction of the hardware suitable for application developers. Such abstraction also enables the sharing of FPGA resources among multiple users and accelerator applications, which is important because, traditionally, FPGAs have been mostly used in single-user, single-embedded-application scenarios. There are many works in the field of FPGA virtualization covering different aspects and targeting different application areas. In this article, we review the system architectures used in the literature for FPGA virtualization. In addition, we identify the primary objectives of FPGA virtualization, based on which we summarize the techniques for realizing FPGA virtualization. This article helps researchers to efficiently learn about FPGA virtualization research by providing a comprehensive review of the existing literature. Masudul Hassan Quraishi, Erfan Bank Tavakoli, Fengbo Ren |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2020 | Learning in the Frequency DomainabstractDeep neural networks have achieved remarkable success in computer vision tasks. Existing neural networks mainly operate in the spatial domain with fixed input sizes. For practical applications, images are usually large and have to be downsampled to the predetermined input size of neural networks. Even though the downsampling operations reduce computation and the required communication bandwidth, it removes both redundant and salient information obliviously, which results in accuracy degradation. Inspired by digital signal processing theories, we analyze the spectral bias from the frequency perspective and propose a learning-based frequency selection method to identify the trivial frequency components which can be removed without accuracy loss. The proposed method of learning in the frequency domain leverages identical structures of the well-known neural networks, such as ResNet-50, MobileNetV2, and Mask R-CNN, while accepting the frequency-domain information as the input. Experiment results show that learning in the frequency domain with static channel selection can achieve higher accuracy than the conventional spatial downsampling approach and meanwhile further reduce the input data size. Specifically for ImageNet classification with the same input size, the proposed method achieves 1.60% and 0.63% top-1 accuracy improvements on ResNet-50 and MobileNetV2, respectively. Even with half input size, the proposed method still improves the top-1 accuracy on ResNet-50 by 1.42%. In addition, we observe a 0.8% average precision improvement on Mask R-CNN for instance segmentation on the COCO dataset. Kai Xu 0007, Minghai Qin, Fei Sun 0002, Yuhao Wang 0002, Yen-Kuang Chen, Fengbo Ren |
CVPR | 6 |
| 2020 | Systolic-CNN: An OpenCL-defined Scalable Run-time-flexible FPGA Accelerator Architecture for Accelerating Convolutional Neural Network Inference in Cloud/Edge ComputingabstractThis paper presents Systolic-CNN, an OpenCLdefined scalable, run-time-flexible FPGA accelerator architecture, optimized for performing the low-latency, energy-efficient inference of various convolutional neural networks (CNNs) in the context of multi-tenancy cloud/edge computing. Systolic-CNN adopts a highly pipelined and parallelized 1-D systolic array architecture, which efficiently explores both spatial and temporal parallelism for accelerating CNN inference on FPGAs. SystolicCNN is highly scalable and parameterized, which can be easily adapted by users to achieve 100% utilization of the coarsegrained computation resources (i.e., DSP blocks) for a given FPGA. In addition, Systolic-CNN is run-time-flexible, which can be time-shared, in the context of multi-tenancy cloud or edge computing, to accelerate a variety of CNN models at run time without the need of recompiling the FPGA kernel hardware nor reprogramming the FPGA. The experiment results based on an Intel Arria 10 GX FPGA Development board show that Systolic-CNN, when mapped with a single-precision data format, can achieve 100% utilization of the DSP block resource and an average inference latency of 10ms, 84ms, 1615ms, and 990ms per image for accelerating AlexNet, ResNet-50, RetinaNet, and Light-weight RetinaNet, respectively. The peak computational throughput is measured at 80–170 GFLOPS/s across the acceleration of different CNN models. Codes are available at https://github.com/PSCLab-ASU/SystolicCNN. Akshay Dua, Yixing Li, Fengbo Ren |
FCCM | 3 |
| 2020 | Cra: A Generic Compression Ratio Adapter for End-To-End Data-Driven Image Compressive Sensing Reconstruction FrameworksabstractEnd-to-end data-driven image compressive sensing reconstruction (EDCSR) frameworks achieve state-of-the-art reconstruction performance in terms of reconstruction speed and accuracy. However, due to their end-to-end nature, existing EDCSR frameworks can not adapt to a variable compression ratio (CR). For applications that desire a variable CR, existing EDCSR frameworks must be trained from scratch at each CR, which is computationally costly and time-consuming. This paper presents a generic compression ratio adapter (CRA) framework that addresses the variable compression ratio (CR) problem for existing EDCSR frameworks with no modification to given reconstruction models nor enormous rounds of training needed. CRA exploits an initial reconstruction network to generate an initial estimate of reconstruction results based on a small portion of the acquired measurements. Subsequently, CRA approximates full measurements for the main reconstruction network by complementing the sensed measurements with resensed initial estimate. Our experiments based on two public image datasets (CIFAR10 and Set5) show that CRA provides an average of 13.02 dB and 5.38 dB PSNR improvement across the CRs from 5 to 30 over a naive zero-padding approach and the AdaptiveNN approach(a prior work), respectively. CRA addresses the fixed-CR limitation of existing EDCSR frameworks and makes them suitable for resource-constrained compressive sensing applications. Kai Xu 0007, Fengbo Ren |
ICASSP | 3 |
| 2020 | MoNet3D: Towards Accurate Monocular 3D Object Localization in Real TimeabstractMonocular multi-object detection and localization in 3D space has been proven to be a challenging task. The MoNet3D algorithm is a novel and effective framework that can predict the 3D position of each object in a monocular image, and draw a 3D bounding box on each object. The MoNet3D method incorporates the prior knowledge of spatial geometric correlation of neighboring objects into the deep neural network training process, in order to improve the accuracy of 3D object localization. Experiments over the KITTI data set show that the accuracy of predicting the depth and horizontal coordinate of the object in 3D space can reach 96.25% and 94.74%, respectively. Meanwhile, the method can realize the real-time image processing capability of 27.85 FPS. Our code is publicly available at https://github.com/CQUlearningsystemgroup/YicongPeng Xichuan Zhou, Yicong Peng, Chunqiao Long, Fengbo Ren, Cong Shi 0003 |
ICML | 4 |
| 2020 | Build a compact binary neural network through bit-level sensitivity and data pruning
Yixing Li, Xichuan Zhou, Fengbo Ren |
Neurocomputing | 4 |
| 2019 | A Deep Learning Approach for Targeted Contrast-Enhanced Ultrasound Based Prostate Cancer DetectionabstractThe important role of angiogenesis in cancer development has driven many researchers to investigate the prospects of noninvasive cancer diagnosis based on the technology of contrast-enhanced ultrasound (CEUS) imaging. This paper presents a deep learning framework to detect prostate cancer in the sequential CEUS images. The proposed method uniformly extracts features from both the spatial and the temporal dimensions by performing three-dimensional convolution operations, which captures the dynamic information of the perfusion process encoded in multiple adjacent frames for prostate cancer detection. The deep learning models were trained and validated against expert delineations over the CEUS images recorded using two types of contrast agents, i.e., the anti-PSMA based agent targeted to prostate cancer cells and the non-targeted blank agent. Experiments showed that the deep learning method achieved over 91 percent specificity and 90 percent average accuracy over the targeted CEUS images for prostate cancer detection, which was superior ( ) than previously reported approaches and implementations. Fan Yang 0021, Xichuan Zhou, Yanli Guo, Fang Tang, Fengbo Ren, Jishun Guo, Shuiwang Ji |
IEEE ACM Trans. Comput. Biol. Bioinform. | 6 |
| 2018 | SqueezedText: A Real-Time Scene Text Recognition by Binary Convolutional Encoder-Decoder NetworkabstractA new approach for real-time scene text recognition is proposed in this paper. A novel binary convolutional encoder-decoder network (B-CEDNet) together with a bidirectional recurrent neural network (Bi-RNN). The B-CEDNet is engaged as a visual front-end to provide elaborated character detection, and a back-end Bi-RNN performs character-level sequential correction and classification based on learned contextual knowledge. The front-end B-CEDNet can process multiple regions containing characters using a one-off forward operation, and is trained under binary constraints with significant compression. Hence it leads to both remarkable inference run-time speedup as well as memory usage reduction. With the elaborated character detection, the back-end Bi-RNN merely processes a low dimension feature sequence with category and spatial information of extracted characters for sequence correction and classification. By training with over 1,000,000 synthetic scene text images, the B-CEDNet achieves a recall rate of 0.86, precision of 0.88 and F-score of 0.87 on ICDAR-03 and ICDAR-13. With the correction and classification by Bi-RNN, the proposed real-time scene text recognition achieves state-of-the-art accuracy while only consumes less than 1-ms inference run-time. The flow processing flow is realized on GPU with a small network size of 1.01 MB for B-CEDNet and 3.23 MB for Bi-RNN, which is much faster and smaller than the existing solutions. Zichuan Liu, Yixing Li, Fengbo Ren, Wang Ling Goh, Hao Yu 0001 |
AAAI | 3 |
| 2018 | LAPRAN: A Scalable Laplacian Pyramid Reconstructive Adversarial Network for Flexible Compressive Sensing Reconstruction
Kai Xu 0007, Fengbo Ren |
ECCV (10) | 3 |
| 2018 | CSVideoNet: A Real-Time End-to-End Learning Framework for High-Frame-Rate Video Compressive SensingabstractThis paper addresses the real-time encoding-decoding problem for high-frame-rate video compressive sensing (CS). Unlike prior works that perform reconstruction using iterative optimization-based approaches, we propose a noniterative model, named "CSVideoNet", which directly learns the inverse mapping of CS and reconstructs the original input in a single forward propagation. To overcome the limitations of existing CS cameras, we propose a multi-rate CNN and a synthesizing RNN to improve the trade-o. between compression ratio (CR) and spatial-temporal resolution of the reconstructed videos. the experiment results demonstrate that CSVideoNet signi.cantly outperforms state-of-the-art approaches. Without any pre/post-processing, we achieve a 25dB Peak signal-to-noise ratio (PSNR) recovery quality at 100x CR, with a frame rate of 125 fps on a Titan X GPU. Due to the feedforward and high-data-concurrency natures of CSVideoNet, it can take advantage of GPU acceleration to achieve three orders of magnitude speed-up over conventional iterative-based approaches. We share the source code at https://github.com/PSCLab-ASU/CSVideoNet. Kai Xu 0007, Fengbo Ren |
WACV | 2 |
| 2018 | A GPU-Outperforming FPGA Accelerator Architecture for Binary Convolutional Neural NetworksabstractFPGA-based hardware accelerators for convolutional neural networks (CNNs) have received attention due to their higher energy efficiency than GPUs. However, it is challenging for FPGA-based solutions to achieve a higher throughput than GPU counterparts. In this article, we demonstrate that FPGA acceleration can be a superior solution in terms of both throughput and energy efficiency when a CNN is trained with binary constraints on weights and activations. Specifically, we propose an optimized fully mapped FPGA accelerator architecture tailored for bitwise convolution and normalization that features massive spatial parallelism with deep pipelines stages. A key advantage of the FPGA accelerator is that its performance is insensitive to data batch size, while the performance of GPU acceleration varies largely depending on the batch size of the data. Experiment results show that the proposed accelerator architecture for binary CNNs running on a Virtex-7 FPGA is 8.3× faster and 75× more energy-efficient than a Titan X GPU for processing online individual requests in small batch sizes. For processing static data in large batch sizes, the proposed solution is on a par with a Titan X GPU in terms of throughput while delivering 9.5× higher energy efficiency. Yixing Li, Zichuan Liu, Kai Xu 0007, Hao Yu 0001, Fengbo Ren |
ACM J. Emerg. Technol. Comput. Syst. | 5 |
| 2017 | A 7.663-TOPS 8.2-W Energy-efficient FPGA Accelerator for Binary Convolutional Neural Networks (Abstract Only)
Yixing Li, Zichuan Liu, Kai Xu 0007, Hao Yu 0001, Fengbo Ren |
FPGA | 5 |
| 2017 | A data-driven compressive sensing framework tailored for energy-efficient wearable sensingabstractCompressive sensing (CS) is a promising technology for realizing energy-efficient wireless sensors for long-term health monitoring. However, conventional model-driven CS frameworks suffer from limited compression ratio and reconstruction quality when dealing with physiological signals due to inaccurate models and the overlook of individual variability. In this paper, we propose a data-driven CS framework that can learn signal characteristics and personalized features from any individual recording of physiologic signals to enhance CS performance with a minimized number of measurements. Such improvements are accomplished by a co-training approach that optimizes the sensing matrix and the dictionary towards improved restricted isometry property and signal sparsity, respectively. Experimental results upon ECG signals show that the proposed method, at a compression ratio of 10×, successfully reduces the isometry constant of the trained sensing matrices by 86% against random matrices and improves the overall reconstructed signal-to-noise ratio by 15dB over conventional model-driven approaches. Kai Xu 0007, Yixing Li, Fengbo Ren |
ICASSP | 3 |
| 2016 | An energy-efficient compressive sensing framework incorporating online dictionary learning for long-term wireless health monitoringabstractWireless body area network (WBAN) is emerging in the mobile healthcare area to replace the traditional wire-connected monitoring devices. As wireless data transmission dominates power cost of sensor nodes, it is beneficial to reduce the data size without much information loss. Compressive sensing (CS) is a perfect candidate to achieve this goal compared to existing compression techniques. In this paper, we proposed a general framework that utilize CS and online dictionary learning (ODL) together. The learned dictionary carries individual characteristics of the original signal, under which the signal has an even sparser representation compared to pre-determined dictionaries. As a consequence, the compression ratio is effectively improved by 2-4× comparing to prior works. Besides, the proposed framework offloads pre-processing from sensor nodes to the server node prior to dictionary learning, providing further reduction in hardware costs. As it is data driven, the proposed framework has the potential to be used with a wide range of physiological signals. Kai Xu 0007, Yixing Li, Fengbo Ren |
ICASSP | 3 |
| 2016 | A Compressive-sensing based Testing Vehicle for 3D TSV Pre-bond and Post-bond Testing DataabstractOnline testing vehicle is required for 3D TSV pre-bond and post-bond testing due to high probability of TSV failures. It has become a challenge to deal with large sets of generated testing data with limited probing when transmitting the data out. In this paper, a lossless compressive-sensing based testing vehicle is developed for online testing of TSVs. By exploring sparsity of the testing data under constraint of failure bound of TSV, sparse-representation based encoding can be deployed by XOR and AND network on chip to deal with large volume of testing data. Experimental results (with benchmarks) have shown that 89.70% pre-bond data compression rate can be achieved under 0.5% probability of failures; and 88.18% post-bond data compression rate can be achieved with 5% probability of failures. Hantao Huang, Hao Yu 0001, Cheng Zhuo, Fengbo Ren |
ISPD | 4 |
| 2014 | A scalable sparse matrix-vector multiplication kernel for energy-efficient sparse-blas on FPGAsabstractSparse Matrix-Vector Multiplication (SpMxV) is a widely used mathematical operation in many high-performance scientific and engineering applications. In recent years, tuned software libraries for multi-core microprocessors (CPUs) and graphics processing units (GPUs) have become the status quo for computing SpMxV. However, the computational throughput of these libraries for sparse matrices tends to be significantly lower than that of dense matrices, mostly due to the fact that the compression formats required to efficiently store sparse matrices mismatches traditional computing architectures. This paper describes an FPGA-based SpMxV kernel that is scalable to efficiently utilize the available memory bandwidth and computing resources. Benchmarking on a Virtex-5 SX95T FPGA demonstrates an average computational efficiency of 91.85%. The kernel achieves a peak computational efficiency of 99.8%, a >50x improvement over two Intel Core i7 processors (i7-2600 and i7-4770) and showing a >300x improvement over two NVIDA GPUs (GTX 660 and GTX Titan), when running the MKL and cuSPARSE sparse-BLAS libraries, respectively. In addition, the SpMxV FPGA kernel is able to achieve higher performance than its CPU and GPU counterparts, while using only 64 single-precision processing elements, with an overall 38-50x improvement in energy efficiency. Richard Dorrance, Fengbo Ren, Dejan Markovic |
FPGA | 2 |
| 2014 | Statistical timing and power analysis of VLSI considering non-linear dependence
Lerong Cheng, Wenyao Xu, Fengbo Ren, Fang Gong, Puneet Gupta 0001, Lei He 0001 |
Integr. | 3 |
| 2013 | A single-precision compressive sensing signal reconstruction engine on FPGAsabstractCompressive sensing (CS) is a promising technology for the low-power and cost-effective data acquisition in wireless healthcare systems. However, its efficient realtime signal reconstruction is still challenging, and there is a clear demand for hardware acceleration. In this paper, we present the first single-precision floating-point CS reconstruction engine implemented a Kintex-7 FPGA using the orthogonal matching pursuit (OMP) algorithm. In order to achieve high performance with maximum hardware utilization, we propose a highly parallel architecture that shares the computing resources among different tasks of OMP by using configurable processing elements (PEs). By fully utilizing the FPGA recourses, our implementation has 128 PEs in parallel and operates at 53.7 MHz. In addition, it can support 2x larger problem size and 10x more sparse coefficients than prior work, which enables higher reconstruction accuracy by adding finer details to the recovered signal. Hardware results from the ECG reconstruction tests show the same level of accuracy as the double-precision C program. Compared to the execution time of a 2.27 GHz CPU, the FPGA reconstruction achieves an average speed-up of 41x. Fengbo Ren, Richard Dorrance, Wenyao Xu, Dejan Markovic |
FPL | 1 |