Haibin Shen

dblp:95/4224 · DBLP profile ↗
← Back
33ranked-venue papers
3as first author
21since 2021 · last 2026
0000-0002-5431-609XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 9 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Security and privacy · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-authorTheory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2026 A Bit-Level Loosely Coupled Spiking Neural Network Accelerator With Fast Inference and Hybrid Early Termination
abstract
Deep spiking neural networks (SNNs) tend to suffer from long inference latency. While existing encoding schemes are insufficient, this work applies a bit-level loosely coupled (BLC) SNN in which output bits are generated sequentially from the current input bit. Consequently, the time steps can be further reduced with minor accuracy loss. The BLC SNN is optimized for hardware implementation and configured based on quantized artificial neural networks (ANNs). A hybrid early termination (ET) scheme is employed to skip redundant computation cycles without accuracy degradation. A pipelined digital accelerator architecture is designed to implement the BLC SNN, in which line buffers are utilized to maximize data reuse. Simulation results show that the proposed hybrid ET scheme reduces computation cycles by 28.31% on LeNet-5. Simulated in SMIC 28 nm technology, the accelerator consumes$0.10~\mu $J/img at 500 MHz with 30.49 TSOPS/W and is adaptable to 8/4 time steps of inference.
Junchuan Gu, Yiwen Gu, Haibin Shen, Kejie Huang
IEEE Trans. Very Large Scale Integr. Syst.4
2025 MVQ: Towards Efficient DNN Compression and Acceleration with Masked Vector Quantization
abstract
Vector quantization(VQ) is a hardware-friendly DNN compression method that can reduce the storage cost and weight-loading datawidth of hardware accelerators. However, conventional VQ techniques lead to significant accuracy loss because the important weights are not well preserved. To tackle this problem, a novel approach called MVQ is proposed, which aims at better approximating important weights with a limited number of codewords. At the algorithm level, our approach removes the less important weights through N:M pruning and then minimizes the vector clustering error between the remaining weights and codewords by the masked k-means algorithm. Only distances between the unpruned weights and the codewords are computed, which are then used to update the codewords. At the architecture level, our accelerator implements vector quantization on an EWS (Enhanced weight stationary) CNN accelerator and proposes a sparse systolic array design to maximize the benefits brought by masked vector quantization.
Shuaiting Li, Chengxuan Wang, Juncan Deng, Zeyu Wang 0010, Zewen Ye, Zongsheng Wang, Haibin Shen, Kejie Huang
ASPLOS (1)7
2025 A FeFET-Based Compute-in-Memory Architecture on FPGA for Neural Network Inference
abstract
Implementing compute-in-memory (CIM) architectures on FPGA offers an effective solution to the von Neumann bottleneck by enabling fast configuration and computation directly within memory. Traditional custom solutions rely on the modification of block RAM (BRAM) to implement memory computing. However, single-word-line activation of BRAM results in low parallelism, and the need for additional adder trees to accumulate partial sums further limits efficiency. To overcome these limitations, we propose a CIM core based on a 2T1C structure as a replacement for BRAM units. This core utilizes a charge redistribution mechanism and reuse of ADC capacitors, achieving high parallelism, low power consumption, and a compact area. By incorporating computational capabilities within a single cell, our design enables dual parallelism, further enhancing performance and efficiency. In addition, we present an automated deployment and mapping tool for deep neural networks (DNNs) on FPGA, allowing users to rapidly develop FPGA-based solutions for different network architectures. Compared to state-of-the-art solutions, our design achieves a peak throughput improvement of 4.5× and a reduction in area by 53%.
Minghan Jiang, Yonggen Li, Rui Xiao 0003, Haibin Shen, Kejie Huang
FCCM4
2025 SSVQ: Unleashing the Potential of Vector Quantization with Sign-Splitting
abstract
Vector Quantization (VQ) has emerged as a prominent weight compression technique, showcasing substantially lower quantization errors than uniform quantization across diverse models, particularly in extreme compression scenarios. However, its efficacy during fine-tuning is limited by the constraint of the compression format, where weight vectors assigned to the same codeword are restricted to updates in the same direction. Consequently, many quantized weights are compelled to move in directions contrary to their local gradient information. To mitigate this issue, we introduce a novel VQ paradigm, Sign-Splitting VQ (SSVQ), which decouples the sign bit of weights from the codebook. Our approach involves extracting the sign bits of uncompressed weights and performing clustering and compression on all-positive weights. We then introduce latent variables for the sign bit and jointly optimize both the signs and the codebook. Additionally, we implement a progressive freezing strategy for the learnable sign to ensure training stability. Extensive experiments on various modern models and tasks demonstrate that SSVQ achieves a significantly superior compression-accuracy trade-off compared to conventional VQ. Furthermore, we validate our algorithm on a hardware accelerator, showing that SSVQ achieves a 3$\times$ speedup over the 8-bit compressed model by reducing memory access. Our code is available at https://github.com/list0830/SSVQ.
Shuaiting Li, Juncan Deng, Chengxuan Wang, Kedong Xu, Rongtao Deng, Haibin Shen, Kejie Huang
ICCV7
2025 A 1FeFET-1T-1C based Compute-in-Memory Macro with Capacitor Reused Pipeline SAR ADC
abstract
Computing-in-memory (CIM) significantly reduces latency and power consumption by combining computation and memory, typically utilizing non-volatile memories (NVM). However, device manufacturing non-uniformity on NVMs can cause output deviations. Additionally, the necessity for bit-shifting circuits and Analog-to-Digital Converters (ADC) increases the area and power overhead. To tackle these challenges, we propose a high-density 1FeFET-1T-1C based CIM macro, integrated with a pipeline Successive-Approximation-Register (SAR) ADC. The design introduces a capacitor structure that counters the non-uniformity issues inherent in FeFET devices. Also, the capacitor array is reused as charge-redistribution and ADCs, substantially minimizing the area and power overhead. Moreover, the pipeline architecture accelerates the conversion process, achieving high speed and high precision. The design is implemented using SMIC 55nm PDK. The energy efficiency (EF) and area efficiency (AF) of the proposed macro are 80.9 TOPS/W and 1.161 TOPS/mm2, respectively. The inference accuracy reaches 91.2% on the CIFAR-10 dataset.
Minghan Jiang, Rui Xiao 0003, Shuaiting Li, Yishu Zhang, Haibin Shen, Kejie Huang
ISCAS7
2025 A Robust Computing-in-Memory Macro With 2T1R1C Cells and Reused Capacitors for Successive-Approximation ADC
abstract
Computing-in-memory (CIM) has emerged as a practical paradigm to bypass the von Neumann bottleneck. However, traditional CIM schemes face challenges due to the nonideal characteristics of nonvolatile memory (NVM). To address this issue, this work provides a resistive random access memory (RRAM)-based CIM macro employing two-transistor-one-RRAM–one-capacitor (2T1R1C) cells, with capacitors reused for the successive-approximation analog-to-digital converter (SAR ADC). Single-level RRAM is utilized to mitigate resistance variation. The multiply-accumulate (MAC) operation is performed via the charge and discharge of capacitors, enhancing robustness across different process, voltage, and temperature (PVT) corners. The capacitors in 2T1R1C cells are repurposed as sampling capacitors to integrate the ADC with the array. A precision-adjustable SAR (PA-SAR) logic is proposed to generate partial sums at varying precision levels aligned with different input bits, optimizing energy efficiency while maintaining reliability. Our proposed 2T1R1C array features an average area of$3.403~\mu $m2 for each cell, which accounts for 87.46% of the total macro area. The total macro area is 1.020 mm2 with a capacity of 256 Kb, achieving an energy density of 0.201 TOPS/mm2. The PA-SAR logic boosts energy efficiency to 44.71 TOPS/W, marking a 38.55% improvement over conventional full-precision schemes.
Rui Xiao 0003, Minghan Jiang, Haibin Shen, Kejie Huang
IEEE Trans. Very Large Scale Integr. Syst.4
2024 A Folded Computation-in-Memory Accelerator for Fast Polynomial Multiplication in BIKE
Chuhui Wang, Zewen Ye, Haibin Shen, Kejie Huang
Euro-Par (2)3
2024 Bridging partial-gated convolution with transformer for smooth-variation image inpainting
Zeyu Wang 0010, Haibin Shen, Kejie Huang
Multim. Tools Appl.2
2024 An All-digital Compute-in-memory FPGA Architecture for Deep Learning Acceleration
abstract
Field Programmable Gate Array (FPGA) is a versatile and programmable hardware platform, which makes it a promising candidate for accelerating Deep Neural Networks (DNNs). However, FPGA’s computing energy efficiency is low due to the domination of energy consumption by interconnect data movement. In this article, we propose an all-digital Compute-in-memory FPGA architecture for deep learning acceleration. Furthermore, we present a bit-serial computing circuit of the Digital CIM core for accelerating vector-matrix multiplication (VMM) operations. A Network-CIM-deployer ( NCIMD ) is also developed to support automatic deployment and mapping of DNN networks. NCIMD provides a user-friendly API of DNN models in Caffe format. Meanwhile, we introduce a Weight-stationary dataflow and describe the method of mapping a single layer of the network to the CIM array in the architecture. We conduct experimental tests on the proposed FPGA architecture in the field of Deep Learning (DL), as well as in non-DL fields, using different architectural layouts and mapping strategies. We also compare the results with the conventional FPGA architecture. The experimental results show that compared to the conventional FPGA architecture, the energy efficiency can achieve a maximum speedup of 16.1×, while the latency can decrease up to 40% in our proposed CIM FPGA architecture.
Yonggen Li, Xin Li 0177, Haibin Shen, Jicong Fan 0002, Yanfeng Xu, Kejie Huang
ACM Trans. Reconfigurable Technol. Syst.3
2023 Thermal Infrared Image Inpainting Via Edge-Aware Guidance
abstract
Image inpainting has achieved fundamental advances with deep learning. However, almost all existing inpainting methods aim to process natural images, while few target Thermal Infrared (TIR) images, which have widespread applications. When applied to TIR images, conventional inpainting methods usually generate distorted or blurry content. In this paper, we propose a novel task—Thermal Infrared Image Inpainting, which aims to reconstruct missing regions of TIR images. Crucially, we propose a novel deep-learning-based model TIR-Fill. We adopt the edge generator to complete the canny edges of broken TIR images. The completed edges are projected to the normalization weights and biases to enhance edge awareness of the model. In addition, a refinement network based on gated convolution is employed to improve TIR image consistency. The experiments demonstrate that our method outperforms state-of-the-art image inpainting approaches on FLIR thermal dataset.
Zeyu Wang 0010, Haibin Shen, Changyou Men, Kejie Huang
ICASSP2
2023 WaveIPT: Joint Attention and Flow Alignment in the Wavelet domain for Pose Transfer
abstract
Human pose transfer aims to synthesis a new image of the source person in a target pose. Among the various existing methods, attention and flow have emerged as two of the most popular and effective approaches. Attention excels at preserving the semantic structure of the source image, which is more reflected in the low-frequency domain. Contrastively, flow is better at retaining fine-grained texture details in the high-frequency domain. To leverage the advantages of both attention and flow simultaneously, this paper proposes Wavelet-aware Image-based Pose Transfer (WaveIPT) as a novel approach to fuse the attention and flow in the wavelet domain. To improve the fusion effect and avoid interference from irrelevant information across different frequencies, WaveIPT first applies Intra-scale Local Correlation (ILC) to adaptively fuse attention and flow in the same scale according to their strengths in low and high-frequency domains. Subsequently, WaveIPT employs Inter-scale Feature Interaction (IFI) to explore inter-scale frequency features, facilitating effective information transfer across different scales. Furthermore, we introduce Progressive Flow Regularization (PFR), an effective method that alleviates the challenges of flow estimation under large pose differences. The experiments on the DeepFashion dataset demonstrate that WaveIPT achieves a new state-of-the-art in terms of both FID and LPIPS, with improvements of 4.97% and 3.89%, respectively.
Tingwei Gao, Haitian Jiang, Haibin Shen, Kejie Huang
ICCV4
2023 OAW-GAN: occlusion-aware warping GAN for unified human video synthesis
Dongxu Wei, Kejie Huang, Jiashen Hua, Baisheng Lai, Haibin Shen
Appl. Intell.6
2023 MsVRL: Self-Supervised Multiscale Visual Representation Learning via Cross-Level Consistency for Medical Image Segmentation
abstract
Automated medical image segmentation for organs or lesions plays an essential role in clinical diagnoses and treatment plannings. However, training an accurate and robust segmentation model is still a long-standing challenge due to the time-consuming and expertise-intensive annotations for training data, especially 3-D medical images. Recently, self-supervised learning emerges as a promising approach for unsupervised visual representation learning, showing great potential to alleviate the expertise annotations for medical images. Although global representation learning has attained remarkable results on iconic datasets, such as ImageNet, it can not be applied directly to medical image segmentation, because the segmentation task is non-iconic, and the targets always vary in physical scales. To address these problems, we propose a Multi-scale Visual Representation self-supervised Learning (MsVRL) model, to perform finer-grained representation and deal with different target scales. Specifically, a multi-scale representation conception, a canvas matching method, an embedding pre-sampling module, a center-ness branch, and a cross-level consistent loss are introduced to improve the performance. After pre-trained on unlabeled datasets (RibFrac and part of MSD), MsVRL performs downstream segmentation tasks on labeled datasets (BCV, spleen of MSD, and KiTS). Results of the experiments show that MsVRL outperforms other state-of-the-art works on these medical image segmentation tasks.
Ruifeng Zheng, Senxiang Yan, Hongcheng Sun, Haibin Shen, Kejie Huang
IEEE Trans. Medical Imaging5
2023 FDA-GAN: Flow-Based Dual Attention GAN for Human Pose Transfer
abstract
Human pose transfer aims at transferring the appearance of the source person to the target pose. Existing methods utilizing flow-based warping for non-rigid human image generation have achieved great success. However, they fail to preserve the appearance details in synthesized images since the spatial correlation between the source and target is not fully exploited. To this end, we propose the Flow-based Dual Attention GAN (FDA-GAN) to apply occlusion- and deformation-aware feature fusion for higher generation quality. Specifically, deformable local attention and flow similarity attention, constituting the dual attention mechanism, can derive the output features responsible for deformable- and occlusion-aware fusion, respectively. Besides, to maintain the pose and global position consistency in transferring, we design a pose normalization network for learning adaptive normalization from the target pose to the source person. Both qualitative and quantitative results show that our method outperforms state-of-the-art models in public iPER and DeepFashion datasets.
Kejie Huang, Dongxu Wei, Zhaoyan Ming, Haibin Shen
IEEE Trans. Multim.5
2023 A Low-Power In-Memory Multiplication and Accumulation Array With Modified Radix-4 Input and Canonical Signed Digit Weights
abstract
Data transfer between the processing and storage units has become a significant bottleneck in modern von Neumann computing systems for artificial intelligence (AI) tasks. Computing in memory (CIM) has emerged as a promising candidate for lowering latency and power consumption. However, the conventional analog CIM schemes are suffering from reliability issues, which may significantly degenerate the accuracy of the computation. Recently, digitized input data and weights have been utilized for high-reliable in-memory computing. However, the properties of the digital memory and input data are not fully utilized. This article presents a novel low-power CIM scheme to further reduce the power consumption by using a modified radix-4 (M-RD4) booth algorithm at the input and a modified canonical signed digit (M-CSD) for the network weights. The simulation results show that M-RD4 and M-CSD reduce the number of nonzero activation bits by 24.2% and the number of nonzero weight bits by 36.0% in AlexNet, respectively. The power consumption can be reduced by 41.6% on average. The computing-power ratio at the fixed-point 8 bit is 60.7 tera operations per second per watt (TOPS/W), and the density is 0.177 TOPS/mm2.
Rui Xiao 0003, Yewei Zhang, Bo Wang 0020, Yanfeng Xu, Jicong Fan 0002, Haibin Shen, Kejie Huang
IEEE Trans. Very Large Scale Integr. Syst.6
2022 A Reconfigurable Convolution-in-Pixel CMOS Image Sensor Architecture
abstract
The separation of the data capture and analysis in modern vision systems has led to a massive amount of data transfer between the end devices and cloud computers, resulting in long latency, slow response, and high power consumption. Efficient hardware architectures are under focused development to enable Artificial Intelligence (AI) at the resource-limited sensing devices. One of the most promising solutions is to enable Processing-in-Pixel (PIP) scheme. However, the conventional schemes suffer from the low fill-factor issue. This paper proposes a PIP based Complementary Metal-Oxide-Semiconductor (CMOS) sensor architecture, which allows convolution operation before the column readout circuit to significantly reduce the overall power consumption while improving the resource utilization of the succeeding deep learning accelerator. The simulation results show that the proposed architecture could support the computing efficiency up to 3.37 TOPS/W at the 8-bit weight configuration, which is four times as high as the conventional schemes after normalization. The transistors required for each pixel are only 3.5T, significantly improving the fill-factor.
Ruibing Song, Kejie Huang, Zongsheng Wang, Haibin Shen
IEEE Trans. Circuits Syst. Video Technol.4
2022 An 8-Bit in Resistive Memory Computing Core With Regulated Passive Neuron and Bitline Weight Mapping
abstract
The rapid development of artificial intelligence (AI) and Internet of Things (IoT) increase the requirement for edge computing with low power and relatively high processing speed devices. The computing-in-memory (CIM) schemes based on emerging resistive nonvolatile memory (NVM) show great potential in reducing the power consumption for AI computing. However, the inconsistency of the NVM may significantly degenerate the performance of the neural network. In this article, we propose a low power resistive RAM (RRAM)-based CIM core to not only achieve high computing efficiency but also greatly enhance the robustness by bit line (BL) regulator and BL weight mapping algorithm. The simulation results show that the power consumption of our proposed 8-bit CIM core is only 12.6 mW ($256\times 256$at 8b). The spurious-free dynamic range (SFDR) and signal to noise and distortion ratio (SNDR) of the CIM core achieve 62.64 and 45.92 dB, respectively. The proposed BL weight mapping scheme improves the top-1 accuracy by 2.46% and 3.47% for AlexNet and VGG16 on ImageNet Large Scale Visual Recognition Competition 2012 (ILSVRC 2012) in 8-bit mode, respectively.
Yewei Zhang, Kejie Huang, Rui Xiao 0003, Bo Wang 0020, Yanfeng Xu, Jicong Fan 0002, Haibin Shen
IEEE Trans. Very Large Scale Integr. Syst.7
2021 C2F-FWN: Coarse-to-Fine Flow Warping Network for Spatial-Temporal Consistent Motion Transfer
abstract
Human video motion transfer (HVMT) aims to synthesize videos that one person imitates other persons' actions. Although existing GAN-based HVMT methods have achieved great success, they either fail to preserve appearance details due to the loss of spatial consistency between synthesized and exemplary images, or generate incoherent video results due to the lack of temporal consistency among video frames. In this paper, we propose Coarse-to-Fine Flow Warping Network (C2F-FWN) for spatial-temporal consistent HVMT. Particularly, C2F-FWN utilizes coarse-to-fine flow warping and Layout-Constrained Deformable Convolution (LC-DConv) to improve spatial consistency, and employs Flow Temporal Consistency (FTC) Loss to enhance temporal consistency. In addition, provided with multi-source appearance inputs, C2F-FWN can support appearance attribute editing with great flexibility and efficiency. Besides public datasets, we also collected a large-scale HVMT dataset named SoloDance for evaluation. Extensive experiments conducted on our SoloDance dataset and the iPER dataset show that our approach outperforms state-of-art HVMT methods in terms of both spatial and temporal consistency. Source code and the SoloDance dataset are available at https://github.com/wswdx/C2F-FWN.
Dongxu Wei, Xiaowei Xu 0004, Haibin Shen, Kejie Huang
AAAI3
2021 DualPathGAN: Facial reenacted emotion synthesis
abstract
Abstract Facial reenactment has developed rapidly in recent years, but few methods have been built upon reenacted face in videos. Facial‐reenacted emotion synthesis can make the process of facial reenactment more practical. A facial‐reenacted emotion synthesis method is proposed that includes a dual‐path generative adversarial network (GAN) for emotion synthesis and a residual‐mask network to impose structural restrictions to preserve the mouth shape of the source person. To train the dual‐path GAN more effectively, a learning strategy based on separated discriminators is proposed. The method is trained and tested on a very challenging imbalanced dataset to evaluate the ability to deal with complex practical scenarios. Compared with general emotion synthesis methods, the proposed method can generate more realistic facial emotion synthesised images or videos with higher quality while retaining the expression contents of the original videos. The DualPathGAN achieves a Fréchet inception distance (FID) score of 9.20, which is lower than the FID score of 11.37 achieved with state‐of‐the‐art methods.
Jiahui Kong, Haibin Shen, Kejie Huang
IET Comput. Vis.2
2021 Foreground-Background Parallel Compression With Residual Encoding for Surveillance Video
abstract
The data storage has been one of the bottlenecks in surveillance systems. The conventional video compression schemes such as H.264 and H.265 do not fully utilize the low information density characteristic of the surveillance video, and they attach equal importance to foreground and background when performing compression. In this article, we propose a novel video compression scheme that compresses the foreground and background of the surveillance video separately. The compression ratio is greatly improved by sharing background information among adjacent frames through an adaptive background updating and interpolation module. Besides, we present two different schemes to compress the foreground and compare their performance in the ablation study to show the importance of temporal information for video compression. In the decoding end, a coarse-to-fine two-stage module is applied to achieve the composition of the foreground and background and the enhancements of frame quality. The experimental results show that our proposed method requires 49.75% less bpp (bits per pixel) than the conventional algorithm H.265 to achieve the same PSNR (36 dB) on the HEVC dataset.
Lirong Wu, Kejie Huang, Haibin Shen, Lianli Gao
IEEE Trans. Circuits Syst. Video Technol.3
2021 GAC-GAN: A General Method for Appearance-Controllable Human Video Motion Transfer
abstract
Human video motion transfer has a wide range of applications in multimedia, computer vision, and graphics. Recently, due to the rapid development of Generative Adversarial Networks (GANs), there has been significant progress in the field. However, almost all existing GAN-based works are prone to address the mapping from human motions to video scenes, with scene appearances encoded individually in the trained models. Therefore, each trained model can only generate videos with a specific scene appearance, and new models are required to be trained to generate new appearances. Besides, existing works lack the capability of appearance control. For example, users have to provide video records of wearing new clothes or performing in new backgrounds to enable clothes or background changing in their synthetic videos, which greatly limits the application flexibility. In this paper, we propose General Appearance-Controllable GAN (GAC-GAN), a general method for appearance-controllable human video motion transfer. To enable general-purpose appearance synthesis, we propose to include appearance information in the conditioning inputs. Thus, once trained, our model can generate new appearances by altering the input appearance information. To achieve appearance control, we first obtain the appearance-controllable conditioning inputs, and then utilize a two-stage GAC-GAN to generate the corresponding appearance-controllable outputs, where we utilize an Appearance-Consistency GAN (ACGAN) loss, and a shadow extraction module for output foreground, and background appearance control respectively. We further build a solo dance dataset containing a large number of dance videos for training, and evaluation. Experimental results on our solo dance dataset, and iPER dataset show that our proposed GAC-GAN can not only support appearance-controllable human video motion transfer but also achieve higher video quality than state-of-art methods.
Dongxu Wei, Xiaowei Xu 0004, Haibin Shen, Kejie Huang
IEEE Trans. Multim.3
2020 A GAN-based Tunable Image Compression System
abstract
The method of importance map has been widely adopted in DNN-based lossy image compression to achieve bit allocation according to the importance of image contents. However, insufficient allocation of bits in non-important regions often leads to severe distortion at low bpp (bits per pixel), which hampers the development of efficient content-weighted image compression systems. This paper rethinks content-based compression by using Generative Adversarial Network (GAN) to reconstruct the non-important regions. Moreover, multiscale pyramid decomposition is applied to both the encoder and the discriminator to achieve global compression of high-resolution images. A tunable compression scheme is also proposed in this paper to compress an image to any specific compression ratio without retraining the model. The experimental results show that our proposed method improves MS-SSIM by more than 10.3% compared to the recently reported GAN-based method [3] to achieve the same low bpp (0.05) on the Kodak dataset.
Lirong Wu, Kejie Huang, Haibin Shen
WACV3
2020 An Efficient Hardware Accelerator for Structured Sparse Convolutional Neural Networks on FPGAs
abstract
Deep convolutional neural networks (CNNs) have achieved state-of-the-art performance in a wide range of applications. However, deeper CNN models, which are usually computation consuming, are widely required for complex artificial intelligence (AI) tasks. Though recent research progress on network compression, such as pruning, has emerged as a promising direction to mitigate computational burden, existing accelerators are still prevented from completely utilizing the benefits of leveraging sparsity due to the irregularity caused by pruning. On the other hand, field-programmable gate arrays (FPGAs) have been regarded as a promising hardware platform for CNN inference acceleration. However, most existing FPGA accelerators focus on dense CNN and cannot address the irregularity problem. In this article, we propose a sparsewise dataflow to skip the cycles of processing multiply-and-accumulates (MACs) with zero weights and exploit data statistics to minimize energy through zeros gating to avoid unnecessary computations. The proposed sparsewise dataflow leads to a low bandwidth requirement and high data sharing. Then, we design an FPGA accelerator containing a vector generator module (VGM) that can match the index between sparse weights and input activations according to the proposed dataflow. Experimental results demonstrate that our implementation can achieve 987-, 46-, and 57-imag/s performance for AlexNet, VGG-16, and ResNet-50 on Xilinx ZCU102, respectively, which provides 1.5×-6.7× speedup and 2.0×-6.0× energy efficiency over previous CNN FPGA accelerators.
Kejie Huang, Shuyuan Yang 0003, Haibin Shen
IEEE Trans. Very Large Scale Integr. Syst.6
2018 Accurate iris center localization method using facial landmark, snakuscule, circle fitting and binary connected component
Kejie Huang, Yue Qiu 0003, Haibin Shen
Multim. Tools Appl.4
2018 Composable Worst-Case Delay Bound Analysis Using Network Calculus
abstract
Performance analysis is playing an indispensable role in design and evaluation for on-chip networks. In former studies, the end-to-end delay bound is calculated by the equivalent service curve method based on network calculus when resource sharing happens. However, in this paper, we propose a composable method to get the bound. This method uses the aggregated local arrival curve to get the local delay bound first, then calculates the end-to-end bound by summing up local bounds. This method solves the scalability problem and largely decreases the computation complexity compared with the former method.
Yanchen Long, Zhonghai Lu, Haibin Shen
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.3
2014 Preservation of local linearity by neighborhood subspace scaling for solving the pre-image problem
abstract
An important issue involved in kernel methods is the pre-image problem. However, it is an ill-posed problem, as the solution is usually nonexistent or not unique. In contrast to direct methods aimed at minimizing the distance in feature space, indirect methods aimed at constructing approximate equivalent models have shown outstanding performance. In this paper, an indirect method for solving the pre-image problem is proposed. In the proposed algorithm, an inverse mapping process is constructed based on a novel framework that preserves local linearity. In this framework, a local nonlinear transformation is implicitly conducted by neighborhood subspace scaling transformation to preserve the local linearity between feature space and input space. By extending the inverse mapping process to test samples, we can obtain pre-images in input space. The proposed method is non-iterative, and can be used for any kernel functions. Experimental results based on image denoising using kernel principal component analysis (PCA) show that the proposed method outperforms the state-of-the-art methods for solving the pre-image problem.
Sheng-Kai Yang, Jian-Yi Meng, Haibin Shen
J. Zhejiang Univ. Sci. C3
2012 Largemargin classification for combating disguise attacks on spam filters
abstract
This paper addresses the challenge of large margin classification for spam filtering in the presence of an adversary who disguises the spam mails to avoid being detected. In practice, the adversary may strategically add good words indicative of a legitimate message or remove bad words indicative of spam. We assume that the adversary could afford to modify a spam message only to a certain extent, without damaging its utility for the spammer. Under this assumption, we present a large margin approach for classification of spam messages that may be disguised. The proposed classifier is formulated as a second-order cone programming optimization. We performed a group of experiments using the TREC 2006 Spam Corpus. Results showed that the performance of the standard support vector machine (SVM) degrades rapidly when more words are injected or removed by the adversary, while the proposed approach is more stable under the disguise attack.
Xichuan Zhou, Haibin Shen
J. Zhejiang Univ. Sci. C2
2011 Integrating outlier filtering in large margin training
abstract
Large margin classifiers such as support vector machines (SVM) have been applied successfully in various classification tasks. However, their performance may be significantly degraded in the presence of outliers. In this paper, we propose a robust SVM formulation which is shown to be less sensitive to outliers. The key idea is to employ an adaptively weighted hinge loss that explicitly incorporates outlier filtering in the SVM training, thus performing outlier filtering and classification simultaneously. The resulting robust SVM formulation is non-convex. We first relax it into a semi-definite programming which admits a global solution. To improve the efficiency, an iterative approach is developed. We have performed experiments using both synthetic and real-world data. Results show that the performance of the standard SVM degrades rapidly when more outliers are included, while the proposed robust SVM training is more stable in the presence of outliers.
Xichuan Zhou, Haibin Shen, Jieping Ye
J. Zhejiang Univ. Sci. C2
2010 A parallel and scalable digital architecture for training support vector machines
abstract
To facilitate the application of support vector machines (SVMs) in embedded systems, we propose and test a parallel and scalable digital architecture based on the sequential minimal optimization (SMO) algorithm for training SVMs. By taking advantage of the mature and popular SMO algorithm, the numerical instability issues that may exist in traditional numerical algorithms are avoided. The error cache updating task, which dominates the computation time of the algorithm, is mapped into multiple processing units working in parallel. Experiment results show that using the proposed architecture, SVM training problems can be solved effectively with inexpensive fixed-point arithmetic and good scalability can be achieved. This architecture overcomes the drawbacks of the previously proposed SVM hardware that lacks the necessary flexibility for embedded applications, and thus is more suitable for embedded use, where scalability is an important concern.
Kuikang Cao, Haibin Shen
J. Zhejiang Univ. Sci. C2
2010 Notifiable infectious disease surveillance with data collected by search engine
abstract
Notifiable infectious diseases are a major public health concern in China, causing about five million illnesses and twelve thousand deaths every year. Early detection of disease activity, when followed by a rapid response, can reduce both social and medical impact of the disease. We aim to improve early detection by monitoring health-seeking behavior and disease-related news over the Internet. Specifically, we counted unique search queries submitted to the Baidu search engine in 2008 that contained disease-related search terms. Meanwhile we counted the news articles aggregated by Baidu’s robot programs that contained disease-related keywords. We found that the search frequency data and the news count data both have distinct temporal association with disease activity. We adopted a linear model and used searches and news with 1–200-day lead time as explanatory variables to predict the number of infections and deaths attributable to four notifiable infectious diseases, i.e., scarlet fever, dysentery, AIDS, and tuberculosis. With the search frequency data and news count data, our approach can quantitatively estimate up-to-date epidemic trends 10–40 days ahead of the release of Chinese Centers for Disease Control and Prevention (Chinese CDC) reports. This approach may provide an additional tool for notifiable infectious disease surveillance.
Xichuan Zhou, Haibin Shen
J. Zhejiang Univ. Sci. C2
2008 Low complexity bit parallel multiplier for GF(2m) generated by equally-spaced trinomials
Haibin Shen, Yier Jin
Inf. Process. Lett.1
2006 Interconnect Estimation for Mesh-Based Reconfigurable Computing
Haibin Shen, Rongquan You, Yier Jin, Aiming Ji
EUC1
2006 Securing C Programs by Dynamic Type Checking
Haibin Shen, Jimin Wang, Lingdi Ping
ISPEC1