Peining Zhen

dblp:252/0735 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
9since 2021 · last 2025
0000-0002-8439-876XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 4 since 2021Computer networks · 2 · 2 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2025 MemATr: An Efficient and Lightweight Memory-augmented Transformer for Video Anomaly Detection
abstract
Anomaly detection in videos is a long-standing and challenging problem. Previous methods often adopt deep and large neural networks to achieve the best detection accuracy; however, the high computational costs prevent them from being used in real-world applications with constrained computational resources. In this paper, we develop a mem ory- a ugmented tr ansformer named MemATr, which is capable of detecting video anomalies effectively. The proposed network is lightweight and can be easily deployed on mobile devices. Furthermore, we propose a memory transformer module to make predictions that are closer to normal inputs, thereby leading to a higher error for abnormal input patterns. Memory-attention is the main component of the proposed memory transformer, which can retrieve the features from learnable values rather than from the backbone like previous methods. Extensive experiments on the UCSD Ped2, CUHK Avenue, and ShanghaiTech benchmarks can demonstrate that our model has a significantly smaller model size while still achieving competitive detection accuracy. Our model has only 1/12 the number of parameters of the baseline model. Besides, our model achieves a 4.6% increase in accuracy on the ShanghaiTech dataset and has roughly the same accuracy compared with the baseline on the other two datasets. We validate the performance of the proposed model on the mobile device and the result shows it only has 49.8ms latency. The effectiveness of the proposed method on mobile devices is further supported by experimental results. A new quantitative parameter AMD (Applicability for Mobile Devices) is proposed to offer a novel approach to assist in making trade-offs for mobile devices. The proposed model obtains state-of-the-art results in terms of AMD.
Jingjing Chang, Peining Zhen, Xiaotao Yan, Yixin Yang 0004, Ziyang Gao, Haibao Chen
ACM Trans. Embed. Comput. Syst.2
2023 Multilayer Perceptron-Based Stress Evolution Analysis Under DC Current Stressing for Multisegment Wires
abstract
Electromigration (EM) is one of the major concerns in the reliability analysis of very large-scale integration (VLSI) systems due to the continuous technology scaling. Accurately predicting the time-to-failure of integrated circuits (ICs) becomes increasingly important for modern IC design. However, traditional methods are often not sufficiently accurate, leading to undesirable over-design especially in advanced technology nodes. In this article, we propose an approach using multilayer perceptrons (MLPs) to compute stress evolution in the interconnect trees during the void nucleation phase. The availability of a customized trial function for neural network training holds the promise of finding dynamic mesh-free stress evolution on complex interconnect trees under time-varying temperatures. Specifically, we formulate a new objective function considering the EM-induced coupled partial differential equations (PDEs), boundary conditions (BCs), and initial conditions to enforce the physics-based constraints in the spatial–temporal domain. The proposed model avoids meshing and reduces temporal iterations compared with conventional numerical approaches like finite element method. Numerical results confirm its advantages on accuracy and computational performance.
Tianshu Hou, Peining Zhen, Ngai Wong 0001, Quan Chen 0007, Guoyong Shi, Haibao Chen
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2023 A Highly Compressed Accelerator With Temporal Optical Flow Feature Fusion and Tensorized LSTM for Video Action Recognition on Terminal Device
abstract
Deep learning-based action recognition has become ubiquitous in the video analysis area; however, large neural networks require enormous computations to achieve high performance, which hinder them from mobile applications that are tightly constrained by hardware resources. In this work, we introduce a highly compact and fast neural network-based action recognition accelerator named ARA on the terminal device. We build an LSTM-based spatio-temporal action recognition model with extracted time-series features from RGB frames and flow features from optical flow fields. Then the LSTM-based spatio-temporal model is deeply compressed with tensor decomposition to further reduce redundant parameters and lessen computation overhead. Based on the datasets UCF-11, UCF-101, and HMDB51, our proposed method achieves 95.87%, 94.08%, and 75.71% classification accuracy, being comparable with other state-of-the-art methods. In particular, our proposed method significantly compresses the parameter of the LSTM model$215\times $on the UCF-101 dataset. The proposed system can also achieve a fast running speed of 157.7 FPS on GPU. Furthermore, we validate the performance of the proposed system on an ARM-based terminal device; the results show it only has 0.017-s latency and 4.73-W power consumption.
Peining Zhen, Xiaotao Yan, Haibao Chen
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2023 Toward Compact Transformers for End-to-End Object Detection With Decomposed Chain Tensor Structure
abstract
DEtection TRansformer (DETR) is a recently proposed method that streamlines the detection pipeline and achieves competitive results against two-stage detectors such as Faster-RCNN. The DETR models get rid of complex anchor generation and post-processing procedures thereby making the detection pipeline more intuitive. However, the numerous redundant parameters in transformers make the computation and storage of the DETR models intensive, which seriously hinder them to be deployed on the resources-constrained devices. In this paper, to obtain a compact end-to-end detection framework, we propose to deeply compress the transformers with low-rank tensor decomposition. The basic idea of our tensor-based compression method is to represent the large-scale weight matrix in one network layer with a chain of low-order matrices. Furthermore, we show that redundant attention heads will hinder the performance of detection transformers. We thus propose a gated multi-head attention (GMHA) module to suppress the redundant attention information by normalizing the attention heads. In GMHA, each attention head has an independent gate to determine the passed attention value, thereby down-weighting the uninformative heads. The accuracy drop of the tensor-compressed DETR models can be mitigated by applying GMHA modules. Lastly, to obtain fully compressed DETR models, a low-bitwidth quantization technique is introduced for further reducing the model storage size. Based on the proposed methods, we can achieve significant parameter and model size reduction while maintaining high detection performance. We conduct extensive experiments on the COCO and PASCAL VOC datasets to validate the effectiveness of our tensor-compressed (tensorized) DETR models. The experimental results on the COCO benchmark show that we can attain$3.7\times $full model compression with$482\times $feed forward network (FFN) parameter reduction and only 0.6 points accuracy drop.
Peining Zhen, Xiaotao Yan, Tianshu Hou, Haibao Chen
IEEE Trans. Circuits Syst. Video Technol.1
2023 A Deep Learning Framework for Solving Stress-based Partial Differential Equations in Electromigration Analysis
abstract
The electromigration-induced reliability issues (EM) in very large scale integration (VLSI) circuits have attracted continuous attention due to technology scaling. Traditional EM methods lead to inaccurate results incompatible with the advanced technology nodes. In this article, we propose a learning-based model by enforcing physical constraints of EM kinetics to solve the EM reliability problem. The method aims at solving stress-based partial differential equations (PDEs) to obtain the hydrostatic stress evolution on interconnect trees during the void nucleation phase, considering varying atom diffusivity on each segment, which is one of the EM random characteristics. The approach proposes a crafted neural network-based framework customized for the EM phenomenon and provides mesh-free solutions benefiting from the employment of automatic differentiation (AD). Experimental results obtained by the proposed model are compared with solutions obtained by competing methods, showing satisfactory accuracy and computational savings.
Tianshu Hou, Peining Zhen, Zhigang Ji, Haibao Chen
ACM Trans. Design Autom. Electr. Syst.2
2023 Towards Accurate Oriented Object Detection in Aerial Images with Adaptive Multi-level Feature Fusion
abstract
Detecting objects in aerial images is a long-standing and challenging problem since the objects in aerial images vary dramatically in size and orientation. Most existing neural network based methods are not robust enough to provide accurate oriented object detection results in aerial images since they do not consider the correlations between different levels and scales of features. In this paper, we propose a novel two-stage network-based detector with a daptive f eature f usion towards highly accurate oriented object det ection in aerial images, named AFF-Det . First, a multi-scale feature fusion module (MSFF) is built on the top layer of the extracted feature pyramids to mitigate the semantic information loss in the small-scale features. We also propose a cascaded oriented bounding box regression method to transform the horizontal proposals into oriented ones. Then the transformed proposals are assigned to all feature pyramid network (FPN) levels and aggregated by the weighted RoI feature aggregation (WRFA) module. The above modules can adaptively enhance the feature representations in different stages of the network based on the attention mechanism. Finally, a rotated decoupled-RCNN head is introduced to obtain the classification and localization results. Extensive experiments are conducted on the DOTA and HRSC2016 datasets to demonstrate the advantages of our proposed AFF-Det. The best detection results can achieve 80.73% mAP and 90.48% mAP, respectively, on these two datasets, outperforming recent state-of-the-art methods.
Peining Zhen, Suming Zhang, Xiaotao Yan, Zhigang Ji, Haibao Chen
ACM Trans. Multim. Comput. Commun. Appl.1
2022 Deeply Tensor Compressed Transformers for End-to-End Object Detection
abstract
DEtection TRansformer (DETR) is a recently proposed method that streamlines the detection pipeline and achieves competitive results against two-stage detectors such as Faster-RCNN. The DETR models get rid of complex anchor generation and post-processing procedures thereby making the detection pipeline more intuitive. However, the numerous redundant parameters in transformers make the DETR models computation and storage intensive, which seriously hinder them to be deployed on the resources-constrained devices. In this paper, to obtain a compact end-to-end detection framework, we propose to deeply compress the transformers with low-rank tensor decomposition. The basic idea of the tensor-based compression is to represent the large-scale weight matrix in one network layer with a chain of low-order matrices. Furthermore, we propose a gated multi-head attention (GMHA) module to mitigate the accuracy drop of the tensor-compressed DETR models. In GMHA, each attention head has an independent gate to determine the passed attention value. The redundant attention information can be suppressed by adopting the normalized gates. Lastly, to obtain fully compressed DETR models, a low-bitwidth quantization technique is introduced for further reducing the model storage size. Based on the proposed methods, we can achieve significant parameter and model size reduction while maintaining high detection performance. We conduct extensive experiments on the COCO dataset to validate the effectiveness of our tensor-compressed (tensorized) DETR models. The experimental results show that we can attain 3.7 times full model compression with 482 times feed forward network (FFN) parameter reduction and only 0.6 points accuracy drop.
Peining Zhen, Ziyang Gao, Tianshu Hou, Haibao Chen
AAAI1
2022 FASSST: Fast Attention Based Single-Stage Segmentation Net for Real-Time Instance Segmentation
abstract
Real-time instance segmentation is crucial in various AI applications. This work designs a network named Fast Attention based Single-Stage Segmentation NeT (FASSST) that performs instance segmentation with video-grade speed. Using an instance attention module (IAM), FASSST quickly locates target instances and segments with region of interest (ROI) feature fusion (RFF) aggregating ROI features from pyramid mask layers. The module employs an efficient single-stage feature regression, straight from features to instance coordinates and class probabilities. Experiments on COCO and CityScapes datasets show that FASSST achieves state-of-the-art performance under competitive accuracy: real-time inference of 47.5FPS on a GTX1080Ti GPU and 5.3FPS on a Jetson Xavier NX board with only 71.6 GFLOPs.
Peining Zhen, Tianshu Hou, Chiu Wa Ng, Haibao Chen, Hao Yu 0001, Ngai Wong 0001
WACV3
2021 Fast Video Facial Expression Recognition by a Deeply Tensor-Compressed LSTM Neural Network for Mobile Devices
abstract
Mobile devices usually suffer from limited computation and storage resources, which seriously hinders them from deep neural network applications. In this article, we introduce a deeply tensor-compressed long short-term memory (LSTM) neural network for fast video-based facial expression recognition on mobile devices. First, a spatio-temporal facial expression recognition LSTM model is built by extracting time-series feature maps from facial clips. The LSTM-based spatio-temporal model is further deeply compressed by means of quantization and tensorization for mobile device implementation. Based on datasets of Extended Cohn-Kanade (CK+), MMI, and Acted Facial Expression in Wild 7.0, experimental results show that the proposed method achieves 97.96%, 97.33%, and 55.60% classification accuracy and significantly compresses the size of network model up to 221× with reduced training time per epoch by 60%. Our work is further implemented on the RK3399Pro mobile device with a Neural Process Engine. The latency of the feature extractor and LSTM predictor can be reduced 30.20× and 6.62× , respectively, on board with the leveraged compression methods. Furthermore, the spatio-temporal model costs only 57.19 MB of DRAM and 5.67W of power when running on the board.
Peining Zhen, Haibao Chen, Zhigang Ji, Hao Yu 0001
ACM Trans. Internet Things1
2020 An Anomaly Comprehension Neural Network for Surveillance Videos on Terminal Devices
abstract
Anomaly comprehension in surveillance videos is more challenging than detection. This work introduces the design of a lightweight and fast anomaly comprehension neural network. For comprehension, a spatio-temporal LSTM model is developed based on the structured, tensorized time-series features extracted from surveillance videos. Deep compression of network size is achieved by tensorization and quantization for the implementation on terminal devices. Experiments on large-scale video anomaly dataset UCF-Crime demonstrate that the proposed network can achieve an impressive inference speed of 266 FPS on a GTX-1080Ti GPU, which is 4.29 faster than ConvLSTM-based method; a 3.34% AUC improvement with 5.55% accuracy niche versus the 3D-CNN based approach; and at least 15k× parameter reduction and 228× storage compression over the RNN-based approaches. Moreover, the proposed framework has been realized on an ARM-core based IOT board with only 2.4W power consumption.
Guangtai Huang, Peining Zhen, Haibao Chen, Ngai Wong 0001, Hao Yu 0001
DATE3