Yuan Zheng 0002

dblp:29/2670-2 · DBLP profile ↗
← Back
9ranked-venue papers
1as first author
5since 2021 · last 2024
0000-0002-7632-6846ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorArtificial intelligence and machine learning · 1
YearPublicationVenuePosition
2024 CAG-FPN: Channel Self-Attention Guided Feature Pyramid Network for Object Detection
abstract
Feature Pyramid Network (FPN) plays a critical role and is indispensable for object detection methods. In recent years, attention mechanism has been utilized to improve FPN due to its excellent performance. Existing attention-based FPN methods generally work with a complex structure, resulting in an increase of computational costs. In view of this, we propose a novel Channel Self-Attention Guided Feature Pyramid Network (CAG-FPN), which not only has a simple structure but also consistently improves detection accuracy. We observe that introducing channel self-attention to the features at the highest level is helpful for object detection, since modeling long-range dependencies between channels triggers an implicit clustering of the same categories of objects, enhancing the semantic continuity. Moreover, our CAG-FPN can be readily plugged into both one-stage and two-stage FPN-based detectors. Experiments on MS COCO dataset verify the superiority and generalization ability of our CAG-FPN. Code is available at https://github.com/ZY-IMU-CV/CAGFPN_CJ_2023.
Huhe Dai, Yuan Zheng 0002
ICASSP3
2024 CENet: Content-Aware Enhanced Network for Practical Scene Parsing
abstract
Attention mechanisms are widely adopted in existing scene parsing methods due to their excellent performance, especially spatial self-attention. However, spatial self-attention suffers from high computational complexity, which limits the practical applications of the scene parsing methods on mobile devices with limited resources. In view of this, we propose a simple yet effective spatial attention module, namely Content-Aware Attention Module (CA2M). CA2M is a lightweight spatial attention module that consists of several convolution and pooling operations, compared to various spatial self-attention modules. Moreover, it is able to adaptively select spatial pixel information which is helpful for scene parsing task. With CA2M, we present a Content-aware Enhanced Network for scene parsing (CENet), where CA2M is introduced into the lateral connections at four different scales, resulting in a semantic alignment at adjacent scales and an effective semantic propagation. To validate the performance of the proposed CA2M and CENet, we conduct extensive experiments and achieve consistently improved performances on three popular benchmarks. Furthermore, we verify their generalization ability when using different baseline models and backbone networks. Code is available at https://github.com/ZY-IMU-CV/CENET_SK_2023.
Zhengtan Wang, Huhe Dai, Yuan Zheng 0002
ICASSP4
2024 RAFNet: Reparameterizable Across-Resolution Fusion Network for Real-Time Image Semantic Segmentation
abstract
The demand to implement semantic segmentation networks on mobile devices has increased dramatically. However, existing real-time semantic segmentation methods still suffer from a large number of network parameters, unsuitable for mobile devices with limited memory resources. The reason mainly arises from the fact that most existing methods take the backbone networks (e.g., ResNet-18 and MobileNet) as an encoder. To alleviate this problem, we propose a novel Reparameterizable Channel & Dilation (RCD) block and construct a considerably lightweight yet effective encoder by stacking several RCD blocks according to three guidelines. The strengths of the proposed encoder result in the abilities not only to extract discriminative feature representations via channel convolutions and dilated convolutions, but also to reduce computational burdens while maintaining segmentation accuracy with the help of re-parameterization technique. Except for encoder, we also present a simple but effective decoder that adopts an across-resolution fusion strategy to fuse multi-scale feature maps generated from the encoder instead of a bottom-up pathway fusion. With such an encoder and a decoder, we provide a Reparameterizable Across-resolution Fusion Network (RAFNet) for real-time semantic segmentation. Extensive experiments demonstrate that our RAFNet achieves a promising trade-off between segmentation accuracy, inference speed and network parameters. Specifically, our RAFNet with only 0.96M parameters obtains 75.3% mIoU at 107 FPS and 75.8% mIoU at 195 FPS on Cityscapes and CamVid test sets for full-resolution inputs, respectively. After quantization and deployment on a Xilinx ZCU104 device, our RAFNet obtains a favorable segmentation performance with only 1.4W power.
Huhe Dai, Yuan Zheng 0002
IEEE Trans. Circuits Syst. Video Technol.3
2023 ICANet: A Lightweight Increasing Context Aided Network for Real-Time Image Semantic Segmentation
abstract
Although existing real-time segmentation methods have reduced the network parameters, they still suffer from slow inference speed. To alleviate this problem, we propose a novel lightweight image segmentation method, called Increasing Context Aided Network (ICANet), which brings a significant increase in inference speed while maintaining competitive segmentation accuracy. To do this, we elaborately design both encoder and decoder. Considering that a large reduction in channel numbers brings poor performance, we propose a Inverted Depthwise Separable convolution block (IDS block) that enriches semantic information in channel and spatial dimensions. By stacking several such IDS blocks, we build an efficient encoder that can capture an increasing context by using different dilation rates in IDS blocks to further improve accuracy. Besides, we present a simple decoder that adopts a bottom-up pathway to effectively fuse multiscale feature maps from encoder. Extensive experiments show our ICANet achieves a SOTA balance between accuracy, speed and parameters.
Huhe Dai, Yuan Zheng 0002
ICME3
2023 Pose-Aided Video-Based Person Re-Identification via Recurrent Graph Convolutional Network
abstract
Existing methods for video-based person re- identification (ReID) mainly learn the appearance feature of a given pedestrian via a feature extractor and a feature aggregator. However, the appearance models would fail to learn a large inter-class variance when different pedestrians have similar appearances. Considering that different pedestrians have different walking postures and body proportions, we propose to learn the discriminative pose feature beyond the appearance feature for video retrieval. Specifically, we implement a two-branch architecture to separately learn the appearance feature and pose feature, and then concatenate them together for inference. To learn the pose feature, we first detect the pedestrian pose in each frame through an off-the-shelf pose detector, and construct a temporal graph using the pose sequence. We then exploit a recurrent graph convolutional network (RGCN) to learn the node embeddings of the temporal pose graph, which devises a global information propagation mechanism to simultaneously achieve the neighborhood aggregation of intra-frame nodes and message passing among inter-frame graphs. Finally, we propose a dual-attention method (DAM) consisting of node-attention and time-attention to obtain the temporal graph representation from the node embeddings, where the self-attention mechanism is employed to learn the importance of each node and each frame. We verify the proposed method on three video-based ReID datasets, i.e., Mars, DukeMTMC and iLIDS-VID, whose experimental results demonstrate that the learned pose feature can effectively improve the performance of existing appearance models.
Honghu Pan, Qiao Liu 0001, Yongyong Chen, Yunqi He, Yuan Zheng 0002, Feng Zheng 0001, Zhenyu He 0001
IEEE Trans. Circuits Syst. Video Technol.5
2020 PTB-TIR: A Thermal Infrared Pedestrian Tracking Benchmark
abstract
Thermal infrared (TIR) pedestrian tracking is one of the important components among numerous applications of computer vision, which has a major advantage: it can track pedestrians in total darkness. The ability to evaluate the TIR pedestrian tracker fairly, on a benchmark dataset, is significant for the development of this field. However, there is not a benchmark dataset. In this paper, we develop a TIR pedestrian tracking dataset for the TIR pedestrian tracker evaluation. The dataset includes 60 thermal sequences with manual annotations. Each sequence has nine attribute labels for the attribute based evaluation. In addition to the dataset, we carry out the large-scale evaluation experiments on our benchmark dataset using nine publicly available trackers. The experimental results help us understand the strengths and weaknesses of these trackers. In addition, in order to gain more insight into the TIR pedestrian tracker, we divide its functions into three components: feature extractor, motion model, and observation model. Then, we conduct three comparison experiments on our benchmark dataset to validate how each component affects the tracker's performance. The findings of these experiments provide some guidelines for future research.
Qiao Liu 0001, Zhenyu He 0001, Xin Li 0034, Yuan Zheng 0002
IEEE Trans. Multim.4
2016 An accurate and practical calibration method for roadside camera using two vanishing points
Xinhua You, Yuan Zheng 0002
Neurocomputing2
2014 A Practical Roadside Camera Calibration Method Based on Least Squares Optimization
abstract
In this paper, we propose a more practical and accurate method for calibrating the roadside camera used in traffic surveillance systems. Considering the characteristics of the traffic scenes, we propose a minimum calibration condition that consists of two vanishing points and a vanishing line, which can be easily satisfied in most traffic scenes. Based on the minimum calibration condition, we provide a calibration method to estimate camera intrinsic parameters and rotation angles, which employs least squares optimization instead of closed-form computation. Compared with the existing calibration methods, our method is suitable for more traffic scenes and is able to accurately determine more camera parameters including the principal point. By making full use of video information, multiple observations of the vanishing points are available from different objects. For more accurate calibration, we present a dynamic calibration method using these observations to correct camera parameters. As for the estimation of the camera translation vector, known lengths in the road or known heights above the road are exploited. The experimental results on synthetic data and real traffic images demonstrate the accuracy, robustness, and practicability of the proposed calibration method.
Yuan Zheng 0002, Silong Peng
IEEE Trans. Intell. Transp. Syst.1
2013 Lighting Estimation of a Convex Lambertian Object Using Redundant Spherical Harmonic Frames
Wen-Yong Zhao, Shaolin Chen, Yuan Zheng 0002, Silong Peng
J. Comput. Sci. Technol.3