Zhen Zuo

dblp:63/6940 · DBLP profile ↗
← Back
24ranked-venue papers
5as first author
8since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 1 since 2021Artificial intelligence and machine learning · 10 · 3 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Databases, data management, data science and information retrieval · 2Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 first-author
YearPublicationVenuePosition
2025 GSFNet: Gyro-Aided Spatial-Frequency Network for Motion Deblurring of UAV Infrared Images
abstract
Unmanned aerial vehicles (UAVs) with thermal imaging cameras are widely used for target tracking, reconnaissance, and search operations. However, rapid thermal camera rotations during field-of-view adjustments introduce significant motion blur, impairing real-time image detection and tracking. While deep learning has been a dominant approach for image deblurring, its application to infrared image motion deblurring (IRMD) remains limited owing to the lack of publicly available datasets and challenges in handling large motion blur or maintaining real-time performance. This study addresses these gaps by constructing a large-scale UAV infrared motion deblurring (U2IRD) benchmark dataset, incorporating gyroscopic steering rate information. Additionally, we propose a gyro-aided spatial frequency network (GSFNet) that uses spatial and frequency domain features for UAV IRMD. The input data converts the gimbal steering rate information into a pixel distribution intensity map as a priori information. Specifically, the designed spatial depth residual attention module captures critical spatial domain details, while the multiple frequency domain feature recovery module extracts frequency domain features for effective deblurring. Extensive evaluations on U2IRD and synthetic thermal blurred image datasets demonstrate that the proposed method achieves state-of-the-art deblurring performance. The new IRMD dataset, available at https://github.com/aurora-sea/U2IRD, is anticipated to facilitate advancements in UAV IRMD research and applications.
Xiaozhong Tong, Zhen Zuo, Shaojing Su, Peng Wu 0025, Junyu Wei, Runze Guo
IEEE Trans. Geosci. Remote. Sens.3
2024 ST-Trans: Spatial-Temporal Transformer for Infrared Small Target Detection in Sequential Images
abstract
The detection of small infrared targets with a low signal-to-noise ratio and low contrast in high-noise backgrounds is challenging due to the lack of spatial features of the targets and the scarcity of real-world datasets. Most existing methods are based on single-frame images, which are prone to numerous false alarms and missed detections. This paper proposes ST-Trans that provides an efficient end-to-end solution for the detection of small infrared targets in the complex context of sequential images. First, the detection of small infrared targets in complex backgrounds relying only on a single image has been significantly difficult due to the lack of available spatial features. The temporal and motion information of the sequence image was found to effectively improve target detection performance. Therefore, we used the C2FDark backbone to learn the spatial features associated with small targets, and the spatial-temporal transformer module to learn the spatiotemporal dependencies between successive frames of small infrared targets. This improved the detection performance in challenging scenes. Second, due to the lack of publicly available infrared small target sequence datasets for training, we annotated a set of small infrared targets for challenging scenes and published them as the sequential infrared small target detection (SIRSTD) dataset. Finally, we performed extensive ablation experiments on the SIRSTD dataset and compared its performance with that of state-of-the-art methods to demonstrate the superiority of the proposed method. The results revealed that ST-Trans outperformed other models and can effectively improve the detection performance for small infrared targets. The SIRSTD dataset is available at https://github.com/aurora-sea/SIRSTD.
Xiaozhong Tong, Zhen Zuo, Shaojing Su, Junyu Wei, Peng Wu 0025, Zongqing Zhao
IEEE Trans. Geosci. Remote. Sens.2
2023 MSAFFNet: A Multiscale Label-Supervised Attention Feature Fusion Network for Infrared Small Target Detection
abstract
The detection of small infrared targets with a low signal-to-noise ratios and contrasts in noisy and cluttered backgrounds is challenging and therefore a domain of active research. Traditional methods result in a large number of false alarms and missed detections. In the case of convolutional neural network-based methods, it may not be possible to identify deep small targets, or the details of the target’s edge contours may not be appropriately considered. Therefore, this paper proposes MSAFFNet to perform infrared small target detection based on an encoder-decoder framework. In the encoder stage, small target features are extracted using a resnet-20 backbone network, and the global contextual features of small targets are extracted using an atrous spatial pyramid pooling module. In the decoding stage, a dual-attention module is used to selectively enhance the spatial details of the target at the shallow level and representative features of the semantic information at the deep level. Multi-scale feature maps are then concatenated to achieve superior feature fusion. Additionally, multi-scale labels are constructed to focus on the details of the target contour and internal features based on edge information and an internal feature aggregation module. Experiments conducted on the NUAA-SIRST, NUDT-SIRST and XDU-SIRST datasets revealed that the proposed approach outperforms the representative methods and achieves an improved detection performance.
Xiaozhong Tong, Shaojing Su, Peng Wu 0025, Runze Guo, Junyu Wei, Zhen Zuo, Bei Sun
IEEE Trans. Geosci. Remote. Sens.6
2023 RISTrack: Robust Infrared Ship Tracking With Modified Appearance Feature Extraction and Matching Strategy
abstract
Infrared (IR) ship tracking is becoming increasingly important in various applications. However, it remains a challenging task as the information that can be obtained from infrared images is limited. Aiming at enhancing IR ship tracking accuracy, we propose an innovative approach by presenting feature integration module (FIM) and backup matching module (BMM). FIM takes appearance feature, complete intersection over union (CIoU), and motion direction metrics into account. Regarding appearance feature extraction, an end-to-end characteristic learning strategy with a cross-guided multi-granularity fusion network is proposed to obtain more integral appearance features and enhance re-identification accuracy, which helps to distinguish individual IR ship targets better. Besides, a backup matching strategy is then used to match the unmatched tracks and detections after cascaded matching. Virtual trajectories are generated for the matched tracks to optimize parameters by parameter optimization module (POM). The accumulation of errors caused by the lack of observations in the Kalman filter is reduced. Thus, the position of IR ships can be estimated more accurately, and more robust IR ship tracking can be achieved. In addition, we present a sequential frame IR ship tracking dataset, providing the first public benchmark for testing IR ship tracking performance. Experimental results indicate that the MOTA, MOTP and IDs of the proposed method are 73.441, 80.826, and 32, respectively, outperforming other state-of-the-art methods. This demonstrates the superior robustness of the proposed method, particularly when the IR ships are occluded or the target texture information is lacking. Our dataset is available at https://github.com/echo-sky/SFIST.
Peng Wu 0025, Shaojing Su, Zhen Zuo, Bei Sun, Junyu Wei, Runze Guo, Xiaozhong Tong, Jiaju Zhang, Honghe Huang
IEEE Trans. Geosci. Remote. Sens.3
2022 Data-driven Kalman Filter with Kernel-based Koopman Operators for Nonlinear Robot Systems
abstract
Designing the Kalman filter for nonlinear robot systems with theoretical guarantees is challenging, especially when the dynamics model is unavailable. This paper proposes a data-driven Kalman filter algorithm using kernel-based Koop-man operators for unknown nonlinear robot systems. First, the Koopman operator using sparse kernel-based extended dynamic decomposition (EDMD) is presented to learn the unknown dynamics with input-output datasets. Unlike classic EDMD, which requires manual selection of kernel functions, our approach automatically constructs kernel functions using an approximate linear dependency analysis method. The resulting Koopman model is a linear dynamic evolution in the kernel space, enabling us to address the nonlinear filtering problem using the standard linear Kalman filter design process. Despite this, our approach generates a nonlinear filtering law thanks to the adopted nonlinear kernel functions. Finally, the effectiveness of the proposed approach is validated by simulated experiments.
Wei Jiang 0006, Zhen Zuo, Meiping Shi, Shaojing Su
IROS3
2022 Context-guided feature enhancement network for automatic check-out
Yihan Sun 0002, Tiejian Luo, Zhen Zuo
Neural Comput. Appl.3
2022 SRCANet: Stacked Residual Coordinate Attention Network for Infrared Ship Detection
abstract
The inability of conventional algorithms to detect infrared (IR) ship targets in complex scenes led to the development of detection methods based on convolutional neural networks (CNNs). In this study, we propose a CNN-based stacked residual coordinate attention network (SRCANet) for detecting IR ship targets. Three-directional stacked interaction modules and a full-scale skip connection feature fusion scheme are introduced. The proposed network maintains and integrates sufficient contextual information of IR ship targets and obtains clear target boundary information. A cascaded residual coordinate attention module (CRCAM) is designed as the basic node in the SRCANet. Additionally, a residual coordinate attention module (RCAM) is introduced, which combines a two-dimensional convolution layer with batch normalisation and rectified linear unit (CBR), a coordination attention module, and a residual connection. The RCAM enhances the input feature map and improves the representability of objects of interest. The CRCAM comprises several cascading RCAMs that deepen the feature extraction layers. Furthermore, because there is no publicly available IR ship target dataset for segmentation, pixel-level annotations are performed on a set of IR ship target images and released as a single-frame IR ship detection (SISD) dataset. Extensive experiments were conducted on the SISD dataset and the widely used single-frame IR small target dataset to demonstrate the superiority of the proposed method. The results indicate that the SRCANet outperforms the state-of-the-art models, and it is more robust when target texture information is lacking. The SISD dataset is available at https://github.com/echo-sky/SISD.
Peng Wu 0025, Honghe Huang, Hanxiang Qian, Shaojing Su, Bei Sun, Zhen Zuo
IEEE Trans. Geosci. Remote. Sens.6
2021 Two-Stream Hybrid Attention Network for Multimodal Classification
abstract
On modern e-commerce platforms like Amazon, the number of products is fast growing, precise and efficient product classification becomes a key lever to great customer shopping experience. To tackle the large-scale product classification problem, a major challenge is how to leverage multimodal product information (e.g., image, text). One of the most successful directions is the attention-based deep multimodal learning, where there are mainly two types of frameworks: 1) keyless attention, which learns the importance of features within each modal; and 2) key-based attention, which learns the importance of features using other modalities. In this paper, we propose a novel Two-stream Hybrid Attention Network (HANet), which leverages both key-based and keyless attention mechanisms to capture the key information across product image and title modalities. We experimentally show that our HANet achieves state-of-the-art performance on Amazon-scale product classification problem.
Qipin Chen, Zhen Zuo, Jinmiao Fu
ICIP3
2019 IP Over SONET/SDH Link-Layer Processing with Multiprocessor
abstract
With the recent increase in IP traffic owing to fiber communication, previous schemes have become inadequate for link-layer processing of IP over SONET/SDH(POS). In this study, a proposal based on [Formula: see text] processors to provide mapping or demapping of IP datagrams from or into SONET/SDH is presented, and the value of [Formula: see text] is decided based on the link layer rate of POS. Further, the mathematic model of proposed architecture are presented in detail. Then the realization procedures are implemented in a Field-Programmable Gate Array (FPGA). Both theoretical analysis and experimental test prove that the proposed scheme is efficient, portable, cost-efficient and has a lower hardware resources consumption.
Zhen Zuo, Jiangyi Qin, Zhiping Huang, Shaojing Su
Int. J. Pattern Recognit. Artif. Intell.1
2018 Scene Segmentation with DAG-Recurrent Neural Networks
abstract
In this paper, we address the challenging task of scene segmentation. In order to capture the rich contextual dependencies over image regions, we propose Directed Acyclic Graph-Recurrent Neural Networks (DAG-RNN) to perform context aggregation over locally connected feature maps. More specifically, DAG-RNN is placed on top of pre-trained CNN (feature extractor) to embed context into local features so that their representative capability can be enhanced. In comparison with plain CNN (as in Fully Convolutional Networks-FCN), DAG-RNN is empirically found to be significantly more effective at aggregating context. Therefore, DAG-RNN demonstrates noticeably performance superiority over FCNs on scene segmentation. Besides, DAG-RNN entails dramatically less parameters as well as demands fewer computation operations, which makes DAG-RNN more favorable to be potentially applied on resource-constrained embedded devices. Meanwhile, the class occurrence frequencies are extremely imbalanced in scene segmentation, so we propose a novel class-weighted loss to train the segmentation network. The loss distributes reasonably higher attention weights to infrequent classes during network training, which is essential to boost their parsing performance. We evaluate our segmentation network on three challenging public scene segmentation benchmarks: Sift Flow, Pascal Context and COCO Stuff. On top of them, we achieve very impressive segmentation performance.
Bing Shuai, Zhen Zuo, Bing Wang 0003, Gang Wang 0012
IEEE Trans. Pattern Anal. Mach. Intell.2
2018 Multimodal Recurrent Neural Networks With Information Transfer Layers for Indoor Scene Labeling
abstract
This paper proposes a new method called multimodal recurrent neural networks (RNNs) for RGB-D scene semantic segmentation. It is optimized to classify image pixels given two input sources: RGB color channels and depth maps. It simultaneously performs training of two RNNs that are crossly connected through information transfer layers, which are learnt to adaptively extract relevant cross-modality features. Each RNN model learns its representations from its own previous hidden states and transferred patterns from the other RNNs previous hidden states; thus, both model-specific and cross-modality features are retained. We exploit the structure of quad-directional 2D-RNNs to model the short- and long-range contextual information in the 2D input image. We carefully designed various baselines to efficiently examine our proposed model structure. We test our multimodal RNNs method on popular RGB-D benchmarks and show how it outperforms previous methods significantly and achieves competitive results with other state-of-the-art works.
Abrar H. Abdulnabi, Bing Shuai, Zhen Zuo, Lap-Pui Chau, Gang Wang 0012
IEEE Trans. Multim.3
2016 DAG-Recurrent Neural Networks for Scene Labeling
abstract
In image labeling, local representations for image units are usually generated from their surrounding image patches, thus long-range contextual information is not effectively encoded. In this paper, we introduce recurrent neural networks (RNNs) to address this issue. Specifically, directed acyclic graph RNNs (DAG-RNNs) are proposed to process DAG-structured images, which enables the network to model long-range semantic dependencies among image units. Our DAG-RNNs are capable of tremendously enhancing the discriminative power of local representations, which significantly benefits the local classification. Meanwhile, we propose a novel class weighting function that attends to rare classes, which phenomenally boosts the recognition accuracy for non-frequent classes. Integrating with convolution and deconvolution layers, our DAG-RNNs achieve new state-of-the-art results on the challenging SiftFlow, CamVid and Barcelona benchmarks.
Bing Shuai, Zhen Zuo, Bing Wang 0003, Gang Wang 0012
CVPR2
2016 Scene Parsing With Integration of Parametric and Non-Parametric Models
abstract
We adopt convolutional neural networks (CNNs) to be our parametric model to learn discriminative features and classifiers for local patch classification. Based on the occurrence frequency distribution of classes, an ensemble of CNNs (CNN-Ensemble) are learned, in which each CNN component focuses on learning different and complementary visual patterns. The local beliefs of pixels are output by CNN-Ensemble. Considering that visually similar pixels are indistinguishable under local context, we leverage the global scene semantics to alleviate the local ambiguity. The global scene constraint is mathematically achieved by adding a global energy term to the labeling energy function, and it is practically estimated in a non-parametric framework. A large margin-based CNN metric learning method is also proposed for better global belief estimation. In the end, the integration of local and global beliefs gives rise to the class likelihood of pixels, based on which maximum marginal inference is performed to generate the label prediction maps. Even without any post-processing, we achieve the state-of-the-art results on the challenging SiftFlow and Barcelona benchmarks.
Bing Shuai, Zhen Zuo, Gang Wang 0012, Bing Wang 0003
IEEE Trans. Image Process.2
2016 Learning Contextual Dependence With Convolutional Hierarchical Recurrent Neural Networks
abstract
Deep convolutional neural networks (CNNs) have shown their great success on image classification. CNNs mainly consist of convolutional and pooling layers, both of which are performed on local image areas without considering the dependence among different image regions. However, such dependence is very important for generating explicit image representation. In contrast, recurrent neural networks (RNNs) are well known for their ability of encoding contextual information in sequential data, and they only require a limited number of network parameters. Thus, we proposed the hierarchical RNNs (HRNNs) to encode the contextual dependence in image representation. In HRNNs, each RNN layer focuses on modeling spatial dependence among image regions from the same scale but different locations. While the cross RNN scale connections target on modeling scale dependencies among regions from the same location but different scales. Specifically, we propose two RNN models: 1) hierarchical simple recurrent network (HSRN), which is fast and has low computational cost and 2) hierarchical long-short term memory recurrent network, which performs better than HSRN with the price of higher computational cost. In this paper, we integrate CNNs with HRNNs, and develop end-to-end convolutional hierarchical RNNs (C-HRNNs) for image classification. C-HRNNs not only utilize the discriminative representation power of CNNs, but also utilize the contextual dependence learning ability of our HRNNs. On four of the most challenging object/scene image classification benchmarks, our C-HRNNs achieve the state-of-the-art results on Places 205, SUN 397, and MIT indoor, and the competitive results on ILSVRC 2012.
Zhen Zuo, Bing Shuai, Gang Wang 0012, Bing Wang 0003, Yushi Chen 0002
IEEE Trans. Image Process.1
2015 Integrating parametric and non-parametric models for scene labeling
abstract
We adopt Convolutional Neural Networks (CNN) as our parametric model to learn discriminative features and classifiers for local patch classification. As visually similar pixels are indistinguishable from local context, we alleviate such ambiguity by introducing a global scene constraint. We estimate the global potential in a non-parametric framework. Furthermore, a large margin based CNN metric learning method is proposed for better global potential estimation. The final pixel class prediction is performed by integrating local and global beliefs. Even without any post-processing, we achieve state-of-the-art performance on SiftFlow and competitive results on Stanford Background benchmark.
Bing Shuai, Gang Wang 0012, Zhen Zuo, Bing Wang 0003, Lifan Zhao
CVPR3
2015 B-spline-based shape coding with accurate distortion measurement using analytical model
Zhongyuan Lai, Zhen Zuo, Zhijun Yao, Wenyu Liu 0001
Neurocomputing3
2015 Exemplar based Deep Discriminative and Shareable Feature Learning for scene image classification
Zhen Zuo, Gang Wang 0012, Bing Shuai, Lifan Zhao, Qingxiong Yang
Pattern Recognit.1
2015 Quaddirectional 2D-Recurrent Neural Networks For Image Labeling
abstract
We adopt Convolutional Neural Networks (CNN) to learn discriminative features for local patch classification. We further introduce quaddirectional 2D Recurrent Neural Networks to model the long range dependencies among pixels. Our quaddirectional 2D-RNN is able to embed the global image context into the compact local representation, which significantly enhance their discriminative power. Our experiments demonstrate that the integration of CNN and quaddirectional 2D-RNN achieves very promising results which are comparable to state-of-the-art on real-world image labeling benchmarks.
Bing Shuai, Zhen Zuo, Gang Wang 0012
IEEE Signal Process. Lett.2
2014 Learning Discriminative and Shareable Features for Scene Classification
Zhen Zuo, Gang Wang 0012, Bing Shuai, Lifan Zhao, Qingxiong Yang, Xudong Jiang 0001
ECCV (1)1
2012 A two-hop clustered image transmission scheme for maximizing network lifetime in wireless multimedia sensor networks
Zhen Zuo, Qin Lu 0001, Wusheng Luo
Comput. Commun.1
2011 A Hybrid Admissible Distortion Checking Algorithm for the B-Spline-Based Operational Rate-Distortion Optimal Shape Coding
abstract
Admissible distortion checking algorithm plays a very important role in both rate distortion performance and computational efficiency of B-spline-based operational rate-distortion optimal shape coding framework under the minimum-maximum criterion. Existing distortion measurement using chord-length parameterization (DMCLP) is fast but results in extra bit-rate problem. In contrast, the up to date accurate distortion measurement using analytical model (ADMAM) can achieve the smallest bit-rate but is very time consuming. It motivates us to develop a hybrid admissible checking algorithm that can take full use of each advantage. Recalling the definitions of both DMCLP and ADMAM for each associated contour point, the authors show that DMCLP is the distance from the parameterized B-spline point while ADMAM is the shortest distance from the approximating B-splines.
Zhongyuan Lai, Zhen Zuo, Wenyu Liu 0001
DCC2
2011 Accurate Distortion Measurement Using Analytical Model for the B-Spline-Based Shape Coding
abstract
Summary form only given. Existing distortion measurements for the B-spline-based shape coding include ap proximation, quantization, or parameterization process, so they are approximate techniques. They may inaccurately predict the actual distortion value, which motivates us to construct a model that can accurately measure the actual distortion. It was reported that the actual distortion for reconstruction quality assessment is the minimal Euclidean distance between each associated contour point and the reconstruction contour.
Zhongyuan Lai, Zhen Zuo, Wenyu Liu 0001
DCC2
2011 Accurate distortion measurement for B-spline-based shape coding
abstract
In this paper, we present a new contour point distortion measurement, called accurate distortion measurement for B-spline-based shape coding (ADMBSC). Different from existing distortion measurements containing approximation, quantization or parameterization, our distortion is defined as the shortest distance from the original B-spline to the associated contour point. This is in line with the subjective-based objective quality metric. Geometric relationships are introduced to simplify computation, followed by a hybrid admissible distortion checking algorithm to reduce execution time. Theoretical analysis and experimental results demonstrate that when the operational rate-distortion optimal shape coding framework under the minimum-maximum criterion is applied, the ADMBSC can lead to the smallest bit-rate among all the distortion measurements that can guarantee the admissible distortion. Moreover, if the original contour has NCpoints, it takes only O(NC) time for segment distortion measuring paradigms, whose computational complexity is the same as the lowest one among the existing distortion measurements.
Zhongyuan Lai, Zhen Zuo, Zhijun Yao, Wenyu Liu 0001
ICIP2
2000 Detection of Sea Surface Small Targets in Infrared Images Based on Multilevel Filter and Minimum Risk Bayes Test
abstract
This paper discusses the research in small target detection in infrared images with heavy clutter background. For most infrared images, ship objects are rather dim in the relative dark sea surface background. The existence of scan line disturbance and noise also increases the difficulty in proper detection. Dim objects must be distinguished from a dark background. On the other hand, the small targets must also be distinguished from clutters. Through analysis of the targets and background, we build characteristic models of small ship objects, noise and sea backgrounds respectively, and indicate their differences in spatial and frequency domains among them. Based on the principles of signal processing, pattern recognition and artificial intelligence, we propose a combined algorithm for detecting sea surface small targets. In this algorithm, components of background and noise are first suppressed by a multilevel filter designed accordingly, meanwhile enhancing the target ones of interest. The pixels of the candidate targets are then discriminated by minimum risk Bayes test. Finally, according to a priori knowledge about the targets such as the ranges of their sizes, the targets of interest can be detected. In particular, the related probability distributions used by statistic decision are obtained by offline learning of typical training samples. Experiments show that the algorithm is excellent for such kinds of target detection and is robust to noise.
Yiu Sang Moon, Tianxu Zhang, Zhengrong Zuo, Zhen Zuo
Int. J. Pattern Recognit. Artif. Intell.4