EDBT 2026 Demo / reviewers in the wild / expert
Yunzuo Zhang
dblp:200/9699
· DBLP profile ↗
39ranked-venue papers
38as first author
38since 2021 · last 2026
0000-0001-7499-4835ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 16 first-author · 15 since 2021Artificial intelligence and machine learning · 15 · 14 first-author · 15 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 9 first-author · 9 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Combining contrastive frame selector and context-attentive adversarial networks for unsupervised video summarization
Yunzuo Zhang, Liye Xue, Yaoge Xiao, Weiqi Lian |
Image Vis. Comput. | 1 |
| 2026 | ORSI Salient Object Detection via Progressive Interaction and Saliency-Guided EnhancementabstractAlthough great progress has been made in Salient Object Detection (SOD) in Optical Remote Sensing Images (ORSIs), it still faces critical challenges, particularly in dealing with irregular topological structures and complex contextual relationships. To address these issues, we propose a Progressive Interaction and Saliency-guided Enhancement Network (PISENet). Specifically, a Progressive Interaction Encoder (PIE) is proposed, which adopts a dual-path heterogeneous fusion architecture and a hierarchical progressive interaction mechanism to capture global irregular topological structures and local fine-grained image details. Meanwhile, it mitigates semantic gaps across multi-scale features and achieves effective cross-level feature fusion. Subsequently, a Global Context Enhancement Module (GCEM) is designed, which incorporates non-local blocks to enhance spatial correlations among features. A parallel multi-branch structure is further utilized to capture multi-level contextual information ranging from local details to long-range semantics, thereby strengthening the modeling of global context. Finally, a Multi-scale Progressive Attention Enhancement Decoder (MPAED) is devised, which adopts a saliency-guided attention mechanism to jointly model spatial and channel-wise dependencies, enhance responses in salient regions and boundaries, and progressively decode and aggregate deep semantic and shallow detailed features. Extensive experiments on three benchmark datasets demonstrate that our method achieves significant superiority over state-of-the-art approaches. Yunzuo Zhang, Liye Xue, Weiqi Lian, Ran Tao 0003 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2026 | Structure-Aware Dual Semantic Augmentation Alignment for Unsupervised Person Re-IdentificationabstractIn recent years, the inherent noise labels in unsupervised training can easily lead to structural distortion in the feature space and the mutual reinforcement of pseudo-label noise. Existing methods mostly focus on single-scale feature alignment, lacking collaborative perception and alignment of local topological structure and cross-camera distribution structure. To address this, we proposes a Structure-Aware Dual Semantic Augmentation Alignment network (SDAA). Firstly, we design an Adaptive Corrected Spatial Attention (ACSA) to achieve adaptive fusion of local detail structure and global semantic structure through multi-scale hybrid pooling and progressive spatial upsampling. Furthermore, we propose a Structural Semantic Decoupling module (SSD) to suppress microstructural noise by perceiving and reconstructing the local topological structure of the feature space. Finally, a Domain Semantic Alignment module (DSA) achieves explicit alignment of cross-camera distribution structure through camera-aware surrogate memory. Experiments on multiple public datasets, including Market-1501, MSMT17, and PersonX, demonstrate that our proposed method outperforms existing methods, validating the effectiveness and generality of the proposed module. Yunzuo Zhang, Weiqi Lian, Zhiwei Tu, Liye Xue, Ran Tao 0003 |
IEEE Signal Process. Lett. | 1 |
| 2026 | BANet: Bidirectional Feature Aggregation and Adaptive Multi-Scene Perception-Based Lane Detection for Autonomous DrivingabstractLane detection is a hot topic in the field of autonomous driving, providing essential assistance to vehicles. Due to the lack of effective integration among hierarchical features and insufficient adaptability to variations of lanes under diverse background conditions, accurate lane detection remains challenging. To address the aforementioned issues, we propose BANet, an efficient lane detection network based on bidirectional feature aggregation and adaptive multi-scene perception, aiming to improve the lane detection accuracy under different backgrounds. Firstly, we propose a Bidirectional Feature Aggregation Module (BFAM) that, via the proposed composition built from PG2f, effectively integrates complementary information across scales, improving anchor localization. Secondly, we propose an Adaptive Multi-scene Perception Module (AMPM) that learns scene-aware spatial information to enhance lane-relevant cues, addressing the no-visual-clue problem. Finally, we propose an Edge Refinement Attention (ERA), which models spatial and channel information in parallel to refine lane representations. The experimental results on the CULane, CurveLanes, and TuSimple datasets show that the proposed network outperforms existing methods and exhibits excellent performance in the most challenging scenes. Yunzuo Zhang, Zhiwei Tu, Weiqi Lian, Yubo Hu, Shibo Sun, Yaoge Xiao, Yu Cheng 0028 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2025 | Video saliency prediction via single feature enhancement and temporal recurrence
Yunzuo Zhang, Yaoge Xiao, Yuekui Zhang |
Eng. Appl. Artif. Intell. | 1 |
| 2025 | An Efficient Perceptual Video Compression Scheme Based on Deep Learning-Assisted Video Saliency and Just Noticeable Distortion
Yunzuo Zhang, Tian Zhang 0008, Shuangshuang Wang, Puze Yu |
Eng. Appl. Artif. Intell. | 1 |
| 2025 | Spatiotemporal dual-branch feature-guided fusion network for driver attention prediction
Yuekui Zhang, Yunzuo Zhang, Yaoge Xiao |
Expert Syst. Appl. | 2 |
| 2025 | Adaptive Downsampling and Scale Enhanced Detection Head for Tiny Object Detection in Remote Sensing ImageabstractIn recent years, the detection for tiny objects in remote sensing images has become a hot research topic. Tiny objects contain a limited number of pixels and are easily confused with the background, which leads to low detection accuracy. To the end, this letter proposes a tiny object detection method based on adaptive downsampling and scale enhanced detection head (SEDH) to improve the accuracy of detection without increasing the model parameters. First, the dynamic feature extraction module (DFEM) is proposed. The module can obtain the context information of tiny objects. Second, the adaptive downsampling module (ADM) is designed to capture local details of tiny objects. Finally, the scale enhanced detection head is constructed which improves the sensitivity to tiny objects, while reducing the number of parameters of the model. To verify the effectiveness of the proposed method, a series of experiments are conducted on the challenging AI-TOD dataset. The experimental results demonstrate that the proposed method effectively trade-offs the relationship between detection accuracy and the number of model parameters. Yunzuo Zhang, Jiawen Zhen, Yaoxing Kang, Yu Cheng 0028 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2025 | Adaptive Differentiation Siamese Fusion Network for Remote Sensing Change DetectionabstractRemote sensing images’ change detection is crucial for disaster monitoring, urban planning, and environmental surveillance. Despite recent advancements in deep learning enhancing change detection methods, challenges persist in detecting small object changes and distinguishing pseudochanges. To address these issues, we propose the adaptive differentiation Siamese fusion network (ADSFNet). ADSFNet features a Siamese encoder-decoder architecture designed to overcome these challenges. The innovative Siamese encoder replaces traditional self-attention with dilated neighborhood attention, enhancing the detection of small objects. In addition, the specific feature enhancer (SFE) captures differences between bitemporal features, improving performance when pseudochanges are present. The multilevel differential feature adaptive fusion module (MDFAFM) integrates differential features across various levels, facilitating the prediction of accurate change maps. Experimental results on benchmark datasets demonstrate that the proposed ADSFNet significantly outperforms existing state-of-the-art (SOTA) methods in accuracy. Yunzuo Zhang, Jiawen Zhen, Yuehui Yang, Yu Cheng 0028 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2025 | Edge aware adaptive fusion network for video salient object detection
Yunzuo Zhang, Shuangshuang Wang, Jiawen Zhen, Puze Yu |
Pattern Recognit. Lett. | 1 |
| 2025 | Cross-erasure enhanced network for occluded person re-identification
Yunzuo Zhang, Yuehui Yang, Weili Kang, Jiawen Zhen |
Pattern Recognit. Lett. | 1 |
| 2025 | Pyramid-structured multi-scale transformer for efficient semi-supervised video object segmentation with adaptive fusion
Yunzuo Zhang, Puze Yu, Yaoge Xiao, Shuangshuang Wang |
Pattern Recognit. Lett. | 1 |
| 2025 | Interleaved Dynamic Fusion Network for Occluded Person Re-IdentificationabstractMost existing occluded person re-identification methods use a part-based approach to extract pedestrian features. The extracted part features are isolated from each other, resulting in insufficient information exchange between part features. To address this issue, we propose an interleaved dynamic fusion network (IDFNet) for occluded person re-identification. Initially, an interleaved feature pyramid module (IFPM) was designed, which recursively transmits rich semantic information from high-level feature maps to the bottom layer through interleaved connections, achieving the extraction of multi-scale information. Secondly, a multi-scale feature dynamic fusion module (MDFM) to effectively integrate multi-scale information in IFPM and achieve cross-scale feature fusion. It allows the network to dynamically select the most suitable features for fusion based on pedestrian characteristics and size. Finally, the designed feature interaction module (FIM) uses different semantic part features as graph nodes, allowing information transfer between nodes, suppressing the transfer of meaningless feature information such as occlusion, promoting the transfer of semantic feature information, and effectively alleviating occlusion problems. Extensive experimental results on both occluded and holistic datasets demonstrate the efficacy of our approach. Yunzuo Zhang, Weiqi Lian, Yuehui Yang, Shuangshuang Wang, Jiawen Zhen |
IEEE Signal Process. Lett. | 1 |
| 2024 | Contextual Correspondence Matters: Bidirectional Graph Matching for Video Summarization
Yunzuo Zhang, Yameng Liu |
ECCV (87) | 1 |
| 2024 | M2SUM: Multi-Granularity Scale-Adaptive Video Summarizer towards Informative Context Representation LearningabstractVideo summarization intends to automatically select meaningful segments from untrimmed videos. Although previous efforts have achieved remarkable progress, they still struggle to robustly aggregate and effectively process multi-granularity contextual information within videos, which hinders understanding towards video content. To address these issues, we propose M2SUM, which is composed of three dominant components including the embedding learning attention (ELA) module, multi-granularity aggregator (MGA), and semantic scale-adaption (SSA) module. ELA dynamically enhances pre-trained visual features by considering the similarity relationship across frame-level and video-level embeddings. MGA incorporates self-attention and temporal convolution into a unified learnable module, robustly learning long-range and short-range multi-granularity temporal dependencies. SSA is exploited to adaptively perform representation fusion after deep interaction across multi-granularity temporal dependencies. According to the fused representations, M2SUM predicts importance scores and generates video summaries. Extensive experiments on standard datasets have proved the effectiveness and superiority of our method in F-score and rank-based evaluations. Yunzuo Zhang, Yameng Liu, Weili Kang |
ICASSP | 1 |
| 2024 | ECPNet: An Enhanced Curve Perception Network for Lane DetectionabstractLane detection methods based on anchors have received increasing attention, but fixed-shape anchors make it difficult to model complex lane line shapes. To solve this problem, we propose an Enhanced Curve Perception Network (ECPNet). Specifically, we propose a Layer-by-layer Context Fusion (LCF) module to fully utilize both high-level and low-level features in lane detection by establishing short hop connections across feature layers of diverse scales. Then, we propose a novel Structural Correction Prediction (SCP) module, which enhances the detection ability of the model on the curve structure lane by dynamically guiding the selection of anchor classification patterns. In addition, ECPNet adaptive calibration pays attention to channel features through the Cross-Channel Attention (CCA) mechanism. Experiments on the two most representative datasets demonstrate that the proposed method achieves state-of-the-art performance in a variety of environments, especially in curvy lanes. Yunzuo Zhang, Cunyu Wu, Yameng Liu |
ICASSP | 1 |
| 2024 | Attention-guided multi-granularity fusion model for video summarization
Yunzuo Zhang, Yameng Liu, Cunyu Wu |
Expert Syst. Appl. | 1 |
| 2024 | CFFM: Multi-task lane object detection method based on cross-layer feature fusion
Yunzuo Zhang, Zhiwei Tu, Cunyu Wu, Tian Zhang 0008 |
Expert Syst. Appl. | 1 |
| 2024 | Surveillance video synopsis framework base on tube set
Yunzuo Zhang, Puze Yu |
J. Vis. Commun. Image Represent. | 1 |
| 2024 | Key frame extraction method for lecture videos based on spatio-temporal subtitles
Yunzuo Zhang, Shui Lam |
Multim. Tools Appl. | 1 |
| 2024 | Multi-scale occlusion suppression network for occluded person re-identification
Yunzuo Zhang, Yuehui Yang, Weili Kang, Jiawen Zhen |
Pattern Recognit. Lett. | 1 |
| 2024 | VSS-Net: Visual Semantic Self-Mining Network for Video SummarizationabstractVideo summarization, with the target to detect valuable segments given untrimmed videos, is a meaningful yet understudied topic. Previous methods primarily consider inter-frame and inter-shot temporal dependencies, which might be insufficient to pinpoint important content due to limited valuable information that can be learned. To address this limitation, we elaborate on a Visual Semantic Self-mining Network (VSS-Net), a novel summarization framework motivated by the widespread success of cross-modality learning tasks. VSS-Net initially adopts a two-stream structure consisting of a Context Representation Graph (CRG) and a Video Semantics Encoder (VSE). They are jointly exploited to establish the groundwork for further boosting the capability of content awareness. Specifically, CRG is constructed using an edge-set strategy tailored to the hierarchical structure of videos, enriching visual features with local and non-local temporal cues from temporal order and visual relationship perspectives. Meanwhile, by learning visual similarity across features, VSE adaptively acquires an instructive video-level semantic representation of the input video from coarse to fine. Subsequently, the two streams converge in a Context-Semantics Interaction Layer (CSIL) to achieve sophisticated information exchange across frame-level temporal cues and video-level semantic representation, guaranteeing informative representations and boosting the sensitivity to important segments. Eventually, importance scores are predicted utilizing a prediction head, followed by key shot selection. We evaluate the proposed framework and demonstrate its effectiveness and superiority against state-of-the-art methods on the widely used benchmarks. Yunzuo Zhang, Yameng Liu, Weili Kang, Ran Tao 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | SFSANet: Multiscale Object Detection in Remote Sensing Image Based on Semantic Fusion and Scale AdaptabilityabstractIn the field of computer vision, remote sensing image object detection plays an important role. Although the object detection algorithm has made significant progress, there are still problems in detecting objects with multi-scale in remote sensing image. Due to the insufficient utilization of object feature information, the detection accuracy of multi-scale objects is very low. To address the aforementioned issues, this paper proposes an effective object detection algorithm for remote sensing image based on semantic fusion and scale adaptability, known as SFSANet. Firstly, in view of the problem that the existing methods ignore the semantic differences between different scale feature maps, the semantic fusion (SF) module is proposed to enrich the semantic information and improve the ability to classify and locate objects. Next, to address the issue of the objects being easily interfered in complex background and the detection performance is poor, the spatial location attention (SLA) module is constructed to suppress background information and make key objects more prominent. Additionally, the scale adaptability module (SA) is designed to enrich the expression of feature information, realize the integration of global and local information, and ensure the integrity of image structure. Finally, we adopt the SIoU loss function as the localization loss to expedite model convergence. In order to verify the effectiveness of the proposed method, we conduct experiments on the mainstream datasets DIOR and NWPU VHR-10, which fully demonstrate the superiority of the proposed method. Yunzuo Zhang, Puze Yu, Shuangshuang Wang, Ran Tao 0003 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Full-Scale Feature Aggregation and Grouping Feature Reconstruction-Based UAV Image Target DetectionabstractUnmanned Aerial Vehicle (UAV) image target detection holds significant value for a wide range of applications in modern society. However, due to the variable flight altitude of UAV, the captured images often exhibit significant differences at the target scale and contain a large number of small targets. The existing methods are difficult to adapt to these changes, resulting in a decrease in detection accuracy. To address this issue, this paper proposes a new method for UAV image object detection based on full scale feature aggregation and grouped feature reconstruction FFAGRNet. Firstly, existing feature fusion methods are hindered by the layer-by-layer transfer structure, which limits effective information exchange between feature maps of different scales. In response, we propose the Full-scale Feature Aggregation module (FFA), which performs scale adaptation and information aggregation across multiple sets of feature maps, producing high-quality aggregated feature maps. Secondly, to further refine aggregation features and eliminate redundancy, we introduce the Grouping Feature Reconstruction module (GFR). This module subdivides aggregation features into multiple sub-level features, allowing them to autonomously learn channel and spatial layouts of target features. Lastly, we present the Parallel Super-resolution Semantic Enhancement module (PSSE) to reconstruct deep feature maps and incorporate spatial contextual information, effectively increasing the proportion of semantic information and enhancing the model’s ability to classify ambiguous targets. To validate the effectiveness of our proposed method, extensive experiments were conducted on the VisDrone2021 and UAVDT datasets. The results demonstrate that compared to the baseline, our method achieves a significant improvement in mAP50, with increases of 7.6% and 4.6% respectively, showcasing excellent performance compared to existing methods. Yunzuo Zhang, Cunyu Wu, Tian Zhang 0008 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Multi-Scale Spatiotemporal Feature Fusion Network for Video Saliency PredictionabstractRecently, video saliency prediction has attracted increasing attention, yet the improvement of its accuracy is still subject to the insufficient use of multi-scale spatiotemporal features. To address this issue, we propose a 3D convolutional Multi-scale Spatiotemporal Feature Fusion Network (MSFFNet) to achieve the full utilization of spatiotemporal features. Specifically, we propose a Bi-directional Temporal-Spatial Feature Pyramid (BiTSFP), the first application of bi-directional fusion architectures in this field, which adds the flow of shallow location information on the basis of the previous flow of deep semantic information. Then, different from simple addition and concatenation, we design an Attention-Guided Fusion (AGF) mechanism that can adaptively learn the fusion weights of adjacent features to integrate them appropriately. Moreover, a Framewise Attention (FA) module is introduced to selectively emphasize the useful frames, augmenting the multi-scale temporal features to be fused. Our model is simple but effective, and it can run in real-time. Experimental results on the DHF1K, Hollywood-2, and UCF-sports datasets demonstrate that the proposed MSFF-Net outperforms existing state-of-the-art methods in accuracy. Yunzuo Zhang, Tian Zhang 0008, Cunyu Wu, Ran Tao 0003 |
IEEE Trans. Multim. | 1 |
| 2023 | Joint Multi-Level Feature Network for Lightweight Person Re-IdentificationabstractLearning fine-grained features is crucial to the performance improvement of person re-identification (Re-ID). Although existing methods have made significant progress, utilizing multi-level information to obtain fine-grained features has not been explored in this field. To alleviate this issue, we propose a lightweight person Re-ID method named Joint Multi-Level Feature Network (JMLFNet) to obtain robust feature representation for the Re-ID task. Specifically, we design a Multi-Attention Block (MAB) and embed it into the lightweight backbone network to improve performance, which can make the network focus on the key parts of pedestrian images. Meanwhile, we propose a Multi-Level Feature Extraction (MLFE) method to extract multi-granularity features of high-level semantic information and low-level detail information, which can effectively capture the feature diversity of pedestrian images. Furthermore, we design a Feature Fusion Block (FFB), which is fused the fine-grained features of high-level and low-level information to better obtain the discriminative feature representation of pedestrian images. Extensive experiments conducted on popular datasets Market1501 and DukeMTMC-reID demonstrate that the proposed JMLFNet has competitive performance compared with the state-of-the-art methods. Yunzuo Zhang, Weili Kang, Yameng Liu, Pengfei Zhu 0005 |
ICASSP | 1 |
| 2023 | Hierarchical Spatiotemporal Feature Fusion Network For Video Saliency PredictionabstractCurrent video saliency prediction methods have made great progress relying on the feature extraction capability of CNN, but there are still many defects in hierarchical feature fusion, limiting the further improvement of accuracy. To address this issue, we propose a 3D convolutional Hierarchical Spatiotemporal Feature Fusion Network (HSFF-Net). Specifically, we propose a Bi-directional Temporal-Spatial Feature Pyramid (BiTSFP), the first application of bi-directional fusion architectures in this field, which adds the flow of shallow location information on the flow of deep semantic information. Then, different from addition and concatenation, we design a Hierarchical Adaptive Fusion (HAF) mechanism that can adaptively learn the fusion weights of adjacent features. Moreover, a Frame-wise Attention (FA) module is introduced to augment the temporal features to be fused. Our model is simple yet effective and can run in real-time. Experimental results on the three video saliency benchmarks demonstrate that the HSFF-Net outperforms existing state-of-the-art methods in accuracy. Yunzuo Zhang, Cunyu Wu |
ICASSP | 1 |
| 2023 | Object interaction-based surveillance video synopsis
Yunzuo Zhang |
Appl. Intell. | 1 |
| 2023 | Multi-Scale Semantic and Detail Extraction Network for Lightweight Person Re-Identification
Yunzuo Zhang, Weili Kang, Yameng Liu, Pengfei Zhu 0005 |
Comput. Vis. Image Underst. | 1 |
| 2023 | Key frame extraction based on quaternion Fourier transform with multiple features fusion
Yunzuo Zhang, Ruixue Liu, Pengfei Zhu 0005, Yameng Liu |
Expert Syst. Appl. | 1 |
| 2023 | Accurate video saliency prediction via hierarchical fusion and temporal recurrence
Yunzuo Zhang, Cunyu Wu |
Image Vis. Comput. | 1 |
| 2023 | Self-Attention Guidance and Multiscale Feature Fusion-Based UAV Image Object DetectionabstractObject detection on UAV images is a recent research hotspot. Existing object detection methods have achieved good results on general scenes, but there are inherent challenges with UAV images. The detection accuracy of UAV images is limited by complex backgrounds, significant scale differences, and densely arranged small objects. To solve these problems, we propose a UAV image object detection network based on Self-attention Guidance and Multi-scale Feature fusion (SGMFNet). Firstly, we design a Global-Local Feature Guidance module (GLFG). This module can effectively combine local information and global information, which makes the model focus on the object area and reduces the impact of complex background. Secondly, an improved Parallel Sampling Feature Fusion module (PSFF) is designed to efficiently fuse multi-scale features. Thirdly, we design an Inverse-residual Feature Enhancement module (IFE), which is embedded in the front of the newly added detection head to enhance feature extraction on small objects. Finally, we conduct a large number of experiments on the VisDrone2019 dataset. The results show that the proposed SGMFNet outperforms other popular methods, and has achieved good results in many scenarios. Yunzuo Zhang, Cunyu Wu, Tian Zhang 0008, Yameng Liu |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2023 | Enhancement multi-module network for few-shot leaky cable fixture detection in railway tunnel
Yunzuo Zhang, Zhouchen Song |
Signal Process. Image Commun. | 1 |
| 2023 | MAR-Net: Motion-Assisted Reconstruction Network for Unsupervised Video SummarizationabstractVideo summarization targets to extract the most important segments from a video by spatiotemporal analysis. Previous methods primarily learn content within videos based on appearance information, with a rare discussion on the effective utilization of motion information, which is equally essential to video understanding. In this letter, we expound upon a Motion-Assisted Reconstruction Network (MAR-Net), which synergistically models appearance and motion information within videos for unsupervised video summarization without any manual annotations. MAR-Net notably comprises a Bidirectional Modality Encoder (BiME) and a Video Context Navigator (VCN). By integrating uni-modal and cross-modal feature aggregation into a unified module, BiME allows for exploring sophisticated dependency relationships among features through a bidirectional attention mechanism. VCN can promote the semantic consistency between the cross-modal contexts and the input video by a consistency loss term, alleviating the noisy impact within the motion stream. Empirical results conducted on benchmark datasets demonstrate that MAR-Net outperforms other state-of-the-art methods. Yunzuo Zhang, Yameng Liu, Weili Kang |
IEEE Signal Process. Lett. | 1 |
| 2023 | FANet: An Arbitrary Direction Remote Sensing Object Detection Network Based on Feature Fusion and Angle ClassificationabstractHigh-precision remote sensing image object detection has broad application prospects in military defense, disaster emergency, urban planning, and other fields. However, the arbitrary orientation, dense arrangement, and small size of objects in remote sensing images lead to poor detection accuracy of existing methods. To achieve accurate detection, this paper proposes an arbitrary directional remote sensing object detection method, called FANet, based on feature fusion and angle classification. Initially, the angle prediction branch is introduced, and the circular smooth label method is used to transform the angle regression problem into a classification problem, which solves the difficult problem of abrupt changes in the boundaries of the rotating frame while realizing the object frame rotation. Subsequently, to extract robust remote sensing objects, innovative introduce pure convolutional model as a backbone network, while Conv is replaced by GSConv to reduce the number of parameters in the model along with ensuring detection accuracy. Finally, the strengthen connection feature pyramid network (SC-FPN) is proposed to redesign the lateral connection part for deep and shallow layer feature fusion, and add jump connections between the input and output of the same level feature map to enrich the feature semantic information. In addition, add a variable parameter to the original localization loss function to satisfy the bounding box regression accuracy under different IoU thresholds, and thus obtain more accurate object detection. The comprehensive experimental results on two public datasets for rotated object detection DOTA and HRSC2016 demonstrate the effectiveness of our method. Yunzuo Zhang, Cunyu Wu, Wei Li 0032, Ran Tao 0003 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | CFANet: Efficient Detection of UAV Image Based on Cross-Layer Feature AggregationabstractWith the rapid development of the unmanned aerial vehicle (UAV) industry, UAV image object detection technology has become a hotspot. However, due to a large number of dense small objects in UAV images, quickly and effectively detecting objects and achieving accurate classification is still a challenge. With this observation, we propose an efficient object detection network for UAV images based on cross-layer feature aggregation (CFANet). Firstly, we design a novel cross-feature aggregation module (CFA) to aggregate features at different scales on the basis of avoiding semantic gaps, so as to replace common features for feature fusion and achieve accurate detection. This method makes up for the defect that the layer-by-layer feature transfer method only focuses on the features of the previous layer and cannot fully integrate spatial and semantic information. Secondly, a layered associative spatial pyramid pooling module (LASPP) is proposed to capture context information while maintaining the sensitivity of feature maps at different layers to detail information. Thirdly, the alpha-IoU loss function is introduced to accelerate the convergence speed of the model and improve the detection accuracy. Finally, an adaptive overlapping slice (AOS) for high-resolution images is proposed to protect the integrity of the object when slicing. To verify the effectiveness of the proposed method, extensive experiments on challenge datasets for object detection in UAV images VisDrone2021 and UAVDT datasets are carried out. The results show that, compared with the other most advanced detectors, the proposed method can achieve significant performance on the basis of ensuring real-time detection. Yunzuo Zhang, Cunyu Wu, Tian Zhang 0008, Wei Li 0032 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2022 | Adaptive Spatio-Temporal Tube for Fast Motion Segments Extraction of VideosabstractExisting motion segments extraction methods suffer from the problem of high computation complexity. To address this issue, we propose a method called adaptive spatio-temporal tube for fast motion segments extraction of videos. Firstly, initial spatio-temporal flow of sub-videos divided from input video is computed by adopting a novel Area-adjusted Spatio-Temporal Tunnel (A-STT) to screen preliminarily moiton segments. Secondly, the Sampling-line Adjustment Mechanism (SAM) is presented to avoid processing the entire amount of video spatial data and reduce computational complexity. The SAM is created by analyzing object consistency to produce a Sampling-line Adjustment Factor (SAF) which is used to dynamically obtain the sampling-line of various sub-videos. Finally, the adaptive spatio-temporal tubes are generated by integrating the initial spatio-temporal flow and SAF, which ensures the robustness of the proposed method. The proposed method is experimented on the public datasets VISOR, CAVIR and self-collected dataset. The experimental results demonstrate that the proposed method outperforms the state-of-the-art methods in terms of both computing speed and accuracy. Yunzuo Zhang, Kaina Guo, Ran Tao 0003 |
IEEE Signal Process. Lett. | 1 |
| 2022 | Joint Reinforcement and Contrastive Learning for Unsupervised Video SummarizationabstractThis letter presents a joint Reinforcement and Contrastive Learning framework termed RCL for unsupervised video summarization, aiming at addressing the existing two shortcomings: (i) poor feature representation, and (ii) inefficient context modeling capability. Concretely, the proposed framework consists of an Optimized Coding Module (OCM) and a Dissimilarity-Guided Attention Graph (DGAG). The OCM is grounded on Gate Recurrent Unit (GRU), which encodes the content within shots into concise representations. Different from the existing approaches, contrastive learning is introduced to promote discriminative and informative feature learning. Afterward, the DGAG adaptively performs feature aggregation by evaluating the semantic dissimilarity across shots to eliminate chaotic message passing for accurate context modeling. Finally, the procedure of importance score prediction is formulated as a node classification task, and these scores are utilized for a summary generation. Extensive experiments on the benchmark datasets demonstrate the superior performance of the proposed method. Yunzuo Zhang, Yameng Liu, Pengfei Zhu 0005, Weili Kang |
IEEE Signal Process. Lett. | 1 |
| 2017 | Motion-State-Adaptive Video Summarization via Spatiotemporal AnalysisabstractWith the explosive growth of video data, managing and browsing videos in a timely and effective manner has become an urgent problem, particularly in surveillance applications. Video summarization as a feasible solution is considerably attracting more attention. In this paper, we propose a novel motion-state-adaptive video summarization method based on spatiotemporal analysis. To overcome the low efficiency of traditional video summarization, the proposed method utilizes spatiotemporal slices to analyze object motion trajectories and selects motion state changes as a metric to summarize videos. Initially, a motion-active segment is detected using motion power. Subsequently, motion state changes are modeled as a collinear segment on a spatiotemporal slice (STS-CS) and an attention curve based on the STS-CS model is formed to extract the key frames. Finally, a visually distinguishing mechanism is employed to refine the key frames. The experimental results demonstrate that the proposed method outperforms the existing state-of-the-art methods in terms of both computational efficiency and detailed video motion dynamic maintenance. This is accomplished with a comparable subjective performance. Yunzuo Zhang, Ran Tao 0003, Yue Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |