Shangdong Zheng

dblp:254/0296 · DBLP profile ↗
← Back
10ranked-venue papers
5as first author
9since 2021 · last 2026
0000-0002-7444-3459ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 6 · 4 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Data-Driven Bidirectional Spatial-Adaptive Network for Weakly Supervised Object Detection in Remote Sensing Images
abstract
Weakly-supervised object detection (WSOD) learns detectors with only image-level classification annotations. Without precise instance-level labels, most previous WSOD methods in remote sensing images (RSIs) select the highest-scoring proposals as the final detection results, which are confronted by two major challenges: (1) instances with small scale or rare poses are easily neglected; (2) optimizing network by the top-scoring region inevitably overlooks many valuable candidate proposals. To mitigate the above-mentioned challenges, we propose a data-driven bidirectional spatial-adaptive network (BSANet). It contains a forward-reverse spatial dropout (FRSD) module to reduce instance ambiguity induced from extreme scales and poses, as well as crowded scene, and to better excavate the entire instances. From attention learning perspective, the proposed FRSD is conceptually similar to a data-driven hard attention mechanism, which adaptively samples and reconstructs the spatially related regions for mining more latent feature responses. Meanwhile, our FRSD effectively alleviates the inherent problem that non-parametric hard attention learning fashion cannot adapt to different datasets. In addition, we build a soft attention branch to simultaneously model soft pixel-level and hard region-level attention information for exploring the complementary benefit between soft and hard attention learning. We evaluate our BSANet on the challenging NWPU VHR-10.v2 and DIOR datasets. Experimental results demonstrate that our method sets a new state-of-the-art.
Zebin Wu 0001, Shangdong Zheng, Yang Xu 0006, Le Wang 0003, Zhihui Wei, Gang Hua 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 Detector With Classifier2: An End-to-End Multi-Stream Feature Aggregation Network for Fine-Grained Object Detection in Remote Sensing Images
abstract
Fine-grained object detection (FGOD) fundamentally comprises two primary tasks: object detection and fine-grained classification. In natural scenes, most FGOD methods benefit from higher instance resolution and fewer environmental variation, attributing more commonly associated with the latter task. In this paper, we propose a unified paradigm named Detector with Classifier2 (DC2), which provides a holistic paradigm by explicitly considering the end-to-end integration of object detection and fine-grained classification tasks, rather than prioritizing one aspect. Initially, our detection sub-network is restricted to only determining whether the proposal is a coarse-category and does not delve into the specific sub-categories. Moreover, in order to reduce redundant pixel-level calculation, we propose an instance-level feature enhancement (IFE) module to model the semantic similarities among proposals, which poses great potential for locating more instances in remote sensing images (RSIs). After obtaining the coarse detection predictions, we further construct a classification sub-network, which is built on top of the former branch to determine the specific sub-categories of the aforementioned predictions. Importantly, the detection network is performed on the complete image, while the classification network conducts secondary modeling for the detected regions. These operations can be denoted as the global contextual information and local intrinsic cues extractions for each instance. Therefore, we propose a multi-stream feature aggregation (MSFA) module to integrate global-stream semantic information and local-stream discriminative cues. Our whole DC2 network follows an end-to-end learning fashion, which effectively excavates the internal correlation between detection and fine-grained classification networks. We evaluate the performance of our DC2 network on two benchmarks SAT-MTB and HRSC2016 datasets. Importantly, our method achieves the new state-of-the-art results compared with recent works (approximately 7% mAP gains on SAT-MTB) and improves baseline by a significant margin (43.2% $v.s.~36.7$ %) without any complicated post-processing strategies. Source codes of the proposed methods are available at https://github.com/zhengshangdong/DC2.
Shangdong Zheng, Zebin Wu 0001, Yang Xu 0006, Chengxun He, Zhihui Wei
IEEE Trans. Image Process.1
2024 Patchwise Temporal-Spatial Feature Aggregation Network for Object Detection in Satellite Video
abstract
In this letter, we propose a patchwise temporal-spatial feature aggregation (PTFA) network for object detection in satellite video. First, the feature extractor processes the key frame (KF) along with its support frames to ensure comprehensive spatial coverage of potential objects. Subsequently, we model the semantic similarities among instance-level proposals to exploring robust interaction between temporally adjacent support frames and KF. Furthermore, due to the extremely small size of objects in satellite video, we crop the input frames to different patches by the fixed criterion. Then, the temporal-spatial feature aggregation (TSFA) operations are performed on instance-level RoI features, which attains more nuanced and comprehensive descriptors from the explicit high-resolution temporal-spatial features. The patch features are reconstructed to the original one for complementing more valid feature responses. Finally, we compare our PTFA network with many recent works on the SAT-MTB dataset. Extensive experiments demonstrate that our method achieves the state-of-the-art performance than various static image and video object detection (VID) approaches.
Shangdong Zheng, Zebin Wu 0001, Yang Xu 0006, Pengfei Liu 0002, Zhihui Wei
IEEE Geosci. Remote. Sens. Lett.1
2024 Hyperspectral Images Single-Source Domain Generalization Based on Nonlinear Sample Generation
abstract
In hyperspectral cross-scene classification tasks, it is often challenging to obtain target domain samples during the training phase. Therefore, models need to be trained on one or multiple source domains and achieve good generalization performance on unknown target domains, known as domain generalization. The presence of domain shift limits the model’s generalization across different domains, while the unknown target domain makes it difficult to accurately characterize the distribution differences between domains. To address this issue, we propose a generalization network based on nonlinear sample generation. The network divides the sample features into invariant features and variant features and generates samples by applying nonlinear transformations to the variant features. To ensure the quality of the generated samples, we introduce contrastive learning into the model. It ensures consistency in similarity between the generated samples and the source samples while maintaining a certain degree of dissimilarity. Experiments conducted on four cross-domain adaptive scenarios demonstrate the superior performance of our proposed method.
Biqi Wang, Yang Xu 0006, Zebin Wu 0001, Shangdong Zheng, Zhihui Wei, Jocelyn Chanussot
IEEE Trans. Geosci. Remote. Sens.4
2024 Accelerating Hyperspectral Anomaly Detection With Enhanced Multivariate Gaussianization Based on FPGA
abstract
Hyperspectral anomaly detection (AD), as a frontier research topic in the field of remotely sensed data processing, aims to identify targets of interest from complex and vast images. Existing AD methods typically involve complex models and many parameters, posing challenges in meeting the requirements of computational efficiency in hyperspectral AD. To address this issue, this article presents an AD acceleration algorithm based on the multivariate Gaussian model as well as its field programmable gate array (FPGA) implementation. By exploiting the parallel processing capabilities of FPGA, we introduce an innovative spectral dimensionality reduction method in which the data processing flow can be accomplished in a distributed manner. Then, we employ an improved linear rotation strategy based on correlation coefficients to accelerate the convergence rate of the proposed AD algorithm. The rotation of Gaussianization in the improved strategy is independent of eigenvalue decomposition, thereby substantially reducing the computational complexity involved during the rotation procedure. Furthermore, we apply a pipeline parallel mechanism to facilitate the FPGA implementation of the AD algorithm and to significantly enhance the computational efficiency. Experimental results on an embedded FPGA platform demonstrate that the FPGA implementation of the hyperspectral AD algorithm proposed in this article achieves a significant acceleration rate with guaranteed high detection accuracy.
Zebin Wu 0001, Jin Sun 0001, Yi Zhang 0025, Yang Xu 0006, Zhihui Wei, Shangdong Zheng
IEEE Trans. Geosci. Remote. Sens.7
2024 Oriented Object Detection for Remote Sensing Images via Object-Wise Rotation-Invariant Semantic Representation
abstract
Oriented object detection (OOD) in remote sensing images (RSIs) remains a challenging work due to an arbitrary orientation of instance. Learning rotation-invariant features is critical in modeling a fixed descriptor for instances with its rotated variants. However, most existing methods construct the descriptor from the perspectives of data or feature augmentation, but ignore the exploration of potentially useful supervision information inside the detection algorithm. In this paper, we propose an object-wise rotation-invariant semantic representation (ORSR) framework, which synergizes the exploration of latent supervision, rotation-invariant learning, and guided attention mechanism into a unified network to boost the performance of OOD in RSIs. First, supervised by our constructed pseudo ground truth of segmentation masks, a semantic segmentation branch is built along with the detection algorithm to refine the representation of backbone features. Moreover, a consistency loss function is proposed to encourage the segmentation branch to make the fixed predictions for backbone features with its rotated variants. Considering that segmentation predictions remain the same affine transformations before and after rotating, we further construct a Kullback-Leibler (KL) Divergence based similarity loss function for encouraging the network to model the rotation-invariant features. Finally, we separate the ”object” descriptor from the segmentation predictions to extend the implicit constraint in our proposed semantic segmentation branch. The separated ”object” descriptor not only involves the spatial regularizer to emphasize the high-responsive regions in image, but also can be guided by the constructed consistency loss function. We evaluate our proposed ORSR on the challenging DOTA, DIOR-R, and HRSC2016 datasets. Extensive experiments demonstrate that the proposed ORSR achieves competitive performance compared to other single-scale and multi-scale detection methods.
Shangdong Zheng, Zebin Wu 0001, Qian Du 0001, Yang Xu 0006, Zhihui Wei
IEEE Trans. Geosci. Remote. Sens.1
2024 Multi-Dimensional Visual Data Restoration: Uncovering the Global Discrepancy in Transformed High-Order Tensor Singular Values
abstract
The recently proposed high-order tensor algebraic framework generalizes the tensor singular value decomposition (t-SVD) induced by the invertible linear transform from order-3 to order-d ( ). However, the derived order-d t-SVD rank essentially ignores the implicit global discrepancy in the quantity distribution of non-zero transformed high-order singular values across the higher modes of tensors. This oversight leads to suboptimal restoration in processing real-world multi-dimensional visual datasets. To address this challenge, in this study, we look in-depth at the intrinsic properties of practical visual data tensors, and put our efforts into faithfully measuring their high-order low-rank nature. Technically, we first present a novel order-d tensor rank definition. This rank function effectively captures the aforementioned discrepancy property observed in real visual data tensors and is thus called the discrepant t-SVD rank. Subsequently, we introduce a nonconvex regularizer to facilitate the construction of the corresponding discrepant t-SVD rank minimization regime. The results show that the investigated low-rank approximation has the closed-form solution and avoids dilemmas caused by the previous convex optimization approach. Based on this new regime, we meticulously develop two models for typical restoration tasks: high-order tensor completion and high-order tensor robust principal component analysis. Numerical examples on order-4 hyperspectral videos, order-4 color videos, and order-5 light field images substantiate that our methods outperform state-of-the-art tensor-represented competitors. Finally, taking a fundamental order-3 hyperspectral tensor restoration task as an example, we further demonstrate the effectiveness of our new rank minimization regime for more practical applications. The source codes of the proposed methods are available at https://github.com/CX-He/DTSVD.git.
Chengxun He, Yang Xu 0006, Zebin Wu 0001, Shangdong Zheng, Zhihui Wei
IEEE Trans. Image Process.4
2023 Instance-Aware Spatial-Frequency Feature Fusion Detector for Oriented Object Detection in Remote-Sensing Images
abstract
In recent years, fusing multi-type features poses great potential for oriented object detection (OOD) in remote sensing images (RSIs). Due to the inexplicit operation of modeling orientation variations, convolutional neural networks (CNNs) are difficult to perceive objects under different transformations (angles and scales). In this paper, we propose a novel instance-aware spatial-frequency feature fusion detector (SFFD) for oriented object detection in remote sensing images. First, a layer-wise frequency-domain analysis (L-FDA) module is built along with CNN layers to extract frequency features. Getting rid of the constrains such as horizontal rectangular kernel in CNNs, our L-FDA possesses outstanding ability of locating mutational signals from frequency space. These mutational signals record the scale and angle information of the oriented instances in images. Subsequently, CNN and frequency features are sent into RoI Pooling layer to obtain multi-type instance-level RoI features. Moreover, the proposed instance-aware cross feature fusion (CFF) module explores the interaction between these diverse features which provides an explicit indicator to compensate the orientation information ignored by instance-level CNN features. Finally, our SFFD unifies the proposed L-FDA module and CFF module into the detection network to localize oriented instances in RSIs. We compare our method with many state-of-the-art methods on DOTA, HRSC2016, and NWPU VHR-10 datasets. Experimental results verify the validity of modeling instance-level object relations from frequency-domain and CNNs for OOD.
Shangdong Zheng, Zebin Wu 0001, Yang Xu 0006, Zhihui Wei
IEEE Trans. Geosci. Remote. Sens.1
2022 Learning Orientation Information From Frequency-Domain for Oriented Object Detection in Remote Sensing Images
abstract
Object detection in remote sensing images (RSIs) poses great difficulties due to arbitrary orientations, various scales and dense location of the targets over the ground. Recent evidence suggests that encoding the orientation information is of great use for training an accurate object detector for oriented object detection (OOD). In this paper, we propose a new frequency-domain orientation learning (FDOL) module with two main components: the frequency domain feature extraction (FFE) network and an orientation enhanced self-attention layer (OES-Layer). The FFE network models the interactions among spatial locations in the frequency domain to determine the frequency of spatial features. Then, these features are fed into our OES-Layer to learn the orientation information. Moreover, the orientation weights are adopted to guide the feature selection in a self-attention architecture, using them as a control gate to emphasize the spatial responses of target instances. Considering that the original similarity weights (calculated by the self-attention algorithm) do not distinctly model the orientation variation, the considered orientation weights provide an efficient asset to emphasize the orientation of objects. Extensive experiments on the DOTA and HRSC2016 datasets demonstrate that our method achieves state-of-the-art performance among single-scale methods, while achieving competitive performance over multi-scale methods.
Shangdong Zheng, Zebin Wu 0001, Yang Xu 0006, Zhihui Wei, Antonio Plaza
IEEE Trans. Geosci. Remote. Sens.1
2019 A Stackable Attention-Guided Multi-scale CNN for Number Plate Detection
Shangdong Zheng, Yang Xu 0006, Tianming Zhan, Zhihui Wei, Zebin Wu 0001
ICIG (1)2