Dong Yang 0012

dblp:33/412-12 · DBLP profile ↗
← Back
21ranked-venue papers
3as first author
10since 2021 · last 2025
0000-0003-2861-5421ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 15 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 since 2021Artificial intelligence and machine learning · 1
YearPublicationVenuePosition
2025 Elaborate Information Refinement Network for Fine-Grained Object Detection in Remote Sensing Images
abstract
0pt Fine-grained object detection aims to localize and classify subcategories of objects by extracting more discriminative semantic features, which is particularly challenging for remote sensing images due to their complex backgrounds and arbitrarily oriented objects. Accurate localization with bounding box regression typically relies on detailed texture and edge information to delineate object boundaries, while fine-grained classification requires more elaborate semantic information. However, existing methods often share the same input features across the model, resulting in a mismatch between the requirements of the localization task and those of the fine-grained classification task. To address this problem, we propose a Elaborate Information Refinement Network (EIRNet), which not only effectively separates features for localization and fine-grained classification but also refines these features according to the specific requirements of each task. For fine-grained classification, we propose a Fine-grained Context Fusion Module (FCFM) to enhance the ability to extract discriminative features by expanding the receptive field. For localization, we introduce an Edge Information Sensing Module (EISM) to extract scale-invariant features by combining high-dimensional information with detailed edge information, thereby improving the network’s ability to accurately locate objects. Additionally, to extract richer fine-grained semantic information, we present a Feature Injection Module (FIM), which collects and fuses information across different levels using a unified structure. This module then distributes the refined features to appropriate levels, effectively reducing inherent information loss and enhancing the local information fusion capability of the network. Through extensive experiments, the proposed method demonstrates a mean average precision (mAP) of 80.23% on ShipRSImageNet and 50.52% on FAIR1M-v1.0, marking improvements of 4.10% and 2.54% over the current state-of-the-art approaches, respectively.
Xi Yang 0011, Dong Yang 0012
IEEE Trans. Geosci. Remote. Sens.3
2025 SCIR: A Weakly Supervised Contextual Instance Refinement Method for Remote Sensing Object Detection
abstract
Most weakly supervised object detection (WSOD) methods currently prioritize the top-scoring object instance from proposals to train the corresponding object detector. The detector tends to focus on the entire object by analyzing the contextual information around the most discriminative activation regions. However, the traditional selective search algorithm yields top-scoring instance proposals that cover only a portion of the object, thereby diminishing the detector’s sensitivity to objects with a wide range of scale variations in remote sensing images (RSIs), particularly tiny object clusters. To address this issue, this paper proposes a novel WSOD method called SAM-guided proposal generation with Contextual Instance Refinement (SCIR) method for detecting objects with significant scale variations in RSIs. Specifically, a SAM-guided proposal generator (SPG) module is designed to generate high-quality proposals based on SAM masks rather than the traditional top-scoring one. Our SPG combines WSOD’s advantage of mining classification clues through inexact supervision with SAM’s capability of pre-learned world knowledge to provide automatic prompts for SAM. Meanwhile, we propose a contextual enhanced feature extractor (CEFE) module to capture the global context of the visual scene, which further activates the feature representation of the entire object. Finally, the feature map from CEFE and proposals generated by SPG are fed into a context-perceived instance refinement (CPIR) module. Our CPIR aims to shift the attention of the detection network from the local feature portion to the entire object by integrating local and global contextual information. Extensive experiments on the challenging DOTA and DIOR datasets demonstrate that our proposed SCIR achieves state-of-the-art performance and is quite effective on multi-scale object issues.
Xi Yang 0011, Zhongyuan Zhou, Songsong Duan, Dong Yang 0012
IEEE Trans. Geosci. Remote. Sens.4
2025 SFDN: A Novel Semantic Feature Decouple Network for Fine-Grained Remote Sensing Object Detection
abstract
Fine-grained object detection (FGOD) aims to identify subcategories of objects by extracting more discriminative semantic information. Bounding box regression typically requires detailed texture and edge information to accurately delineate object boundaries, while classification requires richer semantic information. However, existing methods use the same input features in the model, resulting in an imbalance between the localization task and the fine-grained classification task. To address this issue, we propose a novel Semantic Feature Decouple Network (SFDN) that effectively separates semantic information for localization and fine-grained classification. For the localization task, we propose a Regression Feature Fusion Module (RFFM) to extract feature maps with more edge information. To enhance classification one, we propose a Fine-grained Feature Diversification Module (FFDM) to capture discriminative semantic information from feature maps by introducing fake attention maps. Aiming to extract richer fine-grained semantic information, we propose an Adaptive Local Perception Module (ALPM) to deeply extract multiscale semantic feature information by using dilated convolution at varying dilation rates. Extensive experiments demonstrate that the proposed network respectively achieves the mean average precision of 80.08% and 49.44% on the ShipRSImageNet and FAIR1M-v1.0 datasets, outperforming SOTA methods by 3.95% and 1.46%.
Xi Yang 0011, Zhongyuan Zhou, Dong Yang 0012
IEEE Trans. Geosci. Remote. Sens.3
2024 Domain-Aware Generalized Meta-Learning for Space Target Recognition
abstract
As the exploration and utilization of outer space persist, the proliferation of space targets has significantly increased, underscoring the growing importance of space situational awareness. However, space target images encounter numerous challenges, including overexposure, excessive shadowing, star noise, and motion blur, distinct from natural images. While existing models can address specific issues in space target recognition images, their ability for generalizing to unseen data remains relatively weak. Furthermore, the uniform background and minimal interclass differences in space target images impose significant constraints on recognition accuracy. To tackle these challenges, we propose a domain-aware generalized meta-learning for space target recognition. In the meta-training phase, we introduce a distillation module to generalize the prior knowledge of auxiliary domains. This module distills features and predictions from auxiliary domains, providing prior information to develop a model capable of generalization across diverse domains. In the meta-testing phase, the frozen generalized embedding function is connected with a feature bias module to mitigate domain bias issues. Building on the advanced awareness of the space target domain, which is marked by substantial intraclass variations and minimal interclass variations, we introduce a feature refinement module. This module resolves fine-grained issues by reconstructing features and augmenting the proto loss to narrow the intraclass data distance. In practice, our method is evaluated under out-of-distribution settings on the BUAA-SID-share1.0 dataset, achieving an impressive accuracy of 96.0%, surpassing existing space target recognition algorithms.
Xi Yang 0011, Dechen Kong, Dong Yang 0012
IEEE Trans. Geosci. Remote. Sens.3
2024 Adaptive Mid-Level Feature Attention Learning for Fine-Grained Ship Classification in Optical Remote Sensing Images
abstract
Ship classification in optical remote sensing images is a critical task for various maritime applications, including anti-smuggling, maritime traffic control, and maritime rescue. However, fine-grained ship classification (FGSC) is challenging due to the complex background, intraclass similarity, and interclass difference. In this article, we propose a novel mid-level feature attention learning method for FGSC. Our method incorporates mid-level feature casual attention (MFCA) and mid-level channel attention (MCA) to identify discriminative regions and local features corresponding to subtle visual features. The MFCA constrains the learning process of mid-level features through comparison with attention maps and counterfactual attention maps, while the MCA uses a discriminative component to extract discriminative features from channel information and a diversity component to focus feature channels on more obvious feature regions. Besides, an adaptive weight is added to dynamically adjust the influence of MFCA and MCA in the model. Our method can be trained end-to-end and requires no annotations other than category information. Extensive experiments on two large-scale FGSC datasets, FGSC-23 and FGSCR-42, demonstrate that the proposed method achieves state-of-the-art performance, outperforming existing methods by a significant margin.
Xi Yang 0011, Zilong Zeng, Dong Yang 0012
IEEE Trans. Geosci. Remote. Sens.3
2024 SCSP: An Unsupervised Image-to-Image Translation Network Based on Semantic Cooperative Shape Perception
abstract
This paper introduces a novel approach to unsupervised image-to-image translation, aiming to overcome the limitations of existing methods in accurately capturing the shape of the source domain and the style of the target domain. The proposed method, called Semantic Cooperative Shape Perception (SCSP), focuses on enhancing the quality of generated images by addressing two key aspects. Firstly, the SCSP model employs a fusion generator that divides the mapping process into a unique texture part and a shared semantic part. By using different network structures and constraints, each part learns specific information. The unique texture generator emphasizes the style and texture details of the target domain, while the shared semantic generator focuses on the semantic information present in the source domain. This separation enables the sub-generators to extract and restore different aspects of the target domain more effectively. Secondly, a shape perception loss is introduced to improve the similarity of semantic images. It enhances the shared semantic generator's ability to perceive semantic information related to the same object by imposing constraints on the semantic graph of both the generated and input images. Therefore, the proposed method ensures semantic consistency during the translation process, leading to improved authenticity and image quality. Experimental results on four datasets, including horse2zebra, tiger2leopard, summer2winter, and photo2vangogh, demonstrate that the SCSP model achieves state-of-the-art visualization results and favorable evaluation metrics.
Xi Yang 0011, Dong Yang 0012
IEEE Trans. Multim.4
2024 Improving Cross-Modal Constraints: Text Attribute Person Search With Graph Attention Networks
abstract
Nowadays, video surveillance systems are widely deployed in public areas. However, in the unreachable corner of surveillance cameras, it still seems impossible to find the suspects only depending on eyewitness memory. Therefore, the technology that can detect particular pedestrians only by text-based attributes, or text-attribute person search, attracts lots of attention from academia. Most existing text-attribute person search methods focus on learning better feature representations by designing better network structures or using local information but lack direct constraints between modalities. This paper proposes a feature embedding motivated and graph attention network-based model, optimizing the feature extraction process by its attention mechanism. Meanwhile, this paper studies the effectiveness of the attention mechanism in feature alignment, and thus redesigns the cross-attention module, simplifying the complexity of the model and constraining the inter-modality gap in maximum by the self-attention mechanism of the graph attention network. In this way, the method simultaneously offsets the influence of modal-specific features and optimizes the number of parameters. Thus, the method improves performance and reduces time costs. Meanwhile, according to the inherent feature of attributes, this article introduces a novel embedding space, which effectively enhances the discrimination ability of the model. Extensive experiments illustrate the superiority of our model in two widely used text-attribute person search benchmarks among the state-of-the-art methods.
Xi Yang 0011, Dong Yang 0012
IEEE Trans. Multim.3
2022 An Efficient and Lightweight CNN Model With Soft Quantification for Ship Detection in SAR Images
abstract
Convolutional neural networks have been widely used for synthetic aperture radar (SAR) target detection. Typical methods based on convolutional neural network have obtained favorable detection accuracy at the cost of high model complexity, and thus are difficult to be directly applied to real-time satellites on board as well as maritime rescue. To deal with this problem, this paper proposes an efficient and lightweight target detection network incorporating soft quantization. Firstly, to compensate for the lack of accuracy caused by lightweight networks, a feature fusion module called split bidirectional feature pyramid network is proposed to alleviate the interference of complex background on SAR images. Meanwhile, to adapt the lightweight network and the feature fusion module, a linear transformation module is presented to enhance the linear representation of the model via learnable parameters. Eventually, to make the model size smaller, a soft quantization algorithm is proposed to reduce the accuracy degradation caused by quantization errors. We validate the robustness of the model in several publicly available datasets. Experimental results show that our model achieves 97.0% detection accuracy on SAR ship detection dataset, with a 0.9% accuracy improvement compared to mainstream methods using less than 15x the number of parameters and less than 6x the number of flops.
Xi Yang 0011, Chengzeng Chen, Dong Yang 0012
IEEE Trans. Geosci. Remote. Sens.4
2022 FG-GAN: A Fine-Grained Generative Adversarial Network for Unsupervised SAR-to-Optical Image Translation
abstract
Synthetic aperture radar (SAR) and optical sensing are two important means of Earth observation. SAR can be used for all-day and all-weather Earth observation, but it has the disadvantages of speckle noise and geometric distortion, which are not conducive to human eye recognition. Optical image conforms to the characteristics of human visual observation, but it is easily affected by climate and time. Therefore, to integrate the advantages of the two, researchers have carried out extensive work on SAR-to-optical (S2O) image translation. Most of the existing methods for S2O image translation are supervised and need paired training samples, limiting its large-scale application in remote sensing field. Thus, we give priority to an unsupervised S2O image translation method. Meanwhile, we find that the images generated by unsupervised methods suffer from significant detail deficiencies. To solve this problem, we propose a fine-grained generative adversarial network (FG-GAN) introducing three strategies to enhance the detailed information in generated optical images. First, we design an unbalanced generator (UBG) with complex encoder networks and relatively simple decoder networks. The complex encoder extracts abundant feature information, while the decoder obtains key details by filtering these features. Second, to match the learning ability of the generator, we present a multiscale discriminator (MSD) to enhance the discriminant ability of the network. Third, we propose a comprehensive normalization group (CNG) to promote the physical representation consistency of SAR and optical images. Extensive experiments have been conducted, and the results show that our method is superior to the state-of-the-art (SOTA) methods on both subjective and objective evaluation indicators. Moreover, our FG-GAN has a significant improvement on classification accuracy, indicating its potential in facilitating the performance of practical remote sensing tasks.
Xi Yang 0011, Dong Yang 0012
IEEE Trans. Geosci. Remote. Sens.4
2022 A Universal Ship Detection Method With Domain-Invariant Representations
abstract
Although ship detection methods based on deep learning have achieved remarkable progress, the design of the universal ship detection (USD) system is rarely studied. The main challenge of USD lies in the notorious domain bias and shift problem across multiple domains. This article implements USD based on domain-invariant representations to alleviate this issue. Specifically, the proposed method integrates a multilevel domain classification network (MDCN) and a domain-centric cut-paste module (DCM). First, the backbone network is facilitated to learn domain-independent image features from multiple domains through MDCN, thereby reducing the disturbance of domain-specific features to universal detector. Furthermore, the proposed method combines the domain-related synthetic samples generated by DCM to provide MDCN with strong supervision information, which further motivates the network to be more attentive to the domain-invariant representations at the instance level. Finally, we conduct experiments on multiple ship datasets in the synthetic aperture radar (SAR) and optical domains to verify the effectiveness of the method. The results show that the proposed method outperforms baseline by around 2.95% average precision (AP50), which achieves an effective USD system by complementing the information between domain-invariant representations.
Xin Zhang 0129, Xi Yang 0011, Dong Yang 0012, Xinbo Gao 0001
IEEE Trans. Geosci. Remote. Sens.3
2019 T-SCNN: A Two-Stage Convolutional Neural Network for Space Target Recognition
abstract
Space target recognition plays an important role in the field of space security and exploration. With the rapid development of artificial intelligence technique and explosive increase of image dataset, object recognition based on deep learning has achieved favorable performance. However, the recognition of deep space targets in visible spectrum images still remains in the traditional manual interpretation approach, thus leading to low efficiency and inevitable subjective errors. In this paper, we propose an artificial intelligence method for space target recognition, called Two-Stage Convolutional Neural Network (T-SCNN). Our T-SCNN is composed of two stages, i.e., target locating and target recognition. In the stage of target locating, we first detect all suspected targets from the total image dataset by presenting a minimum bounding rectangle with threshold (MBRT) approach, then cut out all regions encompassing targets to generate target images for training. In the stage of target recognition, we send target images to the well-trained recognition network for identification. Additionally, data augmentation is conducted in the CNN training to satisfy its data quantity requirement. Extensive experiments are performed on our synthetic space target image dataset, and the result demonstrate that the proposed method achieves high accuracy within a short time.
Tan Wu, Xi Yang 0011, Bin Song 0001, Nannan Wang 0001, Xinbo Gao 0001, Liyang Kuang, Xiaoting Nan, Dong Yang 0012
IGARSS9
2019 CNN with spatio-temporal information for fast suspicious object detection and recognition in THz security images
Xi Yang 0011, Tan Wu, Lei Zhang 0019, Dong Yang 0012, Nannan Wang 0001, Bin Song 0001, Xinbo Gao 0001
Signal Process.4
2018 Saliency Deep Embedding for Aurora Image Search
abstract
Deep neural networks have achieved remarkable success in the field of image search. However, the state-of-the-art algorithms are trained and tested for natural images captured with ordinary cameras. In this paper, we aim to explore a new search method for images captured with circular fisheye lens, especially the aurora images. To reduce the interference from uninformative regions and focus on the most interested regions, we propose a saliency proposal network (SPN) to replace the region proposal network (RPN) in the recent Mask R-CNN. In our SPN, the centers of the anchors are not distributed in a rectangular meshing manner, but exhibit spherical distortion. Additionally, the directions of the anchors are along the deformation lines perpendicular to the magnetic meridian, which perfectly accords with the imaging principle of circular fisheye lens. Extensive experiments are performed on the big aurora data, demonstrating the superiority of our method in both search accuracy and efficiency.
Xi Yang 0011, Xinbo Gao 0001, Bin Song 0001, Nannan Wang 0001, Dong Yang 0012
ICME5
2018 Aurora image search with contextual CNN feature
Xi Yang 0011, Xinbo Gao 0001, Bin Song 0001, Dong Yang 0012
Neurocomputing4
2018 GMTI and Parameter Estimation for MIMO SAR System via Fast Interferometry RPCA Method
abstract
Multiple-input multiple-output synthetic aperture radar (MIMO SAR) system has drawn considerable attention because of its extra degrees of freedom for high resolution and wide swath compared with the traditional multichannel SAR system. But how to extract the matched signal without the unmatched interferences is the foremost task for MIMO SAR system. In this paper, by using the orthogonal frequency division multiplexing chirp signals as the transmitted signals, it is demonstrated that the robust principal component analysis (RPCA) method can be successfully employed for ground moving target indication (GMTI) with no need for separating the matched signal and unmatched interferences. It is because the unmatched interference is proven to have low-rank property and noise-level magnitude, which can be separated apart from the matched signal with the RPCA method. However, the traditional RPCA methods may be restricted by the high computational burden due to the complex decompositions and multiple iterations. Hence, a fast interferometry RPCA method is proposed specially for GMTI mode, which takes full advantage of the characteristics of along-track interferometry SAR system. It can improve the probability of detection under low signal-to-clutter-and-noise ratio. Additionally, it will dramatically shorten the computational time. Furthermore, the proposed method can also estimate the radial velocities of the moving targets simultaneously. The results by applying the proposed method into a set of real SAR data are consistent with the analysis presented in this paper.
Yan Huang 0018, Guisheng Liao, Jingwei Xu 0002, Jie Li 0027, Dong Yang 0012
IEEE Trans. Geosci. Remote. Sens.5
2016 Ground moving target detection in MIMO-SAR system
abstract
Multiple-input multiple-output synthetic aperture radar (MIMO-SAR) has drawn widely attention for its increased degrees of freedom. This enhanced architecture offers not only the opportunity to map wider images swaths with improved spatial resolution, but also enables novel SAR modes to resolve some of the contradicting user requirements. In this paper, the performance of ground moving target indication (GMTI) in MIMO-SAR system is analyzed. It becomes evident that the virtual channels can be utilized to fulfill both the image swath and the required signal-to-noise-radio. An analysis on orthogonal frequency division multiplexing (OFDM) chirp signal designing is proposed and a novel GMTI method based on low-rank property is also proposed. Simulation and experimental results show its good performance, which implies that MIMO-SAR would be a good choice for future multichannel systems.
Dong Yang 0012, Xi Yang 0011, Xiaomin Tan, Hongxing Dang
IGARSS1
2015 Strong Clutter Suppression via RPCA in Multichannel SAR/GMTI System
abstract
Clutter suppression and ground moving target indication are challenging tasks in multichannel synthetic aperture radar (SAR) systems. In recent years, robust principal component analysis (RPCA) has attracted much attention for its good performance in distinguishing the different parts from a set of correlative database. Therefore, we propose a fast RPCA-based detection method for multichannel SAR under a strong clutter background in this letter even with channel unbalance or platform motion error. Subsequently, as the existing space-time adaptive processing (STAP) method would fail when the training samples are contaminated by the moving target, we apply the RPCA-based method in the range-Doppler domain to improve the performance of STAP. Since the regions of targets can be detected via RPCA, the remaining samples, which can be regarded as only clutter, are used to estimate the covariance matrix for further processing. The final experiments based on real measured data set show its good performance under the strong clutter background. Although the RPCA-based result differs from that of the STAP method, they can work cooperatively to get a more robust detection performance.
Dong Yang 0012, Xi Yang 0011, Guisheng Liao, Shengqi Zhu 0001
IEEE Geosci. Remote. Sens. Lett.1
2015 Efficient Compressed Sensing Method for Moving-Target Imaging by Exploiting the Geometry Information of the Defocused Results
abstract
Compressed sensing (CS) has been increasingly used in the synthetic aperture radar ground moving-target indication system, particularly for imaging the moving targets, which satisfy the sparse precondition of the CS method. However, efficient moving-target imaging is a key challenge for current CS methods, since the redundant basis brings heavy computation load. In this letter, by exploiting the geometry information of the defocused results, we present an efficient fractional Fourier transform (FRFT) to estimate the Doppler rate and image the moving targets by only two times FRFT rather than time-consuming searching operation. Then, the concept is extended into an efficient CS (ECS) imaging method by two bases consisting of two discrete FRFT matrices rather than the redundant basis. Simulations and real-data process are provided to demonstrate the effectiveness of the ECS method. The proposed ECS method can achieve accurate parameter estimation and imaging performance with low computational complexity.
Xuepan Zhang, Guisheng Liao, Shengqi Zhu 0001, Dong Yang 0012, Wentao Du
IEEE Geosci. Remote. Sens. Lett.4
2014 SAR Imaging With Undersampled Data via Matrix Completion
abstract
High-resolution synthetic aperture radar (SAR) imagery of a wide area of surveillance is a difficult large-data problem. In the past few years, researchers have applied compressive sensing (CS) to SAR, as it exploits redundancy in signals. To further extend the sparse problem from the vector to the matrix, a new theory called matrix completion (MC) has attracted much attention, which can complete a matrix from a small set of corrupted entries based on the assumption that the matrix is essentially of low rank. Inspired by this technique, a novel SAR imaging algorithm is proposed in this letter to deal with the undersampled data. After representing the data of a range cell as a matrix, the phase is compensated to keep the matrix holding the property of low rank. Subsequently, MC can be utilized to recover the full-aperture data in the new constructed matrix. Since the data are completely unsampled in the corresponding azimuth cells, the proposed method has effectively conquered the restriction of previous applications that each received channel must have a small number of samples. The final results in both simulation and real-data experiments show that the targets can be well focused even in the scenario of discarding a large percentage of the received pulses. Moreover, when compared with CS, the method is not required to design the complicated measurement matrix.
Dong Yang 0012, Guisheng Liao, Shengqi Zhu 0001, Xi Yang 0011, Xuepan Zhang
IEEE Geosci. Remote. Sens. Lett.1
2014 A New Method for Radar High-Speed Maneuvering Weak Target Detection and Imaging
abstract
Weak-target detection and imaging are the challenging problems of airborne or spaceborne early warning radar. The envelope of a high-speed weak target after range compression spreads over range during the long observation period. To finely refocus a high-speed weak maneuvering target, motion parameters should be accurately obtained for compensating the envelope. This letter proposes a new imaging approach for high-speed maneuvering targets without a priori knowledge of their motion parameters. In this method, the azimuth compression function is constructed in a range and azimuth 2-D frequency domain, which can eliminate the coupling effect between range and azimuth. Theoretical analysis confirms that the methodology can precisely focus targets. Simulation results show that the proposed algorithm improves the performance for detecting and imaging high-speed maneuvering targets.
Shengqi Zhu 0001, Guisheng Liao, Dong Yang 0012, Haihong Tao
IEEE Geosci. Remote. Sens. Lett.3
2013 A Persymmetric GLRT for Adaptive Detection in Compound-Gaussian Clutter With Random Texture
abstract
We focus on the problem of detecting a signal in compound-Gaussian clutter, where the texture is a random variable with Gamma or inverse Gamma distribution. The persymmetric structure of the covariance matrix is exploited and a persymmetric generalized likelihood ratio test (Per-GLRT) using a three-step procedure is proposed. In addition, we prove that the Per-GLRT ensures constant false alarm rate (CFAR) property with respect to the covariance matrix. Finally, the detector is assessed by Monte Carlo simulations. Performance comparison of the Per-GLRT with the traditional GLRT shows that the former improves the detection performance in training-limited scenarios.
Yongchan Gao, Guisheng Liao, Shengqi Zhu 0001, Dong Yang 0012
IEEE Signal Process. Lett.4