VLDB 2026 Research / reviewers in the wild / expert
Mingyang Ma 0004
dblp:151/7524-4
· DBLP profile ↗
37ranked-venue papers
10as first author
31since 2021 · last 2026
0000-0002-2944-628XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 23 · 3 first-author · 22 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 6 first-author · 7 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Keyframe selection from motion capture data with dual-agent reinforcement learningabstractAnimation production workflows centred around motion capture techniques require animators to edit motions based on a set of keyframes. However, most existing keyframe selection methods are optimisation-based, which suffer from the issues of flexibility and efficiency. In this paper, a novel deep reinforcement learning method with dual agents are proposed for unsupervised keyframe selection. First, an S-Agent and an R-Agent evaluate the actions of selection and refinement, respectively. A deep spatio-temporal network, namely graph keyframe evaluation network (GKEN), is proposed for the agents. Then, an animation specified reward is devised based on reconstruction, which fulfills three important properties of the animation workflow: incremental reward, order insensitivity and non-diminishing returns. During the inference, it is no longer necessary to compute the reconstruction, which significantly decreases the run-time latency. Experiments on the CMU MoCap dataset demonstrate the efficiency of the proposed method without clearly compromising the effectiveness compared with the state-of-the-art methods. • A deep reinforcement learning with dual-agent to identify motion keyframes. • A spatio-temporal deep agent with graph convolutions and transformers. • Comprehensive experiments and human demonstrations for MoCap keyframing. Kun Hu 0008, Clinton Mo, Mingyang Ma 0004, Shaohui Mei, Zhiyong Wang 0001 |
Pattern Recognit. | 4 |
| 2025 | MSFD: Multiscale Feature Decomposition for Cross-Modality Visible-to-Infrared Drone Image TranslationabstractIn the global landscape of the Internet of Things (IoT), drone IoT technology has gained widespread application. This technology can monitor and analyze land use and land cover more quickly and more accurately. Currently, the images collected by drone IoT technology are mostly visible images, which are highly susceptible to external environmental factors, while the acquisition of infrared images is relatively more challenging. Visible-to-infrared drone image translation seeks to convert visible drone images into their corresponding infrared counterparts. Although existing GAN-based image-to-image translation methods have demonstrated impressive results in the domain of natural images, they still face challenges in generating highly realistic infrared drone images. Therefore, a novel Multi-Scale Feature Decomposition (MSFD) method is introduced for visible-to-infrared drone image translation. The proposed approach accomplishes the translation through spectral feature disentanglement and cross-modal recombination. In our model, spectral feature disentanglement is based on the separation of modality-specific spectral information and modality-invariant shared structural content from the image representation. Subsequently, the spectral features and underlying content from different modalities can be recombined by generators to facilitate cross-modality image translation. To enhance the quality of generated images, our method integrates a multi-scale spectral feature encoder to address significant spectral discrepancies between targets and backgrounds in drone images by extracting and fusing spectral features at different scales. Additionally, the strategy of multi-scale generators and discriminators further enhances the generation quality of infrared drone images. The experimental results highlight the superior performance of our model in visible-to-infrared drone image translation. Zhiquan Liu 0001, Zonghao Han, Mingyang Ma 0004, Jian Zhao 0002 |
IEEE Internet Things J. | 4 |
| 2025 | Hyperspectral Tracker With Constrained Object Adaptive Learning and Trajectory ConstructionabstractHyperspectral imaging offers significant potential for precise object tracking, yet the scarcity of dataset volumes specifically tailored for hyperspectral tracking algorithms hinders progress, particularly for deep models with complex structures. Additionally, current deep learning-based hyperspectral trackers typically enhance model accuracy via online or adversarial learning, adversely affecting tracking speed. To address these challenges, this paper introduces the Constrained Object Adaptive Learning hyperspectral Tracker (COALT), an effective parameter-efficient fine-tuning tracker tailored for hyperspectral tracking. COALT integrates Pixel-level Object Constrained Spectral Prompt (POCSP) and Temporal Sequence Trajectory Prompt (TSTP) through Adaptive Learning with Parameter-efficient Fine-tuning (ALPEFT), enabling a transformer-based tracker to capture detailed spectral features and relationships in hyperspectral image sequences through trainable rank decomposition matrices. Specifically, POCSP is designed to retain optimal spectral information with low internal correlation and high object representativeness, enabling rapid image reconstruction. Then, the most representative spectral template and search are fused into a single stream as spectral prompts for the Encoder and Decoder layers. Concurrently, the previous coordinates within the same sequence are tokenized and utilized as temporal prompts by TSTP in the decoder layers. The model is trained with ALPEFT to optimize spectral information learning, which substantially reduces the number of training parameters, alleviating overfitting issues arising from limited data. Meanwhile, the proposed tracker not only retains the ability of pre-trained model to estimate object trajectories in an autoregressive manner but also effectively utilizes spectral information and enhances target location perception during the fine-tuning process. Extensive experiments and evaluations are conducted on two public hyperspectral tracking datasets. The results demonstrate that the proposed COALT tracker achieves satisfactory performance with leading processing speed. The code will be available at https://github.com/PING-CHUANG/COALT. Ye Wang 0020, Mingyang Ma 0004, Ge Zhang 0006, Tao Gao 0001, Shaohui Mei |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Hyperspectral Object Tracking With Context-Aware Learning and Category ConsistencyabstractHyperspectral imaging technology is of crucial importance to improve the performance of object tracking in many remote sensing surveillance areas. Previous methods primarily focused on feature fusion strategies by employing additional enhancement modules. However, these methods commonly lack contextual understanding to distinguish the target from the background and totally ignore the category information of the targets. To address these limitations, a novel hyperspectral object tracker is proposed to incorporate context-aware learning and category consistency tracker (CCTrack), which can adaptively learn context-aware representations in hyperspectral scenarios to obtain global target information with memory storage, while constructing an interframe category consistency constraint to enhance tracking process. Specifically, CCTrack integrates an adaptive context-aware learning (ACL) mechanism, which includes a feature decoupling module (FDM) to extract specific representations from decoupled features, and a Mamba layer to retain and update long-range dependencies. To align with prior knowledge of target recognition and motion patterns, an alignment transformation module (ATM) is employed with the ACL mechanism, fully leveraging spatial-spectral representations. In addition, category consistency constraint modules (C3Ms) are introduced to enforce category consistency across frames by computing the similarities between the target features and the corresponding category name, serving as the constraint to improve tracking performance. Extensive experiments over the hyperspectral object tracking (HOT) benchmark covering various remote sensing scenarios demonstrate that CCTrack outperforms state-of-the-art methods by a significant margin. Ye Wang 0020, Shaohui Mei, Mingyang Ma 0004, Tao Gao 0001, Huiyang Han |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | SAFF-DETR: An End-to-End Object Detection Network for Remote Sensing Images With Targets of Varying Sizes Based on Scale Adaptation and Frequency FusionabstractDeep learning-based object detection algorithms have achieved significant success in the field of computer vision. However, the wide range of target sizes in remote sensing images poses a challenge for single algorithms to detect objects of varying sizes effectively. To address this issue, this paper proposes an end-to-end object detection algorithm for remote sensing images based on Scale Adaptive and Frequency Fusion DETR (SAFF-DETR), which designs a frequency feature enhancement and fusion mechanism to handle targets of varying sizes within a single framework. First, in order to improve the Transformer-based detectors’ ability to small targets, a multibranch representation fusion (MRF) module is proposed to fuse shallow layer frequency representations, boosting the network’s ability to perceive small targets. Furthermore, Cross-layer Spatial and Channel Frequency Attention (CSFA and CCFA) is designed to enable efficient frequency feature interaction across multi-scale features, enhancing the representation capability for targets of different sizes. Moreover, by integrating the two aforementioned attention mechanisms, the Cross-layer Channel-Spatial-wise Frequency Fusion (CCSFF) structure is introduced to realize global feature interactions in one step without repetitive up and down-sampling operations, by which patch division-based Transformer architecture is designed to enhance scale adaptability for object detection. Experimental results over several benchmark datasets demonstrate that the proposed SAFF-DETR can handle extremely varying-size targets and outperforms several SOTA algorithms. The source code will be available at https://github.com/Jianglin-Zhao/SAFF-DETR. Yuanjie Zhi, Jianglin Zhao, Mingyang Ma 0004, Shaohui Mei |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2025 | RIFormer+: Rethinking Rotation-Invariant Feature Learning in Transformer
Yifan Zhang 0006, Mingyang Ma 0004, Shaohui Mei |
IEEE Trans. Multim. | 3 |
| 2025 | HTACPE: A Hybrid Transformer With Adaptive Content and Position Embedding for Sample Learning Efficiency of Hyperspectral TrackerabstractTransformer architecture has demonstrated significant potential in hyperspectral object tracking by leveraging global correlation learning to accurately represent the data distribution. However, existing hyperspectral object trackers based on transformer models typically rely on costly pre-trained models, making them prone to crashing due to overfitting when tuned on small-scale hyperspectral videos, greatly limiting their performance. To address this challenge, in this paper, a Hybrid Transformer with Adaptive Content and Position Embedding (HTACPE) tracker is proposed to improve the learning efficiency of the tracking model, and fully explore the spectral-spatial information. Specifically, an Adaptive Content and Position Embedding Module (ACPEM) is designed to dynamically learn the balance between focusing on positional and content-based information, which allows the model to effectively handle datasets of various sizes. To enhance the spectral-spatial information, a Spectral Grouping Module (SGM) is designed to learn the highfrequency information in complex scenarios, thereby enhancing diversified features. It operates in parallel with the ACPEM feature learning module. Furthermore, a Dynamic Reliability Refinement Module (DRRM) is incorporated to address challenges related to accurate object position perception, iteratively refining prediction parameters to enhance the reliability of the model. Extensive experiments demonstrate that the proposed HTACPE achieves satisfactory tracking performance both qualitatively and quantitatively, especially with insufficient training data. Ye Wang 0020, Shaohui Mei, Mingyang Ma 0004, Yuru Su |
IEEE Trans. Multim. | 3 |
| 2024 | Separable Deep Graph Convolutional Network Integrated With CNN and Prototype Learning for Hyperspectral Image ClassificationabstractGraph convolutional networks (GCNs) have garnered extensive attention in the realm of hyperspectral image (HSI) classification. However, due to the problem of over-smoothing caused by deep GCN, most of the existing GCN-based methods are limited to constructing shallow networks, thus only able to extract superficial features. Moreover, when existing shallow GCNs extend to a more deeper structure, the number of learnable parameters increase linearly, thus leading to poor generalization performance under limited training samples. To address the aforementioned issues, a Separable Deep Graph Convolutional Network Integrated with CNN and Prototype Learning (SDGCP) is proposed for HSI classification, which can extract effective global structural information of HSI without increasing the number of trainable parameters. Specifically, the spectral and spatial features, adaptively selected by the attention module, are encoded into the structure of a graph by the graph encoder with the assistance of the pixel-to-region mapping obtained from the simple linear iterative clustering (SLIC). Then, a separable deep graph convolution module, composed of feature extraction and deep feature propagation, is adopted to capture the long-range contextual relationships from HSI encoded as graph data, which is combined with locally complementary information extracted by CNN after decoding. Finally, to further boost the performance of classification under limited labeled samples, prototype learning with regularization terms is utilized to enhance the intra-class compactness and inter-class separability of feature representations. Extensive experiments on three standard HSI data sets demonstrate the superiority of the proposed SDGCP over the state-of-the-art (SOTA) methods. Yingjie Lu, Shaohui Mei, Fulin Xu, Mingyang Ma 0004 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | DGT: Deformable Graph Transformer for Hyperspectral Image ClassificationabstractTransformers can model global context to enhance the performance of hyperspectral classification. However, the explored global information is generally confined to the spatial neighborhood of target pixels. In order to fully leverage global correlation across broader areas, a deformable graph transformer (DGT) is proposed for hyperspectral classification, in which the global information within an entire image is explored to improve the classification performance. Specifically, DGT layers are designed to adaptively sample virtual nodes at varying distances from an initial graph constructed from an image, by which the global spatial information can be explored using a deformable graph self-attention (DGSA) mechanism. Moreover, a learnable absolute position encoding (LAPE) module is constructed to enhance the spatial context awareness of DGT by integrating positional information into the graph nodes. In addition, graph structure encoding and graph topology encoding are further designed as inductive biases for the graph, by which both local structural information and global topological information of the HSI are captured to enhance the feature extraction capability of the DGT layer. Ultimately, through the stacking of multiple DGT layers, a composite feature fusion learning (CFFL) module is employed to fully utilize the simple low-level and complex abstract high-level features extracted from different layers. Extensive experiments on four datasets demonstrate the superiority and robustness of the proposed DGT over several state-of-the-art (SOTA) methods in terms of various evaluation criteria. Yingjie Lu, Shaohui Mei, Fulin Xu, Mingyang Ma 0004 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Joint Spatial and Spectral Graph-Based Consistent Self-Representation for Unsupervised Hyperspectral Band SelectionabstractBand selection (BS), which effectively reduces spectral dimensionality, stands out as a leading focus within hyperspectral image (HSI) analysis. Self-representation (SR) has surfaced as a favored technique in this domain due to its applicability to BS and unsupervised nature. However, the existing SR-based BS approaches only leverage either spatial or spectral relationships, with few integrating both while concentrating on the representation level rather than the selection level. In addition, employing all spatial pixels for spatial relationship utilization leads to considerable computational complexity. Therefore, this article proposes joint spatial and spectral graph-based consistent SR (JSSGCSR) to more effectively exploit spatial and spectral relationships for BS, which separately conducts SR to handle each view of spatial and spectral graphs to better consider two different structure characteristics, and ultimately integrates two SR results to achieve a unified and robust representative band set by imposing consistent sparsity pattern on their joint representation coefficients. In addition, the spatial and spectral relationships are integrated into different data spaces, that is, spectral graph SR and spatial graph SR are, respectively, conducted in the original HSI and the segmented and pooled HSI, which not only reduces the influence of superpixel segmentation on spectral relationships, but also improves the efficiency of spatial relationship utilization. Experimental results on three benchmark datasets have demonstrated the effectiveness of the proposed JSSGCSR in HSI classification tasks. Mingyang Ma 0004, Fan Li 0003, Zhiyong Wang 0001, Shaohui Mei |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | A Novel Center-Boundary Metric Loss to Learn Discriminative Features for Hyperspectral Image ClassificationabstractLearning discriminative features is of crucial for hyperspectral image (HSI) classification. Though metric learning has been applied to learn effective features in HSI classification tasks, existing metric loss functions only consider distance among features of sample pairs but ignore the feature centers and boundaries in the embedding feature space, which limits the discrimination of learned features. In this paper, a novel metric loss function named center-boundary metric loss (CBML) is proposed to learn more discriminative features so as to improve HSI classification performance. Unlike the existing metric loss functions, CBML not only considers the distance between sample pairs to enhance intra-class similarity and inter-class separability but also pays more attention to the feature centers and boundaries in the embedding feature space that could greatly determine and affect the category of features. Specifically, CBML forces the distance of a sample to its corresponding feature center to be explicitly smaller than that to samples from other classes by a predefined threshold. As a result, the boundaries of different classes will separate an actual distance, which improves the discrimination of learned features. Moreover, in order to improve the training efficiency, a cross mini-batch sampling strategy is further proposed to break through the limitation within the mini-batch by using features between several contiguous mini-batches to sample pairs without increasing the size of the mini-batch. Accordingly, the sampling range of sample pairs is greatly expanded, and the training data is more fully exploited. Experimental results over four benchmark datasets with a typical network for HSI classification demonstrate our proposed method outperforms several state-of-the-arts. Shaohui Mei, Zonghao Han, Mingyang Ma 0004, Fulin Xu |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Hyperspectral Image Reconstruction From RGB Input Through Highlighting Intrinsic PropertiesabstractDozens of spectral bands of hyperspectral images (HSIs) have been successfully reconstructed from only three color band images using deep neural networks according to their powerful nonlinear mapping capability. However, the existing deep-learning-based approaches tend to directly reconstruct HSIs from RGB inputs without emphasizing the discriminative intrinsic properties of different materials, resulting in certain distortion in reconstructed spectra. In this article, an intrinsic image decomposition (IID)-based spectral super-resolution (SSR) framework is proposed to reconstruct spectra of pixels from their reflectance feature and shading feature separately, by which the intrinsic properties can be emphasized during spectral reconstruction. Specifically, a dual hierarchical regression network (DHRNet) is designed for the proposed IID-based SSR task, in which a shading feature extraction module (SFEM) based on dense structure and a reflectance feature extraction module (RFEM) with attention mechanism are first, respectively, designed to reconstruct spectral information from reflectance feature and shading feature, and a feature enhancement module (FEM) is consequently devised to further improve the coarse combined estimation. Ultimately, a novel hybrid loss combining smooth$\boldsymbol {l}_{1}$loss, spectral angel mapper (SAM), and gradient prior is also presented to restrain the spectral distortion while enhancing the sharpness of the reconstructed HSI. Experimental results over three datasets demonstrate the superiority of our proposed framework. Nan Wang 0026, Shaohui Mei, Yifan Zhang 0006, Mingyang Ma 0004, Xiangqing Zhang |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2024 | Adaptive Composite Feature Generation for Object Detection in Remote Sensing ImagesabstractObject detection in remote sensing images identifies and extracts the acquired Earth surface information, providing data support and research basis for multiple fields. Remote sensing image object detection based on knowledge distillation (KD) can transfer the knowledge of a large teacher model to a smaller student model, achieving the effect of low parameter volume and high accuracy. Mainstream methods directly imitate teacher features to improve student performance, ignoring the generation of high-ranking features through teacher features instructing student feature maps in this knowledge transfer process. In this article, an adaptive composite feature generation (ACFG) strategy is proposed to achieve end-to-end trainable KD for object detection in remote sensing images, in which the robustness of feature points under composite masks is improved through adaptive feature mapping. In particular, a composite mask generator (CMG) module is proposed to select student instance-related features and point background features. Furthermore, a global and local projection layer (GLPL) module is proposed to connect the local information and global information of the feature map under the mask generator to adaptively realize the global recovery mapping of the feature map with partial feature points. Finally, balanced decoupling loss (BDL) is improved to handle foreground and background loss separately, so that the two decoupled features can better enable the student model to learn instance-related information. Note that the proposed ACFG is capable of conducting KD for both single-stage and two-stage object detectors. Experimental results using both anchor-based and anchor-free detectors on the DIOR dataset and DOTA dataset demonstrate that the proposed ACFG clearly achieved better performance than several state-of-the-art (SOTA) algorithms for KD. Ziye Zhang 0007, Shaohui Mei, Mingyang Ma 0004, Zonghao Han |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Contextual Adversarial Attack Against Aerial Detection in The Physical WorldabstractDeep Neural Networks (DNNs) have been extensively utilized in aerial detection. However, DNNs are susceptible and vulnerable to adversarial examples Recently, physical attacks have gradually garnered attention due to their effectiveness and practicality, which pose great threats to some security-critical applications. In this paper, we take the first attempt to perform physical attacks in contextual form against aerial detection in the physical world. We propose an innovative contextual attack method against aerial detection in real scenarios, which achieves powerful attack performance and transfers well between various aerial object detectors without smearing or blocking the interested objects. Based on the findings that the targets’ contextual information plays an important role in aerial detection by observing the detectors’ attention maps, we fully use the contextual feature of the interested targets to elaborate background perturbations for the uncovered attacks in physical scenarios. Experiments with proportional scaling are conducted to evaluate the effectiveness of the proposed method, demonstrating its superiority in terms of both attack efficacy and physical practicality. Jiawei Lian, Yuru Su, Mingyang Ma 0004, Shaohui Mei |
IGARSS | 4 |
| 2023 | RIFormer: Learning Rotation-Invariant Features Via TransformerabstractRecently, Transformers have been widely used in many computer vision tasks and have shown promising results. However, like convolutional neural networks (CNNs), Transformers cannot handle rotational variations well, thus hindering its further application in the field of remote sensing. In this paper, we design a rotation-invariant Transformer (RIFormer) to alleviate the abovementioned problem. Moreover, we propose a novel rotation-invariant position embedding (RIPE) to encode positional information of features, and this position-dependent features learned by RIPE is robust to rotations. The experimental results show that proposed RIFormer with RIPE can effectively learn rotation-invariant features compared to the state-of-the-art methods with limited parameters. We provide an open-source implementation of our method. It is publicly available at https://github.com/psychAo/RIFormer. Shaohui Mei, Mingyang Ma 0004 |
IGARSS | 3 |
| 2023 | Multi-scale deep feature fusion based sparse dictionary selection for video summarization
Mingyang Ma 0004, Shuai Wan, Xiuxiu Han, Shaohui Mei |
Signal Process. Image Commun. | 2 |
| 2023 | Hierarchical Feature Fusion of Transformer With Patch Dilating for Remote Sensing Scene ClassificationabstractRecently, the Transformer-based technique has emerged as a promising solution for modeling contextual information in Remote Sensing (RS) scenes and has found widespread applications in RS scene classification. However, how to make full use of intermediate features learned in Transformers is of crucial importance in the RS scene classification tasks. Therefore, this paper proposes a Hierarchical Feature Fusion of Transformer with Patch Dilating (HFFT-PD), which aims to capture rich contextual information from hierarchical features to enhance the performance of RS scene classification. Specifically, the HFFT-PD model consists of a Hierarchical Transformer Merging (HTM) block and a Lightweight Adaptive Channel Compression (LACC) module, in which the HTM is specially designed for the Transformer architecture to bridge the semantic gaps between features from different hierarchical blocks, and the LACC accounts for the significance of distinct channels in the ultimate classification features. In addition, a brand-new Patch Dilating strategy is uniquely designed for the Transformer paradigm, functioning as a reassembly operator predicated on patch features. Contrasting with conventional upsampling techniques, Patch Dilating facilitates upsampling without requiring supplementary information, while concurrently preserving the semantic content of local spatial structure. Extensive and rigorous experiments conducted on the UCM, AID, and NWPU-45 datasets, with training ratios of 80%, 50%, and 20% respectively, demonstrate that our proposed HFFT-PD outperforms the baseline at least by 0.59%, 0.44%, and 0.99% respectively, showcasing the significant superiority of our HFFT-PD over contemporary state-of-the-art methodologies. Mingyang Ma 0004, Yong Li 0036, Shaohui Mei, Zonghao Han, Jian Zhao 0002 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | CBA: Contextual Background Attack Against Optical Aerial Detection in the Physical WorldabstractPatch-based physical attacks have increasingly aroused concerns. However, most existing methods focus on obscuring targets captured on the ground, and some of these methods are simply extended to deceive aerial detectors. They smear the targeted objects in the physical world with the elaborated adversarial patches, which can only slightly sway the aerial detectors’ prediction and with weak attack transferability. To address the above issues, a novel Contextual Background Attack (CBA) framework is proposed to fool aerial detectors in the physical world, which can achieve strong attack efficacy and transferability in real-world scenarios even without smudging the interested objects at all. Specifically, the targets of interest, i.e. the aircraft in aerial images, are adopted to mask adversarial patches. The pixels outside the mask area are optimized to make the generated adversarial patches closely cover the critical contextual background area for detection, which contributes to gifting adversarial patches with more robust and transferable attack potency in the real world. To further strengthen the attack performance, the adversarial patches are forced to be outside targets during training, by which the detected objects of interest, both on and outside patches, benefit the accumulation of attack efficacy. Consequently, the sophisticatedly designed patches are gifted with solid fooling efficacy against objects both on and outside the adversarial patches simultaneously. Extensive proportionally scaled experiments are performed in physical scenarios, demonstrating the superiority and potential of the proposed framework for physical attacks. We expect that the proposed physical attack method will serve as a benchmark for assessing the adversarial robustness of diverse aerial detectors and defense methods. The code has been released at https://github.com/JiaweiLian/CBA. Jiawei Lian, Yuru Su, Mingyang Ma 0004, Shaohui Mei |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Spectral Correlation-Based Diverse Band Selection for Hyperspectral Image ClassificationabstractBand selection which can reduce the spectral dimensionality effectively, has become one of the most popular topics in hyperspectral image (HSI) analysis. Recently, sparse representation based band selection (BS) has emerged as a popular tool. The existing sparse models mainly focus on minimizing reconstruction error and sparsity, while do not fully exploit the unique correlations among hundreds of continuous bands, which may cause representative bands missed and highly-correlated bands selected. Therefore, this paper proposes the spectral correlation based diverse band selection (SCDBS) for HSIs to improve representativeness and diversity of the selected bands. Specifically, a correlation derived weight is used to perform weighted sparse reconstruction to select the bands that are more correlated to the whole HSI, and a correlation minimization term is designed to remove the highly-correlated bands simultaneously. In addition, the proposed method imposes an adjustable sparse constraint by using an ℓ2,0 Mingyang Ma 0004, Shaohui Mei, Fan Li 0003, Yaoyang Ge, Qian Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Rotation-Invariant Feature Learning via Convolutional Neural Network With Cyclic Polar Coordinates Convolutional LayerabstractConvolutional neural networks (CNNs) have been demonstrated to be powerful tools to automatically learn effective features from large datasets. Though features learned in CNNs are approximately scale-, translation-, and position-invariant, and their capacity in dealing with image rotations remains limited. In this article, a novel cyclic polar coordinate convolutional layer (CPCCL) is proposed for CNNs to handle the problem of rotation invariance for feature learning. First, the proposed CPCCL converts rotation variation into translation variation using polar coordinates transformation, which can easily be handled by CNNs. Moreover, cyclic convolution is designed to completely handle the translation variation converted from rotation variation by conducting convolution in a cyclic shift mode. Note that the proposed CPCCL is capable of generalization and can be used as a preprocessing layer for classification CNNs to learn the rotation-invariant feature. Extensive experiments over three benchmark datasets demonstrate that the proposed CPCCL can clearly handle the rotation-sensitive problem in traditional CNNs and outperforms several state-of-the-art rotation-invariant feature learning algorithms. Shaohui Mei, Ruoqiao Jiang, Mingyang Ma 0004 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | Lightweight Multiresolution Feature Fusion Network for Spectral Super-ResolutionabstractSpectral super-resolution (SR), which reconstructs high spatial-resolution hyperspectral images (HSIs) from RGB inputs, has been demonstrated to be one of the effective computational imaging techniques to acquire HSIs. Though deep neural networks have shown their superiority in such a complex mapping problem, existing networks generally involve a very complex structure with huge amounts of parameters, resulting in giant memory occupation. In this article, a lightweight multiresolution feature fusion network (MRFN) is proposed, which adopts a multiresolution feature extraction and fusion framework to fully explore RGB inputs in different scales of resolution. Specifically, a lightweight feature extraction module (LFEM), which adopts cheap convolution and attention mechanisms, is constructed to explore different scales of features under a lightweight structure. Moreover, a hybrid loss function is proposed by encountering not only pixel-value level reconstruction error but also spectral continuity and fidelity. Experiments over three benchmark datasets, i.e., CAVE, Interdisciplinary Computational Vision Laboratory (ICVL), and NTIRE2022 datasets, have demonstrated that the proposed MRFN can reconstruct HSIs from RGB inputs in higher quality with fewer parameters and computational floating-point operations (FLOPs) compared with several state-of-the-art networks. Shaohui Mei, Ge Zhang 0006, Nan Wang 0026, Mingyang Ma 0004, Yifan Zhang 0006, Yan Feng 0005 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Hyperspectral Image Classification Using Hierarchical Spatial-Spectral TransformerabstractIn recent years, convolutional neural networks (CNNs) have been successfully applied in hyperspectral image (HSI) classification tasks. However, the spatial-spectral features within an HSI have not been well explored using convolutions in CNNs. In the paper, a novel end-to-end hierarchical spatial-spectral transformer (HSST) is proposed for HSI classification, in which effective spatial-spectral features are emphasized using multi-head self-attention mechanism (MHSA). MHSA module captures better internal correlation of HSI data than the traditional convolution operation and can compute weighting scores for spatial and spectral context of pixels. Furthermore, a hierarchical architecture is designed to reduce a large number of parameters in the original transformer-style networks while still achieving satisfying classification results. Experimental results over two benchmark HSI datasets demonstrated the proposed HSST obviously outperforms several state-of-the-art deep learning-based HSI classification algorithms. Shaohui Mei, Mingyang Ma 0004, Fulin Xu, Yifan Zhang 0006, Qian Du 0001 |
IGARSS | 3 |
| 2022 | Reconstructing Hyperspectral Images from RGB Inputs Based on Intrinsic Image DecompositionabstractSpectral super-resolution (SR), which generally reconstructs hyperspectral images (HSIs) from RGB inputs, has attracted lots of attention recently. In this paper, a spectral SR algorithm based on intrinsic image decomposition (IID) is proposed, in which RGB images are decomposed into reflectance images and shading images to fully explore RGB features for HSI reconstruction. Considering that features of the reflectance image are only related to the material of objects, the sparsity of material reflectivity is used to reconstruct the reflectance image of HSI. Moreover, an convonlutional neural network (CNN) is constructed to reconstruct shading parts of HSI. Finally, these two reconstructed results are fused to generate the high spectral resolution HSI and an enhancement network is also designed to further improve the recontruction performance. Experimental results with two benchmark datasets, ICVL and CAVE, demonstrate that the performance of the proposed algorithm is superior to several state-of-the-art spectral SR algorithms. Nan Wang 0026, Shaohui Mei, Yifan Zhang 0006, Mingyang Ma 0004, Xiangqing Zhang |
IGARSS | 5 |
| 2022 | Benchmarking Adversarial Patch Against Aerial DetectionabstractDeep neural networks (DNNs) have become essential for aerial detection. However, DNNs are vulnerable to adversarial examples, which pose great security concerns for security-critical systems. Researchers recently devised adversarial patches to evaluate the vulnerability of DNNs-based aerial detection methods physically. Nonetheless, adversarial patches generated by existing algorithms are not strong enough and extremely time-consuming. Moreover, the complicated physical factors are not accommodated well during the optimization process. In this paper, a novel adaptive-patch-based physical attack (AP-PA) framework is proposed to alleviate the above problems, which achieves state-of-the-art performance in both accuracy and efficiency. Specifically, the AP-PA aims to generate adversarial patches that are adaptive in both physical dynamics and varying scales, and by which the particular targets can be hidden from being detected. Furthermore, the adversarial patch is also gifted with attack effectiveness against all targets of the same class with a patch outside the target (No need to smear targeted objects) and robust enough in the physical world. In addition, a new loss is devised to consider more available information of detected objects to optimize the adversarial patch, which can significantly improve the patch’s attack efficacy (Average precision drop up to 87.86% and 85.48% in white-box and black-box settings, respectively) and optimizing efficiency. We also establish one of the first comprehensive, coherent, and rigorous benchmarks to evaluate the attack efficacy of adversarial patches on aerial detection tasks. Finally, several proportionally scaled experiments are performed physically to demonstrate that the elaborated adversarial patches can successfully deceive aerial detection algorithms in dynamic physical circumstances. The code is available at https://github.com/JiaweiLian/AP-PA. Jiawei Lian, Shaohui Mei, Mingyang Ma 0004 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Hyperspectral Image Classification Using Group-Aware Hierarchical TransformerabstractHyperspectral image (HSI) classification is a critical task with numerous applications in the field of remote sensing. Although convolutional neural networks have achieved remarkable success in computer vision, they are still limited in the ability to model long-term dependencies due to small receptive fields. Recently, vision transformers have been used in HSI classification, where multi-head self-attention (MHSA), as the key feature extractor of transformers, learns global dependencies in long-range positions and bands of HSI pixels. Existing vision transformers for classifying HSIs with a large number of bands, however, have some limitations in that features extracted by MHSA may exhibit over-dispersion. In this article, we propose a Group-Aware Hierarchical Transformer (GAHT) for HSI classification, which confines MHSA to the local spatial–spectral context by introducing a new grouped pixel embedding (GPE) module. The GPE emphasizes local relationships within HSI spectral channels, resulting in a global–local fashion from a spatial–spectral context for HSI classification. In addition, we construct our transformer in a hierarchical manner, which can significantly improve classification accuracy with only a few parameters. Extensive experiments on four benchmark HSI datasets demonstrate that the proposed method outperforms state-of-the-art HSI classification algorithms. The source code is available athttps://github.com/MeiShaohui/Group-Aware-Hierarchical-Transformer. Shaohui Mei, Mingyang Ma 0004, Fulin Xu |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Spectral Variability Augmented Sparse Unmixing of Hyperspectral ImagesabstractSpectral unmixing expresses the mixed pixels existing in hyperspectral images as the product of endmembers and their corresponding fractional abundances, which has been widely used in hyperspectral imagery analysis. However, the endmember spectra even for pixels from the same material of an image may include variability due to the influence of lighting conditions and inherent properties of materials within different pixels. Though thein situspectral library has been used to accommodate such variability by using multiplein situspectra to represent each kind of material, the performance improvement may be restricted due to the limited number of endmembers for each material. Therefore, in this article, spectral variability is directly extracted from anin situendmember library and considered to be transferable among different endmembers for the first time. Furthermore, such a spectral variability is further used to augment sparse unmixing by synchronously performing endmember-based reconstruction and spectral variability-augmented reconstruction in the sparse unmixing model. By, respectively, imposing sparse and smoothness regularization over abundances and variability coefficients, a convex optimization-based spectral variability augmented sparse unmixing (SVASU) is finally proposed, and its convergence performance is also analyzed. Experiments conducted over synthetic and real-world datasets demonstrate that the proposed SVASU method not only significantly improves the unmixing performance of conventional spectral library-based unmixing but also outperforms several state-of-the-art sparse unmixing algorithms. Ge Zhang 0006, Shaohui Mei, Bobo Xie, Mingyang Ma 0004, Yifan Zhang 0006, Yan Feng 0005, Qian Du 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | Graph Convolutional Dictionary Selection With L₂, ₚ Norm for Video SummarizationabstractVideo Summarization (VS) has become one of the most effective solutions for quickly understanding a large volume of video data. Dictionary selection with self representation and sparse regularization has demonstrated its promise for VS by formulating the VS problem as a sparse selection task on video frames. However, existing dictionary selection models are generally designed only for data reconstruction, which results in the neglect of the inherent structured information among video frames. In addition, the sparsity commonly constrained by$L_{2,1}$norm is not strong enough, which causes the redundancy of keyframes, i.e., similar keyframes are selected. Therefore, to address these two issues, in this paper we propose a general framework called graph convolutional dictionary selection with$L_{2,p}$($0< p\leq 1$) norm (GCDS$_{2,p}$) for both keyframe selection and skimming based summarization. Firstly, we incorporate graph embedding into dictionary selection to generate the graph embedding dictionary, which can take the structured information depicted in videos into account. Secondly, we propose to use$L_{2,p}$($0< p\leq 1$) norm constrained row sparsity, in which$p$can be flexibly set for two forms of video summarization. For keyframe selection,$0< p< 1$can be utilized to select diverse and representative keyframes; and for skimming,$p=1$can be utilized to select key shots. In addition, an efficient iterative algorithm is devised to optimize the proposed model, and the convergence is theoretically proved. Experimental results including both keyframe selection and skimming based summarization on four benchmark datasets demonstrate the effectiveness and superiority of the proposed method. Mingyang Ma 0004, Shaohui Mei, Shuai Wan, Zhiyong Wang 0001, Xian-Sheng Hua 0001, David Dagan Feng |
IEEE Trans. Image Process. | 1 |
| 2021 | Similarity Based Block Sparse Subset Selection for Video SummarizationabstractVideo summarization (VS) is generally formulated as a subset selection problem where a set of representative keyframes or key segments is selected from an entire video frame set. Though many sparse subset selection based VS algorithms have been proposed in the past decade, most of them adopt linear sparse formulation in the explicit feature vector space of video frames, and don’t consider the local or global relationships among frames. In this paper, we first extend the conventional sparse subset selection for VS into kernel block sparse subset selection (KBS3) to utilize the advantage of kernel sparse coding and introduce a local inter-frame relationship through packing of frame blocks. Going a step further, we propose a similarity based block sparse subset selection (SB2S3) model by applying a specially designed transformation matrix on the KBS3 model in order to introduce a kind of global inter-frame relationship through the similarity. Finally, a greedy pursuit based algorithm is devised for the proposed NP-hard model optimization. The proposed SB2S3 has the following advantages: 1) through the similarity between each frame and any other frame, the global relationship among all frames can be considered; 2) through block sparse coding, the local relationship of adjacent frames is further considered; and 3) it has a wider application, since features can derive similarity, but not vice versa. It is believed that the effect of modeling such global and local relationships among frames in this paper, is similar to that of modeling the long-range and short-range dependencies among frames in deep learning based methods. Experimental results on three benchmark datasets have demonstrated that the proposed approach is superior to not only other sparse subset selection based VS methods but also most unsupervised deep-learning based VS methods. Mingyang Ma 0004, Shaohui Mei, Shuai Wan, Zhiyong Wang 0001, David Dagan Feng, Mohammed Bennamoun |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | Rotation-Invariant Feature Learning in VHR Optical Remote Sensing Images via Nested Siamese Structure With Double Center LossabstractRotation-invariant features are of great importance for object detection and image classification in very-high-resolution (VHR) optical remote sensing images. Though multibranch convolutional neural network (mCNN) has been demonstrated to be very effective for rotation-invariant feature learning, how to effectively train such a network is still an open problem. In this article, a nested Siamese structure (NSS) is proposed for training the mCNN to learn effective rotation-invariant features, which consists of an inner Siamese structure to enhance intraclass cohesion and an outer Siamese structure to enlarge interclass margin. Moreover, a double center loss (DCL) function, in which training samples from the same class are mapped closer to each other while those from different classes are mapped far away to each other, is proposed to train the proposed NSS even with a small amount of training samples. Experimental results over three benchmark data sets demonstrate that the proposed NSS trained by DCL is very effective to encounter rotation varieties when learning features for image classification and outperforms several state-of-the-art rotation-invariant feature learning algorithms even when a small amount of training samples are available. Ruoqiao Jiang, Shaohui Mei, Mingyang Ma 0004 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2021 | Keyframe Extraction From Laparoscopic Videos via Diverse and Weighted Dictionary SelectionabstractLaparoscopic videos have been increasingly acquired for various purposes including surgical training and quality assurance, due to the wide adoption of laparoscopy in minimally invasive surgeries. However, it is very time consuming to view a large amount of laparoscopic videos, which prevents the values of laparoscopic video archives from being well exploited. In this paper, a dictionary selection based video summarization method is proposed to effectively extract keyframes for fast access of laparoscopic videos. Firstly, unlike the low-level feature used in most existing summarization methods, deep features are extracted from a convolutional neural network to effectively represent video frames. Secondly, based on such a deep representation, laparoscopic video summarization is formulated as a diverse and weighted dictionary selection model, in which image quality is taken into account to select high quality keyframes, and a diversity regularization term is added to reduce redundancy among the selected keyframes. Finally, an iterative algorithm with a rapid convergence rate is designed for model optimization, and the convergence of the proposed method is also analyzed. Experimental results on a recently released laparoscopic dataset demonstrate the clear superiority of the proposed methods. The proposed method can facilitate the access of key information in surgeries, training of junior clinicians, explanations to patients, and archive of case files. Mingyang Ma 0004, Shaohui Mei, Shuai Wan, Zhiyong Wang 0001, ZongYuan Ge, Vincent Lam, David Dagan Feng |
IEEE J. Biomed. Health Informatics | 1 |
| 2021 | Patch Based Video Summarization With Block Sparse RepresentationabstractIn recent years, sparse representation has been successfully utilized for video summarization (VS). However, most of the sparse representation based VS methods characterize each video frame with global features. As a result, some important local details could be neglected by global features, which may compromise the performance of summarization. In this paper, we propose to partition each video frame into a number of patches and characterize each patch with global features. Instead of concatenating the features of each patch and utilizing conventional sparse representation, we formulate the VS problem with such video frame representation as block sparse representation by considering each video frame as a block containing a number of patches. By taking the reconstruction constraint into account, we devise a simultaneous version of block-based OMP (Orthogonal Matching Pursuit) algorithm, namely SBOMP, to solve the proposed model. The proposed model is further extended to a neighborhood based model which considers temporally adjacent frames as a super block. This is one of the first sparse representation based VS methods taking both spatial and temporal contexts into account with blocks. Experimental results on two widely used VS datasets have demonstrated that our proposed methods present clear superiority over existing sparse representation based VS methods and are highly comparable to some deep learning ones requiring supervision information for extra model training. Shaohui Mei, Mingyang Ma 0004, Shuai Wan, Junhui Hou, Zhiyong Wang 0001, David Dagan Feng |
IEEE Trans. Multim. | 2 |
| 2020 | Video summarization via block sparse dictionary selection
Mingyang Ma 0004, Shaohui Mei, Shuai Wan, Junhui Hou, Zhiyong Wang 0001, David Dagan Feng |
Neurocomputing | 1 |
| 2019 | Fusing Deep Local and Global Features for Remote Sensing Image Scene ClassificationabstractHigh-resolution remote sensing image scene classification problem has attracted lots of attentions due to its crucial role in a wide range of applications. Recently, many convolutional neural network (CNN) based methods have significantly boosted the performance of image scene classification. However, most of these algorithms only make use of global features learned in the fully connected layers of a CNN, neglecting local features learned in the convolutional layers that is of crucial importance for some remote sensing scenes. Therefore, a sparse representation framework is proposed to make use of both local and global features learned in a CNN for classification of remote sensing scenes by balancing sparse representation based classifiers using these two kinds of features. Specially, in order to reduce redundancy of local features learned in a convolutional layer, the most effective local feature is generated from each convolutional layer using global average pooling and these selected local features of different convolutional layers are cascaded to form local feature representation of the scene. Finally, experimental results on UC-Merced and WHU-RS19 datasets demonstrate that fusing global and local features in a CNN using the proposed sparse representation framework can certainly improve the performance of classification using single kind of features. Moreover, the proposed global average pooling strategy is very effective to fuse local features select most representative features from convolutional layers of a CNN. Keli Yan, Shaohui Mei, Mingyang Ma 0004 |
IGARSS | 3 |
| 2019 | Robust video summarization using collaborative representation of adjacent frames
Mingyang Ma 0004, Shaohui Mei, Shuai Wan, Zhiyong Wang 0001, David Dagan Feng |
Multim. Tools Appl. | 1 |
| 2018 | Video Summarization via Weighted Neighborhood Based RepresentationabstractThe recent explosive growth of multimedia data has posed a new set of challenges in computer vision, and video summarization (VS) techniques are increasingly important to automatically summarize a large amount of multimedia data in an effective and efficient manner. Recent years have witnessed the rise and developments of sparse representation based approaches for VS. While the existing methods select keyframes according to the information contained in the single frame, and such a selection based solely on single-frame information may not be robust. Therefore, in this paper, the information of the single frame's neighborhood is taken into consideration, and different weights are assigned to these neighbouring frames. We formulate the VS problem as a weighted neighborhood based representation model, and design a greedy pursuit algorithm to extract keyframes. Experimental results on a benchmark dataset demonstrate that the proposed method can outperform the state of the arts. Mingyang Ma 0004, Shaohui Mei, Shuai Wan, Zhiyong Wang 0001, Ah Chung Tsoi, David Dagan Feng |
ICIP | 1 |
| 2017 | Exploring the influence of feature representation for dictionary selection based video summarizationabstractDictionary selection based video summarization (VS) algorithms, in which keyframes are considered as a dictionary to reconstruct all the video frames, have been demonstrated to be effective and efficient for video summarization. It has been noticed that the feature representation of video plays a great impact of the performance of VS. In this paper, the influence of feature representation of video frames on the performance of dictionary selection-based VS is for the first time investigated. In addition to the traditional hand-crafted features used in VS, such as color histogram, the deep features learned through deep neural networks are firstly used to represent video frames for dictionary selection-based VS. The impact of dimensionality reduction to the high-dimensional deep learning features on VS is further discussed. Experimental results on a benchmark video dataset demonstrate that deep learning features are able to achieve better performance than traditional hand-crafted features for dictionary selection-based VS. Moreover, the dimensionality of deep learning features can be reduced to decrease the computational cost without the degradation of VS performance. Mingyang Ma 0004, Shaohui Mei, Jingyu Ji, Shuai Wan, Zhiyong Wang 0001, David Dagan Feng |
ICIP | 1 |
| 2017 | Nonlinear kernel sparse dictionary selection for video summarizationabstractSparse dictionary selection (SDS) has demonstrated to be an effective solution for keyframe based video summarization (VS), which generally assumes a linear relation among similar video frames. However, such a linear assumption is not always true for videos. In this paper, the nonlinearity among frames is taken into consideration and a nonlinear SDS model is formulated for VS, in which the nonlinearity is transformed to linearity by projecting a video to a high dimensional feature space induced by a kernel function. Moreover, a kernel simultaneous orthogonal matching pursuit (KSOMP) is proposed to solve the problem. In order to achieve an intuitive and flexible configuration of the VS process, an adaptive criterion is devised to produce video summaries with different lengths for different video content. Experimental results on benchmark video datasets demonstrate that the proposed algorithm outperforms several state-of-the-art VS algorithms. Mingyang Ma 0004, Shaohui Mei, Junhui Hou, Shuai Wan, Zhiyong Wang 0001, David Dagan Feng |
ICME | 1 |