Yanzhou Su

dblp:241/9746 · DBLP profile ↗
← Back
27ranked-venue papers
4as first author
24since 2021 · last 2026
0000-0002-3377-469XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 17 · 3 first-author · 15 since 2021Artificial intelligence and machine learning · 9 · 1 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 8 since 2021
YearPublicationVenuePosition
2026 GMAI-VL & GMAI-VL-5.5M: A Large Vision-Language Model and a Comprehensive Multimodal Dataset Towards General Medical AI
abstract
Despite significant advancements in general AI, its effectiveness in the medical domain is limited by the lack of specialized medical knowledge. To address this, we formulate GMAI-VL-5.5M, a multimodal medical dataset created by converting hundreds of specialized medical datasets with various annotations into high-quality image-text pairs. This dataset offers comprehensive task coverage, diverse modalities, and rich image-text data. Building upon this dataset, we develop GMAI-VL, a 7B-parameter general medical vision-language model, with a three-stage training strategy that enhances the integration of visual and textual information. This approach significantly improves the model's ability to process multimodal data, supporting accurate diagnoses and clinical decision-making. Experiments show that GMAI-VL achieves state-of-the-art performance across various multimodal medical tasks, including visual question answering and medical image diagnosis.
Tianbin Li, Yanzhou Su, Wei Li 0320, Zhe Chen 0017, Ziyan Huang, Guoan Wang, Chenglong Ma 0002, Yanjun Li 0007, Shixiang Tang, Xiaowei Hu 0001, Zhongying Deng, Yuanfeng Ji, Jin Ye 0002, Yu Qiao 0001, Junjun He
AAAI2
2026 MedP-CLIP: Medical CLIP with region-aware prompt integration
Jiahui Peng, He Yao, Yanzhou Su, Sibo Ju, Hongchun Lu, Xue Li 0008, Lincheng Jiang, Min Zhu 0005, Junlong Cheng
Medical Image Anal.4
2026 WideTopo: Improving foresight neural network pruning through training dynamics preservation and wide topologies exploration
Changjian Deng, Jian Cheng 0003, Yanzhou Su, Zeyu An, Ziying Xia, Shiguang Wang
Neural Networks3
2025 Interactive Medical Image Segmentation: A Benchmark Dataset and Baseline
abstract
Interactive Medical Image Segmentation (IMIS) has long been constrained by the limited availability of large-scale, diverse, and densely annotated datasets, which hinders model generalization and consistent evaluation across different models. In this paper, we introduce the IMed-361M benchmark dataset, a significant advancement in general IMIS research. First, we collect and standardize over 6.4 million medical images and their corresponding ground truth masks from multiple data sources. Then, leveraging the strong object recognition capabilities of a vision foundational model, we automatically generated dense interactive masks for each image and ensured their quality through rigorous quality control and granularity management. Unlike previous datasets, which are limited by specific modalities or sparse annotations, IMed-361M spans 14 modalities and 204 segmentation targets, totaling 361 million masks—an average of 56 masks per image. Finally, we developed an IMIS baseline network on this dataset that supports high-quality mask generation through interactive inputs, including clicks, bounding boxes, text prompts, and their combinations. We evaluate its performance on medical image segmentation tasks from multiple perspectives, demonstrating superior accuracy and scalability compared to existing interactive segmentation models. To facilitate research on foundational models in medical computer vision, we release the IMed-361M and model at https://github.com/uni-medical/IMIS-Bench.
Junlong Cheng, Jin Ye 0002, Guoan Wang, Tianbin Li, Haoyu Wang 0010, He Yao, Yanzhou Su, Min Zhu 0005, Junjun He
CVPR11
2025 Adaptive Semantic Alignment for Automated Radiology Report Generation via Cross-Modal Knowledge Integration
abstract
The increasing volume of radiology examinations has created an urgent need for automated report generation that is comparable to those written by radiologists. The major challenge is achieving precise semantic alignment between images and text, ensuring that generated reports accurately capture and describe the visual findings in medical images. Current approaches often struggle with this alignment, compromising diagnostic accuracy and clinical utility. To address these challenges, we present Adaptive Semantic Alignment Method (ASAM), a novel framework that enhances cross-modal semantic alignment through two innovations. First, we introduce a gated disease knowledge base by memory matrix that provides structured medical context to guide the mapping between image and report modalities. Second, we develop a cross-modal pre-trained visual encoder that enriches feature representation through improved understanding of medical imaging characteristics. Extensive experiments demonstrate that ASAM achieves state-of-the-art performance on two public chest X-ray datasets in BLEU-n metrics.
Sibo Ju, Zhaozhen Chen, Yulong Xiao, Yiqing Shen 0003, Yanzhou Su, Xiangwen Liao
ICME5
2025 Leveraging Semantic Asymmetry for Accurate Gross Tumor Volume Segmentation of Nasopharyngeal Carcinoma in Planning CT
Zeli Chen, Yanzhou Su, Tai Ma, Tony C. W. Mok, Yan-Jie Zhou, Yunhao Bai, Zhilin Zheng, Le Lu 0001, Yirui Wang 0002, Jia Ge, Senxiang Yan, Xianghua Ye, Dakai Jin
MICCAI (2)4
2025 Multi-modal MRI Translation via Evidential Regression and Distribution Calibration
Jiyao Liu, Shangqi Gao, Zhaohu Xing, Junzhi Ning, Yanzhou Su, Xiao-Yong Zhang, Junjun He, Ningsheng Xu, Xiahai Zhuang
MICCAI (8)8
2025 RetinaLogos: Fine-Grained Synthesis of High-Resolution Retinal Images Through Captions
Junzhi Ning, Cheng Tang 0003, Kaijing Zhou, Diping Song, Wei Li 0320, Yanzhou Su, Tianbin Li, Jiyao Liu, Jin Ye 0002, Yuanfeng Ji, Junjun He
MICCAI (16)9
2025 FPN-in-FPN: A Nested Multi-scale Aggregation Network for Polyp Segmentation
Jin Ye 0002, Yanzhou Su, Yicheng Wu 0001, Junjun He, Bohan Zhuang, Zhaolin Chen, Jianfei Cai 0001
MICCAI (11)2
2025 A survey for large language models in biomedicine
Chong Wang 0027, Junjun He, Zhongruo Wang, Erfan Darzi, Jin Ye 0002, Tianbin Li, Yanzhou Su, Jing Ke, Kaili Qu, Pietro Liò, Tianyun Wang, Yu Guang Wang 0001, Yiqing Shen 0003
Artif. Intell. Medicine9
2025 A-Eval: A benchmark for cross-dataset and cross-modality evaluation of abdominal multi-organ segmentation
Ziyan Huang, Zhongying Deng, Jin Ye 0002, Haoyu Wang 0010, Yanzhou Su, Tianbin Li, Junlong Cheng, Jianpin Chen, Junjun He, Yun Gu, Shaoting Zhang 0001, Lixu Gu, Yu Qiao 0001
Medical Image Anal.5
2025 SegRap2023: A benchmark of organs-at-risk and gross tumor volume Segmentation for Radiotherapy Planning of Nasopharyngeal Carcinoma
Xiangde Luo, Yunxin Zhong, Shuolin Liu, Mehdi Astaraki, Simone Bendazzoli, Iuliana Toma-Dasu, Yiwen Ye, Ziyang Chen 0003, Yong Xia 0001, Yanzhou Su, Jin Ye 0002, Junjun He, Zhaohu Xing, Hongqiu Wang, Lei Zhu 0003, Kaixiang Yang 0004, Zhiwei Wang 0002, Chan Woong Lee, Sang Joon Park, Jaehee Chun, Constantin Ulrich, Klaus H. Maier-Hein, Nchongmaje Ndipenoch, Alina Dana Miron, Yongmin Li 0001, Chengyang An, Lisheng Wang, Kaiwen Huang 0002, Yunqi Gu, Tao Zhou 0002, Mu Zhou, Shichuan Zhang, Wenjun Liao, Guotai Wang, Shaoting Zhang 0001
Medical Image Anal.12
2025 TransCLIP: Transferring Vision-Language Models for Efficient Video Action Recognition
abstract
Transferring contrastive vision-language pretrained models, such as contrastive language–image pretraining (CLIP), to video recognition task has attracted much attention. Recent studies in this area utilized prompt learning within either text or vision branches, or employed an end-to-end CLIP fine-tuning approach. However, these methods do not fully leverage the learning potential of two branches and may compromise zero-shot generalization. In this work, we present a multimodal framework TransCLIP, aiming to adapt vision–language models by integrating both adapter and prompt tuning techniques for the vision and text encoders. Specifically, we incorporate learnable prompt tokens into each transformer encoder layer’s input of vision and text branches, and integrate lightweight adapters into the key and value matrices of the multi-head self-attention modules, enhancing the model’s capability to capture more related video-specific features. To effectively leverage temporal information in videos, we implement a temporal difference attention module (TemDiff attn) that explicitly computes differences between adjacent frame embeddings and conducts difference-level attention to encode motion-related temporal dependency in videos. In addition, a coarse-and-fine contrastive leaning strategy is employed to better align the video and text branches, enhancing the learning capability of the whole framework. Across different evaluation settings, our model consistently outperforms previous State-of-the-Art methods on several video action recognition benchmarks.
Wen Wang 0012, Yanzhou Su, Jason Gu
IEEE Trans. Ind. Informatics2
2025 SAM-Med3D: A Vision Foundation Model for General-Purpose Segmentation on Volumetric Medical Images
abstract
Existing volumetric medical image segmentation models are typically task-specific, excelling at specific targets but struggling to generalize across anatomical structures or modalities. This limitation restricts their broader clinical use. In this article, we introduce segment anything model (SAM)-Med3D, a vision foundation model (VFM) for general-purpose segmentation on volumetric medical images. Given only a few 3-D prompt points, SAM-Med3D can accurately segment diverse anatomical structures and lesions across various modalities. To achieve this, we gather and preprocess a large-scale 3-D medical image segmentation dataset, SA-Med3D-140K, from 70 public datasets and 8K licensed private cases from hospitals. This dataset includes 22K 3-D images and 143K corresponding masks. SAM-Med3D, a promptable segmentation model characterized by its fully learnable 3-D structure, is trained on this dataset using a two-stage procedure and exhibits impressive performance on both seen and unseen segmentation targets. We comprehensively evaluate SAM-Med3D on 16 datasets covering diverse medical scenarios, including different anatomical structures, modalities, targets, and zero-shot transferability to new/unseen tasks. The evaluation demonstrates the efficiency and efficacy of SAM-Med3D, as well as its promising application to diverse downstream tasks as a pretrained model. Our approach illustrates that substantial medical resources can be harnessed to develop a general-purpose medical AI for various potential applications. Our dataset, code, and models are available at: https://github.com/uni-medical/SAM-Med3D.
Haoyu Wang 0010, Sizheng Guo, Jin Ye 0002, Zhongying Deng, Junlong Cheng, Tianbin Li, Jianpin Chen, Yanzhou Su, Ziyan Huang, Yiqing Shen 0003, Shaoting Zhang 0001, Junjun He
IEEE Trans. Neural Networks Learn. Syst.8
2024 A Fine-tuning Dataset and Benchmark for Large Language Models for Protein Understanding
abstract
The high similarities between protein sequences and natural language, particularly in their sequential data structures, have driven parallel advancements in deep learning models for both domains. In natural language processing (NLP), large language models (LLMs) have achieved remarkable success in tasks such as text generation, translation, and conversational agents, owing to their extensive training on diverse datasets that enable them to capture complex language patterns and generate human-like text. Inspired by these advancements, researchers have attempted to adapt LLMs for protein understanding by integrating a protein sequence encoder with a pre-trained LLM, following designs like LLaVa. However, this adaptation raises a fundamental question: "Can LLMs, originally designed for NLP, effectively comprehend protein sequences as a form of language?" Current datasets fall short in addressing this question due to the lack of a direct correlation between protein sequences and corresponding text descriptions, limiting the ability to train and evaluate LLMs for protein understanding effectively. To bridge this gap, we introduce ProteinLMDataset, a dataset specifically designed for further self-supervised pretraining and supervised fine-tuning (SFT) of LLMs to enhance their capability for protein sequence comprehension. Specifically, ProteinLMDataset includes 17.46 billion tokens for pretraining and 893K instructions for SFT. Additionally, we present ProteinLMBench, the first benchmark dataset consisting of 944 manually verified multiple-choice questions for assessing the protein understanding capabilities of LLMs. ProteinLMBench incorporates protein-related details and sequences in multiple languages, establishing a new standard for evaluating LLMs’ abilities in protein comprehension. The large language model InternLM2-7B, pretrained and fine-tuned on the ProteinLMDataset, outperforms GPT-4 on ProteinLMBench, achieving the highest accuracy score. The dataset and the benchmark are available at https://huggingface. co/datasets/tsynbio/ProteinLMDataset/ and https://huggingface.co/datasets/tsynbio/ProteinLMBench. The code is available at https://github.com/tsynbio/ProteinLMDataset/.
Yiqing Shen 0003, Michail Mamalakis, Luhan He, Tianbin Li, Yanzhou Su, Junjun He, Yu Guang Wang 0001
BIBM7
2024 TourSynbio: A Multi-Modal Large Model and Agent Framework to Bridge Text and Protein Sequences for Protein Engineering
abstract
The structural similarities between protein sequences and natural languages have led to parallel advancements in deep learning across both domains. While large language models (LLMs) have achieved much progress in the domain of natural language processing, their potential in protein engineering remains largely unexplored. Previous approaches have equipped LLMs with protein understanding capabilities by incorporating external protein encoders, but this fails to fully leverage the inherent similarities between protein sequences and natural languages, resulting in sub-optimal performance and increased model complexity. To address this gap, we present TourSynbio-7B, the first multi-modal large model specifically designed for protein engineering tasks without external protein encoders. TourSynbio-7B demonstrates that LLMs can inherently learn to understand proteins as language. The model is post-trained and instruction fine-tuned on InternLM2-7B using ProteinLM-Dataset, a dataset comprising 17.46 billion tokens of text and protein sequence for self-supervised pretraining and 893K instructions for supervised fine-tuning. TourSynbio7B outperforms GPT-4 on the ProteinLMBench, a benchmark of 944 manually verified multiple-choice questions, with 62.18% accuracy. Leveraging TourSynbio-7B’s enhanced protein sequence understanding capability, we introduce TourSynbioAgent, an innovative framework capable of performing various protein engineering tasks, including mutation analysis, inverse folding, protein folding, and visualization. TourSynbio-Agent integrates previously disconnected deep learning models in the protein engineering domain, offering a unified conversational user interface for improved usability. Finally, we demonstrate the efficacy of TourSynbio-7B and TourSynbio-Agent through two wet lab case studies on vanilla key enzyme modification and steroid compound catalysis. Our results show that this combination facilitates protein engineering tasks in wet labs, leading to higher positive rates, improved mutations, shorter delivery times, and increased automation. The model weights are available at https://huggingface.co/tsynbio/Toursynbio and codes at https://github.com/tsynbio/TourSynbio.
Yiqing Shen 0003, Michail Mamalakis, Yungeng Liu, Tianbin Li, Yanzhou Su, Junjun He, Pietro Liò, Yu Guang Wang 0001
BIBM6
2024 A Tiny Efficient U-Net with Gated Linear Attention for Medical Image Segmentation
abstract
Medical image segmentation is crucial for diagnosis and treatment planning. While recent advancements in deep learning, particularly UNet variants, have improved segmentation performance, they often result in increased model complexity, which limits their real-time applicability on resource-constrained devices in clinical settings. To address this challenge, we present the Tiny Efficient U-Net (TE-UNet), a novel lightweight model balancing efficiency and accuracy. TE-UNet uses a U-shaped encoder-decoder framework with a gated linear attention mechanism to process low-level and high-level features, preserving details and reducing complexity. It employs depth-wise separable convolutions for higher-level processing, enhancing efficiency without losing performance. Additionally, skip connections improve multi-scale feature extraction and information flow. Experiments on two public datasets across different modalities demonstrate that TE-UNet outperforms ten state-of-the-art methods, maintaining a parameter size under 40KB and low computational cost. TE-UNet makes real-time segmentation more accessible for various clinical applications.
Sibo Ju, Zhaozhen Chen, Xiangwen Liao, Yiqing Shen 0003, Junjun He, Yanzhou Su
BIBM6
2024 GMAI-MMBench: A Comprehensive Multimodal Evaluation Benchmark Towards General Medical AI
abstract
Large Vision-Language Models (LVLMs) are capable of handling diverse data types such as imaging, text, and physiological signals, and can be applied in various fields. In the medical field, LVLMs have a high potential to offer substantial assistance for diagnosis and treatment. Before that, it is crucial to develop benchmarks to evaluate LVLMs' effectiveness in various medical applications. Current benchmarks are often built upon specific academic literature, mainly focusing on a single domain, and lacking varying perceptual granularities. Thus, they face specific challenges, including limited clinical relevance, incomplete evaluations, and insufficient guidance for interactive LVLMs. To address these limitations, we developed the GMAI-MMBench, the most comprehensive general medical AI benchmark with well-categorized data structure and multi-perceptual granularity to date. It is constructed from 284 datasets across 38 medical image modalities, 18 clinical-related tasks, 18 departments, and 4 perceptual granularities in a Visual Question Answering (VQA) format. Additionally, we implemented a lexical tree structure that allows users to customize evaluation tasks, accommodating various assessment needs and substantially supporting medical AI research and applications. We evaluated 50 LVLMs, and the results show that even the advanced GPT-4o only achieves an accuracy of 53.96\%, indicating significant room for improvement. Moreover, we identified five key insufficiencies in current cutting-edge LVLMs that need to be addressed to advance the development of better medical applications. We believe that GMAI-MMBench will stimulate the community to build the next generation of LVLMs toward GMAI.
Jin Ye 0002, Guoan Wang, Yanjun Li 0007, Zhongying Deng, Wei Li 0320, Tianbin Li, Haodong Duan, Ziyan Huang, Yanzhou Su, Benyou Wang, Shaoting Zhang 0001, Jianfei Cai 0001, Bohan Zhuang, Eric J. Seibel, Junjun He, Yu Qiao 0001
NeurIPS10
2024 Global Adaptive Second-Order Transformer for Remote Sensing Image Semantic Segmentation
abstract
In the domain of remote sensing (RS) image analysis, capturing global context is the key for precise semantic segmentation. Current vision transformer (ViT) advance this field by addressing convolutional neural network’s (CNN) local receptive field limitations. However, ViT predominantly rely on the first-order information in image to establish global relationships, often overlooking the potential of second-order information, which is crucial for enhancing the discrimination of ground objects that exhibit high similarity and constant changes. To address this issue, we propose a global adaptive second-order transformer network (GASOT-Net). Specifically, the proposed global adaptive second-order transformer (GASOT) enhances the existing ViT structure by mining second-order information and adaptively fusing it with the first-order information during the process of establishing global dependency relationships. This approach enables the extraction of more discriminative features, thereby enriching the representation of global features. In addition, the local feature aggregation module (LFAM) is proposed to effectively aggregate features from different stages of CNN as input to the GASOT blocks. Moreover, to refine boundaries of complex ground objects, the global feature enhancement module (GFEM) is used in the decoder stage. In particular, GFEM includes two sub modules—feature shift module (FSM) and hierarchical feature fusion module (HFFM). FSM is used to enhance the local feature representation at first, and then, HFFM hierarchically aggregates local and global features from different stages. We conduct extensive experiments on four benchmark RS datasets, and the results show that our GASOT-Net outperforms other state-of-the-art methods. The code will be available at:https://github.com/j136812832/GASOT-Net.
Jian Cheng 0003, Yanzhou Su, Changjian Deng, Ziying Xia, Nyima Tashi
IEEE Trans. Geosci. Remote. Sens.3
2023 Revisiting Feature Propagation and Aggregation in Polyp Segmentation
Yanzhou Su, Yiqing Shen 0003, Jin Ye 0002, Junjun He, Jian Cheng 0003
MICCAI (5)1
2023 Accurate polyp segmentation through enhancing feature fusion and boosting boundary performance
Yanzhou Su, Jian Cheng 0003, Chuqiao Zhong, Chengzhi Jiang, Jin Ye 0002, Junjun He
Neurocomputing1
2022 Semantic Segmentation for High-Resolution Remote-Sensing Images via Dynamic Graph Context Reasoning
abstract
Semantic segmentation for high-resolution remote-sensing (HRRS) images is one of the most challenging tasks in remote-sensing images understanding. Capturing long-range dependencies in feature representations is crucial for semantic segmentation. Recent graph-based global reasoning networks (GloRe) focus on modeling the global contextual relationship between latent nodes based on fully connected graph in interaction space. However, such a dense operation is susceptible to redundant features. Most importantly, it treats each node equally, ignoring the contextual relationship between nodes in graphs. In this work, we propose to explore more effective contextual representations in semantic segmentation by introducing dynamic graph contextual reasoning module overGloRe, dubbed DGCR. It incorporates local semantic information that represents the relationships between nodes to perform long-range contextual reasoning. More specifically, to provide effectively and flexible reasoning in graph-based reasoning approaches, we construct$k$-nearest neighbor (KNN) graphs rather than fully connected graphs using only the$k$closest nodes depends on pairwise semantic distance. Extensive experiments on the International Society for Photogrammetry and Remote Sensing (ISPRS) Vaihingen and Potsdam datasets demonstrate the effectiveness and superiority of our proposed DGCR module over other state-of-the-art methods.
Yanzhou Su, Jian Cheng 0003, Wen Wang 0012, Haiwei Bai, Haijun Liu 0001
IEEE Geosci. Remote. Sens. Lett.1
2021 Semantic Segmentation for High-Resolution Remote Sensing Images by Light-Weight Network
abstract
Accurate segmentation of high-resolution remote sensing images is increasingly demanded, yet poses significant challenges for algorithm efficiency. Most current approaches pursue accuracy by employing global context information to enhance the overall consistency or utilize multi-scale features or attention mechanisms to optimize object details, without considering the network complexity uniformly. In this paper, we propose a light-weight semantic segmentation network for HRRS images by way of explicitly supervising the objects' body and edge features to optimize the overall consistency and object details of semantic segmentation at the same time. Furthermore, we introduce a score-based feature fusion module to establish the long-range dependency between pixels in the final stage of feature fusion (to combine the body and edge features) effectively. Experiments on ISPRS Vaihingen dataset show an obvious advantage of the proposed approach compared with the existing approaches. Specifically, it achieves 89.60% overall accuracy with only 2.83M parameters and 2.37GFLOPs computation costs.
Changjian Deng, Leikun Liang, Yanzhou Su, Changtao He, Jian Cheng 0003
IGARSS3
2021 TTPP: Temporal Transformer with Progressive Prediction for efficient action anticipation
Wen Wang 0012, Xiaojiang Peng, Yanzhou Su, Yu Qiao 0001, Jian Cheng 0003
Neurocomputing3
2020 Enhancing the discriminative feature learning for visible-thermal cross-modality person re-identification
Haijun Liu 0001, Jian Cheng 0003, Wen Wang 0012, Yanzhou Su, Haiwei Bai
Neurocomputing4
2019 Semantic Segmentation of High Resolution Remote Sensing Image Based on Batch-Attention Mechanism
abstract
Deep convolution neural network has been widely used in recent works for semantic segmentation of High Resolution Remote Sensing(HRRS) images. Because of the limitation of GPU memory, HRRS images are usually split into several sub-images for training convolutional neural networks. For each sub-image, the segmentation model may not have enough information to predict the segmentation map very well. In order to alleviate this problem, we propose to apply a batch-attention module to capture the discriminative information from similar objects, which come from other sub-images in a mini-batch. We also utilize global attention upsample module as the decoder to provide global context and fuse high and low level information better. We evaluate our model on the Potsdam dataset and achieve 88.30% pixAcc and 73.78% mIoU.
Yanzhou Su, Feng Wang 0015, Jian Cheng 0003
IGARSS1
2019 A Discriminatively Learned CNN Embedding For Remote Sensing Image Scene Classification
abstract
In this work, a discriminatively learned CNN embedding is proposed for remote sensing image scene classification. Our proposed siamese network simultaneously computes the classification loss function and the metric learning loss function of the two input images. Specifically, for the classification loss, we use the standard cross-entropy loss function to predict the classes of the images. For the metric learning loss, our siamese network learns to map the intra-class and inter-class input pairs to a feature space where intra-class inputs are close and inter-class inputs are separated by a margin. Concretely, for remote sensing image scene classification, we would like to map images from the same scene to feature vectors that are close, and map images from different scenes to feature vectors that are widely separated. Experiments are conducted on three different remote sensing image datasets to evaluate the effectiveness of our proposed approach. The results demonstrate that the proposed method achieves an excellent classification performance.
Wen Wang 0012, Lijun Du, Yinxing Gao, Yanzhou Su, Feng Wang 0015, Jian Cheng 0003
IGARSS4