EDBT 2026 Demo / reviewers in the wild / expert
Guangzhe Zhao
dblp:223/9562
· DBLP profile ↗
29ranked-venue papers
10as first author
21since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 14 · 4 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 4 first-author · 4 since 2021Security and privacy · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 4 · 4 since 2021Computer networks · 2Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HGAvatar: One-shot high-quality 3D Gaussian head avatar
Feihu Yan, Guangzhe Zhao |
Comput. Vis. Image Underst. | 4 |
| 2026 | Progressive densification of 3D Gaussians for high-fidelity talking head synthesis
Xueni Guo, Feihu Yan, Guangzhe Zhao |
Frontiers Comput. Sci. | 5 |
| 2026 | Fast online learning algorithm based on modified hierarchical Unimodal Thompson Sampling
Tianchi Zhao 0001, Jing Li 0114, Hongyin Shi, Guangzhe Zhao |
Pattern Recognit. | 6 |
| 2025 | An Unified Stochastic Fusion of Dual Diffusion Paths for Faithful Architectural Image SynthesisabstractArchitectural image generation aims to automate the creation of visually and semantically coherent architectural visuals based on inputs such as text prompts or sketches. While diffusion models have recently achieved remarkable success in this domain, they still struggle with preserving structural distortions and semantic misalignment under complex or multi-modal prompts. In this paper, we propose a lightweight and modular enhancement to existing diffusion pipelines that improves consistency and architectural plausibility without requiring external conditioning modules or fixed fine-tuning strategies. Our method introduces a text-image-image conditional generation framework, where two image generated from the same prompt are fused in latent space under unified stochastic control, enabling coherent sampling and minimizing semantic degradation. This approach is model-agnostic and compatible with optional adaptation modules, making it highly extensible and adaptable across different architectures and tasks. Experiments demonstrate that our method consistently produces images with improved semantic fidelity, structural consistency, and stylistic harmony. Furthermore, its modular design allows for efficient integration into existing workflows, offers a promising foundation for controllable and reliable architectural image synthesis. Guangzhe Zhao, Shilong Yang, Feihu Yan |
ICTAI | 1 |
| 2025 | Tri-Plane Dynamic Neural Radiance Fields for High-Fidelity Talking Portrait SynthesisabstractABSTRACT Neural radiation field (NeRF) has been widely used in the field of talking portrait synthesis. However, the inadequate utilisation of audio information and spatial position leads to the inability to generate images with high audio‐lip consistency and realism. This paper proposes a novel tri‐plane dynamic neural radiation field (Tri‐NeRF) that employs an implicit radiation field to study the impacts of audio on facial movements. Specifically, Tri‐NeRF propose tri‐plane offset network (TPO‐Net) to offset spatial positions in three 2D planes guided by audio. This allows for sufficient learning of audio features from image features in a low dimensional state to generate more accurate lip movements. In order to better preserve facial texture details, we innovatively propose a new gated attention fusion module (GAF) to dynamically fuse features based on strong and weak correlation of cross‐modal features. Extensive experiments have demonstrated that Tri‐NeRF can generate talking portraits with audio‐lip consistency and realism. Xueni Guo, Feihu Yan, Guangzhe Zhao |
IET Image Process. | 6 |
| 2025 | StableID: Multimodal learning for stable identity in personalized Text-to-Face generation
Feihu Yan, Guangzhe Zhao |
Pattern Recognit. Lett. | 5 |
| 2025 | Syn-Net: A Synchronous Frequency-Perception Fusion Network for Breast Tumor Segmentation in Ultrasound ImagesabstractAccurate breast tumor segmentation in ultrasound images is a crucial step in medical diagnosis and locating the tumor region. However, segmentation faces numerous challenges due to the complexity of ultrasound images, similar intensity distributions, variable tumor morphology, and speckle noise. To address these challenges and achieve precise segmentation of breast tumors in complex ultrasound images, we propose a Synchronous Frequency-perception Fusion Network (Syn-Net). Initially, we design a synchronous dual-branch encoder to extract local and global feature information simultaneously from complex ultrasound images. Secondly, we introduce a novel Frequency- perception Cross-Feature Fusion (FrCFusion) Block, which utilizes Discrete Cosine Transform (DCT) to learn all-frequency features and effectively fuse local and global features while mitigating issues arising from similar intensity distributions. In addition, we develop a Full-Scale Deep Supervision method that not only corrects the influence of speckle noise on segmentation but also effectively guides decoder features towards the ground truth. We conduct extensive experiments on three publicly available ultrasound breast tumor datasets. Comparison with 14 state-of-the-art deep learning segmentation methods demonstrates that our approach exhibits superior sensitivity to different ultrasound images, variations in tumor size and shape, speckle noise, and similarity in intensity distribution between surrounding tissues and tumors. On the BUSI and Dataset B datasets, our method achieves better Dice scores compared to state-of-the-art methods, indicating superior performance in ultrasound breast tumor segmentation. Guangzhe Zhao, Xingguo Zhu, Feihu Yan, Maozu Guo 0001 |
IEEE J. Biomed. Health Informatics | 1 |
| 2025 | Weakening the Dominant Role of Text: CMOSI Dataset and Multimodal Semantic Enhancement NetworkabstractMultimodal sentiment analysis (MSA) is important for quickly and accurately understanding people's attitudes and opinions about an event. However, existing sentiment analysis methods suffer from the dominant contribution of text modality in the dataset; this is called text dominance. In this context, we emphasize that weakening the dominant role of text modality is important for MSA tasks. To solve the above two problems, from the perspective of datasets, we first propose the Chinese multimodal opinion-level sentiment intensity (CMOSI) dataset. Three different versions of the dataset were constructed: manually proofreading subtitles, generating subtitles using machine speech transcription, and generating subtitles using human cross-language translation. The latter two versions radically weaken the dominant role of the textual model. We randomly collected 144 real videos from the Bilibili video site and manually edited 2557 clips containing emotions from them. From the perspective of network modeling, we propose a multimodal semantic enhancement network (MSEN) based on a multiheaded attention mechanism by taking advantage of the multiple versions of the CMOSI dataset. Experiments with our proposed CMOSI show that the network performs best with the text-unweakened version of the dataset. The loss of performance is minimal on both versions of the text-weakened dataset, indicating that our network can fully exploit the latent semantics in nontext patterns. In addition, we conducted model generalization experiments with MSEN on MOSI, MOSEI, and CH-SIMS datasets, and the results show that our approach is also very competitive and has good cross-language robustness. Ming Yan 0005, Guangzhe Zhao, Guixuan Zhang, Shuwu Zhang |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | Instance-level cross-attention learning for fine-grained customizable face generation
Feihu Yan, Guangzhe Zhao |
Vis. Comput. | 5 |
| 2024 | FTMSNet: Towards boundary-aware polyp segmentation framework based on hybrid Fourier Transform and Multi-scale SubtractionabstractColorectal cancer is one of the diseases with the highest incidence and mortality rates worldwide, posing a severe threat to human life and health. Polyps are the primary cause of colorectal cancer, Colonoscopy is the gold standard for diagnosing colorectal polyps. Accurate polyp segmentation is crucial for patient diagnosis, treatment, and prognosis. However, the irregular shapes and low contrast between colorectal polyps and normal tissue make polyp segmentation a challenging task. Although deep learning-based methods have achieved promising results in this task, few approaches focus on the boundary information of colorectal polyps. In this study, to address the issue of edge blurring in polyp segmentation, we propose a boundary-aware polyp segmentation method based on a hybrid Fourier Transform and Multi-scale Convolutional neural network, referred to as FTMSNet. We introduce the Fourier Transform Module (FTM), which utilizes the Fourier Transform to retain only high-frequency information (such as boundary) in the frequency domain. By leveraging boundary information, our method achieves more precise and clearer boundary delineation for colorectal polyp segmentation. Simultaneously, the Multi-scale Feature Denoising Decoder (MFDD) we introduced is devised to mitigate noise interference during multi-scale information fusion. We validated the performance of FTMSNet on five publicly available datasets. Extensive experimental results demonstrate that our approach surpasses the state-of-the-art polyp segmentation methods. Guangzhe Zhao, Feihu Yan, Maozu Guo 0001 |
BIBM | 1 |
| 2024 | YMamba: A Dual-Branch Network Fusing State Space Model and CNN for Medical Image SegmentationabstractIntelligent technologies like deep learning have significantly improved medical image segmentation, enhancing clinical decision-making and reducing healthcare costs. However, CNN-based methods face limitations due to constrained receptive fields and semantic information loss in deeper layers. Consequently, to address these issues, we innovatively propose the YMamba model which employs a parallel hybrid architecture of CNN and VMamba, enhancing the modeling capability of distant features while maintaining local feature detail textures in medical images without introducing additional parameters. Additionally, our proposed DBFM module employs an enhanced attention strategy to integrate the strengths of both methods more effectively, reinforcing image feature representation and mitigating background noise. Finally, the MCFFD module receives shallow and deep features from the fusion module, addressing the challenge of size variation in target segmentation regions within medical images. Extensive experiments demonstrate that YMamba achieves state-of-the-art results on four medical image datasets BUSI, DDTI, TN3K, and ISIC2016. Guangzhe Zhao, Feihu Yan, Maozu Guo 0001 |
BIBM | 1 |
| 2024 | CMFF-Face: Attention-Based Cross-Modal Feature Fusion for High-Quality Audio-Driven Talking Face GenerationabstractAudio-driven talking face generation creates lip-synchronized and high-quality face videos from given audio and target face images, which is a challenging task due to the inherent modal gap between audio and face images. To address this issue, we propose an attention-based Cross-Modal Feature Fusion network for talking Face generation, called CMFF-Face. Specifically, we introduce a cross-modal feature fusion generator, which incorporates a fusion process in each convolutional encoder layer, allowing for layer-wise fusing of interactive audio and face features to generate high-quality talking faces. Additionally, a lip synchronization discriminator is designed to improve audio-lip synchronization, which uses a two-branch cross-attention mechanism to capture the associations between synchronized audio and face more effectively. Finally, we employ a CLIP-based audio-lip synchronization loss that helps distinguish between positive and negative sample pairs to enhance the lip synchronization. Comprehensive experiments on the LRS2 and LRW datasets demonstrate that our method outperforms the state-of-the-arts in terms of lip synchronization and visual quality. Guangzhe Zhao, Feihu Yan |
ICMR | 1 |
| 2024 | Expression-aware neural radiance fields for high-fidelity talking portrait synthesis
Xueni Guo, Jiahe Li 0007, Feihu Yan, Guangzhe Zhao, Caiyong Wang |
Image Vis. Comput. | 7 |
| 2024 | PMANet: Progressive multi-stage attention networks for skin disease classification
Guangzhe Zhao, Benwang Lin, Feihu Yan |
Image Vis. Comput. | 1 |
| 2023 | Sclera Segmentation and Joint Recognition Benchmarking Competition: SSRBC 2023abstractThis paper presents the summary of the Sclera Segmentation and Joint Recognition Benchmarking Competition (SSRBC 2023) held in conjunction with IEEE International Joint Conference on Biometrics (IJCB 2023). Different from the previous editions of the competition, SSRBC 2023 not only explored the performance of the latest and most advanced sclera segmentation models, but also studied the impact of segmentation quality on recognition performance. Five groups took part in SSRBC 2023 and submitted a total of six segmentation models and one recognition technique for scoring. The submitted solutions included a wide variety of conceptually diverse deep-learning models and were rigorously tested on three publicly available datasets, i.e., MASD, SBVPI and MOBIUS. Most of the segmentation models achieved encouraging segmentation and recognition performance. Most importantly, we observed that better segmentation results always translate into better verification performance. Abhijit Das 0001, Saurabh Atreya, Aritra Mukherjee, Matej Vitek, Caiyong Wang, Guangzhe Zhao, Fadi Boutros, Patrick Siebke, Jan Niklas Kolf, Naser Damer, Sun Ye, Lu Hexin, Fan Aobo, You Sheng, Sabari Nathan, R. Suganya 0001, Rampriya Rajendran Shanthi, Geetanjali Sharma, P. Priyanka, Aditya Nigam, Peter Peer, Umapada Pal 0001, Vitomir Struc |
IJCB | 7 |
| 2023 | SynFacePAD 2023: Competition on Face Presentation Attack Detection Based on Privacy-aware Synthetic Training DataabstractThis paper presents a summary of the Competition on Face Presentation Attack Detection Based on Privacy-aware Synthetic Training Data (SynFacePAD 2023) held at the 2023 International Joint Conference on Biometrics (IJCB 2023). The competition attracted a total of 8 participating teams with valid submissions from academia and industry. The competition aimed to motivate and attract solutions that target detecting face presentation attacks while considering synthetic-based training data motivated by privacy, legal and ethical concerns associated with personal data. To achieve that, the training data used by the participants was limited to synthetic data provided by the organizers. The submitted solutions presented innovations and novel approaches that led to outperforming the considered baseline in the investigated benchmarks. Meiling Fang, Marco Huber, Julian Fierrez, Ramachandra Raghavendra, Naser Damer, Alhasan Alkhaddour, Maksim Kasantcev, Vasiliy Pryadchenko, Ziyuan Yang 0001, Huijie Huangfu, Yi Zhang 0018, Junjun Jiang, Xianming Liu 0005, Xianyun Sun, Caiyong Wang, Zhaohua Chang, Guangzhe Zhao, Juan E. Tapia, Lázaro J. González Soler, Carlos M. Aravena, Daniel Schulz |
IJCB | 20 |
| 2023 | Sclera-TransFuse: Fusing Swin Transformer and CNN for Accurate Sclera SegmentationabstractSclera segmentation is a crucial step in sclera recognition, which has been greatly advanced by Convolutional Neural Networks (CNNs). However, when dealing with non-ideal eye images, many existing CNN-based approaches are still prone to failure. One major reason is that due to the limited range of receptive fields, CNNs are difficult to effectively model global semantic relevance and thus robustly resist noise interference. To solve this problem, this paper proposes a novel two-stream hybrid model, named Sclera-TransFuse, to integrate classical ResNet-34 and recently emerging Swin Transformer encoders. Specially, the self-attentive Swin Transformer has shown a strong ability in capturing long-range spatial dependencies and has a hierarchical structure similar to CNNs. The dual encoders firstly extract coarse- and fine-grained feature representations at hierarchical stages, separately. Then a novel Cross-Domain Fusion (CDF) module based on information interaction and self-attention mechanism is introduced to efficiently fuse the multi-scale features extracted from dual encoders. Finally, the fused features are progressively upsampled and aggregated to predict the sclera masks in the decoder meanwhile deep supervision strategies are employed to learn intermediate feature representations better and faster. Experimental results show that Sclera-TransFuse achieves state-of-the-art performance on various sclera segmentation benchmarks. Additionally, a UBIRIS.v2 subset of 683 eye images with manually labeled sclera masks, and our codes are publicly available to the community through https://github.com/Ihqqq/Sclera-TransFuse. Caiyong Wang, Guangzhe Zhao, Zhaofeng He 0001, Yunlong Wang 0003, Zhenan Sun |
IJCB | 3 |
| 2023 | Iris Liveness Detection Competition (LivDet-Iris) - The 2023 EditionabstractThis paper describes the results of the 2023 edition of the “LivDet” series of iris presentation attack detection (PAD) competitions. New elements in this fifth competition include (1) GAN-generated iris images as a category of presentation attack instruments (PAI), and (2) an evaluation of human accuracy at detecting PAI as a reference benchmark. Clarkson University and the University of Notre Dame contributed image datasets for the competition, composed of samples representing seven different PAI categories, as well as baseline PAD algorithms. Fraunhofer IGD, Beijing University of Civil Engineering and Architecture, and Hochschule Darmstadt contributed results for a total of eight PAD algorithms to the competition. Accuracy results are analyzed by different PAI types, and compared to human accuracy. Overall, the Fraunhofer IGD algorithm, using an attention-based pixel-wise binary supervision network, showed the best-weighted accuracy results (average classification error rate of 37.31%), while the Beijing University of Civil Engineering and Architecture’s algorithm won when equal weights for each PAI were given (average classification rate of 22.15%). These results suggest that iris PAD is still a challenging problem. Patrick Tinsley, Sandip Purnapatra, Mahsa Mitcheff, Aidan Boyd, Colton R. Crum, Kevin W. Bowyer, Patrick J. Flynn, Stephanie Schuckers, Adam Czajka, Meiling Fang, Naser Damer, Caiyong Wang, Xianyun Sun, Zhaohua Chang, Guangzhe Zhao, Juan E. Tapia, Christoph Busch 0001, Carlos M. Aravena, Daniel Schulz |
IJCB | 17 |
| 2023 | MetaScleraSeg: an effective meta-learning framework for generalized sclera segmentation
Caiyong Wang, Wenhui Ma, Guangzhe Zhao, Zhaofeng He 0001 |
Neural Comput. Appl. | 4 |
| 2022 | Research on fatigue detection based on visual featuresabstractAbstract The high incidence of traffic accidents brings immeasurable losses to life. In order to avoid such crises, researchers and automakers have used many methods to solve this problem. Among them, technology based on visual features is widely used in driver fatigue detection. As fatigue detection plays a vital role in the driving process, the high accuracy of fatigue monitoring is very important. This paper focuses on the method based on convolutional neural network to detect driver fatigue. First, in the face detection part, the Single‐Shot Multi‐Box Detector algorithm is used to improve the speed and accuracy of face detection to extract the eye and mouth regions; second, the VGG16 network is used to learn fatigue features, which is performed on the NTHU‐Drowsy Driver Detection (NTHU‐DDD) data set and the other two modified data sets Training test. The main result of this work is that the accuracy of fatigue monitoring is higher than other methods including the original method, with an accuracy rate of over 90%. And it has better generalization ability than the multi‐physical feature fusion detection method. At the same time, we propose the fatigue detection method based on convolutional neural network to improve the advanced driver assistance system (ADAS) to make it more robust and reliable decision making. Guangzhe Zhao, Yanqing He, Hanting Yang |
IET Image Process. | 1 |
| 2022 | Falling motion detection algorithm based on deep learningabstractAbstract Falling is a significant cause of injuries and even death in the elderly. The timely detection of the fall action helps to rescue people who may have physical health problems due to the fall, so fall detection is necessary. The traditional fall detection methods are mostly based on wearable devices, which need to be worn all the time, and the cost of the device is high. In recent years, the fall detection method based on computer vision has become a research hot spot. This paper proposes a framework for falling motion detection based on deep learning. To quickly and accurately classify human movements, a method using bone key points as the feature descriptors of human movements is proposed. The OpenPose algorithm is used to extract the human skeleton point information as the primary human body feature, and then use the deep learning method to classify further and recognise our action features. In this paper, four types of daily actions, such as falling and walking, are classified and recognised. The results show that the algorithm achieves an accuracy of 99.4% on our dataset. Simultaneously, 86.1% accuracy is reached in the public dataset fall detection dataset. Guangzhe Zhao, Zhexue Jin |
IET Image Process. | 2 |
| 2020 | Dual-Stage Construction of Probability for Hyperspectral Image ClassificationabstractRecently, feature extraction-based methods have received increasing attention in the hyperspectral image. In this letter, to ensure a more powerful discriminative ability of extracted features, a dual-stage construction of probability (DSCP) method is proposed for hyperspectral image classification. Specifically, the extended multi-attribute profiles (EMAP) method is applied to extract the shape feature of hyperspectral remote sensing image (HSI) to obtain a more accurate initial probability map. Considering that there are still some noises in the boundaries of the initial probability map, an effective edge-preserving filter-based approach named rolling guidance filter is used for probability post-optimization. Consequently, the class label of each pixel can be determined according to the optimized probability maps. Experiments demonstrate significantly the efficiency of the proposed method in comparison with other advanced methods. Bing Tu, Guangzhe Zhao, Guoyun Zhang |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2020 | Density Peak Covariance Matrix for Feature Extraction of Hyperspectral ImageabstractThe clustering methods have a good application in many aspects, in which the density peak (DP) clustering can effectively cluster similar neighboring pixels so that the features can be extracted well for hyperspectral images (HSIs) classification. In this work, a DP based covariance matrix (DPCM) method is proposed for the feature extraction of HSIs, which not only can effectively extract features but also can reduce the within-class variations and the between-class interference. The proposed method consists of the following steps: First, maximum noise fraction is employed on the original HSI to reduce the computational complexity and eliminate noise. Second, the local densities of the sample are calculated by the DP clustering. Therefore, a reconstructed image can be obtained in which each pixel has a density feature vector. Then, the covariance matrix between each density pixel in the density map is calculated. Last, the extracted covariance matrices are fed back to the support vector machine based on the logarithm Euclidean kernel for label assignment. Experiments are conducted on the Indian pine data set, in which each of the five randomly selected marker data are selected as the training sample. The experimental results show that the method can effectively improve the classification accuracy and is superior to other classification methods. Guangzhe Zhao, Nanying Li, Bing Tu, Guoyun Zhang, Wei He 0021 |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2019 | Multiple convolutional layers fusion framework for hyperspectral image classification
Guangzhe Zhao, Guangyun Liu, Leyuan Fang, Bing Tu, Pedram Ghamisi |
Neurocomputing | 1 |
| 2019 | Iterative fusion convolutional neural networks for classification of optical coherence tomography images
Leyuan Fang, Yuxuan Jin, Laifeng Huang, Guangzhe Zhao |
J. Vis. Commun. Image Represent. | 5 |
| 2018 | Angle-Domain Frequency Synchronization for Massive MIMO Uplink with Adaptive MUI SuppressionabstractIn this paper, we design a novel angle-domain adaptive filtering (ADAF)-based frequency synchronization method for massive multiple-input multiple-output (MIMO) multiuser uplink, which is suitable for users with either separate or overlapped angle-of-arrival (AoA) regions. We first introduce the angle-constraining matrix (ACM), which consists of a set of selected match-filter (MF) beamformers pointing to the AoAs of the interested user. Then, the adaptive beamformer can be acquired by appropriately designing the ADAF vector. Such beamformer can achieve a two-stage adaptive multiuser interference (MUI) suppression, i.e., inherently suppressing the MUI from distant users by ACM in the first stage and substantially mitigating the MUI from adjacent overlapping users by ADAF vector in the second stage. For both separate and mutually overlapping users, the carrier frequency offset (CFO) estimation and subsequent data detection can be performed individually for each user, which reduces the computational complexity. Moreover, ADAF is rather insensitive to imperfect AoA knowledge, making itself promising for multiuser uplink transmission. Numerical results are provided to show the effectiveness of the proposed method, as well as its superiority over existing competitors. Yinghao Ge, Weile Zhang, Feifei Gao 0001, Pengcheng Mu, Guangzhe Zhao |
GLOBECOM | 5 |
| 2018 | Spatially Sparse Code Multiplexing for the Massive MIMO NetworksabstractIn this paper, we investigate a spatially sparse code multiplexing (SCM) transmission scheme for the massive multiple-input multiple-output (MIMO) networks to enhance the access connectivity. We construct a non- orthogonal transmission policy over both power and angle domains to fully utilize the limited angle-domain degree of freedom (DoF). Firstly, the mapping structure in the angle domain and the detection method are presented. Then, we formulate an optimization problem to seek an optimal transmission policy for the proposed SCM framework, where both the design of the mapping matrix and the power allocation are concerned. To simplify the non-convex problem, we solve the problem with three steps. During the first step, we allocate different angle-domain beams for users to obtain a sparse mapping matrix; during the second step, the prime optimization is transformed as a convex power allocation problem. Finally, we pursue a suboptimal transmission strategy for the multiple clusters with iterative power allocation. Simulation results verify that the SCM scheme exhibits significant performance gain in terms of sum rate. Weidong Shao, Shun Zhang 0003, Hongyan Li 0001, Jianpeng Ma 0002, Guangzhe Zhao, Xiushe Zhang |
GLOBECOM | 5 |
| 2018 | Spatial-spectral classification of hyperspectral image via group tensor decomposition
Guangzhe Zhao, Bing Tu, Hongyan Fei, Nanying Li, Xianchang Yang |
Neurocomputing | 1 |
| 2018 | Hyperspectral Image Classification via Superpixel Spectral Metrics RepresentationabstractThis letter proposes a new hyperspectral classification method that fuses superpixel spectral metrics and joint sparse representation (JSR), which is termed as superpixel spectral metrics representation (SSMR). Recently, superpixel segmentation has proven to be a powerful tool to exploit the spatial information of hyperspectral images (HSIs), since the size and shape of each superpixel can be adaptively changed in different structural textures. Moreover, spectral information divergence (SID) has superiority compared to other distance-based similarity measures, particularly when using with a JSR classifier. Taking the aforemementioned advantages into account, superpixel segmentation, SID, and JSR are availably combined to effectively utilize the spectralspatial information of the HSI. The proposed SSMR method includes the following main steps. First, superpixel segmentation is utilized to divide the original map into several superpixels. Second, similarity metric SID among test samples in all superpixels and training samples are calculated. Next, the JSR model is employed to obtain the reconstruction residuals of each class. Then, a regularization parameter λ is introduced to attain balance between JSR and SID. Finally, pixel's label is determined by the minimal total residual. Experimental results on the Indian Pines dataset show better performance than several well-known classification methods. Bing Tu, Wenlan Kuang, Guangzhe Zhao, Hongyan Fei |
IEEE Signal Process. Lett. | 3 |