Kang Ni

dblp:226/5883 · DBLP profile ↗
← Back
24ranked-venue papers
12as first author
20since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 14 · 9 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2026 DenoDet V2: Phase-Amplitude Cross Denoising for SAR Object Detection
abstract
One of the primary challenges in Synthetic Aperture Radar (SAR) object detection lies in the pervasive influence of coherent noise. As a common practice, most existing methods, whether handcrafted approaches or deep learning-based methods, employ the analysis or enhancement of object spatial-domain characteristics to achieve implicit denoising. In this paper, we propose DenoDet V2, which explores a completely novel and different perspective to deconstruct and modulate the features in the transform domain via a carefully designed attention architecture. Compared to DenoDet V1, DenoDet V2 is a major advancement that exploits the complementary nature of amplitude and phase information through a band-wise mutual modulation mechanism, which enables a reciprocal enhancement between phase and amplitude spectra. Extensive experiments on various SAR datasets demonstrate the state-of-the-art performance of DenoDet V2. Notably, DenoDet V2 achieves a significant 0.8% improvement on SARDet-100K dataset compared to DenoDet V1, while reducing the model complexity by half.
Kang Ni, Minrui Zou, Yuxuan Li 0004, Xiang Li 0041, Kehua Guo, Ming-Ming Cheng, Yimian Dai
AAAI1
2025 Learning optical flow from spiking camera with direction disassembly
abstract
Abstract Conventional optical flow estimation methods typically recover two‐dimensional motion from RGB image sequences. Recently, due to the rise and widespread use of spike cameras, learning optical flow from spiking cameras has become a hot topic in the field of two‐dimensional motion estimation. Although existing methods have been designed to learn optical flow by designing feature processing methods for spike streams, there is still insufficient consideration for flow field post‐processing, resulting in limited accuracy of optical flow estimation. To address this problem, an optical flow estimation method based on directional disassembly is proposed. Specifically, the estimated flow fields along the horizontal and vertical directions are disassembled and the motion vectors along the two directions are denoised separately to reduce the burden of post‐processing for complex two‐dimensional motion information. In addition, contextual information is introduced in the post‐processing so that the scene information can effectively contribute to the results of the flow post‐processing. Experimental results show that this proposed method is capable of achieving comparable performance on spike‐based public datasets.
Mingliang Zhai, Xuezhi Xiang, Kang Ni, Hao Gao 0005
IET Image Process.4
2025 Selective Spectral-Spatial Aggregation Transformer for Hyperspectral and LiDAR Classification
abstract
Convolutional neural networks (CNNs) and transformers have achieved excellent classification performances in hyperspectral imagery (HSI) and light detection and ranging (LiDAR) land cover classification. However, for complex land covers, effectively characterizing the contextual information and spectral-spatial interaction features of HSI and LiDAR is crucial for improving classification accuracy. Motivated by this, this letter is dedicated to selective convolutional kernel mechanisms and spectral-spatial interactive transformer feature learning style, proposing a selective spectral-spatial aggregation transformer network, named S2ATNet. A convolution feature selected module (CFSM), which can dynamically capture the contextual features of various land covers, is first utilized in both of HSI and LiDAR branches. Afterward, a cascaded spatial-spectral learning and interactive fusion (CSLIF) block is designed for acquiring the nonlocal spatial-spectral characteristics in an interactive feature learning style. The learned features are fed into the max-average classification head (MACH) to obtain the final classification results. The effectiveness of the proposed S2ATNet is validated on two publicly available datasets. Codes are available athttps://github.com/RSIP-NJUPT/S2ATNet.git.
Kang Ni, Zirun Li, Chunyang Yuan, Zhizhong Zheng, Peng Wang 0030
IEEE Geosci. Remote. Sens. Lett.1
2025 MoCoLSK: Modality-Conditioned High-Resolution Downscaling for Land Surface Temperature
abstract
Land surface temperature (LST) is a critical parameter for environmental studies, but directly obtaining high spatial resolution LST data remains challenging due to the spatiotemporal tradeoff in satellite remote sensing. Guided LST downscaling has emerged as an alternative solution to overcome these limitations, but current methods often neglect spatial nonstationarity, and there is a lack of an open-source ecosystem for deep learning methods. In this article, we propose the modality-conditioned large selective kernel (MoCoLSK) network, a novel architecture that dynamically fuses multimodal data through modality-conditioned projections. MoCoLSK achieves a confluence of dynamic receptive field adjustment and multimodal feature fusion, leading to enhanced LST prediction accuracy. Furthermore, we establish the GrokLST Project, a comprehensive open-source ecosystem featuring the GrokLST dataset, a high-resolution (HR) benchmark, and the GrokLST toolkit, an open-source PyTorch-based toolkit encapsulating MoCoLSK alongside 40+ state-of-the-art approaches. Extensive experimental results validate MoCoLSK’s effectiveness in capturing complex dependencies and subtle variations within multispectral data, outperforming existing methods in LST downscaling. Our code, dataset, and toolkit are available athttps://github.com/GrokCV/GrokLST.
Qun Dai, Chunyang Yuan, Yimian Dai, Yuxuan Li 0004, Xiang Li 0041, Kang Ni, Jianhui Xu, Xiangbo Shu, Jian Yang 0003
IEEE Trans. Geosci. Remote. Sens.6
2025 IDNet: Intensity-Constrained Detail-Enhanced Network for Hyperspectral and LiDAR Collaborative Classification
abstract
The complementarity of hyperspectral image (HSI) and light detection and ranging (LiDAR) data provides significant advantages in land cover classification tasks. Currently, most HSI and LiDAR classification algorithms focus on local-global feature learning frameworks. However, high-frequency characteristics play a vital role in distinguishing different land cover classes. For instance, the high-frequency information in the spectral curve corresponds to spectral features characterized by rapid variations, as well as detailed features (edges and boundaries) in the spatial land cover information. Therefore, embedding high-frequency detail information can enhance the collaborative classification of HSI and LiDAR data. Motivated by this, this paper proposes an Intensity-Constrained Detail-Enhanced Network, named IDNet. This network embeds Laplacian high-frequency features into a Transformer network architecture and designs Transformer blocks with dynamic range awareness and similarity-based spatial-spectral feature grouping interactions to efficiently learn local-global high-frequency spatial-spectral features in the HSI and LiDAR branches. During the feature fusion stage, differential convolution is utilized to construct an enhanced detail expert that emphasizes high-frequency feature information, enabling adaptive dynamic feature fusion to improve the discriminability of the fused features. Experimental results show that compared with the current state-of-the-art algorithms such as S2ENet, HCT-Net and MHST, the proposed IDNet is improved by 0.55%-2.72%, 0.17% -1.31%, 0.41%-4.92%, and 2.58%-13.13% on the Houston 2013, Trento, MUUFL, and Augsburg dataset with OA as the evaluation index, respectively. Code of this project is at https://github.com/ZhaoYuQing01/IDNet.
Kang Ni
IEEE Trans. Geosci. Remote. Sens.4
2025 Coarse-to-Fine High-Order Network for Hyperspectral and LiDAR Classification
abstract
The fusion of hyperspectral and light detection and ranging (LiDAR) data could significantly improve land-cover classification performance. Most existing feature fusion methods focus on “late fusion” or “halfway feature interaction fusion” methods, which treat LiDAR and hyperspectral data as model inputs while overlooking the redundancy between hyperspectral imagery (HSI) and LiDAR features. In addition, the unique characteristics of HSI and LiDAR data, combined with the complexity of land-cover backgrounds, make it challenging to accurately describe their properties. High-order deep features, as a deep statistical representation, could effectively capture the statistical characteristics of these land covers. Based on this, this article focuses on the unique characteristics of HSI and LiDAR data, as well as the distinguishability of features, and designs a progressive hyperspectral and LiDAR collaborative classification method, named coarse-to-fine high-order network (CHNet). In the coarse stage, HSI data redundancy reduction and LiDAR feature reconstruction are performed in either the frequency domain or the spatial domain to ensure the effectiveness of subsequent feature fusion. The fine stage focuses primarily on selective feature fusion and discriminability enhancement, introducing gating mechanisms, deep expert systems, and high-order feature statistics. This approach enhances the discriminability of the fused features while simultaneously reducing their dimensionality. The proposed method, based on a “redundancy removal-feature learning” mechanism, captures more effective deep features by accounting for the different imaging mechanisms of multisource data and the complex background of land covers, ultimately improving the effectiveness of land-cover classification. Experimental results on three public datasets and one self-constructed dataset demonstrate that CHNet achieves superior performance. The code is available athttps://github.com/RSIP-NJUPT/CHNet.
Kang Ni, Yunan Xie, Guofeng Zhao 0004, Zhizhong Zheng, Peng Wang 0030, Tongwei Lu
IEEE Trans. Geosci. Remote. Sens.1
2024 Scene flow estimation from 3D point clouds based on dual-branch implicit neural representations
abstract
Abstract Recently, online optimisation‐based scene flow estimation has attracted significant attention due to its strong domain adaptivity. Although online optimisation‐based methods have made significant advances, the performance is far from satisfactory as only flow priors are considered, neglecting scene priors that are crucial for the representations of dynamic scenes. To address this problem, the authors introduce a dual‐branch MLP‐based architecture to encode implicit scene representations from a source 3D point cloud, which can additionally synthesise a target 3D point cloud. Thus, the mapping function between the source and synthesised target 3D point clouds is established as an extra implicit regulariser to capture scene priors. Moreover, their model infers both flow and scene priors in a stronger bidirectional manner. It can effectively establish spatiotemporal constraints among the synthesised, source, and target 3D point clouds. Experiments on four challenging datasets, including KITTI scene flow, FlyingThings3D, Argoverse, and nuScenes, show that our method can achieve potential and comparable results, proving its effectiveness and generality.
Mingliang Zhai, Kang Ni, Jiucheng Xie, Hao Gao 0005
IET Comput. Vis.2
2024 Multimode Structural Nonconvex Tensor Low-Rank Regularized Hyperspectral Image Destriping and Denoising
abstract
In this paper, we propose an effective multi-mode structural nonconvex tensor low rank (M2SNTLR) regularized hyperspectral image (HSI) destriping and denoising method in a unified framework of tensor representation. Firstly, we exploit and model the tensor low-tubal-rank properties of the HSI and its spectral gradient along spectral mode simultaneously via the tensor tubal rank functions. Secondly, based on the speciality and directionality of stripe noise, we particularly exploit and model the tensor low-tubal-rank properties of stripe noise and its unidirectional vertical gradient along spatial vertical mode simultaneously via also the tensor tubal rank functions. Thirdly, by using the log tensor nuclear norm (logTNN)-based nonconvex surrogate on those tensor tubal rank functions for a closer approximation, we thus propose the logTNN-based multi-mode structural nonconvex tensor low rank priors of HSI and stripe noise. Furthermore, we solve the proposed M2SNTLR model via the alternating direction method of multipliers. At last, extensive experiments validate that the proposed M2SNTLR method performs better denoising results than various low rank-based HSI denoising methods.
Pengfei Liu 0002, Haijian Long, Kang Ni, Zhizhong Zheng
IEEE Geosci. Remote. Sens. Lett.3
2024 Learning graph-based representations for scene flow estimation
Mingliang Zhai, Hao Gao 0005, Ye Liu 0005, Jianhui Nie, Kang Ni
Multim. Tools Appl.5
2024 Remote Sensing Scene Classification via Second-Order Differentiable Token Transformer Network
abstract
The vision transformer has been widely applied in remote sensing image scene classification due to its excellent ability to capture global features. However, remote sensing scene images involve challenges such as scene complexity and small inter-class differences. Directly utilizing the global tokens of transformer for feature learning may increase computational complexity. Therefore, constructing a distinguishable transformer network which adaptively selects tokens can effectively improve the classification performance of remote sensing scene images while considering computational complexity. Based on this, a second-order differentiable token transformer network (SDT2Net) is proposed for considering the efficacy of distinguishable statistical features and non-redundant learnable tokens of remote sensing scene images. A novel transformer block, including an efficient attention block (EAB) and differentiable token compression (DTC) mechanism, is inserted into SDT2Net for acquiring selectable token features of each scene image guided by sparse shift local features and token compression rate learning style. Furthermore, a fast token fusion (FTF) module is developed for acquiring more distinguishable token feature representations. This module utilizes the fast global covariance pooling algorithm to acquire high-order visual tokens and validates the effectiveness of classification tokens and high-order visual tokens for scene classification. Compared with other recent methods, SDT2Net achieves the most advanced performance with comparable FLOP-s (Floating Point Operations Per Second). The code will be available at https://github.com/RSIP-NJUPT/SDT2Net.
Kang Ni, Qianqian Wu 0009, Sichan Li, Zhizhong Zheng, Peng Wang 0030
IEEE Trans. Geosci. Remote. Sens.1
2024 Hyperspectral and LiDAR Classification via Frequency Domain-Based Network
abstract
Local-global feature learning method based on deep learning has significantly improved the collaborative classification of hyperspectral imagery (HSI) and light detection and ranging (LiDAR) data. However, HSI encompasses numerous bands with significant interband correlations. Addressing how to efficiently capture spatial, spectral, and elevation information from hyperspectral and LiDAR data while considering data redundancy and enhancing feature representation of land cover will contribute to enhancing classification effectiveness. Combining frequency feature learning methods with convolutional neural networks (CNNs), transformers, and other architectures to construct an end-to-end feature learning network framework is an effective method. Therefore, this article proposes a frequency domain-based network (FDNet) for the classification of HSI and LiDAR data using a frequency local feature learning framework and a self-attention mechanism based on fast Fourier transform (FFT). FDNet could effectively capture local efficient frequency features of spatial, spectral, and elevation information in HSI and LiDAR data in an adaptive feature learning style, and embedding convolutional offsets into the frequency domain-based transformer network not only enhances local features but also effectively captures global semantic characteristics of land covers while reducing computational complexity. We validated the efficacy of FDNet across three publicly available datasets and a particularly challenging self-constructed dataset, denoted as the Yancheng dataset. The source codes will be available athttps://github.com/RSIP-NJUPT/FDNet.
Kang Ni, Guofeng Zhao 0004, Zhizhong Zheng, Peng Wang 0030
IEEE Trans. Geosci. Remote. Sens.1
2023 Spike-Based Optical Flow Estimation Via Contrastive Learning
abstract
Spiking cameras have shown promising advantages for optical flow estimation in high-speed scenarios. The recent work SCFlow [1] attempts to train an optical flow model using spike frames based on a multi-scale flow reconstruction loss. However, only using the flow reconstruction loss is unable to effectively deal with the details of motion, which may lead to noise and blur in the estimated flow fields. To address this issue, we introduce a contrastive loss into spike-based optical flow estimation, which exploits both the information of positive samples and negative samples. Moreover, we propose a refinement step with flexible reception fields to effectively refine the initial flow fields. Experiments on the spiking optical flow dataset PHM demonstrate that the proposed network is effective for spike-based optical flow estimation. In addition, our method achieves competitive performance compared to recent spike-based, frame-based, and event-based methods.
Mingliang Zhai, Kang Ni, Jiucheng Xie, Hao Gao 0005
ICASSP2
2023 Cross-Modal Optical Flow Estimation via Modality Compensation and Alignment
abstract
Cross-modal optical flow estimation aims to predict motion fields between two frames collected from different modalities, recently attracting intensive attention. However, a substantial yet challenging problem is how to match images across a large modal discrepancy. In this paper, we propose a modality compensation module (MCM) to extract complementary features from different modalities adaptively. Moreover, a cross-modal feature alignment loss is introduced into our network, pulling the compensative features of two cross-modal frames closer and effectively reducing the modal discrepancy. The experimental results demonstrate that our method can achieve competitive performance on the cross-modal optical flow dataset CrossKITTI. Moreover, we experimentally verify that the proposed MCM and cross-modal feature alignment loss are effective for cross-modal optical flow estimation.
Mingliang Zhai, Kang Ni, Jiucheng Xie, Hao Gao 0005
ICASSP2
2023 Learning Scene Flow from 3d Point Clouds with Cross-Transformer and Global Motion Cues
abstract
Scene flow estimation is critical for real-world vision problems such as autonomous driving and augmented reality. Due to the popularity of 3D LiDAR sensors, scene flow estimation from 3D point clouds arouses increasing attention. Existing methods usually use a flow embedding-based layer to find correspondences between point pairs. However, only using a flow embedding-based layer is not enough to model the global mutual relationship between two features due to local matching. In this paper, we introduce a cross-transformer to capture more reliable dependencies for point pairs. Moreover, a global motion-aware module is adopted to learn large displacements with a non-local approach. The experimental results demonstrate that the proposed method achieves comparable performance on public datasets and confirm the effectiveness of exploiting the cross-transformer and global motion cues for scene flow estimation.
Mingliang Zhai, Kang Ni, Jiucheng Xie, Hao Gao 0005
ICASSP2
2023 Scene Flow Estimation from Point Clouds with Contrastive Loss and Dual Pseudo Labels
abstract
Scene flow estimation aims to extract the 3D motion vector between each surface point in two consecutive point clouds. Pseudo-label-based approaches usually exploit point-to-point relations and 3D geometry information to generate the pseudo label for self-supervised learning. However, unreasonable results are still obtained due to the unexploited information of negative samples. Moreover, previous approaches are limited by the fact that pseudo labels are only generated along the forward direction, ignoring the backward direction that has strong spatiotemporal correlations with the forward direction. In this paper, we address these issues in a simple yet effective manner. Specifically, we introduce a contrastive loss to exploit both the information of positive samples and negative samples. Furthermore, we design a dual pseudo labels generation strategy to provide a bidirectional self-supervision for scene flow estimation. Experiments on FlyingThings3D and KITTI datasets show that our method can achieve competitive performance compared to recent self-supervised methods.
Mingliang Zhai, Kang Ni, Jiucheng Xie, Xuezhi Xiang, Hao Gao 0005
ICIP2
2023 DJSPNet: Deep Joint Statistical-Spatial Pooling Network for High-Resolution SAR Image Classification
abstract
The previous approaches based on statistical features or spatial features have achieved promising performance on pixel-wise high-resolution (HR) synthetic aperture radar (SAR) image classification, but these methods always cannot capture local spatial features and global statistical properties efficiently because of the complex spatial structural patterns and statistical nature in SAR patches. Inspired by this, we propose a deep joint statistical–spatial pooling network (DJSPNet), for HR SAR image classification, which combines a group second-order statistical feature learning (GSFL) block and an efficient feature-fusion style (EFS) into an end-to-end feature learning block. GSFL block is designed with a group second-order feature learning method in two steps, where the first step divides convolutional channels into several semantic groups. The second step collects second-order feature statistics by calculating pairwise feature interactions within each group. EFS models second-order attentional statistics between statistical characteristics and spatial features by polynomial kernel approximation and guides the discriminative feature activations in SAR patches. More specifically, both GSFL and EFS are stacked and plugged into the encoder stage of conventional U-Net for distinguishable feature learning. Experimental results suggest that the proposed DJSPNet gives better classification performance compared with related deep feature learning networks on a real TerraSAR-X dataset.
Kang Ni, Mingliang Zhai, Minrui Zou, Qianqian Wu 0009, Peng Wang 0030
IEEE Geosci. Remote. Sens. Lett.1
2022 SAR Image Segmentation Based on Super-Pixel and Kernel-Improved CV Model
abstract
The paper focuses on the effect of pixel intensity random variation in the homogeneous region of Synthetic Aperture Radar (SAR) images. Considering the shortage of CV (Chan-Vese) model to describe the energy variation inside and outside the curve, a SAR image segmentation algorithm based on super-pixel and kernel-improved CV model is proposed. The Simple Linear Iterative Clustering 0 (SLICO) algorithm is employed for pixel reconstruction. Simultaneously, our improved model combines a Laplacian kernel function with 12 fitting energy to construct the novel fitting energy, which can describe the energy inside and outside the curve clearly. Specifically, a distance regularized term is embedded into our proposed model to avoid the reinitialization of level set and accelerate the convergence of the improved model. Finally, experimental results on the Moving and Stationary Target Acquisition and Recognition (MSTAR) data set illustrates our method achieves competitive performances compared with other related algorithms.
Kang Ni
IGARSS1
2022 High-Resolution SAR Image Classification Using Subspace Wavelet Encoding Network
abstract
The feature learning methods based on convolutional neural networks (CNNs) have produced tremendous achievements in high-resolution (HR) synthetic aperture radar (SAR) image classification. However, the inherent speckle noise could weaken the effectiveness of the convolutional feature statistics. To effectively characterize the features of SAR land-covers under speckle noise, we propose a subspace wavelet encoding network (SWENet) trainable end-to-end and based on an encoder–decoder architecture for modeling the robust feature statistics in individual feature subspaces. We introduce a subspace encoder block at the end of the encoder stage and divide the entire feature space into a set of subspaces; the second-order statistics of all subspaces are concatenated. Then, the wavelet pooling block, suppressing the noise and keeping the structures of learned features well, decomposes the features into low-frequency (storing the basic object structures) and high-frequency components by Haar wavelet layer (HWL), and this block reconstructs the processed components using inverse IHWL during the upsampling stage. Especially, the wavelet pooling block is defined in each subspace for powerful feature learning. Experimental results on a TerraSAR-X image classification dataset suggest that our proposed SWENet yields a performance boost over its competitors.
Kang Ni, Pengfei Liu 0002, Peng Wang 0030
IEEE Geosci. Remote. Sens. Lett.1
2022 Subpixel Flood Inundation Mapping Based on Spatial-Spectral Information in Irregular Regions
abstract
Sub-pixel flood inundation mapping (SPFIM) could handle the mixed pixels to obtain the spatial distribution of flood inundation at sub-pixel scale. However, the spatial-spectral information of flood inundation used by SPFIM is usually constructed in a specified rectangular local window, and the number of spectral bands utilized by SPFIM is also little. To solve this problem, this paper proposes SPFIM based on spatial-spectral information in irregular regions (SSIIIR). In SSIIR, the extended random walker algorithm is used to calculate the spatial correlation in irregular regions to obtain spatial information, and the normalized model is constructed to calculate the all spectral bands in irregular regions to yield spectral information. With the help of spatial-spectral information in irregular regions, the flood inundation mapping result is improved. The experimental results show that SSIIIR produces the better flood inundation mapping result than the traditional SPFIM methods.
Peng Wang 0030, Zhongchen He, Kang Ni
IEEE Geosci. Remote. Sens. Lett.4
2022 Subpixel Mapping Based on Multisource Remote Sensing Fusion Data for Land-Cover Classes
abstract
Subpixel mapping (SPM) based on multisource remote sensing fusion data (MRSFD) for land-cover classes, called SPM-MRSFD, is proposed in this letter. First, the original hyperspectral image and the auxiliary panchromatic image are fused to produce the high spatial and spectral resolution fused image by pan-sharpening technology. Second, the fused image with spatial–spectral information and the auxiliary digital surface model (DSM) of light detection and ranging (LiDAR) with elevation information are fused to obtain the MRSFD with spatial–spectral elevation information by feature fusion. Finally, the fractional images with the proportions of subpixels belonging to land-cover classes are derived by unmixing the MRSFD, and the classes labels are allocated to subpixels to obtain the final SPM result according to these proportions’ information. The main contribution of this work is that multiple types of auxiliary information (i.e., spatial–spectral elevation information) from MRSFD is fully utilized and the accuracy of SPM result is improved. Experimental results show that SPM-MRSFD obtains more accurate mapping results than state-of-the-art SPM methods.
Peng Wang 0030, Yulan Wang 0001, Lei Zhang 0110, Kang Ni
IEEE Geosci. Remote. Sens. Lett.4
2020 Synthetic Aperture Radar Scene Classification Using Multiview Cross Correlation Attention Network
abstract
The second-order pooling manner, exploring higher feature statistics than the first-order pooling, has achieved impressive performance in scene classification. However, the object not only presents similarity but also exhibits diversified singularity on the synthetic aperture radar (SAR) image. These make the second-order pooling approaches to explore the single-view second-order feature statistics less adaptable for SAR scene classification. To solve this issue, an end-to-end training framework based on the multiview cross correlation attention network (MCAN) is proposed. The spatial and channelwise self-attention modules are first employed to model the interdependences between the spatial and channel dimensions of their convolutional features. Subsequently, the global spatial and channelwise covariance pooling layers are drawn into the MCAN, and then, they learn the spatial and channel cross correlations within the feature statistics, respectively. Finally, an iterative matrix square-root normalization layer, which owns the capability to fast computing approximate square root of the covariance matrix, is introduced for making the feature representation more discriminative. Experiments on the SAR data set from the TerraSAR-X images for scene classification demonstrate that the MCAN performs better than other related works.
Kang Ni, Peng Wang 0030
IEEE Geosci. Remote. Sens. Lett.1
2019 High-Order Generalized Orderless Pooling Networks for Synthetic-Aperture Radar Scene Classification
abstract
Fixed coding style in bag of visual words (BOVW) model and strong spatial information in convolutional neural network (CNN) feature representation make the feature vector less adaptable for scene classification. With the purpose of extracting the learnable orderless feature for SAR scene classification, the high-order generalized orderless pooling network trained by backpropagation is proposed for learning the high-order vector of locally aggregated descriptors (VLADs) and locality constrained affine subspace coding (LASC), compared with the first-order feature coding style, the proposed network could learn high-order coding features by outer product automatically. Subsequently, for making the feature representation more powerful, the matrix normalization (square root) whose gradients are computed via singular value decomposition (SVD) and elementwise normalization are introduced into the proposed network. Finally, experiments on the SAR scene classification data set from TerraSAR-X image show the proposed networks achieve better performance than the state-of-the-art approaches.
Kang Ni, Peng Wang 0030
IEEE Geosci. Remote. Sens. Lett.1
2019 Synthetic aperture radar river image segmentation using improved localizing region-based active contour model
Kang Ni
Pattern Anal. Appl.1
2018 Adaptive patched L0 gradient minimisation model applied on image smoothing
abstract
L0 gradient minimisation model, one of edge‐aware image smoothing method, also suffers from the stair‐casing effect and images with strong textures cannot be smoothed effectively and weak edges or structures will be smoothed overly. The authors propose a method to overcome these drawbacks above. To begin with, the image is subjected to non‐subsampled shearlet transform to obtain high‐frequency component, and combine all high‐frequency component by maximum local energy rules to obtain the high‐frequency decomposition image, afterwards, introducing the data term associated with high‐frequency decomposition image to keep the similarity of edge and structure between the input and smoothed image. Secondly, the patched L0 gradient minimisation model is presented for improving the description of local information, since different size of the patches has the different texture, exploiting the coefficient of variation to define the size of patch. Finally, defining the adaptive smoothing coefficient based on the gradient to make sure that the smoothing effect of the patch is optimal. The proposed model is applied to image smoothing with desirable results successfully, and the comparisons with other state‐of‐the‐art edge‐preserving image smoothing algorithms demonstrate the great performances of edge‐preserving and texture smoothing.
Kang Ni
IET Image Process.1