Ming Tong

dblp:50/7774 · DBLP profile ↗
← Back
18ranked-venue papers
14as first author
7since 2021 · last 2025
0009-0002-7302-652XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 8 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 first-authorSystems, architecture and hardware · 1
YearPublicationVenuePosition
2025 Terrain Segmentation in PolSAR Images via Statistical Learning and Uncertainty Perception
abstract
Recently, deep learning network is introduced to terrain segmentation application in polarimetric synthetic aperture radar (PolSAR) images and achieves remarkable performance. However, interclass and intraclass variety caused by nonlinear characteristic and intrinsic randomness introduced by speckle is still a challenge for segmentation methods. In this article, a segmentation framework based on nonlinear statistical learning is introduced to bridge the gap between terrain segmentation in PolSAR and optical images. Firstly, a learnable nonlinear statistical description module is proposed to represent the nonlinear characteristic introduced by coherent speckle, which supplements the lack of discriminative linear feature in terrain blocks. Secondly, a conditional diffusion pipeline is introduced to model latent distributions and perceive intrinsic randomness that causes intraclass variety. Finally, the proposed framework adopts the statistical features learned from nonlinear statistical description module as condition to align with the process of distribution modeling in the conditional diffusion segmentation pipeline. Extensive experiments are conducted on two authoritative PolSAR terrain segmentation datasets, which present competitive performance.
Ming Tong, Xiaoxiao Fang, Jiu Jiang, Chu He
IEEE Geosci. Remote. Sens. Lett.1
2024 Uncertainty-Driven Multi-scale Feature Fusion Network for Real-Time Image Deraining
Ming Tong
ICIC (4)1
2024 Semi-UFormer: Semi-supervised Uncertainty-aware Transformer for Image Dehazing
abstract
Image dehazing is fundamental yet not well-solved in computer vision. Most cutting-edge models are trained in synthetic data, leading to the poor performance on real-world hazy scenarios. Besides, they commonly give deterministic dehazed images while neglecting to mine their uncertainty. To bridge the domain gap and enhance the dehazing performance, we propose a novel semi-supervised uncertainty-aware transformer network, called Semi-UFormer. Semi-UFormer can well leverage both the real-world hazy images and their uncertainty guidance information. Specifically, Semi-UFormer builds itself on the knowledge distillation framework. Such teacher-student networks effectively absorb real-world haze information for quality dehazing. Furthermore, an uncertainty estimation block is introduced into the model to estimate the pixel uncertainty representations, which is then used as a guidance signal to help the student network produce haze-free images more accurately. Extensive experiments demonstrate that Semi-UFormer generalizes well from synthetic to real-world images.
Ming Tong, Xuefeng Yan 0001, Yongzhen Wang 0001, Mingqiang Wei
IJCNN1
2024 Wavelet Tree Transformer: Multihead Attention With Frequency-Selective Representation and Interaction for Remote Sensing Object Detection
abstract
Vision Transformer has achieved remarkable success in image recognition tasks owing to its global modeling ability. However, the quadratic computational complexity becomes a prominent issue when dealing with high-resolution remote sensing images. Numerous studies have explored the potential of spectral analysis to reveal for more discriminative features. However, neural network exhibit frequency tendency, and different features are interested in different frequencies. Unfortunately, there is no well-established criterion for selecting appropriate frequency representations. To address these issues, a novel wavelet tree head attention (WTHA-ViT) model is proposed which combines a tree structure on the wavelet frequencies with multihead attention in the Transformer encoder, possessing the ability to interact with cross-combinations of short and long-range as well as high and low-frequency components. First, we construct a wavelet tree reduction module (WTRM) based on the wavelet tree structure, utilizing the wavelet decomposition to retain frequency features suitable for each patch, which enables global modeling with various frequency components while reducing computational complexity. Second, guided by channel correlations, we propose the channel lifting scheme multihead attention (CLSMHA) to model the importance on the heads of multihead attention and focus on the more salient head features. Finally, our WTHA-ViT can replace the backbone of detection networks for dense prediction tasks. Extensive experiments on DOTA-V1.0 and HRSID datasets demonstrate that our model exhibits superior performance and robustness compared to state-of-the-art networks. Besides, we evaluate the transferability of the model on DIOR and LEVIR datasets and verify its generalization ability. The code is available athttps://github.com/conquer-pan/WTHA-ViT.
Chu He, Wei Huang 0059, Jidong Cao, Ming Tong
IEEE Trans. Geosci. Remote. Sens.5
2023 A Statistical-Texture Feature Learning Network for PolSAR Image Classification
abstract
Both traditional and deep learning-based methods have limitations in extracting statistical features from Polarimetric Synthetic Aperture Radar (PolSAR) images that contain regions with different levels of heterogeneity. To address this issue, we present a Statistical-Texture feature Learning Network (STLNet) for PolSAR image classification. Our approach includes several strategies. Firstly, we propose a novelNth-order Statistical feature Learning (N-SL) module as the statistical modeling interface to be combined with the network. In addition, we propose a Multi-level high-order Statistical feature Learning (MSL) module based on theN-SL module to represent the statistical characteristics of PolSAR images. Secondly, we propose a Texture feature Learning (TL) module to explore the spatial relationships among pixels and supplement the learned statistical features. Experimental results on E-SAR and AIRSAR datasets demonstrate that the proposed MSL and TL modules can effectively improve classification performance. Furthermore, STLNet outperforms other networks of comparable size.
Chu He, Xiaoxiao Fang, Ming Tong, Bokun He
IEEE Geosci. Remote. Sens. Lett.4
2023 Learning Scattering Similarity and Texture-Based Attention With Convolutional Neural Networks for PolSAR Image Classification
abstract
The varying polarimetric orientation angles (POAs) result in scattering diversity, leading to ambiguity in the interpretation of polarimetric synthetic aperture radar (PolSAR) images. Exploring the scattering characteristics in the polarimetric rotation domain (PRD) and the complementary features can help overcome the ambiguity. To address this, we propose a novel PolSAR image classification algorithm called learning scattering similarity and texture-based attention with convolutional neural networks (LSTCNNs). Three strategies are included in the proposed method. First, a pixel-level scattering similarity learning (SSL) module is proposed to analyze the scattering components of radar targets by learning the mapping from PolSAR data in the PRD to typical scattering models, with rotation angles as learnable parameters to utilize scattering diversity and avoid ambiguity. Second, a neighborhood-level texture-based attention (TA) module is proposed to learn the spatially enhanced features of PolSAR images, with the attention module design guided by the physical meaning of texture and consideration of channel and position importance. Finally, the proposed LSTCNN, which includes the SSL module, the TA module, and the classification module, combines pixel-level scattering features in the PRD and neighborhood-level texture features to increase the discriminability of features. The experimental results on three PolSAR images acquired by airborne SAR (AIRSAR) and experimental SAR (E-SAR) demonstrate the robustness and excellence of LSTCNN.
Chu He, Bokun He, Ming Tong
IEEE Trans. Geosci. Remote. Sens.4
2021 DM-CTSA: a discriminative multi-focused and complementary temporal/spatial attention framework for action recognition
Ming Tong, Kaibo Yan, Xing Yue
Neural Comput. Appl.1
2020 PTL-LTM model for complex action recognition using local-weighted NMF and deep dual-manifold regularized NMF with sparsity constraint
Ming Tong, He Bai 0010, Xing Yue, Haili Bu
Neural Comput. Appl.1
2020 NMF with local constraint and Deep NMF with temporal dependencies constraint for action recognition
Ming Tong, Yiran Chen 0002, He Bai 0010, Xing Yue
Neural Comput. Appl.1
2020 DKD-DAD: a novel framework with discriminative kinematic descriptor and deep attention-pooled descriptor for action recognition
Ming Tong, He Bai 0010, Mengao Zhao
Neural Comput. Appl.1
2019 D3-LND: A two-stream framework with discriminant deep descriptor, linear CMDT and nonlinear KCMDT descriptors for action recognition
Ming Tong, Mengao Zhao, Yiran Chen 0002, Houyi Wang
Neurocomputing1
2019 Non-negative enhanced discriminant matrix factorization method with sparsity regularization
Ming Tong, Haili Bu, Mengao Zhao, Shengnan Xi
Neural Comput. Appl.1
2019 A deep discriminative and robust nonnegative matrix factorization network method with soft label constraint
Ming Tong, Yiran Chen 0002, Mengao Zhao, Haili Bu, Shengnan Xi
Neural Comput. Appl.1
2018 A new framework of action recognition with discriminative parts, spatio-temporal and causal interaction descriptors
Ming Tong, Yiran Chen 0002, Mengao Zhao, Weijuan Tian
J. Vis. Commun. Image Represent.1
2018 A compact discriminant hierarchical clustering approach for action recognition
Ming Tong, Weijuan Tian, Houyi Wang
Multim. Tools Appl.1
2017 Action recognition new framework with robust 3D-TCCHOGAC and 3D-HOOFGAC
Ming Tong, Houyi Wang, Weijuan Tian, Shulin Yang
Multim. Tools Appl.1
2016 Independent detection and self-recovery video authentication mechanism using extended NMF with different sparseness constraints
Ming Tong, Jinyu Guo, Shichang Tao, Yangcheng Wu
Multim. Tools Appl.1
1994 A distributed-memory implementation of the MCHF atomic structure package
Charlotte Froese Fischer, Ming Tong, Murry Bentley, Zuchang Shen, C. Ravimohan
J. Supercomput.2