Hongbing Ma

dblp:95/1724 · DBLP profile ↗
← Back
38ranked-venue papers
0as first author
29since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 9 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 10 since 2021Artificial intelligence and machine learning · 7 · 7 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Integrating retinex theory with atmospheric scattering model for mixed-degradation image recovery
Yonglong Jiang, Bo Jiang 0001, Jiahe Zhu, Zelong Tan, Zehua Ji, Hongbing Ma
Expert Syst. Appl.6
2026 FSC-MAE: Feature structure coordinated mask autoencoder
Enguang Zuo, Chen Chen 0078, Xiaoyi Lv, Ruishuang Sun, Yinhong Li, Hongbing Ma
Neurocomputing8
2026 PatchFusionMLP: A scalable multi-resolution MLP framework for time series prediction
Xinyu Bi, Xiaoyi Lv, Junyu Zhu, Hongbing Ma, Enguang Zuo
Pattern Recognit.6
2025 The Application of SSB Frequency Offset in Low-Altitude Network
abstract
During the During the National People's Congress and the Chinese Political Consultative Conference in 2024, the ‘low-altitude economy’ was included in the government work report as a significant factor driving new quality productive forces. In the New Radio (NR) network, when a terminal is accessed, the base station uses SSB (Synchronization Signal Block) beam sweeping to detect the optimal beam for the terminal. After the terminal accesses and obtains the configuration information of the reference signal, it feeds back the channel state information (CSI), and the base station uses the optimal beam from CSI-RS (Channel State Information Reference Signal) beam sweeping. In low-altitude communications, 5G antennas flexibly configure the number of beams, considering horizontal and vertical dimensions. Combining SSB frequency offset technology, the SSB frequency points of the low-altitude network can be staggered with the configuration of the ground network, forming a virtual airground heterogeneous frequency network. This approach enhances performance by reducing handover times and interference.
Zixiang Di, Tian Xiao, Zhaoning Wang, Feibi Lv, Hongbing Ma, Jiajia Zhu 0005, Guanghai Liu 0002, Lexi Xu, Xiaomeng Zhu 0001
HPCC7
2025 Efficient Mamba-Attention Network for Remote Sensing Image Super-Resolution
abstract
Lightweight remote sensing image super-resolution (RSISR) methods aim to reconstruct remote sensing images (RSIs) while reducing computational complexity. Previous lightweight model development has primarily focused on the design of convolutional neural networks (CNNs). While CNNs excel at capturing local features, they are limited in establishing long-range dependencies. Mamba, as a model for long-range modeling, has linear computational complexity, making it a viable option for lightweight models. Based on these considerations, this paper proposes an efficient mamba-attention network (EMAN) that can efficiently capture the intricate details and broader semantic information in RSIs. Specifically, we designed a multi-scale detail extraction unit (MDEU) and a multi-dimensional mamba-attention (MDMA). In MDEU, we introduced a multi-scale mechanism and local variance to focus on structural information in RSIs. In MDMA, we integrated spatial expansion and an atrous-based selective scan mechanism to design an efficient scanning method. This method ensures the lightweight nature of the model while establishing global correlations. Additionally, MDMA establishes inter-channel correlations to enhance information exchange. We conducted a comprehensive evaluation of the proposed method on two remote sensing datasets and five benchmark super-resolution (SR) datasets. Extensive experiments demonstrate that our method can achieve superior performance while maintaining a model complexity similar to other lightweight models.
Tianren Wu, Rundong Zhao, Ming Lv, Zhenhong Jia, Liangliang Li 0001, Minqin Liu, Xiaobin Zhao, Hongbing Ma, Gemine Vivone
IEEE Trans. Geosci. Remote. Sens.8
2025 DARI: Transformer-Based Data Augmentation and Rotation Invariance for UAV Person Re-Identification
abstract
The rapid development of Uncrewed Aerial Vehicles (UAVs) and their unique vantage points present both new opportunities and challenges for person Re-Identification (ReID). Uncertain rotations and scale variations of targets in UAV images, coupled with complex environmental factors, hinder existing methods from extracting robust feature representations. Some methods either make minor modifications to the traditional model architecture or apply simple image rotations but still fail to effectively address the challenges of UAV person ReID. To overcome these limitations, we propose a novel Data Augmentation and Rotation Invariance (DARI) algorithm. First, rotation-invariant convolution is introduced to adaptively extract features, mitigating the uncertainty caused by target rotation. Second, a refined data augmentation correction strategy is employed to reduce noise interference by increasing the richness of global features at different stages. Additionally, considering that multiple features of the same identity should yield consistent recognition result, invariant constraints are designed to enhance the clustering effect. We conducted extensive experiments on both UAV and fixed-camera datasets. The results on PRAI-1581 demonstrate a 5.6% and 6.1% improvement in mAP and Rank-1, respectively, compared to baseline. These findings highlight the model’s effectiveness in addressing the challenges of UAV ReID, demonstrating its robustness and superiority.
Fuzeng Zhang, Eksan Firkat, Hongbing Ma, Jihong Zhu 0001, Askar Hamdulla
IEEE Trans. Multim.3
2024 Refining 3D Human Mesh via Model-Free Offsets Estimation
abstract
3D human mesh reconstruction from a single RGB image is a challenging task. Existing methods either utilize parametric mesh models to restrain the 3D human structures or directly regress the 3D coordinates of the mesh vertices. The former ones, called model-based methods, usually fail to recover the high variance of human mesh due to limited capacity of parametric models, whereas the latter ones called model-free methods suffer from unrealistic human structures because of lack of 3D priors. To mitigate the drawbacks of them, we propose that the model-based reconstruction can serve as a good starting point for the model-free refinement, so that the 3D structure priors of the parametric model and the high representation capability of model-free methods can be both inherited. By building a model-free refinement head upon a pretrained model-based regressor, our method reduces the reconstruction errors of 3D human mesh on public datasets H36M and 3DPW, demonstrating the advantage of combining model-based and model-free methods together.
Youze Xue, Hongbing Ma, Huimin Ma 0001
ICASSP3
2024 Dual Guidance Enhancing Camouflaged Object Detection via Focusing Boundary and Localization Representation
abstract
Camouflaged object detection (COD) aims to segment objects that blend into their surrounding environment. However, low-level features in the shallow layers of neural networks, although rich in edge information, often contain a significant amount of redundant information, making it difficult to represent boundary details accurately. On the other hand, deep high-level features retain semantic information for object localization, but the gradual decrease in resolution can introduce biases in representing localization information. To address this issue, we propose a novel boundary and localization representation network (BLR-Net) that guides high-level features to focus on representing localization information while directing low-level features to emphasize boundary details. Firstly, we propose a multi-scale enhanced feature module (MEFM) to capture multi-scale information from backbone features and obtain aggregated feature representations. Next, we propose an extraction boundary module (EBM) that models object boundary features, providing essential boundary information. Subsequently, we introduce a guided learning module (GLM) that utilizes localization features to guide high-level features toward localization representation learning and boundary features to guide low-level features toward boundary representation learning. Finally, we propose a cross-level feature fusion module (CFFM) that aggregates contextual semantic information and gradually fuses multi-level fusion features from the bottom to the top to predict camouflaged objects. Extensive experiments on four benchmark COD datasets demonstrate that BLR-Net outperforms other state-of-the-art COD models.
Zhe Li 0030, Hongbing Ma, Jiabao Sheng
ICME4
2024 Cross Teaching between Single-Spectral and Multi-Spectral Detection Transformers for Remote Sensing Object Detection
abstract
In recent years, to enhance all-weather observation capabilities, remote sensing platforms have been increasingly equipped with thermal infrared (TIR) sensors in addition to visible spectrum (RGB) sensors. Consequently, remote sensing object detection methods started utilizing images from both modalities to improve detection accuracy. However, the inconsistency of target visibility across the two spectrums would introduce confusion of the multi-spectral model, leading to lower detection performance compared to the single-spectral TIR model. In this study, a fine-grained cross-teaching method is proposed to mitigate confusion caused by visibility inconsistency. In detail, leveraging the one-to-one matching mechanism of Detection Transformer (DETR), if a target is visible in both images, knowledge is distilled from a multi-spectral DETR to a TIR-only DETR. Conversely, if an object is only visible in the TIR image, reverse distillation is conducted. Experiments show that cross teaching aids both the single-spectral and multi-spectral models, achieving the state-of-the-art performance on the DroneVehicle dataset.
Jiahe Zhu, Kaiyue Zhou, Shengjin Wang, Hongbing Ma
IGARSS5
2024 FRCE: Transformer-based feature reconstruction and cross-enhancement for occluded person re-identification
Fuzeng Zhang, Hongbing Ma, Jihong Zhu 0001, Askar Hamdulla
Expert Syst. Appl.2
2024 Lightweight Remote Sensing Image Super-Resolution via Background-Based Multiscale Feature Enhancement Network
abstract
In the field of remote sensing image super-resolution (RSISR), most methods based on convolutional neural networks (CNNs) tend to focus on high-weight features in the convolutional kernels, thus overlooking low-weight background features. This bias may result in the neglect of some important information in the background. To address this challenge, we propose a background-based multiscale feature enhancement network (BMFENet), which can extract and supplement missing features from different scale backgrounds to improve the reconstruction of remote sensing images (RSIs). Specifically, we constructed a large kernel feature supplement block (LFSB). The LFSB uses large kernel attention mechanism and multiscale mechanism to expand the receptive field, aggregating global information. Meanwhile, it generates background feature weights to increase the attention to neglected information, thereby reducing the distortion of detailed features. Furthermore, to enhance the nonlinear expression capability of the model, we designed a lattice gated unit (LGU). The LGU removes redundant information through a gating mechanism, efficiently aggregates useful channel information through interchannel interactions and attention mechanisms, and introduces directional convolution to make the model more adaptable to super-resolution (SR) tasks in complex scenes. We validated our method on two remote sensing and four SR benchmark datasets, and the results show that our approach achieves a good balance between performance and complexity.
Tianren Wu, Rundong Zhao, Ming Lv, Zhenhong Jia, Liangliang Li 0001, Zheyuan Wang, Hongbing Ma
IEEE Geosci. Remote. Sens. Lett.7
2024 DIAFNet: A Dynamic Interactive Adaptive Fusion Network Based on Enhanced Differential Features
abstract
Remote Sensing Change Detection (RSCD) seeks to identify areas of interest with changes in photos from multi-temporal remote sensing that are spatially co-registered, thereby monitoring land surface changes. Identifying imbalanced differences between foreground and background categories is crucial when dealing with limited samples and significant interference. Our paper proposes a dynamic interaction and adaptive fusion network (DIAFNet) designed to focus on changes efficiently.DIAFNet utilizes the ResNet18 backbone network for feature extraction. It incorporates the DIAM module to enable the dynamic interaction of features from dual-temporal images, establishing a global feature distribution for autonomous learning of dependencies between diverse features. Additionally, GSConv is introduced to maintain channel correlations, thereby enhancing feature representation. The design of the MFAF module uses the abstract semantic information of deep features to guide the learning process of shallow features, resulting in more precise edge information and comprehensive change areas through adaptive weighted fusion features. Finally, skip connections are employed to minimize fine-grained information loss. Quantitative evaluations on CDD, SYSU-CD, and LEVIR-CD datasets show that our strategy outperforms other state-of-the-art techniques.
Hongbing Ma
IEEE Geosci. Remote. Sens. Lett.2
2023 SUCOLA: Self-adaptive structure refinement unsupervised contrastive learning framework for food safety risk early warning
Enguang Zuo, Junyi Yan, Alimjan Aysa, Chen Chen 0078, Hongbing Ma, Xiaoyi Lv, Kurban Ubul
Eng. Appl. Artif. Intell.6
2023 Infrared Small Target Detection Based on Multidirectional Cumulative Measure
abstract
Robustness of small target detection is a researchable hotspot in infrared surveillance system. The residual phenomenon of background clutter is universal in current local comparison methods. Algorithm of sparse low-rank decomposition restoration cannot be applied to the actual situations due to the long time consumption. This letter proposes a multi-directional cumulative measure (MDCM) to enhance saliency and effectiveness of weak-small target detection. Firstly,multi-directional cumulative mean difference is implemented in central layer and background layer to estimate the background, while multi-directional cumulative derivative multiplying is calculated in central-active layer to characterize overall target’s heterogeneity, then technology of image fusion is adopted to eliminate interference of false target. Finally,a simple adjudicative technology is employed toward separated target region from complex scenes. Compared to up to date existing approaches, extensive simulational testing on four public datasets prove that proposed approach is capable of separating small targets efficiently from an irregular background in a single-scale window and achieve a comparable or even better accuracy.
Guofeng Zhang 0016, Askar Hamdulla, Hongbing Ma
IEEE Geosci. Remote. Sens. Lett.3
2023 A2M: An Amplification-Arbitrary Module for Remote Sensing Image Super-Resolution
abstract
Remote-sensing (RS) image super-resolution (SR) aims to recover high-resolution (HR) images from the corresponding low-resolution (LR) images. In recent years, the SR methods based on convolutional neural networks (CNNs) have achieved incredible performance in case of fixed scale factors (e.g., ×2, ×3, and ×4). However, these methods need to train a single model for each scale factor, and fail to directly reconstruct the HR image of decimal factors. To solve the lack of research on arbitrary scale of RS image SR, we propose a novel amplification module called amplification-arbitrary module (A2M). A2M can be easily embedded in the tail of the previous SR networks, so that the previous networks can also achieve end-to-end arbitrary scale SR. Specifically, we first utilize the combination of convolutional and pixelshuffle layers to zoom in the deep feature matrix 2×, 3×, and 4× along spatial dimension. Information cross transmission (ICT) is then utilized to gather information of multiple spatial sizes. ICT is not only beneficial to enrich the diversity of information, but also can avoid training only a single branch in the training stage. To make better use of multi-scale features, we designed an efficient signal weighting unit (SWU) to generate a correlation matrix at a small cost, and then the signals of multi-scale features at the same position are fused according to the correlation matrix. Experimental results on RS and generic datasets demonstrate that our method with single pre-training model can perform well at any scale factors.
Yuan Xue 0007, Zheyuan Wang, Liangliang Li 0001, Hongbing Ma
IEEE Geosci. Remote. Sens. Lett.4
2023 Infrared Small Target Detection With Patch Tensor Collaborative Sparse and Total Variation Constraint
abstract
Sparse and low-rank modeling has shown the powerful describing abilities to express small targets; however, low-rank model exists the problem of insufficient rank approximation deviation ability and excessive shrinkage, which will lead to inaccurate background estimation. In this letter, a new nonconvex approximation function using the Gaussian model is built toward deeply excavating low-rank information of the background as much as possible. In the sparse collaborative representation, the local prior confidence (LPC) is integrated into the target structure tensor to maximum to distinguish target region and background edge more accurately. In the process of background restoration, total variation constraint (TVC) is employed apropos of describing gray variation of small targets in complex backgrounds more precisely and improving the accuracy of background restoration to further perfectly recover small targets. The low-rank and sparse recovery algorithm engages alternating direction multiplier method (ADMM) for iterative calculation and solution. Compared with advanced optimal algorithms, a great number of experimental results show that the proposed model improves the adaptability and robustness of the detection algorithm to a variety of complex scenes, and has lasting vitality and high-application value.
Guofeng Zhang 0016, Askar Hamdulla, Hongbing Ma
IEEE Geosci. Remote. Sens. Lett.3
2023 Transformer Based Remote Sensing Object Detection With Enhanced Multispectral Feature Extraction
abstract
As a convention, satellites and drones are equipped with sensors of both the visible light spectrum and the infrared (IR) spectrum. However, existing remote sensing object detection methods mostly use RGB images captured by the visible light camera while ignoring IR images. Even for algorithms that take RGB-IR image pairs as input, they may fail to extract all potential features in both spectrums. This letter proposes Multispectral DETR, a remote sensing object detector based on the deformable attention mechanism. To enhance multispectral feature extraction and attention, DropSpectrum and SwitchSpectrum methods are further proposed. DropSpectrum facilitates the extraction of multispectral features by requiring the model to detection some of the targets with only one spectrum. SwitchSpectrum eliminates the level bias caused by the fixed order of RGB-IR feature maps and enhances attention on multispectral features. Experiments on the VEDAI dataset show the state-of-the-art performance of Multispectral DETR and the effectiveness of both DropSpectrum and SwitchSpectrum.
Jiahe Zhu, Huan Zhang 0013, Zelong Tan, Shengjin Wang, Hongbing Ma
IEEE Geosci. Remote. Sens. Lett.6
2022 3D Human Mesh Reconstruction by Learning to Sample Joint Adaptive Tokens for Transformers
abstract
Reconstructing 3D human mesh from a single RGB image is a challenging task due to the inherent depth ambiguity. Researchers commonly use convolutional neural networks to extract features and then apply spatial aggregation on the feature maps to explore the embedded 3D cues in the 2D image. Recently, two methods of spatial aggregation, the transformers and the spatial attention, are adopted to achieve the state-of-the-art performance, whereas they both have limitations. The use of transformers helps modelling long-term dependency across different joints whereas the grid tokens are not adaptive for the positions and shapes of human joints in different images. On the contrary, the spatial attention focuses on joint-specific features. However, the non-local information of the body is ignored by the concentrated attention maps. To address these issues, we propose a Learnable Sampling module to generate joint adaptive tokens and then use transformers to aggregate global information. Feature vectors are sampled accordingly from the feature maps to form the tokens of different joints. The sampling weights are predicted by a learnable network so that the model can learn to sample joint-related features adaptively. Our adaptive tokens are explicitly correlated with human joints, so that more effective modeling of global dependency among different human joints can be achieved. To validate the effectiveness of our method, we conduct experiments on several popular datasets including Human3.6M and 3DPW. Our method achieves lower reconstruction errors in terms of both the vertex-based metric and the joint-based metric compared to previous state of the arts. The codes and the trained models are released at https://github.com/thuxyz19/Learnable-Sampling.
Youze Xue, Jiansheng Chen 0001, Yudong Zhang 0008, Huimin Ma 0001, Hongbing Ma
ACM Multimedia6
2022 Automatic Association of Cross-Domain Network Topology
abstract
Future networks are towards autonomous, with a high level of automatic and intelligent abilities. There are several domains and layers in operator networks, malfunctions can be transmitted from lower layers to upper layers, and from one domain to another domain. At present, cross-domain network malfunctions are mainly relied on the operation and maintenance staff of each professional network to analyze and dispatch orders, resulting in repeated orders and increased human cost. The first and important step of malfunction diagnosis is the construction of network topology. However, cross-domain network topology cannot be associated automatically at present. Based on the performance data, a new method using AI technologies is proposed in this paper, which can associate the connecting cross-domain network ports automatically. The principle is that a same time sequence similarity is shared by the connected ports. Taking the data from real networks and comparing with the existing topology, the connecting relations can be 100% correctly recognized. This method can be widely used to any cross-domain networks, without changing current network equipment.
Sai Han, Guangquan Wang, Qiukeng Fang, Hongbing Ma, Lexi Xu
TrustCom5
2022 Attribute assisted teacher-critical training strategies for image captioning
Jiansheng Chen 0001, Huimin Ma 0001, Hongbing Ma, Wanli Ouyang
Neurocomputing4
2022 Deep Unsupervised Learning for Joint Antenna Selection and Hybrid Beamforming
abstract
In this paper, we propose a novel deep unsupervised learning-based approach that jointly optimizes antenna selection and hybrid beamforming to improve the hardware and spectral efficiencies of massive multiple-input-multiple-output (MIMO) downlink systems. By employing ResNet to extract features from the channel matrices, two neural networks, i.e., the antenna selection network (ASNet) and the hybrid beamforming network (BFNet), are respectively proposed for dynamic antenna selection and hybrid beamformer design. Furthermore, a deep probabilistic subsampling trick and a specially designed quantization function are respectively developed for ASNet and BFNet to preserve the differentiability while embedding discrete constraints into the network structures. With the aid of a flexibly designed loss function, ASNet and BFNet are jointly trained in a phased unsupervised way, which avoids the prohibitive computational cost of acquiring training labels in supervised learning. Simulation results demonstrate the advantage of the proposed approach over conventional optimization-based algorithms in terms of both the achieved rate and the computational complexity.
Zhiyan Liu, Yuwen Yang, Feifei Gao 0001, Hongbing Ma
IEEE Trans. Commun.5
2022 Hyperspectral Estimation of Soil Copper Concentration Based on Improved TabNet Model in the Eastern Junggar Coalfield
abstract
China is the largest coal consumer in the world. The massive exploitation and utilization of coal resources has resulted in serious problems of heavy metal pollution and environmental contamination, such as soil degradation, water pollution, crop damage, and even threatening human lives. Therefore, monitoring soil heavy metal pollution quickly and in real time is an urgent task at present. This research not only formulated a new preprocessing method enlightened by few-shot learning for soil hyperspectral data, but also combined it with other soil-related auxiliary information to extract effective information from the soil hyperspectrum, at the end of which different regression methods were adopted to predict soil heavy metal contamination. This test used 168 actual soil samples from the Eastern Junggar coalfield in Xinjiang for verification. Since copper in the soil is a trace element and the corresponding spectral characteristics are affected by other impurities, improper use of hyperspectral preprocessing methods may introduce interference information or may delete useful information, which makes the model effect unsatisfied. To effectively address the above problems, the preprocessing method of this experiment combined the second-order differential derivation, data enhancement method together with the addition of auxiliary information to allow more effective features to be entered into the model. Next, the Attentive Interpretable Tabular Learning (TabNet) model was improved in three different ways using the original TabNet model and three improved TabNet models to create regression models. One of the improved TabNet models had the best effect, with a list of the top 30 features according to the degree of importance. Meanwhile, the regression prediction of Cu content using four different convolutional neural networks (CNN) revealed that the model with the residual block was the strongest and slightly outperformed the improved TabNet model, but lacked interpretation of the input data. Besides, this experiment also employed different pre-processing methods for regression prediction on various models, and found that the traditional pre-processing methods performed best in traditional regression models (e.g., PLSR) and underperformed in deep learning models. The selected optimal model was compared with partial least square regression (PLSR), and convolutional neural network (CNN) models. The results indicated that both the improved TabNet model and improved CNN model had better performance using the new preprocessing approach proposed in this paper, with improved TabNet yielding a coefficient of determination (R2), root mean square error (RMSE) and ratio of performance to interquartile range (RPIQ) of 0.94, 1.341 and 4.474, respectively. The improved CNN model had a coefficient of determination of 0.942, a root mean square error of 1.324 and an interquartile range of 4.531 in the test dataset.
Yuan Wang 0034, Abdugheni Abliz, Hongbing Ma, Li Liu 0002, Alishir Kurban, Ümüt Halik, Matti Pietikäinen
IEEE Trans. Geosci. Remote. Sens.3
2022 FeNet: Feature Enhancement Network for Lightweight Remote-Sensing Image Super-Resolution
abstract
In the field of remote sensing, due to memory consumption and computational burden, the single-image super-resolution (SISR) methods based on deep convolution neural networks (CNNs) are limited in practical application. To address this problem, we propose a lightweight feature enhancement network (FeNet) for accurate remote-sensing image super-resolution (SR). Considering the existence of equipment with extremely poor hardware facilities, we further design a lighter FeNet-baseline with about 158K parameters. Specifically, inspired by lattice structure, we construct a lightweight lattice block (LLB) as a nonlinear feature extraction function to improve the expression ability. Here, channel separation operation makes the upper and lower branches of the LLB only responsible for half of the features, and the weight coefficients calculated through the attention mechanism enable the upper and lower branches to communicate efficiently. Based on LLB, the feature enhancement block (FEB) is designed in a nested manner to obtain expressive features, where different layers are responsible for the features with different texture richness, and then features from different layers are sequentially fused from deep to shallow. Model parameters and multi-adds operations are used to evaluate network complexity, and extensive experiments on two remote-sensing and four SR benchmark test datasets show that our methods can achieve a good tradeoff between complexity and performance. Our code will be available athttps://github.com/wangzheyuan-666/FeNet.
Zheyuan Wang, Liangliang Li 0001, Yuan Xue 0007, Chenchen Jiang, Kaipeng Sun, Hongbing Ma
IEEE Trans. Geosci. Remote. Sens.7
2022 Boosting Monocular 3D Human Pose Estimation With Part Aware Attention
abstract
Monocular 3D human pose estimation is challenging due to depth ambiguity. Convolution-based and Graph-Convolution-based methods have been developed to extract 3D information from temporal cues in motion videos. Typically, in the lifting-based methods, most recent works adopt the transformer to model the temporal relationship of 2D keypoint sequences. These previous works usually consider all the joints of a skeleton as a whole and then calculate the temporal attention based on the overall characteristics of the skeleton. Nevertheless, the human skeleton exhibits obvious part-wise inconsistency of motion patterns. It is therefore more appropriate to consider each part's temporal behaviors separately. To deal with such part-wise motion inconsistency, we propose the Part Aware Temporal Attention module to extract the temporal dependency of each part separately. Moreover, the conventional attention mechanism in 3D pose estimation usually calculates attention within a short time interval. This indicates that only the correlation within the temporal context is considered. Whereas, we find that the part-wise structure of the human skeleton is repeating across different periods, actions, and even subjects. Therefore, the part-wise correlation at a distance can be utilized to further boost 3D pose estimation. We thus propose the Part Aware Dictionary Attention module to calculate the attention for the part-wise features of input in a dictionary, which contains multiple 3D skeletons sampled from the training set. Extensive experimental results show that our proposed part aware attention mechanism helps a transformer-based model to achieve state-of-the-art 3D pose estimation performance on two widely used public datasets. The codes and the trained models are released at https://github.com/thuxyz19/3D-HPE-PAA.
Youze Xue, Jiansheng Chen 0001, Xiangming Gu, Huimin Ma 0001, Hongbing Ma
IEEE Trans. Image Process.5
2021 Semantic Tag Augmented XlanV Model for Video Captioning
abstract
The key of video captioning is to leverage the cross-modal information from both vision and language perspectives. We propose to leverage the semantic tags to bridge the gap between these modalities rather than directly concatenating or attending to the visual and linguistic features as the previous works. The semantic tags are the object tags and the action tags detected in videos, which can be viewed as partial captions for the input video. To effectively exploit the semantic tags, we design a Semantic Tag augmented XlanV (ST-XlanV) model which encodes 4 kinds of visual and semantic features with X-Linear Attention based cross-attention modules. Moreover, tag related tasks are also designed in the pre-training stage to aid the model more fruitfully exploits the cross-modal information. The proposed model reaches the 5th place in the pre-training for video captioning challenge with the help of the semantic tags. Our codes will be available at: https://github.com/RubickH/ST-XlanV.
Hongwei Xue, Huimin Ma 0001, Hongbing Ma
ACM Multimedia5
2021 Improving Adversarial Robustness of Detector via Objectness Regularization
Jiayu Bao, Hongbing Ma, Huimin Ma 0001
PRCV (4)3
2021 Face Anti-spoofing Based on Cooperative Pose Analysis
Poyu Lin, Huimin Ma 0001, Hongbing Ma
PRCV (3)5
2021 A novel multiscale transform decomposition based multi-focus image fusion framework
Liangliang Li 0001, Hongbing Ma, Zhenhong Jia, Yujuan Si
Multim. Tools Appl.2
2021 Single annotated pixel based weakly supervised semantic segmentation under driving scenes
Xi Li 0010, Huimin Ma 0001, Yanxian Chen, Hongbing Ma
Pattern Recognit.5
2020 A novel approach for multi-focus image fusion based on SF-PAPCNN and ISML in NSST domain
Liangliang Li 0001, Yujuan Si, Linli Wang, Zhenhong Jia, Hongbing Ma
Multim. Tools Appl.5
2018 An Adaptive Patch Prior for Single Image Blind Deblurring
abstract
The blind deblurring algorithm aims to restore the blur kernel and sharp image from a degraded image with blurry and noisy artifacts. In this paper, we propose a novel adaptive patch prior model based on local statistics as a constraint term for blur kernel recovery. With this prior, our approach can rebuild the step edge of a patch and enhance low-level features (edges, corners and junctions) by strengthening the guidance to help sharpen edges and texture structures for latent image restoration. Note that our prior is a nonparametric model that does not rely on external statistical image knowledge and only depends on internal patch information for adaptive computation. Moreover, our proposed prior has the ability to alleviate noise and oversharpening artifacts caused by heuristic methods. Experiments on two benchmark datasets and a natural image showed that our approach compares favorably with other state-of-the-art methods for kernel estimation.
Yongde Guo, Hongbing Ma
ICIP2
2017 Fast Classification of Hyperspectral Images Using Globally Regularized Archetypal Representation With Approximate Solution
abstract
Representation learning plays a crucial rule in pattern recognition. Recently, sparse representation (SR) has become a popular technique in high-dimensional signal processing. In this paper, we provide an alternative archetypal representation (AR) to conduct the representation learning. Compared with SR, AR preserves the sparsity of the learned representation, but has a lower complexity and better interpretation. For the classification of hyperspectral image, spatial information plays an important role in improving the classification performance. Instead of representing each sample individually, we propose a globally regularized AR model, which uses graph regularization to include the contextual information into the learned representation. Although the entire model is convex, it is time consuming to obtain the optimal solution, which limits its large-scale applications. To address the computational issue, we further propose an efficient approximate solver based on line search on a feasible solution trajectory. It turns out that the approximate solution is very close the optimal solution, with a relative error less than 5% usually, but speeds up the optimization dozens of times. Experiments demonstrate that the learned representation with the approximate solution is discriminant comparable to the optimal solution. Using the learned representation as a high-level feature, a linear support vector machine works effectively to produce a high-accuracy classification.
Ding Ni, Hongbing Ma
IEEE Trans. Geosci. Remote. Sens.2
2016 A novel generalized assignment framework for the classification of hyperspectral image
abstract
Recently, sparse representation based classification has been widely used in pattern recognition. Most of existing methods exploit the recovered representation coefficients to reconstruct the inputs, and the classwise reconstruction errors are used to identify the class of the sample based on the subspace assumption. Different from the reconstruction pipeline, an assignment framework is built on the representation coefficients in this paper. More specifically, we treat the representation coefficients as soft assignments of the class labels, and the distribution of the assignments reveals the class of the sample. Under this framework, we can easily generalize it to multi-sample and/or multi-feature scenarios, where multiple assignment instances can be directly fused to stabilize the distribution estimation. As such, the estimated distribution pattern can be used as a new discriminative feature for classification. Experiments on the classification of hyperspectral image demonstrate that the generalized assignment framework can effectively combine neighboring samples and multiple features for collaborative classification, which could achieve significantly better results than several state-of-the-arts.
Ding Ni, Hongbing Ma
ICASSP2
2016 Path-based image sequence interpolation guided by feature points
abstract
We present a method of image sequence interpolation, which can generate a sequence of continuous intermediate frames between two input images. This method is based on a path framework that describes the motion information in the images. A path which starts from one input image, and ends at another input image is constructed for each pixel in the images. The main contribution of this paper is that we take the feature points into consideration. By calculating the position deviation out of the feature points, information and guidance can be given to the process of path optimization, making the interpolation result more plausible and natural. We also increase the conditions and restrictions in the optimization procedure, hence the time and memory cost can be effectively decreased.
Yizhou Fan, Nobuki Yoda, Takeo Igarashi, Hongbing Ma
ICIP4
2015 A sample set perspective on the classification of hyperspectral image with weighted affine constraint
abstract
Hyperspectral images (HSIs) provide abundant spectral information of land covers, allowing detailed analysis of the materials on the earth. Besides, the homogeneity of land covers' distribution, that neighboring pixels usually belong to the same class, is also an important spatial information in HSI analysis. In this paper, a novel classification method in a sample set perspective is proposed to jointly exploit the spectral and the spatial information. More specifically, the proposed method treats the neighboring pixels as a sample set from the same class, and then classifies the set in the affine subspace spanned by the samples of the set using sparse representation based classification. In order to reduce the impact of possible outliers in the set, a weighted affine constraint is enforced on the combination coefficients. With this remedy, the proposed method is not only effective in exploiting the complementary information among the neighboring pixels of the same class, but also robust to possible outliers caused by neighbors of other classes. Experiments on two popular benchmarks demonstrate that our algorithm outperforms several state-of-the-art approaches for HSI classification with limited training samples.
Ding Ni, Hongbing Ma
ICIP2
2015 Classification of Hyperspectral Image Based on Sparse Representation in Tangent Space
abstract
In many real-world problems, data always lie in a low-dimensional manifold. Exploiting the manifold can greatly enhance the discrimination between different categories. In this letter, we propose a classification framework based on sparse representation to directly exploit the underlying manifold. Specifically, using the tangent plane to approximate the local manifold of each test sample, the proposed method classifies the sample by sparse representation in tangent space. Unlike several existing sparse-representation-based classification methods, which sparsely represent the test sample itself, the proposed method sparsely represents the local manifold of the test sample by tangent plane approximation. Therefore, it goes beyond the sample itself and is more robust to kinds of variations confronted in hyperspectral image (HSI) such as illustration differences and spectrum mixing. Experimental results show that the proposed algorithm outperforms several state-of-the-art methods for the classification of HSI with limited training samples.
Ding Ni, Hongbing Ma
IEEE Geosci. Remote. Sens. Lett.2
2015 Hyperspectral Image Classification via Sparse Code Histogram
abstract
Sparse representation-based classifier and its variants have been widely adopted for hyperspectral image (HSI) classification recently. However, sparse representation is unstable so that similar features might obtain significantly different sparse codes. Despite the instability, we find that the sparse codes follow a class-dependent distribution under the structured dictionary consisting of training samples from all classes. Based on this observation, a novel discriminative feature, sparse code histogram (SCH), is developed for HSI classification. By counting the SCH of each sample from the sparse codes of its spatial neighbors, we can statistically obtain the distribution pattern of sparse codes of the class to which the sample belongs, and then treat the SCH as a new feature for classification. To reduce the possible outliers among the neighbors, a shape-adaptive neighborhood extractor is also employed to enhance the stability of the histogram feature. Experimental results demonstrate that SCH enjoys a strong discriminative power, which can achieve notably better performance than several state-of-the-art methods for HSI classification with limited training samples.
Ding Ni, Hongbing Ma
IEEE Geosci. Remote. Sens. Lett.2
2013 Super-Resolution Based on Compressive Sensing and Structural Self-Similarity for Remote Sensing Images
abstract
A super-resolution (SR) method based on compressive sensing (CS), structural self-similarity (SSSIM), and dictionary learning is proposed for reconstructing remote sensing images. This method aims to identify a dictionary that represents high resolution (HR) image patches in a sparse manner. Extra information from similar structures which often exist in remote sensing images can be introduced into the dictionary, thereby enabling an HR image to be reconstructed using the dictionary in the CS framework. We use the K-Singular Value Decomposition method to obtain the dictionary and the orthogonal matching pursuit method to derive sparse representation coefficients. To evaluate the effectiveness of the proposed method, we also define a new SSSIM index, which reflects the extent of SSSIM in an image. The most significant difference between the proposed method and traditional sample-based SR methods is that the proposed method uses only a low-resolution image and its own interpolated image instead of other HR images in a database. We simulate the degradation mechanism of a uniform 2 × 2 blur kernel plus a downsampling by a factor of 2 in our experiments. Comparative experimental results with several image-quality-assessment indexes show that the proposed method performs better in terms of the SR effectivity and time efficiency. In addition, the SSSIM index is strongly positively correlated with the SR quality.
Zongxu Pan, Jing Yu 0005, Huijuan Huang 0001, Shaoxing Hu, Aiwu Zhang, Hongbing Ma
IEEE Trans. Geosci. Remote. Sens.6