EDBT 2026 Demo / reviewers in the wild / expert
Pengming Feng
dblp:159/8786
· DBLP profile ↗
30ranked-venue papers
7as first author
16since 2021 · last 2025
0000-0001-5853-8100ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 4 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 3 first-author · 8 since 2021Artificial intelligence and machine learning · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Band Prompting Aided SAR and Multi-Spectral Data Fusion Framework for Local Climate Zone ClassificationabstractLocal climate zone (LCZ) classification is of great value for understanding the complex interactions between urban development and local climate. Recent studies have increasingly focused on the fusion of synthetic aperture radar (SAR) and multi-spectral data to improve LCZ classification performance. However, it remains challenging due to the distinct physical properties of these two types of data and the absence of effective fusion guidance. In this paper, a novel band prompting aided data fusion framework is proposed for LCZ classification, namely BP-LCZ, which utilizes textual prompts associated with band groups to guide the model in learning the physical attributes of different bands and semantics of various categories inherent in SAR and multi-spectral data to augment the fused feature, thus enhancing LCZ classification performance. Specifically, a band group prompting (BGP) strategy is introduced to align the visual representation effectively at the level of band groups, which also facilitates a more adequate extraction of semantic information of different bands with textual information. In addition, a multivariate supervised matrix (MSM) based training strategy is proposed to alleviate the problem of positive and negative sample confusion by completing the supervised information. The experimental results demonstrate the effectiveness and superiority of the proposed data fusion framework. Haiyan Lan, Mingjie Xie, Xuanjia Zhao, Hongning Liu, Pengming Feng, Dongli Xu, Guangjun He, Jian Guan 0001 |
ICASSP | 6 |
| 2025 | Align Your Rhythm: Generating Highly Aligned Dance Poses with Gating-Enhanced Rhythm-Aware Feature RepresentationabstractAutomatically generating natural, diverse and rhythmic human dance movements driven by music is vital for virtual reality and film industries. However, generating dance that naturally follows music remains a challenge, as existing methods lack proper beat alignment and exhibit unnatural motion dynamics. In this paper, we propose Danceba, a novel framework that leverages gating mechanism to enhance rhythm-aware feature representation for music-driven dance generation, which achieves highly aligned dance poses with enhanced rhythmic sensitivity. Specifically, we introduce Phase-Based Rhythm Extraction (PRE) to precisely extract rhythmic information from musical phase data, capitalizing on the intrinsic periodicity and temporal structures of music. Additionally, we propose Temporal-Gated Causal Attention (TGCA) to focus on global rhythmic features, ensuring that dance movements closely follow the musical rhythm. We also introduce Parallel Mamba Motion Modeling (PMMM) architecture to separately model upper and lower body motions along with musical features, thereby improving the naturalness and diversity of generated dance movements. Extensive experiments confirm that Danceba outperforms state-of-the-art methods, achieving significantly better rhythmic alignment and motion diversity. Project page: https://danceba.github.io/ . Congyi Fan, Jian Guan 0001, Xuanjia Zhao, Dongli Xu, Youtian Lin, Pengming Feng, Haiwei Pan |
ICCV | 7 |
| 2025 | Edge-aware Affinity Enhancement for Image Manipulation LocalizationabstractImage manipulation localization (IML) refers to the task of identifying regions in images that have been altered by specific tampering techniques, such as copy-move, splicing, or inpainting. Transformers have been applied to IML tasks due to their ability to model long-range correlations between pixels through the self-attention mechanism. However, the inherent self-attention mechanism may not accurately model or sufficiently enhance these correlations, particularly in detailed edge traces, due to its limitations. In this paper, we propose an Edge-Aware Affinity Enhancement approach for the IML task. Specifically, we introduce an Affinity Regularization Module to establish inter-patch correlations for feature regularization via random walk propagation. Based on the extracted correlation representation, we propose an Edge-Affinity Guidance strategy to further refine the correlation accuracy, particularly in ambiguous edge regions. Extensive experimental results demonstrate that our method outperforms state-of-the-art image manipulation localization techniques in terms of localization accuracy. Tianyi Zhang 0004, Qinglong Lin, Pengming Feng, Rubo Zhang |
ACM Multimedia | 4 |
| 2025 | Airplane State Discrimination From Single-Temporal High-Resolution Remote Sensing ImagesabstractThe absence of temporal information in single-temporal satellite remote sensing images presents a substantial challenge for target state discrimination. In this letter, a pioneering Remote Sensing Airplane State Discrimination Network (RSASDNet) is introduced, by leveraging the relationship between targets and their backgrounds in single-temporal high-resolution remote sensing images. To facilitate the study, we take airplane state discrimination as an example, and a Remote Sensing Airport Panoptic Segmentation with Airplane States Dataset (RSAPS-ASD) is constructed. RSASDNet incorporates two key innovations: 1) a scene knowledge graph generation module that constructs scene knowledge representation by capturing spatial relationships between airplane instances and their surrounding environment (e.g., taxiways and hangars); and 2) a novel graph-image hybrid convolution discrimination module that synergistically integrates structural knowledge and spatial semantic information through dedicated dual-branch learning. The effectiveness of the proposed method is validated using RSAPS-ASD, with experimental results demonstrating that RSASDNet achieves an impressive accuracy of 73.95% in airplane state discrimination. Zizhen Li, Shichao Jin, Guangjun He, Xueliang Zhang 0002, Pengming Feng |
IEEE Geosci. Remote. Sens. Lett. | 5 |
| 2024 | FPN with GMM Based Feature Enhancement Strategy for Object Detection in Remote Sensing ImagesabstractIn the realm of object detection, the age-old challenge of accommodating large variations in target scales, particularly in the intricate domain of remote sensing imagery, has long perplexed computer vision aficionados. Feature Pyramid Network (FPN) family, a widely-used stalwart, strives to tame this scale variation challenge by harmoniously fusing features across different levels. However, this typical feature fusion strategy often leads us astray. Noise introduction and feature smoothing problems due to different semantic information from high/low resolution feature maps, which results in semantic misalignment and inconspicuous gradient discrepancy between targets and background. This, in turn, leads to the difficulty in locating and distinguishing target from complex background in remote sensing images. In this paper, a GMM Feature Enhancement Module (GFEM) is proposed to address the problem by generating and enhancing feature of target with Gaussian Mixture Model (GMM), hence avoiding the gradient smoothing problem. Moreover, we introduce a generic feature fusion network named GFEM-FPN, elevating our approach to the next level. GFEM-FPN extracts multi-scale target enhancement features to enhance the ability of discriminating targets and background. The proposed methods are evaluated on NWPU VHR-10 and DIOR-R datasets, and the outperformance in results verify the effectiveness of the proposed method. Hongning Liu, Pengming Feng, Mingjie Xie, Dongli Xu, Jian Guan 0001, Guangjun He, Rubo Zhang |
ICASSP | 2 |
| 2024 | FastDrag: Manipulate Anything in One StepabstractDrag-based image editing using generative models provides precise control over image contents, enabling users to manipulate anything in an image with a few clicks. However, prevailing methods typically adopt $n$-step iterations for latent semantic optimization to achieve drag-based image editing, which is time-consuming and limits practical applications. In this paper, we introduce a novel one-step drag-based image editing method, i.e., FastDrag, to accelerate the editing process. Central to our approach is a latent warpage function (LWF), which simulates the behavior of a stretched material to adjust the location of individual pixels within the latent space. This innovation achieves one-step latent semantic optimization and hence significantly promotes editing speeds. Meanwhile, null regions emerging after applying LWF are addressed by our proposed bilateral nearest neighbor interpolation (BNNI) strategy. This strategy interpolates these regions using similar features from neighboring areas, thus enhancing semantic integrity. Additionally, a consistency-preserving strategy is introduced to maintain the consistency between the edited and original images by adopting semantic information from the original image, saved as key and value pairs in self-attention module during diffusion inversion, to guide the diffusion sampling. Our FastDrag is validated on the DragBench dataset, demonstrating substantial improvements in processing time over existing methods, while achieving enhanced editing performance. Xuanjia Zhao, Jian Guan 0001, Congyi Fan, Dongli Xu, Youtian Lin, Haiwei Pan, Pengming Feng |
NeurIPS | 7 |
| 2024 | Integration of Global and Local Knowledge for Foreground Enhancing in Weakly Supervised Temporal Action LocalizationabstractWeakly Supervised Temporal Action Localization (WTAL) aims to identify the temporal duration of actions and classify the action categories with only video-level labels in the training stage. Motivated by the intuition that the attention maps generated from various views will assist in enhancing the foreground action temporal segments, in this paper we propose a WTAL pipeline based on a novel attention mechanism that effectively integrates global and local knowledge. Our attention mechanism is mainly composed of a global attention branch and a local attention branch. Specifically, the global attention branch is built on the inter-segment similarity to sparsely mine out the correlation knowledge within the entire video, while the local attention branch is built on the convolutional structure to densely aggregate the information within the fixed local respective field. Experiments on THUMOS14 and ActivityNet v1.3 datasets demonstrate the effectiveness of our proposed WTAL pipeline compared to state-of-the-art methods. Tianyi Zhang 0004, Ronglu Li, Pengming Feng, Rubo Zhang |
IEEE Trans. Multim. | 3 |
| 2023 | Spectral Masked Autoencoder for Few-Shot Hyperspectral Image ClassificationabstractThough deep learning methods have achieved the state-of-the-art performance for hyperspectral image (HSI) classification, they often highly rely on large amount of samples for training, and introduce few-shot challenge due to the lack of labeled samples. In this paper, a self-supervised method is presented for HSI classification in the few-shot scenario, where masked autoencoder is employed to reconstruct the masked bands in spectral domain for model pretraining with limited labeled sample, namely Spectral-MAE. The proposed method not only avoids the overfitting via the pretraining, but also provides the model’s ability for effective feature extraction while avoiding the high spatial redundancy. Experiments conducted verify the effectiveness of the proposed method for HSI classification in few-shot situation as compared with other methods. Pengming Feng, Kaihan Wang, Jian Guan 0001, Guangjun He, Shichao Jin |
IGARSS | 1 |
| 2023 | Polarization-Guided Strategy for Ship Detection in Single-Polarization SAR ImagesabstractHigh-performance target detection algorithms have been proposed for ship detection in Synthetic Aperture Radar (SAR) images in recent years. However, most of them are applied in the situation of single-polarization SAR images and rarely consider the important polarization information in SAR images. In this paper, a polarization-guided strategy is presented to improve the detection performance in single polarization SAR images by predicting polarization type. In which, an extra polarization-guided head is employed following the backbone network to guide the feature maps from the backbone to explore polarization information. Then, the feature map with the polarization information is fused with those from backbone network to enhance feature extraction in feature pyramid networks (FPN), and yields better detection performance. Experiments conducted demonstrate the effectiveness of the proposed method, which can be easily merged with different detectors. Jian Guan 0001, Haotian Yuan 0002, Pengming Feng, Guangjun He, Shichao Jin |
IGARSS | 3 |
| 2023 | Dual-Range Context Aggregation for Efficient Semantic Segmentation in Remote Sensing ImagesabstractAlthough introducing self-attention mechanisms is beneficial to establish long-range dependencies and explore global context information in the task of remote sensing image semantic segmentation, it results in expensive computation and large memory cost. In this letter, we address this dilemma by proposing a lightweight dual-range context aggregation network (LDCANet) for efficient remote sensing image semantic segmentation. First, a dual-range context aggregation module (DCAM) is designed to aggregate the local features and the global semantic context acquired by convolutions and self-attention, respectively, where self-attention is implemented easily by applying two cascaded linear layers to reduce the computational complexity. Furthermore, a simple and lightweight decoder is employed to combine information from different levels, in which a multilayer perceptron (MLP)-based efficient linear block (ELB) is proposed to yield a strong and efficient representation. Experiments conducted on the International Society for Photogrammetry and Remote Sensing (ISPRS) Vaihingen dataset and the Gaofen Image dataset (GID) prove that our LDCANet achieves an excellent trade-off between segmentation accuracy and computational efficiency. In particular, our method achieves 74.12% mean intersection over union (mIoU) on the ISPRS Vaihingen dataset and 61.42% mIoU on the GID with only 4.98-M parameter size. Guangjun He, Pengming Feng, Dilxat Muhtar, Xueliang Zhang 0002 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2023 | EARL: An Elliptical Distribution Aided Adaptive Rotation Label Assignment for Oriented Object Detection in Remote Sensing ImagesabstractLabel assignment is a crucial process in object detection, which significantly influences the detection performance by determining positive or negative samples during training process. However, existing label assignment strategies barely consider the characteristics of targets in remote sensing images (RSIs) thoroughly, e.g., large variations in scales and aspect ratios, leading to insufficient and imbalanced sampling and introducing more low-quality samples, thereby limiting detection performance. To solve the above problems, an Elliptical Distribution aided Adaptive Rotation Label Assignment (EARL) is proposed to select high-quality positive samples adaptively in anchor-free detectors. Specifically, an adaptive scale sampling (ADS) strategy is presented to select samples adaptively among multi-level feature maps according to the scales of targets, which achieves sufficient sampling with more balanced scale-level sample distribution. In addition, a dynamic elliptical distribution aided sampling (DED) strategy is proposed to make the sample distribution more flexible to fit the shapes and orientations of targets, and filter out low-quality samples. Furthermore, a spatial distance weighting (SDW) module is introduced to integrate the adaptive distance weighting into loss function, which makes the detector more focused on the high-quality samples. Extensive experiments on several popular datasets demonstrate the effectiveness and superiority of our proposed EARL, where without bells and whistles, it can be easily applied to different detectors and achieve state-of-the-art performance. The source code will be available at: https://github.com/Justlovesmile/EARL. Jian Guan 0001, Mingjie Xie, Youtian Lin, Guangjun He, Pengming Feng |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Segmentation-Guided Semantic-Aware Self-Supervised Denoising for SAR ImageabstractSynthetic aperture radar (SAR) images often suffer from speckle noise, which can degrade their visual quality and affect downstream applications. Recently, supervised and self-supervised methods have been proposed by virtue of deep learning with synthetic “noisy-clean” image pairs or only real noisy images as training data, respectively. Among these methods, by avoiding artifact problems in real SAR image denoising, self-supervised methods solve the domain gap problem of supervised methods and hence have attracted significant attention. However, existing self-supervised denoising methods essentially rely on pixel information of images and ignore corresponding semantic information, which makes them challenging to remove speckle noise while retaining detailed features. To this end, we propose a segmentation-guided semantic-aware self-supervised denoising method for SAR images, namely SARDeSeg, where a segmentation network is incorporated with a denoising network and guides it to learn and be aware of the semantic information of the input noisy SAR images. Additionally, a wavelet transform-based connector is introduced to efficiently transmit semantic information between the denoising network and the segmentation network, together with an edge-aware smoothing loss to improve speckle noise suppression while preserving edge features. Experimental results demonstrate that the proposed SARDeSeg outperforms state-of-the-art denoising methods for SAR images, particularly in preserving detailed edge features. Ye Yuan 0011, Yanxia Wu 0001, Pengming Feng, Yulei Wu |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Real-time detection of flying aircraft using hyperspectral satellite imgesabstractReal-time positioning of flying aircraft based on hyperspectral remote sensing images is a novel application of satellite remote sensing. The complex aviation environment proposes huge challenges to accurately obtain features of flying aircraft from hyperspectral images (HSI) based on the traditional remote sensing image processing method. This paper proposes an automatic detection method for flying aircraft by Gaofen-5 (GF-5) HSI. This study firstly acquires the candidate target regions by detecting anomalies among the spectral inter-bands based on the imaging characteristics of the GF-5 remote sensing satellite sensor and the features of flying aircraft in the remote sensing image. The flying aircraft is then detected by confirming its features of a linear combination and proportional displacement. The experimental results demonstrate that the detection accuracy of flying aircraft can reach 99.9% under good weather conditions. Furthermore, the proposed method can effectively reduce the computational cost and realize the real-time detection of flying aircraft, which can be used to detect flying targets in large-scale HSI. Guangjun He, Pengming Feng, Shishuo Liu, Shichao Jin |
ISNCC | 3 |
| 2022 | Multiscale Deep Neural Network With Two-Stage Loss for SAR Target Recognition With Small Training SetabstractDeep learning models have been used recently for target recognition from synthetic aperture radar (SAR) images. However, the performance of these models tends to deteriorate when only a small number of training samples are available due to the problem of overfitting. To address this problem, we propose a two-stage multiscale densely connected convolutional neural networks (TMDC-CNNs). In the proposed TMDC-CNNs, the overfitting issue is addressed with a novel multiscale densely connected network architecture and a two-stage loss function, which integrated the cosine similarity with the prevailing softmax cross-entropy loss. Experiments were conducted on the MSTAR data set, and the results show that our model offers significant recognition accuracy improvements as compared with other state-of-the-art methods, with severely limited training data. The source codes are available athttps://github.com/Stubsx/TMDC-CNNs. Jian Guan 0001, Jiabei Liu, Pengming Feng, Wenwu Wang 0001 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | A Practical Solution for SAR Despeckling With Adversarial Learning Generated Speckled-to-Speckled ImagesabstractIn this letter, we aim to address a synthetic aperture radar (SAR) despeckling problem with the necessity of neither clean (speckle-free) SAR images nor independent speckled image pairs from the same scene, and a practical solution for SAR despeckling (PSD) is proposed. First, an adversarial learning framework is designed to generate speckled-to-speckled (S2S) image pairs from the same scene in the situation where only single speckled SAR images are available. Then, the S2S SAR image pairs are employed to train a modified despeckling Nested-UNet model using the Noise2Noise (N2N) strategy. Moreover, an iterative version of the PSD method (PSDi) is also presented. Experiments are conducted on both synthetic speckled and real SAR data to demonstrate the superiority of the proposed methods compared with several state-of-the-art methods. The results show that our methods can reach a good tradeoff between feature preservation and speckle suppression. Ye Yuan 0011, Jian Guan 0001, Pengming Feng, Yanxia Wu 0001 |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2021 | Low-Dimensional Denoising Embedding Transformer for ECG ClassificationabstractThe transformer based model (e.g., FusingTF) has been employed recently for Electrocardiogram (ECG) signal classification. However, the high-dimensional embedding obtained via 1-D convolution and positional encoding can lead to the loss of the signal’s own temporal information and a large amount of training parameters. In this paper, we propose a new method for ECG classification, called low-dimensional denoising embedding transformer (LDTF), which contains two components, i.e., low-dimensional denoising embedding (LDE) and transformer learning. In the LDE component, a low-dimensional representation of the signal is obtained in the time-frequency domain while preserving its own temporal information. And with the low-dimensional embedding, the transformer learning is then used to obtain a deeper and narrower structure with fewer training parameters than that of the FusingTF. Experiments conducted on the MIT-BIH dataset demonstrates the effectiveness and the superior performance of our proposed method, as compared with state-of-the-art methods. Jian Guan 0001, Pengming Feng, Wenwu Wang 0001 |
ICASSP | 3 |
| 2020 | TOSO: Student's-T Distribution Aided One-Stage Orientation Target Detection in Remote Sensing ImagesabstractIn this paper, a robust Student’s-T distribution aided One-Stage Orientation detector, namely TOSO, is proposed to address orientation target detection in remote sensing images. A one-stage keypoint based network architecture is used to avoid the complicated computation caused by rotation anchor boxes and two main contributions are proposed to enhance the performance. Firstly, a novel geometric transformation method is introduced to provide an orientation bounding box from its surrounding horizontal bounding box, so that the orientation angle is achieved by only regressing the geometric transformation parameters. Secondly, the Student’s-t distribution is used as a joint distribution to associate the classification task with the regression task, which are represented as Gaussian and inverse Gamma distributions, respectively. Experiments on two popular remote sensing public datasets DOTA and HRSC2016 confirm the improvement from our proposed TOSO detector. Pengming Feng, Youtian Lin, Jian Guan 0001, Guangjun He, Huifeng Shi, Jonathon A. Chambers |
ICASSP | 1 |
| 2020 | Meta Metric Learning for Highly Imbalanced Aerial Scene ClassificationabstractClass imbalance is an important factor that affects the performance of deep learning models used for remote sensing scene classification. In this paper, we propose a random finetuning meta metric learning model (RF-MML) to address this problem. Derived from episodic training in meta metric learning, a novel strategy is proposed to train the model, which consists of two phases, i.e., random episodic training and all classes fine-tuning. By introducing randomness into the episodic training and integrating it with fine-tuning for all classes, the few-shot meta-learning paradigm can be successfully applied to class imbalanced data to improve the classification performance. Experiments are conducted to demonstrate the effectiveness of the proposed model on class imbalanced datasets, and the results show the superiority of our model, as compared with other state-of-the-art methods. Jian Guan 0001, Jiabei Liu, Pengming Feng, Tong Shuai, Wenwu Wang 0001 |
ICASSP | 4 |
| 2020 | A Dynamic End-to-End Fusion Filter for Local Climate Zone Classification Using SAR and Multi-Spectrum Remote Sensing DataabstractLocal Climate Zone (LCZ) classification is potentially popular because of its extensive applications. Recently, data from different remote sensors including synthetic aperture radar (SAR) and multi-spectrum are employed for LCZ classification. However, different bands in SAR and multi-spectrum are difficult to fuse because of their various physical properties. In this paper, an dynamic end-to-end fusion filter is proposed. Firstly, a convolutional neural network (CNN) based dynamic filter network (DFN) is introduced to integrate different bands in SAR and multi-spectrum data, which enhances the fusion accuracy by a flexible dynamic operation. Then the filter is used for feature extraction, hence improve the performance of the classifier. The proposed method is evaluated using Sentinel-1 and Sentinel-2 dataset and the improvement of accuracy shows the superiority of the proposed dynamic data fusion approach. Pengming Feng, Youtian Lin, Guangjun He, Jian Guan 0001, Huifeng Shi |
IGARSS | 1 |
| 2020 | Association Loss for Visual Object DetectionabstractConvolutional neural network (CNN) is a popular choice for visual object detection where two sub-nets are often used to achieve object classification and localization separately. However, the intrinsic relation between the localization and classification sub-nets was not exploited explicitly for object detection. In this letter, we propose a novel association loss, namely, the proxy squared error (PSE) loss, to entangle the two sub-nets, thus use the dependency between the classification and localization scores obtained from these two sub-nets to improve the detection performance. We evaluate our proposed loss on the MS-COCO dataset and compare it with the loss in a recent baseline, i.e. the fully convolutional one-stage (FCOS) detector. The results show that our method can improve the AP from 33.8 to 35.4 and AP75 from 35.4 to 37.8, as compared with the FCOS baseline. Dongli Xu, Jian Guan 0001, Pengming Feng, Wenwu Wang 0001 |
IEEE Signal Process. Lett. | 3 |
| 2019 | Embranchment Cnn Based Local Climate Zone Classification Using Sar And Multispectral Remote Sensing DataabstractIn this study, a Local Climate Zone (LCZ) classification framework is established using a Densenet based embranchment Convolutional Neural Network (CNN). Both synthetic aperture radar (SAR) and multispectral data are employed for feature fusion, specifically, considering about the difference in imaging mechanism between SAR and multispectral data, features from both resources are extracted in different branches separately according to the physical properties of each band. Significant accuracy improvement can be achieved when evaluate the proposed method by Sentinel-1 and Sentinel-2 dataset, and the comparison results show the superiority of the proposed embranchment CNN framework over the conventional methods. Pengming Feng, Youtian Lin, Jian Guan 0001, Guangjun He, Zhenghuan Xia, Huifeng Shi |
IGARSS | 1 |
| 2019 | An On-Orbit Ship Detection and Classification Algorithm for Sar SatelliteabstractShip detection using synthetic aperture radar (SAR) plays a vital role in the wide area ocean surveillance. Especially the real-time ship detection and classification, significantly promotes the illegally operating ships monitoring performance. In this study, an on-orbit ship detection and classification method is developed for SAR satellite. An adaptive sliding window is developed to extract the water connected domain and propose the suspected ship target areas. The OpenSARShip dataset and manually selected non-ship slice images are adopted to train a deep learning model, which is applied to classify the proposed ship slice images into three types (cargo ship, other ship and false alarm). The results demonstrate the improvement performance of the proposed method over the constant false alarm rate (CFAR) method, where the detection accuracy improved from 88.5% to 98.4% and the false alarm rate mitigated from 11.5% to 0.7% compared with CFAR respectively. Meanwhile, the proposed method can achieved verification and testing accuracy of 97.2% and 93.4% respectively for ship type classification. Huifeng Shi, Guangjun He, Pengming Feng |
IGARSS | 3 |
| 2019 | Joint Convolutional Neural Network for Small-Scale Ship Classification in SAR ImagesabstractShip classification using synthetic aperture radar (SAR) imagery is a challenge problem in maritime surveillance. Because of the scale limitation of ship targets in SAR image, convolutional neural networks (CNNs) can not achieve similar performance as for natural image classification. In this paper, we propose a joint CNNs framework for small-scale ship targets classification in SAR image, where a generator and a classifier are jointly connected. The generator can reconstruct the small-scale low-resolution (LR) images to large-scale super-resolution (SR) images, and the classifier is used for ship classification. A novel joint loss optimization strategy is introduced to solve the problem, where an MSE-based content loss is employed to generate high quality SR images, and a classification loss is applied to enable the generator and the classifier to be trained in a joint way. Experiments are conducted to demonstrate the superior performance of our proposed method, as compared with the state-of-the-art methods. Yanxia Wu 0001, Ye Yuan 0011, Jian Guan 0001, Libo Yin, Pengming Feng |
IGARSS | 7 |
| 2018 | Polynomial dictionary learning algorithms in sparse representations
Jian Guan 0001, Xuan Wang 0002, Pengming Feng, Jing Dong 0001, Jonathon A. Chambers, Zoe Lin Jiang, Wenwu Wang 0001 |
Signal Process. | 3 |
| 2017 | Particle PHD filter based multi-target tracking using discriminative group-structured dictionary learningabstractStructured sparse representation has been recently found to achieve better efficiency and robustness in exploiting the target appearance model in tracking systems with both holistic and local information. Therefore, to better simultaneously discriminate multi-targets from their background, we propose a novel video-based multi-target tracking system that combines the particle probability hypothesis density (PHD) filter with discriminative group-structured dictionary learning. The discriminative dictionary with group structure learned by the hierarchical K-means clustering algorithm implicitly associates the dictionary atoms with the group labels, simultaneously enforcing the target candidates from the same group (class) to share the same structured sparsity pattern. Furthermore, we propose a new joint likelihood calculation by relating the discriminative sparse codes with the maximum voting technique to enhance the particle PHD updating step. Experimental results on two publicly available benchmark video sequences confirm the improved performance of our proposed method over other state-of-the-art techniques in video-based multi-target tracking. Zeyu Fu, Pengming Feng, Syed M. Naqvi, Jonathon A. Chambers |
ICASSP | 2 |
| 2017 | Matrix of Polynomials Model Based Polynomial Dictionary Learning Method for Acoustic Impulse Response ModelingabstractWe study the problem of dictionary learning for signals that can be represented as polynomials or polynomial matrices, such as convolutive signals with time delays or acoustic impulse responses.Recently, we developed a method for polynomial dictionary learning based on the fact that a polynomial matrix can be expressed as a polynomial with matrix coefficients, where the coefficient of the polynomial at each time lag is a scalar matrix.However, a polynomial matrix can be also equally represented as a matrix with polynomial elements.In this paper, we develop an alternative method for learning a polynomial dictionary and a sparse representation method for polynomial signal reconstruction based on this model.The proposed methods can be used directly to operate on the polynomial matrix without having to access its coefficients matrices.We demonstrate the performance of the proposed method for acoustic impulse response modeling. Jian Guan 0001, Xuan Wang 0002, Pengming Feng, Jing Dong 0001, Wenwu Wang 0001 |
INTERSPEECH | 3 |
| 2017 | Social Force Model-Based MCMC-OCSVM Particle PHD Filter for Multiple Human TrackingabstractVideo-based multiple human tracking often involves several challenges, including target number variation, object occlusions, and noise corruption in sensor measurements. In this paper, we propose a novel method to address these challenges based on probability hypothesis density (PHD) filtering with a Markov chain Monte Carlo (MCMC) implementation. More specifically, a novel social force model (SFM) for describing the interaction between the targets is used to calculate the likelihood within the MCMC resampling step in the prediction step of the PHD filter, and a one class support vector machine (OCSVM) is then used in the update step to mitigate the noise in the measurements, where the SVM is trained with features from both color and oriented gradient histograms. The proposed method is evaluated and compared with state-of-the-art techniques using sequences from the CAVIAR, TUD, and PETS2009 datasets based on the mean Euclidean tracking error on each frame, the optimal subpattern assignment metric, and the multiple object tracking precision metric. The results show improved performance of the proposed method over the baseline algorithms, including the traditional particle PHD filtering method, the traditional SFM-based particle filtering method, multi-Bernoulli filtering, and an online-learning-based tracking method. Pengming Feng, Wenwu Wang 0001, Satnam Singh Dlay, Syed M. Naqvi, Jonathon A. Chambers |
IEEE Trans. Multim. | 1 |
| 2016 | Social force model aided robust particle PHD filter for multiple human trackingabstractIn this paper, we propose a novel robust multiple human tracking approach based upon processing a video signal by utilizing a social force model to enhance the particle probability hypothesis density (PHD) filter. In traditional dynamic models, the states of targets are only predicted by their own history; however, in multiple human tracking, the information from interaction between targets and the intentions of each target can be employed to obtain more robust prediction. Furthermore, such information can mitigate the problems of collision and occlusion. The cardinality of variable number of targets can also be estimated by using the PHD filter, hence improving the overall accuracy of the multiple human tracker. In this work, a background subtraction step has also been employed to identify the new born targets and provide the measurement set for the PHD filter. To evaluate tracking performance, sequences from both the CAVIAR and PETS2009 datasets are employed for evaluation, which shows clear improvement of the proposed method over the conventional particle PHD filter. Pengming Feng, Wenwu Wang 0001, Syed M. Naqvi, Satnam Singh Dlay, Jonathon A. Chambers |
ICASSP | 1 |
| 2016 | Adaptive Retrodiction Particle PHD Filter for Multiple Human TrackingabstractThe probability hypothesis density (PHD) filter is well known for addressing the problem of multiple human tracking for a variable number of targets, and the sequential Monte Carlo implementation of the PHD filter, known as the particle PHD filter, can give state estimates with nonlinear and non-Gaussian models. Recently, Mahler et al. have introduced a PHD smoother to gain more accurate estimates for both target states and number. However, as highlighted by Psiaki in the context of a backward-smoothing extended Kalman filter, with a nonlinear state evolution model the approximation error in the backward filtering requires careful consideration. Psiaki suggests that to minimize the aggregated least-squares error over a batch of data. We instead use the term retrodiction PHD filter to describe the backward filtering algorithm in recognition of the approximation error proposed in the original PHD smoother, and we propose an adaptive recursion step to improve the approximation accuracy. This step combines forward and backward processing through the measurement set and thereby mitigates the problems with the original PHD smoother when the target number changes significantly and the targets appear and disappear randomly. Simulation results show the improved performance of the proposed algorithm and its capability in handling a variable number of targets. Pengming Feng, Wenwu Wang 0001, Syed M. Naqvi, Jonathon A. Chambers |
IEEE Signal Process. Lett. | 1 |
| 2015 | An unsupervised acoustic fall detection system using source separation for sound interference suppressionabstractWe present a novel unsupervised fall detection system that employs the collected acoustic signals (footstep sound signals) from an elderly person׳s normal activities to construct a data description model to distinguish falls from non-falls. The measured acoustic signals are initially processed with a source separation (SS) technique to remove the possible interferences from other background sound sources. Mel-frequency cepstral coefficient (MFCC) features are next extracted from the processed signals and used to construct a data description model based on a one class support vector machine (OCSVM) method, which is finally applied to distinguish fall from non-fall sounds. Experiments on a recorded dataset confirm that our proposed fall detection system can achieve better performance, especially with high level of interference from other sound sources, as compared with existing single microphone based methods. Muhammad Salman Khan 0001, Miao Yu 0001, Pengming Feng, Liang Wang 0001, Jonathon A. Chambers |
Signal Process. | 3 |