Jian Guan 0001

dblp:58/2489-1 · DBLP profile ↗
← Back
48ranked-venue papers
12as first author
28since 2021 · last 2026
0000-0002-0945-1081ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 29 · 9 first-author · 15 since 2021Artificial intelligence and machine learning · 11 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 3 first-author · 6 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Classifier retraining with decoupled federated learning for imbalanced medical image classification
Chunling Chen, Haiwei Pan, Kejia Zhang 0001, Jian Guan 0001
Pattern Recognit.5
2025 Federated Prototype-Aware Pseudo-Labeling for Semi-Supervised Medical Image Classification
abstract
Federated semi-supervised learning (FSSL) enables collaborative training on distributed medical data while preserving privacy, but faces challenges from data heterogeneity and class imbalance. These issues degrade pseudo-labeling quality and introduce confirmation bias. To overcome these limitations, this paper proposes FedPPL, a novel framework for federated medical image classification. FedPPL comprises two key components: Prototype-Aware Thresholding (PAT), which adaptively adjusts pseudo-labeling thresholds using global to mitigate confirmation bias, and Prototype Contrastive Learning (PCL), which enhances feature discriminability to boost accuracy. Experiments on FedISIC2019 and MedMNIST demonstrate that FedPPL achieves more robust and balanced performance than state-of-the-art methods, proving its potential for building reliable and privacy-preserving diagnostic models.
Haiwei Pan, Chunling Chen, Kejia Zhang 0001, Jian Guan 0001
BIBM5
2025 Graph-Enhanced Dual-Stream Feature Fusion with Pre-Trained Model for Acoustic Traffic Monitoring
abstract
Microphone array techniques are widely used in sound source localization and smart city acoustic-based traffic monitoring, but these applications face significant challenges due to the scarcity of labeled real-world traffic audio data and the complexity and diversity of application scenarios. The DCASE Challenge’s Task 10 focuses on using multi-channel audio signals to count vehicles (cars or commercial vehicles) and identify their directions (left-to-right or vice versa). In this paper, we propose a graph-enhanced dual-stream feature fusion network (GEDF-Net) for acoustic traffic monitoring, which simultaneously considers vehicle type and direction to improve detection. We propose a graph-enhanced dual-stream feature fusion strategy which consists of a vehicle type feature extraction (VTFE) branch, a vehicle direction feature extraction (VDFE) branch, and a frame-level feature fusion module to combine the type and direction feature for enhanced performance. A pre-trained model (PANNs) is used in the VTFE branch to mitigate data scarcity and enhance the type features, followed by a graph attention mechanism to exploit temporal relationships and highlight important audio events within these features. The frame-level fusion of direction and type features enables fine-grained feature representation, resulting in better detection performance. Experiments demonstrate the effectiveness of our proposed method. GEDF-Net is our submission that achieved 1st place in the DCASE 2024 Challenge Task 10.
Shitong Fan, Feiyang Xiao, Shuhan Qi, Qiaoxi Zhu, Wenwu Wang 0001, Jian Guan 0001
ICASSP7
2025 Disentangling Hierarchical Features for Anomalous Sound Detection Under Domain Shift
abstract
Anomalous sound detection (ASD) encounters difficulties with domain shift, where the sounds of machines in target domains differ significantly from those in source domains due to varying operating conditions. Existing methods typically employ domain classifiers to enhance detection performance, but they often overlook the influence of domain-unrelated information. This oversight can hinder the model’s ability to clearly distinguish between domains, thereby weakening its capacity to differentiate normal from abnormal sounds. In this paper, we propose a Gradient Reversal-based Hierarchical feature Disentanglement (GRHD) method to address the above challenge. GRHD uses gradient reversal to separate domain-related features from domain-unrelated ones, resulting in more robust feature representations. Additionally, the method employs a hierarchical structure to guide the learning of fine-grained, domain-specific features by leveraging available metadata, such as section IDs and machine sound attributes. Experimental results on the DCASE 2022 Challenge Task 2 dataset demonstrate that the proposed method significantly improves ASD performance under domain shift.
Jian Guan 0001, Jiantong Tian, Qiaoxi Zhu, Feiyang Xiao, Hejing Zhang, Xubo Liu 0001
ICASSP1
2025 Band Prompting Aided SAR and Multi-Spectral Data Fusion Framework for Local Climate Zone Classification
abstract
Local climate zone (LCZ) classification is of great value for understanding the complex interactions between urban development and local climate. Recent studies have increasingly focused on the fusion of synthetic aperture radar (SAR) and multi-spectral data to improve LCZ classification performance. However, it remains challenging due to the distinct physical properties of these two types of data and the absence of effective fusion guidance. In this paper, a novel band prompting aided data fusion framework is proposed for LCZ classification, namely BP-LCZ, which utilizes textual prompts associated with band groups to guide the model in learning the physical attributes of different bands and semantics of various categories inherent in SAR and multi-spectral data to augment the fused feature, thus enhancing LCZ classification performance. Specifically, a band group prompting (BGP) strategy is introduced to align the visual representation effectively at the level of band groups, which also facilitates a more adequate extraction of semantic information of different bands with textual information. In addition, a multivariate supervised matrix (MSM) based training strategy is proposed to alleviate the problem of positive and negative sample confusion by completing the supervised information. The experimental results demonstrate the effectiveness and superiority of the proposed data fusion framework.
Haiyan Lan, Mingjie Xie, Xuanjia Zhao, Hongning Liu, Pengming Feng, Dongli Xu, Guangjun He, Jian Guan 0001
ICASSP9
2025 Attacking Voice Anonymization Systems with Augmented Feature and Speaker Identity Difference
abstract
This study focuses on the First VoicePrivacy Attacker Challenge within the ICASSP 2025 Signal Processing Grand Challenge, which aims to develop speaker verification systems capable of determining whether two anonymized speech signals are from the same speaker. However, differences between feature distributions of original and anonymized speech complicate this task. To address this challenge, we propose an attacker system that combines Data Augmentation enhanced feature representation and Speaker Identity Difference enhanced classifier to improve verification performance, termed DA-SID. Specifically, data augmentation strategies (i.e., data fusion and SpecAugment) are utilized to mitigate feature distribution gaps, while probabilistic linear discriminant analysis (PLDA) is employed to further enhance speaker identity difference. Our system significantly outperforms the baseline, demonstrating exceptional effectiveness and robustness against various voice anonymization systems, ultimately securing a top-5 ranking in the challenge.
Yanzhe Zhang 0001, Zhonghao Bi, Feiyang Xiao, Xuefeng Yang, Qiaoxi Zhu, Jian Guan 0001
ICASSP6
2025 Align Your Rhythm: Generating Highly Aligned Dance Poses with Gating-Enhanced Rhythm-Aware Feature Representation
abstract
Automatically generating natural, diverse and rhythmic human dance movements driven by music is vital for virtual reality and film industries. However, generating dance that naturally follows music remains a challenge, as existing methods lack proper beat alignment and exhibit unnatural motion dynamics. In this paper, we propose Danceba, a novel framework that leverages gating mechanism to enhance rhythm-aware feature representation for music-driven dance generation, which achieves highly aligned dance poses with enhanced rhythmic sensitivity. Specifically, we introduce Phase-Based Rhythm Extraction (PRE) to precisely extract rhythmic information from musical phase data, capitalizing on the intrinsic periodicity and temporal structures of music. Additionally, we propose Temporal-Gated Causal Attention (TGCA) to focus on global rhythmic features, ensuring that dance movements closely follow the musical rhythm. We also introduce Parallel Mamba Motion Modeling (PMMM) architecture to separately model upper and lower body motions along with musical features, thereby improving the naturalness and diversity of generated dance movements. Extensive experiments confirm that Danceba outperforms state-of-the-art methods, achieving significantly better rhythmic alignment and motion diversity. Project page: https://danceba.github.io/ .
Congyi Fan, Jian Guan 0001, Xuanjia Zhao, Dongli Xu, Youtian Lin, Pengming Feng, Haiwei Pan
ICCV2
2025 Preference Aware Item Cold-Start Recommendation With Hierarchical Item Alignment
abstract
Existing cold-start recommendation methods typically use item-level alignment strategies to align the content feature and collaborative feature of warm items during model training. However, these methods are less effective for cold items with low semantic similarity to the warm items when they first appear in the test stage, as they have no historical interactions to obtain the collaborative feature. In this paper, we propose a preference aware recommendation (PARec) model with hierarchical item alignment to solve the item cold-start issue. Our approach exploits user preference from historical records to achieve group-level alignment with item content feature, enhancing recommendation performance. Specifically, our hierarchical item alignment strategy improves recommendations for both high and low similarity cold items by using item-level alignment for high similarity cold items and introducing group-level alignment for low similarity cold items. Low similarity cold items can be successfully recommended through relationships among items, captured by our group-level alignment, based on their co-occurrence possibilities and semantic similarities. For model training, a hierarchical contrastive objective function is presented to balance the performance of warm and cold items, achieving better overall performance. Extensive experiments demonstrate the effectiveness of our method, with results showing its superiority compared to state-of-the-art approaches.
Ben Chen 0004, Bingquan Liu, Lili Shan, Chengjie Sun, Qian Chen 0028, Feiyang Xiao, Jian Guan 0001
IEEE Trans. Knowl. Data Eng.8
2024 Hierarchical Metadata Information Constrained Self-Supervised Learning for Anomalous Sound Detection under Domain Shift
abstract
Self-supervised learning methods have achieved promising performance for anomalous sound detection (ASD) under domain shift by incorporating the metadata of domain shift types and machine sound attributes in feature learning. However, the relation between domain shifts and machine sound attributes has yet to be fully utilised despite their potential benefits for characterising domain shifts. This paper presents a hierarchical metadata information constrained self-supervised ASD method, where the hierarchical relation between domain shift types (section IDs) and attributes is constructed and used as constraints to improve feature representation. In addition, we propose an attribute-group-centre based method for calculating the anomaly score under the domain shift condition. Experiments show improved audio feature learning over the state-of-the-art methods in DCASE 2022 challenge Task 2.
Haiyan Lan, Qiaoxi Zhu, Jian Guan 0001, Yuming Wei, Wenwu Wang 0001
ICASSP3
2024 FPN with GMM Based Feature Enhancement Strategy for Object Detection in Remote Sensing Images
abstract
In the realm of object detection, the age-old challenge of accommodating large variations in target scales, particularly in the intricate domain of remote sensing imagery, has long perplexed computer vision aficionados. Feature Pyramid Network (FPN) family, a widely-used stalwart, strives to tame this scale variation challenge by harmoniously fusing features across different levels. However, this typical feature fusion strategy often leads us astray. Noise introduction and feature smoothing problems due to different semantic information from high/low resolution feature maps, which results in semantic misalignment and inconspicuous gradient discrepancy between targets and background. This, in turn, leads to the difficulty in locating and distinguishing target from complex background in remote sensing images. In this paper, a GMM Feature Enhancement Module (GFEM) is proposed to address the problem by generating and enhancing feature of target with Gaussian Mixture Model (GMM), hence avoiding the gradient smoothing problem. Moreover, we introduce a generic feature fusion network named GFEM-FPN, elevating our approach to the next level. GFEM-FPN extracts multi-scale target enhancement features to enhance the ability of discriminating targets and background. The proposed methods are evaluated on NWPU VHR-10 and DIOR-R datasets, and the outperformance in results verify the effectiveness of the proposed method.
Hongning Liu, Pengming Feng, Mingjie Xie, Dongli Xu, Jian Guan 0001, Guangjun He, Rubo Zhang
ICASSP5
2024 First-Shot Unsupervised Anomalous Sound Detection with Unknown Anomalies Estimated by Metadata-Assisted Audio Generation
abstract
First-shot (FS) unsupervised anomalous sound detection (ASD) is a brand-new task introduced in DCASE 2023 Challenge Task 2, where the anomalous sounds for the target machine types are unseen in training. Existing methods often rely on the availability of normal and abnormal sound data from the target machines. However, due to the lack of anomalous sound data for the target machine types, it becomes challenging when adapting the existing ASD methods to the first-shot task. In this paper, we propose a new framework for the first-shot unsupervised ASD, where metadata-assisted audio generation is used to estimate unknown anomalies, by utilising the available machine information (i.e., metadata and sound data) to fine-tune a text-to-audio generation model for generating the anomalous sounds that contain unique acoustic characteristics accounting for each different machine type. We then use the method of Time-Weighted Frequency domain audio Representation with Gaussian Mixture Model (TWFRGMM) as the backbone to achieve the first-shot unsupervised ASD. Our proposed FS-TWFR-GMM method achieves competitive performance amongst top systems in DCASE 2023 Challenge Task 2, while requiring only 1% model parameters for detection, as validated in our experiments.
Hejing Zhang, Qiaoxi Zhu, Jian Guan 0001, Haohe Liu, Feiyang Xiao, Jiantong Tian, Xinhao Mei, Xubo Liu 0001, Wenwu Wang 0001
ICASSP3
2024 FastDrag: Manipulate Anything in One Step
abstract
Drag-based image editing using generative models provides precise control over image contents, enabling users to manipulate anything in an image with a few clicks. However, prevailing methods typically adopt $n$-step iterations for latent semantic optimization to achieve drag-based image editing, which is time-consuming and limits practical applications. In this paper, we introduce a novel one-step drag-based image editing method, i.e., FastDrag, to accelerate the editing process. Central to our approach is a latent warpage function (LWF), which simulates the behavior of a stretched material to adjust the location of individual pixels within the latent space. This innovation achieves one-step latent semantic optimization and hence significantly promotes editing speeds. Meanwhile, null regions emerging after applying LWF are addressed by our proposed bilateral nearest neighbor interpolation (BNNI) strategy. This strategy interpolates these regions using similar features from neighboring areas, thus enhancing semantic integrity. Additionally, a consistency-preserving strategy is introduced to maintain the consistency between the edited and original images by adopting semantic information from the original image, saved as key and value pairs in self-attention module during diffusion inversion, to guide the diffusion sampling. Our FastDrag is validated on the DragBench dataset, demonstrating substantial improvements in processing time over existing methods, while achieving enhanced editing performance.
Xuanjia Zhao, Jian Guan 0001, Congyi Fan, Dongli Xu, Youtian Lin, Haiwei Pan, Pengming Feng
NeurIPS2
2023 Time-Weighted Frequency Domain Audio Representation with GMM Estimator for Anomalous Sound Detection
abstract
Although deep learning is the mainstream method in unsupervised anomalous sound detection, Gaussian Mixture Model (GMM) with statistical audio frequency representation as input can achieve comparable results with much lower model complexity and fewer parameters. Existing statistical frequency representations, e.g. the log-Mel spectrogram’s average or maximum over time, do not always work well for different machines. This paper presents Time-Weighted Frequency Domain Representation (TWFR) with the GMM method (TWFR-GMM) for anomalous sound detection. The TWFR is a generalized statistical frequency domain representation that can adapt to different machine types, using the global weighted ranking pooling over time-domain. This allows GMM estimator to recognize anomalies, even under domain-shift conditions, as visualized with a Mahalanobis distance-based metric. Experiments on DCASE 2022 Challenge Task2 dataset show that our method has better detection performance than recent deep learning methods. TWFR-GMM is the core of our submission that achieved the 3rd place in DCASE 2022 Challenge Task2.
Jian Guan 0001, Youde Liu, Qiaoxi Zhu, Tieran Zheng, Jiqing Han 0001, Wenwu Wang 0001
ICASSP1
2023 Anomalous Sound Detection Using Audio Representation with Machine ID Based Contrastive Learning Pretraining
abstract
Existing contrastive learning methods for anomalous sound detection refine the audio representation of each audio sample by using the contrast between the samples’ augmentations (e.g., with time or frequency masking). However, they might be biased by the augmented data, due to the lack of physical properties of machine sound, thereby limiting the detection performance. This paper uses contrastive learning to refine audio representations for each machine ID, rather than for each audio sample. The proposed two-stage method uses contrastive learning to pretrain the audio representation model by incorporating machine ID and a self-supervised ID classifier to fine-tune the learnt model, while enhancing the relation between audio features from the same ID. Experiments show that our method outperforms the state-of-the-art methods using contrastive learning or self-supervised classification in overall anomaly detection performance and stability on DCASE 2020 Challenge Task2 dataset.
Jian Guan 0001, Feiyang Xiao, Youde Liu, Qiaoxi Zhu, Wenwu Wang 0001
ICASSP1
2023 Spectral Masked Autoencoder for Few-Shot Hyperspectral Image Classification
abstract
Though deep learning methods have achieved the state-of-the-art performance for hyperspectral image (HSI) classification, they often highly rely on large amount of samples for training, and introduce few-shot challenge due to the lack of labeled samples. In this paper, a self-supervised method is presented for HSI classification in the few-shot scenario, where masked autoencoder is employed to reconstruct the masked bands in spectral domain for model pretraining with limited labeled sample, namely Spectral-MAE. The proposed method not only avoids the overfitting via the pretraining, but also provides the model’s ability for effective feature extraction while avoiding the high spatial redundancy. Experiments conducted verify the effectiveness of the proposed method for HSI classification in few-shot situation as compared with other methods.
Pengming Feng, Kaihan Wang, Jian Guan 0001, Guangjun He, Shichao Jin
IGARSS3
2023 Polarization-Guided Strategy for Ship Detection in Single-Polarization SAR Images
abstract
High-performance target detection algorithms have been proposed for ship detection in Synthetic Aperture Radar (SAR) images in recent years. However, most of them are applied in the situation of single-polarization SAR images and rarely consider the important polarization information in SAR images. In this paper, a polarization-guided strategy is presented to improve the detection performance in single polarization SAR images by predicting polarization type. In which, an extra polarization-guided head is employed following the backbone network to guide the feature maps from the backbone to explore polarization information. Then, the feature map with the polarization information is fused with those from backbone network to enhance feature extraction in feature pyramid networks (FPN), and yields better detection performance. Experiments conducted demonstrate the effectiveness of the proposed method, which can be easily merged with different detectors.
Jian Guan 0001, Haotian Yuan 0002, Pengming Feng, Guangjun He, Shichao Jin
IGARSS1
2023 Anomalous Sound Detection Using Self-Attention-Based Frequency Pattern Analysis of Machine Sounds
Hejing Zhang, Jian Guan 0001, Qiaoxi Zhu, Feiyang Xiao, Youde Liu
INTERSPEECH2
2023 Graph Attention for Automated Audio Captioning
abstract
State-of-the-art audio captioning methods typically use the encoder-decoder structure with pretrained audio neural networks (PANNs) as encoders for feature extraction. However, the convolution operation used in PANNs is limited in capturing the long-time dependencies within an audio signal, thereby leading to potential performance degradation in audio captioning. This letter presents a novel method using graph attention (GraphAC) for encoder-decoder based audio captioning. In the encoder, a graph attention module is introduced after the PANNs to learn contextual association (i.e. the dependency among the audio features over different time frames) through an adjacency graph, and a top-kmask is used to mitigate the interference from noisy nodes. The learnt contextual association leads to a more effective feature representation with feature node aggregation. As a result, the decoder can predict important semantic information about the acoustic scene and events based on the contextual associations learned from the audio signal. Experimental results show that GraphAC outperforms the state-of-the-art methods with PANNs as the encoders, thanks to the incorporation of the graph attention module into the encoder for capturing the long-time dependencies within the audio signal. The source code is available at https://github.com/LittleFlyingSheep/GraphAC.
Feiyang Xiao, Jian Guan 0001, Qiaoxi Zhu, Wenwu Wang 0001
IEEE Signal Process. Lett.2
2023 EARL: An Elliptical Distribution Aided Adaptive Rotation Label Assignment for Oriented Object Detection in Remote Sensing Images
abstract
Label assignment is a crucial process in object detection, which significantly influences the detection performance by determining positive or negative samples during training process. However, existing label assignment strategies barely consider the characteristics of targets in remote sensing images (RSIs) thoroughly, e.g., large variations in scales and aspect ratios, leading to insufficient and imbalanced sampling and introducing more low-quality samples, thereby limiting detection performance. To solve the above problems, an Elliptical Distribution aided Adaptive Rotation Label Assignment (EARL) is proposed to select high-quality positive samples adaptively in anchor-free detectors. Specifically, an adaptive scale sampling (ADS) strategy is presented to select samples adaptively among multi-level feature maps according to the scales of targets, which achieves sufficient sampling with more balanced scale-level sample distribution. In addition, a dynamic elliptical distribution aided sampling (DED) strategy is proposed to make the sample distribution more flexible to fit the shapes and orientations of targets, and filter out low-quality samples. Furthermore, a spatial distance weighting (SDW) module is introduced to integrate the adaptive distance weighting into loss function, which makes the detector more focused on the high-quality samples. Extensive experiments on several popular datasets demonstrate the effectiveness and superiority of our proposed EARL, where without bells and whistles, it can be easily applied to different detectors and achieve state-of-the-art performance. The source code will be available at: https://github.com/Justlovesmile/EARL.
Jian Guan 0001, Mingjie Xie, Youtian Lin, Guangjun He, Pengming Feng
IEEE Trans. Geosci. Remote. Sens.1
2022 Anomalous Sound Detection Using Spectral-Temporal Information Fusion
abstract
Unsupervised anomalous sound detection aims to detect unknown abnormal sounds of machines from normal sounds. However, the state-of-the-art approaches are not always stable and perform dramatically differently even for machines of the same type, making it impractical for general applications. This paper proposes a spectral-temporal fusion based self-supervised method to model the feature of the normal sound, which improves the stability and performance consistency in detection of anomalous sounds from individual machines, even of the same type. Experiments on the DCASE 2020 Challenge Task 2 dataset show that the proposed method achieved 81.39%, 83.48%, 98.22% and 98.83% in terms of the minimum AUC (worst-case detection performance amongst individuals) in four types of real machines (fan, pump, slider and valve), respectively, giving 31.79%, 17.78%, 10.42% and 21.13% improvement compared to the state-of-the-art method, i.e., Glow_Aff. Moreover, the proposed method has improved AUC (average performance of individuals) for all the types of machines in the dataset. The source codes are available at https://github.com/liuyoude/STgram_MFN
Youde Liu, Jian Guan 0001, Qiaoxi Zhu, Wenwu Wang 0001
ICASSP2
2022 Multiscale Deep Neural Network With Two-Stage Loss for SAR Target Recognition With Small Training Set
abstract
Deep learning models have been used recently for target recognition from synthetic aperture radar (SAR) images. However, the performance of these models tends to deteriorate when only a small number of training samples are available due to the problem of overfitting. To address this problem, we propose a two-stage multiscale densely connected convolutional neural networks (TMDC-CNNs). In the proposed TMDC-CNNs, the overfitting issue is addressed with a novel multiscale densely connected network architecture and a two-stage loss function, which integrated the cosine similarity with the prevailing softmax cross-entropy loss. Experiments were conducted on the MSTAR data set, and the results show that our model offers significant recognition accuracy improvements as compared with other state-of-the-art methods, with severely limited training data. The source codes are available athttps://github.com/Stubsx/TMDC-CNNs.
Jian Guan 0001, Jiabei Liu, Pengming Feng, Wenwu Wang 0001
IEEE Geosci. Remote. Sens. Lett.1
2022 A Practical Solution for SAR Despeckling With Adversarial Learning Generated Speckled-to-Speckled Images
abstract
In this letter, we aim to address a synthetic aperture radar (SAR) despeckling problem with the necessity of neither clean (speckle-free) SAR images nor independent speckled image pairs from the same scene, and a practical solution for SAR despeckling (PSD) is proposed. First, an adversarial learning framework is designed to generate speckled-to-speckled (S2S) image pairs from the same scene in the situation where only single speckled SAR images are available. Then, the S2S SAR image pairs are employed to train a modified despeckling Nested-UNet model using the Noise2Noise (N2N) strategy. Moreover, an iterative version of the PSD method (PSDi) is also presented. Experiments are conducted on both synthetic speckled and real SAR data to demonstrate the superiority of the proposed methods compared with several state-of-the-art methods. The results show that our methods can reach a good tradeoff between feature preservation and speckle suppression.
Ye Yuan 0011, Jian Guan 0001, Pengming Feng, Yanxia Wu 0001
IEEE Geosci. Remote. Sens. Lett.2
2022 Distributed Analysis Dictionary Learning Using a Diffusion Strategy
Jing Dong 0001, Liu Yang 0021, Chang Liu 0152, Xiaoqing Luo, Jian Guan 0001
Neural Process. Lett.5
2022 Local Information Assisted Attention-Free Decoder for Audio Captioning
abstract
Automated audio captioning aims to describe audio data with captions using natural language. Existing methods often employ an encoder-decoder structure, where the attention-based decoder (e.g., Transformer decoder) is widely used and achieves state-of-the-art performance. Although this method effectively captures global information within audio data via the self-attention mechanism, it may ignore the event with short time duration, due to its limitation in capturing local information in an audio signal, leading to inaccurate prediction of captions. To address this issue, we propose a method using the pretrained audio neural networks (PANNs) as the encoder and local information assisted attention-free Transformer (LocalAFT) as the decoder. The novelty of our method is in the proposal of the LocalAFT decoder, which allows local information within an audio signal to be captured while retaining the global information. This enables the events of different duration, including short duration, to be captured for more precise caption generation. Experiments show that our method outperforms the state-of-the-art methods in Task 6 of the DCASE 2021 Challenge with the standard attention-based decoder for caption generation.
Feiyang Xiao, Jian Guan 0001, Haiyan Lan, Qiaoxi Zhu, Wenwu Wang 0001
IEEE Signal Process. Lett.2
2021 Low-Dimensional Denoising Embedding Transformer for ECG Classification
abstract
The transformer based model (e.g., FusingTF) has been employed recently for Electrocardiogram (ECG) signal classification. However, the high-dimensional embedding obtained via 1-D convolution and positional encoding can lead to the loss of the signal’s own temporal information and a large amount of training parameters. In this paper, we propose a new method for ECG classification, called low-dimensional denoising embedding transformer (LDTF), which contains two components, i.e., low-dimensional denoising embedding (LDE) and transformer learning. In the LDE component, a low-dimensional representation of the signal is obtained in the time-frequency domain while preserving its own temporal information. And with the low-dimensional embedding, the transformer learning is then used to obtain a deeper and narrower structure with fewer training parameters than that of the FusingTF. Experiments conducted on the MIT-BIH dataset demonstrates the effectiveness and the superior performance of our proposed method, as compared with state-of-the-art methods.
Jian Guan 0001, Pengming Feng, Wenwu Wang 0001
ICASSP1
2021 Robust Estimator for NLOS Error Mitigation in TOA-Based Localization
Jing Dong 0001, Xiaoqing Luo, Jian Guan 0001
WASA (3)3
2021 WRGPruner: A new model pruning solution for tiny salient object detection
Fengwei Jia, Xuan Wang 0002, Jian Guan 0001, Huale Li, Chen Qiu 0003, Shuhan Qi
Image Vis. Comput.3
2021 ARank: Toward specific model pruning via advantage rank for multiple salient objects detection
Fengwei Jia, Xuan Wang 0002, Jian Guan 0001, Huale Li, Chen Qiu 0003, Shuhan Qi
Image Vis. Comput.3
2020 TOSO: Student's-T Distribution Aided One-Stage Orientation Target Detection in Remote Sensing Images
abstract
In this paper, a robust Student’s-T distribution aided One-Stage Orientation detector, namely TOSO, is proposed to address orientation target detection in remote sensing images. A one-stage keypoint based network architecture is used to avoid the complicated computation caused by rotation anchor boxes and two main contributions are proposed to enhance the performance. Firstly, a novel geometric transformation method is introduced to provide an orientation bounding box from its surrounding horizontal bounding box, so that the orientation angle is achieved by only regressing the geometric transformation parameters. Secondly, the Student’s-t distribution is used as a joint distribution to associate the classification task with the regression task, which are represented as Gaussian and inverse Gamma distributions, respectively. Experiments on two popular remote sensing public datasets DOTA and HRSC2016 confirm the improvement from our proposed TOSO detector.
Pengming Feng, Youtian Lin, Jian Guan 0001, Guangjun He, Huifeng Shi, Jonathon A. Chambers
ICASSP3
2020 Meta Metric Learning for Highly Imbalanced Aerial Scene Classification
abstract
Class imbalance is an important factor that affects the performance of deep learning models used for remote sensing scene classification. In this paper, we propose a random finetuning meta metric learning model (RF-MML) to address this problem. Derived from episodic training in meta metric learning, a novel strategy is proposed to train the model, which consists of two phases, i.e., random episodic training and all classes fine-tuning. By introducing randomness into the episodic training and integrating it with fine-tuning for all classes, the few-shot meta-learning paradigm can be successfully applied to class imbalanced data to improve the classification performance. Experiments are conducted to demonstrate the effectiveness of the proposed model on class imbalanced datasets, and the results show the superiority of our model, as compared with other state-of-the-art methods.
Jian Guan 0001, Jiabei Liu, Pengming Feng, Tong Shuai, Wenwu Wang 0001
ICASSP1
2020 A Dynamic End-to-End Fusion Filter for Local Climate Zone Classification Using SAR and Multi-Spectrum Remote Sensing Data
abstract
Local Climate Zone (LCZ) classification is potentially popular because of its extensive applications. Recently, data from different remote sensors including synthetic aperture radar (SAR) and multi-spectrum are employed for LCZ classification. However, different bands in SAR and multi-spectrum are difficult to fuse because of their various physical properties. In this paper, an dynamic end-to-end fusion filter is proposed. Firstly, a convolutional neural network (CNN) based dynamic filter network (DFN) is introduced to integrate different bands in SAR and multi-spectrum data, which enhances the fusion accuracy by a flexible dynamic operation. Then the filter is used for feature extraction, hence improve the performance of the classifier. The proposed method is evaluated using Sentinel-1 and Sentinel-2 dataset and the improvement of accuracy shows the superiority of the proposed dynamic data fusion approach.
Pengming Feng, Youtian Lin, Guangjun He, Jian Guan 0001, Huifeng Shi
IGARSS4
2020 Enhanced Gaze Following via Object Detection and Human Pose Estimation
Jian Guan 0001, Liming Yin, Shuhan Qi, Xuan Wang 0002, Qing Liao 0001
MMM (2)1
2020 A mix-supervised unified framework for salient object detection
Fengwei Jia, Jian Guan 0001, Shuhan Qi, Huale Li, Xuan Wang 0002
Appl. Intell.2
2020 Bi-Connect Net for salient object detection
Fengwei Jia, Xuan Wang 0002, Jian Guan 0001, Qing Liao 0001, Jiajia Zhang 0001, Huale Li, Shuhan Qi
Neurocomputing3
2020 Association Loss for Visual Object Detection
abstract
Convolutional neural network (CNN) is a popular choice for visual object detection where two sub-nets are often used to achieve object classification and localization separately. However, the intrinsic relation between the localization and classification sub-nets was not exploited explicitly for object detection. In this letter, we propose a novel association loss, namely, the proxy squared error (PSE) loss, to entangle the two sub-nets, thus use the dependency between the classification and localization scores obtained from these two sub-nets to improve the detection performance. We evaluate our proposed loss on the MS-COCO dataset and compare it with the loss in a recent baseline, i.e. the fully convolutional one-stage (FCOS) detector. The results show that our method can improve the AP from 33.8 to 35.4 and AP75 from 35.4 to 37.8, as compared with the FCOS baseline.
Dongli Xu, Jian Guan 0001, Pengming Feng, Wenwu Wang 0001
IEEE Signal Process. Lett.2
2020 Audio-Visual Particle Flow SMC-PHD Filtering for Multi-Speaker Tracking
abstract
Sequential Monte Carlo probability hypothesis density (SMC-PHD) filtering is a popular method used recently for audio-visual (AV) multi-speaker tracking. However, due to the weight degeneracy problem, the posterior distribution can be represented poorly by the estimated probability, when only a few particles are present around the peak of the likelihood density function. To address this issue, we propose a new framework where particle flow (PF) is used to migrate particles smoothly from the prior to the posterior probability density. We consider both zero and non-zero diffusion particle flows (ZPF/NPF), and developed two new algorithms, AV-ZPF-SMC-PHD and AV-NPF-SMC-PHD, where the speaker states from the previous frames are also considered for particle relocation. The proposed algorithms are compared systematically with several baseline tracking methods using the AV16.3, AVDIAR and CLEAR datasets, and are shown to offer improved tracking accuracy and average effective sample size (ESS).
Yang Liu 0175, Volkan Kilic, Jian Guan 0001, Wenwu Wang 0001
IEEE Trans. Multim.3
2019 Embranchment Cnn Based Local Climate Zone Classification Using Sar And Multispectral Remote Sensing Data
abstract
In this study, a Local Climate Zone (LCZ) classification framework is established using a Densenet based embranchment Convolutional Neural Network (CNN). Both synthetic aperture radar (SAR) and multispectral data are employed for feature fusion, specifically, considering about the difference in imaging mechanism between SAR and multispectral data, features from both resources are extracted in different branches separately according to the physical properties of each band. Significant accuracy improvement can be achieved when evaluate the proposed method by Sentinel-1 and Sentinel-2 dataset, and the comparison results show the superiority of the proposed embranchment CNN framework over the conventional methods.
Pengming Feng, Youtian Lin, Jian Guan 0001, Guangjun He, Zhenghuan Xia, Huifeng Shi
IGARSS3
2019 Joint Convolutional Neural Network for Small-Scale Ship Classification in SAR Images
abstract
Ship classification using synthetic aperture radar (SAR) imagery is a challenge problem in maritime surveillance. Because of the scale limitation of ship targets in SAR image, convolutional neural networks (CNNs) can not achieve similar performance as for natural image classification. In this paper, we propose a joint CNNs framework for small-scale ship targets classification in SAR image, where a generator and a classifier are jointly connected. The generator can reconstruct the small-scale low-resolution (LR) images to large-scale super-resolution (SR) images, and the classifier is used for ship classification. A novel joint loss optimization strategy is introduced to solve the problem, where an MSE-based content loss is employed to generate high quality SR images, and a classification loss is applied to enable the generator and the classifier to be trained in a joint way. Experiments are conducted to demonstrate the superior performance of our proposed method, as compared with the state-of-the-art methods.
Yanxia Wu 0001, Ye Yuan 0011, Jian Guan 0001, Libo Yin, Pengming Feng
IGARSS3
2019 Bi-directional Features Reuse Network for Salient Object Detection
Fengwei Jia, Xuan Wang 0002, Jian Guan 0001, Shuhan Qi, Qing Liao 0001, Huale Li
PRICAI (3)3
2019 Graph-based supervised discrete image hashing
Jian Guan 0001, Xuan Wang 0002, Hainan Zhao, Jiajia Zhang 0001, Zechao Liu, Shuhan Qi
J. Vis. Commun. Image Represent.1
2019 Large scale product search with spatial quantization and deep ranking
Shuhan Qi, Zawlin Kyaw, Xuan Wang 0002, Zoe Lin Jiang, Jian Guan 0001
Multim. Tools Appl.5
2018 Pixel Meets Region: A Pratical Framework for Salient Object Detection
abstract
Due to the development of deep learning and Fully Convolutional Neural Network (FCN), the research on salient object detection has made great progress in recent years. However, such FCN based models are always affected by the scale-space problem, which reduces the saliency detection accuracy and leads to a blurred object boundary. In this paper, we propose a novel saliency object detection method, which predicts the saliency by incorporating both pixel-level and region-level predictions. First, in order to alleviate the scale-space problem, modified dilated convolution layers and short connections are integrated into the FCN model, and the pixel-level saliency maps is generated by a pixel-wise salient classifier. Then, we employ a superpixel based manifold learning algorithm to obtain a better boundary of salient object, by which a region-level saliency map with clearer object boundary is generated. At last, a simple fusion method is utilized to fuse the two saliency maps into a unified saliency map, followed by a DenseCRF post-refinement module to further optimize the final results. Experiments are conducted on two benchmark datasets to demonstrate the effectiveness of our method.
Xuan Wang 0002, Shuhan Qi, Jian Guan 0001, Fengwei Jia, Lin Yao 0004
ICME4
2018 Balance the Loss: Improving Deep Hash via Loss Weighting and Semantic Preserving
abstract
Learning to hash is widely used in approximate nearest-neighbor (ANN) search. However, traditional hash learning methods, which split the hashing into two parts: feature extraction and hash function learning, usually result in a low retrieval accuracy. Although existing deep learning based hashing methods can improve hashing quality by coupling feature learning and hash encoding, they are always affected by the positive-negative sample imbalance problem. It often deteriorates the performance of the generated hash code. In this paper, we propose an end-to-end deep hashing framework, in which a weighted pairwise loss function is employed to alleviate sample imbalance problem. The loss generated by the positive pairs and negative pairs are given different weights automatically. Moreover, we integrate a classification network into the hashing framework, which can preserve the semantic information by making sure the generated hash codes are also optimal for classification. Comparison experiments are conducted on two benchmark datasets to demonstrate the performance of our proposed approach.
Shuhan Qi, Xuan Wang 0002, Jian Guan 0001, Fengwei Jia, Lin Yao 0004
ICME4
2018 Infrared and visible image fusion based on NSCT and stacked sparse autoencoders
Xiaoqing Luo, Shuhan Qi, Jian Guan 0001, Zhancheng Zhang
Multim. Tools Appl.5
2018 Polynomial dictionary learning algorithms in sparse representations
Jian Guan 0001, Xuan Wang 0002, Pengming Feng, Jing Dong 0001, Jonathon A. Chambers, Zoe Lin Jiang, Wenwu Wang 0001
Signal Process.1
2018 Low rank matrix completion using truncated nuclear norm and sparse regularizer
Jing Dong 0001, Zhichao Xue, Jian Guan 0001, Zi-Fa Han, Wenwu Wang 0001
Signal Process. Image Commun.3
2017 Matrix of Polynomials Model Based Polynomial Dictionary Learning Method for Acoustic Impulse Response Modeling
abstract
We study the problem of dictionary learning for signals that can be represented as polynomials or polynomial matrices, such as convolutive signals with time delays or acoustic impulse responses.Recently, we developed a method for polynomial dictionary learning based on the fact that a polynomial matrix can be expressed as a polynomial with matrix coefficients, where the coefficient of the polynomial at each time lag is a scalar matrix.However, a polynomial matrix can be also equally represented as a matrix with polynomial elements.In this paper, we develop an alternative method for learning a polynomial dictionary and a sparse representation method for polynomial signal reconstruction based on this model.The proposed methods can be used directly to operate on the polynomial matrix without having to access its coefficients matrices.We demonstrate the performance of the proposed method for acoustic impulse response modeling.
Jian Guan 0001, Xuan Wang 0002, Pengming Feng, Jing Dong 0001, Wenwu Wang 0001
INTERSPEECH1
2017 Multiple level visual semantic fusion method for image re-ranking
Shuhan Qi, Fanglin Wang, Xuan Wang 0002, Jia Wei 0003, Jian Guan 0001
Multim. Syst.6