Wenwu Wang 0001

dblp:61/5537-1 · DBLP profile ↗
← Back
225ranked-venue papers
2as first author
117since 2021 · last 2026
0000-0002-8393-5703ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 152 · 1 first-author · 69 since 2021Artificial intelligence and machine learning · 81 · 1 first-author · 54 since 2021Databases, data management, data science and information retrieval · 9 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 6 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Controllable timbre cloning and style replication with reference speech examples for multimodal human-computer interaction
Tianwei Lan, Yuhang Guo 0001, Mengyuan Deng, Jing Wang 0037, Wenwu Wang 0001, Chong Feng 0001
Neurocomputing5
2026 Dual-branch normalizing flow for anomaly detection and localization from images
Yao Li 0019, Shiyong Lan, Wenwu Wang 0001, Yixin Qiao, Guonan Deng
Neurocomputing5
2026 Small object detection using multi-scale detail enhancement and decoupled detection head
Yixin Qiao, Xinyuan Zhou, Shiyong Lan, Wenwu Wang 0001, Yao Li 0019, Guonan Deng
Neurocomputing4
2026 STGFMamba: Spatio-temporal graph Fourier-enhanced Mamba for traffic prediction
Xinyuan Zhou, Ruiyi Lu, Zhiang Hou, Yao Ren, Wenwu Wang 0001, Shiyong Lan
Inf. Sci.6
2026 Zero-shot diverse audio captioning with diffusion models
Yonggang Zhu, Yiming Zhang 0025, Li Xiao 0005, Wenwu Wang 0001, Aidong Men
Knowl. Based Syst.4
2026 DSTFGCN: A dynamic spatial-temporal fusion graph convolution network for traffic flow forecasting
abstract
Traffic flow prediction is one of the core technology of Intelligent Transportation System. Its fundamental challenge is to effectively model the complex spatial-temporal dependencies. Although extensive research has been conducted in this field, the limitations of current methods restrict their effectiveness in accurate predictions. For temporal dependence, existing methods based on recurrent neural networks only focus on local dependencies and ignore global dependencies. For spatial dependencies, existing methods use predefined or adaptive adjacency matrices that cannot accurately reflect the relationships between real traffic flow. To overcome these limitations, we propose a dynamic spatial-temporal fusion graph convolution network (DSTFGCN). In the temporal aspect, we introduce gated dilated causal convolution to capture the local dependencies and node-independent temporal graph convolution to capture the global dependencies specific to each node. In the spatial aspect, we propose a dynamic graph convolution block. It can construct dynamic graphs based on the characteristics of the input data and aggregate both local and global spatial dependencies. Experiments on six real-world datasets have shown that DSTFGCN outperforms current mainstream methods. The codes are available at https://github.com/SYLan2019/DSTFGCN.
Tianyi Pan, Xinyuan Zhou, Shiyong Lan, Wenwu Wang 0001, Hongyu Yang 0002, Zhiang Hou, Yao Ren
Neural Networks4
2026 ENVISAGE: An image captioning dataset and conformal uncertainty modeling framework to assist visually impaired individuals
Özkan Çayli, Volkan Kilic, Wenwu Wang 0001
Pattern Recognit.3
2026 Unpaired overwater image defogging using inverted dark channel prior-guided cycle-consistent generative adversarial network
Yaozong Mo, Tuxin Guan, Qiuping Jiang, Wenqi Ren, Wenwu Wang 0001
Pattern Recognit.6
2026 DMAGaze : Gaze estimation using feature disentanglement and multi-scale attention
Haohan Chen, Hongjia Liu, Shiyong Lan, Wenwu Wang 0001, Yixin Qiao, Yao Li 0019, Guonan Deng
Pattern Recognit. Lett.4
2026 Differentiable interacting multiple model particle filtering
John-Joseph Brady, Yuhui Luo, Wenwu Wang 0001, Victor Elvira, Yunpeng Li 0001
Signal Process.3
2026 Soundscape Captioning Using Sound Affective Quality Network and Large Language Model
abstract
We live in a rich and varied acoustic world, which is experienced by individuals or communities as asoundscape. Computational auditory scene analysis, disentangling acoustic scenes by detecting and classifying events, focuses on objective attributes of sounds, such as their category and temporal characteristics, ignoring their effects on people, such as the emotions they evoke within a context. To fill this gap, we propose the affective soundscape captioning (ASSC) task, which enables automated soundscape analysis, thus avoiding labour-intensive subjective ratings and surveys in conventional methods. With soundscape captioning, context-aware descriptions are generated for soundscape by capturing the acoustic scenes (ASs), audio events (AEs) information, and the corresponding human affective qualities (AQs). To this end, we propose an automatic soundscape captioner (SoundSCaper) system composed of an acoustic model, i.e. SoundAQnet, and a large language model (LLM). SoundAQnet simultaneously models multi-scale information about ASs, AEs, and perceived AQs, while the LLM describes the soundscape with captions by parsing the information captured with SoundAQnet. SoundSCaper is assessed by two juries of 32 people. In expert evaluation, the average score of SoundSCaper-generated captions is slightly lower than that of two soundscape experts on the evaluation set D1 and the external mixed dataset D2, but not statistically significant. In layperson evaluation, SoundSCaper outperforms soundscape experts in several metrics on datasets D1 and D2. In addition to human evaluation, compared to other automated audio captioning (AAC) systems with and without LLM, SoundSCaper performs better on the ASSC task in several natural language processing (NLP) based metrics. Overall, SoundSCaper performs well in human subjective evaluation and various objective captioning metrics, and the generated captions are comparable to those annotated by soundscape experts. The model, source code, LLM scripts, human assessment data, instructions, and evaluation statistics are all publicly available.
Yuanbo Hou, Qiaoqiao Ren, Wenwu Wang 0001, Jian Kang 0002, Tony Belpaeme, Dick Botteldooren
IEEE Trans. Multim.4
2026 Seeing With Words: Interpretable Language-Guided Drone Geo-Localization via LLM-Enriched Semantic Attribute Alignment
abstract
Natural language-guided drone geo-localization (DGL) provides an intuitive and scalable mode of human-drone interaction for tasks such as search, rescue, and surveillance. Recent Vision-Language Models (VLMs) can learn semantic correspondences between text and images during fine-tuning. However, their performance in DGL tasks remains constrained, as complex instructions and cluttered scenes often cause semantic dilution and granularity mismatch, leading to weak cross-modal alignment. Consequently, the models struggle with ambiguous targets and suffer from reduced localization accuracy. To address these challenges, we propose SAA-DGL, a framework for interpretable language-guided Drone Geo-Localization that enriches Semantic Attribute Alignment (SAA) with large language models (LLMs). It introduces two parameter-free cross-modal fusion modules: (1) the LLM-driven Cross-modal Semantic Attribute Enrichment (LCSAE) module, which extracts fine-grained attributes (e.g., color, shape, position) from text and embeds them into visual features as explicit semantic anchors, producing semantically enriched cross-modal representations; and (2) the Bidirectional Feature Alignment (BFA) module, which builds fusion relationships between visual and textual features via similarity-driven mechanisms, enabling effective integration of enriched visual and textual information. This design improves cross-modal consistency and interpretability while preserving pretrained alignment priors and enhancing training stability. Experiments on the GeoText-1652 benchmark show that SAA-DGL achieves state-of-the-art performance and strong robustness under complex visual and linguistic disturbances, validating its effectiveness for challenging geo-localization scenarios. We will release the code.
Changsen Yuan, Yang-Hao Zhou, Cunhan Guo, Danjie Han, Ge Shi 0002, Wenwu Wang 0001
IEEE Trans. Multim.6
2025 Towards Controllable Speech Synthesis in the Era of Large Language Models: A Systematic Survey
abstract
Text-to-speech (TTS) has advanced from generating natural-sounding speech to enabling fine-grained control over attributes like emotion, timbre, and style.Driven by rising industrial demand and breakthroughs in deep learning, e.g., diffusion and large language models (LLMs), controllable TTS has become a rapidly growing research area.This survey provides the first comprehensive review of controllable TTS methods, from traditional control techniques to emerging approaches using natural language prompts.We categorize model architectures, control strategies, and feature representations, while also summarizing challenges, datasets, and evaluations in controllable TTS.This survey aims to guide researchers and practitioners by offering a clear taxonomy and highlighting future directions in this fast-evolving field.One can visit https://github.com/imxtx/ awesome-controllabe-speech-synthesis for a comprehensive paper list and updates.
Tianxin Xie, Yan Rong, Pengfei Zhang 0005, Wenwu Wang 0001, Li Liu 0036
EMNLP4
2025 Graph-Enhanced Dual-Stream Feature Fusion with Pre-Trained Model for Acoustic Traffic Monitoring
abstract
Microphone array techniques are widely used in sound source localization and smart city acoustic-based traffic monitoring, but these applications face significant challenges due to the scarcity of labeled real-world traffic audio data and the complexity and diversity of application scenarios. The DCASE Challenge’s Task 10 focuses on using multi-channel audio signals to count vehicles (cars or commercial vehicles) and identify their directions (left-to-right or vice versa). In this paper, we propose a graph-enhanced dual-stream feature fusion network (GEDF-Net) for acoustic traffic monitoring, which simultaneously considers vehicle type and direction to improve detection. We propose a graph-enhanced dual-stream feature fusion strategy which consists of a vehicle type feature extraction (VTFE) branch, a vehicle direction feature extraction (VDFE) branch, and a frame-level feature fusion module to combine the type and direction feature for enhanced performance. A pre-trained model (PANNs) is used in the VTFE branch to mitigate data scarcity and enhance the type features, followed by a graph attention mechanism to exploit temporal relationships and highlight important audio events within these features. The frame-level fusion of direction and type features enables fine-grained feature representation, resulting in better detection performance. Experiments demonstrate the effectiveness of our proposed method. GEDF-Net is our submission that achieved 1st place in the DCASE 2024 Challenge Task 10.
Shitong Fan, Feiyang Xiao, Shuhan Qi, Qiaoxi Zhu, Wenwu Wang 0001, Jian Guan 0001
ICASSP6
2025 Sound-Based Recognition of Touch Gestures and Emotions for Enhanced Human-Robot Interaction
abstract
Emotion recognition and touch gesture decoding are crucial for advancing human-robot interaction (HRI), especially in social environments where emotional cues and tactile perception play important roles. However, many humanoid robots, such as Pepper, Nao, and Furhat, lack full-body tactile skin, limiting their ability to engage in touch-based emotional and gesture interactions. In addition, vision-based emotion recognition methods usually face strict GDPR compliance challenges due to the need to collect personal facial data. To address these limitations and avoid privacy issues, this paper studies the potential of using the sounds produced by touching during HRI to recognise tactile gestures and classify emotions along the arousal and valence dimensions. Using a dataset of tactile gestures and emotional interactions from 28 participants with the humanoid robot Pepper, we design an audio-only lightweight touch gesture and emotion recognition model with only 0.24M parameters, 0.94MB model size, and 0.7G FLOPs. Experimental results show that the proposed model effectively recognises the arousal and valence states of different emotions, as well as various tactile gestures, when the input audio length varies. The proposed model is of low-latency and achieves similar results as well-known pretrained audio neural networks (PANNs), but with much smaller FLOPs, number of parameters, and model size.
Yuanbo Hou, Qiaoqiao Ren, Wenwu Wang 0001, Dick Botteldooren
ICASSP3
2025 Multi-Modal Rhythmic Generative Model for Chinese Cued Speech Gestures Generation
abstract
Cued speech (CS) is a novel visual coding system, which combines lip reading with several specific hand codings to help hearing-impaired people to communicate effectively. This work focuses on the audio/text-driven CS gestures (i.e., continuous lip and hand gestures movements) generation. Previous work used template-based statistical methods for the French CS generation. However, these methods are fragile since they need careful hand-crafted pre-processing to fit models, resulting in poor robustness. Furthermore, the natural rhythm in generated CS gesture sequences, which is essential for a coding system of spoken languages, was overlooked in prior studies. To solve the above-mentioned problems, we innovatively propose a two-branched rhythmic CS gesture generation framework, which contains a multi-modal adversarial semantic generator (MASG) to generate accurate multi-modal CS gestures (i.e., lip, hand shape and hand position movements), and an audio-driven rhythm generator (ARG) to extract the rhythm information. Moreover, we design a new Gesture Audio Difference (GAD) metric to evaluate the rhythm coherence considering the issue of asynchrony between CS hand gestures and lip movements. Extensive experimental results are presented on two datasets of two tasks (a CS dataset named MCCS-2024 and a co-speech TED dataset) with comprehensive ablation analysis and user study, demonstrating the effectiveness of our method. The code and dataset with multi-modal annotations were made public at https://mccs-2024.github.io/.
Li Liu 0036, Wentao Lei, Wenwu Wang 0001
ICASSP3
2025 Bayesian Nonparametric Clustering for Source Counting with a Small Aperture Microphone Array
abstract
Source counting (SC) in an indoor environment is an important problem in computational auditory scene analysis. However, the problem is challenging, especially when reverberation and ambient noise are present in the environment. To address this problem, we propose an augmented Bayesian non-parametric (ABNP) clustering algorithm for source counting based on sound intensity (SI) captured by a small aperture microphone array. The core idea is to incorporate an infinite Gaussian mixture model (IGMM) and a time-frequency (TF) augmented weight selection and update scheme for sound intensity estimation. The use of IGMM enables the exemption of the maximum number of sources assumed in previous methods. Experiments on both simulated and real-world data show the improved performance by the proposed method as compared with the state of the art baseline methods.
Kunkun SongGong, Pufen Zhang, Xiongwei Zhang, Wenwu Wang 0001, Meng Sun 0001, Chong Jia
ICASSP4
2025 FlowSep: Language-Queried Sound Separation with Rectified Flow Matching
abstract
Language-queried audio source separation (LASS) focuses on separating sounds using textual descriptions of the desired sources. Current methods mainly use discriminative approaches, such as time-frequency masking, to separate target sounds and minimize interference from other sources. However, these models face challenges when separating overlapping sound-tracks, which may lead to artifacts such as spectral holes or incomplete separation. Rectified flow matching (RFM), a generative model that establishes linear relations between the distribution of data and noise, offers superior theoretical properties and simplicity, but has not yet been explored in sound separation. In this work, we introduce FlowSep, a new generative model based on RFM for LASS tasks. FlowSep learns linear flow trajectories from noise to target source features within the variational autoencoder (VAE) latent space. During inference, the RFM-generated latent features are reconstructed into a mel-spectrogram via the pre-trained VAE decoder, followed by a pre-trained vocoder to synthesize the waveform. Trained on 1, 680 hours of audio data, FlowSep outperforms the state-of-the-art models across multiple benchmarks, as evaluated with subjective and objective metrics. Additionally, our results show that FlowSep surpasses a diffusion-based LASS model in both separation quality and inference efficiency, highlighting its strong potential for audio source separation tasks. Code, pre-trained models and demos can be found at: https://audio-agi.github.io/FlowSep_demo/.
Xubo Liu 0001, Haohe Liu, Mark D. Plumbley, Wenwu Wang 0001
ICASSP5
2025 Sound-VECaps: Improving Audio Generation with Visually Enhanced Captions
abstract
Generative models have shown significant achievements in audio generation tasks. However, existing models struggle with complex and detailed prompts, leading to potential performance degradation. We hypothesize that this problem stems from the simplicity and scarcity of the training data. This work aims to create a large-scale audio dataset with rich captions for improving audio generation models. We first develop an automated pipeline to generate detailed captions by transforming predicted visual captions, audio captions, and tagging labels into comprehensive descriptions using a Large Language Model (LLM). The resulting dataset, Sound-VECaps, comprises 1.66M high-quality audio-caption pairs with enriched details including audio event orders, occurred places and environment information. We then demonstrate that training the text-to-audio generation models with Sound-VECaps significantly improves the performance on complex prompts. Furthermore, we conduct ablation studies of the models on several downstream audio-language tasks, showing the potential of Sound-VECaps in advancing audio-text representation learning.Dataset and demos are available at https://yyua8222.github.io/Sound-VECaps-demo/.
Dongya Jia, Xiaobin Zhuang, Yuanzhe Chen, Zhuo Chen 0006, Yuping Wang 0005, Yuxuan Wang 0002, Xubo Liu 0001, Xiyuan Kang, Mark D. Plumbley, Wenwu Wang 0001
ICASSP11
2025 MRGNN: Mamba-Register-Based Graph Neural Network for Unsupervised Anomaly Detection in Multivariate Time Series
Danling Meng, Xinyuan Zhou, Shiyong Lan, Wenwu Wang 0001, Weihong Yuan, Ruiyi Lu
ICIC (12)4
2025 EMGPose: An Efficient Multi-Granularity Representation for Human Pose Estimation
abstract
Current Transformer-based methods typically unwisely represent the entire image at a single granularity. A high-resolution representation of the regions of interest can significantly improve the accuracy of human pose estimation while causing unnecessary computational costs for other regions. To overcome this limitation, we propose an efficient two-stage framework using adaptive Multi-Granularity representation for different important image regions for Human pose estimation (EMGPose). In the first stage, the image is split into coarse-grained patches for simple inference. If without sufficient accuracy, important patches will be resplit into multiple finer-grained patches for the second stage of inference. Furthermore, we propose a new token-merge strategy based on token importance and similarity in Transformer, effectively reducing the computational load from low-information background patches. Extensive experiments demonstrate the excellent performance of the proposed method. Specifically, our model EMGPose-Base achieves 76.3 AP (+0.5 AP) and 62.2 AP (+2.6 AP) and higher efficiency than baseline ViTPose-Base on the COCO validation set and OCHuman test set, respectively.
Guonan Deng, Shiyong Lan, Wenwu Wang 0001, Yixin Qiao, Yao Li 0019, Haohan Chen, Hongyu Yang 0002
ICME3
2025 DCGNet: Detail and Context Guided Small Object Detection Network with Decoupled Detection Head
abstract
Small object detection plays a significant role in many fields, yet it still faces numerous challenges. These include the lack of key features due to its tiny size, and its sensitivity to environmental noise. To address these issues, we introduce a novel network termed Detail and Context Guided Small Object Detection Network (DCGNet). Specifically, 1) we design a Detail and Context Enhanced Feature Pyramid Network (DCE-FPN), which transforms the feature from the temporal domain into the frequency domain, followed by noise reduction on high-frequency components to accentuate the detailed information of small objects and leverages multiple branches with different receptive fields to capture multi-scale contextual information to further differentiate small objects from the background. 2) We propose a Enhanced Decoupled Interaction RCNN (EDI-RCNN) that independently enhances the specific features for each task and subsequently facilitates comprehensive feature interaction. This helps avoid ambiguity and sparsity of the features of the small objects extracted with coupled detection head in existing methods. The results on the well-known small object detection datasets (VisDrone2019 and AI-TODv2) show that our proposed method achieves SOTA in both Average Precision (AP) and Average Precision for small objects (APs) metrics.
Yixin Qiao, Shiyong Lan, Wenwu Wang 0001, Haohan Chen, Yao Li 0019, Guonan Deng
ICME3
2025 EnvSDD: Benchmarking Environmental Sound Deepfake Detection
Han Yin, Yang Xiao 0019, Rohan Kumar Das, Jisheng Bai, Haohe Liu, Wenwu Wang 0001, Mark D. Plumbley
INTERSPEECH6
2025 TeMTG: Text-Enhanced Multi-Hop Temporal Graph Modeling for Audio-Visual Video Parsing
abstract
Audio-Visual Video Parsing (AVVP) task aims to parse the event categories and occurrence times from audio and visual modalities in a given video. Existing methods usually focus on implicitly modeling audio and visual features through weak labels, without mining semantic relationships for different modalities and explicit modeling of event temporal dependencies. This makes it difficult for the model to accurately parse event information for each segment under weak supervision, especially when high similarity between segmental modal features leads to ambiguous event boundaries. Hence, we propose a multimodal optimization framework, TeMTG, that combines text enhancement and multi-hop temporal graph modeling. Specifically, we leverage pre-trained multimodal models to generate modality-specific text embeddings, and fuse them with audio-visual features to enhance the semantic representation of these features. In addition, we introduce a multi-hop temporal graph neural network, which explicitly models the local temporal relationships between segments, capturing the temporal continuity of both short-term and long-range events. Experimental results demonstrate that our proposed method achieves state-of-the-art (SOTA) performance in multiple key indicators in the LLP dataset.
Yaru Chen 0003, Peiliang Zhang, Fei Li 0022, Faegheh Sardari, Ruohao Guo, Wenwu Wang 0001
ICMR7
2025 MADFlow: Multimodal difference compensation flow for multimodal anomaly detection
Yao Li 0019, Xinyuan Zhou, Shiyong Lan, Wenwu Wang 0001, Yixin Qiao
Neurocomputing4
2025 Multi-Layer Gated Recurrent Unit-Based Recurrent Neural Network for Image Captioning
abstract
Generating natural language descriptions of an image, namely image captioning, has received much attention in computer vision and natural language processing. Recent image captioning models are mainly based on the encoder-decoder framework in which visual information is extracted by an encoder, e.g. using convolutional neural network (CNN), and captions are generated by a decoder, e.g. using recurrent neural network (RNN). Although this framework is promising for image captioning, there are still issues in the RNN decoder for exploiting the visual information to generate grammatically and semantically correct captions. More specifically, the RNN decoder has limited ability in dealing with long-term complex dependencies, leading to ineffective use of contextual information from the encoded data. To address this issue, in this paper, we introduce a multi-layer gated recurrent unit (ML-GRU) within the conventional RNN decoder, which enables the modulation of the relevant information flow inside the unit, and thus leads to the generation of semantically coherent captions. The proposed ML-GRU-based RNN decoder has been extensively evaluated on the MSCOCO dataset, and experimental results demonstrate the advantage of our proposed approach over the state-of-the-art approaches across multiple performance metrics.
Özkan Çayli, Volkan Kilic, Aytug Onan, Wenwu Wang 0001
Int. J. Pattern Recognit. Artif. Intell.4
2025 TDU-DLNet: A transformer-based deep unfolding network for dictionary learning
Kai Wu 0004, Jing Dong 0001, Guifu Hu, Chang Liu 0152, Wenwu Wang 0001
Signal Process.5
2025 Noise-robust feature extraction for keyword spotting based on supervised adversarial domain adaptation training strategies
Qianhua He, Zunxian Liu, Mingru Yang, Wenwu Wang 0001
Speech Commun.5
2025 SCAN: Selective Contrastive Learning Against Noisy Data for Acoustic Anomaly Detection
abstract
Acoustic Anomaly Detection (AAD) has gained significant attention for the detection of suspicious activities or faults. Contrastive learning-based unsupervised AAD has outperformed traditional models on academic datasets, however, its model training is predominantly based on datasets containing only normal samples. In real industrial settings, a dataset of normal samples can still be corrupted by abnormal samples. Handling such noisy data is a crucial challenge, yet it remains largely unsolved. To address this issue, this paper proposes a Selective Contrastive learning framework Against Noisy data (SCAN) to mitigate the adverse effects of training the AAD model with anomaly-corrupted data. Specifically, SCAN progressively constructs confidence sample pairs based on the Mahalanobis distance, which is derived from the geometric median. These selected pairs are then integrated into the contrastive learning framework to enhance representation learning and model robustness. Extensive experiments under varying levels of label noise (i.e., the proportion of mislabeled abnormal samples in training data) demonstrate that SCAN outperforms state-of-the-art (SOTA) AAD methods on the real-world industrial datasets DCASE2022 and DCASE2024 Task2.
Zhaoyi Liu 0003, Yuanbo Hou, Wenwu Wang 0001, Sam Michiels, Danny Hughes 0001
IEEE Signal Process. Lett.3
2025 Multimodal Fish Feeding Intensity Assessment in Aquaculture
abstract
Fish feeding intensity assessment (FFIA) aims to evaluate fish appetite changes during feeding, which is crucial in industrial aquaculture applications. Existing FFIA methods are limited by their robustness to noise, computational complexity, and the lack of public datasets for developing the models. To address these issues, we first introduce AV-FFIA, a new dataset containing 27,000 labeled audio and video clips that capture different levels of fish feeding intensity. Then, we introduce multi-modal approaches for FFIA by leveraging the models pre-trained on individual modalities and fused with data fusion methods. We perform benchmark studies of these methods on AV-FFIA, and demonstrate the advantages of the multi-modal approach over the single-modality based approach, especially in noisy environments. However, compared to the methods developed for individual modalities, the multimodal approaches may involve higher computational costs due to the need for independent encoders for each modality. To overcome this issue, we further present a novel unified mixed-modality based method for FFIA, termed as U-FFIA. U-FFIA is a single model capable of processing audio, visual, or audio-visual modalities, by leveraging modality dropout during training and knowledge distillation using the models pre-trained with data from single modality. We demonstrate that U-FFIA can achieve performance better than or on par with the state-of-the-art modality-specific FFIA models, with significantly lower computational overhead, enabling robust and efficient FFIA for improved aquaculture management. To encourage further research, we have released the AV-FFIA dataset, the pre-trained model and codes athttps://github.com/FishMaster93/U-FFIA. Note to Practitioners—Feeding is one of the most important costs in aquaculture. However, current feeding machines usually operate with fixed thresholds or human experiences, lacking the ability to automatically adjust to fish feeding intensity. FFIA can evaluate the intensity changes in fish appetite during the feeding process and optimize the control strategies of the feeding machine to avoid inadequate feeding or overfeeding, thereby reducing the feeding cost and improving the well-being of fish in industrial aquaculture. The existing methods have mainly exploited single-modality data, and have a high sensitivity to input noise. Using video and audio offers improved chances to address the challenges brought by various environments. However, compared with processing data from single modalities, using multiple modalities simultaneously often involves increased computational resources, including memory, processing power, and storage. This can impact system performance and scalability. To address these issues, we focus on the efficient unified model, which is capable of processing both multimodal and single-modal input. Our proposed model achieved state-of-the-art (SOTA) performance in FFIA with high computational efficiency.
Xubo Liu 0001, Haohe Liu, Zhuangzhuang Du, Tao Chen 0009, Guoping Lian, Lihui Wang 0002, Wenwu Wang 0001
IEEE Trans Autom. Sci. Eng.8
2025 RemoteDPL: A Semi-Supervised Object Detector With Dense Pseudo-Labels for Remote Sensing
abstract
Deep learning-based object detection has seen substantial advancements, however, its practical deployment is often constrained by the need for large-scale labeled datasets. This limitation becomes even more critical in remote sensing imagery, where objects are densely distributed and exhibit significant scale variations. To address these challenges, we introduce RemoteDPL, a novel semi-supervised object detection (SSOD) framework that leverages dense pseudo-labeling (DPL) and multi-scale learning. RemoteDPL offers three key contributions. First, a fusion module is designed to dynamically integrate spatial and channel features across scales, improving detection across varied object sizes. Second, an instance density prediction branch is introduced to support pseudo-label mining, enhancing detection performance in densely populated regions. Lastly, we propose a two-stage pseudo-label filtering strategy that first selects "pending" class predictions and then refines them using a joint confidence score based on both classification and density information. Extensive experiments on the DOTA-v1.0 and NWPU datasets confirm the effectiveness of RemoteDPL, demonstrating its clear advantage over existing state-of-the-art (SOTA) semi-supervised object detection methods. On the NWPU dataset, RemoteDPL outperforms the SOTA baseline by +3.44%, +1.10%, and +1.62% under the settings of data labelled with 30%, 40%, and 50%, respectively, highlighting its strong capability in low-label remote sensing scenarios.
Yongjie Ma, Xinyuan Zhou, Shiyong Lan, Wenwu Wang 0001, Zicheng Sun, Yixin Qiao
IEEE Trans. Geosci. Remote. Sens.4
2024 Learning Temporal Resolution in Spectrogram for Audio Classification
abstract
The audio spectrogram is a time-frequency representation that has been widely used for audio classification. One of the key attributes of the audio spectrogram is the temporal resolution, which depends on the hop size used in the Short-Time Fourier Transform (STFT). Previous works generally assume the hop size should be a constant value (e.g., 10 ms). However, a fixed temporal resolution is not always optimal for different types of sound. The temporal resolution affects not only classification accuracy but also computational cost. This paper proposes a novel method, DiffRes, that enables differentiable temporal resolution modeling for audio classification. Given a spectrogram calculated with a fixed hop size, DiffRes merges non-essential time frames while preserving important frames. DiffRes acts as a "drop-in" module between an audio spectrogram and a classifier and can be jointly optimized with the classification task. We evaluate DiffRes on five audio classification tasks, using mel-spectrograms as the acoustic features, followed by off-the-shelf classifier backbones. Compared with previous methods using the fixed temporal resolution, the DiffRes-based method can achieve the equivalent or better classification accuracy with at least 25% computational cost reduction. We further show that DiffRes can improve classification accuracy by increasing the temporal resolution of input acoustic features, without adding to the computational cost.
Haohe Liu, Xubo Liu 0001, Qiuqiang Kong, Wenwu Wang 0001, Mark D. Plumbley
AAAI4
2024 Regime Learning for Differentiable Particle Filters
abstract
Differentiable particle filters are an emerging class of models that combine sequential Monte Carlo techniques with the flexibility of neural networks to perform state space inference. This paper concerns the case where the system may switch between a finite set of state-space models, i.e. regimes. No prior approaches effectively learn both the individual regimes and the switching process simultaneously. In this paper, we propose the neural network based regime learning differentiable particle filter (RLPF) to address this problem. We further design a training procedure for the RLPF and other related algorithms. We demonstrate competitive performance compared to the previous state-of-the-art algorithms on a pair of numerical experiments.
John-Joseph Brady, Yuhui Luo, Wenwu Wang 0001, Victor Elvira, Yunpeng Li 0001
FUSION3
2024 Fusion of Audio and Visual Embeddings for Sound Event Localization and Detection
abstract
Sound event localization and detection (SELD) combines two subtasks: sound event detection (SED) and direction of arrival (DOA) estimation. SELD is usually tackled as an audio-only problem, but visual information has been recently included. Few audio-visual (AV)-SELD works have been published and most employ vision via face/object bounding boxes, or human pose keypoints. In contrast, we explore the integration of audio and visual feature embeddings extracted with pre-trained deep networks. For the visual modality, we tested ResNet50 and Inflated 3D ConvNet (I3D). Our comparison of AV fusion methods includes the AV-Conformer and Cross-Modal Attentive Fusion (CMAF) model. Our best models outperform the DCASE 2023 Task3 audio-only and AV baselines by a wide margin on the development set of the STARSS23 dataset, making them competitive amongst state-of-the-art results of the AV challenge, without model ensembling, heavy data augmentation, or prediction post-processing. Such techniques and further pre-training could be applied as next steps to improve performance.
Davide Berghi, Peipei Wu, Jinzheng Zhao, Wenwu Wang 0001, Philip J. B. Jackson
ICASSP4
2024 CM-PIE: Cross-Modal Perception for Interactive-Enhanced Audio-Visual Video Parsing
abstract
Audio-visual video parsing is the task of categorizing a video with weak labels at the segment level, and predicting them as audible or visible events. Recent methods have leveraged the attention mechanism to capture the semantic correlations among the whole video across the audio-visual modalities. However, these approaches may overlook the importance of individual segments and their interrelations within a video, typically relying on a single modality when learning features. In this paper, we propose a novel interactive-enhanced cross-modal perception method (CM-PIE), which can learn fine-grained features by applying a segment-based attention module. In addition, a cross-modal aggregation block is introduced to jointly optimize the semantic representation of audio and visual signals by enhancing inter-modal interactions. Experimental results show that our model offers improved parsing performance on the Look, Listen, and Parse (LLP) dataset compared to other methods.
Yaru Chen 0003, Ruohao Guo, Xubo Liu 0001, Peipei Wu, Guangyao Li 0001, Wenwu Wang 0001
ICASSP7
2024 Multi-Level Graph Learning For Audio Event Classification And Human-Perceived Annoyance Rating Prediction
abstract
WHO’s report on environmental noise estimates that 22 M people suffer from chronic annoyance related to noise caused by audio events (AEs) from various sources. Annoyance may lead to health issues and adverse effects on metabolic and cognitive systems. In cities, monitoring noise levels does not provide insights into noticeable AEs, let alone their relations to annoyance. To create annoyance-related monitoring, this paper proposes a graph-based model to identify AEs in a sound-scape, and explore relations between diverse AEs and human-perceived annoyance rating (AR). Specifically, this paper proposes a lightweight multi-level graph learning (MLGL) based on local and global semantic graphs to simultaneously perform audio event classification (AEC) and human annoyance rating prediction (ARP). Experiments show that: 1) MLGL with 4.1 M parameters improves AEC and ARP results by using semantic node information in local and global context-aware graphs; 2) MLGL captures relations between coarse-and fine-grained AEs and AR well; 3) Statistical analysis of MLGL results shows that some AEs from different sources significantly correlate with AR, which is consistent with previous research on human perception of these sound sources.
Yuanbo Hou, Qiaoqiao Ren, Siyang Song, Wenwu Wang 0001, Dick Botteldooren
ICASSP5
2024 Hierarchical Metadata Information Constrained Self-Supervised Learning for Anomalous Sound Detection under Domain Shift
abstract
Self-supervised learning methods have achieved promising performance for anomalous sound detection (ASD) under domain shift by incorporating the metadata of domain shift types and machine sound attributes in feature learning. However, the relation between domain shifts and machine sound attributes has yet to be fully utilised despite their potential benefits for characterising domain shifts. This paper presents a hierarchical metadata information constrained self-supervised ASD method, where the hierarchical relation between domain shift types (section IDs) and attributes is constructed and used as constraints to improve feature representation. In addition, we propose an attribute-group-centre based method for calculating the anomaly score under the domain shift condition. Experiments show improved audio feature learning over the state-of-the-art methods in DCASE 2022 challenge Task 2.
Haiyan Lan, Qiaoxi Zhu, Jian Guan 0001, Yuming Wei, Wenwu Wang 0001
ICASSP5
2024 Audiosr: Versatile Audio Super-Resolution at Scale
abstract
Audio super-resolution is a fundamental task that predicts high-frequency components for low-resolution audio, enhancing audio quality in digital applications. Previous methods have limitations such as the limited scope of audio types (e.g., music, speech) and specific bandwidth settings they can handle (e.g., 4 kHz to 8 kHz). In this paper, we introduce a diffusion-based generative model, AudioSR, that is capable of performing robust audio super-resolution on versatile audio types, including sound effects, music, and speech. Specifically, AudioSR can upsample any input audio signal within the bandwidth range of 2 kHz to 16 kHz to a high-resolution audio signal at 24 kHz bandwidth with a sampling rate of 48 kHz. Extensive objective evaluation on various audio super-resolution benchmarks demonstrates the strong result achieved by the proposed model. In addition, our subjective evaluation shows that AudioSR can act as a plug-and-play module to enhance the generation quality of a wide range of audio generative models, including AudioLDM, Fastspeech2, and MusicGen. Our code and demo are available at https://audioldm.github.io/audiosr.
Haohe Liu, Ke Chen 0021, Qiao Tian 0001, Wenwu Wang 0001, Mark D. Plumbley
ICASSP4
2024 Multi-Speaker Localization in the Circular Harmonic Domain on Small Aperture Microphone Arrays Using Deep Convolutional Networks
abstract
Acoustic signal processing in the circular harmonic domain (CHD) is an appealing method for speaker localization, since it inherently supports wideband acoustic sources and provides frequency invariant beampatterns. However, the performance of existing circular harmonic direction-of-arrival (DOA) estimation approaches can be degraded by a variety of factors, including background noise and reverberation in the acoustic environments, small aperture size of the circular array and the presence of multiple active sources. This paper addresses these issues by proposing a novel multi-speaker CHD localization method with small-sized microphone arrays using deep convolutional neural networks (CNN). The core idea is to construct circular harmonic features through joining the selected time-frequency (TF) bins of higher power and the operation of a randomization process by mimicking the sparsity property of speech signals. After that, we implement multi-speaker estimation as a multi-label classification task, and propose to use CNN with binary cross-entropy as the loss function. Experimental results show that our method performs significantly better than the baseline methods, on both simulated and real data, in terms of the accuracy of DOA estimation.
Kunkun SongGong, Pufen Zhang, Xiongwei Zhang, Meng Sun 0001, Wenwu Wang 0001
ICASSP5
2024 Retrieval-Augmented Text-to-Audio Generation
abstract
Despite recent progress in text-to-audio (TTA) generation, we show that the state-of-the-art models, such as AudioLDM, trained on datasets with an imbalanced class distribution, such as AudioCaps, are biased in their generation performance. Specifically, they excel in generating common audio classes while underperforming in the rare ones, thus degrading the overall generation performance. We refer to this problem as long-tailed text-to-audio generation. To address this issue, we propose a simple retrieval-augmented approach for TTA models. Specifically, given an input text prompt, we first leverage a Contrastive Language Audio Pretraining (CLAP) model to retrieve relevant text-audio pairs. The features of the retrieved audio-text data are then used as additional conditions to guide the learning of TTA models. We enhance AudioLDM with our proposed approach and denote the resulting augmented system as Re-AudioLDM. On the AudioCaps dataset, Re-AudioLDM achieves a state-of-the-art Frechet Audio Distance (FAD) of 1.37, outperforming the existing approaches by a large margin. Furthermore, we show that Re-AudioLDM can generate realistic audio for complex scenes, rare audio classes, and even unseen audio types, indicating its potential in TTA tasks.
Haohe Liu, Xubo Liu 0001, Qiushi Huang, Mark D. Plumbley, Wenwu Wang 0001
ICASSP6
2024 First-Shot Unsupervised Anomalous Sound Detection with Unknown Anomalies Estimated by Metadata-Assisted Audio Generation
abstract
First-shot (FS) unsupervised anomalous sound detection (ASD) is a brand-new task introduced in DCASE 2023 Challenge Task 2, where the anomalous sounds for the target machine types are unseen in training. Existing methods often rely on the availability of normal and abnormal sound data from the target machines. However, due to the lack of anomalous sound data for the target machine types, it becomes challenging when adapting the existing ASD methods to the first-shot task. In this paper, we propose a new framework for the first-shot unsupervised ASD, where metadata-assisted audio generation is used to estimate unknown anomalies, by utilising the available machine information (i.e., metadata and sound data) to fine-tune a text-to-audio generation model for generating the anomalous sounds that contain unique acoustic characteristics accounting for each different machine type. We then use the method of Time-Weighted Frequency domain audio Representation with Gaussian Mixture Model (TWFRGMM) as the backbone to achieve the first-shot unsupervised ASD. Our proposed FS-TWFR-GMM method achieves competitive performance amongst top systems in DCASE 2023 Challenge Task 2, while requiring only 1% model parameters for detection, as validated in our experiments.
Hejing Zhang, Qiaoxi Zhu, Jian Guan 0001, Haohe Liu, Feiyang Xiao, Jiantong Tian, Xinhao Mei, Xubo Liu 0001, Wenwu Wang 0001
ICASSP9
2024 PFCA-Net: Pyramid Feature Fusion and Cross Content Attention Network for Automated Audio Captioning
Jianyuan Sun, Wenwu Wang 0001, Mark D. Plumbley
INTERSPEECH2
2024 Efficient Audio Captioning with Encoder-Level Knowledge Distillation
Xuenan Xu, Haohe Liu, Mengyue Wu, Wenwu Wang 0001, Mark D. Plumbley
INTERSPEECH4
2024 Semantic-aware normalizing flow with feature fusion for image anomaly detection
Yao Li 0019, Shiyong Lan, Wenwu Wang 0001, Weikang Huang, Wujiang Zhu
Neurocomputing4
2024 Generating Accurate and Diverse Audio Captions Through Variational Autoencoder Framework
abstract
Generating both diverse and accurate descriptions is an essential goal in the audio captioning task. Traditional methods mainly focus on improving the accuracy of the generated captions but ignore their diversity. In contrast, recent methods have considered generating diverse captions for a given audio clip, but with the potential trade-off in caption accuracy. In this work, we propose a new diverse audio captioning method based on a variational autoencoder structure, dubbed AC-VAE, aiming to achieve a better trade-off between the diversity and accuracy of the generated captions. To improve diversity, AC-VAE learns the latent word distribution at each location based on contextual information. To uphold accuracy, AC-VAE incorporates an autoregressive prior module and a global constraint module, which enable precise modeling of word distribution and encourage semantic consistency of captions at the sentence level. We evaluate the proposed AC-VAE on the Clotho dataset. Experimental results show that AC-VAE achieves a better trade-off between diversity and accuracy compared to the state-of-the-art methods. The code is publicly available at https://github.com/XinMing0411/AC-VAE
Yiming Zhang 0025, Ruoyi Du, Zheng-Hua Tan, Wenwu Wang 0001, Zhanyu Ma
IEEE Signal Process. Lett.4
2024 ASiT: Local-Global Audio Spectrogram Vision Transformer for Event Classification
abstract
Transformers, which were originally developed for natural language processing, have recently generated significant interest in the computer vision and audio communities due to their flexibility in learning long-range relationships. Constrained by the data hungry nature of transformers and the limited amount of labelled data, most transformer-based models for audio tasks are finetuned from ImageNet pretrained models, despite the huge gap between the domain of natural images and audio. This has motivated the research in self-supervised pretraining of audio transformers, which reduces the dependency on large amounts of labeled data and focuses on extracting concise representations of audio spectrograms. In this paper, we proposeLocal-GlobalAudioSpectrogram vIsionTransformer, namely ASiT, a novel self-supervised learning framework that captures local and global contextual information by employing group masked model learning and self-distillation. We evaluate our pretrained models on both audio and speech classification tasks, including audio event classification, keyword spotting, and speaker identification. We further conduct comprehensive ablation studies, including evaluations of different pretraining strategies. The proposed ASiT framework significantly boosts the performance on all tasks and sets a new state-of-the-art performance in five audio and speech classification tasks, outperforming recent methods, including the approaches that use additional datasets for pretraining.
Sara Atito Ali Ahmed, Muhammad Awais 0001, Wenwu Wang 0001, Mark D. Plumbley, Josef Kittler
IEEE ACM Trans. Audio Speech Lang. Process.3
2024 Cooperative Scene-Event Modelling for Acoustic Scene Classification
abstract
Acoustic scene classification (ASC) can be helpful for creating context awareness for intelligent robots. Humans naturally use the relations between acoustic scenes (AS) and audio events (AE) to understand and recognize their surrounding environments. However, in most previous works, ASC and audio event classification (AEC) are treated as independent tasks, with a focus primarily on audio features shared between scenes and events, but not their implicit relations. To address this limitation, we propose a cooperative scene-event modelling (cSEM) framework to automatically model the intricate scene-event relation by an adaptive coupling matrix to improve ASC. Compared with other scene-event modelling frameworks, the proposed cSEM offers the following advantages. First, it reduces the confusion between similar scenes by aligning the information of coarsegrained AS and fine-grained AE in the latent space, and reducing the redundant information between the AS and AE embeddings. Second, it exploits the relation information between AS and AE to improve ASC, which is shown to be beneficial, even if the information of AE is derived from unverified pseudo-labels. Third, it uses a regression-based loss function for cooperative modelling of scene-event relations, which is shown to be more effective than classification-based loss functions. Instantiated from four models based on either Transformer or convolutional neural networks, cSEM is evaluated on real-life and synthetic datasets. Experiments show that cSEM-based models work well in reallife scene-event analysis, offering competitive results on ASC as compared with other multi-feature or multi-model ensemble methods.TheASCaccuracyachievedontheTUT2018,TAU2019, and JSSED datasets is 81.0%, 88.9% and 97.2%, respectively
Yuanbo Hou, Bo Kang, Wenwu Wang 0001, Jian Kang 0002, Dick Botteldooren
IEEE ACM Trans. Audio Speech Lang. Process.4
2024 AudioLDM 2: Learning Holistic Audio Generation With Self-Supervised Pretraining
abstract
Although audio generation shares commonalities across different types of audio, such as speech, music, and sound effects, designing models for each type requires careful consideration of specific objectives and biases that can significantly differ from those of other types. To bring us closer to a unified perspective of audio generation, this paper proposes a holistic framework that utilizes the same learning method for speech, music, and sound effect generation. Our framework utilizes a general representation of audio, called “language of audio” (LOA). Any audio can be translated into LOA based on AudioMAE, a self-supervised pre-trained representation learning model. In the generation process, we translate other modalities into LOA by using a GPT-2 model, and we perform self-supervised audio generation learning with a latent diffusion model conditioned on the LOA of audio in our training set. The proposed framework naturally brings advantages such as reusable self-supervised pretrained latent diffusion models. Experiments on the major benchmarks of text-to-audio, text-to-music, and text-to-speech with three AudioLDM 2 variants demonstrate competitive performance of the AudioLDM 2 variants framework against previous approaches. Our code, pretrained model, and demo are available athttps://audioldm.github.io/audioldm2.
Haohe Liu, Xubo Liu 0001, Xinhao Mei, Qiuqiang Kong, Qiao Tian 0001, Yuping Wang 0005, Wenwu Wang 0001, Yuxuan Wang 0002, Mark D. Plumbley
IEEE ACM Trans. Audio Speech Lang. Process.8
2024 Towards Generating Diverse Audio Captions via Adversarial Training
abstract
Automated audio captioning is a cross-modal translation task for describing the content of audio clips with natural language sentences. This task has attracted increasing attention and substantial progress has been made in recent years. Captions generated by existing models are generally faithful to the content of audio clips, however, these machine-generated captions are often deterministic (e.g., generating a fixed caption for a given audio clip), simple (e.g., using common words and simple grammar), and generic (e.g., generating the same caption for similar audio clips). When people are asked to describe the content of an audio clip, different people tend to focus on different sound events and describe an audio clip diversely from various aspects using distinct words and grammar. We believe that an audio captioning system should have the ability to generate diverse captions, either for a fixed audio clip, or across similar audio clips. To this end, we propose an adversarial training framework based on a conditional generative adversarial network (C-GAN) to improve diversity of audio captioning systems. A caption generator and two hybrid discriminators compete and are learned jointly, where the caption generator can be any standard encoder-decoder captioning model used to generate captions, and the hybrid discriminators assess the generated captions from different criteria, such as their naturalness and semantics. We conduct experiments on the Clotho dataset. The results show that our proposed model can generate captions with better diversity as compared to state-of-the-art methods.
Xinhao Mei, Xubo Liu 0001, Jianyuan Sun, Mark D. Plumbley, Wenwu Wang 0001
IEEE ACM Trans. Audio Speech Lang. Process.5
2024 WavCaps: A ChatGPT-Assisted Weakly-Labelled Audio Captioning Dataset for Audio-Language Multimodal Research
abstract
The advancement of audio-language (AL) multimodal learning tasks has been significant in recent years, yet the limited size of existing audio-language datasets poses challenges for researchers due to the costly and time-consuming collection process. To address this data scarcity issue, we introduceWavCaps, the first large-scale weakly-labelled audio captioning dataset, comprising approximately 400 k audio clips with paired captions. We sourced audio clips and their raw descriptions from web sources and a sound event detection dataset. However, the online-harvested raw descriptions are highly noisy and unsuitable for direct use in tasks such as automated audio captioning. To overcome this issue, we propose a three-stage processing pipeline for filtering noisy data and generating high-quality captions, where ChatGPT, a large language model, is leveraged to filter and transform raw descriptions automatically. We conduct a comprehensive analysis of the characteristics of WavCaps dataset and evaluate it on multiple downstream audio-language multimodal learning tasks. The systems trained on WavCaps outperform previous state-of-the-art (SOTA) models by a significant margin. Our aspiration is for the WavCaps dataset we have proposed to facilitate research in audio-language multimodal learning and demonstrate the potential of utilizing large language models (LLMs) to enhance academic research.
Xinhao Mei, Chutong Meng, Haohe Liu, Qiuqiang Kong, Tom Ko, Chengqi Zhao, Mark D. Plumbley, Yuexian Zou, Wenwu Wang 0001
IEEE ACM Trans. Audio Speech Lang. Process.9
2024 Introduction to the Special Issue on AI-Generated Content for Multimedia
abstract
Our world is becoming rapidly dependent on data of increasing complexity, diversity, and volume which calls for robust and powerful tools to process such big data. Probabilistic generative models fulfill this goal by learning latent characteristic data relations, especially for the recent emergence of large-scale deep generative models that are able to create realistic content, namely, artificial intelligence-generated content (AIGC). The applications of AIGC span across various domains, and witness rich potential in multimedia content creation, including dialog generation, text-to-speech conversion, image/video generation, and cross-modal content generation.
Shengxi Li, Xuelong Li 0001, Leonardo Chiariglione, Jiebo Luo 0001, Wenwu Wang 0001, Zhengyuan Yang, Danilo P. Mandic, Hamido Fujita
IEEE Trans. Circuits Syst. Video Technol.5
2024 A Survey of Cross-Modal Visual Content Generation
abstract
Cross-modal content generation has become very popular in recent years. To generate high-quality and realistic content, a variety of methods have been proposed. Among these approaches, visual content generation has attracted significant attention from academia and industry due to its vast potential in various applications. This survey provides an overview of recent advances in visual content generation conditioned on other modalities, such as text, audio, speech, and music, with a focus on their key contributions to the community. In addition, we summarize the existing publicly available datasets that can be used for training and benchmarking cross-modal visual content generation models. We provide an in-depth exploration of the datasets used for audio-to-visual content generation, filling a gap in the existing literature. Various evaluation metrics are also introduced along with the datasets. Furthermore, we discuss the challenges and limitations encountered in the area, such as modality alignment and semantic coherence. Last, we outline possible future directions for synthesizing visual content from other modalities including the exploration of new modalities, and the development of multi-task multi-modal networks. This survey serves as a resource for researchers interested in quickly gaining insights into this burgeoning field.
Fatemeh Nazarieh, Zhenhua Feng 0001, Muhammad Awais 0001, Wenwu Wang 0001, Josef Kittler
IEEE Trans. Circuits Syst. Video Technol.4
2024 Three-Direction Fusion for Accurate Volumetric Liver and Tumor Segmentation
abstract
Biomedical image segmentation of organs, tissues and lesions has gained increasing attention in clinical treatment planning and navigation, which involves the exploration of two-dimensional (2D) and three-dimensional (3D) contexts in the biomedical image. Compared to 2D methods, 3D methods pay more attention to inter-slice correlations, which offer additional spatial information for image segmentation. An organ or tumor has a 3D structure that can be observed from three directions. Previous studies focus only on the vertical axis, limiting the understanding of the relationship between a tumor and its surrounding tissues. Important information can also be obtained from sagittal and coronal axes. Therefore, spatial information of organs and tumors can be obtained from three directions, i.e. the sagittal, coronal and vertical axes, to understand better the invasion depth of tumor and its relationship with the surrounding tissues. Moreover, the edges of organs and tumors in biomedical image may be blurred. To address these problems, we propose a three-direction fusion volumetric segmentation (TFVS) model for segmenting 3D biomedical images from three perspectives in sagittal, coronal and transverse planes, respectively. We use the dataset of the liver task provided by the Medical Segmentation Decathlon challenge to train our model. The TFVS method demonstrates a competitive performance on the 3D-IRCADB dataset. In addition, the t-test and Wilcoxon signed-rank test are also performed to show the statistical significance of the improvement by the proposed method as compared with the baseline methods. The proposed method is expected to be beneficial in guiding and facilitating clinical diagnosis and treatment.
Feng Zhan, Wenwu Wang 0001, Yina Guo, Lidan He
IEEE J. Biomed. Health Informatics2
2024 Labelled Non-Zero Diffusion Particle Flow SMC-PHD Filtering for Multi-Speaker Tracking
abstract
Particle flow (PF) is a method originally proposed for single target tracking, and used recently to address the weight degeneracy problem of the sequential Monte Carlo probability hypothesis density (SMC-PHD) filter for audio-visual (AV) multi-speaker tracking, where the particle flow is calculated by using only the measurements near the particle, assuming that the target is detected, as in a recent method based on non-zero particle flow (NPF), i.e. the AV-NPF-SMC-PHD filter. This, however, can be problematic when occlusion happens and the occluded speaker may not be detected. To address this issue, we propose a new method where the labels of the particles are estimated using the likelihood function, and the particle flow is calculated in terms of the selected particles with the same labels. As a result, the particles associated with detected speakers and undetected speakers are distinguished based on the particle labels. With this novel method, named as AV-LPF-SMC-PHD, the speaker states can be estimated as the weighted mean of the labelled particles, which is computationally more efficient than using a clustering method as in the AV-NPF-SMC-PHD filter. The proposed algorithm is compared systematically with several baseline tracking methods using the AV16.3, AVDIAR and CLEAR datasets, and is shown to offer improved tracking accuracy with a lower computational cost.
Yang Liu 0175, Yong Xu 0004, Peipei Wu, Wenwu Wang 0001
IEEE Trans. Multim.4
2024 A Novel Composite Graph Neural Network
abstract
Graph neural networks (GNNs) have achieved great success in many fields due to their powerful capabilities of processing graph-structured data. However, most GNNs can only be applied to scenarios where graphs are known, but real-world data are often noisy or even do not have available graph structures. Recently, graph learning has attracted increasing attention in dealing with these problems. In this article, we develop a novel approach to improving the robustness of the GNNs, called composite GNN. Different from existing methods, our method uses composite graphs (C-graphs) to characterize both sample and feature relations. The C-graph is a unified graph that unifies these two kinds of relations, where edges between samples represent sample similarities, and each sample has a tree-based feature graph to model feature importance and combination preference. By jointly learning multiaspect C-graphs and neural network parameters, our method improves the performance of semisupervised node classification and ensures robustness. We conduct a series of experiments to evaluate the performance of our method and the variants of our method that only learn sample relations or feature relations. Extensive experimental results on nine benchmark datasets demonstrate that our proposed method achieves the best performance on almost all the datasets and is robust to feature noises.
Zhaogeng Liu, Jielong Yang, Xionghu Zhong, Wenwu Wang 0001, Hechang Chen, Yi Chang 0001
IEEE Trans. Neural Networks Learn. Syst.4
2024 Regression-Based Hyperparameter Learning for Support Vector Machines
abstract
Unification of classification and regression is a major challenge in machine learning and has attracted increasing attentions from researchers. In this article, we present a new idea for this challenge, where we convert the classification problem into a regression problem, and then use the methods in regression to solve the problem in classification. To this end, we leverage the widely used maximum margin classification algorithm and its typical representative, support vector machine (SVM). More specifically, we convert SVM into a piecewise linear regression task and propose a regression-based SVM (RBSVM) hyperparameter learning algorithm, where regression methods are used to solve several key problems in classification, such as learning of hyperparameters, calculation of prediction probabilities, and measurement of model uncertainty. To analyze the uncertainty of the model, we propose a new concept of model entropy, where the leave-one-out prediction probability of each sample is converted into entropy, and then used to quantify the uncertainty of the model. The model entropy is different from the classification margin, in the sense that it considers the distribution of all samples, not just the support vectors. Therefore, it can assess the uncertainty of the model more accurately than the classification margin. In the case of the same classification margin, the farther the sample distribution is from the classification hyperplane, the lower the model entropy. Experiments show that our algorithm (RBSVM) provides higher prediction accuracy and lower model uncertainty, when compared with state-of-the-art algorithms, such as Bayesian hyperparameter search and gradient-based hyperparameter learning algorithms.
Shili Peng, Wenwu Wang 0001, Yinli Chen, Xueling Zhong, Qinghua Hu
IEEE Trans. Neural Networks Learn. Syst.2
2023 Personalized Dialogue Generation with Persona-Adaptive Attention
abstract
Persona-based dialogue systems aim to generate consistent responses based on historical context and predefined persona. Unlike conventional dialogue generation, the persona-based dialogue needs to consider both dialogue context and persona, posing a challenge for coherent training. Specifically, this requires a delicate weight balance between context and persona. To achieve that, in this paper, we propose an effective framework with Persona-Adaptive Attention (PAA), which adaptively integrates the weights from the persona and context information via our designed attention. In addition, a dynamic masking mechanism is applied to the PAA to not only drop redundant information in context and persona but also serve as a regularization mechanism to avoid overfitting. Experimental results demonstrate the superiority of the proposed PAA framework compared to the strong baselines in both automatic and human evaluation. Moreover, the proposed PAA approach can perform equivalently well in a low-resource regime compared to models trained in a full-data setting, which achieve a similar result with only 20% to 30% of data compared to the larger models trained in the full-data setting. To fully exploit the effectiveness of our design, we designed several variants for handling the weighted information in different ways, showing the necessity and sufficiency of our weighting and masking designs.
Qiushi Huang, Yu Zhang 0006, Tom Ko, Xubo Liu 0001, Bo Wu 0018, Wenwu Wang 0001, Lilian Tang
AAAI6
2023 Learning Retrieval Augmentation for Personalized Dialogue Generation
abstract
Personalized dialogue generation, focusing on generating highly tailored responses by leveraging persona profiles and dialogue context, has gained significant attention in conversational AI applications.However, persona profiles, a prevalent setting in current personalized dialogue datasets, typically composed of merely four to five sentences, may not offer comprehensive descriptions of the persona about the agent, posing a challenge to generate truly personalized dialogues.To handle this problem, we propose Learning Retrieval Augmentation for Personalized DialOgue Generation (LAPDOG), which studies the potential of leveraging external knowledge for persona dialogue generation.Specifically, the proposed LAPDOG model consists of a story retriever and a dialogue generator.The story retriever uses a given persona profile as queries to retrieve relevant information from the story document, which serves as a supplementary context to augment the persona profile.The dialogue generator utilizes both the dialogue history and the augmented persona profile to generate personalized responses.For optimization, we adopt a joint training framework that collaboratively learns the story retriever and dialogue generator, where the story retriever is optimized towards desired ultimate metrics (e.g., BLEU) to retrieve content for the dialogue generator to generate personalized responses.Experiments conducted on the CONVAI2 dataset with ROCStory as a supplementary data source show that the proposed LAPDOG method substantially outperforms the baselines, indicating the effectiveness of the proposed method.The LAPDOG model code is publicly available for further exploration.
Qiushi Huang, Xubo Liu 0001, Wenwu Wang 0001, Tom Ko, Yu Zhang 0006, Lilian Tang
EMNLP4
2023 Siamese Network Based on MLP and Multi-head Cross Attention for Visual Object Tracking
Piaoyang Li, Shiyong Lan, Shipeng Sun, Wenwu Wang 0001, Yongyang Gao, Yongyu Yang, Guangyu Yu
ICANN (10)4
2023 GanNeXt: A New Convolutional GAN for Anomaly Detection
Bowei Pu, Shiyong Lan, Wenwu Wang 0001, Caiying Yang, Hongyu Yang 0002
ICANN (3)3
2023 Visual-Haptic-Kinesthetic Object Recognition with Multimodal Transformer
Xinyuan Zhou, Shiyong Lan, Wenwu Wang 0001, Hongyu Yang 0002
ICANN (7)3
2023 Time-Weighted Frequency Domain Audio Representation with GMM Estimator for Anomalous Sound Detection
abstract
Although deep learning is the mainstream method in unsupervised anomalous sound detection, Gaussian Mixture Model (GMM) with statistical audio frequency representation as input can achieve comparable results with much lower model complexity and fewer parameters. Existing statistical frequency representations, e.g. the log-Mel spectrogram’s average or maximum over time, do not always work well for different machines. This paper presents Time-Weighted Frequency Domain Representation (TWFR) with the GMM method (TWFR-GMM) for anomalous sound detection. The TWFR is a generalized statistical frequency domain representation that can adapt to different machine types, using the global weighted ranking pooling over time-domain. This allows GMM estimator to recognize anomalies, even under domain-shift conditions, as visualized with a Mahalanobis distance-based metric. Experiments on DCASE 2022 Challenge Task2 dataset show that our method has better detection performance than recent deep learning methods. TWFR-GMM is the core of our submission that achieved the 3rd place in DCASE 2022 Challenge Task2.
Jian Guan 0001, Youde Liu, Qiaoxi Zhu, Tieran Zheng, Jiqing Han 0001, Wenwu Wang 0001
ICASSP6
2023 Anomalous Sound Detection Using Audio Representation with Machine ID Based Contrastive Learning Pretraining
abstract
Existing contrastive learning methods for anomalous sound detection refine the audio representation of each audio sample by using the contrast between the samples’ augmentations (e.g., with time or frequency masking). However, they might be biased by the augmented data, due to the lack of physical properties of machine sound, thereby limiting the detection performance. This paper uses contrastive learning to refine audio representations for each machine ID, rather than for each audio sample. The proposed two-stage method uses contrastive learning to pretrain the audio representation model by incorporating machine ID and a self-supervised ID classifier to fine-tune the learnt model, while enhancing the relation between audio features from the same ID. Experiments show that our method outperforms the state-of-the-art methods using contrastive learning or self-supervised classification in overall anomaly detection performance and stability on DCASE 2020 Challenge Task2 dataset.
Jian Guan 0001, Feiyang Xiao, Youde Liu, Qiaoxi Zhu, Wenwu Wang 0001
ICASSP5
2023 Gct: Gated Contextual Transformer for Sequential Audio Tagging
abstract
Audio tagging aims to assign predefined tags to audio clips to indicate the class information of audio events. Sequential audio tagging (SAT) means detecting both the class information of audio events, and the order in which they occur within the audio clip. Most existing methods for SAT are based on connectionist temporal classification (CTC). However, CTC cannot effectively capture event connections due to the conditional independence assumption between outputs at different times. The contextual Transformer (cTransformer) addresses this issue by exploiting contextual information in SAT. Nevertheless, cTransformer is also limited in exploiting contextual information as it only uses forward information in inference. This paper proposes a gated contextual Transformer (GCT) with forward-backward inference (FBI). In addition, a gated contextual multi-layer perceptron (GCMLP) block is proposed in GCT to improve the performance of cTransformer structurally. Experiments on the two real-life audio datasets with manually annotated sequential labels show that the proposed GCT with GCMLP and FBI performs better than the CTC-based methods and cTransformer.
Yuanbo Hou, Wenwu Wang 0001, Dick Botteldooren
ICASSP3
2023 Simple Pooling Front-Ends for Efficient Audio Classification
abstract
Recently, there has been increasing interest in building efficient audio neural networks for on-device scenarios. Most existing approaches are designed to reduce the size of audio neural networks using methods such as model pruning. In this work, we show that instead of reducing model size using complex methods, eliminating the temporal redundancy in the input audio features (e.g., mel-spectrogram) could be an effective approach for efficient audio classification. To do so, we proposed a family of simple pooling front-ends (SimPFs) which use simple non-parametric pooling operations to reduce the redundant information within the mel-spectrogram. We perform extensive experiments on four audio classification tasks to evaluate the performance of SimPFs. Experimental results show that SimPFs can achieve a reduction in more than half of the number of floating point operations (FLOPs) for off-the-shelf audio neural networks, with negligible degradation or even some improvements in audio classification performance.
Xubo Liu 0001, Haohe Liu, Qiuqiang Kong, Xinhao Mei, Mark D. Plumbley, Wenwu Wang 0001
ICASSP6
2023 An Improved Optimal Transport Kernel Embedding Method with Gating Mechanism for Singing Voice Separation and Speaker Identification
abstract
Singing voice separation (SVS) and speaker identification (SI) are two classic problems in speech signal processing. Deep neural networks (DNNs) solve these two problems by extracting effective representations of the target signal from the input mixture. Since essential features of a signal can be well reflected on its latent geometric structure of the feature distribution, a natural way to address SVS/SI is to extract the geometry-aware and distribution-related features of the target signal. To do this, this work introduces the concept of optimal transport (OT) to SVS/SI and proposes an improved optimal transport kernel embedding (iOTKE) to extract the target-distribution-related features. The iOTKE learns an OT from the input signal to the target signal on the basis of a reference set learned from all training data. Thus it can maintain the feature diversity and preserve the latent geometric structure of the distribution for the target signal. To further improve the feature selection ability, we extend the proposed iOTKE to a gated version, i.e., gated iOTKE (G-iOTKE), by incorporating a lightweight gating mechanism. The gating mechanism controls effective information flow and enables the proposed method to select important features for a specific input signal. We evaluated the proposed G-iOTKE on SVS/SI. Experimental results showed that the proposed method provided better results than other models.
Weitao Yuan, Yuren Bian, Shengbei Wang, Masashi Unoki, Wenwu Wang 0001
ICASSP5
2023 DLAHSD: Dynamic Label Adopted In Auxiliary Head for SAR Detection
abstract
Ship detection in synthetic aperture radar (SAR) images is a major issue in maritime surveillance and port management. Existing challenges are mainly as follows: (1) Tiny ships are mixed with scattered noise spots on the sea. (2) Ships are present in extreme aspect-ratios and various scales. (3) The land background blurs the outline of coastal ships. To address these problems, we propose an efficient detection neural network (DLAHSD) that integrates the Multi-scale Feature Location Fusion (MFLF) module and the Auxiliary Detection Head (ADH) based CenterNet. In addition, we designed a Dynamic Elliptic Gaussian (DEG) module to label the heatmap of ships. Experimental results on the challenging SSDD dataset show that our model offers improved performance over the baseline methods. The codes will be available at https://github.com/SYLan2019/DLAHSD.
Xiaoxiao Yin, Shiyong Lan, Weikang Huang, Yitong Ma, Wenwu Wang 0001, Hongyu Yang 0002, Yilin Zheng
ICIP5
2023 A Semantics-Aware Normalizing Flow Model for Anomaly Detection
abstract
Anomaly detection in computer vision aims to detect outliers from input image data. Examples include texture defect detection and semantic discrepancy detection. However, existing methods are limited in detecting both types of anomalies, especially for the latter. In this work, we propose a novel semantics-aware normalizing flow model to address the above challenges. First, we employ the semantic features extracted from a backbone network as the initial input of the normalizing flow model, which learns the mapping from the normal data to a normal distribution according to semantic attributes, thus enhances the discrimination of semantic anomaly detection. Second, we design a new feature fusion module in the normalizing flow model to integrate texture features and semantic features, which can substantially improve the fitting of the distribution function with input data, thus achieving improved performance for the detection of both types of anomalies. Extensive experiments on five well-known datasets for semantic anomaly detection show that the proposed method outperforms the state-of-the-art baselines. The codes will be available at https://github.com/SYLan2019/SANF-AD.
Shiyong Lan, Weikang Huang, Wenwu Wang 0001, Hongyu Yang 0002, Yitong Ma, Yongjie Ma
ICME4
2023 AudioLDM: Text-to-Audio Generation with Latent Diffusion Models
abstract
Text-to-audio (TTA) systems have recently gained attention for their ability to synthesize general audio based on text descriptions. However, previous studies in TTA have limited generation quality with high computational costs. In this study, we propose AudioLDM, a TTA system that is built on a latent space to learn continuous audio representations from contrastive language-audio pretraining (CLAP) embeddings. The pretrained CLAP models enable us to train LDMs with audio embeddings while providing text embeddings as the condition during sampling. By learning the latent representations of audio signals without modelling the cross-modal relationship, AudioLDM improves both generation quality and computational efficiency. Trained on AudioCaps with a single GPU, AudioLDM achieves state-of-the-art TTA performance compared to other open-sourced systems, measured by both objective and subjective metrics. AudioLDM is also the first TTA system that enables various text-guided audio manipulations (e.g., style transfer) in a zero-shot fashion. Our implementation and demos are available at https://audioldm.github.io.
Haohe Liu, Zehua Chen 0005, Xinhao Mei, Xubo Liu 0001, Danilo P. Mandic, Wenwu Wang 0001, Mark D. Plumbley
ICML7
2023 Joint Prediction of Audio Event and Annoyance Rating in an Urban Soundscape by Hierarchical Graph Representation Learning
abstract
Sound events in daily life carry rich information about the objective world. The composition of these sounds affects the mood of people in a soundscape. Most previous approaches only focus on classifying and detecting audio events and scenes, but may ignore their perceptual quality that may impact humans' listening mood for the environment, e.g. annoyance. To this end, this paper proposes a novel hierarchical graph representation learning (HGRL) approach which links objective audio events (AE) with subjective annoyance ratings (AR) of the soundscape perceived by humans. The hierarchical graph consists of fine-grained event (fAE) embeddings with single-class event semantics, coarse-grained event (cAE) embeddings with multi-class event semantics, and AR embeddings. Experiments show the proposed HGRL successfully integrates AE with AR for AEC and ARP tasks, while coordinating the relations between cAE and fAE and further aligning the two different grains of AE information with the AR.
Yuanbo Hou, Siyang Song, Qiaoqiao Ren, Weicheng Xie 0001, Jian Kang 0002, Wenwu Wang 0001, Dick Botteldooren
INTERSPEECH8
2023 Adapting Language-Audio Models as Few-Shot Audio Learners
abstract
Contrastive language-audio pretraining (CLAP) has become a new paradigm to learn audio concepts with audio-text pairs. CLAP models have shown unprecedented performance as zero-shot classifiers on downstream tasks. To further adapt CLAP with domain-specific knowledge, a popular method is to finetune its audio encoder with available labelled examples. However, this is challenging in low-shot scenarios, as the amount of annotations is limited compared to the model size. In this work, we introduce a Training-efficient (Treff) adapter to rapidly learn with a small set of examples while maintaining the capacity for zero-shot classification. First, we propose a cross-attention linear model (CALM) to map a set of labelled examples and test audio to test labels. Second, we find initialising CALM as a cosine measurement improves our Treff adapter even without training. The Treff adapter outperforms metric-based methods in few-shot settings and yields competitive results to fully-supervised methods.
Jinhua Liang, Xubo Liu 0001, Haohe Liu, Huy Phan, Emmanouil Benetos, Mark D. Plumbley, Wenwu Wang 0001
INTERSPEECH7
2023 Visually-Aware Audio Captioning With Adaptive Audio-Visual Attention
abstract
Audio captioning aims to generate text descriptions of audio clips.In the real world, many objects produce similar sounds.How to accurately recognize ambiguous sounds is a major challenge for audio captioning.In this work, inspired by inherent human multimodal perception, we propose visuallyaware audio captioning, which makes use of visual information to help the description of ambiguous sounding objects.Specifically, we introduce an off-the-shelf visual encoder to extract video features and incorporate the visual features into an audio captioning system.Furthermore, to better exploit complementary audio-visual contexts, we propose an audio-visual attention mechanism that adaptively integrates audio and visual context and removes the redundant information in the latent space.Experimental results on AudioCaps, the largest audio captioning dataset, show that our proposed method achieves state-of-theart results on machine translation metrics.
Xubo Liu 0001, Qiushi Huang, Xinhao Mei, Haohe Liu, Qiuqiang Kong, Jianyuan Sun, Shengchen Li, Tom Ko, Yu Zhang 0006, Lilian Tang, Mark D. Plumbley, Volkan Kilic, Wenwu Wang 0001
INTERSPEECH13
2023 Ontology-aware Learning and Evaluation for Audio Tagging
abstract
This study defines a new evaluation metric for audio tagging tasks to alleviate the limitation of the mean average precision (mAP) metric.The mAP metric treats different kinds of sound as independent classes without considering their relations.The proposed metric, ontology-aware mean average precision (OmAP), addresses the weaknesses of mAP by utilizing additional ontology during evaluation.Specifically, we reweight the false positive events in the model prediction based on the AudioSet ontology graph distance to the target classes.The OmAP also provides insights into model performance by evaluating different coarse-grained levels in the ontology graph.We conduct a human assessment and show that OmAP is more consistent with human perception than mAP.We also propose an ontology-based loss function (OBCE) that reweights binary cross entropy (BCE) loss based on the ontology distance.Our experiment shows that OBCE can improve both mAP and OmAP metrics on the AudioSet tagging task.
Haohe Liu, Qiuqiang Kong, Xubo Liu 0001, Xinhao Mei, Wenwu Wang 0001, Mark D. Plumbley
INTERSPEECH5
2023 Dual Transformer Decoder based Features Fusion Network for Automated Audio Captioning
abstract
Automated audio captioning (AAC) which generates textual descriptions of audio content.Existing AAC models achieve good results but only use the high-dimensional representation of the encoder.There is always insufficient information learning of high-dimensional methods owing to high-dimensional representations having a large amount of information.In this paper, a new encoder-decoder model called the Lowand High-Dimensional Feature Fusion (LHDFF) is proposed.LHDFF uses a new PANNs encoder called Residual PANNs (RPANNs) to fuse low-and high-dimensional features.Lowdimensional features contain limited information about specific audio scenes.The fusion of low-and high-dimensional features can improve model performance by repeatedly emphasizing specific audio scene information.To fully exploit the fused features, LHDFF uses a dual transformer decoder structure to generate captions in parallel.Experimental results show that LHDFF outperforms existing audio captioning models.
Jianyuan Sun, Xubo Liu 0001, Xinhao Mei, Volkan Kilic, Mark D. Plumbley, Wenwu Wang 0001
INTERSPEECH6
2023 End-to-end translation of human neural activity to speech with a dual-dual generative adversarial network
Yina Guo, Anhong Wang, Wenwu Wang 0001
Knowl. Based Syst.5
2023 Discriminative analysis dictionary learning with adaptively ordinal locality preserving
Jing Dong 0001, Kai Wu 0004, Chang Liu 0152, Xue Mei, Wenwu Wang 0001
Neural Networks5
2023 Audio Event-Relational Graph Representation Learning for Acoustic Scene Classification
abstract
Most deep learning-based acoustic scene classification (ASC) approaches identify scenes based on acoustic features converted from audio clips containing mixed information entangled by polyphonic audio events (AEs). However, these approaches have difficulties in explaining what cues they use to identify scenes. This paper conducts the first study on disclosing the relationship between real-life acoustic scenes and semantic embeddings from the most relevant AEs. Specifically, we propose an event-relational graph representation learning (ERGL) framework for ASC to classify scenes, and simultaneously answer clearly and straightly which cues are used in classifying. In the event-relational graph, embeddings of each event are treated as nodes, while relationship cues derived from each pair of nodes are described by multi-dimensional edge features. Experiments on a real-life ASC dataset show that the proposed ERGL achieves competitive performance on ASC by learning embeddings of only a limited number of AEs. The results show the feasibility of recognizing diverse acoustic scenes based on the audio event-relational graph. Visualizations of graph representations learned by ERGL are available here(https://github.com/Yuanbo2020/ERGL).
Yuanbo Hou, Siyang Song, Chuang Yu 0001, Wenwu Wang 0001, Dick Botteldooren
IEEE Signal Process. Lett.4
2023 Graph Attention for Automated Audio Captioning
abstract
State-of-the-art audio captioning methods typically use the encoder-decoder structure with pretrained audio neural networks (PANNs) as encoders for feature extraction. However, the convolution operation used in PANNs is limited in capturing the long-time dependencies within an audio signal, thereby leading to potential performance degradation in audio captioning. This letter presents a novel method using graph attention (GraphAC) for encoder-decoder based audio captioning. In the encoder, a graph attention module is introduced after the PANNs to learn contextual association (i.e. the dependency among the audio features over different time frames) through an adjacency graph, and a top-kmask is used to mitigate the interference from noisy nodes. The learnt contextual association leads to a more effective feature representation with feature node aggregation. As a result, the decoder can predict important semantic information about the acoustic scene and events based on the contextual associations learned from the audio signal. Experimental results show that GraphAC outperforms the state-of-the-art methods with PANNs as the encoders, thanks to the incorporation of the graph attention module into the encoder for capturing the long-time dependencies within the audio signal. The source code is available at https://github.com/LittleFlyingSheep/GraphAC.
Feiyang Xiao, Jian Guan 0001, Qiaoxi Zhu, Wenwu Wang 0001
IEEE Signal Process. Lett.4
2023 U-Shaped Transformer With Frequency-Band Aware Attention for Speech Enhancement
Yi Li 0047, Yang Sun 0003, Wenwu Wang 0001, Syed M. Naqvi
IEEE ACM Trans. Audio Speech Lang. Process.3
2023 Unsupervised Deep Unfolded Representation Learning for Singing Voice Separation
abstract
Learning effective vocal representations from a waveform mixture is a crucial but challenging task for deep neural network (DNN)-based singing voice separation (SVS). Successful representation learning (RL) depends heavily on well-designed neural architectures and effective general priors. However, DNNs for RL in SVS are mostly built on generic architectures without general priors being systematically considered. To address these issues, we introduce deep unfolding to RL and propose two RL-based models for SVS, deep unfolded representation learning (DURL) and optimal transport DURL (OT-DURL). In both models, we formulate RL as a sequence of optimization problems for signal reconstruction, where three general priors, synthesis, non-negative, and our novel analysis, are incorporated. In DURL and OT-DURL, we take different approaches in penalizing the analysis prior. DURL uses the Euclidean distance as its penalty, while OT-DURL uses a more sophisticated penalty known as the OT distance. We address the optimization problems in DURL and OT-DURL with the first-order operator splitting algorithm and unfold the obtained iterative algorithms to novel encoders, by mapping the synthesis/analysis/non-negative priors to different interpretable sublayers of the encoders. We evaluated these DURL and OT-DURL encoders in the unsupervised informed SVS and supervised Open-Unmix frameworks. Experimental results indicate that (1) the OT-DURL encoder is better than the DURL encoder and (2) both encoders can considerably improve the vocal-signal-separation performance compared with those of the baseline model.
Weitao Yuan, Shengbei Wang, Masashi Unoki, Wenwu Wang 0001
IEEE ACM Trans. Audio Speech Lang. Process.5
2023 ACTUAL: Audio Captioning With Caption Feature Space Regularization
abstract
Audio captioning aims at describing the content of audio clips with human language. Due to the ambiguity of audio content, different people may perceive the same audio clip differently, resulting in caption disparities (i.e., the same audio clip may be described by several captions with diverse semantics). In the literature, the one-to-many strategy is often employed to train the audio captioning models, where a related caption is randomly selected as the optimization target for each audio clip at each training iteration. However, we observe that this can lead to significant variations during the optimization process and adversely affect the performance of the model. In this paper, we address this issue by proposing an audio captioning method, named ACTUAL (Audio Captioning with capTion featUre spAce reguLarization). ACTUAL involves a two-stage training process: (i) in the first stage, we use contrastive learning to construct a proxy feature space where the similarities between captions at the audio level are explored, and (ii) in the second stage, the proxy feature space is utilized as additional supervision to improve the optimization of the model in a more stable direction. We conduct extensive experiments to demonstrate the effectiveness of the proposed ACTUAL method. The results show that proxy caption embedding can significantly improve the performance of the baseline model and the proposed ACTUAL method offers competitive performance on two datasets compared to state-of-the-art methods. The code is publicly available athttps://github.com/PRIS-CV/Caption-Feature-Space-Regularization.
Yiming Zhang 0025, Hong Yu 0006, Ruoyi Du, Zheng-Hua Tan, Wenwu Wang 0001, Zhanyu Ma
IEEE ACM Trans. Audio Speech Lang. Process.5
2023 Audio-Visual Event Localization by Learning Spatial and Semantic Co-Attention
abstract
This work aims to temporally localize events that are both audible and visible in video. Previous methods mainly focused on temporal modeling of events with simple fusion of audio and visual features. In natural scenes, a video records not only the events of interest but also ambient acoustic noise and visual background, resulting in redundant information in the raw audio and visual features. Thus, direct fusion of the two features often causes false localization of the events. In this paper, we propose a co-attention model to exploit the spatial and semantic correlations between the audio and visual features, which helps guide the extraction of discriminative features for better event localization. Our assumption is that in an audio-visual event, shared semantic information between audio and visual features exists and can be extracted by attention learning. Specifically, the proposed co-attention model is composed of a co-spatial attention module and a co-semantic attention module that are used to model the spatial and semantic correlations, respectively. The proposed co-attention model can be applied to various event localization tasks, such as cross-modality localization and multimodal event localization. Experiments on the public audio-visual event (AVE) dataset demonstrate that the proposed method achieves state-of-the-art performance by learning spatial and semantic co-attention.
Xionghu Zhong, Minjie Cai, Hao Chen 0002, Wenwu Wang 0001
IEEE Trans. Multim.5
2022 UAV-enabled Edge Computing for Optimal Task Distribution in Target Tracking
Shidrokh Goudarzi, Wenwu Wang 0001, Pei Xiao 0001, Lyudmila Mihaylova, Simon J. Godsill
FUSION2
2022 Deep Learning for Audio Visual Emotion Recognition
Tassadaq Hussain, Wenwu Wang 0001, Nidhal Bouaynaya, Hassan M. Fathallah-Shaykh, Lyudmila Mihaylova
FUSION2
2022 Face Super-Resolution with Spatial Attention Guided by Multiscale Receptive-Field Features
Weikang Huang, Shiyong Lan, Wenwu Wang 0001, Xuedong Yuan, Hongyu Yang 0002, Piaoyang Li
ICANN (1)3
2022 A Transformer-Based GAN for Anomaly Detection
Caiyin Yang, Shiyong Lan, Weikang Huang, Wenwu Wang 0001, Hongyu Yang 0002, Piaoyang Li
ICANN (2)4
2022 Anomalous Sound Detection Using Spectral-Temporal Information Fusion
abstract
Unsupervised anomalous sound detection aims to detect unknown abnormal sounds of machines from normal sounds. However, the state-of-the-art approaches are not always stable and perform dramatically differently even for machines of the same type, making it impractical for general applications. This paper proposes a spectral-temporal fusion based self-supervised method to model the feature of the normal sound, which improves the stability and performance consistency in detection of anomalous sounds from individual machines, even of the same type. Experiments on the DCASE 2020 Challenge Task 2 dataset show that the proposed method achieved 81.39%, 83.48%, 98.22% and 98.83% in terms of the minimum AUC (worst-case detection performance amongst individuals) in four types of real machines (fan, pump, slider and valve), respectively, giving 31.79%, 17.78%, 10.42% and 21.13% improvement compared to the state-of-the-art method, i.e., Glow_Aff. Moreover, the proposed method has improved AUC (average performance of individuals) for all the types of machines in the dataset. The source codes are available at https://github.com/liuyoude/STgram_MFN
Youde Liu, Jian Guan 0001, Qiaoxi Zhu, Wenwu Wang 0001
ICASSP4
2022 Diverse Audio Captioning Via Adversarial Training
abstract
Audio captioning aims at generating natural language descriptions for audio clips automatically. Existing audio captioning models have shown promising improvement in recent years. However, these models are mostly trained via maximum likelihood estimation (MLE), which tends to make captions generic, simple and deterministic. As different people may describe an audio clip from different aspects using distinct words and grammars, we argue that an audio captioning system should have the ability to generate diverse captions for a fixed audio clip and across similar audio clips. To address this problem, we propose an adversarial training framework for audio captioning based on a conditional generative adversarial network (C-GAN), which aims at improving the naturalness and diversity of generated captions. Unlike processing data of continuous values in a classical GAN, a sentence is composed of discrete tokens and the discrete sampling process is non-differentiable. To address this issue, policy gradient, a reinforcement learning technique, is used to back-propagate the reward to the generator. The results show that our proposed model can generate more diverse captions, as compared to state-of-the-art methods.
Xinhao Mei, Xubo Liu 0001, Jianyuan Sun, Mark D. Plumbley, Wenwu Wang 0001
ICASSP5
2022 Partial Arithmetic Consensus based Distributed Intensity Particle Flow SMC-PHD Filter for Multi-Target Tracking
abstract
Intensity Particle Flow (IPF) SMC-PHD has been proposed recently for multi-target tracking. In this paper, we extend IPF-SMC-PHD filter to distributed setting, and develop a novel consensus method for fusing the estimates from individual sensors, based on Arithmetic Average (AA) fusion. Different from conventional AA method which may be degraded when unreliable estimates are presented, we develop a novel arithmetic consensus method to fuse estimates from each individual IPF-SMC-PHD filter with partial consensus. The proposed method contains a scheme for evaluating the reliability of the sensor nodes and preventing unreliable sensor information to be used in fusion and communication in sensor network, which help improve fusion accuracy and reduce sensor communication costs. Numerical simulations are performed to demonstrate the advantages of the proposed algorithm over the uncooperative IPF-SMC-PHD and distributed particle-PHD with AA fusion.
Peipei Wu, Jinzheng Zhao, Shidrokh Goudarzi, Wenwu Wang 0001
ICASSP4
2022 A Mutual Learning Framework for Few-Shot Sound Event Detection
abstract
Although prototypical network (ProtoNet) has proved to be an effective method for few-shot sound event detection, two problems still exist. Firstly, the small-scaled support set is insufficient so that the class prototypes may not represent the class center accurately. Secondly, the feature extractor is task-agnostic (or class-agnostic): the feature extractor is trained with base-class data and directly applied to unseen-class data. To address these issues, we present a novel mutual learning framework with transductive learning, which aims at iteratively updating the class prototypes and feature extractor. More specifically, we propose to update class prototypes with transductive inference to make the class prototypes as close to the true class center as possible. To make the feature extractor to be task-specific, we propose to use the updated class prototypes to fine-tune the feature extractor. After that, a fine-tuned feature extractor further helps produce better class prototypes. Our method achieves the F-score of 38.4% on the DCASE 2021 Task 5 evaluation set, which won the first place in the few-shot bioacoustic event detection task of Detection and Classification of Acoustic Scenes and Events (DCASE) 2021 Challenge.
Dongchao Yang, Helin Wang, Yuexian Zou, Zhongjie Ye, Wenwu Wang 0001
ICASSP5
2022 Audio-Visual Tracking of Multiple Speakers Via a PMBM Filter
abstract
Audio-visual tracking of multiple speakers requires to estimate the state (e.g. velocity and location) of each speaker by leveraging the information of both audio and visual modalities. Estimating the number of speakers and their states jointly remains a challenging problem. We propose an Audio-Visual Possion Multi-Bernoulli Mixture Filter (AV-PMBM) that can not only predict the number of speakers but also give accurate estimation of their states. We also propose a novel sound source localization technique based on DOA information and a deep learning based object detector to provide reliable audio measurements for the AV tracker. To our knowledge, this represents the first attempt using PMBM for multi-speaker tracking with audio visual modalities. Experiments on the AV16.3 dataset demonstrate that AV-PMBM achieves state-of-the-art performance in optimal sub-pattern assignment (OSPA).
Jinzheng Zhao, Peipei Wu, Xubo Liu 0001, Yong Xu 0004, Lyudmila Mihaylova, Simon J. Godsill, Wenwu Wang 0001
ICASSP7
2022 DSTAGNN: Dynamic Spatial-Temporal Aware Graph Neural Network for Traffic Flow Forecasting
abstract
As a typical problem in time series analysis, traffic flow prediction is one of the most important application fields of machine learning. However, achieving highly accurate traffic flow prediction is a challenging task, due to the presence of complex dynamic spatial-temporal dependencies within a road network. This paper proposes a novel Dynamic Spatial-Temporal Aware Graph Neural Network (DSTAGNN) to model the complex spatial-temporal interaction in road network. First, considering the fact that historical data carries intrinsic dynamic information about the spatial structure of road networks, we propose a new dynamic spatial-temporal aware graph based on a data-driven strategy to replace the pre-defined static graph usually used in traditional graph convolution. Second, we design a novel graph neural network architecture, which can not only represent dynamic spatial relevance among nodes with an improved multi-head attention mechanism, but also acquire the wide range of dynamic temporal dependency from multi-receptive field features via multi-scale gated convolution. Extensive experiments on real-world data sets demonstrate that our proposed method significantly outperforms the state-of-the-art methods.
Shiyong Lan, Yitong Ma, Weikang Huang, Wenwu Wang 0001, Hongyu Yang 0002, Pyang Li
ICML4
2022 Separate What You Describe: Language-Queried Audio Source Separation
abstract
In this paper, we introduce the task of language-queried audio source separation (LASS), which aims to separate a target source from an audio mixture based on a natural language query of the target source (e.g., "a man tells a joke followed by people laughing"). A unique challenge in LASS is associated with the complexity of natural language description and its relation with the audio sources. To address this issue, we proposed LASS-Net, an end-to-end neural network that is learned to jointly process acoustic and linguistic information, and separate the target source that is consistent with the language query from an audio mixture. We evaluate the performance of our proposed system with a dataset created from the AudioCaps dataset. Experimental results show that LASS-Net achieves considerable improvements over baseline methods. Furthermore, we observe that LASS-Net achieves promising generalization results when using diverse human-annotated descriptions as queries, indicating its potential use in real-world scenarios. The separated audio samples and source code are available at https://liuxubo717.github.io/LASS-demopage.
Xubo Liu 0001, Haohe Liu, Qiuqiang Kong, Xinhao Mei, Jinzheng Zhao, Qiushi Huang, Mark D. Plumbley, Wenwu Wang 0001
INTERSPEECH8
2022 On Metric Learning for Audio-Text Cross-Modal Retrieval
abstract
Audio-text retrieval aims at retrieving a target audio clip or caption from a pool of candidates given a query in another modality. Solving such cross-modal retrieval task is challenging because it not only requires learning robust feature representations for both modalities, but also requires capturing the fine-grained alignment between these two modalities. Existing cross-modal retrieval models are mostly optimized by metric learning objectives as both of them attempt to map data to an embedding space, where similar data are close together and dissimilar data are far apart. Unlike other cross-modal retrieval tasks such as image-text and video-text retrievals, audio-text retrieval is still an unexplored task. In this work, we aim to study the impact of different metric learning objectives on the audio-text retrieval task. We present an extensive evaluation of popular metric learning objectives on the AudioCaps and Clotho datasets. We demonstrate that NT-Xent loss adapted from self-supervised learning shows stable performance across different datasets and training settings, and outperforms the popular triplet-based losses. Our code is available at https://github.com/XinhaoMei/audio-text_retrieval.
Xinhao Mei, Xubo Liu 0001, Jianyuan Sun, Mark D. Plumbley, Wenwu Wang 0001
INTERSPEECH5
2022 RaDur: A Reference-aware and Duration-robust Network for Target Sound Detection
abstract
Target sound detection (TSD) aims to detect the target sound from a mixture audio given the reference information.Previous methods use a conditional network to extract a sounddiscriminative embedding from the reference audio, and then use it to detect the target sound from the mixture audio.However, the network performs much differently when using different reference audios (e.g.performs poorly for noisy and shortduration reference audios), and tends to make wrong decisions for transient events (i.e.shorter than 1 second).To overcome these problems, in this paper, we present a reference-aware and duration-robust network (RaDur) for TSD.More specifically, in order to make the network more aware of the reference information, we propose an embedding enhancement module to take into account the mixture audio while generating the embedding, and apply the attention pooling to enhance the features of target sound-related frames and weaken the features of noisy frames.In addition, a duration-robust focal loss is proposed to help model different-duration events.To evaluate our method, we build two TSD datasets based on UrbanSound and Audioset.Extensive experiments show the effectiveness of our methods.
Dongchao Yang, Helin Wang, Zhongjie Ye, Yuexian Zou, Wenwu Wang 0001
INTERSPEECH5
2022 Audio Visual Multi-Speaker Tracking with Improved GCF and PMBM Filter
Jinzheng Zhao, Peipei Wu, Xubo Liu 0001, Shidrokh Goudarzi, Haohe Liu, Yong Xu 0004, Wenwu Wang 0001
INTERSPEECH7
2022 A Hybrid Approach to Blind Video Quality Prediction of User Generated Content
abstract
Increased use of social media platforms has resulted in a vast amount of user-generated video content being released to the internet daily. Measuring and monitoring the perceptual quality of these videos is vital for efficient network and storage management. However, these videos do not have a pristine reference, posing challenges for accurate quality monitoring. In this paper, we introduce a hybrid metric to measure the perceptual quality of user-generated video content using both pixel-level and compression-level features. Our experiments on large-scale databases of user-generated content show that the proposed method performs comparably in predicting the perceptual quality when compared with state-of-the-art metrics.
Buddhiprabha Erabadda, Gosala Kulupana, Thanuja Mallikarachchi, Wenwu Wang 0001, Warnakulasuriya Anil Chandana Fernando
PCS4
2022 Multiscale Deep Neural Network With Two-Stage Loss for SAR Target Recognition With Small Training Set
abstract
Deep learning models have been used recently for target recognition from synthetic aperture radar (SAR) images. However, the performance of these models tends to deteriorate when only a small number of training samples are available due to the problem of overfitting. To address this problem, we propose a two-stage multiscale densely connected convolutional neural networks (TMDC-CNNs). In the proposed TMDC-CNNs, the overfitting issue is addressed with a novel multiscale densely connected network architecture and a two-stage loss function, which integrated the cosine similarity with the prevailing softmax cross-entropy loss. Experiments were conducted on the MSTAR data set, and the results show that our model offers significant recognition accuracy improvements as compared with other state-of-the-art methods, with severely limited training data. The source codes are available athttps://github.com/Stubsx/TMDC-CNNs.
Jian Guan 0001, Jiabei Liu, Pengming Feng, Wenwu Wang 0001
IEEE Geosci. Remote. Sens. Lett.4
2022 Support vector machine embedding discriminative dictionary pair learning for pattern classification
Jing Dong 0001, Liu Yang 0021, Chang Liu 0152, Wei Cheng 0005, Wenwu Wang 0001
Neural Networks5
2022 Local Information Assisted Attention-Free Decoder for Audio Captioning
abstract
Automated audio captioning aims to describe audio data with captions using natural language. Existing methods often employ an encoder-decoder structure, where the attention-based decoder (e.g., Transformer decoder) is widely used and achieves state-of-the-art performance. Although this method effectively captures global information within audio data via the self-attention mechanism, it may ignore the event with short time duration, due to its limitation in capturing local information in an audio signal, leading to inaccurate prediction of captions. To address this issue, we propose a method using the pretrained audio neural networks (PANNs) as the encoder and local information assisted attention-free Transformer (LocalAFT) as the decoder. The novelty of our method is in the proposal of the LocalAFT decoder, which allows local information within an audio signal to be captured while retaining the global information. This enables the events of different duration, including short duration, to be captured for more precise caption generation. Experiments show that our method outperforms the state-of-the-art methods in Task 6 of the DCASE 2021 Challenge with the standard attention-based decoder for caption generation.
Feiyang Xiao, Jian Guan 0001, Haiyan Lan, Qiaoxi Zhu, Wenwu Wang 0001
IEEE Signal Process. Lett.5
2022 Acoustic Source Localization in the Circular Harmonic Domain Using Deep Learning Architecture
abstract
The problem of direction of arrival (DOA) estima- tion with a circular microphone array has been addressed with classical source localization methods, such as the model-based methods and the parametric methods. These methods have an ad- vantage in estimating the DOAs in a blind manner, i.e. with no (or limited) prior knowledge about the sound sources. However, their performance tends to degrade rapidly in noisy and reverberant environments or in the presence of sensor array limitations, such as sensor gain and phase errors. In this paper, we present a new approach by leveraging the strength of a convolutional neural network (CNN)-based deep learning approach. In particular, we design new circular harmonic features that are frequency- invariant as inputs to the CNN architecture, so as to offer improvements in DOA estimation in unseen adverse environments and obtain good adaptation to array imperfections. To our knowledge, such a deep learning approach has not been used in the circular harmonic domain. Experiments performed on both simulated and real-data show that our method gives significantly better performance, than the recent baseline methods, in a variety of noise and reverberation levels, in terms of the accuracy of the DOA estimation.
Kunkun SongGong, Wenwu Wang 0001
IEEE ACM Trans. Audio Speech Lang. Process.2
2021 An Improved Event-Independent Network for Polyphonic Sound Event Localization and Detection
abstract
Polyphonic sound event localization and detection (SELD), which jointly performs sound event detection (SED) and direction-of-arrival (DoA) estimation, detects the type and occurrence time of sound events as well as their corresponding DoA angles simultaneously. We study the SELD task from a multi-task learning perspective. Two open problems are addressed in this paper. Firstly, to detect overlapping sound events of the same type but with different DoAs, we propose to use a trackwise output format and solve the accompanying track permutation problem with permutation-invariant training. Multi-head self-attention is further used to separate tracks. Secondly, a previous finding is that, by using hard parameter-sharing, SELD suffers from a performance loss compared with learning the subtasks separately. This is solved by a soft parameter-sharing scheme. We term the proposed method as Event Independent Network V2 (EINV2), which is an improved version of our previously-proposed method and an end-to-end network for SELD. We show that our proposed EINV2 for joint SED and DoA estimation outperforms previous methods by a large margin, and has comparable performance to state-of-the-art ensemble models.
Yin Cao, Turab Iqbal, Qiuqiang Kong, Fengyan An, Wenwu Wang 0001, Mark D. Plumbley
ICASSP5
2021 Low-Dimensional Denoising Embedding Transformer for ECG Classification
abstract
The transformer based model (e.g., FusingTF) has been employed recently for Electrocardiogram (ECG) signal classification. However, the high-dimensional embedding obtained via 1-D convolution and positional encoding can lead to the loss of the signal’s own temporal information and a large amount of training parameters. In this paper, we propose a new method for ECG classification, called low-dimensional denoising embedding transformer (LDTF), which contains two components, i.e., low-dimensional denoising embedding (LDE) and transformer learning. In the LDE component, a low-dimensional representation of the signal is obtained in the time-frequency domain while preserving its own temporal information. And with the low-dimensional embedding, the transformer learning is then used to obtain a deeper and narrower structure with fewer training parameters than that of the FusingTF. Experiments conducted on the MIT-BIH dataset demonstrates the effectiveness and the superior performance of our proposed method, as compared with state-of-the-art methods.
Jian Guan 0001, Pengming Feng, Wenwu Wang 0001
ICASSP5
2021 Enhancing Audio Augmentation Methods with Consistency Learning
abstract
Data augmentation is an inexpensive way to increase training data diversity, and is commonly achieved via transformations of existing data. For tasks such as classification, there is a good case for learning representations of the data that are invariant to such transformations, yet this is not explicitly enforced by classification losses such as the cross-entropy loss. This paper investigates the use of training objectives that explicitly impose this consistency constraint, and how it can impact downstream audio classification tasks. In the context of deep convolutional neural networks in the supervised setting, we show empirically that certain measures of consistency are not implicitly captured by the cross-entropy loss, and that incorporating such measures into the loss function can improve the performance of tasks such as audio tagging. Put another way, we demonstrate how existing augmentation methods can further improve learning by enforcing consistency.
Turab Iqbal, Karim Helwani, Arvindh Krishnaswamy, Wenwu Wang 0001
ICASSP4
2021 Dimension Selected Subspace Clustering
abstract
Subspace clustering is a popular method for clustering unlabelled data. However, the computational cost of the subspace clustering algorithm can be unaffordable when dealing with a large data set. Using a set of dimension sketched data instead of the original data set can be helpful for mitigating the computational burden. Thus, finding a way for dimension sketching becomes an important problem. In this paper, a new dimension sketching algorithm is proposed, which aims to select informative dimensions that have significant effects on the clustering results. Experimental results reveal that this method can significantly improve subspace clustering performance on both synthetic and real-world datasets, in comparison with two baseline methods.
Shuoyang Li, Yuhui Luo, Jonathon A. Chambers, Wenwu Wang 0001
ICASSP4
2021 A Global-Local Attention Framework for Weakly Labelled Audio Tagging
abstract
Weakly labelled audio tagging aims to predict the classes of sound events within an audio clip, where the onset and offset times of the sound events are not provided. Previous works have used the multiple instance learning (MIL) framework, and exploited the information of the whole audio clip by MIL pooling functions. However, the detailed information of sound events such as their durations may not be considered under this framework. To address this issue, we propose a novel two-stream framework for audio tagging by exploiting the global and local information of sound events. The global stream aims to analyze the whole audio clip in order to capture the local clips that need to be attended using a class-wise selection module. These clips are then fed to the local stream to exploit the detailed information for a better decision. Experimental results on the AudioSet show that our proposed method can significantly improve the performance of audio tagging under different baseline network architectures.
Helin Wang, Yuexian Zou, Wenwu Wang 0001
ICASSP3
2021 Weighted Magnitude-Phase Loss for Speech Dereverberation
abstract
In real rooms, recorded speech usually contains reverberation, which degrades the quality and intelligibility of the speech. It has proven effective to use neural networks to estimate complex ideal ratio masks (cIRMs) using mean square error (MSE) loss for speech dereverberation. However, in some cases, when using MSE loss to estimate complex-valued masks, phase may have a disproportionate effect compared to magnitude. We propose a new weighted magnitude-phase loss function, which is divided into a magnitude component and a phase component, to train a neural network to estimate complex ideal ratio masks. A weight parameter is introduced to adjust the relative contribution of magnitude and phase to the overall loss. We find that our proposed loss function outperforms the regular MSE loss function for speech dereverberation.
Jingshu Zhang, Mark D. Plumbley, Wenwu Wang 0001
ICASSP3
2021 Robust Visual Object Tracking with Spatiotemporal Regularisation and Discriminative Occlusion Deformation
abstract
Spatiotemporal regularized Discriminative Correlation Filters (DCF) have been proposed recently for visual tracking, achieving state-of-the-art performance. However, the tracking performance of the online learning model used in this kind methods is highly dependent on the quality of the appearance feature of the target, and the target feature appearance could be heavily deformed due to the occlusion by other objects or the variations in their dynamic self-appearance. In this paper, we propose a new approach to mitigate these two kinds of appearance deformation. Firstly, we embed the occlusion perception block into the model update stage, then we adaptively adjust the model update according to the situation of occlusion. Secondly, we use the relatively stable colour statistics to deal with the appearance shape changes in large targets, and compute the histogram response scores as a complementary part of final correlation response. Extensive experiments are performed on four well-known datasets, i.e. OTB100, VOT-2018, UAV123, and TC128. The results show that the proposed approach outperforms the baseline DCF method, especially, on the TC128/UAV123 datasets, with a gain of over 4.05 %/2.43% in mean overlap precision. We will release our code at https://github.com/SYLan2019/STD0D.
Shiyong Lan, Shipeng Sun, Wenwu Wang 0001
ICIP5
2021 SAGAN: Skip-Attention GAN For Anomaly Detection
abstract
Generative Adversarial Networks (GANs) have been used recently for anomaly detection from images, where the anomaly scores are obtained by comparing the global difference between the input and generated image. However, the anomalies often appear in local areas of an image scene, and ignoring such information can lead to unreliable detection of anomalies. In this paper, we propose an efficient anomaly detection network Skip-Attention GAN (SAGAN), which adds attention modules to capture local information to improve the accuracy of latent representation of images, and uses depth-wise separable convolutions to reduce the number of parameters in the model. We evaluate the proposed method on the CIFAR-10 dataset and the LBOT dataset (built by ourselves), and show that the performance of our method in terms of area under curve (AUC) on both datasets is improved by more than 10% on average, as compared with three recent baseline methods.
Shiyong Lan, Weikang Huang, Wenwu Wang 0001
ICIP5
2021 Enhancing Alzheimer's Disease Diagnosis via Hierarchical 3D-FCN with Multi-Modal Features
abstract
Alzheimer’s disease (AD) is an incurable, progressive neurological disorder of the human brain related to loss of memory, commonly seen in the elderly population. Accurate detection of AD can help with proper treatment and prevent brain function damage. Existing CNN-based methods need to predetermine informative locations in sMRI, which means the stage of distinguishing lesions is separated from the later stages of feature extraction and classifier construction. In this paper, a novel “two-stage” framework based on a hierarchical 3D fully convolutional network (H-3D-FCN) is proposed to automatically identify discriminative local patches and regions in the sMRI. We further optimize the diagnosis performance by constructing a multi-layer perceptron (MLP) model which combines the multi-modal features (e.g., MMSE score, age, gender, APOE 4) with the risk probability maps (RPMs) generated from the H-3D-FCN model. Experiments on three typical AD datasets, namely, ADNI, AIBL, and NACC, show that our method achieves state-of-the-art performance as compared with recent baselines.
Chao Liu 0001, Dading Chong, Wenwu Wang 0001
ICIP4
2021 SpecAugment++: A Hidden Space Data Augmentation Method for Acoustic Scene Classification
abstract
In this paper, we present SpecAugment++, a novel data augmentation method for deep neural networks based acoustic scene classification (ASC).Different from other popular data augmentation methods such as SpecAugment and mixup that only work on the input space, SpecAugment++ is applied to both the input space and the hidden space of the deep neural networks to enhance the input and the intermediate feature representations.For an intermediate hidden state, the augmentation techniques consist of masking blocks of frequency channels and masking blocks of time frames, which improve generalization by enabling a model to attend not only to the most discriminative parts of the feature, but also the entire parts.Apart from using zeros for masking, we also examine two approaches for masking based on the use of other samples within the minibatch, which helps introduce noises to the networks to make them more discriminative for classification.The experimental results on the DCASE 2018 Task1 dataset and DCASE 2019 Task1 dataset show that our proposed method can obtain 3.6% and 4.7% accuracy gains over a strong baseline without augmentation (i.e.CP-ResNet) respectively, and outperforms other previous data augmentation methods.
Helin Wang, Yuexian Zou, Wenwu Wang 0001
Interspeech3
2021 Crossfire Conditional Generative Adversarial Networks for Singing Voice Extraction
abstract
Generative adversarial networks (GANs) and Conditional GANs (cGANs) have recently been applied for singing voice extraction (SVE), since they can accurately model the vocal distributions and effectively utilize a large amount of unlabelled datasets.However, current GANs/cGANs based SVE frameworks have no explicit mechanism to eliminate the mutual interferences between different sources.In this work, we introduce a novel 'crossfire' criterion into GANs to complement its standard adversarial training, which forms a dual-objective GANs, namely Crossfire GANs (Cr-GANs).In addition, we design a Generalized Projection Method (GPM) for cGANs based frameworks to extract more effective conditional information for SVE.Using the proposed GPM, we extend our Cr-GANs to conditional version, i.e., Crossfire Conditional GANs (Cr-cGANs).The proposed methods were evaluated on the DSD100 and CCMixter datasets.The numerical results have shown that the 'crossfire' criterion and GPM are beneficial to each other and considerably improve the separation performance of existing GANs/cGANs based SVE methods.
Weitao Yuan, Shengbei Wang, Xiangrui Li, Masashi Unoki, Wenwu Wang 0001
Interspeech5
2021 Convolutional fusion network for monaural speech enhancement
Yang Xian, Yang Sun 0003, Wenwu Wang 0001, Syed M. Naqvi
Neural Networks3
2021 Indoor Multi-Speaker Localization Based on Bayesian Nonparametrics in the Circular Harmonic Domain
abstract
Circular microphone arrays have been used for multi-speaker localization in computational auditory scene analysis, for their high flexibility in sound field analysis, including the generation of frequency-invariant eigenbeams for wideband acoustic sources. However, the localization performance of existing circular harmonic approaches, such as circular harmonics beamformer (CHB) depends strongly on the physical characteristics (such as shape) of sensor arrays, and the level of uncertainties presented in acoustic environments (such as background noise, room reverberation, and the number of sources). These uncertainties may limit the performance or practical application of the speaker localization algorithms. To address these issues, in this paper, we present a new indoor multi-speaker localization method in the circular harmonic domain based on the acoustic holography beamforming (AHB) technique and the Bayesian nonparametrics (BNP) method. More specifically, we use the AHB technique, which combines the delay-and-sum beamforming with acoustic-holography-based virtual sensing, to generate direction of arrival (DOA) measurements in the time-frequency (TF) domain, and then design a BNP algorithm based on the infinite Gaussian mixture model (IGMM) to estimate the DOAs of the individual sources without the prior knowledge about the number of sources. These estimates may degrade in the presence of room reverberation and background noise. To address this issue, we develop a robust TF bin selection and permutation method on the basis of mixture weights, using power, power ratio and local variance estimated at each TF bin. Experiments performed on both simulated and real-data show that our method gives significantly better performance, than four recent baseline methods, in a variety of noise and reverberation levels, in terms of the root-mean-square error (RMSE) of the DOA estimation and the source detecting success rate.
Kunkun SongGong, Wenwu Wang 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2021 Multiple Acoustic Source Localization in Microphone Array Networks
abstract
The problem of multiple acoustic source localization using observations from a microphone array network is investigated in this article. Multiple source signals are assumed to be window-disjoint-orthogonal (WDO) on the time-frequency (TF) domain and time delay of arrival (TDOA) measurements are extracted at each TF bin. A Bayesian network model is then proposed to jointly assign the measurements to different sources and estimate the acoustic source locations. Considering that the WDO assumption is usually violated under reverberant and noisy environments, we construct a relational network by coding the distance information between the distributed microphone arrays such that adjacent arrays have higher probabilities of observing the same acoustic source, which is able to mitigate the miss detection issues in adverse environments. A Laplace approximate variational inference method is introduced to estimate the hidden variables in the proposed Bayesian network model. Both simulations and real data experiments are performed. The results show that our proposed method is able to achieve better source localization accuracy than existing methods.
Jielong Yang, Xionghu Zhong, Weiguang Chen, Wenwu Wang 0001
IEEE ACM Trans. Audio Speech Lang. Process.4
2021 Evolving Multi-Resolution Pooling CNN for Monaural Singing Voice Separation
abstract
Monaural singing voice separation (MSVS) is a challenging task and has been extensively studied. Deep neural networks (DNNs) are current state-of-the-art methods for MSVS. However, they are often designed manually, which is time-consuming and error-prone. They are also pre-defined, thus cannot adapt their structures to the training data. To address these issues, we first designed a multi-resolution convolutional neural network (CNN) for MSVS called multi-resolution pooling CNN (MRP-CNN), which uses various-sized pooling operators to extract multi-resolution features. We then introduced Neural Architecture Search (NAS) to extend the MRP-CNN to the evolving MRP-CNN (E-MRP-CNN) to automatically search for effective MRP-CNN structures using genetic algorithms optimized in terms of a single objective taking into account only separation performance and multiple objectives taking into account both separation performance and model complexity. The E-MRP-CNN using the multi-objective algorithm gives a set of Pareto-optimal solutions, each providing a trade-off between separation performance and model complexity. Evaluations on the MIR-1 K, DSD100, and MUSDB18 datasets were used to demonstrate the advantages of the E-MRP-CNN over several recent baselines.
Weitao Yuan, Bofei Dong, Shengbei Wang, Masashi Unoki, Wenwu Wang 0001
IEEE ACM Trans. Audio Speech Lang. Process.5
2021 Short-Term Traffic Flow Prediction With Wavelet and Multi-Dimensional Taylor Network Model
abstract
Accurate prediction of the traffic state has received sustained attention for its ability to provide the anticipatory traffic condition required for people's travel and traffic management. In this paper, we propose a novel short-term traffic flow prediction method based on wavelet transform (WT) and multi-dimensional Taylor network (MTN), which is named as W-MTN. Influenced by the short-term noise disturbance in traffic flow information, the WT is employed to improve prediction accuracy by decomposing the time series of traffic flow. The MTN model, which exploits polynomials to approximate the unknown nonlinear function, makes full use of periodicity and temporal feature without transcendental knowledge and mechanism of the system to be predicted. Our proposed W-MTN model is evaluated on the traffic flow information in a certain area of Shenzhen, China. The experimental results indicate that the proposed W-MTN model offers better prediction performance and temporal correlation, as compared with the corresponding models in the known literature. In addition, the proposed model shows good robustness and generalization ability, when considering data from the different days and locations.
Shanliang Zhu, Yu Zhao 0047, Qingling Li, Wenwu Wang 0001, Shuguo Yang
IEEE Trans. Intell. Transp. Syst.5
2020 A Covert Ultrasonic Phone-to-Phone Communication Scheme
Liming Shi, Limin Yu, Kaizhu Huang, Xu Zhu 0001, Zhi Wang 0003, Xiaofei Li 0001, Wenwu Wang 0001, Xinheng Wang 0001
CollaborateCom (1)7
2020 Meta Metric Learning for Highly Imbalanced Aerial Scene Classification
abstract
Class imbalance is an important factor that affects the performance of deep learning models used for remote sensing scene classification. In this paper, we propose a random finetuning meta metric learning model (RF-MML) to address this problem. Derived from episodic training in meta metric learning, a novel strategy is proposed to train the model, which consists of two phases, i.e., random episodic training and all classes fine-tuning. By introducing randomness into the episodic training and integrating it with fine-tuning for all classes, the few-shot meta-learning paradigm can be successfully applied to class imbalanced data to improve the classification performance. Experiments are conducted to demonstrate the effectiveness of the proposed model on class imbalanced datasets, and the results show the superiority of our model, as compared with other state-of-the-art methods.
Jian Guan 0001, Jiabei Liu, Pengming Feng, Tong Shuai, Wenwu Wang 0001
ICASSP6
2020 Weakly Labelled Audio Tagging Via Convolutional Networks with Spatial and Channel-Wise Attention
abstract
Multiple instance learning (MIL) with convolutional neural networks (CNNs) has been proposed recently for weakly labelled audio tagging. However, features from the various CNN filtering channels and spatial regions are often treated equally, which may limit its performance in event prediction. In this paper, we propose a novel attention mechanism, namely, spatial and channel-wise attention (SCA). For spatial attention, we divide it into global and local submodules with the former to capture the event-related spatial regions and the latter to estimate the onset and offset of the events. Considering the variations in CNN channels, channel-wise attention is also exploited to recognize different sound scenes. The proposed SCA can be employed into any CNNs seamlessly with affordable overheads and is end-to-end trainable fashion. Extensive experiments on weakly labelled dataset Audioset show that the proposed SCA with CNNs achieves a state-of-the-art mean average precision (mAP) of 0.390.
Sixin Hong, Yuexian Zou, Wenwu Wang 0001, Meng Cao 0002
ICASSP3
2020 Learning With Out-of-Distribution Data for Audio Classification
abstract
In supervised machine learning, the assumption that training data is labelled correctly is not always satisfied. In this paper, we investigate an instance of labelling error for classification tasks in which the dataset is corrupted with out-of-distribution (OOD) instances: data that does not belong to any of the target classes, but is labelled as such. We show that detecting and relabelling certain OOD instances, rather than discarding them, can have a positive effect on learning. The proposed method uses an auxiliary classifier, trained on data that is known to be in-distribution, for detection and relabelling. The amount of data required for this is shown to be small. Experiments are carried out on the FSDnoisy18k audio dataset, where OOD instances are very prevalent. The proposed method is shown to improve the performance of convolutional neural networks by a significant margin. Comparisons with other noise-robust techniques are similarly encouraging.
Turab Iqbal, Yin Cao, Qiuqiang Kong, Mark D. Plumbley, Wenwu Wang 0001
ICASSP5
2020 Source Separation with Weakly Labelled Data: an Approach to Computational Auditory Scene Analysis
abstract
Source separation is the task of separating an audio recording into individual sound sources. Source separation is fundamental for computational auditory scene analysis. Previous work on source separation has focused on separating particular sound classes such as speech and music. Much previous work requires mixtures and clean source pairs for training. In this work, we propose a source separation framework trained with weakly labelled data. Weakly labelled data only contains the tags of an audio clip, without the occurrence time of sound events. We first train a sound event detection system with AudioSet. The trained sound event detection system is used to detect segments that are most likely to contain a target sound event. Then a regression is learnt from a mixture of two randomly selected segments to a target segment conditioned on the audio tagging prediction of the target segment. Our proposed system can separate 527 kinds of sound classes from AudioSet within a single system. A U-Net is adopted for the separation system and achieves an average SDR of 5.67 dB over 527 sound classes in AudioSet.
Qiuqiang Kong, Yuxuan Wang 0002, Xuchen Song, Yin Cao, Wenwu Wang 0001, Mark D. Plumbley
ICASSP5
2020 An Analytical Solution to Jacobsen Estimator for Windowed Signals
abstract
Interpolated discrete Fourier transform (DFT) is a well-known method for frequency estimation of complex sinusoids. For signals without windowing (or with rectangular-windowing), this has been well investigated and a large number of estimators have been developed. However, very few algorithms have been developed for windowed signals so far. In this paper, we extend the well-known Jacobsen estimator to windowed signals. The extension is deduced from the fact that an arbitrary cosine-sum window functions are composed of complex sinusoids. Consequently, the Jacobsen estimator for windowed signals can be formulated as an algebraic equation with no approximation and thus an analytical solution to the estimator can be obtained. Simulation results show that our approach improves the performance in comparison with the conventional interpolated DFT algorithms for windowed signals.
Takahiro Murakami, Wenwu Wang 0001
ICASSP2
2020 Gated Multi-Head Attention Pooling for Weakly Labelled Audio Tagging
abstract
Multiple instance learning (MIL) has recently been used for weakly labelled audio tagging, where the spectrogram of an audio signal is divided into segments to form instances in a bag, and then the low-dimensional features of these segments are pooled for tagging.The choice of a pooling scheme is the key to exploiting the weakly labelled data.However, the traditional pooling schemes are usually fixed and unable to distinguish the contributions, making it difficult to adapt to the characteristics of the sound events.In this paper, a novel pooling algorithm is proposed for MIL, named gated multi-head attention pooling (GMAP), which is able to attend to the information of events from different heads at different positions.Each head allows the model to learn information from different representation subspaces.Furthermore, in order to avoid the redundancy of multi-head information, a gating mechanism is used to fuse individual head features.The proposed GMAP increases the modeling power of the single-head attention with no computational overhead.Experiments are carried out on Audioset, which is a large-scale weakly labelled dataset, and show superior results to the non-adaptive pooling and the vanilla attention pooling schemes.
Sixin Hong, Yuexian Zou, Wenwu Wang 0001
INTERSPEECH3
2020 Environmental Sound Classification with Parallel Temporal-Spectral Attention
abstract
Convolutional neural networks (CNN) are one of the bestperforming neural network architectures for environmental sound classification (ESC).Recently, temporal attention mechanisms have been used in CNN to capture the useful information from the relevant time frames for audio classification, especially for weakly labelled data where the onset and offset times of the sound events are not applied.In these methods, however, the inherent spectral characteristics and variations are not explicitly exploited when obtaining the deep features.In this paper, we propose a novel parallel temporal-spectral attention mechanism for CNN to learn discriminative sound representations, which enhances the temporal and spectral features by capturing the importance of different time frames and frequency bands.Parallel branches are constructed to allow temporal attention and spectral attention to be applied respectively in order to mitigate interference from the segments without the presence of sound events.The experiments on three environmental sound classification (ESC) datasets and two acoustic scene classification (ASC) datasets show that our method improves the classification performance and also exhibits robustness to noise.
Helin Wang, Yuexian Zou, Dading Chong, Wenwu Wang 0001
INTERSPEECH4
2020 Optimal feasible step-size based working set selection for large scale SVMs training
Shili Peng, Qinghua Hu, Wenwu Wang 0001
Neurocomputing4
2020 Modeling Label Dependencies for Audio Tagging With Graph Convolutional Network
abstract
As a multi-label classification task, audio tagging aims to predict the presence or absence of certain sound events in an audio recording. Existing works in audio tagging do not explicitly consider the probabilities of the co-occurrences between sound events, which is termed as the label dependencies in this study. To address this issue, we propose to model the label dependencies via a graph-based method, where each node of the graph represents a label. An adjacency matrix is constructed by mining the statistical relations between labels to represent the graph structure information, and a graph convolutional network (GCN) is employed to learn node representations by propagating information between neighboring nodes based on the adjacency matrix, which implicitly models the label dependencies. The generated node representations are then applied to the acoustic representations for classification. Experiments on Audioset show that our method achieves a state-of-the-art mean average precision (mAP) of 0.434.
Helin Wang, Yuexian Zou, Dading Chong, Wenwu Wang 0001
IEEE Signal Process. Lett.4
2020 Association Loss for Visual Object Detection
abstract
Convolutional neural network (CNN) is a popular choice for visual object detection where two sub-nets are often used to achieve object classification and localization separately. However, the intrinsic relation between the localization and classification sub-nets was not exploited explicitly for object detection. In this letter, we propose a novel association loss, namely, the proxy squared error (PSE) loss, to entangle the two sub-nets, thus use the dependency between the classification and localization scores obtained from these two sub-nets to improve the detection performance. We evaluate our proposed loss on the MS-COCO dataset and compare it with the loss in a recent baseline, i.e. the fully convolutional one-stage (FCOS) detector. The results show that our method can improve the AP from 33.8 to 35.4 and AP75 from 35.4 to 37.8, as compared with the FCOS baseline.
Dongli Xu, Jian Guan 0001, Pengming Feng, Wenwu Wang 0001
IEEE Signal Process. Lett.4
2020 PANNs: Large-Scale Pretrained Audio Neural Networks for Audio Pattern Recognition
abstract
Audio pattern recognition is an important research topic in the machine learning area, and includes several tasks such as audio tagging, acoustic scene classification, music classification, speech emotion classification and sound event detection. Recently, neural networks have been applied to tackle audio pattern recognition problems. However, previous systems are built on specific datasets with limited durations. Recently, in computer vision and natural language processing, systems pretrained on large-scale datasets have generalized well to several tasks. However, there is limited research on pretraining systems on large-scale datasets for audio pattern recognition. In this paper, we propose pretrained audio neural networks (PANNs) trained on the large-scale AudioSet dataset. These PANNs are transferred to other audio related tasks. We investigate the performance and computational complexity of PANNs modeled by a variety of convolutional neural networks. We propose an architecture called Wavegram-Logmel-CNN using both log-mel spectrogram and waveform as input feature. Our best PANN system achieves a state-of-the-art mean average precision (mAP) of 0.439 on AudioSet tagging, outperforming the best previous system of 0.392. We transfer PANNs to six audio pattern recognition tasks, and demonstrate state-of-the-art performance in several of those tasks. We have released the source code and pretrained models of PANNs:https://github.com/qiuqiangkong/audioset_tagging_cnn.
Qiuqiang Kong, Yin Cao, Turab Iqbal, Yuxuan Wang 0002, Wenwu Wang 0001, Mark D. Plumbley
IEEE ACM Trans. Audio Speech Lang. Process.5
2020 Sound Event Detection of Weakly Labelled Data With CNN-Transformer and Automatic Threshold Optimization
abstract
Sound event detection (SED) is a task to detect sound events in an audio recording. One challenge of the SED task is that many datasets such as the Detection and Classification of Acoustic Scenes and Events (DCASE) datasets are weakly labelled. That is, there are only audio tags for each audio clip without the onset and offset times of sound events. We compare segment-wise and clip-wise training for SED that is lacking in previous works. We propose a convolutional neural network transformer (CNN-Transfomer) for audio tagging and SED, and show that CNN-Transformer performs similarly to a convolutional recurrent neural network (CRNN). Another challenge of SED is that thresholds are required for detecting sound events. Previous works set thresholds empirically, and are not an optimal approaches. To solve this problem, we propose an automatic threshold optimization method. The first stage is to optimize the system with respect to metrics that do not depend on thresholds, such as mean average precision (mAP). The second stage is to optimize the thresholds with respect to metrics that depends on those thresholds. Our proposed automatic threshold optimization system achieves a state-of-the-art audio tagging F1 of 0.646, outperforming that without threshold optimization of 0.629, and a sound event detection F1 of 0.584, outperforming that without threshold optimization of 0.564.
Qiuqiang Kong, Yong Xu 0004, Wenwu Wang 0001, Mark D. Plumbley
IEEE ACM Trans. Audio Speech Lang. Process.3
2020 Joint Raindrop and Haze Removal From a Single Image
abstract
In a recent study, it was shown that, with adversarial training of an attentive generative network, it is possible to convert a raindrop degraded image into a relatively clean one. However, in real world, raindrop appearance is not only formed by individual raindrops, but also by the distant raindrops accumulation and the atmospheric veiling, namely haze. Current methods are limited in extracting accurate features from a raindrop degraded image with background scene, the blurred raindrop regions, and the haze. In this paper, we propose a new model for an image corrupted by the raindrops and the haze, and introduce an integrated multi-task algorithm to address the joint raindrop and haze removal (JRHR) problem by combining an improved estimate of the atmospheric light, a modified transmission map, a generative adversarial network (GAN) and an optimized visual attention network. The proposed algorithm can extract more accurate features for both sky and non-sky regions. Experimental evaluation has been conducted to show that the proposed algorithm significantly outperforms state-of-the-art algorithms on both synthetic and real-world images in terms of both qualitative and quantitative measures.
Yina Guo, Xiaowen Ren, Anhong Wang, Wenwu Wang 0001
IEEE Trans. Image Process.5
2020 Audio-Visual Particle Flow SMC-PHD Filtering for Multi-Speaker Tracking
abstract
Sequential Monte Carlo probability hypothesis density (SMC-PHD) filtering is a popular method used recently for audio-visual (AV) multi-speaker tracking. However, due to the weight degeneracy problem, the posterior distribution can be represented poorly by the estimated probability, when only a few particles are present around the peak of the likelihood density function. To address this issue, we propose a new framework where particle flow (PF) is used to migrate particles smoothly from the prior to the posterior probability density. We consider both zero and non-zero diffusion particle flows (ZPF/NPF), and developed two new algorithms, AV-ZPF-SMC-PHD and AV-NPF-SMC-PHD, where the speaker states from the previous frames are also considered for particle relocation. The proposed algorithms are compared systematically with several baseline tracking methods using the AV16.3, AVDIAR and CLEAR datasets, and are shown to offer improved tracking accuracy and average effective sample size (ESS).
Yang Liu 0175, Volkan Kilic, Jian Guan 0001, Wenwu Wang 0001
IEEE Trans. Multim.4
2019 Acoustic Scene Generation with Conditional Samplernn
abstract
Acoustic scene generation (ASG) is a task to generate waveforms for acoustic scenes. ASG can be used to generate audio scenes for movies and computer games. Recently, neural networks such as SampleRNN have been used for speech and music generation. However, ASG is more challenging due to its wide variety. In addition, evaluating a generative model is also difficult. In this paper, we propose to use a conditional SampleRNN model to generate acoustic scenes conditioned on the input classes. We also propose objective criteria to evaluate the quality and diversity of the generated samples based on classification accuracy. The experiments on the DCASE 2016 Task 1 acoustic scene data show that with the generated audio samples, a classification accuracy of 65.5% can be achieved compared to samples generated by a random model of 6.7% and samples from real recording of 83.1%. The performance of a classifier trained only on generated samples achieves an accuracy of 51.3%, as opposed to an accuracy of 6.7% with samples generated by a random model.
Qiuqiang Kong, Yong Xu 0004, Turab Iqbal, Yin Cao, Wenwu Wang 0001, Mark D. Plumbley
ICASSP5
2019 Generalisation in Environmental Sound Classification: The 'Making Sense of Sounds' Data Set and Challenge
abstract
Humans are able to identify a large number of environmental sounds and categorise them according to high-level semantic categories, e.g. urban sounds or music. They are also capable of generalising from past experience to new sounds when applying these categories. In this paper we report on the creation of a data set that is structured according to the top-level of a taxonomy derived from human judgements and the design of an associated machine learning challenge, in which strong generalisation abilities are required to be successful. We introduce a baseline classification system, a deep convolutional network, which showed strong performance with an average accuracy on the evaluation data of 80.8%. The result is discussed in the light of two alternative explanations: An unlikely accidental category bias in the sound recordings or a more plausible true acoustic grounding of the high-level categories.
Christian Kroos, Oliver Bones, Yin Cao, Lara Harris, Philip J. B. Jackson, William J. Davies, Wenwu Wang 0001, Trevor J. Cox, Mark D. Plumbley
ICASSP7
2019 Enhanced Streaming Based Subspace Clustering Applied to Acoustic Scene Data Clustering
abstract
Labelled data are often required to train an acoustic scene classification system. However, it is time-consuming and expensive to label the data manually. An unsupervised clustering algorithm can be used to facilitate the labelling process by dividing the acoustic data into different categories. Nevertheless, it can be problematic to run a clustering algorithm with growing data volume and dimension due to the sharp increase in the computational and memory costs. We propose a new streaming based subspace clustering algorithm which allows the data to be clustered on the fly, and also resolves data points in the overlapping regions of two subspaces by augmenting the learned low-rank representation with the original data samples. Experimental results show that our method can achieve the clustering objective for overwhelmingly high-volume data in an online fashion, while retaining good accuracy and reducing the memory cost significantly.
Shuoyang Li, Yuantao Gu, Yuhui Luo, Jonathon A. Chambers, Wenwu Wang 0001
ICASSP5
2019 Labelled Non-zero Particle Flow for SMC-PHD Filtering
abstract
The sequential Monte Carlo probability hypothesis density (SMC-PHD) filter assisted by particle flows (PF) has been shown to be promising for audio-visual multi-speaker tracking. A clustering step is often employed for calculating the particle flow, which leads to a substantial increase in the computational cost. To address this issue, we propose an alternative method based on the labelled non-zero particle flow (LNPF) to adjust the particle states. Results obtained from the AV16.3 dataset show improved performance by the proposed method in terms of computational efficiency and tracking accuracy as compared with baseline AV-NPF-SMC-PHD methods.
Yang Liu 0175, Qinghua Hu, Yuexian Zou, Wenwu Wang 0001
ICASSP4
2019 Background Adaptation for Improved Listening Experience in Broadcasting
abstract
The intelligibility of speech in noise can be improved by modifying the speech. But with object-based audio, there is the possibility of altering the background sound while leaving the speech unaltered. This may prove a less intrusive approach, affording good speech intelligibility without overly compromising the perceived sound quality. In this study, the technique of spectral weighting was applied to the background. The frequency-dependent weightings for adaptation were learnt by maximising a weighted combination of two perceptual objective metrics for speech intelligibility and audio quality. The balance between the two objective metrics was determined by the perceptual relationship between intelligibility and quality. A neural network was trained to provide a fast solution for real-time processing. Tested in a variety of background sounds and speech-to-background ratios (SBRs), the proposed method led to a large intelligibility gain over the unprocessed baseline. Compared to an approach using constant weightings, the proposed method was able to dynamically preserve the overall audio quality better with respect to SBR changes.
Trevor J. Cox, Bruno Fazenda, Qingju Liu, Wenwu Wang 0001
ICASSP5
2019 Proximal Deep Recurrent Neural Network for Monaural Singing Voice Separation
abstract
The recent deep learning methods can offer state-of-the-art performance for Monaural Singing Voice Separation (MSVS). In these deep methods, the recurrent neural network (RNN) is widely employed. This work proposes a novel type of Deep RNN (DRNN), namely Proximal DRNN (P-DRNN) for MSVS, which improves the conventional Stacked RNN (S-RNN) by introducing a novel interlayer structure. The interlayer structure is derived from an optimization problem for Monaural Source Separation (MSS). Accordingly, this enables a new hierarchical processing in the proposed P-DRNN with the explicit state transfers between different layers and the skip connections from the inputs, which are efficient for source separation. Finally, the proposed approach is evaluated on the MIR-IK dataset to verify its effectiveness. The numerical results show that the P-DRNN performs better than the conventional S-RNN and several recent MSVS methods.
Weitao Yuan, Shengbei Wang, Xiangrui Li, Masashi Unoki, Wenwu Wang 0001
ICASSP5
2019 Semantic Super-resolution for Extremely Low-resolution Vehicle License Plate
abstract
Vehicle license plate (VLP) super-resolution (SR) is of great demand in intelligent traffic systems. Super-Resolution for extremely low-resolution VLP remains challenging and the state-of-the-art SR methods hardly provide satisfying results for low-resolution (LR) VLPs. In this study, from a new perspective, we develop an effective solution to achieve the super-resolution of the extremely LR VLP images, by using the semantic information of the characters. Specifically, we firstly exploit the pervasive sparse prior for the character recognition in LR condition for VLPs. Then the semantic information extracted from the sparse representation-based classification (SRC) results is employed to alleviate the illness of the SR problem. To maximize the benefit brought by the semantic information from SRC, we employ sparse-coding based super-resolution (SCSR) method to upscale VLP images. In the end, an exponential soft labeling method is designed to reduce the possible bias introduced by character classification. Extensive experiments on the self-built Chinese VLP dataset (VLP100) and public UFPR-ALPR dataset validate the feasibility and effectiveness of our proposed VLP-SR system.
Yuexian Zou, Yi Wang 0033, Wenjie Guan, Wenwu Wang 0001
ICASSP4
2019 Single-Channel Signal Separation and Deconvolution with Generative Adversarial Networks
abstract
Single-channel signal separation and deconvolution aims to separate and deconvolve individual sources from a single-channel mixture. Single-channel signal separation and deconvolution is a challenging problem in which no prior knowledge of the mixing filters is available. Both individual sources and mixing filters need to be estimated. In addition, a mixture may contain non-stationary noise which is unseen in the training set. We propose a synthesizing-decomposition (S-D) approach to solve the single-channel separation and deconvolution problem. In synthesizing, a generative model for sources is built using a generative adversarial network (GAN). In decomposition, both mixing filters and sources are optimized to minimize the reconstruction error of the mixture. The proposed S-D approach achieves a peak-to-noise-ratio (PSNR) of 18.9 dB and 15.4 dB in image inpainting and completion, outperforming a baseline convolutional neural network PSNR of 15.3 dB and 12.2 dB, respectively and achieves a PSNR of 13.2 dB in source separation together with deconvolution, outperforming a convolutive non-negative matrix factorization (NMF) baseline of 10.1 dB.
Qiuqiang Kong, Yong Xu 0004, Philip J. B. Jackson, Wenwu Wang 0001, Mark D. Plumbley
IJCAI4
2019 Multi-channel Convolutional Neural Networks with Multi-level Feature Fusion for Environmental Sound Classification
Dading Chong, Yuexian Zou, Wenwu Wang 0001
MMM (2)3
2019 A Source Counting Method Using Acoustic Vector Sensor Based on Sparse Modeling of DOA Histogram
abstract
The number of sources present in a mixture is crucial information often assumed to be known or detected by source counting. The existing methods for source counting in underdetermined blind speech separation suffer from the overlapping between sources with low W-disjoint orthogonality. To address this issue, we propose to fit the direction-of-arrival (DOA) histogram with multiple von-Mises density (VM) functions directly and form a sparse recovery problem, where all the source clusters and the sidelobes in the DOA histogram are fitted with VM functions of different spatial parameters. We also developed a formula to perform the source counting taking advantage of the values of the sparse source vector to reduce the influence of sidelobes. Experiments are carried out to evaluate the proposed source counting method, and the results show that the proposed method outperforms two well-known baseline methods.
Yang Chen 0019, Wenwu Wang 0001, Bingyin Xia
IEEE Signal Process. Lett.2
2019 A Speech Synthesis Approach for High Quality Speech Separation and Generation
abstract
We propose a new method for source separation by synthesizing the source from a speech mixture corrupted by various environmental noise. Unlike traditional source separation methods which estimate the source from the mixture as a replica of the original source (e.g. by solving an inverse problem), our proposed method is a synthesis-based approach which aims to generate a new signal (i.e. “fake” source) that sounds similar to the original source. The proposed system has an encoder-decoder topology, where the encoder predicts intermediate-level features from the mixture, i.e. Mel-spectrum of the target source, using a hybrid recurrent and hourglass network, while the decoder is a state-of-the-art WaveNet speech synthesis network conditioned on the Mel-spectrum, which directly generates time-domain samples of the sources. Both objective and subjective evaluations were performed on the synthesized sources, and show great advantages of our proposed method for high-quality speech source separation and generation.
Qingju Liu, Philip J. B. Jackson, Wenwu Wang 0001
IEEE Signal Process. Lett.3
2019 A Skip Attention Mechanism for Monaural Singing Voice Separation
abstract
This work proposes a simple but effective attention mechanism, namely Skip Attention (SA), for monaural singing voice separation (MSVS). First, the SA, embedded in the convolutional encoder-decoder network (CEDN), realizes an attention-driven and dependency modeling for the repetitive structures of the music source. Second, the SA, replacing the popular skip connection in the CEDN, effectively controls the flow of the low-level (vocal and musical) features to the output and improves the feature sensitivity and accuracy for MSVS. Finally, we implement the proposed SA on the Stacked Hourglass Network (SHN), namely Skip Attention SHN (SA-SHN). Quantitative and qualitative evaluation results have shown that the proposed SA-SHN achieves significant performance improvement on the MIR-1K dataset (compared to the state-of-the-art SHN) and competitive MSVS performance on the DSD100 dataset (compared to the state-of-the-art DenseNet), even without using any data augmentation methods.
Weitao Yuan, Shengbei Wang, Xiangrui Li, Masashi Unoki, Wenwu Wang 0001
IEEE Signal Process. Lett.5
2019 Sound Event Detection and Time-Frequency Segmentation from Weakly Labelled Data
abstract
Sound event detection (SED) aims to detect when and recognize what sound events happen in an audio clip. Many supervised SED algorithms rely on strongly labelled data that contains the onset and offset annotations of sound events. However, many audio tagging datasets are weakly labelled, that is, only the presence of the sound events is known, without knowing their onset and offset annotations. In this paper, we propose a time–frequency (T–F) segmentation framework trained on weakly labelled data to tackle the sound event detection and separation problem. In training, a segmentation mapping is applied on a T–F representation, such as log mel spectrogram of an audio clip to obtain T–F segmentation masks of sound events. The T–F segmentation masks can be used for separating the sound events from the background scenes in the T–F domain. Then, a classification mapping is applied on the T–F segmentation masks to estimate the presence probabilities of the sound events. We model the segmentation mapping using a convolutional neural network and the classification mapping using a global weighted rank pooling. In SED, predicted onset and offset times can be obtained from the T–F segmentation masks. As a byproduct, separated waveforms of sound events can be obtained from the T–F segmentation masks. We remixed the DCASE 2018 Task 1 acoustic scene data with the DCASE 2018 Task 2 sound events data. When mixing under 0 dB, the proposed method achieved F1 scores of 0.534, 0.398, and 0.167 in audio tagging, frame-wise SED and event-wise SED, outperforming the fully connected deep neural network baseline of 0.331, 0.237, and 0.120, respectively. In T–F segmentation, we achieved an F1 score of 0.218, where previous methods were not able to do T–F segmentation.
Qiuqiang Kong, Yong Xu 0004, Iwona Sobieraj, Wenwu Wang 0001, Mark D. Plumbley
IEEE ACM Trans. Audio Speech Lang. Process.4
2019 Weakly Labelled AudioSet Tagging With Attention Neural Networks
abstract
Audio tagging is the task of predicting the presence or absence of sound classes within an audio clip. Previous work in audio tagging focused on relatively small datasets limited to recognizing a small number of sound classes. We investigate audio tagging on AudioSet, which is a dataset consisting of over 2 million audio clips and 527 classes. AudioSet is weakly labelled, in that only the presence or absence of sound classes is known for each clip, whereas the onset and offset times are unknown. To address the weakly labelled audio tagging problem, we propose attention neural networks as a way to attend the most salient parts of an audio clip. We bridge the connection between attention neural networks and multiple instance learning (MIL) methods, and propose decision-level and feature-level attention neural networks for audio tagging. We investigate attention neural networks modeled by different functions, depths, and widths. Experiments on AudioSet show that the feature-level attention neural network achieves a state-of-the-art mean average precision of 0.369, outperforming the best MIL method of 0.317 and Google's deep neural network baseline of 0.314. In addition, we discover that the audio tagging performance on AudioSet-embedding features has a weak correlation with the number of training samples and the quality of labels of each sound class.
Qiuqiang Kong, Changsong Yu, Yong Xu 0004, Turab Iqbal, Wenwu Wang 0001, Mark D. Plumbley
IEEE ACM Trans. Audio Speech Lang. Process.5
2019 Modeling the Comb Filter Effect and Interaural Coherence for Binaural Source Separation
abstract
Typical methods for binaural source separation consider only the direct sound as the target signal in a mixture. However, in most scenarios, this assumption limits the source separation performance. It is well known that the early reflections interact with the direct sound, producing acoustic effects at the listening position, e.g. the so-called comb filter effect. In this article, we propose a novel source separation model, that utilizes both the direct sound and the first early reflection information to model the comb filter effect. This is done by observing the interaural phase difference obtained from the time-frequency representation of binaural mixtures. Furthermore, a method is proposed to model the interaural coherence of the signals. Including information related to the sound multipath propagation, the performance of the proposed separation method is improved with respect to the baselines that did not use such information, as illustrated by using binaural recordings made in four rooms, having different sizes and reverberation times.
Luca Remaggi, Philip J. B. Jackson, Wenwu Wang 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2019 Two-Stage Monaural Source Separation in Reverberant Room Environments Using Deep Neural Networks
abstract
Deep neural networks (DNNs) have been used for dereverberation and separation in the monaural source separation problem. However, the performance of current state-of-the-art methods is limited, particularly when applied in highly reverberant room environments. In this paper, we propose a two-stage approach with two DNN-based methods to address this problem. In the first stage, the dereverberation of the speech mixture is achieved with the proposed dereverberation mask (DM). In the second stage, the dereverberant speech mixture is separated with the ideal ratio mask (IRM). To realize this two-stage approach, in the first DNN-based method, the DM is integrated with the IRM to generate the enhanced time-frequency (T-F) mask, namely the ideal enhanced mask (IEM), as the training target for the single DNN. In the second DNN-based method, the DM and the IRM are predicted with two individual DNNs. The IEEE and the TIMIT corpora with real room impulse responses and noise from the NOISEX dataset are used to generate speech mixtures for evaluations. The proposed methods outperform the state-of-the-art specifically in highly reverberant room environments.
Yang Sun 0003, Wenwu Wang 0001, Jonathon A. Chambers, Syed M. Naqvi
IEEE ACM Trans. Audio Speech Lang. Process.2
2018 Large-Scale Weakly Supervised Audio Classification Using Gated Convolutional Neural Network
abstract
In this paper, we present a gated convolutional neural network and a temporal attention-based localization method for audio classification, which won the 1st place in the large-scale weakly supervised sound event detection task of Detection and Classification of Acoustic Scenes and Events (DCASE) 2017 challenge. The audio clips in this task, which are extracted from YouTube videos, are manually labelled with one or more audio tags, but without time stamps of the audio events, hence referred to as weakly labelled data. Two subtasks are defined in this challenge including audio tagging and sound event detection using this weakly labelled data. We propose a convolutional recurrent neural network (CRNN) with learnable gated linear units (GLUs) non-linearity applied on the log Mel spectrogram. In addition, we propose a temporal attention method along the frames to predict the locations of each audio event in a chunk from the weakly labelled data. The performances of our systems were ranked the 1st and the 2nd as a team in these two sub-tasks of DCASE 2017 challenge with F value 55.6% and Equal error 0.73, respectively.
Yong Xu 0004, Qiuqiang Kong, Wenwu Wang 0001, Mark D. Plumbley
ICASSP3
2018 Synthesis of Images by Two-Stage Generative Adversarial Networks
abstract
In this paper, we propose a divide-and-conquer approach using two generative adversarial networks (GANs) to explore how a machine can draw colorful pictures (bird) using a small amount of training data. In our work, we simulate the procedure of an artist drawing a picture, where one begins with drawing objects' contours and edges and then paints them different colors. We adopt two GAN models to process basic visual features including shape, texture and color. We use the first GAN model to generate object shape, and then paint the black and white image based on the knowledge learned using the second GAN model. We run our experiments on 600 color images. The experimental results show that the use of our approach can generate good quality synthetic images, comparable to real ones.
Qiang Huang 0007, Philip J. B. Jackson, Mark D. Plumbley, Wenwu Wang 0001
ICASSP4
2018 Intelligent Signal Processing Mechanisms for Nuanced Anomaly Detection in Action Audio-Visual Data Streams
abstract
We consider the problem of anomaly detection in an audiovisual analysis system designed to interpret sequences of actions from visual and audio cues. The scene activity recognition is based on a generative framework, with a high-level inference model for contextual recognition of sequences of actions. The system is endowed with anomaly detection mechanisms, which facilitate differentiation of various types of anomalies. This is accomplished using intelligence provided by a classifier incongruence detector, classifier confidence module and data quality assessment system, in addition to the classical outlier detection module. The paper focuses on one of the mechanisms, the classifier incongruence detector, the purpose of which is to flag situations when the video and audio modalities disagree in action interpretation. We demonstrate the merit of using the Delta divergence measure for this purpose. We show that this measure significantly enhances the incongruence detection rate in the Human Action Manipulation complex activity recognition data set.
Josef Kittler, Ioannis Kaloskampis, Cemre Zor, Yulia Hicks, Wenwu Wang 0001
ICASSP6
2018 Audio Set Classification with Attention Model: A Probabilistic Perspective
abstract
This paper investigates the Audio Set classification. Audio Set is a large scale weakly labelled dataset (WLD) of audio clips. In WLD only the presence of a label is known, without knowing the happening time of the labels. We propose an attention model to solve this WLD problem and explain the attention model from a novel probabilistic perspective. Each audio clip in Audio Set consists of a collection of features. We call each feature as an instance and the collection as a bag following the terminology in multiple instance learning. In the attention model, each instance in the bag has a trainable probability measure for each class. The classification of the bag is the expectation of the classification output of the instances in the bag with respect to the learned probability measure. Experiments show that the proposed attention model achieves a mAP of 0.327 on Audio Set, outperforming the Google's baseline of 0.314.
Qiuqiang Kong, Yong Xu 0004, Wenwu Wang 0001, Mark D. Plumbley
ICASSP3
2018 A Joint Separation-Classification Model for Sound Event Detection of Weakly Labelled Data
abstract
Source separation (SS) aims to separate individual sources from an audio recording. Sound event detection (SED) aims to detect sound events from an audio recording. We propose a joint separation-classification (JSC) model trained only on weakly labelled audio data, that is, only the tags of an audio recording are known but the time of the events are unknown. First, we propose a separation mapping from the time-frequency (T-F) representation of an audio to the T-F segmentation masks of the audio events. Second, a classification mapping is built from each T-F segmentation mask to the presence probability of each audio event. In the source separation stage, sources of audio events and time of sound events can be obtained from the T-F segmentation masks. The proposed method achieves an equal error rate (EER) of 0.14 in SED, outperforming deep neural network baseline of 0.29. Source separation SDR of 8.08 dB is obtained by using global weighted rank pooling (GWRP) as probability mapping, out-performing the global max pooling (GMP) based probability mapping giving SDR at 0.03 dB. Source code of our work is published.
Qiuqiang Kong, Yong Xu 0004, Wenwu Wang 0001, Mark D. Plumbley
ICASSP3
2018 Iterative Deep Neural Networks for Speaker-Independent Binaural Blind Speech Separation
abstract
In this paper, we propose an iterative deep neural network (DNN)-based binaural source separation scheme, for recovering two concurrent speech signals in a room environment. Besides the commonly-used spectral features, the DNN also takes non-linearly wrapped binaural spatial features as input, which are refined iteratively using parameters estimated from the DNN output via a feedback loop. Different DNN structures have been tested, including a classic multilayer perception regression architecture as well as a new hybrid network with both convolutional and densely-connected layers. Objective evaluations in terms of PESQ and STOI showed consistent improvement over baseline methods using traditional binaural features, especially when the hybrid DNN architecture was employed. In addition, our proposed scheme is robust to mismatches between the training and testing data.
Qingju Liu, Yong Xu 0004, Philip J. B. Jackson, Wenwu Wang 0001, Philip Coleman
ICASSP4
2018 Non-Zero Diffusion Particle Flow SMC-PHD Filter for Audio-Visual Multi-Speaker Tracking
abstract
The sequential Monte Carlo probability hypothesis density (SMC-PHD) filter has been shown to be promising for audio-visual multi-speaker tracking. Recently, the zero diffusion particle flow (ZPF) has been used to mitigate the weight degeneracy problem in the SMC-PHD filter. However, this leads to a substantial increase in the computational cost due to the migration of particles from prior to posterior distribution with a partial differential equation. This paper proposes an alternative method based on the non-zero diffusion particle flow (NPF) to adjust the particle states by fitting the particle distribution with the posterior probability density using the non-zero diffusion. This property allows efficient computation of the migration of particles. Results from the AV16.3 dataset demonstrate that we can significantly mitigate the weight degeneracy problem with a smaller computational cost as compared with the ZPF based SMC-PHD filter.
Yang Liu 0175, Adrian Hilton 0001, Jonathon A. Chambers, Yuxin Zhao 0001, Wenwu Wang 0001
ICASSP5
2018 Bayesian Inference for Multi-Line Spectra in Linear Sensor Array
abstract
For a linear sensor array, using line spectra is a common technique for estimating directions of arrival (DOA) of single-tone sources. Yet, very few papers consider multitone sources. For the first time, we provide the optimal Bayesian inference for multi-line spectra, i.e. a superposition of line spectra, and estimate the DOAs of the multi-tone sources. For tractable computation via fast Fourier transform, we apply a grid-based method, in which source's tones and sensor's array measure are both uncorrelated. Exploiting this method, we interpret the superposition of sensor's data as a complex Gaussian mixture of multi-tone signals. We then estimate DOA via conjugate Von-Mises, also known as circular Gaussian distribution. Our simulation shows that the multitone method is superior to traditional single-tone method for detecting multi-tone source's frequencies, particularly for the sources with overlapping frequencies. The posterior DOA's resolution can be tuned via Von-Mises' parameter a priori, which enhances the sparsity of DOA's estimation.
Viet Hung Tran, Wenwu Wang 0001, Yuhui Luo, Jonathon A. Chambers
ICASSP2
2018 Spatially Regularized Low Rank Tensor Optimization for Visual Data Completion
abstract
Low-rank tensor completion is a recent method for estimating the values of the missing elements in tensor data by minimizing the tensor rank. However, with only the low rank prior, the local piecewise smooth structure that is important for visual data is not used effectively. To address this problem, we define a new spatial regularization S-norm for tensor completion in order to exploit the local spatial smoothness structure of visual data. More specifically, we introduce the S-norm to the tensor completion model based on a non-convex LogDet function. The S-norm helps to drive the neighborhood elements towards similar values. We utilize the Alternating Direction Method of Multiplier (ADMM) to optimize the proposed model. Experimental results in visual data demonstrate that our method outperforms the state-of-the-art tensor completion models.
Jianchao Gao, Wenwu Wang 0001
ICIP3
2018 Cascade Deep Networks for Sparse Linear Inverse Problems
abstract
Sparse deep networks have been widely used in many linear inverse problems, such as image super-resolution and signal recovery. Its performance is as good as deep learning at the same time its parameters are much less than deep learning. However, when the linear inverse problems involve several linear transformations or the ratio of input dimension to output dimension is large, the performance of a single sparse deep network is poor. In this paper, we propose a cascade sparse deep network to address the above problem. In our model, we trained two cascade sparse networks based on Gregor and LeCun's “learned ISTA” and “learned CoD”. The cascade structure can effectively improve the performance as compared to the non-cascade model. We use the proposed methods in image sparse code prediction and signal recovery. The experimental results show that both algorithms perform favorably against a single sparse network.
Huan Zhang 0001, Wenwu Wang 0001
ICPR3
2018 LD-CNN: A Lightweight Dilated Convolutional Neural Network for Environmental Sound Classification
abstract
Environmental Sound Classification (ESC) plays a vital role in machine auditory scene perception. Deep learning based ESC methods., such as the Dilated Convolutional Neural Network (D-CNN)., have achieved the state-of-art results on public datasets. However., the D-CNN ESC model size is often larger than 100MB and is only suitable for the systems with powerful GPUs., which prevents their applications in handheld devices. In this study., we take the D-CNN ESC framework and focus on reducing the model size while maintaining the ESC performance. As a result., a lightweight D-CNN (termed as LD-CNN) ESC system is developed. Our work lies on twofold. First., we propose into reduce the number of parameters in the convolution layers by factorizing a two-dimensional convolution filters (L ×W) to two separable one-dimensional convolution filters ( L ×1 and 1×W). Second., we propose to replace the first fully connection layer (FCL) by a Feature Sum layer (FSL) to further reduce the number of parameters. This is motivated by our finding that the features of the environmental sounds have weak absolute locality property and a global sum operation can be applied to compress the feature map. Experiments on three public datasets (ESC50., UrbanSound8K., and CICESE) show that the proposed system offers comparable classification performance but with a much smaller model size. For example., the model size of our proposed system is about 2.05MB., which is 50 times smaller than the original D-CNN model., but at a loss of only 1%-2 % classification accuracy.
Yuexian Zou, Wenwu Wang 0001
ICPR3
2018 Standard-independent I/Q imbalance estimation and compensation scheme inOFDM
abstract
Direct-conversion transceivers are gaining increasing attention due to their low power consumption. However, they suffer from a serious in- and quadrature-phase (I/Q) imbalance problem. The I/Q imbalance can severely limit the achievable operating signal-to-noise ratio (SNR) at the receiver and, consequently, the supported constellation sizes and data rates. In this paper, we first investigate the effects of I/Q imbalance on orthogonal frequency division multiplexing (OFDM) receivers, and then propose a new I/Q imbalance compensation scheme. In the proposed method, a new statistic, which is robust against channel distortion, is used to estimate the I/Q imbalance parameters, and then the I/Q imbalance is corrected in the frequency domain. Simulations are performed to verify the effectiveness of the proposed method for I/Q imbalance compensation. The results show that the proposed I/Q imbalance compensation method can achieve bit error rate (BER) performance close to that in the ideal case without I/Q imbalance in additive white Gaussian noise (AWGN) or multipath environments. Furthermore, because no pilot information is required, this method can be applied in various standard communication systems.
Fanglin Gu, Shan Wang 0005, Wenwu Wang 0001
Frontiers Inf. Technol. Electron. Eng.3
2018 Error sensitivity analysis of Delta divergence - a novel measure for classifier incongruence detection
abstract
The state of classifier incongruence in decision making systems incorporating multiple classifiers is often an indicator of anomaly caused by an unexpected observation or an unusual situation. Its assessment is important as one of the key mechanisms for domain anomaly detection. In this paper, we investigate the sensitivity of Delta divergence, a novel measure of classifier incongruence, to estimation errors. Statistical properties of Delta divergence are analysed both theoretically and experimentally. The results of the analysis provide guidelines on the selection of threshold for classifier incongruence detection based on this measure.
Josef Kittler, Cemre Zor, Ioannis Kaloskampis, Yulia Hicks, Wenwu Wang 0001
Pattern Recognit.5
2018 Efficient multi-modal geometric mean metric learning
Jianqing Liang, Qinghua Hu, Pengfei Zhu 0001, Wenwu Wang 0001
Pattern Recognit.4
2018 Polynomial dictionary learning algorithms in sparse representations
Jian Guan 0001, Xuan Wang 0002, Pengming Feng, Jing Dong 0001, Jonathon A. Chambers, Zoe Lin Jiang, Wenwu Wang 0001
Signal Process.7
2018 Multi-source phase retrieval from multi-channel phaseless STFT measurements
Yina Guo, Anhong Wang, Wenwu Wang 0001
Signal Process.3
2018 Momentum fractional LMS for power signal parameter estimation
Syed Zubair, Naveed Ishtiaq Chaudhary, Zeshan Aslam Khan, Wenwu Wang 0001
Signal Process.4
2018 A non-intrusive method for estimating binaural speech intelligibility from noise-corrupted signals captured by a pair of microphones
abstract
A non-intrusive method is introduced to predict binaural speech intelligibility in noise directly from signals captured using a pair of microphones. The approach combines signal processing techniques in blind source separation and localisation, with an intrusive objective intelligibility measure (OIM). Therefore, unlike classic intrusive OIMs, this method does not require a clean reference speech signal and knowing the location of the sources to operate. The proposed approach is able to estimate intelligibility in stationary and fluctuating noises, when the noise masker is presented as a point or diffused source, and is spatially separated from the target speech source on a horizontal plane. The performance of the proposed method was evaluated in two rooms. When predicting subjective intelligibility measured as word recognition rate, this method showed reasonable predictive accuracy with correlation coefficients above 0.82, which is comparable to that of a reference intrusive OIM in most of the conditions. The proposed approach offers a solution for fast binaural intelligibility prediction, and therefore has practical potential to be deployed in situations where on-site speech intelligibility is a concern.
Qingju Liu, Wenwu Wang 0001, Trevor J. Cox
Speech Commun.3
2018 Low rank matrix completion using truncated nuclear norm and sparse regularizer
Jing Dong 0001, Zhichao Xue, Jian Guan 0001, Zi-Fa Han, Wenwu Wang 0001
Signal Process. Image Commun.5
2018 An RIP-Based Performance Guarantee of Covariance-Assisted Matching Pursuit
abstract
An OMP-like covariance-assisted matching pursuit (CAMP) method has recently been proposed. Given a prior knowledge of the covariance and mean of the sparse coefficients, CAMP balances the least squares estimator and the prior knowledge by leveraging the Gauss–Markov theorem. In this letter, we study the performance of CAMP in the framework of restricted isometry property (RIP). It is shown that under some conditions on RIP and the minimum magnitude of the nonzero elements of the sparse signal, CAMP with sparse level$K$can recover the exact support of the sparse signal from noisy measurements.$l_2$bounded noise and Gaussian noise are considered in our analysis. We also discuss the extreme conditions of noise (e.g., the noise power is infinite) to simply show the stability of CAMP.
Jiayang Wang, Gen Li 0005, Lucas Rencker, Wenwu Wang 0001, Yuantao Gu
IEEE Signal Process. Lett.4
2018 Bootstrap Averaging for Model-Based Source Separation in Reverberant Conditions
abstract
Recently proposed model-based methods use time-frequency (T-F) masking for source separation, where the T-F masks are derived from various cues described by a frequency domain Gaussian mixture model (GMM). These methods work well for separating mixtures recorded in low-to-medium level of reverberation, however, their performance degrades as the level of reverberation is increased. We note that the relatively poor performance of these methods under reverberant conditions can be attributed to the high variance of the frequency-dependent GMM parameter estimates. To address this limitation, a novel bootstrap-based approach is proposed to improve the accuracy of expectation maximization estimates of a frequency-dependent GMM based on an a priori chosen initialization scheme. It is shown how the proposed technique allows us to construct time-frequency masks which lead to improved model-based source separation for reverberant speech mixtures. Experiments and analysis are performed on speech mixtures formed using real room-recorded impulse responses.
Swati Chandna 0002, Wenwu Wang 0001
IEEE ACM Trans. Audio Speech Lang. Process.2
2018 Manifold-Based Visual Object Counting
abstract
Visual object counting (VOC) is an emerging area in computer vision which aims to estimate the number of objects of interest in a given image or video. Recently, object density based estimation method is shown to be promising for object counting as well as rough instance localization. However, the performance of this method tends to degrade when dealing with new objects and scenes. To address this limitation, we propose a manifold-based method for visual object counting (M-VOC), based on the manifold assumption that similar image patches share similar object densities. Firstly, the local geometry of a given image patch is represented linearly by its neighbors using a predefined patch training set, and the object density of this given image patch is reconstructed by preserving the local geometry using locally linear embedding. To improve the characterization of local geometry, additional constraints such as sparsity and non-negativity are also considered via regularization, nonlinear mapping, and kernel trick. Compared with the state-of-the-art VOC methods, our proposed M-VOC methods achieve competitive performance on seven benchmark datasets. Experiments verify that the proposed M-VOC methods have several favorable properties, such as robustness to the variation in the size of training dataset and image resolution, as often encountered in real-world VOC applications.
Yi Wang 0033, Yuexian Zou, Wenwu Wang 0001
IEEE Trans. Image Process.3
2018 Multiple Speaker Tracking in Spatial Audio via PHD Filtering and Depth-Audio Fusion
abstract
In the object-based spatial audio system, positions of the audio objects (e.g., speakers/talkers or voices) presented in the sound scene are required as important metadata attributes for object acquisition and reproduction. Binaural microphones are often used as a physical device to mimic human hearing and to monitor and analyze the scene, including localization and tracking of multiple speakers. The binaural audio tracker, however, is usually prone to the errors caused by room reverberation and background noise. To address this limitation, we present a multimodal tracking method by fusing the binaural audio with depth information (from a depth sensor, e.g., Kinect). More specifically, the probability hypothesis density (PHD) filtering framework is first applied to the depth stream, and a novel clutter intensity model is proposed to improve the robustness of the PHD filter when an object is occluded either by other objects or due to the limited field of view of the depth sensor. To compensate misdetections in the depth stream, a novel gap filling technique is presented to map audio azimuths obtained from the binaural audio tracker to 3D positions, using speaker-dependent spatial constraints learned from the depth stream. With our proposed method, both the errors in the binaural tracker and the misdetections in the depth tracker can be significantly reduced. Real-room recordings are used to show the improved performance of the proposed method in removing outliers and reducing misdetections.
Qingju Liu, Wenwu Wang 0001, Teófilo Emídio de Campos, Philip J. B. Jackson, Adrian Hilton 0001
IEEE Trans. Multim.2
2017 Assessment of musical noise using localization of isolated peaks in time-frequency domain
abstract
Musical noise is a recurrent issue that appears in spectral techniques for denoising or blind source separation. Due to localised errors of estimation, isolated peaks may appear in the processed spectrograms, resulting in annoying tonal sounds after synthesis known as “musical noise”. In this paper, we propose a method to assess the amount of musical noise in an audio signal, by characterising the impact of these artificial isolated peaks on the processed sound. It turns out that because of the constraints between STFT coefficients, the isolated peaks are described as time-frequency “spots” in the spectrogram of the processed audio signal. The quantification of these “spots”, achieved through the adaptation of a method for localisation of significant STFT regions, allows for an evaluation of the amount of musical noise. We believe that this will pave the way to an objective measure and a better understanding of this phenomenon.
Ronan Hamon, Valentin Emiya, Lucas Rencker, Wenwu Wang 0001, Mark D. Plumbley
ICASSP4
2017 Fast tagging of natural sounds using marginal co-regularization
abstract
Automatic and fast tagging of natural sounds in audio collections is a very challenging task due to wide acoustic variations, the large number of possible tags, the incomplete and ambiguous tags provided by different labellers. To handle these problems, we use a co-regularization approach to learn a pair of classifiers on sound and text. The first classifier maps low-level audio features to a true tag list. The second classifier maps actively corrupted tags to the true tags, reducing incorrect mappings caused by low-level acoustic variations in the first classifier, and to augment the tags with additional relevant tags. Training the classifiers is implemented using marginal co-regularization, pair of which draws the two classifiers into agreement by a joint optimization. We evaluate this approach on two sound datasets, Freefield1010 and Task4 of DCASE2016. The results obtained show that marginal co-regularization outperforms the baseline GMM in both efficiency and effectiveness.
Qiang Huang 0007, Yong Xu 0004, Philip J. B. Jackson, Wenwu Wang 0001, Mark D. Plumbley
ICASSP4
2017 A joint detection-classification model for audio tagging of weakly labelled data
abstract
Audio tagging aims to assign one or several tags to an audio clip. Most of the datasets are weakly labelled, which means only the tags of the clip are known, without knowing the occurrence time of the tags. The labeling of an audio clip is often based on the audio events in the clip and no event level label is provided to the user. Previous works have used the bag of frames model assume the tags occur all the time, which is not the case in practice. We propose a joint detection-classification (JDC) model to detect and classify the audio clip simultaneously. The JDC model has the ability to attend to informative and ignore uninformative sounds. Then only informative regions are used for classification. Experimental results on the “CHiME Home” dataset show that the JDC model reduces the equal error rate (EER) from 19.0% to 16.9%. More interestingly, the audio event detector is trained successfully without needing the event level label.
Qiuqiang Kong, Yong Xu 0004, Wenwu Wang 0001, Mark D. Plumbley
ICASSP3
2017 Particle flow for sequential Monte Carlo implementation of probability hypothesis density
abstract
Target tracking is a challenging task and generally no analytical solution is available, especially for the multi-target tracking systems. To address this problem, probability hypothesis density (PHD) filter is used by propagating the PHD instead of the full multi-target posterior. Recently, the particle flow filter based on the log homotopy provides a new way for state estimation. In this paper, we propose a novel sequential Monte Carlo (SMC) implementation for the PHD filter assisted by the particle flow (PF), which is called PF-SMC-PHD filter. Experimental results show that our proposed filter has higher accuracy than the SMC-PHD filter and is computationally cheaper than the Gaussian mixture PHD (GM-PHD) filter.
Yang Liu 0175, Wenwu Wang 0001, Yuxin Zhao 0001
ICASSP2
2017 A greedy algorithm with learned statistics for sparse signal reconstruction
abstract
We address the problem of sparse signal reconstruction from a few noisy samples. Recently, a Covariance-Assisted Matching Pursuit (CAMP) algorithm has been proposed, improving the sparse coefficient update step of the classic Orthogonal Matching Pursuit (OMP) algorithm. CAMP allows the a-priori mean and covariance of the non-zero coefficients to be considered in the coefficient update step. In this paper, we analyze CAMP, which leads to a new interpretation of the update step as a maximum-a-posteriori (MAP) estimation of the non-zero coefficients at each step. We then propose to leverage this idea, by finding a MAP estimate of the sparse reconstruction problem, in a greedy OMP-like way. Our approach allows the statistical dependencies between sparse coefficients to be modelled, while keeping the practicality of OMP. Experiments show improved performance when reconstructing the signal from a few noisy samples.
Lucas Rencker, Wenwu Wang 0001, Mark D. Plumbley
ICASSP2
2017 Convolutional gated recurrent neural network incorporating spatial features for audio tagging
abstract
Environmental audio tagging is a newly proposed task to predict the presence or absence of a specific audio event in a chunk. Deep neural network (DNN) based methods have been successfully adopted for predicting the audio tags in the domestic audio scene. In this paper, we propose to use a convolutional neural network (CNN) to extract robust features from mel-filter banks (MFBs), spectrograms or even raw waveforms for audio tagging. Gated recurrent unit (GRU) based recurrent neural networks (RNNs) are then cascaded to model the long-term temporal structure of the audio signal. To complement the input information, an auxiliary CNN is designed to learn on the spatial features of stereo recordings. We evaluate our proposed methods on Task 4 (audio tagging) of the Detection and Classification of Acoustic Scenes and Events 2016 (DCASE 2016) challenge. Compared with our recent DNN-based method, the proposed structure can reduce the equal error rate (EER) from 0.13 to 0.11 on the development set. The spatial features can further reduce the EER to 0.10. The performance of the end-to-end learning on raw waveforms is also comparable. Finally, on the evaluation set, we get the state-of-the-art performance with 0.12 EER while the performance of the best existing system is 0.15 EER.
Yong Xu 0004, Qiuqiang Kong, Qiang Huang 0007, Wenwu Wang 0001, Mark D. Plumbley
IJCNN4
2017 Matrix of Polynomials Model Based Polynomial Dictionary Learning Method for Acoustic Impulse Response Modeling
abstract
We study the problem of dictionary learning for signals that can be represented as polynomials or polynomial matrices, such as convolutive signals with time delays or acoustic impulse responses.Recently, we developed a method for polynomial dictionary learning based on the fact that a polynomial matrix can be expressed as a polynomial with matrix coefficients, where the coefficient of the polynomial at each time lag is a scalar matrix.However, a polynomial matrix can be also equally represented as a matrix with polynomial elements.In this paper, we develop an alternative method for learning a polynomial dictionary and a sparse representation method for polynomial signal reconstruction based on this model.The proposed methods can be used directly to operate on the polynomial matrix without having to access its coefficients matrices.We demonstrate the performance of the proposed method for acoustic impulse response modeling.
Jian Guan 0001, Xuan Wang 0002, Pengming Feng, Jing Dong 0001, Wenwu Wang 0001
INTERSPEECH5
2017 Attention and Localization Based on a Deep Convolutional Recurrent Model for Weakly Supervised Audio Tagging
abstract
Audio tagging aims to perform multi-label classification on audio \nchunks and it is a newly proposed task in the Detection and \nClassification of Acoustic Scenes and Events 2016 (DCASE \n2016) challenge. This task encourages research efforts to better \nanalyze and understand the content of the huge amounts of \naudio data on the web. The difficulty in audio tagging is that \nit only has a chunk-level label without a frame-level label. This \npaper presents a weakly supervised method to not only predict \nthe tags but also indicate the temporal locations of the occurred \nacoustic events. The attention scheme is found to be effective \nin identifying the important frames while ignoring the unrelated \nframes. The proposed framework is a deep convolutional recurrent \nmodel with two auxiliary modules: an attention module \nand a localization module. The proposed algorithm was evaluated \non the Task 4 of DCASE 2016 challenge. State-of-the-art \nperformance was achieved on the evaluation set with equal error \nrate (EER) reduced from 0.13 to 0.11, compared with the \nconvolutional recurrent baseline system.
Yong Xu 0004, Qiuqiang Kong, Qiang Huang 0007, Wenwu Wang 0001, Mark D. Plumbley
INTERSPEECH4
2017 Binaural and log-power spectra features with deep neural networks for speech-noise separation
abstract
Binaural features of interaural level difference and interaural phase difference have proved to be very effective in training deep neural networks (DNNs), to generate time-frequency masks for target speech extraction in speech-speech mixtures. However, effectiveness of binaural features is reduced in more common speech-noise scenarios, since the noise may over-shadow the speech in adverse conditions. In addition, the reverberation also decreases the sparsity of binaural features and therefore adds difficulties to the separation task. To address the above limitations, we highlight the spectral difference between speech and noise spectra and incorporate the log-power spectra features to extend the DNN input. Tested on two different reverberant rooms at different signal to noise ratios (SNR), our proposed method shows advantages over the baseline method using only binaural features in terms of signal to distortion ratio (SDR) and Short-Time Perceptual Intelligibility (STOI).
Alfredo Zermini, Qingju Liu, Yong Xu 0004, Mark D. Plumbley, Dave Betts, Wenwu Wang 0001
MMSP6
2017 Sparse analysis model based multiplicative noise removal with enhanced regularization
Jing Dong 0001, Zi-Fa Han, Yuxin Zhao 0001, Wenwu Wang 0001, Ales Procházka, Jonathon A. Chambers
Signal Process.4
2017 Sparse ℓ1-Optimal Multiloudspeaker Panning and Its Relation to Vector Base Amplitude Panning
abstract
Panning techniques, such as vector base amplitude panning (VBAP), are a widely used practical approach for spatial sound reproduction using multiple loudspeakers. Although limited to a relatively small listening area, they are very efficient and offer good localization accuracy, timbral quality, as well as a graceful degradation of quality outside the sweet spot. The aim of this paper is to investigate optimal sound reproduction techniques that adopt some of the advantageous properties of VBAP, such as the sparsity and the locality of the active loudspeakers for the reproduction of a single audio object. To this end, we state the task of multiloudspeaker panning as an ℓ1optimization problem. We demonstrate and prove that the resulting solutions are exactly sparse. Moreover, we show the effect of adding a nonnegativity constraint on the loudspeaker gains in order to preserve the locality of the panning solution. Adding this constraint, ℓ1-optimal panning can be formulated as a linear program. Using this representation, we prove that unique ℓ1-optimal panning solutions incorporating a nonnegativity constraint are identical to VBAP using a Delaunay triangulation for the loudspeaker setup. Using results from linear programming and duality theory, we describe properties and special cases, such as solution ambiguity, of the VBAP solution.
Andreas Franck, Wenwu Wang 0001, Filippo Maria Fazi
IEEE ACM Trans. Audio Speech Lang. Process.2
2017 Acoustic Reflector Localization: Novel Image Source Reversion and Direct Localization Methods
abstract
Acoustic reflector localization is an important issue in audio signal processing, with direct applications in spatial audio, scene reconstruction, and source separation. Several methods have recently been proposed to estimate the 3-D positions of acoustic reflectors given room impulse responses (RIRs). In this paper, we categorize these methods as “image-source reversion,” which localizes the image source before finding the reflector position, and “direct localization,” which localizes the reflector without intermediate steps. We present five new contributions. First, an onset detector, called the clustered dynamic programing projected phase-slope algorithm, is proposed to automatically extract the time of arrival for early reflections within the RIRs of a compact microphone array. Second, we propose an image-source reversion method that uses the RIRs from a single loudspeaker. It is constructed by combining an image source locator (the image source direction and range (ISDAR) algorithm), and a reflector locator (using the loudspeaker-image bisection (LIB) algorithm). Third, two variants of it, exploiting multiple loudspeakers, are proposed. Fourth, we present a direct localization method, the ellipsoid tangent sample consensus (ETSAC), exploiting ellipsoid properties to localize the reflector. Finally, systematic experiments on simulated and measured RIRs are presented, comparing the proposed methods with the state-of-the-art. ETSAC generates errors lower than the alternative methods compared through our datasets. Nevertheless, the ISDAR-LIB combination performs well and has a run time 200 times faster than ETSAC.
Luca Remaggi, Philip J. B. Jackson, Philip Coleman, Wenwu Wang 0001
IEEE ACM Trans. Audio Speech Lang. Process.4
2017 Unsupervised Feature Learning Based on Deep Models for Environmental Audio Tagging
abstract
Environmental audio tagging aims to predict only the presence or absence of certain acoustic events in the interested acoustic scene. In this paper, we make contributions to audio tagging in two parts, respectively, acoustic modeling and feature learning. We propose to use a shrinking deep neural network (DNN) framework incorporating unsupervised feature learning to handle the multilabel classification task. For the acoustic modeling, a large set of contextual frames of the chunk are fed into the DNN to perform a multilabel classification for the expected tags, considering that only chunk (or utterance) level rather than frame-level labels are available. Dropout and background noise aware training are also adopted to improve the generalization capability of the DNNs. For the unsupervised feature learning, we propose to use a symmetric or asymmetric deep denoising auto-encoder (syDAE or asyDAE) to generate new data-driven features from the logarithmic Mel-filter banks features. The new features, which are smoothed against background noise and more compact with contextual information, can further improve the performance of the DNN baseline. Compared with the standard Gaussian mixture model baseline of the DCASE 2016 audio tagging challenge, our proposed method obtains a significant equal error rate (EER) reduction from 0.21 to 0.13 on the development set. The proposed asyDAE system can get a relative 6.7% EER reduction compared with the strong DNN baseline on the development set. Finally, the results also show that our approach obtains the state-of-the-art performance with 0.15 EER on the evaluation set of the DCASE 2016 audio tagging task while EER of the first prize of this challenge is 0.17.
Yong Xu 0004, Qiang Huang 0007, Wenwu Wang 0001, Peter Foster, Siddharth Sigtia, Philip J. B. Jackson, Mark D. Plumbley
IEEE ACM Trans. Audio Speech Lang. Process.3
2017 Social Force Model-Based MCMC-OCSVM Particle PHD Filter for Multiple Human Tracking
abstract
Video-based multiple human tracking often involves several challenges, including target number variation, object occlusions, and noise corruption in sensor measurements. In this paper, we propose a novel method to address these challenges based on probability hypothesis density (PHD) filtering with a Markov chain Monte Carlo (MCMC) implementation. More specifically, a novel social force model (SFM) for describing the interaction between the targets is used to calculate the likelihood within the MCMC resampling step in the prediction step of the PHD filter, and a one class support vector machine (OCSVM) is then used in the update step to mitigate the noise in the measurements, where the SVM is trained with features from both color and oriented gradient histograms. The proposed method is evaluated and compared with state-of-the-art techniques using sequences from the CAVIAR, TUD, and PETS2009 datasets based on the mean Euclidean tracking error on each frame, the optimal subpattern assignment metric, and the multiple object tracking precision metric. The results show improved performance of the proposed method over the baseline algorithms, including the traditional particle PHD filtering method, the traditional SFM-based particle filtering method, multi-Bernoulli filtering, and an online-learning-based tracking method.
Pengming Feng, Wenwu Wang 0001, Satnam Singh Dlay, Syed M. Naqvi, Jonathon A. Chambers
IEEE Trans. Multim.2
2017 Semisupervised Online Multikernel Similarity Learning for Image Retrieval
abstract
Metric learning plays a fundamental role in the fields of multimedia retrieval and pattern recognition. Recently, an online multikernel similarity (OMKS) learning method has been presented for content-based image retrieval (CBIR), which was shown to be promising for capturing the intrinsic nonlinear relations within multimodal features from large-scale data. However, the similarity function in this method is learned only from labeled images. In this paper, we present a new framework to exploit unlabeled images and develop a semisupervised OMKS algorithm. The proposed method is a multistage algorithm consisting of feature selection, selective ensemble learning, active sample selection, and triplet generation. The novel aspects of our work are the introduction of classification confidence to evaluate the labeling process and select the reliably labeled images to train the metric function, and a method for reliable triplet generation, where a new criterion for sample selection is used to improve the accuracy of label prediction for unlabeled images. Our proposed method offers advantages in challenging scenarios, in particular, for a small set of labeled images with high-dimensional features. Experimental results demonstrate the effectiveness of the proposed method as compared with several baseline methods.
Jianqing Liang, Qinghua Hu, Wenwu Wang 0001, Yahong Han
IEEE Trans. Multim.3
2016 Social force model aided robust particle PHD filter for multiple human tracking
abstract
In this paper, we propose a novel robust multiple human tracking approach based upon processing a video signal by utilizing a social force model to enhance the particle probability hypothesis density (PHD) filter. In traditional dynamic models, the states of targets are only predicted by their own history; however, in multiple human tracking, the information from interaction between targets and the intentions of each target can be employed to obtain more robust prediction. Furthermore, such information can mitigate the problems of collision and occlusion. The cardinality of variable number of targets can also be estimated by using the PHD filter, hence improving the overall accuracy of the multiple human tracker. In this work, a background subtraction step has also been employed to identify the new born targets and provide the measurement set for the PHD filter. To evaluate tracking performance, sequences from both the CAVIAR and PETS2009 datasets are employed for evaluation, which shows clear improvement of the proposed method over the conventional particle PHD filter.
Pengming Feng, Wenwu Wang 0001, Syed M. Naqvi, Satnam Singh Dlay, Jonathon A. Chambers
ICASSP2
2016 Identity association using PHD filters in multiple head tracking with depth sensors
abstract
The work on 3D human pose estimation has been through a significant amount of progress in recent years, particularly due to the widespread availability of commodity depth sensors. However, most pose estimation methods follow a tracking-as-detection approach which does not explicitly handle occlusions, thus introducing outliers and identity association issues when multiple targets are involved. To address these issues, we propose a new method based on Probability Hypothesis Density (PHD) filter. In this method, the PHD filter with a novel clutter intensity model is used to remove outliers in the 3D head detection results, followed by an identity association scheme with occlusion detection for the targets. Experimental results show that our proposed method greatly mitigates the outliers, and correctly associates identities to individual detections with low computational cost.
Qingju Liu, Teófilo Emídio de Campos, Wenwu Wang 0001, Adrian Hilton 0001
ICASSP3
2016 Predicting Binaural Speech Intelligibility from Signals Estimated by a Blind Source Separation Algorithm
abstract
State-of-the-art binaural objective intelligibility measures (OIMs) require individual source signals for making intelligibility predictions, limiting their usability in real-time online operations. This limitation may be addressed by a blind source separation (BSS) process, which is able to extract the underlying sources from a mixture. In this study, a speech source is presented with either a stationary noise masker or a fluctuating noise masker whose azimuth varies in a horizontal plane, at two speech-to-noise ratios (SNRs). Three binaural OIMs are used to predict speech intelligibility from the signals separated by a BSS algorithm. The model predictions are compared with listeners' word identification rate in a perceptual listening experiment. The results suggest that with SNR compensation to the BSS-separated speech signal, the OIMs can maintain their predictive power for individual maskers compared to their performance measured from the direct signals. It also reveals that the errors in SNR between the estimated signals are not the only factors that decrease the predictive accuracy of the OIMs with the separated signals. Artefacts or distortions on the estimated signals caused by the BSS algorithm may also be concerns.
Qingju Liu, Philip J. B. Jackson, Wenwu Wang 0001
INTERSPEECH4
2016 Higher-Order Circularity Based I/Q Imbalance Compensation in Direct-Conversion Receivers
abstract
In-phase and quadrature-phase (I/Q) imbalance is a critical issue limit the achievable operating signal-to-noise ratio (SNR) at the receiver in direct conversion architecture. In recent literatures, the second-and fourth-order circularity property of communication signals have been used for designing compensator to eliminate the I/Q imbalance. In this paper, we investigate whether moment circularity of an order higher than four can be used in receiver I/Q imbalance compensation. It is shown that the sixth-order moment E[z4z*2] is a suitable statistic for measuring the sixth-order circularity of representative communication signals such as M-QAM and M-PSK with M > 2. Two blind algorithms are then proposed to update the coefficients of I/Q imbalance compensator by restoring the sixth-order circularity of the compensator output signal. Simulation results show that the new proposed methods based on sixth-order statistic converges faster or gives lower steady-state variance than the reference methods that are based on second-and fourth-order statistics.
Fanglin Gu, Shan Wang 0005, Jibo Wei, Wenwu Wang 0001
VTC Fall4
2016 Audio head pose estimation using the direct to reverberant speech ratio
abstract
Head pose is an important cue in many applications such as, speech recognition and face recognition. Most approaches to head pose estimation to date have focussed on the use of visual information of a subject’s head. These visual approaches have a number of limitations such as, an inability to cope with occlusions, changes in the appearance of the head, and low resolution images . We present here a novel method for determining coarse head pose orientation purely from audio information, exploiting the direct to reverberant speech energy ratio (DRR) within a reverberant room environment. Our hypothesis is that a speaker facing towards a microphone will have a higher DRR and a speaker facing away from the microphone will have a lower DRR. This method has the advantage of actually exploiting the reverberations within a room rather than trying to suppress them. This also has the practical advantage that most enclosed living spaces, such as meeting rooms or offices are highly reverberant environments. In order to test this hypothesis we also present a new data set featuring 56 subjects recorded in three different rooms, with different acoustic properties , adopting 8 different head poses in 4 different room positions captured with a 16 element microphone array . As far as the authors are aware this data set is unique and will make a significant contribution to further work in the area of audio head pose estimation . Using this data set we demonstrate that our proposed method of using the DRR for audio head pose estimation provides a significant improvement over previous methods.
Mark Barnard, Wenwu Wang 0001
Speech Commun.2
2016 Adaptive Retrodiction Particle PHD Filter for Multiple Human Tracking
abstract
The probability hypothesis density (PHD) filter is well known for addressing the problem of multiple human tracking for a variable number of targets, and the sequential Monte Carlo implementation of the PHD filter, known as the particle PHD filter, can give state estimates with nonlinear and non-Gaussian models. Recently, Mahler et al. have introduced a PHD smoother to gain more accurate estimates for both target states and number. However, as highlighted by Psiaki in the context of a backward-smoothing extended Kalman filter, with a nonlinear state evolution model the approximation error in the backward filtering requires careful consideration. Psiaki suggests that to minimize the aggregated least-squares error over a batch of data. We instead use the term retrodiction PHD filter to describe the backward filtering algorithm in recognition of the approximation error proposed in the original PHD smoother, and we propose an adaptive recursion step to improve the approximation accuracy. This step combines forward and backward processing through the measurement set and thereby mitigates the problems with the original PHD smoother when the target number changes significantly and the targets appear and disappear randomly. Simulation results show the improved performance of the proposed algorithm and its capability in handling a variable number of targets.
Pengming Feng, Wenwu Wang 0001, Syed M. Naqvi, Jonathon A. Chambers
IEEE Signal Process. Lett.2
2016 Mean-Shift and Sparse Sampling-Based SMC-PHD Filtering for Audio Informed Visual Speaker Tracking
abstract
The probability hypothesis density (PHD) filter based on sequential Monte Carlo (SMC) approximation (also known as SMC-PHD filter) has proven to be a promising algorithm for multispeaker tracking. However, it has a heavy computational cost as surviving, spawned, and born particles need to be distributed in each frame to model the state of the speakers and to estimate jointly the variable number of speakers with their states. In particular, the computational cost is mostly caused by the born particles as they need to be propagated over the entire image in every frame to detect the new speaker presence in the view of the visual tracker. In this paper, we propose to use the audio data to improve the visual SMC-PHD (V-SMC-PHD) filter by using the direction of arrival angles of the audio sources to determine when to propagate the born particles and reallocate the surviving and spawned particles. The tracking accuracy of the audio-visual SMC-PHD (AV-SMC-PHD) algorithm is further improved by using a modified mean-shift algorithm to search and climb density gradients iteratively to find the peak of the probability distribution, and the extra computational complexity introduced by mean-shift is controlled with a sparse sampling technique. These improved algorithms, named as AVMS-SMC-PHD and sparse-AVMS-SMC-PHD, respectively, are compared systematically with AV-SMC-PHD and V-SMC-PHD based on the AV16.3, AMI, and CLEAR datasets.
Volkan Kilic, Mark Barnard, Wenwu Wang 0001, Adrian Hilton 0001, Josef Kittler
IEEE Trans. Multim.3
2015 A 3D model for room boundary estimation
abstract
Estimating the geometric properties of an indoor environment through acoustic room impulse responses (RIRs) is useful in various applications, e.g., source separation, simultaneous localization and mapping, and spatial audio. Previously, we developed an algorithm to estimate the reflector's position by exploiting ellipses as projection of 3D spaces. In this article, we present a model for full 3D reconstruction of environments. More specifically, the three components of the previous method, respectively, MUSIC for direction of arrival (DOA) estimation, numerical search adopted for reflector estimation and the Hough transform to refine the results, are extended for 3D spaces. A variation is also proposed using RANSAC instead of the numerical search and the Hough transform wich significantly reduces the run time. Both methods are tested on simulated and measured RIR data. The proposed methods perform better than the baseline, reducing the estimation error.
Luca Remaggi, Philip J. B. Jackson, Wenwu Wang 0001, Jonathon A. Chambers
ICASSP3
2015 Audio informed visual speaker tracking with SMC-PHD filter
abstract
Sequential Monte Carlo probability hypothesis density (SMC-PHD) filter has received much interest in the field of nonlinear non-Gaussian visual tracking due to its ability to handle a variable number of speakers. The SMC-PHD filter employs surviving, spawned and born particles to model the state of the speakers and jointly estimates the variable number of speakers with their states. The born particles play a critical role in the detection of new speakers, which makes it necessary to propagate them in each frame. However, this increases the computational cost of the visual tracker. Here, we propose to use audio data to determine when to propagate the born particles and re-allocate the surviving and spawned particles. In our framework, we employ audio data as an aid to visual SMC-PHD (V-SMC-PHD) filter by using the direction of arrival (DOA) angles of the audio sources to reshape the distribution of the particles. Experimental results on the AV16:3 dataset with multi-speaker sequences show that our proposed audio-visual SMC-PHD (AV-SMC-PHD) filter improves the tracking performance in terms of estimation accuracy and computational efficiency.
Volkan Kilic, Mark Barnard, Wenwu Wang 0001, Adrian Hilton 0001, Josef Kittler
ICME3
2015 Reverberant speech separation with probabilistic time-frequency masking for B-format recordings
Xiaoyi Chen 0002, Wenwu Wang 0001, Yingmin Wang, Xionghu Zhong, Atiyeh Alinaghi
Speech Commun.2
2015 Audio Assisted Robust Visual Tracking With Adaptive Particle Filtering
abstract
The problem of tracking multiple moving speakers in indoor environments has received much attention. Earlier techniques were based purely on a single modality, e.g., vision. Recently, the fusion of multi-modal information has been shown to be instrumental in improving tracking performance, as well as robustness in the case of challenging situations like occlusions (by the limited field of view of cameras or by other speakers). However, data fusion algorithms often suffer from noise corrupting the sensor measurements which cause non-negligible detection errors. Here, a novel approach to combining audio and visual data is proposed. We employ the direction of arrival angles of the audio sources to reshape the typical Gaussian noise distribution of particles in the propagation step and to weight the observation model in the measurement step. This approach is further improved by solving a typical problem associated with the PF, whose efficiency and accuracy usually depend on the number of particles and noise variance used in state estimation and particle propagation. Both parameters are specified beforehand and kept fixed in the regular PF implementation which makes the tracker unstable in practice. To address these problems, we design an algorithm which adapts both the number of particles and noise variance based on tracking error and the area occupied by the particles in the image. Experiments on the AV16.3 dataset show the advantage of our proposed methods over the baseline PF method and an existing adaptive PF algorithm for tracking occluded speakers with a significantly reduced number of particles.
Volkan Kilic, Mark Barnard, Wenwu Wang 0001, Josef Kittler
IEEE Trans. Multim.3
2015 Heterogeneous Feature Selection With Multi-Modal Deep Neural Networks and Sparse Group LASSO
abstract
Heterogeneous feature representations are widely used in machine learning and pattern recognition, especially for multimedia analysis. The multi-modal, often also high- dimensional , features may contain redundant and irrelevant information that can deteriorate the performance of modeling in classification. It is a challenging problem to select the informative features for a given task from the redundant and heterogeneous feature groups. In this paper, we propose a novel framework to address this problem. This framework is composed of two modules, namely, multi-modal deep neural networks and feature selection with sparse group LASSO. Given diverse groups of discriminative features, the proposed technique first converts the multi-modal data into a unified representation with different branches of the multi-modal deep neural networks. Then, through solving a sparse group LASSO problem, the feature selection component is used to derive a weight vector to indicate the importance of the feature groups. Finally, the feature groups with large weights are considered more relevant and hence are selected. We evaluate our framework on three image classification datasets. Experimental results show that the proposed approach is effective in selecting the relevant feature groups and achieves competitive classification performance as compared with several recent baseline methods.
Qinghua Hu, Wenwu Wang 0001
IEEE Trans. Multim.3
2014 Audio-visual tracking of a variable number of speakers with a random finite set approach
Volkan Kilic, Xionghu Zhong, Mark Barnard, Wenwu Wang 0001, Josef Kittler
FUSION4
2014 A Bayesian performance bound for time-delay of arrival based acoustic source tracking in a reverberant environment
Xionghu Zhong, Wenwu Wang 0001, Syed M. Naqvi, Chng Eng Siong
FUSION2
2014 Analysis SimCO: A new algorithm for analysis dictionary learning
abstract
We consider the dictionary learning problem for the analysis model based sparse representation. A novel algorithm is proposed by adapting the synthesis model based simultaneous codeword optimisation (SimCO) algorithm to the analysis model. This algorithm assumes that the analysis dictionary contains unit Ł2-norm atoms and trains the dictionary by the optimisation on manifolds. This framework allows one to update multiple dictionary atoms in each iteration, leading to a computationally efficient optimisation process. We demonstrate the competitive performance of the proposed algorithm using experiments on both synthetic and real data, as compared with three baseline algorithms, Analysis K-SVD, analysis operator learning (AOL) and learning overcomplete sparsifying transforms (LOST), respectively.
Jing Dong 0001, Wenwu Wang 0001, Wei Dai 0001
ICASSP2
2014 Joint Mixing Vector and Binaural Model Based Stereo Source Separation
abstract
In this paper the mixing vector (MV) in the statistical mixing model is compared to the binaural cues represented by interaural level and phase differences (ILD and IPD). It is shown that the MV distributions are quite distinct while binaural models overlap when the sources are close to each other. On the other hand, the binaural cues are more robust to high reverberation than MV models. According to this complementary behavior we introduce a new robust algorithm for stereo speech separation which considers both additive and convolutive noise signals to model the MV and binaural cues in parallel and estimate probabilistic time-frequency masks. The contribution of each cue to the final decision is also adjusted by weighting the log-likelihoods of the cues empirically. Furthermore, the permutation problem of the frequency domain blind source separation (BSS) is addressed by initializing the MVs based on binaural cues. Experiments are performed systematically on determined and underdetermined speech mixtures in five rooms with various acoustic properties including anechoic, highly reverberant, and spatially-diffuse noise conditions. The results in terms of signal-to-distortion-ratio (SDR) confirm the benefits of integrating the MV and binaural cues, as compared with two state-of-the-art baseline algorithms which only use MV or the binaural cues.
Atiyeh Alinaghi, Philip J. B. Jackson, Qingju Liu, Wenwu Wang 0001
IEEE ACM Trans. Audio Speech Lang. Process.4
2014 Robust Multi-Speaker Tracking via Dictionary Learning and Identity Modeling
abstract
We investigate the problem of visual tracking of multiple human speakers in an office environment. In particular, we propose novel solutions to the following challenges: (1) robust and computationally efficient modeling and classification of the changing appearance of the speakers in a variety of different lighting conditions and camera resolutions; (2) dealing with full or partial occlusions when multiple speakers cross or come into very close proximity; (3) automatic initialization of the trackers, or re-initialization when the trackers have lost lock caused by e.g. the limited camera views. First, we develop new algorithms for appearance modeling of the moving speakers based on dictionary learning (DL), using an off-line training process. In the tracking phase, the histograms (coding coefficients) of the image patches derived from the learned dictionaries are used to generate the likelihood functions based on Support Vector Machine (SVM) classification. This likelihood function is then used in the measurement step of the classical particle filtering (PF) algorithm. To improve the computational efficiency of generating the histograms, a soft voting technique based on approximate Locality-constrained Soft Assignment (LcSA) is proposed to reduce the number of dictionary atoms (codewords) used for histogram encoding. Second, an adaptive identity model is proposed to track multiple speakers whilst dealing with occlusions. This model is updated online using Maximum a Posteriori (MAP) adaptation, where we control the adaptation rate using the spatial relationship between the subjects. Third, to enable automatic initialization of the visual trackers, we exploit audio information, the Direction of Arrival (DOA) angle, derived from microphone array recordings. Such information provides, a priori, the number of speakers and constrains the search space for the speaker's faces. The proposed system is tested on a number of sequences from three publicly available and challenging data corpora (AV16.3, EPFL pedestrian data set and CLEAR) with up to five moving subjects.
Mark Barnard, Piotr Koniusz, Wenwu Wang 0001, Josef Kittler, Syed M. Naqvi, Jonathon A. Chambers
IEEE Trans. Multim.3
2014 Interference Reduction in Reverberant Speech Separation With Visual Voice Activity Detection
abstract
The visual modality, deemed to be complementary to the audio modality, has recently been exploited to improve the performance of blind source separation (BSS) of speech mixtures, especially in adverse environments where the performance of audio-domain methods deteriorates steadily. In this paper, we present an enhancement method to audio-domain BSS with the integration of voice activity information, obtained via a visual voice activity detection (VAD) algorithm. Mimicking aspects of human hearing, binaural speech mixtures are considered in our two-stage system. Firstly, in the off-line training stage, a speaker-independent voice activity detector is formed using the visual stimuli via the adaboosting algorithm. In the on-line separation stage, interaural phase difference (IPD) and interaural level difference (ILD) cues are statistically analyzed to assign probabilistically each time-frequency (TF) point of the audio mixtures to the source signals. Next, the detected voice activity cues (found via the visual VAD) are integrated to reduce the interference residual. Detection of the interference residual takes place gradually, with two layers of boundaries in the correlation and energy ratio map. We have tested our algorithm on speech mixtures generated using room impulse responses at different reverberation times and noise levels. Simulation results show performance improvement of the proposed method for target speech extraction in noisy and reverberant environments, in terms of signal-to-interference ratio (SIR) and perceptual evaluation of speech quality (PESQ).
Qingju Liu, Andrew J. Aubrey, Wenwu Wang 0001
IEEE Trans. Multim.3
2013 Audio-visual face detection for tracking in a meeting room environment
Mark Barnard, Wenwu Wang 0001, Josef Kittler, Syed M. Naqvi, Jonathon A. Chambers
FUSION2
2013 Acoustic source tracking in a reverberant environment using a pairwise synchronous microphone network
Xionghu Zhong, Arash Mohammadi 0001, Wenwu Wang 0001, A. Benjamin Premkumar, Amir Asif
FUSION3
2013 Spatial and coherence cues based time-frequency masking for binaural reverberant speech separation
abstract
Most of the binaural source separation algorithms only consider the dissimilarities between the recorded mixtures such as interaural phase and level differences (IPD, ILD) to classify and assign the time-frequency (T-F) regions of the mixture spectrograms to each source. However, in this paper we show that the coherence between the left and right recordings can provide extra information to label the T-F units from the sources. This also reduces the effect of reverberation which contains random reflections from different directions showing low correlation between the sensors. Our algorithm assigns the T-F regions into original sources based on weighted combination of IPD, ILD, the mixing vector models and the estimated interaural coherence (IC) between the left and right recordings. The binaural room impulse responses measured in four rooms with various acoustic conditions have been used to evaluate the performance of the proposed method which shows an average improvement of more than 2.23 dB in signal-to-distortion ratio (SDR) in room D with T60= 0.89 s over the state-of-the-art algorithms.
Atiyeh Alinaghi, Wenwu Wang 0001, Philip J. B. Jackson
ICASSP2
2013 Audio head pose estimation using the direct to reverberant speech ratio
abstract
Head pose is an important cue in many applications such as, speech recognition and face recognition. Most approaches to head pose estimation to date have used visual information to model and recognise a subject's head in different configurations. These approaches have a number of limitations such as, inability to cope with occlusions, changes in the appearance of the head, and low resolution images. We present here a novel method for determining coarse head pose orientation purely from audio information, exploiting the direct to reverberant speech energy ratio (DRR) within a highly reverberant meeting room environment. Our hypothesis is that a speaker facing towards a microphone will have a higher DRR and a speaker facing away from the microphone will have a lower DRR. This hypothesis is confirmed by experiments conducted on the publicly available AV16.3 database.
Mark Barnard, Wenwu Wang 0001, Josef Kittler
ICASSP2
2013 Audio constrained particle filter based visual tracking
abstract
We present a robust and efficient audio-visual (AV) approach to speaker tracking in a room environment. A challenging problem with visual tracking is to deal with occlusions (caused by the limited field of view of cameras or by other speakers). Another challenge is associated with the particle filtering (PF) algorithm, commonly used for visual tracking, which requires a large number of particles to ensure the distribution is well modelled. In this paper, we propose a new method of fusing audio into the PF based visual tracking. We use the direction of arrival angles (DOAs) of the audio sources to reshape the typical Gaussian noise distribution of particles in the propagation step and to weight the observation model in the measurement step. Experiments on AV16.3 datasets show the advantage of our proposed method over the baseline PF method for tracking occluded speakers with a significantly reduced number of particles.
Volkan Kilic, Mark Barnard, Wenwu Wang 0001, Josef Kittler
ICASSP3
2013 Sparse coding with adaptive dictionary learning for underdetermined blind speech separation
Tao Xu 0038, Wenwu Wang 0001, Wei Dai 0001
Speech Commun.2
2013 Video-Aided Model-Based Source Separation in Real Reverberant Rooms
abstract
Source separation algorithms that utilize only audio data can perform poorly if multiple sources or reverberation are present. In this paper we therefore propose a video-aided model-based source separation algorithm for a two-channel reverberant recording in which the sources are assumed static. By exploiting cues from video, we first localize individual speech sources in the enclosure and then estimate their directions. The interaural spatial cues, the interaural phase difference and the interaural level difference, as well as the mixing vectors are probabilistically modeled. The models make use of the source direction information and are evaluated at discrete time-frequency points. The model parameters are refined with the well-known expectation-maximization (EM) algorithm. The algorithm outputs time-frequency masks that are used to reconstruct the individual sources. Simulation results show that by utilizing the visual modality the proposed algorithm can produce better time-frequency masks thereby giving improved source estimates. We provide experimental results to test the proposed algorithm in different scenarios and provide comparisons with both other audio-only and audio-visual algorithms and achieve improved performance both on synthetic and real data. We also include dereverberation based pre-processing in our algorithm in order to suppress the late reverberant components from the observed stereo mixture and further enhance the overall output of the algorithm. This advantage makes our algorithm a suitable candidate for use in under-determined highly reverberant settings where the performance of other audio-only and audio-visual methods is limited.
Muhammad Salman Khan 0001, Syed M. Naqvi, Ata ur-Rehman, Wenwu Wang 0001, Jonathon A. Chambers
IEEE Trans. Speech Audio Process.4
2012 A dictionary learning approach to tracking
abstract
The problem of tracking people using multiple cameras is of much current interest as a means of providing cues for audio-visual blind source separation in dynamic environments. Here we investigate the use of one of the current state-of-the-art techniques in object recognition combined with one of the most popular methods of modelling object motion, particle filters, for tracking people. The dictionary learning or Bag-of-Words approach to object recognition has proved to be very effective in recent years, as shown in a number of large comparisons such as the PASCAL Visual Object recognition Challenge (VOC). In this paper we use this proven object recognition method within the framework of a particle filter. This provides a more accurate and robust tracking of people in a multiple camera environment. We also demonstrate that the dictionary learning approach can provide a principled method for the fusion of multiple features.
Mark Barnard, Wenwu Wang 0001, Josef Kittler, Syed M. Naqvi, Jonathon A. Chambers
ICASSP2
2012 Dictionary learning and update based on simultaneous codeword optimization (SimCO)
abstract
Dictionary learning aims to adapt elementary codewords directly from training data so that each training signal can be best approximated by a linear combination of only a few codewords. Following the two-stage iterative processes: sparse coding and dictionary update, that are commonly used, for example, in the algorithms of MOD and K-SVD, we propose a novel framework that allows one to update an arbitrary set of codewords and the corresponding sparse coefficients simultaneously, hence termed simultaneous codeword optimization (SimCO). Under this framework, we have developed two algorithms, namely the primitive and the regularized SimCO. Simulations are provided to show the advantages of our approach over the K-SVD algorithm in terms of both learning performance and running speed.
Wei Dai 0001, Tao Xu 0038, Wenwu Wang 0001
ICASSP3
2012 Multimodal (audio-visual) source separation exploiting multi-speaker tracking, robust beamforming and time-frequency masking
abstract
A novel multimodal source separation approach is proposed for physically moving and stationary sources which exploits a circular microphone array, multiple video cameras, robust spatial beamforming and time-frequency masking. The challenge of separating moving sources, including higher reverberation time (RT) even for physically stationary sources, is that the mixing filters are time varying; as such the unmixing filters should also be time varying but these are difficult to determine from only audio measurements. Therefore in the proposed approach, visual modality is used to facilitate the separation for both stationary and moving sources. The movement of the sources is detected by a three-dimensional tracker based on a Markov Chain Monte Carlo particle filter. The audio separation is performed by a robust least squares frequency invariant data-independent beamformer. The uncertainties in source localisation and direction of arrival information obtained from the 3D video-based tracker are controlled by using a convex optimisation approach in the beamformer design. In the final stage, the separated audio sources are further enhanced by applying a binary time-frequency masking technique in the cepstral domain. Experimental results show that using the visual modality, the proposed algorithm cannot only achieve performance better than conventional frequency-domain source separations algorithms, but also provide acceptable separation performance for moving sources.
Syed M. Naqvi, Wenwu Wang 0001, Muhammad Salman Khan 0001, Mark Barnard, Jonathon A. Chambers
IET Signal Process.2
2012 Use of bimodal coherence to resolve the permutation problem in convolutive BSS
Qingju Liu, Wenwu Wang 0001, Philip J. B. Jackson
Signal Process.2
2011 Integrating binaural cues and blind source separation method for separating reverberant speech mixtures
abstract
This paper presents a new method for reverberant speech separation, based on the combination of binaural cues and blind source separation (BSS) for the automatic classification of the time-frequency (T-F) units of the speech mixture spectrogram. The main idea is to model interaural phase difference, interaural level difference and frequency bin-wise mixing vectors by Gaussian mixture models for each source and then evaluate that model at each T-F point and assign the units with high probability to that source. The model parameters and the assigned regions are refined iteratively using the Expectation-Maximization (EM) algorithm. The proposed method also addresses the permutation problem of the frequency domain BSS by initializing the mixing vectors for each frequency channel. The EM algorithm starts with binaural cues and after a few iterations the estimated probabilistic mask is used to initialize and re-estimate the mixing vector model parameters. We performed experiments on speech mixtures, and showed an average of about 0.8 dB improvement in signal-to-distortion (SDR) over the binaural only baseline.
Atiyeh Alinaghi, Wenwu Wang 0001, Philip J. B. Jackson
ICASSP2
2011 A multistage approach to blind separation of convolutive speech mixtures
Tariqullah Jan, Wenwu Wang 0001, DeLiang Wang
Speech Commun.2
2010 A block-based compressed sensing method for underdetermined blind speech separation incorporating binary mask
abstract
A block-based compressed sensing approach coupled with binary time-frequency masking is presented for the underdetermined speech separation problem. The proposed algorithm consists of multiple steps. First, the mixed signals are segmented to a number of blocks. For each block, the unknown mixing matrix is estimated in the transform domain by a clustering algorithm. Using the estimated mixing matrix, the sources are recovered by a compressed sensing approach. The coarsely separated sources are then used to estimate the time-frequency binary masks which are further applied to enhance the separation performance. The separated source components from all the blocks are concatenated to reconstruct the whole signal. Numerical experiments are provided to show the improved separation performance of the proposed algorithm, as compared with two recent approaches. The block-based operation has the advantage in improving considerably the computational efficiency of the compressed sensing algorithm without degrading its separation performance.
Tao Xu 0038, Wenwu Wang 0001
ICASSP2
2010 Bimodal coherence based scale ambiguity cancellation for target speech extraction and enhancement
abstract
We present a novel method for extracting target speech from au-ditory mixtures using bimodal coherence, which is statistically characterised by a Gaussian mixture modal (GMM) in the off-line training process, using the robust features obtained from the audio-visual speech. We then adjust the ICA-separated spectral components using the bimodal coherence in the time-frequency domain, to mitigate the scale ambiguities in different frequency bins. We tested our algorithm on the XM2VTS database, and the results show the performance improvement with our pro-posed algorithm in terms of signal to interference ratio (SIR) measurements. Index Terms: speech extraction, bimodal coherence, audio-visual, Gaussian mixture model (GMM), independent compo-
Qingju Liu, Wenwu Wang 0001, Philip J. B. Jackson
INTERSPEECH2
2009 A multistage approach for blind separation of convolutive speech mixtures
abstract
In this paper, we propose a novel algorithm for the separation of convolutive speech mixtures using two-microphone recordings, based on the combination of independent component analysis (ICA) and ideal binary mask (IBM), together with a post-filtering process in the cepstral domain. Essentially, the proposed algorithm consists of three steps. First, a constrained convolutive ICA algorithm is applied to separate the source signals from two-microphone recordings. In the second step, we estimate the IBM by comparing the energy of corresponding time-frequency (T-F) units from the separated sources obtained with the convolutive ICA algorithm. The last step is to reduce musical noise caused typically by T-F masking using cepstral smoothing. The performance of the proposed approach is evaluated based on both reverberant mixtures generated using a simulated room model and real recordings. The proposed algorithm offers considerably higher efficiency, together with improved speech quality while producing similar separation performance as compared with a recent approach.
Tariqullah Jan, Wenwu Wang 0001, DeLiang Wang
ICASSP2
2008 Convolutive non-negative sparse coding
abstract
Non-negative sparse coding (NSC) is a powerful technique for low-rank data approximation, and has found several successful applications in signal processing. However, the temporal dependency, which is a vital clue for many realistic signals, has not been taken into account in its conventional model. In this paper, we propose a general framework, i.e., convolutive non-negative sparse coding (CNSC), by considering a convolutive model for the low-rank approximation of the original data. Using this model, we have developed an effective learning algorithm based on the multiplicative adaptation of the reconstruction error function defined by the squared Euclidean distance. The proposed algorithm is applied to the separation of music audio objects in the magnitude spectrum domain. Interesting numerical results are provided to demonstrate its advantages over both the conventional NSC and an existing convolutive coding method.
Wenwu Wang 0001
IJCNN1
2007 A New Variable Step-Size LMS Algorithm with Robustness to Nonstationary Noise
abstract
A new variable step-size least-mean-square (VSSLMS) algorithm is presented in this paper for applications in which the desired response contains nonstationary noise with high variance. The step size of the proposed VSSLMS algorithm is controlled by the normalized square Euclidean norm of the averaged gradient vector, and is henceforth referred to as the NSVSSLMS algorithm. As shown by the analysis and simulation results, the proposed algorithm has both fast convergence rate and robustness to high-variance noise signals, and performs better than Greenburg's sum method, which is a robust algorithm for applications with nonstationary noise.
Yonggang Zhang 0001, Jonathon A. Chambers, Wenwu Wang 0001, Paul Kendrick, Trevor J. Cox
ICASSP (3)3
2005 Video assisted speech source separation
abstract
We investigate the problem of integrating the complementary audio and visual modalities for speech separation. Rather than using independence criteria suggested in most blind source separation (BSS) systems, we use visual features from a video signal as additional information to optimize the unmixing matrix. We achieve this by using a statistical model characterizing the nonlinear coherence between audio and visual features as a separation criterion for both instantaneous and convolutive mixtures. We acquire the model by applying the Bayesian framework to the fused feature observations based on a training corpus. We point out several key existing challenges to the success of the system. Experimental results verify the proposed approach, which outperforms the audio only separation system in a noisy environment, and also provides a solution to the permutation problem.
Wenwu Wang 0001, Darren Cosker, Yulia Hicks, Saeid Sanei, Jonathon A. Chambers
ICASSP (5)1
2005 Variable step-size sign natural gradient algorithm for sequential blind source separation
abstract
A novel variable step-size sign natural gradient algorithm (VS-S-NGA) for online blind separation of independent sources is presented. A sign operator for the adaptation of the separation model is obtained from the derivation of a generalized dynamic separation model. A variable step size is also derived to better match the dynamics of the input signals and unmixing matrix. The proposed sign algorithm is appealing in practice due to its computational simplicity. Experimental results verify the superior convergence performance over conventional NGA in both stationary and nonstationary environments.
Lianxi Yuan, Wenwu Wang 0001, Jonathon A. Chambers
IEEE Signal Process. Lett.2
2004 A coupled HMM for solving the permutation problem in frequency domain BSS
abstract
Permutation of the outputs at different frequency bins remains as a major problem in convolutive blind source separation (BSS). A coupled hidden Markov model (CHMM) effectively exploits the psychoacoustic characteristics of signals to mitigate such permutations. A joint diagonalization algorithm has been used for convolutive BSS; it incorporates a non-unitary penalty term within the cross-power spectrum-based cost function in the frequency domain. The proposed CHMM system couples a number of conventional HMMs, equivalent to the number of outputs, by making state transitions in each model dependent, not only on its own previous state, but also on some aspects of the state of the other models. Using this method, the permutation effect is substantially reduced; it is demonstrated using a number of simulation studies.
Saeid Sanei, Wenwu Wang 0001, Jonathon A. Chambers
ICASSP (5)2