VLDB 2026 Research / reviewers in the wild / expert
Xubo Liu 0001
dblp:235/1970-1
· DBLP profile ↗
29ranked-venue papers
3as first author
29since 2021 · last 2025
0009-0004-9950-2672ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 24 · 3 first-author · 24 since 2021Artificial intelligence and machine learning · 16 · 2 first-author · 16 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Disentangling Hierarchical Features for Anomalous Sound Detection Under Domain ShiftabstractAnomalous sound detection (ASD) encounters difficulties with domain shift, where the sounds of machines in target domains differ significantly from those in source domains due to varying operating conditions. Existing methods typically employ domain classifiers to enhance detection performance, but they often overlook the influence of domain-unrelated information. This oversight can hinder the model’s ability to clearly distinguish between domains, thereby weakening its capacity to differentiate normal from abnormal sounds. In this paper, we propose a Gradient Reversal-based Hierarchical feature Disentanglement (GRHD) method to address the above challenge. GRHD uses gradient reversal to separate domain-related features from domain-unrelated ones, resulting in more robust feature representations. Additionally, the method employs a hierarchical structure to guide the learning of fine-grained, domain-specific features by leveraging available metadata, such as section IDs and machine sound attributes. Experimental results on the DCASE 2022 Challenge Task 2 dataset demonstrate that the proposed method significantly improves ASD performance under domain shift. Jian Guan 0001, Jiantong Tian, Qiaoxi Zhu, Feiyang Xiao, Hejing Zhang, Xubo Liu 0001 |
ICASSP | 6 |
| 2025 | FlowSep: Language-Queried Sound Separation with Rectified Flow MatchingabstractLanguage-queried audio source separation (LASS) focuses on separating sounds using textual descriptions of the desired sources. Current methods mainly use discriminative approaches, such as time-frequency masking, to separate target sounds and minimize interference from other sources. However, these models face challenges when separating overlapping sound-tracks, which may lead to artifacts such as spectral holes or incomplete separation. Rectified flow matching (RFM), a generative model that establishes linear relations between the distribution of data and noise, offers superior theoretical properties and simplicity, but has not yet been explored in sound separation. In this work, we introduce FlowSep, a new generative model based on RFM for LASS tasks. FlowSep learns linear flow trajectories from noise to target source features within the variational autoencoder (VAE) latent space. During inference, the RFM-generated latent features are reconstructed into a mel-spectrogram via the pre-trained VAE decoder, followed by a pre-trained vocoder to synthesize the waveform. Trained on 1, 680 hours of audio data, FlowSep outperforms the state-of-the-art models across multiple benchmarks, as evaluated with subjective and objective metrics. Additionally, our results show that FlowSep surpasses a diffusion-based LASS model in both separation quality and inference efficiency, highlighting its strong potential for audio source separation tasks. Code, pre-trained models and demos can be found at: https://audio-agi.github.io/FlowSep_demo/. Xubo Liu 0001, Haohe Liu, Mark D. Plumbley, Wenwu Wang 0001 |
ICASSP | 2 |
| 2025 | Sound-VECaps: Improving Audio Generation with Visually Enhanced CaptionsabstractGenerative models have shown significant achievements in audio generation tasks. However, existing models struggle with complex and detailed prompts, leading to potential performance degradation. We hypothesize that this problem stems from the simplicity and scarcity of the training data. This work aims to create a large-scale audio dataset with rich captions for improving audio generation models. We first develop an automated pipeline to generate detailed captions by transforming predicted visual captions, audio captions, and tagging labels into comprehensive descriptions using a Large Language Model (LLM). The resulting dataset, Sound-VECaps, comprises 1.66M high-quality audio-caption pairs with enriched details including audio event orders, occurred places and environment information. We then demonstrate that training the text-to-audio generation models with Sound-VECaps significantly improves the performance on complex prompts. Furthermore, we conduct ablation studies of the models on several downstream audio-language tasks, showing the potential of Sound-VECaps in advancing audio-text representation learning.Dataset and demos are available at https://yyua8222.github.io/Sound-VECaps-demo/. Dongya Jia, Xiaobin Zhuang, Yuanzhe Chen, Zhuo Chen 0006, Yuping Wang 0005, Yuxuan Wang 0002, Xubo Liu 0001, Xiyuan Kang, Mark D. Plumbley, Wenwu Wang 0001 |
ICASSP | 8 |
| 2025 | Contextual Reasoning for Robust Composed Image Retrieval with Vision-Language ModelsabstractComposed Image Retrieval (CIR) combines a reference image with modification text for precise and flexible searches. However, existing methods face two key challenges: first, the limited information in modification text hampers the model's ability to understand user intent, leading to reduced accuracy and diversity; second, reliance on unidirectional constraints overlooks the complementary role of reference and target captions. In this paper, we propose CR-CIR a novel framework that leverages Contextual Reasoning and vision-language models to enhance CIR. Specifically, we use a VLM (e.g., BLIP2) to address the scarcity of textual annotations in existing datasets by generating descriptive captions for both reference and target images. In addition, we enhance the modification text with contextual information using a VLM (e.g., MiniCPM), enriching the model's understanding of user intent. Then our method incorporates a Dual Reasoning Modification Module, which imposes bidirectional constraints by integrating both image and text modalities. Additionally, we introduce a Modality Shift Regularization Loss that assumes symmetry and correlation between text and image domain transformations in the latent space. This new loss function enforces consistent modality shifts, significantly enhancing the model's interpretative and generalization abilities. Experimental results on benchmark CIR datasets demonstrate that the proposed method achieves state-of-the-art (SOTA) performance. Our code and dataset will be available at https://github.com/kola1124/CR-CIR.git. Yujian Lee, Xubo Liu 0001, Hui Zhang 0062, Zailong Chen, Yiyang Hu, Guquan Jing, Yunting Lai |
ICMR | 3 |
| 2025 | Noise-Robust Sound Event Detection and Counting via Language-Queried Sound Separation
Yuanjian Chen, Yang Xiao 0019, Han Yin, Yadong Guan, Xubo Liu 0001 |
IEEE Signal Process. Lett. | 5 |
| 2025 | Multimodal Fish Feeding Intensity Assessment in AquacultureabstractFish feeding intensity assessment (FFIA) aims to evaluate fish appetite changes during feeding, which is crucial in industrial aquaculture applications. Existing FFIA methods are limited by their robustness to noise, computational complexity, and the lack of public datasets for developing the models. To address these issues, we first introduce AV-FFIA, a new dataset containing 27,000 labeled audio and video clips that capture different levels of fish feeding intensity. Then, we introduce multi-modal approaches for FFIA by leveraging the models pre-trained on individual modalities and fused with data fusion methods. We perform benchmark studies of these methods on AV-FFIA, and demonstrate the advantages of the multi-modal approach over the single-modality based approach, especially in noisy environments. However, compared to the methods developed for individual modalities, the multimodal approaches may involve higher computational costs due to the need for independent encoders for each modality. To overcome this issue, we further present a novel unified mixed-modality based method for FFIA, termed as U-FFIA. U-FFIA is a single model capable of processing audio, visual, or audio-visual modalities, by leveraging modality dropout during training and knowledge distillation using the models pre-trained with data from single modality. We demonstrate that U-FFIA can achieve performance better than or on par with the state-of-the-art modality-specific FFIA models, with significantly lower computational overhead, enabling robust and efficient FFIA for improved aquaculture management. To encourage further research, we have released the AV-FFIA dataset, the pre-trained model and codes athttps://github.com/FishMaster93/U-FFIA. Note to Practitioners—Feeding is one of the most important costs in aquaculture. However, current feeding machines usually operate with fixed thresholds or human experiences, lacking the ability to automatically adjust to fish feeding intensity. FFIA can evaluate the intensity changes in fish appetite during the feeding process and optimize the control strategies of the feeding machine to avoid inadequate feeding or overfeeding, thereby reducing the feeding cost and improving the well-being of fish in industrial aquaculture. The existing methods have mainly exploited single-modality data, and have a high sensitivity to input noise. Using video and audio offers improved chances to address the challenges brought by various environments. However, compared with processing data from single modalities, using multiple modalities simultaneously often involves increased computational resources, including memory, processing power, and storage. This can impact system performance and scalability. To address these issues, we focus on the efficient unified model, which is capable of processing both multimodal and single-modal input. Our proposed model achieved state-of-the-art (SOTA) performance in FFIA with high computational efficiency. Xubo Liu 0001, Haohe Liu, Zhuangzhuang Du, Tao Chen 0009, Guoping Lian, Lihui Wang 0002, Wenwu Wang 0001 |
IEEE Trans Autom. Sci. Eng. | 2 |
| 2024 | Learning Temporal Resolution in Spectrogram for Audio ClassificationabstractThe audio spectrogram is a time-frequency representation that has been widely used for audio classification. One of the key attributes of the audio spectrogram is the temporal resolution, which depends on the hop size used in the Short-Time Fourier Transform (STFT). Previous works generally assume the hop size should be a constant value (e.g., 10 ms). However, a fixed temporal resolution is not always optimal for different types of sound. The temporal resolution affects not only classification accuracy but also computational cost. This paper proposes a novel method, DiffRes, that enables differentiable temporal resolution modeling for audio classification. Given a spectrogram calculated with a fixed hop size, DiffRes merges non-essential time frames while preserving important frames. DiffRes acts as a "drop-in" module between an audio spectrogram and a classifier and can be jointly optimized with the classification task. We evaluate DiffRes on five audio classification tasks, using mel-spectrograms as the acoustic features, followed by off-the-shelf classifier backbones. Compared with previous methods using the fixed temporal resolution, the DiffRes-based method can achieve the equivalent or better classification accuracy with at least 25% computational cost reduction. We further show that DiffRes can improve classification accuracy by increasing the temporal resolution of input acoustic features, without adding to the computational cost. Haohe Liu, Xubo Liu 0001, Qiuqiang Kong, Wenwu Wang 0001, Mark D. Plumbley |
AAAI | 2 |
| 2024 | CM-PIE: Cross-Modal Perception for Interactive-Enhanced Audio-Visual Video ParsingabstractAudio-visual video parsing is the task of categorizing a video with weak labels at the segment level, and predicting them as audible or visible events. Recent methods have leveraged the attention mechanism to capture the semantic correlations among the whole video across the audio-visual modalities. However, these approaches may overlook the importance of individual segments and their interrelations within a video, typically relying on a single modality when learning features. In this paper, we propose a novel interactive-enhanced cross-modal perception method (CM-PIE), which can learn fine-grained features by applying a segment-based attention module. In addition, a cross-modal aggregation block is introduced to jointly optimize the semantic representation of audio and visual signals by enhancing inter-modal interactions. Experimental results show that our model offers improved parsing performance on the Look, Listen, and Parse (LLP) dataset compared to other methods. Yaru Chen 0003, Ruohao Guo, Xubo Liu 0001, Peipei Wu, Guangyao Li 0001, Wenwu Wang 0001 |
ICASSP | 3 |
| 2024 | Audio Prompt Tuning for Universal Sound SeparationabstractUniversal sound separation (USS) is a task to separate arbitrary sounds from an audio mixture. Existing USS systems are capable of separating arbitrary sources, given a few examples of the target sources as queries. However, separating arbitrary sounds with a single system is challenging, and the robustness is not always guaranteed. In this work, we propose audio prompt tuning (APT), a simple yet effective approach to enhance existing USS systems. Specifically, APT improves the separation performance of specific sources through training a small number of prompt parameters with limited audio samples, while maintaining the generalization of the USS model by keeping its parameters frozen. We evaluate the proposed method on MUSDB18 and ESC-50 datasets. Compared with the baseline model, APT can improve the signal-to-distortion ratio performance by 0.67 dB and 2.06 dB using the full training set of two datasets. Moreover, APT with only 5 audio samples even outperforms the baseline systems utilizing full training data on the ESC-50 dataset, indicating the great potential of few-shot APT. Yuzhuo Liu, Xubo Liu 0001, Yan Zhao 0010, Pingchuan Tain, Yuxuan Wang 0002 |
ICASSP | 2 |
| 2024 | Retrieval-Augmented Text-to-Audio GenerationabstractDespite recent progress in text-to-audio (TTA) generation, we show that the state-of-the-art models, such as AudioLDM, trained on datasets with an imbalanced class distribution, such as AudioCaps, are biased in their generation performance. Specifically, they excel in generating common audio classes while underperforming in the rare ones, thus degrading the overall generation performance. We refer to this problem as long-tailed text-to-audio generation. To address this issue, we propose a simple retrieval-augmented approach for TTA models. Specifically, given an input text prompt, we first leverage a Contrastive Language Audio Pretraining (CLAP) model to retrieve relevant text-audio pairs. The features of the retrieved audio-text data are then used as additional conditions to guide the learning of TTA models. We enhance AudioLDM with our proposed approach and denote the resulting augmented system as Re-AudioLDM. On the AudioCaps dataset, Re-AudioLDM achieves a state-of-the-art Frechet Audio Distance (FAD) of 1.37, outperforming the existing approaches by a large margin. Furthermore, we show that Re-AudioLDM can generate realistic audio for complex scenes, rare audio classes, and even unseen audio types, indicating its potential in TTA tasks. Haohe Liu, Xubo Liu 0001, Qiushi Huang, Mark D. Plumbley, Wenwu Wang 0001 |
ICASSP | 3 |
| 2024 | First-Shot Unsupervised Anomalous Sound Detection with Unknown Anomalies Estimated by Metadata-Assisted Audio GenerationabstractFirst-shot (FS) unsupervised anomalous sound detection (ASD) is a brand-new task introduced in DCASE 2023 Challenge Task 2, where the anomalous sounds for the target machine types are unseen in training. Existing methods often rely on the availability of normal and abnormal sound data from the target machines. However, due to the lack of anomalous sound data for the target machine types, it becomes challenging when adapting the existing ASD methods to the first-shot task. In this paper, we propose a new framework for the first-shot unsupervised ASD, where metadata-assisted audio generation is used to estimate unknown anomalies, by utilising the available machine information (i.e., metadata and sound data) to fine-tune a text-to-audio generation model for generating the anomalous sounds that contain unique acoustic characteristics accounting for each different machine type. We then use the method of Time-Weighted Frequency domain audio Representation with Gaussian Mixture Model (TWFRGMM) as the backbone to achieve the first-shot unsupervised ASD. Our proposed FS-TWFR-GMM method achieves competitive performance amongst top systems in DCASE 2023 Challenge Task 2, while requiring only 1% model parameters for detection, as validated in our experiments. Hejing Zhang, Qiaoxi Zhu, Jian Guan 0001, Haohe Liu, Feiyang Xiao, Jiantong Tian, Xinhao Mei, Xubo Liu 0001, Wenwu Wang 0001 |
ICASSP | 8 |
| 2024 | AudioLDM 2: Learning Holistic Audio Generation With Self-Supervised PretrainingabstractAlthough audio generation shares commonalities across different types of audio, such as speech, music, and sound effects, designing models for each type requires careful consideration of specific objectives and biases that can significantly differ from those of other types. To bring us closer to a unified perspective of audio generation, this paper proposes a holistic framework that utilizes the same learning method for speech, music, and sound effect generation. Our framework utilizes a general representation of audio, called “language of audio” (LOA). Any audio can be translated into LOA based on AudioMAE, a self-supervised pre-trained representation learning model. In the generation process, we translate other modalities into LOA by using a GPT-2 model, and we perform self-supervised audio generation learning with a latent diffusion model conditioned on the LOA of audio in our training set. The proposed framework naturally brings advantages such as reusable self-supervised pretrained latent diffusion models. Experiments on the major benchmarks of text-to-audio, text-to-music, and text-to-speech with three AudioLDM 2 variants demonstrate competitive performance of the AudioLDM 2 variants framework against previous approaches. Our code, pretrained model, and demo are available athttps://audioldm.github.io/audioldm2. Haohe Liu, Xubo Liu 0001, Xinhao Mei, Qiuqiang Kong, Qiao Tian 0001, Yuping Wang 0005, Wenwu Wang 0001, Yuxuan Wang 0002, Mark D. Plumbley |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2024 | Towards Generating Diverse Audio Captions via Adversarial TrainingabstractAutomated audio captioning is a cross-modal translation task for describing the content of audio clips with natural language sentences. This task has attracted increasing attention and substantial progress has been made in recent years. Captions generated by existing models are generally faithful to the content of audio clips, however, these machine-generated captions are often deterministic (e.g., generating a fixed caption for a given audio clip), simple (e.g., using common words and simple grammar), and generic (e.g., generating the same caption for similar audio clips). When people are asked to describe the content of an audio clip, different people tend to focus on different sound events and describe an audio clip diversely from various aspects using distinct words and grammar. We believe that an audio captioning system should have the ability to generate diverse captions, either for a fixed audio clip, or across similar audio clips. To this end, we propose an adversarial training framework based on a conditional generative adversarial network (C-GAN) to improve diversity of audio captioning systems. A caption generator and two hybrid discriminators compete and are learned jointly, where the caption generator can be any standard encoder-decoder captioning model used to generate captions, and the hybrid discriminators assess the generated captions from different criteria, such as their naturalness and semantics. We conduct experiments on the Clotho dataset. The results show that our proposed model can generate captions with better diversity as compared to state-of-the-art methods. Xinhao Mei, Xubo Liu 0001, Jianyuan Sun, Mark D. Plumbley, Wenwu Wang 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2023 | Personalized Dialogue Generation with Persona-Adaptive AttentionabstractPersona-based dialogue systems aim to generate consistent responses based on historical context and predefined persona. Unlike conventional dialogue generation, the persona-based dialogue needs to consider both dialogue context and persona, posing a challenge for coherent training. Specifically, this requires a delicate weight balance between context and persona. To achieve that, in this paper, we propose an effective framework with Persona-Adaptive Attention (PAA), which adaptively integrates the weights from the persona and context information via our designed attention. In addition, a dynamic masking mechanism is applied to the PAA to not only drop redundant information in context and persona but also serve as a regularization mechanism to avoid overfitting. Experimental results demonstrate the superiority of the proposed PAA framework compared to the strong baselines in both automatic and human evaluation. Moreover, the proposed PAA approach can perform equivalently well in a low-resource regime compared to models trained in a full-data setting, which achieve a similar result with only 20% to 30% of data compared to the larger models trained in the full-data setting. To fully exploit the effectiveness of our design, we designed several variants for handling the weighted information in different ways, showing the necessity and sufficiency of our weighting and masking designs. Qiushi Huang, Yu Zhang 0006, Tom Ko, Xubo Liu 0001, Bo Wu 0018, Wenwu Wang 0001, Lilian Tang |
AAAI | 4 |
| 2023 | Learning Retrieval Augmentation for Personalized Dialogue GenerationabstractPersonalized dialogue generation, focusing on generating highly tailored responses by leveraging persona profiles and dialogue context, has gained significant attention in conversational AI applications.However, persona profiles, a prevalent setting in current personalized dialogue datasets, typically composed of merely four to five sentences, may not offer comprehensive descriptions of the persona about the agent, posing a challenge to generate truly personalized dialogues.To handle this problem, we propose Learning Retrieval Augmentation for Personalized DialOgue Generation (LAPDOG), which studies the potential of leveraging external knowledge for persona dialogue generation.Specifically, the proposed LAPDOG model consists of a story retriever and a dialogue generator.The story retriever uses a given persona profile as queries to retrieve relevant information from the story document, which serves as a supplementary context to augment the persona profile.The dialogue generator utilizes both the dialogue history and the augmented persona profile to generate personalized responses.For optimization, we adopt a joint training framework that collaboratively learns the story retriever and dialogue generator, where the story retriever is optimized towards desired ultimate metrics (e.g., BLEU) to retrieve content for the dialogue generator to generate personalized responses.Experiments conducted on the CONVAI2 dataset with ROCStory as a supplementary data source show that the proposed LAPDOG method substantially outperforms the baselines, indicating the effectiveness of the proposed method.The LAPDOG model code is publicly available for further exploration. Qiushi Huang, Xubo Liu 0001, Wenwu Wang 0001, Tom Ko, Yu Zhang 0006, Lilian Tang |
EMNLP | 3 |
| 2023 | Simple Pooling Front-Ends for Efficient Audio ClassificationabstractRecently, there has been increasing interest in building efficient audio neural networks for on-device scenarios. Most existing approaches are designed to reduce the size of audio neural networks using methods such as model pruning. In this work, we show that instead of reducing model size using complex methods, eliminating the temporal redundancy in the input audio features (e.g., mel-spectrogram) could be an effective approach for efficient audio classification. To do so, we proposed a family of simple pooling front-ends (SimPFs) which use simple non-parametric pooling operations to reduce the redundant information within the mel-spectrogram. We perform extensive experiments on four audio classification tasks to evaluate the performance of SimPFs. Experimental results show that SimPFs can achieve a reduction in more than half of the number of floating point operations (FLOPs) for off-the-shelf audio neural networks, with negligible degradation or even some improvements in audio classification performance. Xubo Liu 0001, Haohe Liu, Qiuqiang Kong, Xinhao Mei, Mark D. Plumbley, Wenwu Wang 0001 |
ICASSP | 1 |
| 2023 | AudioLDM: Text-to-Audio Generation with Latent Diffusion ModelsabstractText-to-audio (TTA) systems have recently gained attention for their ability to synthesize general audio based on text descriptions. However, previous studies in TTA have limited generation quality with high computational costs. In this study, we propose AudioLDM, a TTA system that is built on a latent space to learn continuous audio representations from contrastive language-audio pretraining (CLAP) embeddings. The pretrained CLAP models enable us to train LDMs with audio embeddings while providing text embeddings as the condition during sampling. By learning the latent representations of audio signals without modelling the cross-modal relationship, AudioLDM improves both generation quality and computational efficiency. Trained on AudioCaps with a single GPU, AudioLDM achieves state-of-the-art TTA performance compared to other open-sourced systems, measured by both objective and subjective metrics. AudioLDM is also the first TTA system that enables various text-guided audio manipulations (e.g., style transfer) in a zero-shot fashion. Our implementation and demos are available at https://audioldm.github.io. Haohe Liu, Zehua Chen 0005, Xinhao Mei, Xubo Liu 0001, Danilo P. Mandic, Wenwu Wang 0001, Mark D. Plumbley |
ICML | 5 |
| 2023 | Adapting Language-Audio Models as Few-Shot Audio LearnersabstractContrastive language-audio pretraining (CLAP) has become a new paradigm to learn audio concepts with audio-text pairs. CLAP models have shown unprecedented performance as zero-shot classifiers on downstream tasks. To further adapt CLAP with domain-specific knowledge, a popular method is to finetune its audio encoder with available labelled examples. However, this is challenging in low-shot scenarios, as the amount of annotations is limited compared to the model size. In this work, we introduce a Training-efficient (Treff) adapter to rapidly learn with a small set of examples while maintaining the capacity for zero-shot classification. First, we propose a cross-attention linear model (CALM) to map a set of labelled examples and test audio to test labels. Second, we find initialising CALM as a cosine measurement improves our Treff adapter even without training. The Treff adapter outperforms metric-based methods in few-shot settings and yields competitive results to fully-supervised methods. Jinhua Liang, Xubo Liu 0001, Haohe Liu, Huy Phan, Emmanouil Benetos, Mark D. Plumbley, Wenwu Wang 0001 |
INTERSPEECH | 2 |
| 2023 | Visually-Aware Audio Captioning With Adaptive Audio-Visual AttentionabstractAudio captioning aims to generate text descriptions of audio clips.In the real world, many objects produce similar sounds.How to accurately recognize ambiguous sounds is a major challenge for audio captioning.In this work, inspired by inherent human multimodal perception, we propose visuallyaware audio captioning, which makes use of visual information to help the description of ambiguous sounding objects.Specifically, we introduce an off-the-shelf visual encoder to extract video features and incorporate the visual features into an audio captioning system.Furthermore, to better exploit complementary audio-visual contexts, we propose an audio-visual attention mechanism that adaptively integrates audio and visual context and removes the redundant information in the latent space.Experimental results on AudioCaps, the largest audio captioning dataset, show that our proposed method achieves state-of-theart results on machine translation metrics. Xubo Liu 0001, Qiushi Huang, Xinhao Mei, Haohe Liu, Qiuqiang Kong, Jianyuan Sun, Shengchen Li, Tom Ko, Yu Zhang 0006, Lilian Tang, Mark D. Plumbley, Volkan Kilic, Wenwu Wang 0001 |
INTERSPEECH | 1 |
| 2023 | Ontology-aware Learning and Evaluation for Audio TaggingabstractThis study defines a new evaluation metric for audio tagging tasks to alleviate the limitation of the mean average precision (mAP) metric.The mAP metric treats different kinds of sound as independent classes without considering their relations.The proposed metric, ontology-aware mean average precision (OmAP), addresses the weaknesses of mAP by utilizing additional ontology during evaluation.Specifically, we reweight the false positive events in the model prediction based on the AudioSet ontology graph distance to the target classes.The OmAP also provides insights into model performance by evaluating different coarse-grained levels in the ontology graph.We conduct a human assessment and show that OmAP is more consistent with human perception than mAP.We also propose an ontology-based loss function (OBCE) that reweights binary cross entropy (BCE) loss based on the ontology distance.Our experiment shows that OBCE can improve both mAP and OmAP metrics on the AudioSet tagging task. Haohe Liu, Qiuqiang Kong, Xubo Liu 0001, Xinhao Mei, Wenwu Wang 0001, Mark D. Plumbley |
INTERSPEECH | 3 |
| 2023 | Dual Transformer Decoder based Features Fusion Network for Automated Audio CaptioningabstractAutomated audio captioning (AAC) which generates textual descriptions of audio content.Existing AAC models achieve good results but only use the high-dimensional representation of the encoder.There is always insufficient information learning of high-dimensional methods owing to high-dimensional representations having a large amount of information.In this paper, a new encoder-decoder model called the Lowand High-Dimensional Feature Fusion (LHDFF) is proposed.LHDFF uses a new PANNs encoder called Residual PANNs (RPANNs) to fuse low-and high-dimensional features.Lowdimensional features contain limited information about specific audio scenes.The fusion of low-and high-dimensional features can improve model performance by repeatedly emphasizing specific audio scene information.To fully exploit the fused features, LHDFF uses a dual transformer decoder structure to generate captions in parallel.Experimental results show that LHDFF outperforms existing audio captioning models. Jianyuan Sun, Xubo Liu 0001, Xinhao Mei, Volkan Kilic, Mark D. Plumbley, Wenwu Wang 0001 |
INTERSPEECH | 2 |
| 2022 | Diverse Audio Captioning Via Adversarial TrainingabstractAudio captioning aims at generating natural language descriptions for audio clips automatically. Existing audio captioning models have shown promising improvement in recent years. However, these models are mostly trained via maximum likelihood estimation (MLE), which tends to make captions generic, simple and deterministic. As different people may describe an audio clip from different aspects using distinct words and grammars, we argue that an audio captioning system should have the ability to generate diverse captions for a fixed audio clip and across similar audio clips. To address this problem, we propose an adversarial training framework for audio captioning based on a conditional generative adversarial network (C-GAN), which aims at improving the naturalness and diversity of generated captions. Unlike processing data of continuous values in a classical GAN, a sentence is composed of discrete tokens and the discrete sampling process is non-differentiable. To address this issue, policy gradient, a reinforcement learning technique, is used to back-propagate the reward to the generator. The results show that our proposed model can generate more diverse captions, as compared to state-of-the-art methods. Xinhao Mei, Xubo Liu 0001, Jianyuan Sun, Mark D. Plumbley, Wenwu Wang 0001 |
ICASSP | 2 |
| 2022 | Audio-Visual Tracking of Multiple Speakers Via a PMBM FilterabstractAudio-visual tracking of multiple speakers requires to estimate the state (e.g. velocity and location) of each speaker by leveraging the information of both audio and visual modalities. Estimating the number of speakers and their states jointly remains a challenging problem. We propose an Audio-Visual Possion Multi-Bernoulli Mixture Filter (AV-PMBM) that can not only predict the number of speakers but also give accurate estimation of their states. We also propose a novel sound source localization technique based on DOA information and a deep learning based object detector to provide reliable audio measurements for the AV tracker. To our knowledge, this represents the first attempt using PMBM for multi-speaker tracking with audio visual modalities. Experiments on the AV16.3 dataset demonstrate that AV-PMBM achieves state-of-the-art performance in optimal sub-pattern assignment (OSPA). Jinzheng Zhao, Peipei Wu, Xubo Liu 0001, Yong Xu 0004, Lyudmila Mihaylova, Simon J. Godsill, Wenwu Wang 0001 |
ICASSP | 3 |
| 2022 | Neural Vocoder is All You Need for Speech Super-resolutionabstractSpeech super-resolution (SR) is a task to increase speech sampling rate by generating high-frequency components. Existing speech SR methods are trained in constrained experimental settings, such as a fixed upsampling ratio. These strong constraints can potentially lead to poor generalization ability in mismatched real-world cases. In this paper, we propose a neural vocoder based speech super-resolution method (NVSR) that can handle a variety of input resolution and upsampling ratios. NVSR consists of a mel-bandwidth extension module, a neural vocoder module, and a post-processing module. Our proposed system achieves state-of-the-art results on the VCTK multi-speaker benchmark. On 44.1 kHz target resolution, NVSR outperforms WSRGlow and Nu-wave by 8% and 37% respectively on log spectral distance and achieves a significantly better perceptual quality. We also demonstrate that prior knowledge in the pre-trained vocoder is crucial for speech SR by performing mel-bandwidth extension with a simple replication-padding method. Samples can be found in https://haoheliu.github.io/nvsr. Haohe Liu, Woosung Choi, Xubo Liu 0001, Qiuqiang Kong, Qiao Tian 0001, DeLiang Wang |
INTERSPEECH | 3 |
| 2022 | Separate What You Describe: Language-Queried Audio Source SeparationabstractIn this paper, we introduce the task of language-queried audio source separation (LASS), which aims to separate a target source from an audio mixture based on a natural language query of the target source (e.g., "a man tells a joke followed by people laughing"). A unique challenge in LASS is associated with the complexity of natural language description and its relation with the audio sources. To address this issue, we proposed LASS-Net, an end-to-end neural network that is learned to jointly process acoustic and linguistic information, and separate the target source that is consistent with the language query from an audio mixture. We evaluate the performance of our proposed system with a dataset created from the AudioCaps dataset. Experimental results show that LASS-Net achieves considerable improvements over baseline methods. Furthermore, we observe that LASS-Net achieves promising generalization results when using diverse human-annotated descriptions as queries, indicating its potential use in real-world scenarios. The separated audio samples and source code are available at https://liuxubo717.github.io/LASS-demopage. Xubo Liu 0001, Haohe Liu, Qiuqiang Kong, Xinhao Mei, Jinzheng Zhao, Qiushi Huang, Mark D. Plumbley, Wenwu Wang 0001 |
INTERSPEECH | 1 |
| 2022 | VoiceFixer: A Unified Framework for High-Fidelity Speech RestorationabstractSpeech restoration aims to remove distortions in speech signals. Prior methods mainly focus on a single type of distortion, such as speech denoising or dereverberation. However, speech signals can be degraded by several different distortions simultaneously in the real world. It is thus important to extend speech restoration models to deal with multiple distortions. In this paper, we introduce VoiceFixer, a unified framework for high-fidelity speech restoration. VoiceFixer restores speech from multiple distortions (e.g., noise, reverberation, and clipping) and can expand degraded speech (e.g., noisy speech) with a low bandwidth to 44.1 kHz full-bandwidth high-fidelity speech. We design VoiceFixer based on (1) an analysis stage that predicts intermediate-level features from the degraded speech, and (2) a synthesis stage that generates waveform using a neural vocoder. Both objective and subjective evaluations show that VoiceFixer is effective on severely degraded speech, such as real-world historical speech recordings. Samples of VoiceFixer are available at https://haoheliu.github.io/voicefixer. Haohe Liu, Xubo Liu 0001, Qiuqiang Kong, Qiao Tian 0001, Yan Zhao 0010, DeLiang Wang, Chuanzeng Huang, Yuxuan Wang 0002 |
INTERSPEECH | 2 |
| 2022 | On Metric Learning for Audio-Text Cross-Modal RetrievalabstractAudio-text retrieval aims at retrieving a target audio clip or caption from a pool of candidates given a query in another modality. Solving such cross-modal retrieval task is challenging because it not only requires learning robust feature representations for both modalities, but also requires capturing the fine-grained alignment between these two modalities. Existing cross-modal retrieval models are mostly optimized by metric learning objectives as both of them attempt to map data to an embedding space, where similar data are close together and dissimilar data are far apart. Unlike other cross-modal retrieval tasks such as image-text and video-text retrievals, audio-text retrieval is still an unexplored task. In this work, we aim to study the impact of different metric learning objectives on the audio-text retrieval task. We present an extensive evaluation of popular metric learning objectives on the AudioCaps and Clotho datasets. We demonstrate that NT-Xent loss adapted from self-supervised learning shows stable performance across different datasets and training settings, and outperforms the popular triplet-based losses. Our code is available at https://github.com/XinhaoMei/audio-text_retrieval. Xinhao Mei, Xubo Liu 0001, Jianyuan Sun, Mark D. Plumbley, Wenwu Wang 0001 |
INTERSPEECH | 2 |
| 2022 | Audio Visual Multi-Speaker Tracking with Improved GCF and PMBM Filter
Jinzheng Zhao, Peipei Wu, Xubo Liu 0001, Shidrokh Goudarzi, Haohe Liu, Yong Xu 0004, Wenwu Wang 0001 |
INTERSPEECH | 3 |
| 2021 | Token-Level Supervised Contrastive Learning for Punctuation RestorationabstractPunctuation is critical in understanding natural language text. Currently, most automatic speech recognition (ASR) systems do not generate punctuation, which affects the performance of downstream tasks, such as intent detection and slot filling. This gives rise to the need for punctuation restoration. Recent work in punctuation restoration heavily utilizes pre-trained language models without considering data imbalance when predicting punctuation classes. In this work, we address this problem by proposing a token-level supervised contrastive learning method that aims at maximizing the distance of representation of different punctuation marks in the embedding space. The result shows that training with token-level supervised contrastive learning obtains up to 3.2% absolute F1 improvement on the test set. Qiushi Huang, Tom Ko, Lilian Tang, Xubo Liu 0001, Bo Wu 0018 |
Interspeech | 4 |