VLDB 2026 Research / reviewers in the wild / expert
Helin Wang
dblp:180/6754
· DBLP profile ↗
32ranked-venue papers
13as first author
29since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 22 · 10 first-author · 20 since 2021Artificial intelligence and machine learning · 21 · 8 first-author · 20 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Explainable AI in Medicine: A Comprehensive Narrative Review of Methods, Applications, and Future DirectionsabstractABSTRACT Artificial Intelligence (AI) is increasingly utilised in medicine; however, its “black‐box” nature continues to hinder clinical trust, adoption, and validation. Explainable AI (XAI) has emerged as a critical field to address these transparency challenges by making AI‐driven decisions more interpretable and actionable. This narrative review examines the progress of XAI in medicine over the past decade. We first introduce fundamental XAI concepts and describe our review methodology, followed by a comprehensive analysis of key application domains, including medical imaging, electronic health records (EHRs), and multi‐omics data. Methodologically, we categorise XAI techniques into model‐agnostic approaches (e.g., SHAP, LIME, Anchors) and model‐specific approaches (e.g., Grad‐CAM, LRP, TreeSHAP). Beyond summarising their principles, advantages, and limitations, we further provide a systematic analysis of clinical reliability and failure modes associated with each class of methods, highlighting how explanation techniques may produce misleading, unstable, or non‐causal interpretations in real‐world clinical settings. The review then discusses the demonstrated benefits of XAI, including result validation, bias detection, and improved patient–clinician communication, while critically examining persistent challenges such as limited clinical deployment, inconsistent evaluation standards, and the lack of prospective validation. Finally, we outline future research directions, emphasising the need to adapt XAI to large‐scale foundation models and conversational AI systems, as well as to extend its applicability in biomedical and multi‐omics interpretation. We argue that while XAI is essential for improving transparency, trust, and clinical adoption, its reliable and scalable integration into clinical workflows requires both methodological innovation and rigorous, clinically grounded validation frameworks. Helin Wang, Xueyu Liu, Jiashuo Shi, Yu-ang Li, Guanghui Yue 0001, Yongfei Wu |
Expert Syst. J. Knowl. Eng. | 1 |
| 2025 | AdaHet-MKD: An Adaptive Heterogeneous Multi-teacher Knowledge Distillation for Medical Image AnalysisabstractContrastive Language-Image Pre-training (CLIP) has emerged as an effective framework for multi-modal representation learning, achieving notable success in diverse tasks such as medical image analysis. CLIP's growing prominence in medical image applications is restricted by its significant computational demands, creating implementation challenges in resource-constrained clinical environments. While knowledge distillation offers an effective approach for model compression with preserved accuracy, existing methods suffer from two fundamental limitations. Firstly, existing methods focus on learning better information from single models while ignoring the fact that student models can generalize well under the guidance of multiple teachers. Secondly, they overlook the complementary information in the CLIP model where the text encoder and image encoder can be leveraged as heterogeneous information to teach one single modality. To tackle these challenges, we propose an Adaptive Heterogeneous Multi-teacher Knowledge Distillation (AdaHet-MKD) framework for effective knowledge transfer across heterogeneous text-image models and among multiple teacher models. The key innovations include: (i) adaptively determining the contribution of each teacher model to specific instances, thereby generating integrated soft logits, and (ii) enabling the student model to operate independently of the teacher model's architecture, which enhances flexibility in teacher-student pairings. Experimental evaluations on publicly available medical datasets demonstrate that our approach has achieved the state-of-the-art performance compared to baselines. Helin Wang, Wei Du 0010, Ning Liu 0014, Qian Li 0043, Yanyu Xu 0001, Li-Zhen Cui 0001 |
CIKM | 1 |
| 2025 | SoloAudio: Target Sound Extraction with Language-oriented Audio Diffusion TransformerabstractIn this paper, we introduce SoloAudio, a novel diffusion-based generative model for target sound extraction (TSE). Our approach trains latent diffusion models on audio, replacing the previous U-Net backbone with a skip-connected Transformer that operates on latent features. SoloAudio supports both audio-oriented and language-oriented TSE by utilizing a CLAP model as the feature extractor for target sounds. Furthermore, SoloAudio leverages synthetic audio generated by state-of-the-art text-to-audio models for training, demonstrating strong generalization to out-of-domain data and unseen sound events. We evaluate this approach on the FSD Kaggle 2018 mixture dataset and real data from AudioSet, where SoloAudio achieves the state-of-the-art results on both in-domain and out-of-domain data, and exhibits impressive zero-shot and few-shot capabilities. Source code1and demos2are released. Helin Wang, Jiarui Hai, Yen-Ju Lu, Karan Thakkar, Mounya Elhilali, Najim Dehak |
ICASSP | 1 |
| 2025 | SSR-Speech: Towards Stable, Safe and Robust Zero-shot Text-based Speech Editing and SynthesisabstractIn this paper, we introduce SSR-Speech, a neural codec autoregressive model designed for stable, safe, and robust zero-shot text-based speech editing and text-to-speech synthesis. SSR-Speech is built on a Transformer decoder and incorporates classifier-free guidance to enhance the stability of the generation process. A watermark Encodec is proposed to embed frame-level watermarks into the edited regions of the speech so that which parts were edited can be detected. In addition, the waveform reconstruction leverages the original unedited speech segments, providing superior recovery compared to the Encodec model. Our approach achieves state-of-the-art performance in the RealEdit speech editing task and the LibriTTS text-to-speech task, surpassing previous methods. Furthermore, SSR-Speech excels in multi-span speech editing and also demonstrates remarkable robustness to background sounds. The source code1and demos2are released. Helin Wang, Meng Yu 0003, Jiarui Hai, Chen Chen 0075, Rilin Chen, Najim Dehak, Dong Yu 0001 |
ICASSP | 1 |
| 2025 | DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMsabstractIn video Multimodal Large Language Models (video MLLMs), the visual encapsulation process plays a pivotal role in converting video contents into representative tokens for LLM input. While linear projectors are widely employed for encapsulation, they introduce semantic indistinctness and temporal incoherence when applied to videos. Conversely, the structure of resamplers shows promise in tackling these challenges, but an effective solution remains unexplored. Drawing inspiration from resampler structures, we introduce DisCo, a novel visual encapsulation method designed to yield semantically distinct and temporally coherent visual tokens for video MLLMs. DisCo integrates two key components: (1) A Visual Concept Discriminator (VCD) module, assigning unique semantics for visual tokens by associating them in pair with discriminative concepts in the video. (2) A Temporal Focus Calibrator (TFC) module, ensuring consistent temporal focus of visual tokens to video elements across every video frame. Through extensive experiments on multiple video MLLM frameworks, we demonstrate that DisCo remarkably outperforms previous state-of-the-art methods across a variety of video understanding benchmarks, while also achieving higher token efficiency thanks to the reduction of semantic indistinctness. The code: https://github.com/ZJHTerry18/DisCo. Jiahe Zhao, Rongkun Zheng, Yi Wang 0074, Helin Wang, Hengshuang Zhao |
ICCV | 4 |
| 2025 | Audio Large Language Models Can Be Descriptive Speech Quality EvaluatorsabstractAn ideal multimodal agent should be aware of the quality of its input modalities. Recent advances have enabled large language models (LLMs) to incorporate auditory systems for handling various speech-related tasks. However, most audio LLMs remain unaware of the quality of the speech they process. This limitation arises because speech quality evaluation is typically excluded from multi-task training due to the lack of suitable datasets. To address this, we introduce the first natural language-based speech evaluation corpus, generated from authentic human ratings. In addition to the overall Mean Opinion Score (MOS), this corpus offers detailed analysis across multiple dimensions and identifies causes of quality degradation. It also enables descriptive comparisons between two speech samples (A/B tests) with human-like judgment. Leveraging this corpus, we propose an alignment approach with LLM distillation (ALLD) to guide the audio LLM in extracting relevant information from raw speech and generating meaningful responses. Experimental results demonstrate that ALLD outperforms the previous state-of-the-art regression model in MOS prediction, with a mean square error of 0.17 and an A/B test accuracy of 98.6%. Additionally, the generated responses achieve BLEU scores of 25.8 and 30.2 on two tasks, surpassing the capabilities of task-specific models. This work advances the comprehensive perception of speech signals by audio LLMs, contributing to the development of real-world auditory and sensory intelligent agents. Chen Chen 0075, Siyin Wang, Helin Wang, Zhehuai Chen, Chao Zhang 0031, Chao-Han Huck Yang, Chng Eng Siong |
ICLR | 4 |
| 2025 | ALMTokenizer: A Low-bitrate and Semantic-rich Audio Codec Tokenizer for Audio Language ModelingabstractRecent advancements in audio language models have underscored the pivotal role of audio tokenization, which converts audio signals into discrete tokens, thereby facilitating the application of language model architectures to the audio domain. In this study, we introduce ALMTokenizer, a novel low-bitrate and semantically rich audio codec tokenizer for audio language models. Prior methods, such as Encodec, typically encode individual audio frames into discrete tokens without considering the use of context information across frames. Unlike these methods, we introduce a novel query-based compression strategy to capture holistic information with a set of learnable query tokens by explicitly modeling the context information across frames. This design not only enables the codec model to capture more semantic information but also encodes the audio signal with fewer token sequences. Additionally, to enhance the semantic information in audio codec models, we introduce the following: (1) A masked autoencoder (MAE) loss, (2) Vector quantization based on semantic priors, and (3) An autoregressive (AR) prediction loss. As a result, ALMTokenizer achieves competitive reconstruction performance relative to state-of-the-art approaches while operating at a lower bitrate. Within the same audio language model framework, ALMTokenizer outperforms previous tokenizers in audio understanding and generation tasks.[https://dongchaoyang.top/ALMTokenizer/] Dongchao Yang, Songxiang Liu, Haohan Guo, Jiankun Zhao, Helin Wang, Zeqian Ju, Xueyuan Chen, Xu Tan 0003, Xixin Wu, Helen M. Meng |
ICML | 6 |
| 2025 | EzAudio: Enhancing Text-to-Audio Generation with Efficient Diffusion Transformer
Jiarui Hai, Yong Xu 0004, Hao Zhang 0112, Chenxing Li, Helin Wang, Mounya Elhilali, Dong Yu 0001 |
INTERSPEECH | 5 |
| 2024 | Finding Spoken Identifications: Using GPT-4 Annotation for an Efficient and Fast Dataset Creation PipelineabstractThe growing emphasis on fairness in speech-processing tasks requires datasets with speakers from diverse subgroups that allow training and evaluating fair speech technology systems. However, creating such datasets through manual annotation can be costly. To address this challenge, we present a semi-automated dataset creation pipeline that leverages large language models. We use this pipeline to generate a dataset of speakers identifying themself or another speaker as belonging to a particular race, ethnicity, or national origin group. We use OpenaAI’s GPT-4 to perform two complex annotation tasks- separating files relevant to our intended dataset from the irrelevant ones (filtering) and finding and extracting information on identifications within a transcript (tagging). By evaluating GPT-4’s performance using human annotations as ground truths, we show that it can reduce resources required by dataset annotation while barely losing any important information. For the filtering task, GPT-4 had a very low miss rate of 6.93%. GPT-4’s tagging performance showed a trade-off between precision and recall, where the latter got as high as 97%, but precision never exceeded 45%. Our approach reduces the time required for the filtering and tagging tasks by 95% and 80%, respectively. We also present an in-depth error analysis of GPT-4’s performance. Maliha Jahan, Helin Wang, Thomas Thebaud, Yinglun Sun, Giang Ha Le, Zsuzsanna Fagyal, Odette Scharenborg, Mark Hasegawa-Johnson, Laureano Moro-Velázquez, Najim Dehak |
LREC/COLING | 2 |
| 2024 | DPM-TSE: A Diffusion Probabilistic Model for Target Sound ExtractionabstractCommon target sound extraction (TSE) approaches primarily relied on discriminative approaches in order to separate the target sound while minimizing interference from the unwanted sources, with varying success in separating the target from the background. This study introduces DPM-TSE, a generative method based on diffusion probabilistic modeling (DPM) for Target Sound Extraction (TSE), to achieve both cleaner target renderings as well as improved separability from unwanted sounds. The technique also tackles the noise floor of DPM by introducing a correction method for noise schedules and sample steps. This approach is evaluated using both objective and subjective quality metrics on the FSD Kaggle 2018 dataset. The results show that DPM-TSE has a significant improvement in perceived quality in terms of target extraction and purity. Jiarui Hai, Helin Wang, Dongchao Yang, Karan Thakkar, Najim Dehak, Mounya Elhilali |
ICASSP | 2 |
| 2024 | Efficient Reinforcement Learning via Decoupling Exploration and Utilization
Jingpu Yang, Helin Wang, Qirui Zhao, Zhecheng Shi, Zirui Song, Miao Fang 0001 |
ICIC (2) | 2 |
| 2024 | A Paradigm for Generating Operational Seamless Land Surface Temperature ProductsabstractLand surface temperature (LST) is a direct result of earth-atmosphere interactions and has been widely used in earth system science and climate change. Thus, it is recognized as one of the essential climate variables (ECVs). Thermal infrared (TIR) remote sensing is the most effective way to obtain high-quality LST at a large scale. However, TIR cannot penetrate the Cloud and obtain the cloudy sky LST, which seriously hinders its applications. To address this challenge, we proposed a paradigm for estimating the seamless LST at regional and global scales. Validation results showed the root mean square error (RMSE) of the produced seamless LST achieved approximately 3K. The produced seamless LST at China landmass, East Asia, and global have been freely released to the public (https://elite.bnu.edu.cn). Jie Cheng 0001, Shugui Zhou, Xiangchen Meng, Shengyue Dong, Aixia Yang, Qi Zeng 0005, Manqing Liu, Mengfei Guo, Chenze Wu, Helin Wang |
IGARSS | 14 |
| 2024 | DreamVoice: Text-Guided Voice Conversion
Jiarui Hai, Karan Thakkar, Helin Wang, Zengyi Qin, Mounya Elhilali |
INTERSPEECH | 3 |
| 2024 | Noise-robust Speech Separation with Fast Generative Correction
Helin Wang, Jesús Villalba 0001, Laureano Moro-Velázquez, Jiarui Hai, Thomas Thebaud, Najim Dehak |
INTERSPEECH | 1 |
| 2023 | Masked Spectrogram Prediction for Self-Supervised Audio Pre-TrainingabstractTransformer-based models attain excellent results and generalize well when trained on sufficient amounts of data. However, constrained by the limited data available in the audio domain, most transformer-based models for audio tasks are finetuned from pre-trained models in other domains (e.g. image), which has a notable gap with the audio domain. Other methods explore the self-supervised learning approaches directly in the audio domain but currently do not perform well in the downstream tasks. In this paper, we present a novel self-supervised learning method for transformer-based audio models, called masked spectrogram prediction (MaskSpec), to learn powerful audio representations from unlabeled audio data (AudioSet used in this paper). Our method masks random patches of the input spectrogram and reconstructs the masked regions with an encoder-decoder architecture. Experimental results demonstrate MaskSpec reaches the performance of 0.471 (mAP) on AudioSet, 0.854 (mAP) on Open-MIC2018, 0.982 (accuracy) on ESC-50, 0.976 (accuracy) on SCV2, and 0.823 (accuracy) on DCASE2019 Task1A. The source code and pre-trained models have been released.1 Dading Chong, Helin Wang, Peilin Zhou, Qingcheng Zeng |
ICASSP | 2 |
| 2023 | DuTa-VC: A Duration-aware Typical-to-atypical Voice Conversion Approach with Diffusion Probabilistic Model
Helin Wang, Thomas Thebaud, Jesús Villalba 0001, Myra Sydnor, Becky Lammers, Najim Dehak, Laureano Moro-Velázquez |
INTERSPEECH | 1 |
| 2023 | NoreSpeech: Knowledge Distillation based Conditional Diffusion Model for Noise-robust Expressive TTS
Dongchao Yang, Songxiang Liu, Helin Wang, Chao Weng, Yuexian Zou |
INTERSPEECH | 3 |
| 2023 | Benchmarking Large Language Models on CMExam - A comprehensive Chinese Medical Exam DatasetabstractRecent advancements in large language models (LLMs) have transformed the field of question answering (QA). However, evaluating LLMs in the medical field is challenging due to the lack of standardized and comprehensive datasets. To address this gap, we introduce CMExam, sourced from the Chinese National Medical Licensing Examination. CMExam consists of 60K+ multiple-choice questions for standardized and objective evaluations, as well as solution explanations for model reasoning evaluation in an open-ended manner. For in-depth analyses of LLMs, we invited medical professionals to label five additional question-wise annotations, including disease groups, clinical departments, medical disciplines, areas of competency, and question difficulty levels. Alongside the dataset, we further conducted thorough experiments with representative LLMs and QA algorithms on CMExam. The results show that GPT-4 had the best accuracy of 61.6% and a weighted F1 score of 0.617. These results highlight a great disparity when compared to human accuracy, which stood at 71.6%. For explanation tasks, while LLMs could generate relevant reasoning and demonstrate improved performance after finetuning, they fall short of a desired standard, indicating ample room for improvement. To the best of our knowledge, CMExam is the first Chinese medical exam dataset to provide comprehensive medical annotations. The experiments and findings of LLM evaluation also provide valuable insights into the challenges and potential solutions in developing Chinese medical QA systems and LLM evaluation pipelines. Junling Liu, Peilin Zhou, Yining Hua, Dading Chong, Zhongyu Tian, Helin Wang, Chenyu You, Zhenhua Guo 0001, Lei Zhu 0017, Michael Lingzhi Li |
NeurIPS | 7 |
| 2023 | Diffsound: Discrete Diffusion Model for Text-to-Sound GenerationabstractGenerating sound effects that people want is an important topic. However, there are limited studies in this area for sound generation. In this study, we investigate generating sound conditioned on a text prompt and propose a novel text-to-sound generation framework that consists of a text encoder, a Vector Quantized Variational Autoencoder (VQ-VAE), a token-decoder, and a vocoder. The framework first uses the token-decoder to transfer the text features extracted from the text encoder to a mel-spectrogram with the help of VQ-VAE, and then the vocoder is used to transform the generated mel-spectrogram into a waveform. We found that the token-decoder significantly influences the generation performance. Thus, we focus on designing a good token-decoder in this study. We begin with th21e traditional autoregressive (AR) token-decoder, which has shown state-of-the-art performance in previous sound generation works. However, the AR token-decoder always predicts the mel-spectrogram tokens one by one in order, which may introduce the unidirectional bias and accumulation of errors problems. Moreover, with the AR token-decoder, the sound generation time increases linearly with the sound duration. To overcome the shortcomings introduced by AR token-decoders, we propose a non-autoregressive token-decoder based on the discrete diffusion model, named Diffsound. Specifically, the Diffsound model predicts all of the mel-spectrogram tokens in one step and then refines the predicted tokens in the next step, so the best-predicted results can be obtained by iteration. Our experiments show that our proposed Diffsound model not only produces better text-to-sound generation results when compared with the AR token-decoder but also has a faster generation speed,i.e., MOS: 3.56v.s2.786, and the generation speed is five times faster than the AR decoder. Furthermore, to automatically assess the quality of generated samples, we define three different objective evaluation metricsi.e., Fréchet Inception Distance (FID), Kullback-Leibler (KL), and audio caption loss, which can comprehensively assess the relevance and fidelity of the generated samples. Dongchao Yang, Jianwei Yu 0001, Helin Wang, Chao Weng, Yuexian Zou, Dong Yu 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2022 | A Mutual Learning Framework for Few-Shot Sound Event DetectionabstractAlthough prototypical network (ProtoNet) has proved to be an effective method for few-shot sound event detection, two problems still exist. Firstly, the small-scaled support set is insufficient so that the class prototypes may not represent the class center accurately. Secondly, the feature extractor is task-agnostic (or class-agnostic): the feature extractor is trained with base-class data and directly applied to unseen-class data. To address these issues, we present a novel mutual learning framework with transductive learning, which aims at iteratively updating the class prototypes and feature extractor. More specifically, we propose to update class prototypes with transductive inference to make the class prototypes as close to the true class center as possible. To make the feature extractor to be task-specific, we propose to use the updated class prototypes to fine-tune the feature extractor. After that, a fine-tuned feature extractor further helps produce better class prototypes. Our method achieves the F-score of 38.4% on the DCASE 2021 Task 5 evaluation set, which won the first place in the few-shot bioacoustic event detection task of Detection and Classification of Acoustic Scenes and Events (DCASE) 2021 Challenge. Dongchao Yang, Helin Wang, Yuexian Zou, Zhongjie Ye, Wenwu Wang 0001 |
ICASSP | 2 |
| 2022 | Improving Target Sound Extraction with Timestamp InformationabstractTarget sound extraction (TSE) aims to extract the sound part of a target sound event class from a mixture audio with multiple sound events.The previous works mainly focus on the problems of weakly-labelled data, jointly learning and new classes, however, no one cares about the onset and offset times of the target sound event, which has been emphasized in the auditory scene analysis.In this paper, we study to utilize such timestamp information to help extract the target sound via a target sound detection network and a target-weighted time-frequency loss function.More specifically, we use the detection result of a target sound detection (TSD) network as the additional information to guide the learning of target sound extraction network.We also find that the result of TSE can further improve the performance of the TSD network, so that a mutual learning framework of the target sound detection and extraction is proposed.In addition, a target-weighted time-frequency loss function is designed to pay more attention to the temporal regions of the target sound during training.Experimental results on the synthesized data generated from the Freesound Datasets show that our proposed method can significantly improve the performance of TSE. Helin Wang, Dongchao Yang, Chao Weng, Yuexian Zou |
INTERSPEECH | 1 |
| 2022 | RaDur: A Reference-aware and Duration-robust Network for Target Sound DetectionabstractTarget sound detection (TSD) aims to detect the target sound from a mixture audio given the reference information.Previous methods use a conditional network to extract a sounddiscriminative embedding from the reference audio, and then use it to detect the target sound from the mixture audio.However, the network performs much differently when using different reference audios (e.g.performs poorly for noisy and shortduration reference audios), and tends to make wrong decisions for transient events (i.e.shorter than 1 second).To overcome these problems, in this paper, we present a reference-aware and duration-robust network (RaDur) for TSD.More specifically, in order to make the network more aware of the reference information, we propose an embedding enhancement module to take into account the mixture audio while generating the embedding, and apply the attention pooling to enhance the features of target sound-related frames and weaken the features of noisy frames.In addition, a duration-robust focal loss is proposed to help model different-duration events.To evaluate our method, we build two TSD datasets based on UrbanSound and Audioset.Extensive experiments show the effectiveness of our methods. Dongchao Yang, Helin Wang, Zhongjie Ye, Yuexian Zou, Wenwu Wang 0001 |
INTERSPEECH | 2 |
| 2022 | Calibrate and Refine! A Novel and Agile Framework for ASR Error Robust Intent DetectionabstractThe past ten years have witnessed the rapid development of textbased intent detection, whose benchmark performances have already been taken to a remarkable level by deep learning techniques.However, automatic speech recognition (ASR) errors are inevitable in real-world applications due to the environment noise, unique speech patterns and etc, leading to sharp performance drop in state-of-the-art text-based intent detection models.Essentially, this phenomenon is caused by the semantic drift brought by ASR errors and most existing works tend to focus on designing new model structures to reduce its impact, which is at the expense of versatility and flexibility.Different from previous one-piece model, in this paper, we propose a novel and agile framework called CR-ID for ASR error robust intent detection with two plug-and-play modules, namely semantic drift calibration module (SDCM) and phonemic refinement module (PRM), which are both model-agnostic and thus could be easily integrated to any existing intent detection models without modifying their structures.Experimental results on SNIPS dataset show that, our proposed CR-ID framework achieves competitive performance and outperform all the baseline methods on ASR outputs, which verifies that CR-ID can effectively alleviate the semantic drift caused by ASR errors. Peilin Zhou, Dading Chong, Helin Wang, Qingcheng Zeng |
INTERSPEECH | 3 |
| 2021 | Audio-Oriented Multimodal Machine Comprehension via Dynamic Inter- and Intra-modality AttentionabstractWhile Machine Comprehension (MC) has attracted extensive research interests in recent years, existing approaches mainly belong to the category of Machine Reading Comprehension task which mines textual inputs (paragraphs and questions) to predict the answers (choices or text spans). However, there are a lot of MC tasks that accept audio input in addition to the textual input, e.g. English listening comprehension test. In this paper, we target the problem of Audio-Oriented Multimodal Machine Comprehension, and its goal is to answer questions based on the given audio and textual information. To solve this problem, we propose a Dynamic Inter- and Intra-modality Attention (DIIA) model to effectively fuse the two modalities (audio and textual). DIIA can work as an independent component and thus be easily integrated into existing MC models. Moreover, we further develop a Multimodal Knowledge Distillation (MKD) module to enable our multimodal MC model to accurately predict the answers based only on either the text or the audio. As a result, the proposed approach can handle various tasks including: Audio-Oriented Multimodal Machine Comprehension, Machine Reading Comprehension and Machine Listening Comprehension, in a single model, making fair comparisons possible between our model and the existing unimodal MC models. Experimental results and analysis prove the effectiveness of the proposed approaches. First, the proposed DIIA boosts the baseline models by up to 21.08% in terms of accuracy; Second, under the unimodal scenarios, the MKD module allows our multimodal MC model to significantly outperform the unimodal models by up to 18.87%, which are trained and tested with only audio or textual data. Zhiqi Huang 0001, Xian Wu 0001, Shen Ge, Helin Wang, Wei Fan 0001, Yuexian Zou |
AAAI | 5 |
| 2021 | A Global-Local Attention Framework for Weakly Labelled Audio TaggingabstractWeakly labelled audio tagging aims to predict the classes of sound events within an audio clip, where the onset and offset times of the sound events are not provided. Previous works have used the multiple instance learning (MIL) framework, and exploited the information of the whole audio clip by MIL pooling functions. However, the detailed information of sound events such as their durations may not be considered under this framework. To address this issue, we propose a novel two-stream framework for audio tagging by exploiting the global and local information of sound events. The global stream aims to analyze the whole audio clip in order to capture the local clips that need to be attended using a class-wise selection module. These clips are then fed to the local stream to exploit the detailed information for a better decision. Experimental results on the AudioSet show that our proposed method can significantly improve the performance of audio tagging under different baseline network architectures. Helin Wang, Yuexian Zou, Wenwu Wang 0001 |
ICASSP | 1 |
| 2021 | Contrastive Self-Supervised Learning for Text-Independent Speaker VerificationabstractCurrent speaker verification models rely on supervised training with massive annotated data. But the collection of labeled utterances from multiple speakers is expensive and facing privacy issues. To open up an opportunity for utilizing massive unlabeled utterance data, our work exploits a contrastive self-supervised learning (CSSL) approach for text-independent speaker verification task. The core principle of CSSL lies in minimizing the distance between the embeddings of augmented segments truncated from the same utterance as well as maximizing those from different utterances. We proposed channel-invariant loss to prevent the network from encoding the undesired channel information into the speaker representation. Bearing these in mind, we conduct intensive experiments on VoxCeleb1&2 datasets. The self-supervised thin-ResNet34 fine-tuned with only 5% of the labeled data can achieve comparable performance to the fully supervised model, which is meaningful to economize lots of manual annotation. Yuexian Zou, Helin Wang |
ICASSP | 3 |
| 2021 | TeCANet: Temporal-Contextual Attention Network for Environment-Aware Speech DereverberationabstractIn this paper, we exploit the effective way to leverage contextual information to improve the speech dereverberation performance in real-world reverberant environments. We propose a temporal-contextual attention approach on the deep neural network (DNN) for environment-aware speech dereverberation, which can adaptively attend to the contextual information. More specifically, a FullBand based Temporal Attention approach (FTA) is proposed, which models the correlations between the fullband information of the context frames. In addition, considering the difference between the attenuation of high frequency bands and low frequency bands (high frequency bands attenuate faster than low frequency bands) in the room impulse response (RIR), we also propose a SubBand based Temporal Attention approach (STA). In order to guide the network to be more aware of the reverberant environments, we jointly optimize the dereverberation network and the reverberation time (RT60) estimator in a multi-task manner. Our experimental results indicate that the proposed method outperforms our previously proposed reverberation-time-aware DNN and the learned attention weights are fully physical consistent. We also report a preliminary yet promising dereverberation and recognition experiment on real test data. Helin Wang, Bo Wu 0011, Lianwu Chen, Meng Yu 0003, Jianwei Yu 0001, Yong Xu 0004, Shixiong Zhang 0001, Chao Weng, Dan Su 0002, Dong Yu 0001 |
Interspeech | 1 |
| 2021 | SpecAugment++: A Hidden Space Data Augmentation Method for Acoustic Scene ClassificationabstractIn this paper, we present SpecAugment++, a novel data augmentation method for deep neural networks based acoustic scene classification (ASC).Different from other popular data augmentation methods such as SpecAugment and mixup that only work on the input space, SpecAugment++ is applied to both the input space and the hidden space of the deep neural networks to enhance the input and the intermediate feature representations.For an intermediate hidden state, the augmentation techniques consist of masking blocks of frequency channels and masking blocks of time frames, which improve generalization by enabling a model to attend not only to the most discriminative parts of the feature, but also the entire parts.Apart from using zeros for masking, we also examine two approaches for masking based on the use of other samples within the minibatch, which helps introduce noises to the networks to make them more discriminative for classification.The experimental results on the DCASE 2018 Task1 dataset and DCASE 2019 Task1 dataset show that our proposed method can obtain 3.6% and 4.7% accuracy gains over a strong baseline without augmentation (i.e.CP-ResNet) respectively, and outperforms other previous data augmentation methods. Helin Wang, Yuexian Zou, Wenwu Wang 0001 |
Interspeech | 1 |
| 2021 | Unsupervised Multi-Target Domain Adaptation for Acoustic Scene ClassificationabstractIt is well known that the mismatch between training (source) and test (target) data distribution will significantly decrease the performance of acoustic scene classification (ASC) systems.To address this issue, domain adaptation (DA) is one solution and many unsupervised DA methods have been proposed.These methods focus on a scenario of single source domain to single target domain.However, we will face such problem that test data comes from multiple target domains.This problem can be addressed by producing one model per target domain, but this solution is too costly.In this paper, we propose a novel unsupervised multi-target domain adaption (MTDA) method for ASC, which can adapt to multiple target domains simultaneously and make use of the underlying relation among multiple domains.Specifically, our approach combines traditional adversarial adaptation with two novel discriminator tasks that learns a common subspace shared by all domains.Furthermore, we propose to divide the target domain into the easy-to-adapt and hard-to-adapt domain, which enables the system to pay more attention to hard-to-adapt domain in training.The experimental results on the DCASE 2020 Task 1-A dataset and the DCASE 2019 Task 1-B dataset show that our proposed method significantly outperforms the previous unsupervised DA methods. Dongchao Yang, Helin Wang, Yuexian Zou |
Interspeech | 2 |
| 2020 | Environmental Sound Classification with Parallel Temporal-Spectral AttentionabstractConvolutional neural networks (CNN) are one of the bestperforming neural network architectures for environmental sound classification (ESC).Recently, temporal attention mechanisms have been used in CNN to capture the useful information from the relevant time frames for audio classification, especially for weakly labelled data where the onset and offset times of the sound events are not applied.In these methods, however, the inherent spectral characteristics and variations are not explicitly exploited when obtaining the deep features.In this paper, we propose a novel parallel temporal-spectral attention mechanism for CNN to learn discriminative sound representations, which enhances the temporal and spectral features by capturing the importance of different time frames and frequency bands.Parallel branches are constructed to allow temporal attention and spectral attention to be applied respectively in order to mitigate interference from the segments without the presence of sound events.The experiments on three environmental sound classification (ESC) datasets and two acoustic scene classification (ASC) datasets show that our method improves the classification performance and also exhibits robustness to noise. Helin Wang, Yuexian Zou, Dading Chong, Wenwu Wang 0001 |
INTERSPEECH | 1 |
| 2020 | Modeling Label Dependencies for Audio Tagging With Graph Convolutional NetworkabstractAs a multi-label classification task, audio tagging aims to predict the presence or absence of certain sound events in an audio recording. Existing works in audio tagging do not explicitly consider the probabilities of the co-occurrences between sound events, which is termed as the label dependencies in this study. To address this issue, we propose to model the label dependencies via a graph-based method, where each node of the graph represents a label. An adjacency matrix is constructed by mining the statistical relations between labels to represent the graph structure information, and a graph convolutional network (GCN) is employed to learn node representations by propagating information between neighboring nodes based on the adjacency matrix, which implicitly models the label dependencies. The generated node representations are then applied to the acoustic representations for classification. Experiments on Audioset show that our method achieves a state-of-the-art mean average precision (mAP) of 0.434. Helin Wang, Yuexian Zou, Dading Chong, Wenwu Wang 0001 |
IEEE Signal Process. Lett. | 1 |
| 2017 | Gait generation and control of biped robot with moving torso based on virtual constraintabstractThis paper presents a novel control method of extended virtual constraint to mimic human movement for a three-link planar robot with moving torso. Inspired by two supine yoga movements, the dynamic model of bipedal walker is modified accordingly, which enlarges application of generalized planar biped robot and provides a more stable and robust walking gait. Due to the continuity of kinematics and discreteness of impact, the walking motion is regarded as a hybrid system, whose zero dynamics determines the stable state of robot. Hence, a within-stride feedback controller is designed based on input-output linearization and extension of virtual constraint. Moreover, poincaré return map is adopted to analyze the stability of walking gait. The simulation results demonstrate the validity of proposed control law, leading to asymptotically stable walking with modified planar biped robot. Helin Wang, Hao Zhang 0008 |
SMC | 1 |