EDBT 2026 Demo / reviewers in the wild / expert
Haizhou Li 0001
dblp:36/4118
· DBLP profile ↗
731ranked-venue papers
17as first author
274since 2021 · last 2026
0000-0001-9158-9401ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 488 · 7 first-author · 158 since 2021Artificial intelligence and machine learning · 474 · 13 first-author · 175 since 2021Databases, data management, data science and information retrieval · 12 · 4 since 2021Human-computer interaction and ubiquitous computing · 12 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 1 first-author · 2 since 2021Systems, architecture and hardware · 7 · 2 since 2021Security and privacy · 4 · 3 since 2021Computer networks · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards Training-Free and Accurate ANN-to-SNN Conversion via Activation-Aware RedistributionabstractConversion represents an effective approach for obtaining low-power models by transforming Artificial Neural Networks (ANNs) into event-driven Spiking Neural Networks (SNNs) without additional training. However, existing training-free conversion methods often incur substantial conversion errors. Here, we first reveal that these conversion errors primarily arise from a distributional mismatch, as the activation distributions of ANNs exhibit channel-wise shifts and scaling, whereas spike rates lack corresponding channel-specific characteristics. To address this limitation, we propose Adaptive Integrate-and-Fire (AIF) neurons with channel-specific thresholds and membrane-potential offsets that dynamically adjust spike rates. These parameters are optimized to jointly minimize conversion errors and maximize information entropy, enabling AIF neurons to capture the activation distribution characteristics of the original ANN. Moreover, AIF neurons can be seamlessly integrated into Transformer architectures with only negligible additional computational cost. Our method achieves state-of-the-art results on multiple vision and natural language processing benchmarks, in particular attaining a notable top-1 accuracy of 85.52% on ImageNet-1K. Honglin Cao, Shuai Wang 0058, Zijian Zhou 0005, Ammar Belatreche, Wenjie Wei, Malu Zhang, Haizhou Li 0001 |
AAAI | 10 |
| 2026 | CATCH: A Controllable Theme Detection Framework with Contextualized Clustering and Hierarchical GenerationabstractTheme detection is a fundamental task in user-centric dialogue systems, aiming to identify the latent topic of each utterance without relying on predefined schemas. Unlike intent induction, which operates within fixed label spaces, theme detection requires cross-dialogue consistency and alignment with personalized user preferences, posing significant challenges. Existing methods often struggle with sparse, short utterances for accurate topic representation and fail to capture user-level thematic preferences across dialogues. To address these challenges, we propose CATCH (Controllable Theme Detection with Contextualized Clustering and Hierarchical Generation), a unified framework that integrates three core components: (1) context-aware topic representation, which enriches utterance-level semantics using surrounding topic segments; (2) preference-guided topic clustering, which jointly models semantic proximity and personalized feedback to align themes across dialogue; and (3) a hierarchical theme generation mechanism designed to suppress noise and produce robust, coherent topic labels. Experiments on a multi-domain customer dialogue benchmark (DSTC-12) demonstrate the effectiveness of CATCH with 8B LLM in both theme clustering and topic generation quality. Rui Ke, Shenghao Yang 0001, Kuang Wang, Feng Jiang 0007, Haizhou Li 0001 |
AAAI | 6 |
| 2026 | NaturalSloth: Revisiting Denial-of-Service Attacks on Large Language ModelsabstractLLM serving is limited by provider-side resources: longer generations consume more GPU time, increase latency, and reduce throughput in multi-tenant systems.This creates a denial-of-service (DoS) risk, where attackers degrade service by inducing excessive generation.Prior work on LLM DoS primarily relies on adversarial perturbations that delay end-of-sequence termination.We show perturbations are often unnecessary: natural, benignlooking instructions that specify impractical and meaningless tasks can already trigger excessive generation.To study this overlooked vulnerability, we introduce NaturalSloth, an adversarial dataset of natural, instruction-based DoS prompts.Starting from a human-curated seed set spanning diverse attack categories, we design a multi-agent synthesis framework to scale the dataset while preserving malicious intent and increasing semantic diversity.Experiments across a wide range of proprietary and open-source LLMs show that NaturalSloth consistently induces excessive generation, with attack effectiveness further amplified when combined with jailbreak techniques.Our analysis also reveals significant limitations of existing defenses, highlighting the need for dedicated protections against natural DoS attacks. 1 Yiming Chen 0010, Zexin Li 0001, Xianghu Yue, Robby T. Tan, Haizhou Li 0001 |
ACL (1) | 5 |
| 2026 | S2S-Arena: Evaluating Paralinguistic Instruction Following in Speech-to-Speech ModelsabstractFeng Jiang, Zhiyu Lin, Yiyang Liu, Liumeng Xue, Fan Bu, Yuhao Du, Xiangying Chen, Benyou Wang, Haizhou Li. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Feng Jiang 0007, Liumeng Xue, Xiangying Chen, Benyou Wang, Haizhou Li 0001 |
ACL (1) | 9 |
| 2026 | UniMoE-Audio: Unified Speech and Music Generation with Dynamic-Capacity Mixture-of-ExpertsabstractZhenyu Liu, Yunxin li, Xuanyu Zhang, Qixun Teng, Shenyuan Jiang, Xinyu Chen, Haoyuan Shi, Haolan Chen, Fanbo Meng, Mingjun Zhao, Yu Xu, Yancheng He, Baotian Hu, Haizhou Li, Min Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yunxin Li, Xuanyu Zhang 0006, Qixun Teng, Shenyuan Jiang, Xinyu Chen 0003, Haolan Chen, Mingjun Zhao, Yancheng He, Baotian Hu, Haizhou Li 0001, Min Zhang 0005 |
ACL (1) | 14 |
| 2026 | Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMsabstractZhenyu Liu, Xuanyu Zhang, Yunxin li, Qixun Teng, Shenyuan Jiang, Haolan Chen, Mingjun Zhao, Fanbo Meng, Yu Xu, Yancheng He, Baotian Hu, Haizhou Li, Min Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Xuanyu Zhang 0006, Yunxin Li, Qixun Teng, Shenyuan Jiang, Haolan Chen, Mingjun Zhao, Yancheng He, Baotian Hu, Haizhou Li 0001, Min Zhang 0005 |
ACL (1) | 12 |
| 2026 | Improving Conversational Recommendation with Contextual Adaptation of External Recommenders and LLM-Based Reranking
Chuang Li 0006, Weida Liang, Hengchang Hu, See-Kiong Ng, Min-Yen Kan, Haizhou Li 0001, Yang Deng 0002 |
ECIR (2) | 6 |
| 2026 | Spiking neural networks for EEG signal analysis: From theory to practice
Siqi Cai 0002, Zheyuan Lin, Wenjie Wei, Shuai Wang 0058, Malu Zhang, Tanja Schultz, Haizhou Li 0001 |
Neural Networks | 8 |
| 2026 | Closed-loop correction reprogramming for fine-grained visual prompting
Xueyi Zhang 0001, Siqi Cai 0002, Mingrui Lao, Haizhou Li 0001 |
Neural Networks | 5 |
| 2026 | PAL: Prompting analytic learning with missing modality for multi-modal class-incremental learning
Xianghu Yue, Yiming Chen 0010, Xueyi Zhang 0001, Xiaoxue Gao, Mengling Feng, Mingrui Lao, Huiping Zhuang, Haizhou Li 0001 |
Pattern Recognit. | 8 |
| 2026 | The effect of speech representations on EEG-based auditory attention detection
Siqi Cai 0002, Haizhou Li 0001 |
Pattern Recognit. Lett. | 4 |
| 2026 | Emphasis rendering for conversational text-to-speech with multi-modal multi-scale context modeling
Rui Liu 0008, Zhenqi Jia, Yifan Hu 0004, Haizhou Li 0001 |
Speech Commun. | 5 |
| 2026 | TPEech: Target Speaker Extraction and Noise Suppression With Historical Dialogue Text CuesabstractIn complex multi-speaker scenarios with significant speaker overlap and background noise, extracting the target speaker's speech remains a major challenge. This capability is crucial for dialogue-based applications such as AI speech assistants, where downstream tasks such as speech recognition depend on clean speech. A potential solution to address these challenges is Target Speaker Extraction (TSE), which leverages auxiliary information to extract target speech from mixed and noisy speech, thus overcoming the limitations of Speech Separation (SS) and Speech Enhancement (SE). In particular, we propose a multi-modal TSE network, namely Text Prompt Extractor with echo cue block (TPEech), which uses historical dialogue text as cues for extraction and incorporates the echo cue block (ECB) to further exploit this cue and enhance TSE performance. The experiments show the excellent extraction and denoising capabilities of our proposed network. TPEech achieves an SI-SDRi of 9.632 dB, an SDR of 13.045 dB, a PESQ of 2.814, and a STOI of 0.885, outperforming competitive baselines. Additionally, we experimentally verify that TPEech is robust against semantically incomplete textual prompts. Dataset and source code will be publicly available. Ziyang Jiang, Shuai Wang 0016, Xinyuan Qian 0001, Haizhou Li 0001 |
IEEE Signal Process. Lett. | 5 |
| 2026 | VoiceBench: Benchmarking LLM-Based Voice AssistantsabstractAbstract Recent advancements in large language models (LLMs) like GPT-4o have enabled real-time speech interactions through LLM-based voice assistants, offering an improved user experience over text-based interactions. However, a suitable benchmark to rigorously evaluate such speech interactions systems is currently lacking. To bridge this gap, we introduce VoiceBench, the first benchmark specifically designed to assess LLM-based voice assistants. VoiceBench comprises 6,783 synthetic and real spoken instructions recorded from diverse speakers across eight distinct tasks. These instructions are meticulously crafted to assess three crucial capability areas: general knowledge, instruction-following, and safety compliance. Furthermore, VoiceBench systematically incorporates realistic variations common in spoken interactions, including differences in speaker characteristics (e.g., accents), heterogeneous environmental conditions (e.g., reverberation), and content complexities such as mispronunciations. Extensive experiments reveal the limitations of current LLM-based voice assistant models and offer valuable insights for future research and development in this field.1 Yiming Chen 0010, Xianghu Yue, Chen Zhang 0020, Xiaoxue Gao, Robby T. Tan, Haizhou Li 0001 |
Trans. Assoc. Comput. Linguistics | 6 |
| 2026 | Lightweight and Personalized Single-Eye Emotion Recognition via CNN-SNN Spatiotemporal Learning and Memory-Inferred Event FeaturesabstractEmotion recognition is essential for improving user experience and interaction quality in human-centered applications. While recent studies have leveraged both event and traditional cameras to enhance eye-based emotion recognition, their practical deployment is hindered by the scarcity of event cameras and the complexity of dual-modality frameworks. Personalization, which is critical for handling individual differences in emotional expression, is also affected by these factors, resulting in reduced performance and adaptation efficiency. To address these challenges, we propose a lightweight and personalized single-eye emotion recognition network, called LPSEER. LPSEER introduces a novel hybrid neural architecture that integrates a convolutional neural network (CNN) and a spiking neural network (SNN) to capture spatiotemporal features from video frames and events, respectively. Additionally, we design a memorybased event feature inference (MEFI) module that recalls event features from video frames, eliminating the reliance on event cameras during inference and personalization while retaining the discriminative advantages of event-based representations. Experimental results demonstrate that LPSEER achieves state-of- the-art recognition accuracy while maintaining the smallest model size and lowest computational cost. Further experiments confirm the strong generalization capabilities and the ability to achieve faster, more accurate personalization. These advantages collectively enable lightweight, accurate, and efficient emotion recognition for real-world human-centered applications. Qianhui Liu, Jiqing Zhang, Yang Wang 0106, Malu Zhang, Xin Yang 0011, Gang Pan 0001, Haizhou Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2026 | IML-Spikeformer: Input-Aware Multilevel Spiking Transformer for Speech ProcessingabstractSpiking neural networks (SNNs), inspired by biological neural mechanisms, represent a promising neuromorphic computing paradigm that offers energy-efficient alternatives to traditional artificial neural networks (ANNs). Despite proven effectiveness, SNN architectures have struggled to achieve competitive performance on large-scale speech processing tasks. Two key challenges hinder progress: 1) the high computational overhead during training caused by multitimestep spike firing and 2) the absence of large-scale SNN architectures tailored to speech processing tasks. To overcome the issues, we introduce the input-aware multilevel spikeformer (IML-Spikeformer), a spiking transformer architecture specifically designed for large-scale speech processing. Central to our design is the input-aware multilevel spike (IMLS) mechanism, which simulates multitimestep spike firing within a single timestep using an adaptive, input-aware thresholding scheme. IML-Spikeformer further integrates a reparameterized spiking self-attention (RepSSA) module with a hierarchical decay mask (HDM), forming the HD-RepSSA module. This module enhances the precision of attention maps and enables modeling of multiscale temporal dependencies in speech signals. Experiments demonstrate that IML-Spikeformer achieves word error rates (WERs) of 6.0% on AiShell-1 and 3.4% on Librispeech-960, comparable to conventional ANN transformers while reducing theoretical inference energy consumption by $4.64\times $ and $4.32\times $ , respectively. IML-Spikeformer marks an advance of scalable SNN architectures for large-scale speech processing in both task performance and energy efficiency. Our source code and model checkpoints are publicly available at github.com/Pooookeman/IML-Spikeformer. Zeyang Song, Yuhong Chou, Jibin Wu, Haizhou Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2026 | SNN-FT: Temporal-Coded Spiking Neural Networks for Fourier TransformabstractThe Fourier transform (FT) stands as a fundamental tool in modern signal processing with widespread applications across various scientific and engineering fields. Therefore, there remains a need for continued research efforts to devise energy-efficient implementations of the FT. Due to their inherent energy efficiency, biologically plausible spiking neural networks (SNNs) emerge as a promising alternative solution. However, current SNN implementations of the FT suffer from two key shortcomings, namely, high latency and reduced accuracy. In this article, we analyze the underlying causes of these limitations and highlight deficiencies in the existing spike-based encoding mechanisms and spiking neuron models. We then propose a new SNN-based FT (SNN-FT) based on a logarithmically polarized time-to-first-spike (TTFS) encoding method (called LP-TTFS) along with a novel piecewise spiking neuron (PTSN) model based on ternary spikes (referred to as PTSN). The resulting SNN-FT is mathematically equivalent to the conventional FT and demonstrates superior performance in accuracy as well as reduced latency. We assess the performance of the proposed SNN-FT alternative through extensive experiments on FT-based applications, such as radar and audio signal processing, and the obtained results demonstrate the efficacy of SNN-FT and its superiority over the existing approaches. This study unveils a novel energy-efficient neuromorphic computing technique with great potential for FT applications across diverse scientific and engineering domains. Shuai Wang 0058, Haorui Zheng, Ammar Belatreche, Guoqing Wang 0001, Yeying Jin, Jibin Wu, Malu Zhang, Yang Yang 0002, Haizhou Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 10 |
| 2025 | Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-SpeechabstractVisual Text-to-Speech (VTTS) aims to take the environmental image as the prompt to synthesize the reverberant speech for the spoken content. The challenge of this task lies in understanding the spatial environment from the image. Many attempts have been made to extract global spatial visual information from the RGB space of an spatial image. However, local and depth image information are crucial for understanding the spatial environment, which previous works have ignored. To address the issues, we propose a novel multi-modal and multi-scale spatial environment understanding scheme to achieve immersive VTTS, termed M2SE-VTTS. The multi-modal aims to take both the RGB and Depth spaces of the spatial image to learn more comprehensive spatial information, and the multi-scale seeks to model the local and global spatial knowledge simultaneously. Specifically, we first split the RGB and Depth images into patches and adopt the Gemini-generated environment captions to guide the local spatial understanding. After that, the multi-modal and multi-scale features are integrated by the local-aware global spatial understanding. In this way, M2SE-VTTS effectively models the interactions between local and global spatial contexts in the multi-modal spatial environment. Objective and subjective evaluations suggest that our model outperforms the advanced baselines in environmental speech generation. Rui Liu 0008, Shuwei He, Yifan Hu 0004, Haizhou Li 0001 |
AAAI | 4 |
| 2025 | Aligning Language Models Using Follow-up Likelihood as Reward SignalabstractIn natural human-to-human conversations, participants often receive feedback signals from one another based on their follow-up reactions. These reactions can include verbal responses, facial expressions, changes in emotional state, and other non-verbal cues. Similarly, in human-machine interactions, the machine can leverage the user's follow-up utterances as feedback signals to assess whether it has appropriately addressed the user's request. Therefore, we propose using the likelihood of follow-up utterances as rewards to differentiate preferred responses from less favored ones, without relying on human or commercial LLM-based preference annotations. Our proposed reward mechanism, ``Follow-up Likelihood as Reward" (FLR), matches the performance of strong reward models trained on large-scale human or GPT-4 annotated data on 8 pairwise-preference and 4 rating-based benchmarks. Building upon the FLR mechanism, we propose to automatically mine preference data from the online generations of a base policy model. The preference data are subsequently used to boost the helpfulness of the base model through direct alignment from preference (DAP) methods, such as direct preference optimization (DPO). Lastly, we demonstrate that fine-tuning the language model that provides follow-up likelihood with natural language feedback significantly enhances FLR's performance on reward modeling benchmarks and effectiveness in aligning the base policy model's helpfulness. Chen Zhang 0055, Dading Chong, Feng Jiang 0007, Chengguang Tang, Anningzhe Gao, Guohua Tang, Haizhou Li 0001 |
AAAI | 7 |
| 2025 | SongEditor: Adapting Zero-Shot Song Generation Language Model as a Multi-Task EditorabstractThe emergence of novel generative modeling paradigms, particularly audio language models, has significantly advanced the field of song generation. Although state-of-the-art models are capable of synthesizing both vocals and accompaniment tracks up to several minutes long concurrently, research about partial adjustments or editing of existing songs is still underexplored, which allows for more flexible and effective production. In this paper, we present SongEditor, the first song editing paradigm that introduces the editing capabilities into language-modeling song generation approaches, facilitating both segment-wise and track-wise modifications. SongEditor offers the flexibility to adjust lyrics, vocals, and accompaniments, as well as synthesizing songs from scratch. The core components of SongEditor include a music tokenizer, an autoregressive language model, and a diffusion generator, enabling generating an entire section, masked lyrics, or even separated vocals and background music. Extensive experiments demonstrate that the proposed SongEditor achieves exceptional performance in end-to-end song editing, as evidenced by both objective and subjective metrics. Shuai Wang 0016, Hangting Chen, Jianwei Yu 0001, Wei Tan 0011, Rongzhi Gu, Yaoxun Xu, Yizhi Zhou, Haina Zhu, Haizhou Li 0001 |
AAAI | 10 |
| 2025 | UniCodec: Unified Audio Codec with Single Domain-Adaptive CodebookabstractYidi Jiang, Qian Chen, Shengpeng Ji, Yu Xi, Wen Wang, Chong Zhang, Xianghu Yue, ShiLiang Zhang, Haizhou Li. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Yidi Jiang, Qian Chen 0003, Shengpeng Ji, Yu Xi, Wen Wang 0001, Chong Zhang 0003, Xianghu Yue, Shiliang Zhang, Haizhou Li 0001 |
ACL (1) | 9 |
| 2025 | Know You First and Be You Better: Modeling Human-Like User Simulators via Implicit ProfilesabstractUser simulators are crucial for replicating human interactions with dialogue systems, supporting both collaborative training and automatic evaluation, especially for large language models (LLMs). However, current role-playing methods face challenges such as a lack of utterance-level authenticity and user-level diversity, often hindered by role confusion and dependence on predefined profiles of well-known figures. In contrast, direct simulation focuses solely on text, neglecting implicit user traits like personality and conversation-level consistency. To address these issues, we introduce the User Simulator with Implicit Profiles (USP), a framework that infers implicit user profiles from human-machine interactions to simulate personalized and realistic dialogues. We first develop an LLM-driven extractor with a comprehensive profile schema, then refine the simulation using conditional supervised fine-tuning and reinforcement learning with cycle consistency, optimizing at both the utterance and conversation levels. Finally, a diverse profile sampler captures the distribution of real-world user profiles. Experimental results show that USP outperforms strong baselines in terms of authenticity and diversity while maintaining comparable consistency. Additionally, using USP to evaluate LLM on dynamic multi-turn aligns well with mainstream benchmarks, demonstrating its effectiveness in real-world applications. Kuang Wang, Xianfei Li, Shenghao Yang 0001, Li Zhou 0010, Feng Jiang 0007, Haizhou Li 0001 |
ACL (1) | 6 |
| 2025 | Soundwave: Less is More for Speech-Text Alignment in LLMsabstractExisting end-to-end speech large language models (LLMs) usually rely on large-scale annotated data for training, while data-efficient training has not been discussed in depth.We focus on two fundamental problems between speech and text: the representation space gap and sequence length inconsistency.We propose Soundwave, which utilizes an efficient training strategy and a novel architecture to address these issues.Results show that Soundwave outperforms other advanced speech LLMs in speech translation and AIR-Bench speech tasks with only a fraction of the training data.Further analysis shows that Soundwave still retains its intelligence during conversation. Ruiyu Zhang, Benyou Wang, Haizhou Li 0001 |
ACL (1) | 6 |
| 2025 | Second Language (Arabic) Acquisition of LLMs via Progressive Vocabulary ExpansionabstractJianqing Zhu, Huang Huang, Zhihang Lin, Juhao Liang, Zhengyang Tang, Khalid Almubarak, Mosen Alharthi, Bang An, Juncai He, Xiangbo Wu, Fei Yu, Junying Chen, Ma Zhuoheng, Yuhao Du, He Zhang, Saied Alshahrani, Emad A. Alghamdi, Lian Zhang, Ruoyu Sun, Haizhou Li, Benyou Wang, Jinchao Xu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Jianqing Zhu, Zhihang Lin, Juhao Liang, Zhengyang Tang, Khalid Almubarak, Mosen Alharthi, Bang An 0004, Juncai He 0001, Xiangbo Wu, Fei Yu 0017, Zhuoheng Ma, Saied Alshahrani, Emad A. Alghamdi, Ruoyu Sun 0001, Haizhou Li 0001, Benyou Wang, Jinchao Xu |
ACL (1) | 20 |
| 2025 | Benchmarking Large Language Models Under Data Contamination: A Survey from Static to Dynamic EvaluationabstractSimin Chen, Yiming Chen, Zexin Li, Yifan Jiang, Zhongwei Wan, Yixin He, Dezhi Ran, Tianle Gu, Haizhou Li, Tao Xie, Baishakhi Ray. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Yiming Chen 0010, Zexin Li 0001, Zhongwei Wan, Yixin He 0002, Dezhi Ran, Tianle Gu, Haizhou Li 0001, Tao Xie 0001, Baishakhi Ray |
EMNLP | 9 |
| 2025 | From Word to World: Evaluate and Mitigate Culture Bias in LLMs via Word Association TestabstractThe human-centered word association test (WAT) serves as a cognitive proxy, revealing sociocultural variations through culturally shared semantic expectations and implicit linguistic patterns shaped by lived experiences.We extend this test into an LLM-adaptive, freerelation task to assess the alignment of large language models (LLMs) with cross-cultural cognition.To address culture preference, we propose CultureSteer, an innovative approach that moves beyond superficial cultural prompting by embedding cultural-specific semantic associations directly within the model's internal representation space.Experiments show that current LLMs exhibit significant bias toward Western (notably American) schemas at the word association level.In contrast, our model substantially improves cross-cultural alignment, capturing diverse semantic associations.Further validation on culture-sensitive downstream tasks confirms its efficacy in fostering cognitive alignment across cultures.This work contributes a novel methodological paradigm for enhancing cultural awareness in LLMs, advancing the development of more inclusive language technologies. Xunlian Dai, Li Zhou 0010, Benyou Wang, Haizhou Li 0001 |
EMNLP | 4 |
| 2025 | Multimodal Fine-grained Context Interaction Graph Modeling for Conversational Speech SynthesisabstractConversational Speech Synthesis (CSS) aims to generate speech with natural prosody by understanding the multimodal dialogue history (MDH).The latest work predicts the accurate prosody expression of the target utterance by modeling the utterance-level interaction characteristics of MDH and the target utterance.However, MDH contains fine-grained semantic and prosody knowledge at the word level.Existing methods overlook the fine-grained semantic and prosodic interaction modeling.To address this gap, we propose MFCIG-CSS, a novel Multimodal Fine-grained Context Interaction Graph-based CSS system.Our approach constructs two specialized multimodal fine-grained dialogue interaction graphs: a semantic interaction graph and a prosody interaction graph.These two interaction graphs effectively encode interactions between word-level semantics, prosody, and their influence on subsequent utterances in MDH.The encoded interaction features are then leveraged to enhance synthesized speech with natural conversational prosody.Experiments on the DailyTalk dataset demonstrate that MFCIG-CSS outperforms all baseline models in terms of prosodic expressiveness.Code and speech samples are available at https://github.com/AI-S2-Lab/MFCIG-CSS. Zhenqi Jia, Rui Liu 0008, Berrak Sisman, Haizhou Li 0001 |
EMNLP | 4 |
| 2025 | Hanfu-Bench: A Multimodal Benchmark on Cross-Temporal Cultural Understanding and TranscreationabstractCulture is a rich and dynamic domain that evolves across both geography and time.However, existing studies on cultural understanding with vision-language models (VLMs) primarily emphasize geographic diversity, often overlooking the critical temporal dimensions.To bridge this gap, we introduce Hanfu-Bench, a novel, expert-curated multimodal dataset.Hanfu, a traditional garment spanning ancient Chinese dynasties, serves as a representative cultural heritage that reflects the profound temporal aspects of Chinese culture while remaining highly popular in Chinese contemporary society.Hanfu-Bench comprises two core tasks: cultural visual understanding and cultural image transcreation.The former task examines temporal-cultural feature recognition based on single-or multi-image inputs through multiplechoice visual question answering, while the latter focuses on transforming traditional attire into modern designs through cultural element inheritance and modern context adaptation.Our evaluation shows that closed VLMs perform comparably to non-experts on visual cutural understanding but fall short by 10% to human experts, while open VLMs lags further behind non-experts.For the transcreation task, multi-faceted human evaluation indicates that the best-performing model achieves a success rate of only 42%.Our benchmark provides an essential testbed, revealing significant challenges in this new direction of temporal cultural understanding and creative adaptation. Li Zhou 0010, Lutong Yu, Dongchu Xie, Shaohuan Cheng, Wenyan Li 0001, Haizhou Li 0001 |
EMNLP | 6 |
| 2025 | MEIJU - The 1st Multimodal Emotion and Intent Joint Understanding ChallengeabstractMultimodal Emotion and Intent Joint Understanding (MEIJU) aims to decode the semantic information expressed in the multimodal dialogues while inferring the emotions and intents, providing users with a more humanized human-machine interaction experience. However, challenges such as difficulties in data acquisition and imbalance annotations have made it difficult for current methods to meet the demands of practical applications. Therefore, we have organized two tracks focusing on the key themes of semi-supervised learning and class imbalance. Additionally, we have prepared data in two different languages (English and Mandarin) for each track, treating each language as a sub-track, to encourage participants to explore solutions in more diverse linguistic environments. Our code can be found at https://github.com/AI-S2-Lab/MEIJU2025-baseline. Rui Liu 0008, Xiaofen Xing, Zheng Lian 0004, Haizhou Li 0001, Björn W. Schuller, Haolin Zuo |
ICASSP | 4 |
| 2025 | Speech Separation for Low-Resource LanguagesabstractSpeech separation aims to equip machines with the human ability of selective listening, i.e. to focus attention on specific information in spoken communication. Studies have shown that the language spoken in a cocktail party scenario matters. While the development of speech separation models can leverage extensive databases, for the majority of languages only very limited data is available. This work presents the very first study on speech separation for low-resource languages. We choose blind source separation as the task to be studied and analyze three strategies to overcome the data scarcity of two low-resource languages from the GlobalPhoneMS2 database. We show that data from other languages can be used to develop models that work for low-resource languages. Finetuning additionally boosts the performance, and training on multiple languages increases both performance and robustness. We show that dynamic mixing in the development helps to find a trade-off between performance and development time. Marvin Borsdorf, Zexu Pan, Pascal Himmelmann, Haizhou Li 0001, Tanja Schultz |
ICASSP | 4 |
| 2025 | MacST: Multi-Accent Speech Synthesis via Text Transliteration for Accent ConversionabstractIn accented voice conversion or accent conversion, we seek to convert the accent in speech from one another while preserving speaker identity and semantic content. In this study, we formulate a novel method for creating multi-accented speech samples, thus pairs of accented speech samples by the same speaker, through text transliteration for training accent conversion systems. We begin by generating transliterated text with Large Language Models (LLMs), which is then fed into multilingual TTS models to synthesize accented English speech. As a reference system, we built a sequence-to-sequence model on the synthetic parallel corpus for accent conversion. We validated the proposed method for both native and non-native English speakers. Subjective and objective evaluations further validate our dataset’s effectiveness in accent conversion studies. Sho Inoue, Shuai Wang 0016, Wanxing Wang, Pengcheng Zhu 0004, Mengxiao Bi, Haizhou Li 0001 |
ICASSP | 6 |
| 2025 | E1 TTS: Simple and Fast Non-Autoregressive TTSabstractThis paper introduces Easy One-Step Text-to-Speech (E1 TTS), an efficient non-autoregressive zero-shot text-to-speech system based on denoising diffusion pretraining and distribution matching distillation. The training of E1 TTS is straightforward; it does not require explicit monotonic alignment between the text and audio pairs. The inference of E1 TTS is efficient, requiring only one neural network evaluation for each utterance. Despite its sampling efficiency, E1 TTS achieves naturalness and speaker similarity comparable to various strong baseline models. Audio samples are available at e1tts.github.io. Shuai Wang 0016, Pengcheng Zhu 0004, Mengxiao Bi, Haizhou Li 0001 |
ICASSP | 5 |
| 2025 | ATGnet: Adaptive Temporal Graph Network for EEG-enabled Sound Source Tracking in Cocktail Party ScenariosabstractDecoding selective auditory attention from electroencephalography (EEG) signals has gained considerable interest. However, few studies have looked into tracking the dynamic trajectory of moving sound source in complex auditory environments, e.g. with multiple moving speakers. We propose a novel model, namely Adaptive Temporal Graph Network (ATGnet), to continuously track the sound source trajectory using spatial-temporal EEG representations. ATGnet incorporates an adaptive graph topology to extract spatial features, and a graph-convolutional long short-term memory (GC-LSTM) network to capture spatial-temporal dependency. We evaluated ATGnet by performing within-subject leave-one-trial-out cross-validation on EEG signals from 10 participants. Experiment results indicate that ATGnet effectively overcomes the variation of signals across trials and subjects. They further confirm that ATGnet robustly tracks both attended and unattended sound sources, and significantly outperforms traditional methods. ATGnet offers a promising solution to continuous sound source tracking in dynamic conditions, with potential applications in neuro-steered hearing devices. Saurav Pahuja, Gabriel Ivucic, Siqi Cai 0002, Dashanka De Silva, Tanja Schultz, Haizhou Li 0001 |
ICASSP | 6 |
| 2025 | Multi-Level Speaker Representation for Target Speaker ExtractionabstractTarget speaker extraction (TSE) relies on a reference cue of the target to extract the target speech from a speech mixture. While a speaker embedding is commonly used as the reference cue, such embedding pre-trained with a large number of speakers may suffer from confusion of speaker identity. In this work, we propose a multi-level speaker representation approach, from raw features to neural embeddings, to serve as the speaker reference cue. We generate a spectral-level representation from the enrollment magnitude spectrogram as a raw, low-level feature, which significantly improves the model’s generalization capability. Additionally, we propose a contextual embedding feature based on cross-attention mechanisms that integrate frame-level embeddings from a pre-trained speaker encoder. By incorporating speaker features across multiple levels, we significantly enhance the performance of the TSE model. Our approach achieves a 2.74 dB improvement and a 4.94% increase in extraction accuracy on Libri2mix test set over the baseline. Shuai Wang 0016, Yangjie Wei, Yannan Wang, Haizhou Li 0001 |
ICASSP | 7 |
| 2025 | Quantized Spike-driven TransformerabstractSpiking neural networks (SNNs) are emerging as a promising energy-efficient alternative to traditional artificial neural networks (ANNs) due to their spike-driven paradigm.
However, recent research in the SNN domain has mainly focused on enhancing accuracy by designing large-scale Transformer structures, which typically rely on substantial computational resources, limiting their deployment on resource-constrained devices.
To overcome this challenge, we propose a quantized spike-driven Transformer baseline (QSD-Transformer), which achieves reduced resource demands by utilizing a low bit-width parameter.
Regrettably, the QSD-Transformer often suffers from severe performance degradation.
In this paper, we first conduct empirical analysis and find that the bimodal distribution of quantized spike-driven self-attention (Q-SDSA) leads to spike information distortion (SID) during quantization, causing significant performance degradation. To mitigate this issue, we take inspiration from mutual information entropy and propose a bi-level optimization strategy to rectify the information distribution in Q-SDSA.
Specifically, at the lower level, we introduce an information-enhanced LIF to rectify the information distribution in Q-SDSA.
At the upper level, we propose a fine-grained distillation scheme for the QSD-Transformer to align the distribution in Q-SDSA with that in the counterpart ANN.
By integrating the bi-level optimization strategy, the QSD-Transformer can attain enhanced energy efficiency without sacrificing its high-performance advantage.
We validate the QSD-Transformer on various visual tasks, and experimental results indicate that our method achieves state-of-the-art results in the SNN domain.
For instance, when compared to the prior SNN benchmark on ImageNet, the QSD-Transformer achieves 80.3\% top-1 accuracy, accompanied by significant reductions of 6.0$\times$ and 8.1$\times$ in power consumption and model size, respectively. Code is available at https://github.com/bollossom/QSD-Transformer. Xuerui Qiu, Malu Zhang, Jieyuan Zhang, Wenjie Wei, Honglin Cao, Junsheng Guo, Rui-Jie Zhu 0003, Yimeng Shan, Yang Yang 0002, Haizhou Li 0001 |
ICLR | 10 |
| 2025 | MoMuSE: Momentum Multi-modal Target Speaker Extraction for Real-time Scenarios with Impaired Visual CuesabstractAudio-visual Target Speaker Extraction (AV-TSE) aims to isolate the speech of a specific target speaker from an audio mixture using time-synchronized visual cues. In real-world scenarios, visual cues are not always available due to various impairments, which undermines the stability of AV-TSE. Despite this challenge, humans can maintain attentional momentum over time, even when the target speaker is not visible. In this paper, we introduce the Momentum Multi-modal target Speaker Extraction (MoMuSE), which retains a speaker identity momentum in memory, enabling the model to continuously track the target speaker. Designed for real-time inference, MoMuSE extracts the current speech window with guidance from both visual cues and dynamically updated speaker momentum. Experimental results demonstrate that MoMuSE exhibits significant improvement, particularly in scenarios with severe impairment of visual cues. Shuai Wang 0016, Kong-Aik Lee, Man-Wai Mak, Haizhou Li 0001 |
ICME | 6 |
| 2025 | Binary Event-Driven Spiking TransformerabstractTransformer-based Spiking Neural Networks (SNNs) introduce a novel event-driven self-attention paradigm that combines the high performance of Transformers with the energy efficiency of SNNs. However, the larger model size and increased computational demands of the Transformer structure limit their practicality in resource-constrained scenarios. In this paper, we integrate binarization techniques into Transformer-based SNNs and propose the Binary Event-Driven Spiking Transformer, i.e. BESTformer. The proposed BESTformer can significantly reduce storage and computational demands by representing weights and attention maps with a mere 1-bit. However, BESTformer suffers from a severe performance drop from its full-precision counterpart due to the limited representation capability of binarization. To address this issue, we propose a Coupled Information Enhancement (CIE) method, which consists of a reversible framework and information enhancement distillation. By maximizing the mutual information between the binary model and its full-precision counterpart, the CIE method effectively mitigates the performance degradation of the BESTformer. Extensive experiments on static and neuromorphic datasets demonstrate that our method achieves superior performance to other binary SNNs, showcasing its potential as a compact yet high-performance model for resource-limited edge devices. The repository of this paper is available at https://github.com/CaoHLin/BESTFormer. Honglin Cao, Zijian Zhou 0005, Wenjie Wei, Ammar Belatreche, Dehao Zhang, Malu Zhang, Yang Yang 0002, Haizhou Li 0001 |
IJCAI | 9 |
| 2025 | Towards Emotionally Consistent Text-Based Speech Editing: Introducing EmoCorrector and The ECD-TSE Dataset
Rui Liu 0008, Pu Gao, Jiatian Xi, Berrak Sisman, Carlos Busso, Haizhou Li 0001 |
INTERSPEECH | 6 |
| 2025 | Accent Normalization Using Self-Supervised Discrete Tokens with Non-Parallel Data
Qibing Bai, Sho Inoue, Zhongjie Jiang, Yannan Wang, Haizhou Li 0001 |
INTERSPEECH | 6 |
| 2025 | PersonaTAB: Predicting Personality Traits using Textual, Acoustic, and Behavioral Cues in Fully-Duplex Speech Dialogs
Sho Inoue, Shuai Wang 0016, Haizhou Li 0001 |
INTERSPEECH | 3 |
| 2025 | REAL-T: Real Conversational Mixtures for Target Speaker Extraction
Shaole Li, Shuai Wang 0016, Jiangyu Han, Wupeng Wang, Haizhou Li 0001 |
INTERSPEECH | 6 |
| 2025 | SpeechRefiner: Towards Perceptual Quality Refinement for Front-End Algorithms
Zhongjie Jiang, Yannan Wang, Haizhou Li 0001 |
INTERSPEECH | 6 |
| 2025 | Decoding Listener's Identity: Person Identification from EEG Signals Using a Lightweight Spiking TransformerabstractEEG-based person identification enables applications in security, personalized brain-computer interfaces (BCIs), and cognitive monitoring. However, existing techniques often rely on deep learning architectures at high computational cost, limiting their scope of applications. In this study, we propose a novel EEG person identification approach using spiking neural networks (SNNs) with a lightweight spiking transformer for efficiency and effectiveness. The proposed SNN model is capable of handling the temporal complexities inherent in EEG signals. On the EEG-Music Emotion Recognition Challenge dataset, the proposed model achieves 100% classification accuracy with less than 10% energy consumption of traditional deep neural networks. This study offers a promising direction for energy-efficient and high-performance BCIs. The source code is available at https://github.com/PatrickZLin/Decode-ListenerIdentity. Zheyuan Lin, Siqi Cai 0002, Haizhou Li 0001 |
INTERSPEECH | 3 |
| 2025 | GTAnet: Geometry-Guided Temporal Attention for EEG-Based Sound Source Tracking in Cocktail Party Scenarios
Saurav Pahuja, Gabriel Ivucic, Siqi Cai 0002, Dashanka De Silva, Haizhou Li 0001, Tanja Schultz |
INTERSPEECH | 5 |
| 2025 | NeuroSpex+: Dual-Task Training of Neuro-Guided Speaker Extraction with Speech Envelope and Waveform
Dashanka De Silva, Siqi Cai 0002, Saurav Pahuja, Tanja Schultz, Haizhou Li 0001 |
INTERSPEECH | 5 |
| 2025 | Incorporating Linguistic Constraints from External Knowledge Source for Audio-Visual Target Speech Extraction
Shuai Wang 0016, Xixin Wu, Helen M. Meng, Haizhou Li 0001 |
INTERSPEECH | 5 |
| 2025 | TVC-MusicGen: Time-Varying Structure Control for Background Music Generation via Self-Supervised Training
Hangting Chen, Shuai Wang 0016, Haina Zhu, Haizhou Li 0001 |
INTERSPEECH | 5 |
| 2025 | UniTalker: Conversational Speech-Visual SynthesisabstractConversational Speech Synthesis (CSS) is a key task in the user-agent interaction area, aiming to generate more expressive and empathetic speech for users. However, it is well-known that ''listening'' and ''eye contact'' play crucial roles in conveying emotions during real-world interpersonal communication. Existing CSS research is limited to perceiving only text and speech within the dialogue context, which restricts its effectiveness. Moreover, speech-only responses further constrain the interactive experience. To address these limitations, we introduce a Conversational Speech-Visual Synthesis (CSVS) task as an extension of traditional CSS. By leveraging multimodal dialogue context, it provides users with coherent audiovisual responses. To this end, we develop a CSVS system named UniTalker, which is a unified model that seamlessly integrates multimodal perception and multimodal rendering capabilities. Specifically, it leverages a large-scale language model to comprehensively understand multimodal cues in the dialogue context, including speaker, text, speech, and the talking-face animations. After that, it employs multi-task sequence prediction to first infer the target utterance's emotion and then generate empathetic speech and natural talking-face animations. To ensure that the generated speech-visual content remains consistent in terms of emotion, content, and duration, we introduce three key optimizations: 1) Designing a specialized neural landmark codec to tokenize and reconstruct facial expression sequences. 2) Proposing a bimodal speech-visual hard alignment decoding strategy. 3) Applying emotion-guided rendering during the generation stage. Comprehensive objective and subjective experiments demonstrate that our model synthesizes more empathetic speech and provides users with more natural and emotionally consistent talking-face animations. The source code and generated samples are available at: https://github.com/AI-S2-Lab/UniTalker. Yifan Hu 0004, Rui Liu 0008, Yi Ren 0006, Xiang Yin 0006, Haizhou Li 0001 |
ACM Multimedia | 5 |
| 2025 | EventLip: Enhancing Event-Based Lip Reading via Frequency-Aware Spatiotemporal Hypergraph ModelingabstractEvent cameras, with their microsecond-level temporal resolution and sparse visual encoding, provide a transformative paradigm for automatic lip reading (ALR). However, event data inherently lack explicit spatial structure and exhibit a pronounced frequency-domain bias. The low-frequency components fail to capture crucial lip structural information, which fundamentally impedes the modeling of intra-frame topological dependencies and inter-frame semantic evolution-both of which are critical for robust lip reading. To this end, we propose FAST-HG, a Frequency-Aware SpatioTemporal HyperGraph framework specifically designed for event-based lip reading. First, we apply low-frequency perturbation to improve the model's robustness for capturing discriminative features, and integrate adaptive high-frequency filtering to enhance edge-aware representations. Then, we construct a Spatial Region Hypergraph (SRH) and a Temporal Semantic Hypergraph (TSH). The former captures intra-frame topological dependencies among lip regions, while the latter explicitly models inter-frame structural associations throughout the lip movement process, enabling the model to capture discriminative patterns in lip dynamics. Furthermore, we propose a viseme-aware label smoothing strategy, where a novel viseme-level edit distance is designed to quantify visual similarities between classes and guide the construction of soft labels. FAST-HG achieves 79.85% and 84.03% accuracy on the DVS-Lip and DVS-LRW100 datasets, respectively, significantly outperforming prior methods and establishing a new benchmark for event-based lip reading. Xueyi Zhang 0001, Jialu Sun, Xianghu Yue, Tianfang Xiao, Siqi Cai 0002, Mingrui Lao, Haizhou Li 0001 |
ACM Multimedia | 8 |
| 2025 | TrustCLIP: Learning from Noisy Labels via Semantic Label Verification and Trust-aligned Gradient ProjectionabstractPrompt learning has emerged as an efficient adaptation paradigm for vision-language models (VLMs), yet it remains highly vulnerable to label noise, which limits its real-world applicability. We propose TrustCLIP, a noise-robust prompt tuning framework that leverages the inherent semantic structure of CLIP through two key components: Semantic Label Verification (SLV) and Trust-aligned Gradient Projection (TGP). SLV defines a semantic trust boundary based on CLIP's zero-shot predictions to identify reliable samples for standard supervised training. For uncertain samples, TGP projects their gradients into a trust-aligned subspace constructed from the gradients of clean samples, thereby preserving semantically aligned learning signals while suppressing noise-induced optimization drift. Unlike prior approaches, TrustCLIP doesn't require additional parameters, loss reweighting, or uncertainty estimation. Extensive experiments on 7 benchmark datasets with both synthetic and real-world noisy labels demonstrate that TrustCLIP consistently outperforms state-of-the-art methods in terms of both robustness and transferability. Xueyi Zhang 0001, Peiyin Zhu, Mingrui Lao, Siqi Cai 0002, Yanming Guo, Haizhou Li 0001 |
ACM Multimedia | 8 |
| 2025 | Take the essence and discard the dross: A Rethinking on Data Selection for Fine-Tuning Large Language ModelsabstractZiche Liu, Rui Ke, Yajiao Liu, Feng Jiang, Haizhou Li. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Ziche Liu, Rui Ke, Yajiao Liu, Feng Jiang 0007, Haizhou Li 0001 |
NAACL (Long Papers) | 5 |
| 2025 | Does Mapo Tofu Contain Coffee? Probing LLMs for Food-related Cultural KnowledgeabstractLi Zhou, Taelin Karidi, Wanlong Liu, Nicolas Garneau, Yong Cao, Wenyu Chen, Haizhou Li, Daniel Hershcovich. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Li Zhou 0010, Taelin Karidi, Wanlong Liu, Nicolas Garneau, Yong Cao 0001, Wenyu Chen 0001, Haizhou Li 0001, Daniel Hershcovich |
NAACL (Long Papers) | 7 |
| 2025 | Bipolar Self-attention for Spiking TransformersabstractHarnessing the event-driven characteristic, Spiking Neural Networks (SNNs) present a promising avenue toward energy-efficient Transformer architectures. However, existing Spiking Transformers still suffer significant performance gaps compared to their Artificial Neural Network counterparts. Through comprehensive analysis, we attribute this gap to these two factors. First, the binary nature of spike trains limits Spiking Self-attention (SSA)’s capacity to capture negative–negative and positive–negative membrane potential interactions on Querys and Keys. Second, SSA typically omits Softmax functions to avoid energy-intensive multiply-accumulate operations, thereby failing to maintain row-stochasticity constraints on attention scores.
To address these issues, we propose a Bipolar Self-attention (BSA) paradigm, effectively modeling multi-polar membrane potential interactions with a fully spike-driven characteristic. Specifically, we demonstrate that ternary matrix multiplication provides a closer approximation to real-valued computation on both distribution and local correlation, enabling clear differentiation between homopolar and heteropolar interactions. Moreover, we propose a shift-based Softmax approximation named Shiftmax, which efficiently achieves low-entropy activation and partly maintains row-stochasticity without non-linear operation, enabling precise attention allocation.
Extensive experiments show that BSA achieves substantial performance improvements across various tasks, including image classification, semantic segmentation, and event-based tracking. These results establish its potential as a fundamental building block for energy-efficient Spiking Transformers. Shuai Wang 0058, Malu Zhang, Dehao Zhang, Yimeng Shan, Jieyuan Zhang, Yichen Xiao, Honglin Cao, Zeyu Ma 0002, Yang Yang 0002, Haizhou Li 0001 |
NeurIPS | 12 |
| 2025 | S2NN: Sub-bit Spiking Neural NetworksabstractSpiking Neural Networks (SNNs) offer an energy-efficient paradigm for machine intelligence, but their continued scaling poses challenges for resource-limited deployment. Despite recent advances in binary SNNs, the storage and computational demands remain substantial for large-scale networks. To further explore the compression and acceleration potential of SNNs, we propose Sub-bit Spiking Neural Networks (S$^2$NNs) that represent weights with less than one bit. Specifically, we first establish an S$^2$NN baseline by leveraging the clustering patterns of kernels in well-trained binary SNNs. This baseline is highly efficient but suffers from \textit{outlier-induced codeword selection bias} during training. To mitigate this issue, we propose an \textit{outlier-aware sub-bit weight quantization} (OS-Quant) method, which optimizes codeword selection by identifying and adaptively scaling outliers. Furthermore, we propose a \textit{membrane potential-based feature distillation} (MPFD) method, improving the performance of highly compressed S$^2$NN via more precise guidance from a teacher model. Extensive results on vision reveal that S$^2$NN outperforms existing quantized SNNs in both performance and efficiency, making it promising for edge computing applications. Wenjie Wei, Malu Zhang, Jieyuan Zhang, Ammar Belatreche, Shuai Wang 0058, Yimeng Shan, Honglin Cao, Guoqing Wang 0001, Yang Yang 0002, Haizhou Li 0001 |
NeurIPS | 11 |
| 2025 | SongBloom: Coherent Song Generation via Interleaved Autoregressive Sketching and Diffusion RefinementabstractGenerating music with coherent structure, harmonious instrumental and vocal elements remains a significant challenge in song generation. Existing language models and diffusion-based methods often struggle to balance global coherence with local fidelity, resulting in outputs that lack musicality or suffer from incoherent progression and mismatched lyrics. This paper introduces SongBloom, a novel framework for full-length song generation that leverages an interleaved paradigm of autoregressive sketching and diffusion-based refinement. SongBloom employs an autoregressive diffusion model that combines the high fidelity of diffusion models with the scalability of language models.
Specifically, it gradually extends a musical sketch from short to long and refines the details from coarse to fine-grained. The interleaved generation paradigm effectively integrates prior semantic and acoustic context to guide the generation process.
Experimental results demonstrate that SongBloom outperforms existing methods across both subjective and objective metrics and achieves performance comparable to the state-of-the-art commercial music generation platforms.
Audio samples are available on our demo page:
https://cypress-yang.github.io/SongBloom_demo. Shuai Wang 0016, Hangting Chen, Wei Tan 0011, Jianwei Yu 0001, Haizhou Li 0001 |
NeurIPS | 6 |
| 2025 | Listening to the Brain: Multi-Band sEEG Auditory Reconstruction via Dynamic Spatio-Temporal HypergraphsabstractSpeech is a fundamental form of human communication, and speech perception constitutes the initial stage of language comprehension. Although brain-to-speech interface technologies have made significant progress in recent years, most existing studies focus on neural decoding during speech production. Such approaches heavily rely on articulatory motor regions, rendering them unsuitable for individuals with speech motor impairments, such as those with aphasia or locked-in syndrome. To address this limitation, we construct and release NeuroListen, the first publicly available stereo-electroencephalography (sEEG) dataset specifically designed for auditory reconstruction. It contains over 10 hours of neural–speech paired recordings from 5 clinical participants, covering a wide range of semantic categories. Building on this dataset, we propose HyperSpeech, a multi-band neural decoding framework that employs dynamic spatio-temporal hypergraph neural networks to capture high-order dependencies across frequency, spatial, and temporal dimensions. Experimental results demonstrate that HyperSpeech significantly outperforms existing methods across multiple objective speech quality metrics, and achieves superior performance in human subjective evaluations, validating its effectiveness and advancement. This study provides a dedicated dataset and modeling framework for auditory speech decoding, offering foundations for neural language processing and assistive communication systems. Xueyi Zhang 0001, Ruicong Wang, Jialu Sun, Siqi Cai 0002, Haizhou Li 0001 |
NeurIPS | 5 |
| 2025 | Unveiling the Spatial-temporal Effective Receptive Fields of Spiking Neural NetworksabstractSpiking Neural Networks (SNNs) demonstrate significant potential for energy-efficient neuromorphic computing through an event-driven paradigm. While training methods and computational models have greatly advanced, SNNs struggle to achieve competitive performance in visual long-sequence modeling tasks. In artificial neural networks, the effective receptive field (ERF) serves as a valuable tool for analyzing feature extraction capabilities in visual long-sequence modeling. Inspired by this, we introduce the Spatio-Temporal Effective Receptive Field (ST-ERF) to analyze the ERF distributions across various Transformer-based SNNs. Based on the proposed ST-ERF, we reveal that these models suffer from establishing a robust global ST-ERF, thereby limiting their visual feature modeling capabilities. To overcome this issue, we propose two novel channel-mixer architectures: \underline{m}ulti-\underline{l}ayer-\underline{p}erceptron-based m\underline{ixer} (MLPixer) and \underline{s}plash-and-\underline{r}econstruct \underline{b}lock (SRB). These architectures enhance global spatial ERF through all timesteps in early network stages of Transformer-based SNNs, improving performance on challenging visual long-sequence modeling tasks. Extensive experiments conducted on the Meta-SDT variants and across object detection and semantic segmentation tasks further validate the effectiveness of our proposed method. Beyond these specific applications, we believe the proposed ST-ERF framework can provide valuable insights for designing and optimizing SNN architectures across a broader range of tasks. The code is available at \href{https://github.com/EricZhang1412/Spatial-temporal-ERF}{\faGithub~EricZhang1412/Spatial-temporal-ERF}. Jieyuan Zhang, Shuai Wang 0058, Wenjie Wei, Qian Sun 0014, Malu Zhang, Yang Yang 0002, Haizhou Li 0001 |
NeurIPS | 9 |
| 2025 | Dendritic Resonate-and-Fire Neuron for Effective and Efficient Long Sequence ModelingabstractThe explosive growth in sequence length has intensified the demand for effective and efficient long sequence modeling. Benefiting from intrinsic oscillatory membrane dynamics, Resonate-and-Fire (RF) neurons can efficiently extract frequency components from input signals and encode them into spatiotemporal spike trains, making them well-suited for long sequence modeling. However, RF neurons exhibit limited effective memory capacity and a trade-off between energy efficiency and training speed on complex temporal tasks. Inspired by the dendritic structure of biological neurons, we propose a Dendritic Resonate-and-Fire (D-RF) model, which explicitly incorporates a multi-dendritic and soma architecture. Each dendritic branch encodes specific frequency bands by utilizing the intrinsic oscillatory dynamics of RF neurons, thereby collectively achieving comprehensive frequency representation. Furthermore, we introduce an adaptive threshold mechanism into the soma structure. This mechanism adjusts the firing threshold according to historical spiking activity, thereby reducing redundant spikes while maintaining training efficiency in long-sequence tasks. Extensive experiments demonstrate that our method maintains competitive accuracy while substantially ensuring sparse spikes without compromising computational efficiency during training. These results underscore its potential as an effective and efficient solution for long sequence modeling on edge platforms. Dehao Zhang, Malu Zhang, Shuai Wang 0058, Wenjie Wei, Zeyu Ma 0002, Guoqing Wang 0001, Yang Yang 0002, Haizhou Li 0001 |
NeurIPS | 9 |
| 2025 | Boosting Discriminability for Robust Multimodal Entity Linking with Visual Modality MissingabstractMultimodal Entity Linking (MEL) aims to retrieve ambiguous mentions within multimodal contexts to the referent entities in a multimodal knowledge base, typically based on the assumption of modality completeness. However, when deployed in open-world applications, MEL systems may encounter uncertainly missing of visual modalities from user-proposed mentions. In this paper, we propose a novel setting dubbed MEL-MM to simulate the practical challenge, and reveal that the semantic discriminability is a crucial factor to enhance the anti-missingness resilience. To this end, we introduce an innovative yet efficient approach termed Cross-View Introspective Ranking Distillation (CVIRD), which seeks to sufficiently align the linking similarities between teacher and student models trained from modality-complete and incomplete data. To be specific, as the first concept in CVIRD, Missing-Aware Ranking Distillation (MARD) focuses on modeling the discriminability by formulating the similarity rankings between mention and entities in a missing-sensitive and differentiable manner. Moreover, the second concept of Cross-View Distillation with Introspection (CVDI) aims to improve discriminability extraction in MARD through multi-level distillation, considering both cross-view retrieval and self-consistency. Experiments verify the effectiveness and model-agnostic ability of our method, which achieves superior performance in contrast to competitive missingness-resilient strategies. Mingrui Lao, Yanming Guo, Xueyi Zhang 0001, Siqi Cai 0002, Zhaoyun Ding, Haizhou Li 0001 |
SIGIR | 7 |
| 2025 | Beyond Binary: Towards Fine-Grained LLM-Generated Text Detection via Role Recognition and Involvement MeasurementabstractThe rapid development of large language models (LLMs), like ChatGPT, has resulted in the widespread presence of LLM-generated content on social media platforms, raising concerns about misinformation, data biases, and privacy violations, which can undermine trust in online discourse. While detecting LLM-generated content is crucial for mitigating these risks, current methods often focus on binary classification, failing to address the complexities of real-world scenarios like human-LLM collaboration. To move beyond binary classification and address these challenges, we propose a new paradigm for detecting LLM-generated content. This approach introduces two novel tasks: LLM Role Recognition (LLM-RR), a multi-class classification task that identifies specific roles of an LLM in content generation, and LLM Involvement Measurement (LLM-IM), a regression task that quantifies the extent of LLM involvement in content creation. To support these tasks, we propose LLMDetect, a benchmark designed to evaluate detectors' performance on these new tasks. LLMDetect includes the Hybrid News Detection Corpus (HNDC) for training detectors, as well as DetectEval, a comprehensive evaluation suite that considers five distinct cross-context variations and two multi-intensity variations within the same LLM role. This allows for a thorough assessment of detectors' generalization and robustness across diverse contexts. Our empirical validation of 10 baseline detection methods demonstrates that fine-tuned Pre-trained Language Model (PLM)-based models consistently outperform others on both tasks, while advanced LLMs face challenges in accurately detecting their own generated content. Our experimental results and analysis offer insights for developing more effective detection models for LLM-generated content. This research enhances the understanding of LLM-generated content and establishes a foundation for more nuanced detection methodologies. Li Zhou 0010, Feng Jiang 0007, Benyou Wang, Haizhou Li 0001 |
WWW | 5 |
| 2025 | Ego4D: Around the World in 3,600 Hours of Egocentric VideoabstractWe introduce Ego4D, a massive-scale egocentric video dataset and benchmark suite. It offers 3,670 hours of daily-life activity video spanning hundreds of scenarios (household, outdoor, workplace, leisure, etc.) captured by 931 unique camera wearers from 74 worldwide locations and 9 different countries. The approach to collection is designed to uphold rigorous privacy and ethics standards, with consenting participants and robust de-identification procedures where relevant. Ego4D dramatically expands the volume of diverse egocentric video footage publicly available to the research community. Portions of the video are accompanied by audio, 3D meshes of the environment, eye gaze, stereo, and/or synchronized videos from multiple egocentric cameras at the same event. Furthermore, we present a host of new benchmark challenges centered around understanding the first-person visual experience in the past (querying an episodic memory), present (analyzing hand-object manipulation, audio-visual conversation, and social interactions), and future (forecasting activities). By publicly sharing this massive annotated dataset and benchmark suite, we aim to push the frontier of first-person perception. Kristen Grauman, Andrew Westbury, Eugene Byrne, Vincent Cartillier, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang 0007, Devansh Kukreja, Miao Liu 0007, Xingyu Liu 0001, Tushar Nagarajan, Ilija Radosavovic, Santhosh K. Ramakrishnan, Fiona Ryan, Jayant Sharma 0002, Michael Wray, Mengmeng Xu 0006, Eric Zhongcong Xu, Chen Zhao 0002, Siddhant Bansal, Dhruv Batra, Sean Crane, Tien Do, Morrie Doulaty, Akshay Erapalli, Christoph Feichtenhofer, Adriano Fragomeni, Qichen Fu, Abrham Gebreselasie, Cristina González, James Hillis, Xuhua Huang, Yifei Huang 0002, Wenqi Jia 0001, Weslie Khoo, Jáchym Kolár, Satwik Kottur, Anurag Kumar 0003, Federico Landini, Yanghao Li, Zhenqiang Li 0002, Karttikeya Mangalam, Raghava Modhugu, Jonathan Munro, Tullie Murrell, Takumi Nishiyasu, Will Price, Paola Ruiz Puentes, Merey Ramazanova, Leda Sari, Kiran K. Somasundaram, Audrey Southerland, Yusuke Sugano, Ruijie Tao, Minh Vo, Xindi Wu, Takuma Yagi, Ziwei Zhao 0003, Yunyi Zhu, Pablo Andrés Arbeláez, David Crandall, Dima Damen, Giovanni Maria Farinella, Christian Fügen, Bernard Ghanem, Vamsi K. Ithapu, C. V. Jawahar, Hanbyul Joo, Kris Makoto Kitani, Haizhou Li 0001, Richard A. Newcombe, Aude Oliva, Hyun Soo Park, James M. Rehg, Yoichi Sato 0001, Jianbo Shi, Zheng Shou 0001, Antonio Torralba 0001, Lorenzo Torresani, Mingfei Yan, Jitendra Malik |
IEEE Trans. Pattern Anal. Mach. Intell. | 75 |
| 2025 | ExPO: Explainable Phonetic Trait-Oriented Network for Speaker VerificationabstractIn speaker verification, we use computational method to verify if an utterance matches the identity of an enrolled speaker. This task is similar to the manual task of forensic voice comparison, where linguistic analysis is combined with auditory measurements to compare and evaluate voice samples. Despite much success, we have yet to develop a speaker verification system that offers explainable results comparable to those from manual forensic voice comparison. A novel approach, Explainable Phonetic Trait-Oriented (ExPO) network, is proposed in this letter to introduce the speaker'sphonetic traitwhich describes the speaker's characteristics at the phonetic level, resembling what forensic comparison does. ExPO not only generates utterance-level speaker embeddings but also allows for fine-grained analysis and visualization of phonetic traits, offering an explainable speaker verification process. Furthermore, we investigate phonetic traits from within-speaker and between-speaker variation perspectives to determine which trait is most effective for speaker verification, marking an important step towards explainable speaker verification. Shuai Wang 0016, Tianchi Liu 0004, Haizhou Li 0001 |
IEEE Signal Process. Lett. | 4 |
| 2025 | Analytic Class Incremental Learning for Sound Source Localization With Privacy ProtectionabstractSound Source Localization (SSL) enabling technology for applications such as surveillance and robotics. While traditional Signal Processing (SP)-based Sound Source Localization (SSL) methods provide analytic solutions under specific signal and noise assumptions, recent Deep Learning (DL)-based methods have significantly outperformed them. However, their success depends on extensive training data and substantial computational resources. Moreover, they often rely on large-scale annotated spatial data and may struggle to adapt to evolving sound classes. To mitigate these challenges, we propose a novel Class Incremental Learning (CIL) approach, termed SSL-CIL, which avoids serious accuracy degradation due tocatastrophic forgettingby incrementally updating the DL-based SSL model through a closed-form analytic solution. In particular, data privacy is ensured since the learning process does not revisit any historical data (exemplar-free), which is more suitable for smart home scenarios. Empirical results in the public SSLR dataset demonstrate the superior performance of our proposal, achieving a localization accuracy of 90.9%, surpassing other competitive methods. Xinyuan Qian 0001, Xianghu Yue, Huiping Zhuang, Haizhou Li 0001 |
IEEE Signal Process. Lett. | 5 |
| 2025 | Hierarchical Control of Emotion Rendering in Speech SynthesisabstractEmotional text-to-speech synthesis (TTS) aims to generate realistic emotional speech from input text. However, quantitatively controlling multi-level emotion rendering remains challenging. In this paper, we propose a flow-matching based emotional TTS framework with a novel approach for emotion intensity modeling to facilitate fine-grained control over emotion rendering at the phoneme, word, and utterance levels. We introduce a hierarchical emotion distribution (ED) extractor that captures a quantifiable ED embedding across different speech segment levels. Additionally, we explore various acoustic features and assess their impact on emotion intensity modeling. During TTS training, the hierarchical ED embedding effectively captures the variance in emotion intensity from the reference audio and correlates it with linguistic and speaker information. The TTS model not only generates emotional speech during inference, but also quantitatively controls the emotion rendering over the speech constituents. Both objective and subjective evaluations demonstrate the effectiveness of our framework in terms of speech quality, emotional expressiveness, and hierarchical emotion control. Sho Inoue, Kun Zhou 0003, Shuai Wang 0016, Haizhou Li 0001 |
IEEE Trans. Affect. Comput. | 4 |
| 2025 | Connecting Cross-Modal Representations for Compact and Robust Multimodal Sentiment Analysis With Sentiment Word Substitution ErrorabstractMultimodal Sentiment Analysis (MSA) seeks to fuse textual, acoustic, and visual information to predict a speaker’s sentiment states effectively. However, in real-world scenarios, the text modality received by MSA systems is often obtained through automatic speech recognition (ASR) models. Unfortunately, ASR may erroneously recognize sentiment words as phonetically similar neutral alternatives, leading to sentiment degradation in text and impacting MSA accuracy. Recent attempts aim to first identify the sentiment word substitution (SWS) error in ASR results and then refine the corrupted word embeddings using multimodal information for final multimodal fusion. However, such a method includes a burdensome and ambiguous detection operation and ignores the inherent correlations and heterogeneity among different modalities. To address these issues, we propose a more compact system, termedARF-MSAconsisting of three key components to achieving robust MSA with SWS errors: 1)Alignment: we establish connections between the “text-acoustic’ and “text-visual” representations to effectively map the “text-acoustic-visual” data into a unified sentiment space by leveraging their multimodal correlation knowledge; 2)Refinement: we perform fine-grained comparisons between the text modality and the other two modalities in the unified sentiment space, enabling refinement of the sentiment expression within the text modality more concisely; 3)Fusion: Finally, we hierarchically fuse the dominant and non-dominant representation from three heterogeneity modalities to obtain the multimodal feature for MSA. We conduct extensive experiments on the real-world datasets and the results demonstrate the effectiveness of our model. Qiyuan Sun, Haolin Zuo, Rui Liu 0008, Haizhou Li 0001 |
IEEE Trans. Affect. Comput. | 4 |
| 2025 | Human-Inspired Computing for Robust and Efficient Audio-Visual Speech RecognitionabstractHumans excel at audiovisual speech recognition (AVSR), motivating the development of human-inspired computing for robust and efficient AVSR models. Spiking neural networks (SNNs), mimicking the brain’s information-processing mechanisms, offer a promising foundation. However, research on SNN-based AVSR remains limited, with most audio-visual methods focusing on object or digit recognition. These methods oversimplify multimodal fusion, neglecting modality-specific characteristics and interactions. Additionally, they often rely on future information, increasing recognition latency and limiting real-time applicability. Inspired by human speech perception, this paper proposes a novel human-inspired SNN named HI-AVSNN for AVSR, incorporating three computing characteristics: spike activity, cueing interaction, and causal processing. For cueing interaction, we introduce a Spike-Driven Visual-Cued Speech Processing (sVCSP) scheme, where visual features hierarchically guide speech processing to enhance critical features. For causal processing, we align the temporal dimensions of SNN with audio-visual inputs and apply temporal masking to ensure only past and current information is used. For spike activity, in addition to SNNs, we incorporate event cameras to capture lip movements as spikes, efficiently encoding visual data like the human retina. Experiments on two event-based AVSR datasets demonstrate our method outperforms existing audio-visual SNN fusion techniques, showcasing the effectiveness, robustness, and efficiency achieved through our human-inspired computing. Qianhui Liu, Yang Wang 0106, Xin Yang 0011, Gang Pan 0001, Haizhou Li 0001 |
IEEE Trans. Computers | 6 |
| 2025 | Universal Low Bit-Rate Speech Steganalysis Integrating Domain-Specific and Domain-Shared KnowledgeabstractUniversal low bit-rate speech steganalysis is a cutting-edge research task addressing real-world application needs and has garnered significant attention recently. However, the existing methods are still inadequate in extracting available information from various steganographic domains and fail to deliver interpretable forensic results for specific steganographic domains containing embedded information. In view of this, we present a novel universal low bit-rate speech steganalysis approach that seamlessly combines domain-specific and domain-shared information, enabling comprehensive and effective speech steganography detection. This approach comprises two vital components: the Matching Identification Network (MIN) and the Content Alignment Network (CAN). The MIN incorporates three effective separable backbones for capturing informative domain-specific embeddings, inherently unveiling local and global dependencies. In this network, we also design a cross-domain matching module to establish correlations among steganographic domains, thereby enhancing detection performance through multi-domain collaboration and facilitating effective forensics for the embedded domains. Moreover, the CAN acquires more informative domain-shared embeddings by using a metric learning-based Siamese architecture to process pairs of naive and recompressed speech samples. Experimental results demonstrate that the presented method not only significantly surpasses the existing universal steganalysis methods, but also competes with or even surpasses dedicated steganalysis methods in certain cases. In addition, our method can provide accurate forensic results regarding the existence of hidden information within each steganographic domain without relevant supervisory information, marking a significant milestone in pursuit of speech steganalysis. The source code for this work will be publicly available on GitHub. Hui Tian 0002, Yiqin Qiu, Haizhou Li 0001, Xinpeng Zhang 0001, Athanasios V. Vasilakos |
IEEE Trans. Dependable Secur. Comput. | 3 |
| 2025 | Nes2Net: A Lightweight Nested Architecture for Foundation Model Driven Speech Anti-SpoofingabstractSpeech foundation models have significantly advanced various speech-related tasks by providing exceptional representation capabilities. However, their high-dimensional output features often create a mismatch with downstream task models, which typically require lower-dimensional inputs. A common solution is to apply a dimensionality reduction (DR) layer, but this approach increases parameter overhead, computational costs, and risks losing valuable information. To address these issues, we propose Nested Res2Net (Nes2Net), a lightweight back-end architecture designed to directly process high-dimensional features without DR layers. The nested structure enhances multi-scale feature extraction, improves feature interaction, and preserves high-dimensional information. We first validate Nes2Net on CtrSVDD, a singing voice deepfake detection dataset, and report a 22% performance improvement and an 87% back-end computational cost reduction over the state-of-the-art baseline. Additionally, extensive testing across four diverse datasets: ASVspoof 2021, ASVspoof 5, PartialSpoof, and In-the-Wild, covering fully spoofed speech, adversarial attacks, partial spoofing, and real-world scenarios, consistently highlights Nes2Net’s superior robustness and generalization capabilities. The code package and pre-trained models are available at https://github.com/Liu-Tianchi/Nes2Net. Tianchi Liu 0004, Duc-Tuan Truong, Rohan Kumar Das, Kong-Aik Lee, Haizhou Li 0001 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2025 | Spiking Neural Networks With Adaptive Membrane Time Constant for Event-Based TrackingabstractThe brain-inspired Spiking Neural Networks (SNNs) work in an event-driven manner and have an implicit recurrence in neuronal membrane potential to memorize information over time, which are inherently suitable to handle temporal event-based streams. Despite their temporal nature and recent approaches advancements, these methods have predominantly been assessed on event-based classification tasks. In this paper, we explore the utility of SNNs for event-based tracking tasks. Specifically, we propose a brain-inspired adaptive Leaky Integrate-and-Fire neuron (BA-LIF) that can adaptively adjust the membrane time constant according to the inputs, thereby accelerating the leakage of meaningless noise features and reducing the decay of valuable information. SNNs composed of our proposed BA-LIF neurons can achieve high performance without a careful and time-consuming trial-by-error initialization on the membrane time constant. The adaptive capability of our network is further improved by introducing an extra temporal feature aggregator (TFA) that assigns attention weights over the temporal dimension. Extensive experiments on various event-based tracking datasets validate the effectiveness of our proposed method. We further validate the generalization capability of our method by applying it to other event-classification tasks. Jiqing Zhang, Malu Zhang, Yuanchen Wang, Qianhui Liu, Haizhou Li 0001, Xin Yang 0011 |
IEEE Trans. Image Process. | 6 |
| 2025 | Enhancing Real-World Active Speaker Detection With Multi-Modal Extraction Pre-TrainingabstractAudio-visual active speaker detection (AV-ASD) aims to identify which visible face is speaking in a scene with one or more persons. Most existing AV-ASD methods prioritize capturing speech-lip correspondence. However, there is a noticeable gap in addressing the challenges from real-world AV-ASD scenarios. Due to the presence of low-quality noisy videos in such cases, AV-ASD systems without a selective listening ability are short of effectively filtering out disruptive voice components from mixed audio inputs. In this paper, we propose a Multi-modal Speech Extraction-to-Detection framework named ‘MuSED’, which is pre-trained with audio-visual target speech extraction to learn the denoising ability, then it is fine-tuned with the AV-ASD task. Meanwhile, to better capture the multi-modal information and deal with real-world problems such as missing modality, MuSED is modelled on the time domain directly and integrates the multi-modal plus-and-minus augmentation strategy. Our experiments demonstrate that MuSED substantially outperforms the state-of-the-art AV-ASD methods and achieves 95.6% mAP on the AVA-ActiveSpeaker dataset, 98.3% AP on the ASW dataset, and 97.9% F1 on the Columbia AV-ASD dataset, respectively. We will publicly release the code in due course. Ruijie Tao, Xinyuan Qian 0001, Rohan Kumar Das, Xiaoxue Gao, Haizhou Li 0001 |
IEEE Trans. Multim. | 6 |
| 2025 | Tunnel Prescribed Control of Nonlinear Systems With Unknown Control DirectionsabstractThis article solves the entry capture problem (ECP) such that for any initial tracking error, it can be regulated into the prescribed performance constraints within a user-given time. The challenge lies in how to remove the initial condition limitation and to handle the ECP for nonlinear systems under unknown control directions and asymmetric performance constraints. For better tracking performance, we propose a unified tunnel prescribed performance (TPP) providing strict and tight allowable set. With the aid of a scaling function, error self-tuning functions (ESFs) are then developed to make the control scheme suitable to any initial condition (including the initial constraint violation), where the initial values of ESFs always satisfy performance constraints. In lieu of the Nussbaum technique, an orientation function is introduced to deal with unknown control directions while such way is capable of reducing the control peaking problem. Using ESFs, together with TPP and an orientation function, the resulted tunnel prescribed control (TPC) leads to a solution for the underlying ECP, which also exhibits a low complexity level since no command filters or dynamic surface control is required. Finally, simulation results are provided to further demonstrate these theoretical findings. Ruihang Ji, Dongyu Li, Shuzhi Sam Ge, Haizhou Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2025 | Zero-Shot Relation Classification Through Inference on Category AttributesabstractThe goal of relationship classification (RC) is to predict the semantic relationship between two entities in a given sentence. With the advent of deep learning and pretrained language models, RC research has progressed by leaps and bounds. However, the current studies are focused mainly on predicting semantic relationships from a predefined set. How to recognize unseen relationships remains a challenge, which is also known as the zero-shot RC (ZSRC) task. Some ZSRC-related methods directly map relationship categories to numerical indices, constraining the model's ability to autonomously infer and understand these relationships, while others rely heavily on manual definitions. To address these issues and inspired by the way of reasoning in which humans perform RC tasks, we propose a new framework to handle the ZSRC task through inference on category attributes (ICAs). The main idea of ICA is to detect the semantic relationship between promises, which are RC sentences, and hypotheses, which are relational sentences of entities created by templates. Specifically, instead of manual design, we introduce two hypothesis templates derived from the label words (LWs) and descriptions (LDs) associated with each relationship. These templates are used to automatically convert the RC data into the textual entailment (TE) format. Furthermore, they are fine-tuned with a pretrained TE model, facilitating the acquisition of relational knowledge and enabling the generalization of semantic reasoning rules learned from seen classes to unseen classes. Moreover, to implement multirelationship semantic inference for all unseen classes, we propose an entailment difference mechanism to enhance the reasoning capability of the model. Besides the current ZSRC test setting, we also examine our method in an even more challenging setting to deal with data scarcity in real-world applications. The outstanding performance of ICA on the FewRel and Wiki-ZSL datasets demonstrates its effectiveness in the ZSRC task. Yaochu Jin, Bin Wang 0040, Yan Zhang 0004, Kuangrong Hao, Haizhou Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2025 | Toward Building Human-Like Sequential Memory Using Brain-Inspired Spiking Neural ModelsabstractThe brain is able to acquire and store memories of everyday experiences in real-time. It can also selectively forget information to facilitate memory updating. However, our understanding of the underlying mechanisms and coordination of these processes within the brain remains limited. However, no existing artificial intelligence models have yet matched human-level capabilities in terms of memory storage and retrieval. This study introduces a brain-inspired spiking neural model that integrates the learning and forgetting processes of sequential memory. The proposed model closely mimics the distributed and sparse temporal coding observed in the biological neural system. It employs one-shot online learning for memory formation and uses biologically plausible mechanisms of neural oscillation and phase precession to retrieve memorized sequences reliably. In addition, an active forgetting mechanism is integrated into the spiking neural model, enabling memory removal, flexibility, and updating. The proposed memory model not only enhances our understanding of human memory processes but also provides a robust framework for addressing temporal modeling tasks. Malu Zhang, Xiaoling Luo 0001, Jibin Wu, Ammar Belatreche, Siqi Cai 0002, Yang Yang 0002, Haizhou Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2024 | Emotion Rendering for Conversational Speech Synthesis with Heterogeneous Graph-Based Context ModelingabstractConversational Speech Synthesis (CSS) aims to accurately express an utterance with the appropriate prosody and emotional inflection within a conversational setting. While recognising the significance of CSS task, the prior studies have not thoroughly investigated the emotional expressiveness problems due to the scarcity of emotional conversational datasets and the difficulty of stateful emotion modeling. In this paper, we propose a novel emotional CSS model, termed ECSS, that includes two main components: 1) to enhance emotion understanding, we introduce a heterogeneous graph-based emotional context modeling mechanism, which takes the multi-source dialogue history as input to model the dialogue context and learn the emotion cues from the context; 2) to achieve emotion rendering, we employ a contrastive learning-based emotion renderer module to infer the accurate emotion style for the target utterance. To address the issue of data scarcity, we meticulously create emotional labels in terms of category and intensity, and annotate additional emotional information on the existing conversational dataset (DailyTalk). Both objective and subjective evaluations suggest that our model outperforms the baseline models in understanding and rendering emotions. These evaluations also underscore the importance of comprehensive emotional annotations. Code and audio samples can be found at: https://github.com/walker-hyf/ECSS. Rui Liu 0008, Yifan Hu 0004, Yi Ren 0006, Xiang Yin 0006, Haizhou Li 0001 |
AAAI | 5 |
| 2024 | Restoring Speaking Lips from Occlusion for Audio-Visual Speech RecognitionabstractPrior studies on audio-visual speech recognition typically assume the visibility of speaking lips, ignoring the fact that visual occlusion occurs in real-world videos, thus adversely affecting recognition performance. To address this issue, we propose a framework that restores occluded lips in a video by utilizing both the video itself and the corresponding noisy audio. Specifically, the framework aims to achieve these three tasks: detecting occluded frames, masking occluded areas, and reconstruction of masked regions. We tackle the first two issues by utilizing the Class Activation Map (CAM) obtained from occluded frame detection to facilitate the masking of occluded areas. Additionally, we introduce a novel synthesis-matching strategy for the reconstruction to ensure the compatibility of audio features with different levels of occlusion. Our framework is evaluated in terms of Word Error Rate (WER) on the original videos, the videos corrupted by concealed lips, and the videos restored using the framework with several existing state-of-the-art audio-visual speech recognition methods. Experimental results substantiate that our framework significantly mitigates performance degradation resulting from lip occlusion. Under -5dB noise conditions, AV-Hubert's WER increases from 10.62% to 13.87% due to lip occlusion, but rebounds to 11.87% in conjunction with the proposed framework. Furthermore, the framework also demonstrates its capacity to produce natural synthesized images in qualitative assessments. Zexu Pan, Malu Zhang, Robby T. Tan, Haizhou Li 0001 |
AAAI | 5 |
| 2024 | A Comprehensive Analysis of the Effectiveness of Large Language Models as Automatic Dialogue EvaluatorsabstractAutomatic evaluation is an integral aspect of dialogue system research. The traditional reference-based NLG metrics are generally found to be unsuitable for dialogue assessment. Consequently, recent studies have suggested various unique, reference-free neural metrics that better align with human evaluations. Notably among them, large language models (LLMs), particularly the instruction-tuned variants like ChatGPT, are shown to be promising substitutes for human judges. Yet, existing works on utilizing LLMs for automatic dialogue evaluation are limited in their scope in terms of the number of meta-evaluation datasets, mode of evaluation, coverage of LLMs, etc. Hence, it remains inconclusive how effective these LLMs are. To this end, we conduct a comprehensive study on the application of LLMs for automatic dialogue evaluation. Specifically, we analyze the multi-dimensional evaluation capability of 30 recently emerged LLMs at both turn and dialogue levels, using a comprehensive set of 12 meta-evaluation datasets. Additionally, we probe the robustness of the LLMs in handling various adversarial perturbations at both turn and dialogue levels. Finally, we explore how model-level and dimension-level ensembles impact the evaluation performance. All resources are available at https://github.com/e0397123/comp-analysis. Chen Zhang 0020, Luis Fernando D'Haro, Yiming Chen 0010, Malu Zhang, Haizhou Li 0001 |
AAAI | 5 |
| 2024 | TC-LIF: A Two-Compartment Spiking Neuron Model for Long-Term Sequential ModellingabstractThe identification of sensory cues associated with potential opportunities and dangers is frequently complicated by unrelated events that separate useful cues by long delays. As a result, it remains a challenging task for state-of-the-art spiking neural networks (SNNs) to establish long-term temporal dependency between distant cues. To address this challenge, we propose a novel biologically inspired Two-Compartment Leaky Integrate-and-Fire spiking neuron model, dubbed TC-LIF. The proposed model incorporates carefully designed somatic and dendritic compartments that are tailored to facilitate learning long-term temporal dependencies. Furthermore, the theoretical analysis is provided to validate the effectiveness of TC-LIF in propagating error gradients over an extended temporal duration. Our experimental results, on a diverse range of temporal classification tasks, demonstrate superior temporal classification capability, rapid training convergence, and high energy efficiency of the proposed TC-LIF model. Therefore, this work opens up a myriad of opportunities for solving challenging temporal processing tasks on emerging neuromorphic computing systems. Our code is publicly available at https://github.com/ZhangShimin1/TC-LIF. Qu Yang, Chenxiang Ma, Jibin Wu, Haizhou Li 0001, Kay Chen Tan |
AAAI | 5 |
| 2024 | Uncovering the Potential of ChatGPT for Discourse Analysis in Dialogue: An Empirical StudyabstractLarge language models, like ChatGPT, have shown remarkable capability in many downstream tasks, yet their ability to understand discourse structures of dialogues remains less explored, where it requires higher level capabilities of understanding and reasoning. In this paper, we aim to systematically inspect ChatGPT’s performance in two discourse analysis tasks: topic segmentation and discourse parsing, focusing on its deep semantic understanding of linear and hierarchical discourse structures underlying dialogue. To instruct ChatGPT to complete these tasks, we initially craft a prompt template consisting of the task description, output format, and structured input. Then, we conduct experiments on four popular topic segmentation datasets and two discourse parsing datasets. The experimental results showcase that ChatGPT demonstrates proficiency in identifying topic structures in general-domain conversations yet struggles considerably in specific-domain conversations. We also found that ChatGPT hardly understands rhetorical structures that are more complex than topic structures. Our deeper investigation indicates that ChatGPT can give more reasonable topic structures than human annotations but only linearly parses the hierarchical rhetorical structures. In addition, we delve into the impact of in-context learning (e.g., chain-of-thought) on ChatGPT and conduct the ablation study on various prompt components, which can provide a research foundation for future work. The code is available at https://github.com/yxfanSuda/GPTforDDA. Yaxin Fan, Feng Jiang 0007, Peifeng Li 0001, Haizhou Li 0001 |
LREC/COLING | 4 |
| 2024 | Advancing Topic Segmentation and Outline Generation in Chinese Texts: The Paragraph-level Topic Representation, Corpus, and BenchmarkabstractTopic segmentation and outline generation strive to divide a document into coherent topic sections and generate corresponding subheadings, unveiling the discourse topic structure of a document. Compared with sentence-level topic structure, the paragraph-level topic structure can quickly grasp and understand the overall context of the document from a higher level, benefitting many downstream tasks such as summarization, discourse parsing, and information retrieval. However, the lack of large-scale, high-quality Chinese paragraph-level topic structure corpora restrained relative research and applications. To fill this gap, we build the Chinese paragraph-level topic representation, corpus, and benchmark in this paper. Firstly, we propose a hierarchical paragraph-level topic structure representation with three layers to guide the corpus construction. Then, we employ a two-stage man-machine collaborative annotation method to construct the largest Chinese Paragraph-level Topic Structure corpus (CPTS), achieving high quality. We also build several strong baselines, including ChatGPT, to validate the computability of CPTS on two fundamental tasks (topic segmentation and outline generation) and preliminarily verified its usefulness for the downstream task (discourse parsing). Feng Jiang 0007, Xiaomin Chu, Peifeng Li 0001, Qiaoming Zhu, Haizhou Li 0001 |
LREC/COLING | 6 |
| 2024 | CrossTune: Black-Box Few-Shot Classification with Label EnhancementabstractTraining or finetuning large-scale language models (LLMs) requires substantial computation resources, motivating recent efforts to explore parameter-efficient adaptation to downstream tasks. One approach is to treat these models as black boxes and use forward passes (Inference APIs) to interact with them. Current research focuses on adapting these black-box models to downstream tasks using gradient-free prompt optimization, but this often involves an expensive process of searching task-specific prompts. Therefore, we are motivated to study black-box language model adaptation without prompt search. Specifically, we introduce a label-enhanced cross-attention network called CrossTune, which models the semantic relatedness between the input text sequence and task-specific label descriptions. Its effectiveness is examined in the context of few-shot text classification. To improve the generalization of CrossTune, we utilize ChatGPT to generate additional training data through in-context learning. A switch mechanism is implemented to exclude low-quality ChatGPT-generated data. Through extensive experiments on seven benchmark text classification datasets, we demonstrate that our proposed approach outperforms the previous state-of-the-art gradient-free black-box tuning method by 5.7% on average. Even without using ChatGPT-augmented data, CrossTune performs better or comparably than previous black-box tuning methods, suggesting the effectiveness of our approach. Danqing Luo, Chen Zhang 0020, Yan Zhang 0004, Haizhou Li 0001 |
LREC/COLING | 4 |
| 2024 | DynaThink: Fast or Slow? A Dynamic Decision-Making Framework for Large Language ModelsabstractLarge language models (LLMs) have demonstrated emergent capabilities across diverse reasoning tasks via popular Chains-of-Thought (COT) prompting.However, such a simple and fast COT approach often encounters limitations in dealing with complicated problems, while a thorough method, which considers multiple reasoning pathways and verifies each step carefully, results in slower inference.This paper addresses the challenge of enabling LLMs to autonomously select between fast and slow inference methods, thereby optimizing both efficiency and effectiveness.We introduce a dynamic decision-making framework that categorizes tasks into two distinct pathways: 'Fast', designated for tasks where the LLM quickly identifies a high-confidence solution, and 'Slow', allocated for tasks that the LLM perceives as complex and for which it has low confidence in immediate solutions as well as requiring more reasoning paths to verify.Experiments on five popular reasoning benchmarks demonstrated the superiority of the Dyna-Think over baselines. Yan Zhang 0004, Chen Zhang 0020, Zuozhu Liu, Hongwei Wang 0001, Haizhou Li 0001 |
EMNLP | 6 |
| 2024 | Robust Decoding of the Auditory Attention from EEG Recordings Through Graph Convolutional NetworksabstractAuditory attention decoding (AAD) with electroencephalography (EEG) holds great promise in brain-computer interface (BCI). Despite much progress, it remains a research topic on how to effectively evaluate the performance of EEG-based AAD algorithms under an appropriate setting that reflects the use scenarios. It is desired that systems are evaluated under cross-subject and cross-trial settings. However, systems are often reported under same-subject, same-trial settings, where test data are not truly separated from the training data, thus potentially leading to model overfitting due to data leakage. In this paper, we study the robustness of graph convolutional network (GCN), a novel approach to learning the intricate spatial patterns in multi-channel EEG signals, by comparing GCN across cross-subject, cross-trial, and same-subject, same-trial settings. On two publicly available AAD datasets, it is found that GCN exhibits remarkable robustness, outperforming previous conventional convolutional neural network (CNN) solutions. We confirm the superiority of our GCN-based AAD model in terms of generalization and robustness. Siqi Cai 0002, Haizhou Li 0001 |
ICASSP | 3 |
| 2024 | LOCSELECT: Target Speaker Localization with an Auditory Selective Hearing MechanismabstractThe prevailing noise-resistant and reverberation-resistant localization algorithms primarily emphasize separating and providing directional output for each speaker in multi-speaker scenarios, without association with the identity of speakers. In this paper, we present a target speaker localization algorithm with a selective hearing mechanism. Given a reference speech of the target speaker, we first produce a speaker-dependent spectrogram mask to eliminate interfering speakers’ speech. Subsequently, a Long-Short-Term Memory (LSTM) network is employed to extract the target speaker’s location from the filtered spectrogram. Experiments validate the superiority of our proposed method over existing algorithms for different scale-invariant signal-to-noise ratios (SNR) conditions. Specifically, at SNR = -10 dB, our proposed network LocSelect achieves a mean absolute error (MAE) of 3.55° and an accuracy (ACC) of 87.40%. Xinyuan Qian 0001, Zexu Pan, Kainan Chen, Haizhou Li 0001 |
ICASSP | 5 |
| 2024 | Hierarchical Emotion Prediction and Control in Text-to-Speech SynthesisabstractIt remains a challenge to effectively control the emotion rendering in text-to-speech (TTS) synthesis. Prior studies have primarily focused on learning a global prosodic representation at the utterance level, which strongly correlates with linguistic prosody. Our goal is to construct a hierarchical emotion distribution (ED) that effectively encapsulates intensity variations of emotions at various levels of granularity, encompassing phonemes, words, and utterances. During TTS training, the hierarchical ED is extracted from the ground-truth audio and guides the predictor to establish a connection between emotional and linguistic prosody. At run-time inference, the TTS model generates emotional speech and, at the same time, provides quantitative control of emotion over the speech constituents. Both objective and subjective evaluations validate the effectiveness of the proposed framework in terms of emotion prediction and control. Sho Inoue, Kun Zhou 0003, Shuai Wang 0016, Haizhou Li 0001 |
ICASSP | 4 |
| 2024 | Prompt-Driven Target Speech DiarizationabstractWe introduce a novel task named ‘target speech diarization’, which seeks to determine ‘when target event occurred’ within an audio signal. We devise a neural architecture called Prompt-driven Target Speech Diarization (PTSD), that works with diverse prompts that specify the target speech events of interest. We train and evaluate PTSD using sim2spk, sim3spk and sim4spk datasets, which are derived from the Librispeech. We show that the proposed framework accurately localizes target speech events. Furthermore, our framework exhibits versatility through its impressive performance in three diarization-related tasks: target speaker voice activity detection, overlapped speech detection and gender diarization. In particular, PTSD achieves comparable performance to specialized models across these tasks on both real and simulated data. This work serves as a reference benchmark and provides valuable insights into prompt-driven target speech processing. Yidi Jiang, Zhengyang Chen, Ruijie Tao, Liqun Deng, Yanmin Qian, Haizhou Li 0001 |
ICASSP | 6 |
| 2024 | Audio-Visual Active Speaker Extraction for Sparsely Overlapped Multi-Talker SpeechabstractTarget speaker extraction aims to extract the speech of a specific speaker from a multi-talker mixture as specified by an auxiliary reference. Most studies focus on the scenario where the target speech is highly overlapped with the interfering speech. However, this scenario only accounts for a small percentage of real-world conversations. In this paper, we aim at the sparsely overlapped scenarios in which the auxiliary reference needs to perform two tasks simultaneously: detect the activity of the target speaker and disentangle the active speech from any interfering speech. We propose an audio-visual speaker extraction model named ActiveExtract, which leverages speaking activity from audio-visual active speaker detection (ASD). The ASD directly provides the frame-level activity of the target speaker, while its intermediate feature representation is trained to discriminate speech-lip synchronization that could be used for speaker disentanglement. Experimental results show our model outperforms baselines across various overlapping ratios, achieving an average improvement of more than 4 dB in terms of SI-SNR. Ruijie Tao, Zexu Pan, Meng Ge, Shuai Wang 0016, Haizhou Li 0001 |
ICASSP | 6 |
| 2024 | Gradient Weighting for Speaker Verification in Extremely Low Signal-to-Noise RatioabstractSpeaker verification is hampered by background noise, particularly at extremely low Signal-to-Noise Ratio (SNR) under 0 dB. It is difficult to suppress noise without introducing unwanted artifacts, which adversely affects speaker verification. We proposed the mechanism called Gradient Weighting (Grad-W), which dynamically identifies and reduces artifact noise during prediction. The mechanism is based on the property that the gradient indicates which parts of the input the model is paying attention to. Specifically, when the speaker network focuses on a region in the denoised utterance but not on the clean counterpart, we consider it artifact noise and assign higher weights for this region during optimization of enhancement. We validate it by training an enhancement model and testing the enhanced utterance on speaker verification. The experimental results show that our approach effectively reduces artifact noise, improving speaker verification across various SNR levels. Kong-Aik Lee, Ville Hautamäki, Meng Ge, Haizhou Li 0001 |
ICASSP | 5 |
| 2024 | Spiking-Leaf: A Learnable Auditory Front-End for Spiking Neural NetworksabstractBrain-inspired spiking neural networks (SNNs) have demonstrated great potential for temporal signal processing. However, their performance in speech processing remains limited due to the lack of an effective auditory front-end. To address this limitation, we introduce Spiking-LEAF, a learnable auditory front-end meticulously designed for SNN-based speech processing. Spiking-LEAF combines a learnable filter bank with a novel two-compartment spiking neuron model called IHC-LIF. The IHC-LIF neurons draw inspiration from the structure of inner hair cells (IHC) and they leverage segregated dendritic and somatic compartments to effectively capture multi-scale temporal dynamics of speech signals. Additionally, the IHC-LIF neurons incorporate the lateral feedback mechanism along with spike regularization loss to enhance spike encoding efficiency. On keyword spotting and speaker identification tasks, the proposed Spiking-LEAF outperforms both SOTA spiking auditory front-ends and conventional real-valued acoustic features in terms of classification accuracy, noise robustness, and encoding efficiency. Zeyang Song, Jibin Wu, Malu Zhang, Zheng Shou 0001, Haizhou Li 0001 |
ICASSP | 5 |
| 2024 | Leveraging in-the-wild Data for Effective Self-supervised Pretraining in Speaker RecognitionabstractCurrent speaker recognition systems primarily rely on supervised approaches, constrained by the scale of labeled datasets. To boost the system performance, researchers leverage large pretrained models such as WavLM to transfer learned high-level features to the downstream speaker recognition task. However, this approach introduces extra parameters as the pretrained model remains in the inference stage. Another group of researchers directly apply self-supervised methods such as DINO to speaker embedding learning, yet they have not explored its potential on large-scale in-the-wild datasets. In this paper, we present the effectiveness of DINO training on the large-scale WenetSpeech dataset and its transferability in enhancing the supervised system performance on the CNCeleb dataset. Additionally, we introduce a confidence-based data filtering algorithm to remove unreliable data from the pretraining dataset, leading to better performance with less training data. The associated pretrained models, confidence files, pretraining and finetuning scripts will be made available in the Wespeaker toolkit. Shuai Wang 0016, Qibing Bai, Qi Liu 0018, Jianwei Yu 0001, Zhengyang Chen, Bing Han 0008, Yanmin Qian, Haizhou Li 0001 |
ICASSP | 8 |
| 2024 | SVAD: A Robust, Low-Power, and Light-Weight Voice Activity Detection with Spiking Neural NetworksabstractSpeech applications are expected to be low-power and robust under noisy conditions. An effective Voice Activity Detection (VAD) front-end lowers the computational need. Spiking Neural Networks (SNNs) are known to be biologically plausible and power-efficient. However, SNN-based VADs have yet to achieve noise robustness and often require large models for high performance. This paper introduces a novel SNN-based VAD model, referred to as sVAD, which features an auditory encoder with an SNN-based attention mechanism. Particularly, it provides effective auditory feature representation through SincNet and 1D convolution, and improves noise robustness with attention mechanisms. The classifier utilizes Spiking Recurrent Neural Networks (sRNN) to exploit temporal speech information. Experimental results demonstrate that our sVAD achieves remarkable noise robustness and meanwhile maintains low power consumption and a small footprint, making it a promising solution for real-world VAD applications. Qu Yang, Qianhui Liu, Meng Ge, Zeyang Song, Haizhou Li 0001 |
ICASSP | 6 |
| 2024 | An Empirical Study on the Impact of Positional Encoding in Transformer-Based Monaural Speech EnhancementabstractTransformer architecture has enabled recent progress in speech enhancement. Since Transformers are position-agostic, positional encoding is the de facto standard component used to enable Transformers to distinguish the order of elements in a sequence. However, it remains unclear how positional encoding exactly impacts speech enhancement based on Transformer architectures. In this paper, we perform a comprehensive empirical study evaluating five positional encoding methods, i.e., Sinusoidal and learned absolute position embedding (APE), T5-RPE, KERPLE, as well as the Transformer without positional encoding (No-Pos), across both causal and noncausal configurations. We conduct extensive speech enhancement experiments, involving spectral mapping and masking methods. Our findings establish that positional encoding is not quite helpful for the models in a causal configuration, which indicates that causal attention may implicitly incorporate position information. In a noncausal configuration, the models significantly benefit from the use of positional encoding. In addition, we find that among the four position embeddings, relative position embeddings outperform APEs. Qiquan Zhang, Meng Ge, Hongxu Zhu, Eliathamby Ambikairajah, Zhaoheng Ni, Haizhou Li 0001 |
ICASSP | 7 |
| 2024 | Tailored Domain-Specific Summaries: A Two-Stage Method Combining Extractive and Abstractive Summarization Models
Feng Jiang 0007, Lingyi Yang, Haizhou Li 0001 |
ICONIP (9) | 4 |
| 2024 | LitE-SNN: Designing Lightweight and Efficient Spiking Neural Network through Spatial-Temporal Compressive Network Search and Joint Optimization
Qianhui Liu, Malu Zhang, Gang Pan 0001, Haizhou Li 0001 |
IJCAI | 5 |
| 2024 | Apprenticeship-Inspired Elegance: Synergistic Knowledge Distillation Empowers Spiking Neural Networks for Efficient Single-Eye Emotion Recognition
Yang Wang 0106, Haiyang Mei, Qirui Bao, Ziqi Wei 0001, Zheng Shou 0001, Haizhou Li 0001, Bo Dong 0004, Xin Yang 0011 |
IJCAI | 6 |
| 2024 | Target Speech Extraction with Pre-trained AV-HuBERT and Mask-And-Recover StrategyabstractAudio-visual target speech extraction (AV-TSE) is one of the enabling technologies in robotics and many audiovisual applications. One of the challenges of AV-TSE is how to effectively utilize audio-visual synchronization information in the process. AV-HuBERT can be a useful pre-trained model for lip-reading, which has not been adopted by AV-TSE. In this paper, we would like to explore the way to integrate a pretrained AV-HuBERT into our AV-TSE system. We have good reasons to expect an improved performance. To benefit from the inter and intra-modality correlations, we also propose a novel Mask-And-Recover (MAR) strategy for self-supervised learning. The experimental results on the VoxCeleb2 dataset show that our proposed model outperforms the baselines both in terms of subjective and objective metrics, suggesting that the pre-trained AV-HuBERT model provides more informative visual cues for target speech extraction. Furthermore, through a comparative study, we confirm that the proposed Mask-And-Recover strategy is significantly effective. Xueyuan Chen, Xixin Wu, Haizhou Li 0001, Helen M. Meng |
IJCNN | 4 |
| 2024 | How Do Neural Spoofing Countermeasures Detect Partially Spoofed Audio?
Tianchi Liu 0004, Lin Zhang 0054, Rohan Kumar Das, Ruijie Tao, Haizhou Li 0001 |
INTERSPEECH | 6 |
| 2024 | FluentEditor: Text-based Speech Editing by Considering Acoustic and Prosody Consistency
Rui Liu 0008, Jiatian Xi, Ziyue Jiang 0001, Haizhou Li 0001 |
INTERSPEECH | 4 |
| 2024 | wTIMIT2mix: A Cocktail Party Mixtures Database to Study Target Speaker Extraction for Normal and Whispered Speech
Marvin Borsdorf, Zexu Pan, Haizhou Li 0001, Tanja Schultz |
INTERSPEECH | 3 |
| 2024 | Does the Lombard Effect Matter in Speech Separation? Introducing the Lombard-GRID-2mix Dataset
Iva Ewert, Marvin Borsdorf, Haizhou Li 0001, Tanja Schultz |
INTERSPEECH | 3 |
| 2024 | SA-WavLM: Speaker-Aware Self-Supervised Pre-training for Mixture SpeechabstractIt was shown that pre-trained models with self-supervised learning (SSL) techniques are effective in various downstream speech tasks.However, most such models are trained on singlespeaker speech data, limiting their effectiveness in mixture speech.This motivates us to explore pre-training on mixture speech.This work presents SA-WavLM, a novel pre-trained model for mixture speech.Specifically, SA-WavLM follows an "extract-merge-predict" pipeline in which the representations of each speaker in the input mixture are first extracted individually and then merged before the final prediction.In this pipeline, SA-WavLM performs speaker-informed extractions with the consideration of the interactions between different speakers.Furthermore, a speaker shuffling strategy is proposed to enhance the robustness towards the speaker absence.Experiments show that SA-WavLM either matches or improves upon the state-of-the-art pre-trained models. Jingru Lin, Meng Ge, Junyi Ao, Liqun Deng, Haizhou Li 0001 |
INTERSPEECH | 5 |
| 2024 | ASA: An Auditory Spatial Attention Dataset with Multiple Speaking Locations
Zijie Lin, Tianyu He, Siqi Cai 0002, Haizhou Li 0001 |
INTERSPEECH | 4 |
| 2024 | Leveraging Graphic and Convolutional Neural Networks for Auditory Attention Detection with EEG
Saurav Pahuja, Gabriel Ivucic, Pascal Himmelmann, Siqi Cai 0002, Tanja Schultz, Haizhou Li 0001 |
INTERSPEECH | 6 |
| 2024 | ED-sKWS: Early-Decision Spiking Neural Networks for Rapid, and Energy-Efficient Keyword Spotting
Zeyang Song, Qianhui Liu, Qu Yang, Yizhou Peng, Haizhou Li 0001 |
INTERSPEECH | 5 |
| 2024 | WeSep: A Scalable and Flexible Toolkit Towards Generalizable Target Speaker Extraction
Shuai Wang 0016, Shaoxiong Lin, Meng Ge, Jianwei Yu 0001, Yanmin Qian, Haizhou Li 0001 |
INTERSPEECH | 9 |
| 2024 | An Exploration of Length Generalization in Transformer-Based Speech Enhancement
Qiquan Zhang, Hongxu Zhu, Xinyuan Qian 0001, Eliathamby Ambikairajah, Haizhou Li 0001 |
INTERSPEECH | 5 |
| 2024 | Generative Expressive Conversational Speech SynthesisabstractConversational Speech Synthesis (CSS) aims to express a target utterance with the proper speaking style in a user-agent conversation setting. Existing CSS methods employ effective multi-modal context modeling techniques to achieve empathy understanding and expression. However, they often need to design complex network architectures and meticulously optimize the modules within them. In addition, due to the limitations of small-scale datasets containing scripted recording styles, they often fail to simulate real natural conversational styles. To address the above issues, we propose a novel generative expressive CSS system, termed GPT-Talker.We transform the multimodal information of the multi-turn dialogue history into discrete token sequences and seamlessly integrate them to form a comprehensive user-agent dialogue context. Leveraging the power of GPT, we predict the token sequence, that includes both semantic and style knowledge, of response for the agent. After that, the expressive conversational speech is synthesized by the conversation-enriched VITS to deliver feedback to the user.Furthermore, we propose a large-scale Natural CSS Dataset called NCSSD, that includes both naturally recorded conversational speech in improvised styles and dialogues extracted from TV shows. It encompasses both Chinese and English languages, with a total duration of 236 hours. We conducted comprehensive experiments on the reliability of the NCSSD and the effectiveness of our GPT-Talker. Both subjective and objective evaluations demonstrate that our model outperforms other state-of-the-art CSS systems significantly in terms of naturalness and expressiveness. The Code, Dataset, and Pre-trained Model are available at: https://github.com/AI-S2-Lab/GPT-Talker. Rui Liu 0008, Yifan Hu 0004, Yi Ren 0006, Xiang Yin 0006, Haizhou Li 0001 |
ACM Multimedia | 5 |
| 2024 | GROOT: Generating Robust Watermark for Diffusion-Model-Based Audio SynthesisabstractAmid the burgeoning development of generative models like diffusion models, the task of differentiating synthesized audio from its natural counterpart grows more daunting. Deepfake detection offers a viable solution to combat this challenge. Yet, this defensive measure unintentionally fuels the continued refinement of generative models. Watermarking emerges as a proactive and sustainable tactic, preemptively regulating the creation and dissemination of synthesized content. Thus, this paper, as a pioneer, proposes the generative robust audiowatermarking method (Groot), presenting a paradigm for proactively supervising the synthesized audio and its source diffusion models. In this paradigm, the processes of watermark generation and audio synthesis occur simultaneously, facilitated by parameter-fixed diffusion models equipped with a dedicated encoder. The watermark embedded within the audio can subsequently be retrieved by a lightweight decoder. The experimental results highlight Groot's outstanding performance, particularly in terms of robustness, surpassing that of the leading state-of-the-art methods. Beyond its impressive resilience against individual post-processing attacks, Groot exhibits exceptional robustness when facing compound attacks, maintaining an average watermark extraction accuracy of around 95%. Our audio samples are available at https://groot-gaw.github.io/. Yue Li 0041, Dongdong Lin, Hui Tian 0002, Haizhou Li 0001 |
ACM Multimedia | 5 |
| 2024 | ListenFormer: Responsive Listening Head Generation with Non-autoregressive TransformersabstractAs one of the crucial elements in human-robot interaction, responsive listening head generation has attracted considerable attention from researchers. It aims to generate a listening head video based on speaker's audio and video as well as a reference listener image. However, existing methods exhibit two limitations: 1) the generation capability of their models is limited, resulting in generated videos that are far from real ones, and 2) they mostly employ autoregressive generative models, unable to mitigate the risk of error accumulation. To tackle these issues, we propose Listenformer that leverages the powerful temporal modeling capability of transformers for generation. It can perform non-autoregressive prediction with the proposed two-stage training method, simultaneously achieving temporal continuity and overall consistency in the outputs. To fully utilize the information from the speaker inputs, we designed an audio-motion attention fusion module, which improves the correlation of audio and motion features for accurate response. Additionally, a novel decoding method called sliding window with a large shift is proposed for Listenformer, demonstrating both excellent computational efficiency and effectiveness. Extensive experiments show that Listenformer outperforms the existing state-of-the-art methods on ViCo and L2L datasets. And a perceptual user study demonstrates the comprehensive performance of our method in generating diversity, identity preserving, speaker-listener synchronization, and attitude matching. Our code is available at https://liushenme.github.io/ListenFormer.github.io/. Miao Liu 0007, Jing Wang 0037, Xinyuan Qian 0001, Haizhou Li 0001 |
ACM Multimedia | 4 |
| 2024 | Multi-Stage Face-Voice Association Learning with Keynote Speaker DiarizationabstractThe human brain has the capability to associate the unknown person's voice and face by leveraging their general relationship, referred to as "cross-modal speaker verification''. This task poses significant challenges due to the complex relationship between the modalities. In this paper, we propose a "Multi-stage Face-voice Association Learning with Keynote Speaker Diarization''(MFV-KSD) framework. MFV-KSD contains a keynote speaker diarization front-end to effectively address the noisy speech inputs issue. To balance and enhance the intra-modal feature learning and inter-modal correlation understanding, MFV-KSD utilizes a novel three-stage training strategy. Our experimental results demonstrated robust performance, achieving the first rank in the 2024 Face-voice Association in Multilingual Environments (FAME) challenge with an overall Equal Error Rate (EER) of 19.9%. Details can be found in https://github.com/TaoRuijie/MFV-KSD. Ruijie Tao, Yidi Jiang, Duc-Tuan Truong, Chng Eng Siong, Massimo Alioto, Haizhou Li 0001 |
ACM Multimedia | 7 |
| 2024 | MMAL: Multi-Modal Analytic Learning for Exemplar-Free Audio-Visual Class Incremental TasksabstractClass-incremental learning poses a significant challenge under an exemplar-free constraint, leading to catastrophic forgetting and sub-par incremental accuracy. Previous attempts have focused primarily on single-modality tasks, such as image classification or audio event classification. However, in the context of Audio-Visual Class-Incremental Learning (AVCIL), the effective integration and utilization of heterogeneous modalities, with their complementary and enhancing characteristics, remains largely unexplored. To bridge this gap, we propose the Multi-Modal Analytic Learning (MMAL) framework, an exemplar-free solution for AVCIL that employs a closed-form, linear approach. To be specific, MMAL introduces a modality fusion module that re-formulates the AVCIL problem through a Recursive Least-Square (RLS) perspective. Complementing this, a Modality-Specific Knowledge Compensation (MSKC) module is designed to further alleviate the under-fitting limitation intrinsic to analytic learning by harnessing individual knowledge from audio and visual modality in tandem. Comprehensive experimental comparisons with existing methods show that our proposed MMAL demonstrates superior performance with the accuracy of 76.71%, 78.98%, and 76.19% on AVE, Kinetics-Sounds, and VGGSounds100 datasets, respectively, setting new state-of-the-art AVCIL performance. Notably, compared to those memory-based methods, our MMAL, being an exemplar-free approach, provides good data privacy and can better leverage multi-modal information for improved incremental accuracy. Xianghu Yue, Xueyi Zhang 0001, Yiming Chen 0010, Mingrui Lao, Huiping Zhuang, Xinyuan Qian 0001, Haizhou Li 0001 |
ACM Multimedia | 8 |
| 2024 | AceGPT, Localizing Large Language Models in ArabicabstractHuang Huang, Fei Yu, Jianqing Zhu, Xuening Sun, Hao Cheng, Song Dingjie, Zhihong Chen, Mosen Alharthi, Bang An, Juncai He, Ziche Liu, Junying Chen, Jianquan Li, Benyou Wang, Lian Zhang, Ruoyu Sun, Xiang Wan, Haizhou Li, Jinchao Xu. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Fei Yu 0017, Jianqing Zhu, Xuening Sun, Dingjie Song, Mosen Alharthi, Bang An 0004, Juncai He 0001, Ziche Liu, Benyou Wang, Ruoyu Sun 0001, Haizhou Li 0001, Jinchao Xu |
NAACL-HLT | 18 |
| 2024 | CMB: A Comprehensive Medical Benchmark in ChineseabstractXidong Wang, Guiming Chen, Song Dingjie, Zhang Zhiyi, Zhihong Chen, Qingying Xiao, Junying Chen, Feng Jiang, Jianquan Li, Xiang Wan, Benyou Wang, Haizhou Li. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Xidong Wang, Guiming Chen, Dingjie Song, Zhiyi Zhang 0007, Qingying Xiao, Feng Jiang 0007, Benyou Wang, Haizhou Li 0001 |
NAACL-HLT | 12 |
| 2024 | SD-Eval: A Benchmark Dataset for Spoken Dialogue Understanding Beyond WordsabstractSpeech encompasses a wealth of information, including but not limited to content, paralinguistic, and environmental information.This comprehensive nature of speech significantly impacts communication and is crucial for human-computer interaction.Chat-Oriented Large Language Models (LLMs), known for their general-purpose assistance capabilities, have evolved to handle multi-modal inputs, including speech.Although these models can be adept at recognizing and analyzing speech, they often fall short of generating appropriate responses.We argue that this is due to the lack of principles on task definition and model development, which requires open-source datasets and metrics suitable for model evaluation.To bridge the gap, we present SD-Eval, a benchmark dataset aimed at multidimensional evaluation of spoken dialogue understanding and generation.SD-Eval focuses on paralinguistic and environmental information and includes 7,303 utterances, amounting to 8.76 hours of speech data. The data is aggregated from eight public datasets, representing four perspectives: emotion, accent, age, and background sound.To assess the SD-Eval benchmark dataset, we implement three different models and construct a training set following a process similar to that of SD-Eval. The training set contains 1,052.72 hours of speech data and 724.4k utterances. We also conduct a comprehensive evaluation using objective evaluation methods (e.g. BLEU and ROUGE), subjective evaluations and LLM-based metrics for the generated responses.Models conditioned with paralinguistic and environmental information outperform their counterparts in both objective and subjective measures.Moreover, experiments demonstrate that LLM-based metrics show a higher correlation with human evaluation compared to traditional metrics.We open-source SD-Eval at https://github.com/amphionspace/SD-Eval. Junyi Ao, Yuancheng Wang, Xiaohai Tian, Dekun Chen, Jun Zhang 0066, Lu Lu 0015, Yuxuan Wang 0002, Haizhou Li 0001, Zhizheng Wu 0001 |
NeurIPS | 8 |
| 2024 | Alignment at Pre-training! Towards Native Alignment for Arabic LLMsabstractThe alignment of large language models (LLMs) is critical for developing effective and safe language models. Traditional approaches focus on aligning models during the instruction tuning or reinforcement learning stages, referred to in this paper as `\textit{post alignment}'. We argue that alignment during the pre-training phase, which we term 'native alignment', warrants investigation. Native alignment aims to prevent unaligned content from the beginning, rather than relying on post-hoc processing. This approach leverages extensively aligned pre-training data to enhance the effectiveness and usability of pre-trained models. Our study specifically explores the application of native alignment in the context of Arabic LLMs. We conduct comprehensive experiments and ablation studies to evaluate the impact of native alignment on model performance and alignment stability. Additionally, we release open-source Arabic LLMs that demonstrate state-of-the-art performance on various benchmarks, providing significant benefits to the Arabic LLM community. Juhao Liang, Zhenyang Cai, Jianqing Zhu, Kewei Zong, Bang An 0004, Mosen Alharthi, Juncai He 0001, Haizhou Li 0001, Benyou Wang, Jinchao Xu |
NeurIPS | 10 |
| 2024 | Language Without Borders: A Dataset and Benchmark for Code-Switching Lip ReadingabstractLip reading aims at transforming the videos of continuous lip movement into textual contents, and has achieved significant progress over the past decade. It serves as a critical yet practical assistance for speech-impaired individuals, with more practicability than speech recognition in noisy environments. With the increasing interpersonal communications in social media owing to globalization, the existing monolingual datasets for lip reading may not be sufficient to meet the exponential proliferation of bilingual and even multilingual users. However, to our best knowledge, research on code-switching is only explored in speech recognition, while the attempts in lip reading are seriously neglected. To bridge this gap, we have collected a bilingual code-switching lip reading benchmark composed of Chinese and English, dubbed CSLR. As the pioneering work, we recruited 62 speakers with proficient foundations in bothspoken Chinese and English to express sentences containing both involved languages. Through rigorous criteria in data selection, CSLR benchmark has accumulated 85,560 video samples with a resolution of 1080x1920, totaling over 71.3 hours of high-quality code-switching lip movement data. To systematically evaluate the technical challenges in CSLR, we implement commonly-used lip reading backbones, as well as competitive solutions in code-switching speech for benchmark testing. Experiments show CSLR to be a challenging and under-explored lip reading task. We hope our proposed benchmark will extend the applicability of code-switching lip reading, and further contribute to the communities of cross-lingual communication and collaboration. Our dataset and benchmark are accessible at https://github.com/cslr-lipreading/CSLR. Xueyi Zhang 0001, Mingrui Lao, Jun Tang 0001, Yanming Guo, Siqi Cai 0002, Xianghu Yue, Haizhou Li 0001 |
NeurIPS | 8 |
| 2024 | On the Effectiveness of Enrollment Speech Augmentation For Target Speaker ExtractionabstractDeep learning technologies have significantly advanced the performance of target speaker extraction (TSE) tasks. To enhance the generalization and robustness of these algorithms when training data is insufficient, data augmentation is a commonly adopted technique. Unlike typical data augmentation applied to speech mixtures, this work thoroughly investigates the effectiveness of augmenting the enrollment speech space. We found that for both pretrained and jointly optimized speaker encoders, directly augmenting the enrollment speech leads to consistent performance improvement. In addition to conventional methods such as noise and reverberation addition, we propose a novel augmentation method called self-estimated speech augmentation (SSA). Experimental results on the Libri2Mix test set show that our proposed method can achieve an improvement of up to 2.5 dB. Shuai Wang 0016, Haizhou Li 0001, Man-Wai Mak, Kong-Aik Lee |
SLT | 4 |
| 2024 | Neurospex: Neuro-Guided Speaker Extraction With Cross-Modal FusionabstractIn the study of auditory attention, it has been revealed that there exists a robust correlation between attended speech and elicited neural responses, measurable through electroencephalography (EEG). Therefore, it is possible to use the attention information available within EEG signals to guide the extraction of the target speaker in a cocktail party computationally. In this paper, we present a neuro-guided speaker extraction model, i.e. NeuroSpex, using the EEG response of the listener as the sole auxiliary reference cue to extract attended speech from monaural speech mixtures. We propose a novel EEG signal encoder that captures the attention information. Additionally, we propose a cross-attention (CA) mechanism to enhance the speech feature representations, generating a speaker extraction mask. Experimental results on a publicly available dataset demonstrate that our proposed model outperforms two baseline models across various evaluation metrics. Dashanka De Silva, Siqi Cai 0002, Saurav Pahuja, Tanja Schultz, Haizhou Li 0001 |
SLT | 5 |
| 2024 | Enhancing Speaker Extraction Through Rectifying Target ConfusionabstractTarget Speaker Extraction (TSE) aims to extract target speech from mixed audio using clues that identify the target speaker. However, TSE often faces the Target Confusion (TC) problem, where the model extracts the interfering speech instead of the target speech, leading to significant performance degradation. In this paper, we propose a novel model with two branches that enhance target speech extraction by explicitly modeling the interference. Additionally, we propose a Target Confusion Rectification (TCR) method to address the aforementioned TC problem. When the TSE model outputs the wrong speaker, the TCR method performs a rectifying step to ensure the model extracts the correct speaker. Experiments show that under the train-100 subset of Libri2Mix dataset, our proposed method significantly improves the extracting performance in terms of SI-SNRi, PESQ score and extracting accuracy, with that under ‘mix_clean’ subset slightly better than that under ‘mix_both’ subset. Shuai Wang 0016, Yanmin Qian, Haizhou Li 0001 |
SLT | 6 |
| 2024 | Amphion: an Open-Source Audio, Music, and Speech Generation ToolkitabstractAmphion is an open-source toolkit for Audio, Music, and Speech Generation, targeting to ease the way for junior researchers and engineers into these fields. It presents a unified framework that includes diverse generation tasks and models, with the added bonus of being easily extendable for new incorporation. The toolkit is designed with beginner-friendly workflows and pre-trained models, allowing both beginners and seasoned researchers to kick-start their projects with relative ease. The initial release of Amphion v0.1 supports a range of tasks including Text to Speech (TTS), Text to Audio (TTA), and Singing Voice Conversion (SVC), supplemented by essential components like data preprocessing, state-of-the-art vocoders, and evaluation metrics. This paper presents a high-level overview of Amphion. Amphion is open-sourced at https://github.com/open-mmlab/Amphion. Xueyao Zhang, Liumeng Xue, Yicheng Gu, Yuancheng Wang, Jiaqi Li 0030, Haorui He, Chaoren Wang, Songting Liu, Junan Zhang, Zihao Fang, Haopeng Chen, Tze Ying Tang, Lexiao Zou, Mingxuan Wang, Kai Chen 0026, Haizhou Li 0001, Zhizheng Wu 0001 |
SLT | 18 |
| 2024 | Intelligent event-based lip reading word classification with spiking neural networks using spatio-temporal attention features and triplet loss
Qianhui Liu, Meng Ge, Haizhou Li 0001 |
Inf. Sci. | 3 |
| 2024 | Efficient spiking neural network design via neural architecture search
Qianhui Liu, Malu Zhang, Lang Feng 0002, De Ma, Haizhou Li 0001, Gang Pan 0001 |
Neural Networks | 6 |
| 2024 | A Hybrid Neural Coding Approach for Pattern Recognition With Spiking Neural NetworksabstractRecently, brain-inspired spiking neural networks (SNNs) have demonstrated promising capabilities in solving pattern recognition tasks. However, these SNNs are grounded on homogeneous neurons that utilize a uniform neural coding for information representation. Given that each neural coding scheme possesses its own merits and drawbacks, these SNNs encounter challenges in achieving optimal performance such as accuracy, response time, efficiency, and robustness, all of which are crucial for practical applications. In this study, we argue that SNN architectures should be holistically designed to incorporate heterogeneous coding schemes. As an initial exploration in this direction, we propose a hybrid neural coding and learning framework, which encompasses a neural coding zoo with diverse neural coding schemes discovered in neuroscience. Additionally, it incorporates a flexible neural coding assignment strategy to accommodate task-specific requirements, along with novel layer-wise learning methods to effectively implement hybrid coding SNNs. We demonstrate the superiority of the proposed framework on image classification and sound localization tasks. Specifically, the proposed hybrid coding SNNs achieve comparable accuracy to state-of-the-art SNNs, while exhibiting significantly reduced inference latency and energy consumption, as well as high noise robustness. This study yields valuable insights into hybrid neural coding designs, paving the way for developing high-performance neuromorphic systems. Qu Yang, Jibin Wu, Haizhou Li 0001, Kay Chen Tan |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | Advancing speaker embedding learning: Wespeaker toolkit for research and production
Shuai Wang 0016, Zhengyang Chen, Bing Han 0008, Chengdong Liang, Xu Xiang, Wen Ding 0005, Johan Rohdin, Anna Silnova, Yanmin Qian, Haizhou Li 0001 |
Speech Commun. | 12 |
| 2024 | Transferable Adversarial Attacks Against ASRabstractGiven the extensive research and real-world applications of automatic speech recognition (ASR), ensuring the robustness of ASR models against minor input perturbations becomes a crucial consideration for maintaining their effectiveness in real-time scenarios. Previous explorations into ASR model robustness have predominantly revolved around evaluating accuracy on white-box settings with full access to ASR models. Nevertheless, full ASR model details are often not available in real-world applications. Therefore, evaluating the robustness of black-box ASR models is essential for a comprehensive understanding of ASR model resilience. In this regard, we thoroughly study the vulnerability of practical black-box attacks in cutting-edge ASR models and propose to employ two advanced time-domain-based transferable attacks alongside our differentiable feature extractor. We also propose a speech-aware gradient optimization approach (SAGO) for ASR, which forces mistranscription with minimal impact on human imperceptibility through voice activity detection rule and a speech-aware gradient-oriented optimizer. Our comprehensive experimental results reveal performance enhancements compared to baseline approaches across five models on two databases. Xiaoxue Gao, Zexin Li 0001, Yiming Chen 0010, Cong Liu 0005, Haizhou Li 0001 |
IEEE Signal Process. Lett. | 5 |
| 2024 | Selective HuBERT: Self-Supervised Pre-Training for Target Speaker in Clean and Mixture SpeechabstractSelf-supervised pre-trained speech models were shown effective for various downstream speech processing tasks. Since they are mainly pre-trained to map input speech to pseudo-labels, the resulting representations are only effective for the type of pre-train data used, either clean or mixture speech. With the idea of selective auditory attention, we propose a novel pre-training solution called Selective-HuBERT, or SHuBERT, which learns the selective extraction of target speech representations from either clean or mixture speech. Specifically, SHuBERT is trained to predict pseudo labels of a target speaker, conditioned on an enrolment speech from the target speaker. By doing so, SHuBERT is expected to selectively attend to the target speaker in a complex acoustic environment, thus benefiting various downstream tasks. We further introduce a dual-path training strategy and use the cross-correlation constraint between the two branches to encourage the model to generate noise-invariant representation. Experiments on SUPERB benchmark and LibriMix dataset demonstrate the universality and noise-robustness of SHuBERT. Furthermore, we find that our high-quality representation can be easily integrated with conventional supervised learning methods to achieve significant performance, even under extremely low-resource labeled data. Jingru Lin, Meng Ge, Wupeng Wang, Haizhou Li 0001, Mengling Feng |
IEEE Signal Process. Lett. | 4 |
| 2024 | Text-Guided HuBERT: Self-Supervised Speech Pre-Training via Generative Adversarial NetworksabstractHuman language can be expressed in either written or spoken form, i.e. text or speech. Humans can acquire knowledge from text to improve speaking and listening. However, the quest for speech pre-trained models to leverage unpaired text has just started. In this letter, we investigate a new way to pre-train such a joint speech-text model to learn enhanced speech representations and benefit various speech-related downstream tasks. Specifically, we propose a novel pre-training method, text-guided HuBERT, or T-HuBERT, which performs self-supervised learning over speech to derive phoneme-like discrete representations. And these phoneme-like pseudo-label sequences are firstly derived from speech via the generative adversarial networks (GAN) to be statistically similar to those from additional unpaired textual data. In this way, we build a bridge between unpaired speech and text in a unsupervised manner. Extensive experiments demonstrate the significant superiority of our proposed method over various strong baselines, which achieves up to 15.3% relative Word Error Rate (WER) reduction on the LibriSpeech dataset. Duo Ma, Xianghu Yue, Junyi Ao, Xiaoxue Gao, Haizhou Li 0001 |
IEEE Signal Process. Lett. | 5 |
| 2024 | Contrastive Learning Based Modality-Invariant Feature Acquisition for Robust Multimodal Emotion Recognition With Missing ModalitiesabstractMultimodal emotion recognition (MER) aims to understand the way that humans express their emotions by exploring complementary information across modalities. However, it is hard to guarantee that full-modality data is always available in real-world scenarios. To deal with missing modalities, researchers focused on meaningful joint multimodal representation learning during cross-modal missing modality imagination. However, the cross-modal imagination mechanism is highly susceptible to errors due to the “modality gap” issue, which affects the imagination accuracy, thus, the final recognition performance. To this end, we introduce the concept of a modality-invariant feature into the missing modality imagination network, which contains two key modules: 1) a novel contrastive learning-based module to extract modality-invariant features under full modalities; 2) a robust imagination module based on imagined invariant features to reconstruct missing information under missing conditions. Finally, we incorporate imagined and available modalities for emotion recognition. Experimental results on benchmark datasets demonstrate that our proposed method outperforms existing state-of-the-art strategies. Compared with our previous work, our extended version is more effective on multimodal emotion recognition with missing modalities. The code is released athttps://github.com/ZhuoYulang/CIF-MMIN. Rui Liu 0008, Haolin Zuo, Zheng Lian 0004, Björn W. Schuller, Haizhou Li 0001 |
IEEE Trans. Affect. Comput. | 5 |
| 2024 | An Investigation of Time-Frequency Representation Discriminators for High-Fidelity VocodersabstractGenerative Adversarial Network (GAN) based vocoders are superior in both inference speed and synthesis quality when reconstructing an audible waveform from an acoustic representation. This study focuses on improving the discriminator for GAN-based vocoders. Most existing Time-Frequency Representation (TFR)-based discriminators are rooted in Short-Time Fourier Transform (STFT), which owns a constant Time-Frequency (TF) resolution, linearly scaled center frequencies, and a fixed decomposition basis, making it incompatible with signals like singing voices that require dynamic attention for different frequency bands and different time intervals. Motivated by that, we propose a Multi-Scale Sub-Band Constant-Q Transform CQT (MS-SB-CQT) discriminator and a Multi-Scale Temporal-Compressed Continuous Wavelet Transform CWT (MS-TC-CWT) discriminator. Both CQT and CWT have a dynamic TF resolution for different frequency bands. In contrast, CQT has a better modeling ability in pitch information, and CWT has a better modeling ability in short-time transients. Experiments conducted on both speech and singing voices confirm the effectiveness of our proposed discriminators. Moreover, the STFT, CQT, and CWT-based discriminators can be used jointly for better performance. The proposed discriminators can boost the synthesis quality of various state-of-the-art GAN-based vocoders, including HiFi-GAN, BigVGAN, and APNet. Yicheng Gu, Xueyao Zhang, Liumeng Xue, Haizhou Li 0001, Zhizheng Wu 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2024 | Computation and Parameter Efficient Multi-Modal Fusion Transformer for Cued Speech RecognitionabstractCued Speech (CS) is a pure visual coding method used by hearing-impaired people that combines lip reading with several specific hand shapes to make the spoken language visible. Automatic CS recognition (ACSR) seeks to transcribe visual cues of speech into text, which can help hearing-impaired people to communicate effectively. The visual information of CS contains lip reading and hand cueing, thus the fusion of them plays an important role in ACSR. However, most previous fusion methods struggle to capture the global dependency present in long sequence inputs of multi-modal CS data. As a result, these methods generally fail to learn the effective cross-modal relationships that contribute to the fusion. Recently, attentionbased transformers have been a prevalent idea for capturing the global dependency over the long sequence in multi-modal fusion, but existing multi-modal fusion transformers suffer from both poor recognition accuracy and inefficient computation for the ACSR task. To address these problems, we develop a novel computation and parameter efficient multi-modal fusion transformer by proposing a novel Token-Importance-Aware Attention mechanism (TIAA), where a token utilization rate (TUR) is formulated to select the important tokens from the multi-modal streams. More precisely, TIAA firstly models the modality-specific fine-grained temporal dependencies over all tokens of each modality, and then learns the efficient cross-modal interaction for the modality-shared coarse-grained temporal dependencies over the important tokens of different modalities. Besides, a lightweight gated hidden projection is designed to control the feature flows of TIAA. The resulting model, named Economical Cued Speech Fusion Transformer (EcoCued), achieves state-of-the-art performance on all existing CS datasets (i.e., Mandarin Chinese, French, and British CS), compared with existing transformerbased fusion methods and ACSR fusion methods. Notably, our method dramatically reduces the computational complexity from O(T2) to O(T). We will release the source code and data as open source. Lei Liu 0049, Li Liu 0036, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2024 | Golden Gemini is All You Need: Finding the Sweet Spots for Speaker VerificationabstractThe residual neural networks (ResNet) demonstrate the impressive performance in automatic speaker verification (ASV). They treat the time and frequency dimensions equally, following the default stride configuration designed for image recognition, where the horizontal and vertical axes exhibit similarities. This approach ignores the fact that time and frequency are asymmetric in speech representation. We address this issue and postulateGolden-Gemini Hypothesis,which posits the prioritization of temporal resolution over frequency resolution for ASV. The hypothesis is verified by conducting a systematic study on the impact of temporal and frequency resolutions on the performance, using a trellis diagram to represent the stride space. We further identify two optimal points, namelyGolden Gemini, which serves as a guiding principle for designing 2D ResNet-based ASV models. By following the principle, a state-of-the-art ResNet baseline model gains a significant performance improvement on VoxCeleb, SITW, and CNCeleb datasets with 7.70%/11.76% average EER/minDCF reductions, respectively, across different network depths (ResNet18, 34, 50, and 101), while reducing the number of parameters by 16.5% and FLOPs by 4.1%. We refer to it asGeminiResNet. Further investigation reveals the efficacy of the proposedGolden Geminioperating points across various training conditions and architectures. Furthermore, we present a new benchmark, namely theGeminiDF-ResNet, using a cutting-edge model.Codes and pre-trained models are available athttps://github.com/Tianchi-Liu9/Golden-Gemini-for-Speaker-Verification. Tianchi Liu 0004, Kong-Aik Lee, Qiongqiong Wang, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2024 | Controllable Accented Text-to-Speech Synthesis With Fine and Coarse-Grained Intensity RenderingabstractAccented text-to-speech (TTS) synthesis seeks to generate speech with an accent (L2) as a variant of the standard version (L1), which is challenging as L2 is different from L1 in terms of phonetic rendering and prosody pattern (pitch, energy, and duration variance, etc.). Accented TTS has several significant real-world applications, such as language learning, preserving and documenting endangered languages and dialects, etc. that make it an important area of research and development. Moreover, changing the accent intensity of any conversational AI system has the potential to allow specific users to understand its produced speech better. However, there is no intuitive solution for the control of the accent intensity for an utterance at both fine and coarse-grained levels, that are phoneme and utterance levels respectively. In this work, we propose a neural TTS architecture that allows us to control the accent style and its intensity. This is achieved through two novel mechanisms: 1) the front-end and back-end accent knowledge injection mechanism to enhance the accent interpretability of TTS modeling; and 2) an automatic speech recognition (ASR) based accent intensity modeling strategy to quantify the accent intensity in both L2 phoneme and utterance levels. In the front-end, a newaccent variation adaptorseeks to project the accent-aware pitch, energy and duration features at a phoneme level, with the help of the fine-grained accent intensity information; In the back-end, a consistency constraint module that ensures the synthesized L2 speech manifests the expected accent intensity, is injected in the front-end, precisely. Experiments show that the proposed system attains superior performance to the baseline models in terms of accent rendering and intensity control. To our knowledge, this is the first study of accented TTS with explicit intensity control at both fine and coarse-grained levels. Rui Liu 0008, Berrak Sisman, Guanglai Gao, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2024 | NeuroHeed: Neuro-Steered Speaker Extraction Using EEG SignalsabstractHumans possess the remarkable ability to selectively attend to a single speaker amidst competing voices and background noise, known asselective auditory attention. Recent studies in auditory neuroscience indicate a strong correlation between the attended speech signal and the corresponding brain's elicited neuronal activities. In this work, we study such brain activities measured using affordable and non-intrusive electroencephalography (EEG) devices. We present NeuroHeed, a speaker extraction model that leverages the listener's synchronized EEG signals to extract the attended speech signal in a cocktail party scenario, in which the extraction process is conditioned on a neuronal attractor encoded from the EEG signal. We propose both an offline and an online NeuroHeed, with the latter designed for real-time inference. In the online NeuroHeed, we additionally propose an autoregressive speaker encoder, which accumulates past extracted speech signals for self-enrollment of the attended speaker information into an auditory attractor, that retains the attentional momentum over time. Online NeuroHeed extracts the current window of the speech signals with guidance from both attractors. Experimental results on KUL dataset two-speaker scenario demonstrate that NeuroHeed effectively extracts brain-attended speech signals with an average scale-invariant signal-to-noise ratio improvement (SI-SDRi) of 14.3 dB and extraction accuracy of 90.8% in offline settings, and SI-SDRi of 11.2 dB and extraction accuracy of 85.1% in online settings. Zexu Pan, Marvin Borsdorf, Siqi Cai 0002, Tanja Schultz, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2024 | Multi-Agent Deep Learning for the Detection of Multiple Speech Steganography MethodsabstractThe ability to detect multiple steganographic methods in speech streams is an important prerequisite for steganalysis methods to move from theory to practical application, but it is also a challenging problem. To address this challenge, we propose a novel steganalysis method based on multi-agent deep learning, which can effectively detect multiple steganography methods in speech streams. Our method utilizes multiple agents to learn the features of multiple sub-training datasets separately and then fuses the information of each agent through the weight parameter aggregation mechanism to obtain the final weight parameter of the steganalysis model. Experimental results show that our proposed method outperforms the state-of-art steganalysis methods. In particular, for low embedding rates, the presented method increases average detection accuracy by about 9%. Congcong Sun 0002, Hui Tian 0002, Haizhou Li 0001, Zhenxing Qian |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2024 | Overview of Speaker Modeling and Its Applications: From the Lens of Deep Speaker Representation LearningabstractSpeaker individuality information is among the most critical elements within speech signals. By thoroughly and accurately modeling this information, it can be utilized in various intelligent speech applications, such as speaker recognition, speaker diarization, speech synthesis, and target speaker extraction. In this overview, we present a comprehensive review of neural approaches to speaker representation learning from both theoretical and practical perspectives. Theoretically, we discuss speaker encoders ranging from supervised to self-supervised learning algorithms, standalone models to large pretrained models, pure speaker embedding learning to joint optimization with downstream tasks, and efforts toward interpretability. Practically, we systematically examine approaches for robustness and effectiveness, introduce and compare various open-source toolkits in the field. Through the systematic and comprehensive review of the relevant literature, research activities, and resources, we provide a clear reference for researchers in the speaker characterization and modeling field, as well as for those who wish to apply speaker modeling techniques to specific downstream tasks. Shuai Wang 0016, Zhengyang Chen, Kong-Aik Lee, Yanmin Qian, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2024 | Speech Separation With Pretrained Frontend to Minimize Domain MismatchabstractSpeech separation seeks to separate individual speech signals from a speech mixture. Typically, most separation models are trained on synthetic data due to the unavailability of target reference in real-world cocktail party scenarios. As a result, there exists a domain gap between real and synthetic data when deploying speech separation models in real-world applications. In this paper, we propose a self-supervised domain-invariant pretrained (DIP) frontend that is exposed to mixture data without the need for target reference speech. The DIP frontend utilizes a Siamese network with two innovative pretext tasks, mixture predictive coding (MPC) and mixture invariant coding (MIC), to capture shared contextual cues between real and synthetic unlabeled mixtures. Subsequently, we freeze the DIP frontend as a feature extractor when training the downstream speech separation models on synthetic data. By pretraining the DIP frontend with the contextual cues, we expect that the speech separation skills learned from synthetic data can be effectively transferred to real data. To benefit from the DIP frontend, we introduce a novel separation pipeline to align the feature resolution of the separation models. We evaluate the speech separation quality on standard benchmarks and real-world datasets. The results confirm the superiority of our DIP frontend over existing speech separation models. This study underscores the potential of large-scale pretraining to enhance the quality and intelligibility of speech separation in real-world applications. Wupeng Wang, Zexu Pan, Shuai Wang 0016, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2024 | Overview of the Tenth Dialog System Technology Challenge: DSTC10abstractThis article introduces the Tenth Dialog System Technology Challenge (DSTC-10). This edition of the DSTC focuses on applying end-to-end dialog technologies for five distinct tasks in dialog systems, namely 1. Incorporation of Meme images into open domain dialogs, 2. Knowledge-grounded Task-oriented Dialogue Modeling on Spoken Conversations, 3. Situated Interactive Multimodal dialogs, 4. Reasoning for Audio Visual Scene-Aware Dialog, and 5. Automatic Evaluation and Moderation of Open-domainDialogue Systems. This article describes the task definition, provided datasets, baselines, and evaluation setup for each track. We also summarize the results of the submitted systems to highlight the general trends of the state-of-the-art technologies for the tasks. Koichiro Yoshino, Yun-Nung Chen, Paul A. Crook, Satwik Kottur, Jinchao Li, Behnam Hedayatnia, Seungwhan Moon, Zhengcong Fei, Zekang Li, Jinchao Zhang 0001, Yang Feng 0004, Jie Zhou 0016, Seokhwan Kim, Yang Liu 0004, Di Jin 0005, Alexandros Papangelis, Karthik Gopalakrishnan 0001, Dilek Hakkani-Tür, Babak Damavandi, Alborz Geramifard, Chiori Hori, Chen Zhang 0020, Haizhou Li 0001, João Sedoc, Luis Fernando D'Haro, Rafael E. Banchs, Alexander I. Rudnicky |
IEEE ACM Trans. Audio Speech Lang. Process. | 24 |
| 2024 | RefXVC: Cross-Lingual Voice Conversion With Enhanced Reference LeveragingabstractThis paper proposes RefXVC, a method for cross-lingual voice conversion (XVC) that leverages reference information to improve conversion performance. Previous XVC works generally take an average speaker embedding to condition the speaker identity, which does not account for the changing timbre of speech that occurs with different pronunciations. To address this, our method uses both global and local speaker embeddings to capture the timbre changes during speech conversion. Additionally, we observed a connection between timbre and pronunciation in different languages and utilized this by incorporating a timbre encoder and a pronunciation matching network into our model. Furthermore, we found that the variation in tones is not adequately reflected in a sentence, and therefore, we used multiple references to better capture the range of a speaker's voice. The proposed method outperformed existing systems in terms of both speech quality and speaker similarity, highlighting the effectiveness of leveraging reference information in cross-lingual voice conversion. Mingyang Zhang 0003, Yi Zhou 0020, Yi Ren 0006, Chen Zhang 0020, Xiang Yin 0006, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2024 | Accented Text-to-Speech Synthesis With Limited DataabstractThis paper presents an accented text-to-speech (TTS) synthesis framework with limited training data. We study two aspects concerning accent rendering: phonetic (phoneme difference) and prosodic (pitch pattern and phoneme duration) variations. The proposed accented TTS framework consists of two models: an accented front-end for grapheme-to-phoneme (G2P) conversion and an accented acoustic model with integrated pitch and duration predictors for phoneme-to-Mel-spectrogram prediction. The accented front-end directly models the phonetic variation, while the accented acoustic model explicitly controls the prosodic variation. Specifically, both models are first pretrained on a large amount of data, then only the accent-related layers are fine-tuned on a limited amount of data for the target accent. In the experiments, speech data of three English accents, i.e., General American English, Irish English, and British English Received Pronunciation, are used for pre-training. The pretrained models are then fine-tuned with Scottish and General Australian English accents, respectively. Both objective and subjective evaluation results show that the accented TTS frontend fine-tuned with a small accented phonetic lexicon (5k words) effectively handles the phonetic variation of accents, while the accented TTS acoustic model fine-tuned with a limited amount of accented speech data (approximately 3 minutes) effectively improves the prosodic rendering including pitch and duration. The overall accent modeling contributes to improved speech quality and accent similarity. Xuehao Zhou, Mingyang Zhang 0003, Yi Zhou 0020, Zhizheng Wu 0001, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2024 | Audio-Visual Temporal Forgery Detection Using Embedding-Level Fusion and Multi-Dimensional Contrastive LossabstractAudio-visual deepfake detection is the process of identifying and detecting deepfakes that have been generated using both audio and visual content with AI algorithms. Most existing methods primarily focus on the overall authenticity while neglecting the position of forgeries in time. This can be particularly problematic, as even a small alteration in a clip can significantly impact its meaning. Such brand new attacks are dangerous and how to tackle such attacks remains an open question. In this paper, we present a novel neural network-based model to tackle the temporal forgery detection (TFD) problem. It consists of new audio and visual encoders with cross-modal attention for embedding extraction, and an embedding-level fusion mechanism with self-attention for forgery localization. Besides, a multi-dimensional contrastive loss is proposed which helps the model not only to capture audio-visual inconsistency for deepfake detection but also to exploit temporal inconsistency by coherently constraining the extracted embeddings. Extensive experiments on the LAV-DF dataset show that the presented method outperforms several state-of-the-art temporal forgery localization methods by up to 23.4% on [email protected] and 13.8% on AR@100. In addition, we also show the effectiveness of the proposed model on deepfake detection. Miao Liu 0007, Jing Wang 0037, Xinyuan Qian 0001, Haizhou Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Deep Cross-Modal Retrieval Between Spatial Image and Acoustic SpeechabstractCross-modal Retrieval (CMR) is formulated for the scenarios where the queries and retrieval results are of different modalities. Existing Cross-modal Retrieval (CMR) studies mainly focus on the common contextualized information between text transcripts and images, and the synchronized event information in audio-visual recordings. Unlike all previous works, in this article, we investigate the geometric correspondence between images and speech recordings captured in the same space and formulate a novel CMR task, called Spatial Image-Acoustic Retrieval (SIAR). To this end, we first design a novel speech encoder that consists of convolution neural networks and transformer layers, to learn space-aware speech representations. Then, to eliminate the cross-modal inherent discrepancy, we propose the Contrastive Speech Image Retrieval (CSIR) method which uses supervised contrastive learning to attract the same-space cross-modal features while repelling the ones from different spaces. Finally, image and speech features are directly compared and we predict the SIAR result with the maximum similarity. Extensive experiments demonstrate that our proposed speech encoder can recognize space from human speeches with superior performance over the other prevailing networks. It also sets our penultimate goal of speech-to-speech retrieval. Furthermore, our CSIR proposal can successfully perform bi-directional SIAR between spatial images and reverberant speeches with promising results. Code and data will be available. Xinyuan Qian 0001, Wei Xue 0002, Qiquan Zhang, Ruijie Tao, Haizhou Li 0001 |
IEEE Trans. Multim. | 5 |
| 2024 | Few-Shot Contrastive Transfer Learning With Pretrained Model for Masked Face VerificationabstractFace verification has seen remarkable progress that benefits from large-scale publicly available databases. However, it remains a challenge how to generalize a pretrained face verification model to a new scenario with a limited amount of data. In many real-world applications, the training database only contains a limited number of identities with two images for each identity due to the privacy concern. In this article, we propose to transfer knowledge from a pretrained unmasked face verification model to a new model for verification between masked and unmasked faces, to meet the application requirements during the COVID-19 pandemic. To overcome the lack of intra-class diversity resulting from only a pair of masked and unmasked faces for each identity ($\text{i.e.},$two shots for each identity), a static prototype classification function is designed to learn features for masked faces by utilizing unmasked face knowledge from the pretrained model. Meanwhile, a contrastive constrained embedding function is designed to preserve unmasked face knowledge of the pretrained model during the transfer learning process. By combining these two functions, our method uses knowledge acquired from the pretrained unmasked face verification model to proceed with verification between masked and unmasked faces with a limited amount of training data. Extensive experiments demonstrate that our method can perform better than state-of-the-art methods for verification between masked and unmasked faces in the few-shot transfer learning setting. Zhenyu Weng, Huiping Zhuang, Fulin Luo, Haizhou Li 0001, Zhiping Lin 0001 |
IEEE Trans. Multim. | 4 |
| 2024 | A Bio-Inspired Spiking Attentional Neural Network for Attentional Selection in the Listening BrainabstractHumans show a remarkable ability in solving the cocktail party problem. Decoding auditory attention from the brain signals is a major step toward the development of bionic ears emulating human capabilities. Electroencephalography (EEG)-based auditory attention detection (AAD) has attracted considerable interest recently. Despite much progress, the performance of traditional AAD decoders remains to be improved, especially in low-latency settings. State-of-the-art AAD decoders based on deep neural networks generally lack the intrinsic temporal coding ability in biological networks. In this study, we first propose a bio-inspired spiking attentional neural network, denoted as BSAnet, for decoding auditory attention. BSAnet is capable of exploiting the temporal dynamics of EEG signals using biologically plausible neurons and an attentional mechanism. Experiments on two publicly available datasets confirm the superior performance of BSAnet over other state-of-the-art systems across various evaluation conditions. Moreover, BSAnet imitates realistic brain-like information processing, through which we show the advantage of brain-inspired computational models. Siqi Cai 0002, Peiwen Li, Haizhou Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2024 | Event-Triggered Tracking Control for Nonlinear Systems With Prescribed PerformanceabstractThis article addresses the entry capture problem (ECP) of uncertain nonlinear systems under asymmetric performance constraints. We show that such ECP is commonly encountered in practice that has not been well addressed, whose tracking error is free from any performance constraints initially then is driven into the prescribed region in finite time. For better-transient performance, a unified tunnel prescribed performance (TPP) is developed to provide strict and tight allowable set. By utilizing a scaling function, together with an error scaling function (ESF), and a more general error-dependent transformation function (ETF), we propose an event-triggered tracking control strategy leading to a solution for the underlying ECP with various initial conditions and asymmetric performance constraints. This control strategy is of significant simplicity, stemming from that only 1-bit signal is needed for each data transmission between controller and actuator. We also show that the tracking error (including the initial-constraint violation) is regulated into the prescribed region in a given time globally. Finally, simulations are conducted to illustrate the above theoretical findings. Ruihang Ji, Shuzhi Sam Ge, Kai Zhao 0004, Haizhou Li 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2023 | Dynamic Transformers Provide a False Sense of EfficiencyabstractDespite much success in natural language processing (NLP), pre-trained language models typically lead to a high computational cost during inference.Multi-exit is a mainstream approach to address this issue by making a tradeoff between efficiency and accuracy, where the saving of computation comes from an early exit.However, whether such saving from earlyexiting is robust remains unknown.Motivated by this, we first show that directly adapting existing adversarial attack approaches targeting model accuracy cannot significantly reduce inference efficiency.To this end, we propose a simple yet effective attacking framework, SAME, a novel slowdown attack framework on multi-exit models, which is specially tailored to reduce the efficiency of the multi-exit models.By leveraging the multi-exit models' design characteristics, we utilize all internal predictions to guide the adversarial sample generation instead of merely considering the final prediction.Experiments on the GLUE benchmark show that SAME can effectively diminish the efficiency gain of various multi-exit models by 80% on average, convincingly validating its effectiveness and generalization ability. 1 Yiming Chen 0010, Zexin Li 0001, Wei Yang 0013, Cong Liu 0005, Robby T. Tan, Haizhou Li 0001 |
ACL (1) | 7 |
| 2023 | Minimizing the Accumulated Trajectory Error to Improve Dataset DistillationabstractModel-based deep learning has achieved astounding successes due in part to the availability of large-scale real-world data. However, processing such massive amounts of data comes at a considerable cost in terms of computations, storage, training and the search for good neural architectures. Dataset distillation has thus recently come to the fore. This paradigm involves distilling information from large real-world datasets into tiny and compact synthetic datasets such that processing the latter ideally yields similar performances as the former. State-of-the-art methods primarily rely on learning the synthetic dataset by matching the gradients obtained during training between the real and synthetic data. However, these gradient-matching methods suffer from the so-called accumulated trajectory error caused by the discrepancy between the distillation and subsequent evaluation. To mitigate the adverse impact of this accumulated trajectory error, we propose a novel approach that encourages the optimization algorithm to seek a flat trajectory. We show that the weights trained on synthetic data are robust against the accumulated errors perturbations with the regularization towards the flat trajectory. Our method, called Flat Trajectory Distillation (FTD), is shown to boost the performance of gradient-matching methods by up to 4.7% on a subset of images of the ImageNet dataset with higher resolution images. We also validate the effectiveness and generalizability of our method with datasets of different resolutions and demonstrate its applicability to neural architecture search. Code is available at. https://github.com/AngusDujw/FTD-distillation. Jiawei Du 0002, Yidi Jiang, Vincent Y. F. Tan, Joey Tianyi Zhou, Haizhou Li 0001 |
CVPR | 5 |
| 2023 | Seeing What You Said: Talking Face Generation Guided by a Lip Reading ExpertabstractTalking face generation, also known as speech-to-lip generation, reconstructs facial motions concerning lips given coherent speech input. The previous studies revealed the importance of lip-speech synchronization and visual quality. Despite much progress, they hardly focus on the content of lip movements i.e., the visual intelligibility of the spoken words, which is an important aspect of generation quality. To address the problem, we propose using a lipreading expert to improve the intelligibility of the generated lip regions by penalizing the incorrect generation results. Moreover, to compensate for data scarcity, we train the lip-reading expert in an audio-visual self-supervised manner. With a lip-reading expert, we propose a novel contrastive learning to enhance lip-speech synchronization, and a transformer to encode audio synchronically with video, while considering global temporal dependency of audio. For evaluation, we propose a new strategy with two different lip-reading experts to measure intelligibility of the generated videos. Rigorous experiments show that our proposal is superior to other State-of-the-art (SOTA) methods, such as Wav2Lip, in reading intelligibility i.e., over 38% Word Error Rate (WER) on LRS2 dataset and 27.8% accuracy on LRW dataset. We also achieve the SOTA performance in lip-speech synchronization and comparable performances in visual quality. Xinyuan Qian 0001, Malu Zhang, Robby T. Tan, Haizhou Li 0001 |
CVPR | 5 |
| 2023 | Multi-Head Attention and GRU for Improved Match-Mismatch Classification of Speech Stimulus and EEG ResponseabstractThis work is based on the participation by the HyperAttention team in the Auditory EEG Decoding Challenge, 2023 (ICASSP 2023 Signal Processing Grand Challenge) task 1, which deals with the match-mismatch classification of speech stimuli and EEG responses of human listeners. We demonstrate the benefits of using mel-spectrograms instead of speech envelopes as input features as well as the effectiveness of Multi-Head Attention and GRU for EEG and speech processing. With a total score of 79.05 %, we reach the second place in the challenge. Marvin Borsdorf, Saurav Pahuja, Gabriel Ivucic, Siqi Cai 0002, Haizhou Li 0001, Tanja Schultz |
ICASSP | 5 |
| 2023 | Self-Transriber: Few-Shot Lyrics Transcription With Self-TrainingabstractThe current lyrics transcription approaches heavily rely on supervised learning with labeled data, but such data are scarce and manual labeling of singing is expensive. How to benefit from unlabeled data and alleviate limited data problem have not been explored for lyrics transcription. We propose the first semi-supervised lyrics transcription paradigm, Self-Transcriber, by leveraging on unlabeled data using selftraining with noisy student augmentation. We attempt to demonstrate the possibility of lyrics transcription with a few amount of labeled data. Self-Transcriber generates pseudo labels of the unlabeled singing using teacher model, and augments pseudo-labels to the labeled data for student model update with both self-training and supervised training losses. This work closes the gap between supervised and semi- supervised learning as well as opens doors for few-shot learning of lyrics transcription. Our experiments show that our approach using only 12.7 hours of labeled data achieves competitive performance compared with the supervised approaches trained on 149.1 hours of labeled data for lyrics transcription. Xiaoxue Gao, Xianghu Yue, Haizhou Li 0001 |
ICASSP | 3 |
| 2023 | ImagineNet: Target Speaker Extraction with Intermittent Visual Cue Through Embedding InpaintingabstractThe speaker extraction technique seeks to single out the voice of a target speaker from the interfering voices in a speech mixture. Typically an auxiliary reference of the target speaker is used to form voluntary attention. Either a pre-recorded utterance or a synchronized lip movement in a video clip can serve as the auxiliary reference. The use of visual cue is not only feasible, but also effective due to its noise robustness, and becoming popular. However, it is difficult to guarantee that such parallel visual cue is always available in real-world applications where visual occlusion or intermittent communication can occur. In this paper, we study the audio-visual speaker extraction algorithms with intermittent visual cue. We propose a joint speaker extraction and visual embedding inpainting framework to explore the mutual benefits. To encourage the interaction between the two tasks, they are performed alternately with an interlacing structure and optimized jointly. We also propose two types of visual inpainting losses and study our proposed method with two types of popularly used visual embeddings. The experimental results show that we outperform the baseline in terms of signal quality, perceptual quality, and intelligibility. Zexu Pan, Wupeng Wang, Marvin Borsdorf, Haizhou Li 0001 |
ICASSP | 4 |
| 2023 | Speaker Recognition with Two-Step Multi-Modal Deep CleansingabstractNeural network-based speaker recognition has achieved significant improvement in recent years. A robust speaker representation learns meaningful knowledge from both hard and easy samples in the training set to achieve good performance. However, noisy samples (i.e., with wrong labels) in the training set induce confusion and cause the network to learn the incorrect representation. In this paper, we propose a two-step audio-visual deep cleansing framework to eliminate the effect of noisy labels in speaker representation learning. This framework contains a coarse-grained cleansing step to search for the complex samples, followed by a fine-grained cleansing step to filter out the noisy labels. Our study starts from an efficient audio-visual speaker recognition system, which achieves a close to perfect equal-error-rate (EER) of 0.01%, 0.07% and 0.13% on the Vox-O, E and H test sets. With the proposed multi-modal cleansing mechanism, four different speaker recognition networks achieve an average improvement of 5.9%. Code has been made available at: https://github.com/TaoRuijie/AVCleanse. Ruijie Tao, Kong-Aik Lee, Haizhou Li 0001 |
ICASSP | 4 |
| 2023 | Token2vec: A Joint Self-Supervised Pre-Training Framework Using Unpaired Speech and TextabstractSelf-supervised pre-training has been successful in both text and speech processing. Speech and text offer different but complementary information. The question is whether we are able to perform a speech-text joint pre-training on unpaired speech and text. In this paper, we take the idea of self-supervised pre-training one step further and propose token2vec, a novel joint pre-training framework for unpaired speech and text based on discrete representations of speech. Specifically, we introduce two modality-specific tokenizers for speech and text. Based on these tokenizers, we convert speech/text sequences into discrete speech/text token sequences consisting of similar language units, thus mitigating the domain mismatch problem and length mismatch problem, which are caused by the distinct characteristics between speech and text. Finally, we feed the discrete speech and text tokens into a modality-agnostic Transformer encoder and pre-train with token-level masking language modeling (tMLM). Experiments show that token2vec is significantly superior to various speech-only pre-training baselines, with up to 17.7% relative WER reduction. Token2vec model is also validated on a non-ASR task, i.e., spoken intent classification, and shows good transferability. Xianghu Yue, Junyi Ao, Xiaoxue Gao, Haizhou Li 0001 |
ICASSP | 4 |
| 2023 | Ripple Sparse Self-Attention for Monaural Speech EnhancementabstractThe use of Transformer represents a recent success in speech enhancement. However, as its core component, self-attention suffers from quadratic complexity, which is computationally prohibited for long speech recordings. Moreover, it allows each time frame to attend to all time frames, neglecting the strong local correlations of speech signals. This study presents a simple yet effective sparse self-attention for speech enhancement, called ripple attention, which simultaneously performs fine- and coarse-grained modeling for local and global dependencies, respectively. Specifically, we employ local band attention to enable each frame to attend to its closest neighbor frames in a window at fine granularity, while employing dilated attention outside the window to model the global dependencies at a coarse granularity. We evaluate the efficacy of our ripple attention for speech enhancement on two commonly used training objectives. Extensive experimental results consistently confirm the superior performance of the ripple attention design over standard full self-attention, blockwise attention, and dual-path attention (Sep-Former) in terms of speech quality and intelligibility. Qiquan Zhang, Hongxu Zhu, Xinyuan Qian 0001, Zhaoheng Ni, Haizhou Li 0001 |
ICASSP | 6 |
| 2023 | Exploiting Modality-Invariant Feature for Robust Multimodal Emotion Recognition with Missing ModalitiesabstractMultimodal emotion recognition leverages complementary information across modalities to gain performance. However, we cannot guarantee that the data of all modalities are always present in practice. In the studies to predict the missing data across modalities, the inherent difference between heterogeneous modalities, namely the modality gap, presents a challenge. To address this, we propose to use invariant features for a missing modality imagination network (IF-MMIN) which includes two novel mechanisms: 1) an invariant feature learning strategy that is based on the central moment discrepancy (CMD) distance under the full-modality scenario; 2) an invariant feature based imagination module (IF-IM) to alleviate the modality gap during the missing modalities prediction, thus improving the robustness of multimodal joint representation. Comprehensive experiments on the benchmark dataset IEMOCAP demonstrate that the proposed model outperforms all baselines and invariantly improves the overall emotion recognition performance under uncertain missing-modality conditions. We release the code at: https://github.com/ZhuoYulang/IF-MMIN. Haolin Zuo, Rui Liu 0008, Jinming Zhao, Guanglai Gao, Haizhou Li 0001 |
ICASSP | 5 |
| 2023 | Local and Global Context Modeling with Relation Matching Task for Dialog Act RecognitionabstractIn dialog act recognition (DAR) of an utterance in a conversation, the prior studies have focused either on the global context using the whole utterances in the dialog, or the local context using the neighbouring utterance flow in the dialog. However, their methods attempt to deal with all types of dialogs indiscriminately. In this study, we propose a model to extract the local context information by an inter-utterance relation matching task (RMT), and a DAR framework to incorporate the local context information into a hierarchical network to fulfil both local and global context modeling. Extensive evaluations were conducted on a Mandarin dialog corpus and two benchmark English corpora. It is found that the different dialog types possess different window lengths for RMT, which is related to the length of subtopics in a given type of dialog. According to ablation experiments, the global information contributed more to the DAR in the hierarchical framework, while the contribution ratio of the local to the global context information was larger than 0.1. The results demonstrated that the proposed RMT and DAR framework significantly improved the DAR performance. Yuke Si, Yan Zhang 0004, Xiaobao Wang, Longbiao Wang, Jianwu Dang 0001, Chng Eng Siong, Haizhou Li 0001 |
IJCNN | 8 |
| 2023 | Betray Oneself: A Novel Audio DeepFake Detection Model via Mono-to-Stereo Conversion
Rui Liu 0008, Guanglai Gao, Haizhou Li 0001 |
INTERSPEECH | 4 |
| 2023 | Explicit Intensity Control for Accented Text-to-speech
Rui Liu 0008, Haolin Zuo, De Hu, Guanglai Gao, Haizhou Li 0001 |
INTERSPEECH | 5 |
| 2023 | Target Active Speaker Detection with Audio-visual Cues
Yidi Jiang, Ruijie Tao, Zexu Pan, Haizhou Li 0001 |
INTERSPEECH | 4 |
| 2023 | Self-Supervised Acoustic Word Embedding Learning via Correspondence Transformer Encoder
Jingru Lin, Xianghu Yue, Junyi Ao, Haizhou Li 0001 |
INTERSPEECH | 4 |
| 2023 | PIAVE: A Pose-Invariant Audio-Visual Speaker Extraction NetworkabstractIt is common in everyday spoken communication that we look at the turning head of a talker to listen to his/her voice. Humans see the talker to listen better, so do machines. However, previous studies on audio-visual speaker extraction have not effectively handled the varying talking face. This paper studies how to take full advantage of the varying talking face. We propose a Pose-Invariant Audio-Visual Speaker Extraction Network (PIAVE) that incorporates an additional pose-invariant view to improve audio-visual speaker extraction. Specifically, we generate the pose-invariant view from each original pose orientation, which enables the model to receive a consistent frontal view of the talker regardless of his/her head pose, therefore, forming a multi-view visual input for the speaker. Experiments on the multi-view MEAD and in-the-wild LRS3 dataset demonstrate that PIAVE outperforms the state-of-the-art and is more robust to pose variations. Meng Ge, Zhizheng Wu 0001, Haizhou Li 0001 |
INTERSPEECH | 4 |
| 2023 | High-Quality Automatic Voice Over with Accurate Alignment: Supervision through Self-Supervised Discrete Speech Units
Junchen Lu, Berrak Sisman, Mingyang Zhang 0003, Haizhou Li 0001 |
INTERSPEECH | 4 |
| 2023 | CoBERT: Self-Supervised Speech Representation Learning Through Code Representation Learning
Chutong Meng, Junyi Ao, Tom Ko, Mingxuan Wang, Haizhou Li 0001 |
INTERSPEECH | 5 |
| 2023 | EEG-based Auditory Attention Detection with Spatiotemporal Graph and Graph Convolutional Network
Ruicong Wang, Siqi Cai 0002, Haizhou Li 0001 |
INTERSPEECH | 3 |
| 2023 | Speaker Extraction with Detection of Presence and Absence of Target Speakers
Marvin Borsdorf, Zexu Pan, Haizhou Li 0001, Yangjie Wei |
INTERSPEECH | 4 |
| 2023 | Slow-Fast Time Parameter Aggregation Network for Class-Incremental Lip ReadingabstractClass incremental learning has yet to be explored in the field of lip-reading, which can circumvent data privacy issues and avoid the high training costs associated with joint training. In this paper, we introduce a benchmark for Class-Incremental Lip-Reading (CILR). To simultaneously improve the plasticity for new classes and stability for old classes in incremental learning, we propose a Slow-Fast Time Parameter Aggregation Network (TPAN) that decouples representation learning of new and old knowledge, taking into account the task characteristics of lip-reading. The TPAN comprises two dynamically evolving branches: one that uses fast gradient descent and the other employs slow momentum updates to retain old knowledge while adapting to new knowledge. Additionally, to achieve efficient knowledge transfer of the incremental model, we design a Hybrid Sequence-Distribution Distillation (HSDD) strategy to transfer knowledge in temporal feature view and classification probability view. We present a comprehensive comparison of the proposed method and previous state-of-the-art class incremental learning methods on the most commonly used lip-reading datasets LRW and LRW1000. The experimental result show that the proposed method can reduce the effect of catastrophic forgetting and improve the incremental accuracy. Xueyi Zhang 0001, Tao Wang 0074, Jun Tang 0001, Songyang Lao, Haizhou Li 0001 |
ACM Multimedia | 6 |
| 2023 | Disentangling Voice and Content with Self-Supervision for Speaker RecognitionabstractFor speaker recognition, it is difficult to extract an accurate speaker representation from speech because of its mixture of speaker traits and content. This paper proposes a disentanglement framework that simultaneously models speaker traits and content variability in speech. It is realized with the use of three Gaussian inference layers, each consisting of a learnable transition model that extracts distinct speech components. Notably, a strengthened transition model is specifically designed to model complex speech dynamics. We also propose a self-supervision method to dynamically disentangle content without the use of labels other than speaker identities. The efficacy of the proposed framework is validated via experiments conducted on the VoxCeleb and SITW datasets with 9.56\% and 8.24\% average reductions in EER and minDCF, respectively. Since neither additional model training nor data is specifically needed, it is easily applicable in practical use. Tianchi Liu 0004, Kong-Aik Lee, Qiongqiong Wang, Haizhou Li 0001 |
NeurIPS | 4 |
| 2023 | GrammarGPT: Exploring Open-Source LLMs for Native Chinese Grammatical Error Correction with Supervised Fine-Tuning
Yaxin Fan, Feng Jiang 0007, Peifeng Li 0001, Haizhou Li 0001 |
NLPCC (3) | 4 |
| 2023 | Enhancing Subject-Independent EEG-Based Auditory Attention Decoding with WGAN and Pearson Correlation CoefficientabstractElectroencephalography (EEG) related research faces a significant challenge of subject independence due to the variation in brain signals and responses among individuals. While deep learning models hold promise in addressing this challenge, their effectiveness depends on large datasets for training and generalization across participants. To overcome this limitation, we propose a solution to the above limitation by increasing the size and quality of training data for subject-independent auditory attention decoding (AAD) using EEG with deep learning. Specifically, our method employs a Wasserstein Generative Adversarial Network (WGAN) to generate synthetic data, with Pearson correlation filtering the most realistic samples. We evaluated this method on a publicly available dataset of selective auditory attention experiments and showed superior performance in subject-independent AAD performance. The mixed training set, consisting of both real and artificial data generated by the WGAN+Pearson Correlation Coefficient, demonstrated approximately 4% improvement in AAD accuracy for a 1-second window. These results demonstrate that deep learning remains a viable approach to overcoming data scarcity in subject-independent AAD tasks based on EEG. Moreover, the proposed method has the potential to improve the generalization and reliability of EEG classification tasks. Saurav Pahuja, Gabriel Ivucic, Felix Putze, Siqi Cai 0002, Haizhou Li 0001, Tanja Schultz |
SMC | 5 |
| 2023 | DNN controlled adaptive front-end for replay attack detection systemsabstractDeveloping robust countermeasures to protect automatic speaker verification systems against replay spoofing attacks is a well-recognized challenge. Current approaches to spoofing detection are generally based on a fixed front-end, typically a time-invariant filter bank, followed by a machine learning back-end. In this paper, we propose a novel approach whereby the front-end comprises an adaptive filter bank with a deep neural network-based controller, which is jointly trained along with a neural network back-end. Specifically, the deep neural network-based adaptive filter controller tunes the selectivity and sensitivity of the front-end filter bank at every frame to capture replay-related artefacts. We demonstrate the effectiveness of the proposed framework in spoofing attack detection on a synthesized dataset and ASVSpoof 2019 and ASVSpoof 2021 challenge datasets in terms of equal error rate and its ability to capture artefacts that differentiate replayed signals from genuine ones in comparison to conventional non-adaptive front-end. Buddhi Wickramasinghe, Eliathamby Ambikairajah, Vidhyasaharan Sethu, Julien Epps, Haizhou Li 0001, Ting Dang |
Speech Commun. | 5 |
| 2023 | Time-Domain Speech Separation Networks With Graph Encoding AuxiliaryabstractEnd-to-end time-domain speech separation with masking strategy has shown its performance advantage, where a 1-D convolutional layer is used as the speech encoder to encode a sliding window of waveform to a latent feature representation, i.e. an embedding vector. A large window leads to low resolution in the speech processing, on the other hand, a small window offers high resolution but at the expense of high computational cost. In this work, we propose a graph encoding technique to model the fine structural knowledge of speech samples in a window of reasonable size. Specifically, we build a graph representation for each latent representation, and encode the structural details with a graph convolutional network encoder. The encoded graph feature representation complements the original latent feature representation and benefits the separation and reconstruction of speech. Experiments on various models and datasets show that our proposed encoding technique significantly improves the speech quality over other time-domain speech encoders. Zexu Pan, Meng Ge, Zhen Yang 0001, Haizhou Li 0001 |
IEEE Signal Process. Lett. | 5 |
| 2023 | Towards Zero-Shot Multi-Speaker Multi-Accent Text-to-Speech SynthesisabstractThis paper presents a framework towards multi-accent neural text-to-speech synthesis for zero-shot multi-speaker, which employs an encoder-decoder architecture and an accent classifier to control the pronunciation variation from the encoder. The encoder and decoder are pre-trained on a large-scale multi-speaker corpus. The accent-informed encoder outputs are taken by the attention-based decoder to generate accented prosody. This framework allows for fine-tuning with limited training data from multiple accents, and is able to generate accented speech for unseen speakers. Both objective and subjective evaluations confirm the effectiveness of the proposed framework. Mingyang Zhang 0003, Xuehao Zhou, Zhizheng Wu 0001, Haizhou Li 0001 |
IEEE Signal Process. Lett. | 4 |
| 2023 | TTS-Guided Training for Accent Conversion Without Parallel DataabstractAccent Conversion (AC) seeks to change the accent of speech from one (source) to another (target) while preserving the speech content and speaker identity. However, many existing AC approaches rely on source-target parallel speech data during training or reference speech at run-time. We propose a novel accent conversion framework without the need for either parallel data or reference speech. Specifically, a text-to-speech (TTS) system is first pretrained with target-accented speech data. This TTS model and its hidden representations are expected to be associated only with the target accent. Then, a speech encoder is trained to convert the accent of the speech under the supervision of the pretrained TTS model. In doing so, the source-accented speech and its corresponding transcription are forwarded to the speech encoder and the pretrained TTS, respectively. The output of the speech encoder is optimized to be the same as the text embedding in the TTS system. At run-time, the speech encoder is combined with the pretrained speech decoder to convert the source-accented speech toward the target. In the experiments, we converted English with two source accents (Chinese/Indian) to the target accent (American/British/Canadian). Both objective metrics and subjective listening tests successfully validate that the proposed approach generates speech samples that are close to the target accent with high speech quality. Yi Zhou 0020, Zhizheng Wu 0001, Mingyang Zhang 0003, Xiaohai Tian, Haizhou Li 0001 |
IEEE Signal Process. Lett. | 5 |
| 2023 | Emotion Intensity and its Control for Emotional Voice ConversionabstractEmotional voice conversion (EVC) seeks to convert the emotional state of an utterance while preserving the linguistic content and speaker identity. In EVC, emotions are usually treated as discrete categories overlooking the fact that speech also conveys emotions with various intensity levels that the listener can perceive. In this paper, we aim to explicitly characterize and control the intensity of emotion. We propose to disentangle the speaker style from linguistic content and encode the speaker style into a style embedding in a continuous space that forms the prototype of emotion embedding. We further learn the actual emotion encoder from an emotion-labelled database and study the use of relative attributes to represent fine-grained emotion intensity. To ensure emotional intelligibility, we incorporateemotion classification lossandemotion embedding similarity lossinto the training of the EVC network. As desired, the proposed network controls the fine-grained emotion intensity in the output speech. Through both objective and subjective evaluations, we validate the effectiveness of the proposed network for emotional expressiveness and emotion intensity control. Kun Zhou 0003, Berrak Sisman, Rajib Rana, Björn W. Schuller, Haizhou Li 0001 |
IEEE Trans. Affect. Comput. | 5 |
| 2023 | Speech Synthesis With Mixed EmotionsabstractEmotional speech synthesis aims to synthesize human voices with various emotional effects. The current studies are mostly focused on imitating an averaged style belonging to a specific emotion type. In this paper, we seek to generate speech with a mixture of emotions at run-time. We propose a novel formulation that measures the relative difference between the speech samples of different emotions. We then incorporate our formulation into a sequence-to-sequence emotional text-to-speech framework. During the training, the framework does not only explicitly characterize emotion styles but also explores the ordinal nature of emotions by quantifying the differences with other emotions. At run-time, we control the model to produce the desired emotion mixture by manually defining an emotion attribute vector. The objective and subjective evaluations have validated the effectiveness of the proposed framework. To our best knowledge, this research is the first study on modelling, synthesizing, and evaluating mixed emotions in speech. Kun Zhou 0003, Berrak Sisman, Rajib Rana, Björn W. Schuller, Haizhou Li 0001 |
IEEE Trans. Affect. Comput. | 5 |
| 2023 | PoLyScriber: Integrated Fine-Tuning of Extractor and Lyrics Transcriber for Polyphonic MusicabstractLyrics transcription of polyphonic music is challenging as the background music affects lyrics intelligibility. Typically, lyrics transcription can be performed by a two-step pipeline, i.e. a singing vocal extraction front end, followed by a lyrics transcriber back end, where the front end and back end are trained separately. Such a two-step pipeline suffers from both imperfect vocal extraction and mismatch between front end and back end. In this work, we propose a novel end-to-end integrated fine-tuning framework, that we call PoLyScriber, to globally optimize the vocal extractor front end and lyrics transcriber back end for lyrics transcription in polyphonic music. The experimental results show that our proposed PoLyScriber achieves substantial improvements over the existing approaches on publicly available test datasets. Xiaoxue Gao, Chitralekha Gupta, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2023 | Audio-Visual Cross-Attention Network for Robotic Speaker TrackingabstractAudio-visual signals can be used jointly for robotic perception as they complement each other. Such multi-modal sensory fusion has a clear advantage, especially under noisy acoustic conditions. Speaker localization, as an essential robotic function, was traditionally solved as a signal processing problem that now increasingly finds deep learning solutions. The question is how to fuse audio-visual signals in an effective way. Speaker tracking is not only more desirable, but also potentially more accurate than speaker localization because it explores the speaker's temporal motion dynamics for smoothed trajectory estimation. However, due to the lack of large annotated dataset, speaker tracking is not well studied as speaker localization. In this paper, we study robotic speaker Direction of Arrival (DoA) estimation with a focus on audio-visual fusion and tracking methodology. We propose a Cross-Modal Attentive Fusion (CMAF) mechanism, which explores self-attention to learn intra-modal temporal dependencies, and cross-attention mechanism for inter-modal alignment. We also collect a realistic dataset on a robotic platform to support the study. The experimental results demonstrate that our proposed network outperforms the state-of-the-art audio-visual localization and tracking methods under noisy conditions, with an improved accuracy of 5.82% and 3.62% at SNR = −20 dB, respectively. Xinyuan Qian 0001, Zhengdong Wang, Guohui Guan 0001, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2023 | Self-Supervised Training of Speaker Encoder With Multi-Modal Diverse Positive PairsabstractWe study a novel neural speaker encoder and its training strategies for speaker recognition without using any identity labels. The speaker encoder is trained to extract a fixed dimensional speaker embedding from a spoken utterance of variable length. Contrastive learning is a typical self-supervised learning technique. However, the contrastive learning of the speaker encoder depends very much on the sampling strategy of positive and negative pairs. It is common that we sample a positive pair of segments from the same utterance. Unfortunately, such a strategy, denoted as poor-man's positive pairs (PPP), lacks the necessary diversity. In this work, we propose a multi-modal contrastive learning technique with novel sampling strategies. By cross-referencing between speech and face data, we find diverse positive pairs (DPP) for contrastive learning, thus improving the robustness of speaker encoder. We train the speaker encoder on the VoxCeleb2 dataset without any speaker labels, and achieve an equal error rate (EER) of 2.89%, 3.17% and 6.27% under the proposed progressive clustering strategy, and an EER of 1.44%, 1.77% and 3.27% under the two-stage learning strategy with pseudo labels, on the three test sets of VoxCeleb1. This novel solution outperforms the state-of-the-art self-supervised learning methods by a large margin, at the same time, achieves comparable results with the supervised learning counterpart. We also evaluate our self-supervised learning technique on the LRS2 and LRW datasets, where speaker information is unavailable. All experiments suggest that the proposed neural architecture and sampling strategies are robust across datasets. Ruijie Tao, Kong-Aik Lee, Rohan Kumar Das, Ville Hautamäki, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2023 | STFF-SM: Steganalysis Model Based on Spatial and Temporal Feature Fusion for Speech StreamsabstractThe real-time detection of speech steganography in Voice-over-Internet-Protocol (VoIP) scenarios remains an open problem, as it requires steganalysis methods to perform for low-intensity embeddings and short-sample inputs, as well as provide rapid detection results. To address these challenges, this paper presents a novel steganalysis model based on spatial and temporal feature fusion (STFF-SM). Differing from the existing methods, we take both the integer and fractional pitch delays as input, and design subframe-stitch module to organically integrate subframe-wise integer delays and frame-wise fractional pitch delays. Further, we design a spatial fusion module based on pre-activation residual convolution to extract the pitch spatial features and gradually increase their dimensions to discover finer steganographic distortions to enhance the detection effect, where a Group-Squeeze-Weighting block is introduced to alleviate the information loss in the process of increasing the feature dimension. In addition, we design a temporal fusion module to extract pitch temporal features using the stacked LSTM, where a Gated Feed-Forward Network is introduced to learn the interaction between different feature maps while suppressing the features that are not useful for detection. We evaluated the performance of STFF-SM through comprehensive experiments and comparisons with the state-of-the-art solutions. The experimental results demonstrate that STFF-SM can well meet the needs of real-time detection of speech steganography in VoIP streams, and outperforms the existing methods in detection performance, especially with low embedding strengths and short window sizes. Hui Tian 0002, Yiqin Qiu, Wojciech Mazurczyk, Haizhou Li 0001, Zhenxing Qian |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2023 | PoE: A Panel of Experts for Generalized Automatic Dialogue AssessmentabstractChatbots are expected to be knowledgeable across multiple domains, e.g. for daily chit-chat, exchange of information, and grounding in emotional situations. To effectively measure the quality of such conversational agents, a model-based automatic dialogue evaluation metric (ADEM) is expected to perform well across multiple domains. Despite significant progress, existing ADEMs tend to perform well only on data that are similar to its training data (overfit to its training domain). This calls for a domain-generalized metric that can assess dialogues of different characteristics. To this end, we propose aPanel of Experts(PoE), a multitask network that consists of a shared transformer encoder and a collection of lightweight adapters. The shared encoder captures the general knowledge of dialogues across domains, while each adapter specializes in one specific domain and serves as a domain expert. To validate the idea, we construct a high-quality multi-domain dialogue dataset leveraging data augmentation and pseudo-labeling. The PoE network is comprehensively assessed on 16 dialogue evaluation datasets spanning a wide range of dialogue domains. It achieves state-of-the-art performance in terms of mean Spearman correlation over all the evaluation datasets. It exhibits better zero-shot generalization than existing state-of-the-art ADEMs and the ability to easily adapt to new domains with few-shot transfer learning. Chen Zhang 0055, Luis Fernando D'Haro, Qiquan Zhang, Thomas Friedrichs, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2023 | A Time-Frequency Attention Module for Neural Speech EnhancementabstractSpeech enhancement plays an essential role in a wide range of speech processing applications. Recent studies on speech enhancement tend to investigate how to effectively capture the long-term contextual dependencies of speech signals to boost performance. However, these studies generally neglect the time-frequency (T-F) distribution information of speech spectral components, which is equally important for speech enhancement. In this paper, we propose a simple yet very effective network module, which we term the T-F attention (TFA) module, that uses two parallel attention branches, i.e., time-frame attention and frequency-channel attention, to explicitly exploit position information to generate a 2-D attention map to characterise the salient T-F speech distribution. We validate our TFA module as part of two widely used backbone networks (residual temporal convolution network and Transformer) and conduct speech enhancement with four most popular training objectives. Our extensive experiments demonstrate that our proposed TFA module consistently leads to substantial enhancement performance improvements in terms of the five most widely used objective metrics, with negligible parameter overheads. In addition, we further evaluate the efficacy of speech enhancement as a front-end for a downstream speech recognition task. Our evaluation results show that the TFA module significantly improves the robustness of the system to noisy conditions. Qiquan Zhang, Xinyuan Qian 0001, Zhaoheng Ni, Aaron Nicolson, Eliathamby Ambikairajah, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2023 | Optimization of Cross-Lingual Voice Conversion With Linguistics Losses to Reduce Foreign AccentsabstractCross-lingual voice conversion (XVC) transforms the speaker identity of a source speaker to that of a target speaker who speaks a different language. Due to the intrinsic differences between languages, the converted speech may carry an unwanted foreign accent. In this paper, we first investigate the intelligibility of the converted speech and confirm the performance degradation caused by the accent/intelligibility issue. With the goal of generating native-sounding speech, this paper further proposes a novel training scheme with two additional linguistic losses for speech waveform generation: 1) a frame-wise phonetic content loss derived from bottleneck features, and 2) an automatic speech recognition loss on characters. Experiments were conducted between English and Mandarin Chinese conversions. The experimental results confirmed that the generated speech sounds more natural with the proposed linguistic losses and the proposed solution significantly improves speech intelligibility. Yi Zhou 0020, Zhizheng Wu 0001, Xiaohai Tian, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2023 | Online Multi-Face Tracking With Multi-Modality Cascaded MatchingabstractTracking multiple faces online in unconstrained videos is a challenging problem as faces may appear drastically different over time and identities can be inferred only based on information available from past frames. Previous tracking methods focus on face information without reference to other modality information such as a person’s overall body appearance, leading to suboptimal performance. In this paper, we propose a new online multi-face tracking method, called online multi-face tracking with multi-modality cascaded matching (OMTMCM), to improve the tracking performance by using both face and body information. The proposed OMTMCM consists of two stages, namely detection alignment and detection association. In the first stage, a detection alignment module is designed to align face detection with body detection from the same person for the subsequent detection association. In the second stage, a cascaded matching module is designed to associate face detections across frames to locate trajectory of each target face by using both face and body information. Specifically, aligned face-body detections in the current frame are matched in a cascade manner with body and face features that are selected from past frames and stored in the designed feature memory. In this way, our method can track multiple faces online with both face and body information while eliminating the possibility of face detection and body detection from the same person being separately assigned with different identities. Experimental results demonstrate our method is on par with or better than other online tracking methods for multi-face tracking. Zhenyu Weng, Huiping Zhuang, Haizhou Li 0001, Balakrishnan Ramalingam, Mohan Rajesh Elara, Zhiping Lin 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2023 | Separable Convolution Network With Dual-Stream Pyramid Enhanced Strategy for Speech SteganalysisabstractSteganography based on fixed codebook has become one of the most important branches of speech steganography due to its high imperceptibility and having the largest available carrier space. As its countermeasure technique, this paper presents a novel steganalysis method based on separable convolution network (SepSteNet) with dual-stream pyramid enhanced strategy (DPES). Specifically, to better acquire discriminative representations, we design the pulse-aware separable block to capture the pulse correspondence along independent levels of pulse positions, where the pulse-aware excitation module is plugged to avoid noisy clue accumulation by adaptively emphasizing the salient part. Moreover, the global attending block is introduced to enhance correspondence features through calculating global responses at distinct subframes. In addition, to eliminate the negative impact of sample content, DPES is leveraged to incorporate cross-domain coherence features by the inverted connected dual-stream branches. With the original and calibration speech samples, two branches enable the correspondence of two detection feature domains to interact with each other to generate coherence features independent of sample content, thereby improving the detection performance. The performance of the presented method is comprehensively evaluated and compared with the state of the arts. The experimental results demonstrate that the presented method significantly outperforms the existing ones. Furthermore, DPES is shown to be a general enhancement strategy that can effectively improve the performance of the existing deep neural network for speech steganalysis. The source code for this work will be publicly available on GitHub. Yiqin Qiu, Hui Tian 0002, Haizhou Li 0001, Chin-Chen Chang 0001, Athanasios V. Vasilakos |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2023 | A Tandem Learning Rule for Effective Training and Rapid Inference of Deep Spiking Neural NetworksabstractSpiking neural networks (SNNs) represent the most prominent biologically inspired computing model for neuromorphic computing (NC) architectures. However, due to the nondifferentiable nature of spiking neuronal functions, the standard error backpropagation algorithm is not directly applicable to SNNs. In this work, we propose a tandem learning framework that consists of an SNN and an artificial neural network (ANN) coupled through weight sharing. The ANN is an auxiliary structure that facilitates the error backpropagation for the training of the SNN at the spike-train level. To this end, we consider the spike count as the discrete neural representation in the SNN and design an ANN neuronal activation function that can effectively approximate the spike count of the coupled SNN. The proposed tandem learning rule demonstrates competitive pattern recognition and regression capabilities on both the conventional frame- and event-based vision datasets, with at least an order of magnitude reduced inference time and total synaptic operations over other state-of-the-art SNN implementations. Therefore, the proposed tandem learning rule offers a novel solution to training efficient, low latency, and high-accuracy deep SNNs with low computing resources. Jibin Wu, Yansong Chua, Malu Zhang, Guoqi Li 0002, Haizhou Li 0001, Kay Chen Tan |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2022 | MDD-Eval: Self-Training on Augmented Data for Multi-Domain Dialogue EvaluationabstractChatbots are designed to carry out human-like conversations across different domains, such as general chit-chat, knowledge exchange, and persona-grounded conversations. To measure the quality of such conversational agents, a dialogue evaluator is expected to conduct assessment across domains as well. However, most of the state-of-the-art automatic dialogue evaluation metrics (ADMs) are not designed for multi-domain evaluation. We are motivated to design a general and robust framework, MDD-Eval, to address the problem. Specifically, we first train a teacher evaluator with human-annotated data to acquire a rating skill to tell good dialogue responses from bad ones in a particular domain and then, adopt a self-training strategy to train a new evaluator with teacher-annotated multi-domain data, that helps the new evaluator to generalize across multiple domains. MDD-Eval is extensively assessed on six dialogue evaluation benchmarks. Empirical results show that the MDD-Eval framework achieves a strong performance with an absolute improvement of 7% over the state-of-the-art ADMs in terms of mean Spearman correlation scores across all the evaluation benchmarks. Chen Zhang 0055, Luis Fernando D'Haro, Thomas Friedrichs, Haizhou Li 0001 |
AAAI | 4 |
| 2022 | Just Rank: Rethinking Evaluation with Word and Sentence SimilaritiesabstractWord and sentence embeddings are useful feature representations in natural language processing.However, intrinsic evaluation for embeddings lags far behind, and there has been no significant update since the past decade.Word and sentence similarity tasks have become the de facto evaluation method.It leads models to overfit to such evaluations, negatively impacting embedding models' development.This paper first points out the problems using semantic similarity as the gold standard for word and sentence embedding evaluations.Further, we propose a new intrinsic evaluation method called EvalRank, which shows a much stronger correlation with downstream tasks.Extensive experiments are conducted based on 60+ models and popular datasets to certify our judgments.Finally, the practical evaluation toolkit is released for future benchmarking purposes.1 Bin Wang 0040, C.-C. Jay Kuo, Haizhou Li 0001 |
ACL (1) | 3 |
| 2022 | M3ED: Multi-modal Multi-scene Multi-label Emotional Dialogue DatabaseabstractJinming Zhao, Tenggan Zhang, Jingwen Hu, Yuchen Liu, Qin Jin, Xinchao Wang, Haizhou Li. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Jinming Zhao, Tenggan Zhang, Jingwen Hu 0003, Yuchen Liu 0003, Qin Jin, Xinchao Wang, Haizhou Li 0001 |
ACL (1) | 7 |
| 2022 | Ego4D: Around the World in 3, 000 Hours of Egocentric VideoabstractWe introduce Ego4D, a massive-scale egocentric video dataset and benchmark suite. It offers 3,670 hours of dailylife activity video spanning hundreds of scenarios (household, outdoor, workplace, leisure, etc.) captured by 931 unique camera wearers from 74 worldwide locations and 9 different countries. The approach to collection is designed to uphold rigorous privacy and ethics standards, with consenting participants and robust de-identification procedures where relevant. Ego4D dramatically expands the volume of diverse egocentric video footage publicly available to the research community. Portions of the video are accompanied by audio, 3D meshes of the environment, eye gaze, stereo, and/or synchronized videos from multiple egocentric cameras at the same event. Furthermore, we present a host of new benchmark challenges centered around understanding the first-person visual experience in the past (querying an episodic memory), present (analyzing hand-object manipulation, audio-visual conversation, and social interactions), and future (forecasting activities). By publicly sharing this massive annotated dataset and benchmark suite, we aim to push the frontier of first-person perception. Project page: https://ego4d-data.org/ Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang 0007, Miao Liu 0007, Xingyu Liu 0001, Tushar Nagarajan, Ilija Radosavovic, Santhosh K. Ramakrishnan, Fiona Ryan, Jayant Sharma 0002, Michael Wray, Mengmeng Xu 0006, Eric Zhongcong Xu, Chen Zhao 0002, Siddhant Bansal, Dhruv Batra, Vincent Cartillier, Sean Crane, Tien Do, Morrie Doulaty, Akshay Erapalli, Christoph Feichtenhofer, Adriano Fragomeni, Qichen Fu, Abrham Gebreselasie, Cristina González, James Hillis, Xuhua Huang, Yifei Huang 0002, Wenqi Jia 0001, Weslie Khoo, Jáchym Kolár, Satwik Kottur, Anurag Kumar 0003, Federico Landini, Yanghao Li, Zhenqiang Li 0002, Karttikeya Mangalam, Raghava Modhugu, Jonathan Munro, Tullie Murrell, Takumi Nishiyasu, Will Price, Paola Ruiz Puentes, Merey Ramazanova, Leda Sari, Kiran K. Somasundaram, Audrey Southerland, Yusuke Sugano, Ruijie Tao, Minh Vo, Xindi Wu, Takuma Yagi, Ziwei Zhao 0003, Yunyi Zhu, Pablo Andrés Arbeláez, David Crandall, Dima Damen, Giovanni Maria Farinella, Christian Fügen, Bernard Ghanem, Vamsi K. Ithapu, C. V. Jawahar, Hanbyul Joo, Kris Makoto Kitani, Haizhou Li 0001, Richard A. Newcombe, Aude Oliva, Hyun Soo Park, James M. Rehg, Yoichi Sato 0001, Jianbo Shi, Zheng Shou 0001, Antonio Torralba 0001, Lorenzo Torresani, Mingfei Yan, Jitendra Malik |
CVPR | 74 |
| 2022 | Generate, Discriminate and Contrast: A Semi-Supervised Sentence Representation Learning FrameworkabstractMost sentence embedding techniques heavily rely on expensive human-annotated sentence pairs as the supervised signals.Despite the use of large-scale unlabeled data, the performance of unsupervised methods typically lags far behind that of the supervised counterparts in most downstream tasks.In this work, we propose a semi-supervised sentence embedding framework, GenSE, that effectively leverages large-scale unlabeled data.Our method include three parts: 1) Generate: A generator/discriminator model is jointly trained to synthesize sentence pairs from open-domain unlabeled corpus; 2) Discriminate: Noisy sentence pairs are filtered out by the discriminator to acquire high-quality positive and negative sentence pairs; 3) Contrast: A prompt-based contrastive approach is presented for sentence representation learning with both annotated and synthesized data.Comprehensive experiments show that GenSE achieves an average correlation score of 85.19 on the STS datasets and consistent performance improvement on four domain adaptation tasks, significantly surpassing the state-of-the-art methods and convincingly corroborating its effectiveness and generalization ability. 1 Yiming Chen 0010, Yan Zhang 0004, Bin Wang 0040, Zuozhu Liu, Haizhou Li 0001 |
EMNLP | 5 |
| 2022 | Analyzing and Evaluating Faithfulness in Dialogue SummarizationabstractDialogue summarization is abstractive in nature, making it suffer from factual errors.The factual correctness of summaries has the highest priority before practical applications.Many efforts have been made to improve faithfulness in text summarization.However, there is a lack of systematic study on dialogue summarization systems.In this work, we first perform the fine-grained human analysis on the faithfulness of dialogue summaries and observe that over 35% of generated summaries are faithfully inconsistent respective the source dialogues.Furthermore, we present a new model-level faithfulness evaluation method.It examines generation models with multi-choice questions created by rule-based transformations.Experimental results show that our evaluation schema is a strong proxy for the factual correctness of summarization models.The humanannotated faithfulness samples and the evaluation toolkit are released to facilitate future research toward faithful dialogue summarization.Code Bin Wang 0040, Chen Zhang 0020, Yan Zhang 0004, Yiming Chen 0010, Haizhou Li 0001 |
EMNLP | 5 |
| 2022 | FineD-Eval: Fine-grained Automatic Dialogue-Level EvaluationabstractRecent model-based reference-free metrics for open-domain dialogue evaluation exhibit promising correlations with human judgment 1 .However, they either perform turn-level evaluation or look at a single dialogue quality dimension.One would expect a good evaluation metric to assess multiple quality dimensions at the dialogue level.To this end, we are motivated to propose a multi-dimensional dialogue-level metric, which consists of three sub-metrics with each targeting a specific dimension.The submetrics are trained with novel self-supervised objectives and exhibit strong correlations with human judgment for their respective dimensions.Moreover, we explore two approaches to combine the sub-metrics: metric ensemble and multitask learning.Both approaches yield a holistic metric that significantly outperforms individual sub-metrics.Compared to the existing state-of-the-art metric, the combined metrics achieve around 16% relative improvement on average across three high-quality dialoguelevel evaluation benchmarks. Chen Zhang 0055, Luis Fernando D'Haro, Qiquan Zhang, Thomas Friedrichs, Haizhou Li 0001 |
EMNLP | 5 |
| 2022 | Experts Versus All-Rounders: Target Language Extraction for Multiple Target LanguagesabstractTarget language extraction (TLE) is a novel task in the field of selective auditory attention, which seeks to extract all speech signals that are spoken in a target language from other sources in a multilingual cocktail party. In our prior studies, a TLE model was trained to extract a predefined, single target language, referred to as Single-TLE. In this paper, we extend the Single-TLE framework to Multi-TLE. Multi-TLE models can also extract all speech signals of one specific target language, but they are optimized on a set of multiple target languages during training. As such, they learn the characteristics of several target languages and can replace multiple Single-TLE models without retraining. We perform experiments on the GlobalPhoneMCP database and incorporate a dynamic language mixing scheme for training. The Multi-TLE model does not only outperform Single-TLE models, but when given a language ID as additional input, it is also able to extract the speech of a specific target language from a mixture which contains multiple learned target languages. Marvin Borsdorf, Kevin Scheck, Haizhou Li 0001, Tanja Schultz |
ICASSP | 3 |
| 2022 | Genre-Conditioned Acoustic Models for Automatic Lyrics Transcription of Polyphonic MusicabstractLyrics transcription of polyphonic music is challenging not only because the singing vocals are corrupted by the background music, but also because the background music and the singing style vary across music genres, such as pop, metal, and hip hop, which affects lyrics intelligibility of the song in different ways. In this work, we propose to transcribe the lyrics of polyphonic music using a novel genre-conditioned network. The proposed network adopts pre-trained model parameters, and incorporates the genre adapters between layers to capture different genre peculiarities for lyrics-genre pairs, thereby only requiring lightweight genre-specific parameters for training. Our experiments show that the proposed genre-conditioned network outperforms the existing lyrics transcription systems. Xiaoxue Gao, Chitralekha Gupta, Haizhou Li 0001 |
ICASSP | 3 |
| 2022 | L-SpEx: Localized Target Speaker ExtractionabstractSpeaker extraction aims to extract the target speaker’s voice from a multi-talker speech mixture given an auxiliary reference utterance. Recent studies show that speaker extraction benefits from the location or direction of the target speaker. However, these studies assume that the target speaker’s location is known in advance or detected by an extra visual cue, e.g., face image or video. In this paper, we propose an end-to-end localized target speaker extraction on pure speech cues, that is called L-SpEx. Specifically, we design a speaker localizer driven by the target speaker’s embedding to extract the spatial features, including direction-of-arrival (DOA) of the target speaker and beamforming output. Then, the spatial cues and target speaker’s embedding are both used to form a top-down auditory attention to the target speaker. Experiments on the multi-channel reverberant dataset called MCLibri2Mix show that our L-SpEx approach significantly outperforms the baseline system. Meng Ge, Chenglin Xu, Longbiao Wang, Chng Eng Siong, Jianwu Dang 0001, Haizhou Li 0001 |
ICASSP | 6 |
| 2022 | MFA: TDNN with Multi-Scale Frequency-Channel Attention for Text-Independent Speaker Verification with Short UtterancesabstractThe time delay neural network (TDNN) represents one of the state-of-the-art of neural solutions to text-independent speaker verification. However, they require a large number of filters to capture the speaker characteristics at any local frequency region. In addition, the performance of such systems may degrade under short utterance scenarios. To address these issues, we propose a multi-scale frequency-channel attention (MFA), where we characterize speakers at different scales through a novel dual-path design which consists of a convolutional neural network and TDNN. We evaluate the proposed MFA on the VoxCeleb database and observe that the proposed framework with MFA can achieve state-of-the-art performance while reducing parameters and computation complexity. Further, the MFA mechanism is found to be effective for speaker verification with short test utterances. Tianchi Liu 0004, Rohan Kumar Das, Kong-Aik Lee, Haizhou Li 0001 |
ICASSP | 4 |
| 2022 | Visualtts: TTS with Accurate Lip-Speech Synchronization for Automatic Voice OverabstractIn this paper, we formulate a novel task to synthesize speech in sync with a silent pre-recorded video, denoted as automatic voice over (AVO). Unlike traditional speech synthesis, AVO seeks to generate not only human-sounding speech, but also perfect lip-speech synchronization. A natural solution to AVO is to condition the speech rendering on the temporal progression of lip sequence in the video. We propose a novel text-to-speech model that is conditioned on visual input, named VisualTTS, for accurate lip-speech synchronization. The proposed VisualTTS adopts two novel mechanisms that are 1) textual-visual attention, and 2) visual fusion strategy during acoustic decoding, which both contribute to forming accurate alignment between the input text content and lip motion in input lip sequence. Experimental results show that VisualTTS achieves accurate lip-speech synchronization and outperforms all baseline systems. Junchen Lu, Berrak Sisman, Rui Liu 0008, Mingyang Zhang 0003, Haizhou Li 0001 |
ICASSP | 5 |
| 2022 | Self-Supervised Speaker Recognition with Loss-Gated LearningabstractIn self-supervised learning for speaker recognition, pseudo labels are useful as the supervision signals. It is a known fact that a speaker recognition model doesn’t always benefit from pseudo labels due to their unreliability. In this work, we observe that a speaker recognition network tends to model the data with reliable labels faster than those with unreliable labels. This motivates us to study a loss-gated learning (LGL) strategy, which extracts the reliable labels through the fitting ability of the neural network during training. With the proposed LGL, our speaker recognition model obtains a 46.3% performance gain over the system without it. Further, the proposed self-supervised speaker recognition with LGL trained on the VoxCeleb2 dataset without any labels achieves an equal error rate of 1.66% on the VoxCeleb1 original test set. Ruijie Tao, Kong-Aik Lee, Rohan Kumar Das, Ville Hautamäki, Haizhou Li 0001 |
ICASSP | 5 |
| 2022 | A Hybrid Learning Framework for Deep Spiking Neural Networks with One-Spike Temporal CodingabstractBio-inspired spiking neural networks (SNNs) are compelling candidates for spatio-temporal information processing on ultra-low power neuromorphic computing chips. However, the existing SNN training methods have not fully exploited the temporal information of spikes that plays a critical role in sparse information representation and communication. Hereby, we present a hybrid learning framework for deep SNNs with one-spike temporal coding to make full utilization of the spike timing. We first propose a novel ANN-to-SNN conversion method based on forward propagation mechanisms of ANNs and SNNs to offer a good initialization for SNNs. The performance of the converted SNN is further improved by training with a timing-based backpropagation (BP) method. Experimental results demonstrate that the proposed hybrid learning framework can achieve competitive accuracies on both visual and audio recognition tasks with significantly improved training efficiency over direct SNN BP methods. Jibin Wu, Malu Zhang, Qi Liu 0005, Haizhou Li 0001 |
ICASSP | 5 |
| 2022 | ADD 2022: the first Audio Deep Synthesis Detection ChallengeabstractAudio deepfake detection is an emerging topic, which was included in the ASVspoof 2021. However, the recent shared tasks have not covered many real-life and challenging scenarios. The first Audio Deep synthesis Detection challenge (ADD) was motivated to fill in the gap. The ADD 2022 includes three tracks: low-quality fake audio detection (LF), partially fake audio detection (PF) and audio fake game (FG). The LF track focuses on dealing with bona fide and fully fake utterances with various real-world noises etc. The PF track aims to distinguish the partially fake audio from the real. The FG track is a rivalry game, which includes two tasks: an audio generation task and an audio fake detection task. In this paper, we describe the datasets, evaluation metrics, and protocols. We also report major findings that reflect the recent advances in audio deepfake detection tasks. Jiangyan Yi, Ruibo Fu, Jianhua Tao 0001, Shuai Nie 0001, Haoxin Ma, Chenglong Wang 0001, Tao Wang 0074, Zhengkun Tian, Ye Bai 0001, Cunhang Fan, Shan Liang 0007, Shuai Zhang 0014, Xinrui Yan, Zhengqi Wen, Haizhou Li 0001 |
ICASSP | 17 |
| 2022 | Time-Frequency Attention for Monaural Speech EnhancementabstractMost studies on speech enhancement generally don’t explicitly consider the energy distribution of speech in time-frequency (T-F) representation, which is important for accurate prediction of mask or spectra. In this paper, we present a simple yet effective T-F attention (TFA) module, where a 2-D attention map is produced to provide differentiated weights to the spectral components of T-F representation. To validate the effectiveness of our proposed TFA module, we use the residual temporal convolution network (ResTCN) as the backbone network and conduct extensive experiments on two commonly used training targets. Our experiments demonstrate that applying our TFA module significantly improves the performance in terms of five objective evaluation metrics with negligible parameter overhead. The evaluation results show that the proposed ResTCN with the TFA module (ResTCN+TFA) consistently outperforms other baselines by a large margin. Qiquan Zhang, Zhaoheng Ni, Aaron Nicolson, Haizhou Li 0001 |
ICASSP | 5 |
| 2022 | Memobert: Pre-Training Model with Prompt-Based Learning for Multimodal Emotion RecognitionabstractMultimodal emotion recognition study is hindered by the lack of labelled corpora in terms of scale and diversity, due to the high annotation cost and label ambiguity. In this paper, we propose a multimodal pre-training model MEmoBERT for multimodal emotion recognition, which learns multimodal joint representations through self-supervised learning from a self-collected large-scale unlabeled video data that come in sheer volume. Furthermore, unlike the conventional "pre-train, finetune" paradigm, we propose a prompt-based method that reformulates the downstream emotion classification task as a masked text prediction one, bringing the downstream task closer to the pre-training. Extensive experiments on two benchmark datasets, IEMOCAP and MSP-IMPROV, show that our proposed MEmoBERT significantly enhances emotion recognition performance. Jinming Zhao, Qin Jin, Xinchao Wang, Haizhou Li 0001 |
ICASSP | 5 |
| 2022 | Accurate Emotion Strength Assessment for Seen and Unseen Speech Based on Data-Driven Deep LearningabstractEmotion classification of speech and assessment of the emotion strength are required in applications such as emotional text-to-speech and voice conversion. The emotion attribute ranking function based on Support Vector Machine (SVM) was proposed to predict emotion strength for emotional speech corpus. However, the trained ranking function doesn't generalize to new domains, which limits the scope of applications, especially for out-of-domain or unseen speech. In this paper, we propose a data-driven deep learning model, i.e. StrengthNet, to improve the generalization of emotion strength assessment for seen and unseen speech. This is achieved by the fusion of emotional data from various domains. We follow a multi-task learning network architecture that includes an acoustic encoder, a strength predictor, and an auxiliary emotion predictor. Experiments show that the predicted emotion strength of the proposed StrengthNet is highly correlated with ground truth scores for both seen and unseen speech. We release the source codes at: https://github.com/ttslr/StrengthNet. Rui Liu 0008, Berrak Sisman, Björn W. Schuller, Guanglai Gao, Haizhou Li 0001 |
INTERSPEECH | 5 |
| 2022 | Pre-Training Transformer Decoder for End-to-End ASR Model with Unpaired Speech DataabstractThis paper studies a novel pre-training technique with unpaired speech data, Speech2C, for encoder-decoder based automatic speech recognition (ASR).Within a multi-task learning framework, we introduce two pre-training tasks for the encoderdecoder network using acoustic units, i.e., pseudo codes, derived from an offline clustering model.One is to predict the pseudo codes via masked language modeling in encoder output, like HuBERT model, while the other lets the decoder learn to reconstruct pseudo codes autoregressively instead of generating textual scripts.In this way, the decoder learns to reconstruct original speech information with codes before learning to generate correct text.Comprehensive experiments on the LibriSpeech corpus show that the proposed Speech2C can relatively reduce the word error rate (WER) by 19.2% over the method without decoder pre-training, and also outperforms significantly the state-of-the-art wav2vec 2.0 and Hu-BERT on fine-tuning subsets of 10h and 100h.We release our code and model at https://github.com/microsoft/SpeechT5/tree/main/Speech2C. Junyi Ao, Shujie Liu 0001, Haizhou Li 0001, Tom Ko, Li-Rong Dai 0001, Jinyu Li 0001, Yao Qian, Furu Wei |
INTERSPEECH | 5 |
| 2022 | Blind Language Separation: Disentangling Multilingual Cocktail Party Voices by Language
Marvin Borsdorf, Kevin Scheck, Haizhou Li 0001, Tanja Schultz |
INTERSPEECH | 3 |
| 2022 | Disentanglement of Emotional Style and Speaker Identity for Expressive Voice ConversionabstractExpressive voice conversion performs identity conversion for emotional speakers by jointly converting speaker identity and emotional style. Due to the hierarchical structure of speech emotion, it is challenging to disentangle the emotional style for different speakers. Inspired by the recent success of speaker disentanglement with variational autoencoder (VAE), we propose an any-to-any expressive voice conversion framework, that is called StyleVC. StyleVC is designed to disentangle linguistic content, speaker identity, pitch, and emotional style information. We study the use of style encoder to model emotional style explicitly. At run-time, StyleVC converts both speaker identity and emotional style for arbitrary speakers. Experiments validate the effectiveness of our proposed framework in both objective and subjective evaluations. Zongyang Du, Berrak Sisman, Kun Zhou 0003, Haizhou Li 0001 |
INTERSPEECH | 4 |
| 2022 | A Hybrid Continuity Loss to Reduce Over-Suppression for Time-domain Target Speaker ExtractionabstractThe speaker extraction algorithm extracts the target speech from a mixture speech containing interference speech and background noise. The extraction process sometimes over-suppresses the extracted target speech, which not only creates artifacts during listening but also harms the performance of downstream automatic speech recognition algorithms. We propose a hybrid continuity loss function for time-domain speaker extraction algorithms to settle the over-suppression problem. On top of the waveform-level loss used for superior signal quality, i.e., SI-SDR, we introduce a multi-resolution delta spectrum loss in the frequency-domain, to ensure the continuity of an extracted speech signal, thus alleviating the over-suppression. We examine the hybrid continuity loss function using a time-domain audio-visual speaker extraction algorithm on the YouTube LRS2-BBC dataset. Experimental results show that the proposed loss function reduces the over-suppression and improves the word error rate of speech recognition on both clean and noisy two-speakers mixtures, without harming the reconstructed speech quality. Zexu Pan, Meng Ge, Haizhou Li 0001 |
INTERSPEECH | 3 |
| 2022 | Knowledge distillation for In-memory keyword spotting model
Zeyang Song, Qi Liu 0005, Qu Yang, Haizhou Li 0001 |
INTERSPEECH | 4 |
| 2022 | LightHuBERT: Lightweight and Configurable Speech Representation Learning with Once-for-All Hidden-Unit BERTabstractSelf-supervised speech representation learning has shown promising results in various speech processing tasks.However, the pre-trained models, e.g., HuBERT, are storage-intensive Transformers, limiting their scope of applications under lowresource settings.To this end, we propose LightHuBERT, a once-for-all Transformer compression framework, to find the desired architectures automatically by pruning structured parameters.More precisely, we create a Transformer-based supernet that is nested with thousands of weight-sharing subnets and design a two-stage distillation strategy to leverage the contextualized latent representations from HuBERT.Experiments on automatic speech recognition (ASR) and the SU-PERB benchmark show the proposed LightHuBERT enables over 10 9 architectures concerning the embedding dimension, attention dimension, head number, feed-forward network ratio, and network depth.LightHuBERT outperforms the original HuBERT on ASR and five SUPERB tasks with the Hu-BERT size, achieves comparable performance to the teacher model in most tasks with a reduction of 29% parameters, and obtains a 3.5× compression ratio in three SUPERB tasks, e.g., automatic speaker verification, keyword spotting, and intent classification, with a slight accuracy loss.The code and pre-trained models are available at https://github.com/mechanicalsea/lighthubert. Rui Wang 0073, Qibing Bai, Junyi Ao, Zhixiang Xiong, Zhihua Wei 0001, Yu Zhang 0006, Tom Ko, Haizhou Li 0001 |
INTERSPEECH | 9 |
| 2022 | Deep residual spiking neural network for keyword spotting in low-resource settings
Qu Yang, Qi Liu 0005, Haizhou Li 0001 |
INTERSPEECH | 3 |
| 2022 | DDAM '22: 1st International Workshop on Deepfake Detection for Audio MultimediaabstractOver the last few years, the technology of speech synthesis and voice conversion has made significant improvement with the development of deep learning. The models can generate realistic and human-like speech. It is difficult for most people to distinguish the generated audio from the real. However, this technology also poses a great threat to the global political economy and social stability if some attackers and criminals misuse it with the intent to cause harm. In this workshop, we aim to bring together researchers from the fields of audio deepfake detection, audio deep synthesis, audio fake game and adversarial attacks to further discuss recent research and future directions for detecting deepfake and manipulated audios in multimedia. Jianhua Tao 0001, Jiangyan Yi, Cunhang Fan, Ruibo Fu, Shan Liang 0007, Pengyuan Zhang, Haizhou Li 0001, Helen M. Meng, Dong Yu 0001, Masato Akagi |
ACM Multimedia | 7 |
| 2022 | Training Spiking Neural Networks with Local Tandem LearningabstractSpiking neural networks (SNNs) are shown to be more biologically plausible and energy efficient over their predecessors. However, there is a lack of an efficient and generalized training method for deep SNNs, especially for deployment on analog computing substrates. In this paper, we put forward a generalized learning rule, termed Local Tandem Learning (LTL). The LTL rule follows the teacher-student learning approach by mimicking the intermediate feature representations of a pre-trained ANN. By decoupling the learning of network layers and leveraging highly informative supervisor signals, we demonstrate rapid network convergence within five training epochs on the CIFAR-10 dataset while having low computational complexity. Our experimental results have also shown that the SNNs thus trained can achieve comparable accuracies to their teacher ANNs on CIFAR-10, CIFAR-100, and Tiny ImageNet datasets. Moreover, the proposed LTL rule is hardware friendly. It can be easily implemented on-chip to perform fast parameter calibration and provide robustness against the notorious device non-ideality issues. It, therefore, opens up a myriad of opportunities for training and deployment of SNN on ultra-low-power mixed-signal neuromorphic computing chips. Qu Yang, Jibin Wu, Malu Zhang, Yansong Chua, Xinchao Wang, Haizhou Li 0001 |
NeurIPS | 6 |
| 2022 | Noise-robust voice conversion with domain adversarial training
Hongqiang Du, Lei Xie 0001, Haizhou Li 0001 |
Neural Networks | 3 |
| 2022 | Progressive Tandem Learning for Pattern Recognition With Deep Spiking Neural NetworksabstractSpiking neural networks (SNNs) have shown clear advantages over traditional artificial neural networks (ANNs) for low latency and high computational efficiency, due to their event-driven nature and sparse communication. However, the training of deep SNNs is not straightforward. In this paper, we propose a novel ANN-to-SNN conversion and layer-wise learning framework for rapid and efficient pattern recognition, which is referred to as progressive tandem learning. By studying the equivalence between ANNs and SNNs in the discrete representation space, a primitive network conversion method is introduced that takes full advantage of spike count to approximate the activation value of ANN neurons. To compensate for the approximation errors arising from the primitive network conversion, we further introduce a layer-wise learning method with an adaptive training scheduler to fine-tune the network weights. The progressive tandem learning framework also allows hardware constraints, such as limited weight precision and fan-in connections, to be progressively imposed during training. The SNNs thus trained have demonstrated remarkable classification and regression capabilities on large-scale object recognition, image reconstruction, and speech separation tasks, while requiring at least an order of magnitude reduced inference time and synaptic operations than other state-of-the-art SNN implementations. It, therefore, opens up a myriad of opportunities for pervasive mobile and embedded devices with a limited power budget. Jibin Wu, Chenglin Xu, Daquan Zhou, Malu Zhang, Haizhou Li 0001, Kay Chen Tan |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2022 | Emotional voice conversion: Theory, databases and ESDabstractIn this paper, we first provide a review of the state-of-the-art emotional voice conversion research, and the existing emotional speech databases. We then motivate the development of a novel emotional speech database (ESD) that addresses the increasing research need. With this paper, the ESD database1 is now made available to the research community. The ESD database consists of 350 parallel utterances spoken by 10 native English and 10 native Chinese speakers and covers 5 emotion categories (neutral, happy, angry, sad and surprise). More than 29 h of speech data were recorded in a controlled acoustic environment. The database is suitable for multi-speaker and cross-lingual emotional voice conversion studies. As case studies, we implement several state-of-the-art emotional voice conversion systems on the ESD database. This paper provides a reference study on ESD in conjunction with its release. Kun Zhou 0003, Berrak Sisman, Rui Liu 0008, Haizhou Li 0001 |
Speech Commun. | 4 |
| 2022 | Discriminative speaker embedding with serialized multi-layer multi-head attention
Hongning Zhu, Kong-Aik Lee, Haizhou Li 0001 |
Speech Commun. | 3 |
| 2022 | Neural Acoustic-Phonetic Approach for Speaker Verification With Phonetic Attention MaskabstractTraditional acoustic-phonetic approach makes use of both spectral and phonetic information when comparing the voice of speakers. While phonetic units are not equally informative, the phonetic context of speech plays an important role in speaker verification (SV). In this paper, we propose a neural acoustic-phonetic approach that learns to dynamically assign differentiated weights to spectral features for SV. Such differentiated weights form a phonetic attention mask (PAM). The neural acoustic-phonetic framework consists of two training pipelines, one for SV and another for speech recognition. Through the PAM, we leverage the phonetic information for SV. We evaluate the proposed neural acoustic-phonetic framework on the RSR2015 database Part III corpus, that consists of random digit strings. We show that the proposed framework with PAM consistently outperforms baseline with an equal error rate reduction of 13.45% and 10.20% for female and male data, respectively. Tianchi Liu 0004, Rohan Kumar Das, Kong-Aik Lee, Haizhou Li 0001 |
IEEE Signal Process. Lett. | 4 |
| 2022 | Speaker Extraction With Co-Speech Gestures CueabstractSpeaker extraction seeks to extract the clean speech of a target speaker from a multi-talker mixture speech. There have been studies to use a pre-recorded speech sample or face image of the target speaker as the speaker cue. In human communication, co-speech gestures that are naturally timed with speech also contribute to speech perception. In this work, we explore the use of co-speech gestures sequence, e.g. hand and body movements, as the speaker cue for speaker extraction, which could be easily obtained from low-resolution video recordings, thus more available than face recordings. We propose two networks using the co-speech gestures cue to perform attentive listening on the target speaker, one that implicitly fuses the co-speech gestures cue in the speaker extraction process, the other performs speech separation first, followed by explicitly using the co-speech gestures cue to associate a separated speech to the target speaker. The experimental results show that the co-speech gestures cue is informative in associating with the target speaker. Zexu Pan, Xinyuan Qian 0001, Haizhou Li 0001 |
IEEE Signal Process. Lett. | 3 |
| 2022 | Automatic Lyrics Transcription of Polyphonic Music With Lyrics-Chord Multi-Task LearningabstractLyrics are the words that make up a song, while chords are harmonic sets of multiple notes in music. Lyrics and chords are generally essential information in music, i.e. unaccompanied singing vocals mixed with instrumental music, representing important components in polyphonic music. In a traditional lyrics transcription task, we first extract the singing vocals from the polyphonic music and then transcribe the resulting singing vocals, where the two steps are optimized independently. In this paper, we propose novel end-to-end network architectures that are designed to disentangle lyrics from chords in polyphonic music for effective lyrics transcription in a single step, where we consider chords as musical words, analogously to lexical words as lyrics intuitively. We start by studying a single-task lyrics transcriber as the reference baseline and the initial model to develop the multi-task lyrics transcription solutions. The main idea is to take advantage of chord transcription available in the training data through multi-task training to improve lyrics transcription. The experiments show that the proposed multitask lyrics transcriber significantly outperforms other competing solutions, with a word error rate (WER) of 31.82% on a standard test dataset. Xiaoxue Gao, Chitralekha Gupta, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2022 | Deep Learning Approaches in Topics of Singing Information ProcessingabstractSinging, the vocal production of musical tones, is one of the most important elements of music. Addressing the needs of real-world applications, the study of technologies related to singing voices has become an increasingly active area of research. In this paper, we provide a comprehensive overview of the recent developments in the field of singing information processing, specifically in the topics of singing skill evaluation, singing voice synthesis, singing voice separation, and lyrics synchronization and transcription. We will especially focus on deep learning approaches including modern representation learning techniques for singing voices. We will also provide an overview of contributions in public datasets for singing voice research. Chitralekha Gupta, Haizhou Li 0001, Masataka Goto |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2022 | Decoding Knowledge Transfer for Neural Text-to-Speech TrainingabstractNeural end-to-end text-to-speech (TTS) is superior to conventional statistical methods in many ways. However, the exposure bias problem, that arises from the mismatch between the training and inference process in autoregressive models, remains an issue. It often leads to performance degradation in face of out-of-domain test data. To address this problem, we study a novel decoding knowledge transfer strategy, and propose a multi-teacher knowledge distillation (MT-KD) network for Tacotron2 TTS model. The idea is to pre-train two Tacotron2 TTS teacher models in teacher forcing and scheduled sampling modes, and transfer the pre-trained knowledge to a student model that performs free running decoding. We show that the MT-KD network provides an adequate platform for neural TTS training, where the student model learns to emulate the behaviors of the two teachers, at the same time, minimizing the mismatch between training and run-time inference. Experiments on both Chinese and English data show that MT-KD system consistently outperforms the competitive baselines in terms of naturalness, robustness and expressiveness for in-domain and out-of-domain test data. Furthermore, we show that knowledge distillation outperforms adversarial learning and data augmentation in addressing the exposure bias problem. Rui Liu 0008, Berrak Sisman, Guanglai Gao, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2022 | USEV: Universal Speaker Extraction With Visual CueabstractA speaker extraction algorithm seeks to extract the target speaker's speech from a multi-talker speech mixture. The prior studies focus mostly on speaker extraction from a highly overlapped multi-talker speech mixture. However, the target-interference speaker overlapping ratios could vary over a wide range from 0% to 100% in natural speech communication, furthermore, the target speaker could be absent in the speech mixture, the speech mixtures in such universal multi-talker scenarios are described asgeneral speech mixtures. The speaker extraction algorithm requires an auxiliary reference, such as a video recording or a pre-recorded speech, to form top-down auditory attention on the target speaker. We advocate that a visual cue, i.e., lip movement, is more informative than an audio cue, i.e., pre-recorded speech, to serve as the auxiliary reference for speaker extraction in disentangling the target speaker from ageneral speech mixture. In this paper, we propose a universal speaker extraction network with a visual cue, that works for all multi-talker scenarios. In addition, we propose a scenario-aware differentiated loss function for network training, to balance the network performance over different target-interference speaker pairing scenarios. The experimental results show that our proposed method outperforms various competitive baselines forgeneral speech mixturesin terms of signal fidelity. Zexu Pan, Meng Ge, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2022 | Selective Listening by Synchronizing Speech With LipsabstractA speaker extraction algorithm seeks to extract the speech of a target speaker from a multi-talker speech mixture when given a cue that represents the target speaker, such as a pre-enrolled speech utterance, or an accompanying video track. Visual cues are particularly useful when a pre-enrolled speech is not available. In this work, we don’t rely on the target speaker’s pre-enrolled speech, but rather use the target speaker’s face track as the speaker cue, that is referred to as the auxiliary reference, to form an attractor towards the target speaker. We advocate that the temporal synchronization between the speech and its accompanying lip movements is a direct and dominant audio-visual cue. Therefore, we propose a self-supervised pre-training strategy, to exploit the speech-lip synchronization cue for target speaker extraction, which allows us to leverage abundant unlabeled in-domain data. We transfer the knowledge from the pre-trained model to the attractor encoder of the speaker extraction network. We show that the proposed speaker extraction network outperforms various competitive baselines in terms of signal quality, perceptual quality, and intelligibility, achieving state-of-the-art performance. Zexu Pan, Ruijie Tao, Chenglin Xu, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2022 | EEG-Based Auditory Attention Detection via Frequency and Channel Neural AttentionabstractHumans have the ability to pay attention to one of the sound sources in a multispeaker acoustic environment. Auditory attention detection (AAD) seeks to detect the attended speaker from one’s brain signals that will enable many innovative human–machine systems. However, effective representation learning of electroencephalography (EEG) signals remains a challenge. In this article, we propose a neural attention mechanism that dynamically assigns differentiated weights to the subbands and the channels of EEG signals to derive discriminative representations for AAD. In the nutshell, we would like to build a computational attention mechanism, i.e., neural attention, to model the auditory attention in human brain. We incorporate the proposed neural attention into an AAD system, and validate the neural attention mechanism through comprehensive experiments on two publicly available datasets. The experimental results demonstrate that the proposed system significantly outperforms the state-of-the-art reference baselines. Siqi Cai 0002, Enze Su, Longhan Xie, Haizhou Li 0001 |
IEEE Trans. Hum. Mach. Syst. | 4 |
| 2022 | Rectified Linear Postsynaptic Potential Function for Backpropagation in Deep Spiking Neural NetworksabstractSpiking neural networks (SNNs) use spatiotemporal spike patterns to represent and transmit information, which are not only biologically realistic but also suitable for ultralow-power event-driven neuromorphic implementation. Just like other deep learning techniques, deep SNNs (DeepSNNs) benefit from the deep architecture. However, the training of DeepSNNs is not straightforward because the well-studied error backpropagation (BP) algorithm is not directly applicable. In this article, we first establish an understanding as to why error BP does not work well in DeepSNNs. We then propose a simple yet efficient rectified linear postsynaptic potential function (ReL-PSP) for spiking neurons and a spike-timing-dependent BP (STDBP) learning algorithm for DeepSNNs where the timing of individual spikes is used to convey information (temporal coding), and learning (BP) is performed based on spike timing in an event-driven manner. We show that DeepSNNs trained with the proposed single spike time-based learning algorithm can achieve the state-of-the-art classification accuracy. Furthermore, by utilizing the trained model parameters obtained from the proposed STDBP learning algorithm, we demonstrate ultralow-power inference operations on a recently proposed neuromorphic inference accelerator. The experimental results also show that the neuromorphic hardware consumes 0.751 mW of the total power consumption and achieves a low latency of 47.71 ms to classify an image from the Modified National Institute of Standards and Technology (MNIST) dataset. Overall, this work investigates the contribution of spike timing dynamics for information encoding, synaptic plasticity, and decision-making, providing a new perspective to the design of future DeepSNNs and neuromorphic hardware. Malu Zhang, Jibin Wu, Ammar Belatreche, Burin Amornpaisannon, Venkata Pavan Kumar Miriyala, Hong Qu 0002, Yansong Chua, Trevor E. Carlson, Haizhou Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 11 |
| 2021 | Bootstrapped Unsupervised Sentence Representation LearningabstractYan Zhang, Ruidan He, Zuozhu Liu, Lidong Bing, Haizhou Li. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Yan Zhang 0004, Ruidan He, Zuozhu Liu, Lidong Bing, Haizhou Li 0001 |
ACL/IJCNLP (1) | 5 |
| 2021 | DynaEval: Unifying Turn and Dialogue Level EvaluationabstractChen Zhang, Yiming Chen, Luis Fernando D’Haro, Yan Zhang, Thomas Friedrichs, Grandee Lee, Haizhou Li. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Chen Zhang 0055, Yiming Chen 0010, Luis Fernando D'Haro, Yan Zhang 0004, Thomas Friedrichs, Grandee Lee, Haizhou Li 0001 |
ACL/IJCNLP (1) | 7 |
| 2021 | Target Language Extraction at Multilingual Cocktail PartiesabstractTypically, target speaker extraction seeks to extract a target speaker's contribution according to his or her individual voice characteristics. In a “multilingual cocktail party” however, listeners may desire to extract speaker contributions spoken in a particular language, regard-less of the number of contributing speakers. In this paper, we pro-pose a novel task called “target language extraction” (TLE) which extracts voices based on the spoken language rather than on individ-ual speaker characteristics. We introduce a new database for TLE which simulates the multilingual cocktail party problem in mixtures of two and four speakers with German as the target language. The database is derived from the GlobalPhone 2000 Speaker Package and is called “GlobalPhone Multilingual Cocktail Party - German” (GlobalPhoneMCP-GE). Our experimental results show that our approach to TLE achieves very good performance regardless of the number of speakers in the mixture and that TLE generalizes well to unseen speakers and interfering languages. This work represents the first attempt at target language extraction. Marvin Borsdorf, Haizhou Li 0001, Tanja Schultz |
ASRU | 2 |
| 2021 | Expressive Voice Conversion: A Joint Framework for Speaker Identity and Emotional Style TransferabstractTraditional voice conversion (VC) has been focused on speaker identity conversion for speech with a neutral expression. We note that emotional expression plays an essential role in daily communication, and the emotional style of speech can be speaker-dependent. In this paper, we study a technique to jointly convert the speaker identity and speaker-dependent emotional style, that is called expressive voice conversion. We propose a StarGAN-based framework to learn a many-to-many mapping across different speakers, that takes into account speaker-dependent emotional style without the need for parallel data. To this end, we condition the generator on emotional style encoding derived from a pre-trained speech emotion recognition (SER) model. The experiments validate the effectiveness of our proposed framework in both objective and subjective evaluations. To our best knowledge, this is the first study on expressive voice conversion. Zongyang Du, Berrak Sisman, Kun Zhou 0003, Haizhou Li 0001 |
ASRU | 4 |
| 2021 | PL-EESR: Perceptual Loss Based End-to-End Robust Speaker Representation ExtractionabstractSpeech enhancement aims to improve the perceptual quality of the speech signal by suppression of the background noise. However, excessive suppression may lead to speech distortion and speaker information loss, which degrades the performance of speaker embedding extraction. To alleviate this problem, we propose an end-to-end deep learning framework, dubbed PL-EESR, for robust speaker representation extraction. This framework is optimized based on the feedback of the speaker identification task and the high-level perceptual deviation between the raw speech signal and its noisy version. We conducted speaker verification tasks in both noisy and clean environment respectively to evaluate our system. Compared to the baseline, our method shows better performance in both clean and noisy environments, which means our method can not only enhance the speaker relative information but also avoid adding distortions. Kong-Aik Lee, Ville Hautamäki, Haizhou Li 0001 |
ASRU | 4 |
| 2021 | DEEPA: A Deep Neural Analyzer for Speech and Singing VocodingabstractConventional vocoders are commonly used as analysis tools to provide interpretable features for downstream tasks such as speech synthesis and voice conversion. They are built under certain assumptions about the signals following signal processing principle, therefore, not easily generalizable to different audio, for example, from speech to singing. In this paper, we propose a deep neural analyzer, denoted as DeepA – a neural vocoder that extracts F0 and timbre/aperiodicity encoding from the input speech that emulate those defined in conventional vocoders. Therefore, the resulting parameters are more interpretable than other latent neural representations. At the same time, as the deep neural analyzer is learnable, it is expected to be more accurate for signal reconstruction and manipulation, and generalizable from speech to singing. The proposed neural analyzer is built based on a variational autoencoder (VAE) architecture. We show that DeepA improves F0 estimation over the conventional vocoder (WORLD). To our best knowledge, this is the first study dedicated to the development of a neural framework for extracting learnable vocoder-like parameters. Sergey Nikonorov, Berrak Sisman, Mingyang Zhang 0003, Haizhou Li 0001 |
ASRU | 4 |
| 2021 | Exploring Teacher-Student Learning Approach for Multi-Lingual Speech-to-Intent ClassificationabstractEnd-to-end speech-to-intent classification has shown its advantage in harvesting information from both text and speech. In this paper, we study a technique to develop such an end-to-end system that supports multiple languages. To overcome the scarcity of multi-lingual speech corpus, we exploit knowledge from a pre-trained multi-lingual natural language processing model. Multi-lingual bidirectional encoder representations from transformers (mBERT) models are trained on multiple languages and hence expected to perform well in the multi-lingual scenario. In this work, we employ a teacher-student learning approach to sufficiently extract information from an mBERT model to train a multi-lingual speech model. In particular, we use synthesized speech generated from an English-Mandarin text corpus for analysis and training of a multi-lingual intent classification model. We also demonstrate that the teacher-student learning approach obtains an improved performance (91.02%) over the traditional end-to-end (89.40%) intent classification approach in a practical multi-lingual scenario. Bidisha Sharma, Maulik C. Madhavi, Xuehao Zhou, Haizhou Li 0001 |
ASRU | 4 |
| 2021 | Revisiting Self-training for Few-shot Learning of Language ModelabstractAs unlabeled data carry rich task-relevant information, they are proven useful for fewshot learning of language model.The question is how to effectively make use of such data.In this work, we revisit the self-training technique for language model fine-tuning and present a state-of-the-art prompt-based fewshot learner, SFLM.Given two views of a text sample via weak and strong augmentation techniques, SFLM generates a pseudo label on the weakly augmented version.Then, the model predicts the same pseudo label when fine-tuned with the strongly augmented version.This simple approach is shown to outperform other state-of-the-art supervised and semi-supervised counterparts on six sentence classification and six sentence-pair classification benchmarking tasks.In addition, SFLM only relies on a few in-domain unlabeled data.We conduct a comprehensive analysis to demonstrate the robustness of our proposed approach under various settings, including augmentation techniques, model scale, and fewshot knowledge transfer across tasks. Yiming Chen 0010, Yan Zhang 0004, Chen Zhang 0020, Grandee Lee, Haizhou Li 0001 |
EMNLP (1) | 6 |
| 2021 | Graphspeech: Syntax-Aware Graph Attention Network for Neural Speech SynthesisabstractAttention-based end-to-end text-to-speech synthesis (TTS) is superior to conventional statistical methods in many ways. Transformer-based TTS is one of such successful implementations. While Transformer TTS models the speech frame sequence well with a self-attention mechanism, it does not associate input text with output utterances from a syntactic point of view at sentence level. We propose a novel neural TTS model, denoted as GraphSpeech, that is formulated under graph neural network framework. GraphSpeech encodes explicitly the syntactic relation of input lexical tokens in a sentence, and incorporates such information to derive syntactically motivated character embeddings for TTS attention mechanism. Experiments show that GraphSpeech consistently outperforms the Transformer TTS baseline in terms of spectrum and prosody rendering of utterances. Rui Liu 0008, Berrak Sisman, Haizhou Li 0001 |
ICASSP | 3 |
| 2021 | Data Augmentation with Signal Companding for Detection of Logical Access AttacksabstractThe recent advances in voice conversion (VC) and text-to-speech (TTS) make it possible to produce natural sounding speech that poses threat to automatic speaker verification (ASV) systems. To this end, research on spoofing countermeasures has gained attention to protect ASV systems from such attacks. While the advanced spoofing countermeasures are able to detect known nature of spoofing attacks, they are not that effective under unknown attacks. In this work, we propose a novel data augmentation technique using a-law and mu-law based signal companding. We believe that the proposed method has an edge over traditional data augmentation by adding small perturbation or quantization noise. The studies are conducted on ASVspoof 2019 logical access corpus using light convolutional neural network based system. We find that the proposed data augmentation technique based on signal companding outperforms the state-of-the-art spoofing countermeasures showing ability to handle unknown nature of attacks. Rohan Kumar Das, Haizhou Li 0001 |
ICASSP | 3 |
| 2021 | Multi-Stage Speaker Extraction with Utterance and Frame-Level Reference SignalsabstractSpeaker extraction requires a sample speech from the target speaker as the reference. However, enrolling a speaker with a long speech is not practical. We propose a speaker extraction technique, that performs in multiple stages to take full advantage of short reference speech sample. The extracted speech in early stages is used as the reference speech for late stages. For the first time, we use frame-level sequential speech embedding as the reference for target speaker. This is a departure from the traditional utterance-based speaker embedding reference. In addition, a signal fusion scheme is proposed to combine the decoded signals in multiple scales with automatically learned weights. Experiments on WSJ0-2mix and its noisy versions (WHAM! and WHAMR!) show that SpEx++ consistently outperforms other state-of-the-art baselines. Meng Ge, Chenglin Xu, Longbiao Wang, Chng Eng Siong, Jianwu Dang 0001, Haizhou Li 0001 |
ICASSP | 6 |
| 2021 | Representation Learning with Spectro-Temporal-Channel Attention for Speech Emotion RecognitionabstractConvolutional neural network (CNN) is found to be effective in learning representation for speech emotion recognition. CNNs do not explicitly model the associations or relative importance of features in the spectral/temporal/channel-wise axes. In this paper, we propose an attention module, named spectro-temporal-channel (STC) attention module that is integrated with CNN to improve representation learning ability. Our module infers an attention map along the three dimensions, namely time, frequency, and CNN channel. Experiments are conducted on the IEMOCAP database to evaluate the effectiveness of the proposed representation learning method. The results demonstrate that the proposed method outperforms the traditional CNN method by an absolute increase of 3.13% in terms of F1 score. Lili Guo 0001, Longbiao Wang, Chenglin Xu, Jianwu Dang 0001, Chng Eng Siong, Haizhou Li 0001 |
ICASSP | 6 |
| 2021 | Learning Disentangled Feature Representations for Speech Enhancement Via Adversarial TrainingabstractNeural speech enhancement degrades significantly in face of unseen noise. To address such mismatch, we propose to learn noise-agnostic feature representations by disentanglement learning, which removes the unspecified noise factor, while keeping the specified factors of variation associated with the clean speech. Specifically, a discriminator module is introduced to distinguish the type of noises, which is referred to as the disentangler. With the adversarial training strategy, a gradient reversal layer seeks to disentangle the noise factor and remove it from the feature representation. Experiment results show that the proposed approach achieves 5.8% and 5.2% relative improvements over the best baseline in terms of perceptual evaluation of the speech quality (PESQ) and segmental signal-to-noise ratio (SSNR), respectively. The ablation study indicates that the proposed disentangler module is also effective in other encoder-decoder-like structures. Nana Hou, Chenglin Xu, Chng Eng Siong, Haizhou Li 0001 |
ICASSP | 4 |
| 2021 | Muse: Multi-Modal Target Speaker Extraction with Visual CuesabstractSpeaker extraction algorithm relies on the speech sample from the target speaker as the reference point to focus its attention. Such a reference speech is typically pre-recorded. On the other hand, the temporal synchronization between speech and lip movement also serves as an informative cue. Motivated by this idea, we study a novel technique to use speech-lip visual cues to extract reference target speech directly from mixture speech during inference time, without the need of pre-recorded reference speech. We propose a multi-modal speaker extraction network, named MuSE, that is conditioned only on a lip image sequence. MuSE not only outperforms other competitive baselines in terms of SI-SDR and PESQ, but also shows consistent improvement in cross-dataset evaluations. Zexu Pan, Ruijie Tao, Chenglin Xu, Haizhou Li 0001 |
ICASSP | 4 |
| 2021 | Multi-Target DoA Estimation with an Audio-Visual Fusion MechanismabstractMost of the prior studies in the spatial Direction of Arrival (DoA) domain focus on a single modality. However, humans use auditory and visual senses to detect the presence of sound sources. With this motivation, we propose to use neural networks with audio and visual signals for multi-speaker localization. The use of heterogeneous sensors can provide complementary information to overcome uni-modal challenges, such as noise, reverberation, illumination variations, and occlusions. We attempt to address these issues by introducing an adaptive weighting mechanism for audio-visual fusion. We also propose a novel video simulation method that generates visual features from noisy target 3D annotations that are synchronized with acoustic features. Experimental results confirm that audio-visual fusion consistently improves the performance of speaker DoA estimation, while the adaptive weighting mechanism shows clear benefits. Xinyuan Qian 0001, Maulik C. Madhavi, Zexu Pan, Haizhou Li 0001 |
ICASSP | 5 |
| 2021 | Leveraging Acoustic and Linguistic Embeddings from Pretrained Speech and Language Models for Intent ClassificationabstractIntent classification is a task in spoken language understanding. An intent classification system is usually implemented as a pipeline process, with a speech recognition module followed by text processing that classifies the intents. There are also studies of end-to-end system that take acoustic features as input and classifies the intents directly. Such systems don’t take advantage of relevant linguistic information, and suffer from limited training data. In this work, we propose a novel intent classification framework that employs acoustic features extracted from a pretrained speech recognition system and linguistic features learned from a pretrained language model. We use knowledge distillation technique to map the acoustic embeddings towards linguistic embeddings. We perform fusion of both acoustic and linguistic embeddings through cross-attention approach to classify intents. With the pro-posed method, we achieve 90.86% and 99.07% accuracy on ATIS and Fluent speech corpus, respectively. Bidisha Sharma, Maulik C. Madhavi, Haizhou Li 0001 |
ICASSP | 3 |
| 2021 | The Multi-Speaker Multi-Style Voice Cloning Challenge 2021abstractThe Multi-speaker Multi-style Voice Cloning Challenge (M2VoC) aims to provide a common sizable dataset as well as a fair testbed for the benchmarking of the popular voice cloning task. Specifically, we formulate the challenge to adapt an average TTS model to the stylistic target voice with limited data from target speaker, evaluated by speaker identity and style similarity. The challenge consists of two tracks, namely few-shot track and one-shot track, where the participants are required to clone multiple target voices with 100 and 5 samples respectively. There are also two sub-tracks in each track. For sub-track a, to fairly compare different strategies, the participants are allowed to use only the training data provided by the organizer strictly. For sub-track b, the participants are allowed to use any data publicly available. In this paper, we present a detailed explanation on the tasks and data used in the challenge, followed by a summary of submitted systems and evaluation results. Qicong Xie, Xiaohai Tian, Guanghou Liu, Lei Xie 0001, Zhiyong Wu 0001, Haizhou Li 0001, Fen Hong, Hui Bu |
ICASSP | 9 |
| 2021 | Seen and Unseen Emotional Style Transfer for Voice Conversion with A New Emotional Speech DatasetabstractEmotional voice conversion aims to transform emotional prosody in speech while preserving the linguistic content and speaker identity. Prior studies show that it is possible to disentangle emotional prosody using an encoder-decoder network conditioned on discrete representation, such as one-hot emotion labels. Such networks learn to remember a fixed set of emotional styles. In this paper, we propose a novel framework based on variational auto-encoding Wasserstein generative adversarial network (VAW-GAN), which makes use of a pre-trained speech emotion recognition (SER) model to transfer emotional style during training and at run-time inference. In this way, the network is able to transfer both seen and unseen emotional style to a new utterance. We show that the proposed framework achieves remarkable performance by consistently outperforming the baseline framework. This paper also marks the release of an emotional speech dataset (ESD) for voice conversion, which has multiple speakers and languages. Kun Zhou 0003, Berrak Sisman, Rui Liu 0008, Haizhou Li 0001 |
ICASSP | 4 |
| 2021 | Accumulated Decoupled Learning with Gradient Staleness Mitigation for Convolutional Neural NetworksabstractGradient staleness is a major side effect in decoupled learning when training convolutional neural networks asynchronously. Existing methods that ignore this effect might result in reduced generalization and even divergence. In this paper, we propose an accumulated decoupled learning (ADL), which includes a module-wise gradient accumulation in order to mitigate the gradient staleness. Unlike prior arts ignoring the gradient staleness, we quantify the staleness in such a way that its mitigation can be quantitatively visualized. As a new learning scheme, the proposed ADL is theoretically shown to converge to critical points in spite of its asynchronism. Extensive experiments on CIFAR-10 and ImageNet datasets are conducted, demonstrating that ADL gives promising generalization results while the state-of-the-art methods experience reduced generalization and divergence. In addition, our ADL is shown to have the fastest training speed among the compared methods. Huiping Zhuang, Zhenyu Weng, Fulin Luo, Kar-Ann Toj, Haizhou Li 0001, Zhiping Lin 0001 |
ICML | 5 |
| 2021 | GCC-PHAT with Speech-oriented Attention for Robotic Sound Source LocalizationabstractRobotic audition is a basic sense that helps robots perceive the surroundings and interact with humans. Sound Source Localization (SSL) is an essential module for a robotic system. However, the performance of most sound source localization techniques degrades in noisy and reverberant environments due to inaccurate Time Difference of Arrival (TDoA) estimation. In robotic sound source localization, we are more interested in detecting the arrival of human speech than other sound sources. Ideally, we expect an effective TDoA estimation to respond only to speech signals, while masking off other interferences. In this paper, we propose a novel technique that learns to attend to speech fundamental frequency and harmonics while suppressing noise interference and reverberation. The novel TDoA feature is referred to as Generalized Cross Correlation with Phase Transform and Speech Mask (GCC-PHAT-SM). We perform sound source localization experiments on real-world data captured from a robotic platform. Experiments show that GCC-PHAT-SM feature significantly outperforms traditional Generalized Cross Correlation (GCC) feature in noisy and reverberant acoustic environments. Xinyuan Qian 0001, Zihan Pan, Malu Zhang, Haizhou Li 0001 |
ICRA | 5 |
| 2021 | Rethinking Benchmarks for Neuromorphic Learning AlgorithmsabstractWe rely on benchmarking datasets to monitor the research progress. However, recent studies have cast doubts on the effectiveness of current neuromorphic benchmarking datasets; and the debate remains largely unsettled. In this paper, we assess the richness and usefulness of temporal information embedded in these benchmarking datasets for SNN decision making. To this end, we propose a segregated spatio-temporal learning framework that allows us to selectively control the information flow along both spatial and temporal directions during feedforward and backward propagation. Leveraging on this framework, we conduct a comprehensive study on seven widely used neuromorphic audio and vision datasets. Our findings are threefold. First, the existing neuromorphic benchmarks only make limited contributions in highlighting the temporal processing capability of spiking neurons. Second, the temporal credit assignment is redundant for tasks that only require short-range temporal dependency. Third, we recommend the neuromorphic research community to develop novel benchmarks that require both short-range and long-range temporal dependencies. Such appropriate benchmark datasets would be helpful in guiding the development of powerful SNN-based learning algorithms and computational models. Qu Yang, Jibin Wu, Haizhou Li 0001 |
IJCNN | 3 |
| 2021 | Reinforcement Learning for Emotional Text-to-Speech Synthesis with Improved Emotion DiscriminabilityabstractEmotional text-to-speech synthesis (ETTS) has seen much progress in recent years.However, the generated voice is often not perceptually identifiable by its intended emotion category.To address this problem, we propose a new interactive training paradigm for ETTS, denoted as i-ETTS, which seeks to directly improve the emotion discriminability by interacting with a speech emotion recognition (SER) model.Moreover, we formulate an iterative training strategy with reinforcement learning to ensure the quality of i-ETTS optimization.Experimental results demonstrate that the proposed i-ETTS outperforms the state-of-the-art baselines by rendering speech with more accurate emotion style.To our best knowledge, this is the first study of reinforcement learning in emotional text-to-speech synthesis. Rui Liu 0008, Berrak Sisman, Haizhou Li 0001 |
Interspeech | 3 |
| 2021 | Universal Speaker Extraction in the Presence and Absence of Target Speakers for Speech of One and Two Talkers
Marvin Borsdorf, Chenglin Xu, Haizhou Li 0001, Tanja Schultz |
Interspeech | 3 |
| 2021 | GlobalPhone Mix-To-Separate Out of 2: A Multilingual 2000 Speakers Mixtures Database for Speech Separation
Marvin Borsdorf, Chenglin Xu, Haizhou Li 0001, Tanja Schultz |
Interspeech | 3 |
| 2021 | Diagnosis of COVID-19 Using Auditory Acoustic CuesabstractCOVID-19 can be pre-screened based on symptoms and confirmed using other laboratory tests.The cough or speech from patients are also studied in the recent time for detection of COVID-19 as they are indicators of change in anatomy and physiology of the respiratory system.Along this direction, the diagnosis of COVID-19 using acoustics (DiCOVA) challenge aims to promote such research by releasing publicly available cough/speech corpus.We participated in the Track-1 of the challenge, which deals with COVID-19 detection using cough sounds from individuals.In this challenge, we use a few novel auditory acoustic cues based on long-term transform, equivalent rectangular bandwidth spectrum and gammatone filterbank.We evaluate these representations using logistic regression, random forest and multilayer perceptron classifiers for detection of COVID-19.On the blind test set, we obtain an area under the ROC curve (AUC) of 83.49% for the best system submitted to the challenge.It is worth noting that the submitted system ranked among the top few systems on the leaderboard and outperformed the challenge baseline by a large margin. Rohan Kumar Das, Maulik C. Madhavi, Haizhou Li 0001 |
Interspeech | 3 |
| 2021 | Knowledge Distillation from BERT Transformer to Speech Transformer for Intent ClassificationabstractEnd-to-end intent classification using speech has numerous advantages compared to the conventional pipeline approach using automatic speech recognition (ASR), followed by natural language processing modules.It attempts to predict intent from speech without using an intermediate ASR module.However, such end-to-end framework suffers from the unavailability of large speech resources with higher acoustic variation in spoken language understanding.In this work, we exploit the scope of the transformer distillation method that is specifically designed for knowledge distillation from a transformer based language model to a transformer based speech model.In this regard, we leverage the reliable and widely used bidirectional encoder representations from transformers (BERT) model as a language model and transfer the knowledge to build an acoustic model for intent classification using the speech.In particular, a multilevel transformer based teacher-student model is designed, and knowledge distillation is performed across attention and hidden sub-layers of different transformer layers of the student and teacher models.We achieve an intent classification accuracy of 99.10% and 88.79% for Fluent speech corpus and ATIS database, respectively.Further, the proposed method demonstrates better performance and robustness in acoustically degraded condition compared to the baseline method. Yidi Jiang, Bidisha Sharma, Maulik C. Madhavi, Haizhou Li 0001 |
Interspeech | 4 |
| 2021 | Neural Speaker Extraction with Speaker-Speech Cross-Attention Network
Wupeng Wang, Chenglin Xu, Meng Ge, Haizhou Li 0001 |
Interspeech | 4 |
| 2021 | Phonetically Motivated Self-Supervised Speech Representation Learning
Xianghu Yue, Haizhou Li 0001 |
Interspeech | 2 |
| 2021 | Temporal Convolutional Network with Frequency Dimension Adaptive Attention for Speech EnhancementabstractDespite much progress, most temporal convolutional networks (TCN) based speech enhancement models are mainly focused on modeling the long-term temporal contextual dependencies of speech frames, without taking into account the distribution information of speech signal in frequency dimension. In this study, we propose a frequency dimension adaptive attention (FAA) mechanism to improve TCNs, which guides the model selectively emphasize the frequency-wise features with important speech information and also improves the representation capability of network. Our extensive experimental investigation demonstrates that the proposed FAA mechanism is able to consistently provide significant improvements in terms of speech quality (PESQ), intelligibility (STOI) and three other composite metrics. More promisingly, it has better generalization ability to real-world noisy environment. Qiquan Zhang, Aaron Nicolson, Haizhou Li 0001 |
Interspeech | 5 |
| 2021 | Multi-Level Transfer Learning from Near-Field to Far-Field Speaker VerificationabstractIn far-field speaker verification, the performance of speaker embeddings is susceptible to degradation when there is a mismatch between the conditions of enrollment and test speech.To solve this problem, we propose the feature-level and instancelevel transfer learning in the teacher-student framework to learn a domain-invariant embedding space.For the feature-level knowledge transfer, we develop the contrastive loss to transfer knowledge from teacher model to student model, which can not only decrease the intra-class distance, but also enlarge the inter-class distance.Moreover, we propose the instance-level pairwise distance transfer method to force the student model to preserve pairwise instances distance from the well optimized embedding space of the teacher model.On FFSVC 2020 evaluation set, our EER on Full-eval trials is relatively reduced by 13.9% compared with the fusion system result on Partialeval trials of Task2.On Task1, compared with the winner's DenseNet result on Partial-eval trials, our minDCF on Full-eval trials is relatively reduced by 6.3%.On Task3, the EER and minDCF of our proposed method on Full-eval trials are very close to the result of the fusion system on Partial-eval trials.Our results also outperform other competitive domain adaptation methods. Li Zhang 0084, Qing Wang 0039, Kong-Aik Lee, Lei Xie 0001, Haizhou Li 0001 |
Interspeech | 5 |
| 2021 | Limited Data Emotional Voice Conversion Leveraging Text-to-Speech: Two-Stage Sequence-to-Sequence TrainingabstractEmotional voice conversion (EVC) aims to change the emotional state of an utterance while preserving the linguistic content and speaker identity. In this paper, we propose a novel 2-stage training strategy for sequence-to-sequence emotional voice conversion with a limited amount of emotional speech data. We note that the proposed EVC framework leverages text-to-speech (TTS) as they share a common goal that is to generate high-quality expressive voice. In stage 1, we perform style initialization with a multi-speaker TTS corpus, to disentangle speaking style and linguistic content. In stage 2, we perform emotion training with a limited amount of emotional speech data, to learn how to disentangle emotional style and linguistic information from the speech. The proposed framework can perform both spectrum and prosody conversion and achieves significant improvement over the state-of-the-art baselines in both objective and subjective evaluation. Kun Zhou 0003, Berrak Sisman, Haizhou Li 0001 |
Interspeech | 3 |
| 2021 | Cross-Lingual Voice Conversion with a Cycle Consistency Loss on Linguistic Representation
Yi Zhou 0020, Xiaohai Tian, Zhizheng Wu 0001, Haizhou Li 0001 |
Interspeech | 4 |
| 2021 | Serialized Multi-Layer Multi-Head Attention for Neural Speaker EmbeddingabstractThis paper proposes a serialized multi-layer multi-head attention for neural speaker embedding in text-independent speaker verification.In prior works, frame-level features from one layer are aggregated to form an utterance-level representation.Inspired by the Transformer network, our proposed method utilizes the hierarchical architecture of stacked self-attention mechanisms to derive refined features that are more correlated with speakers.Serialized attention mechanism contains a stack of self-attention modules to create fixed-dimensional representations of speakers.Instead of utilizing multi-head attention in parallel, the proposed serialized multi-layer multi-head attention is designed to aggregate and propagate attentive statistics from one layer to the next in a serialized manner.In addition, we employ an input-aware query for each utterance with the statistics pooling.With more layers stacked, the neural network can learn more discriminative speaker embeddings.Experiment results on VoxCeleb1 dataset and SITW dataset show that our proposed method outperforms other baseline methods, including x-vectors and other x-vectors + conventional attentive pooling approaches by 9.7% in EER and 8.1% in DCF10 -2 . Hongning Zhu, Kong-Aik Lee, Haizhou Li 0001 |
Interspeech | 3 |
| 2021 | Is Someone Speaking?: Exploring Long-term Temporal Features for Audio-visual Active Speaker DetectionabstractActive speaker detection (ASD) seeks to detect who is speaking in a visual scene of one or more speakers. The successful ASD depends on accurate interpretation of short-term and long-term audio and visual information, as well as audio-visual interaction. Unlike the prior work where systems make decision instantaneously using short-term features, we propose a novel framework, named TalkNet, that makes decision by taking both short-term and long-term features into consideration. TalkNet consists of audio and visual temporal encoders for feature representation, audio-visual cross-attention mechanism for inter-modality interaction, and a self-attention mechanism to capture long-term speaking evidence. The experiments demonstrate that TalkNet achieves 3.5% and 2.2% improvement over the state-of-the-art systems on the AVA-ActiveSpeaker dataset and Columbia ASD dataset, respectively. Code has been made available at: https://github.com/TaoRuijie/TalkNet_ASD. Ruijie Tao, Zexu Pan, Rohan Kumar Das, Xinyuan Qian 0001, Zheng Shou 0001, Haizhou Li 0001 |
ACM Multimedia | 6 |
| 2021 | Proceedings of the 22nd Annual Meeting of the Special Interest Group on Discourse and Dialogue
Haizhou Li 0001, Gina-Anne Levow, Chitralekha Gupta, Berrak Sisman, Siqi Cai 0002, David Vandyke, Nina Dethlefs, Yan Wu 0002, Junyi Jessy Li |
SIGDIAL | 1 |
| 2021 | Optimizing Voice Conversion Network with Cycle Consistency Loss of Speaker IdentityabstractWe propose a novel training scheme to optimize voice conversion network with a speaker identity loss function. The training scheme not only minimizes frame-level spectral loss, but also speaker identity loss. We introduce a cycle consistency loss that constrains the converted speech to maintain the same speaker identity as reference speech at utterance level. While the proposed training scheme is applicable to any voice conversion networks, we formulate the study under the average model voice conversion framework in this paper. Experiments conducted on CMU-ARCTIC and CSTR-VCTK corpus confirm that the proposed method outperforms baseline methods in terms of speaker similarity. Hongqiang Du, Xiaohai Tian, Lei Xie 0001, Haizhou Li 0001 |
SLT | 4 |
| 2021 | Vaw-Gan For Disentanglement And Recomposition Of Emotional Elements In SpeechabstractEmotional voice conversion (EVC) aims to convert the emotion of speech from one state to another while preserving the linguistic content and speaker identity. In this paper, we study the disentanglement and recomposition of emotional elements in speech through variational autoencoding Wasserstein generative adversarial network (VAW-GAN). We propose a speaker-dependent EVC framework based on VAW-GAN, that includes two VAW-GAN pipelines, one for spectrum conversion, and another for prosody conversion. We train a spectral encoder that disentangles emotion and prosody (F0) information from spectral features; we also train a prosodic encoder that disentangles emotion modulation of prosody (affective prosody) from linguistic prosody. At run-time, the decoder of spectral VAW-GAN is conditioned on the output of prosodic VAW-GAN. The vocoder takes the converted spectral and prosodic features to generate the target emotional speech. Experiments validate the effectiveness of our proposed method in both objective and subjective evaluations. Kun Zhou 0003, Berrak Sisman, Haizhou Li 0001 |
SLT | 3 |
| 2021 | HuRAI: A brain-inspired computational model for human-robot auditory interface
Jibin Wu, Qi Liu 0005, Malu Zhang, Zihan Pan, Haizhou Li 0001, Kay Chen Tan |
Neurocomputing | 5 |
| 2021 | FastTalker: A neural text-to-speech architecture with shallow and group autoregression
Rui Liu 0008, Berrak Sisman, Yixing Lin, Haizhou Li 0001 |
Neural Networks | 4 |
| 2021 | Factorized WaveNet for voice conversion with limited data
Hongqiang Du, Xiaohai Tian, Lei Xie 0001, Haizhou Li 0001 |
Speech Commun. | 4 |
| 2021 | An adaptive transmission line cochlear model based front-end for replay attack detection
Tharshini Gunendradasan, Eliathamby Ambikairajah, Julien Epps, Vidhyasaharan Sethu, Haizhou Li 0001 |
Speech Commun. | 5 |
| 2021 | NHSS: A speech and singing parallel database
Bidisha Sharma, Xiaoxue Gao, Karthika Vijayan, Xiaohai Tian, Haizhou Li 0001 |
Speech Commun. | 5 |
| 2021 | Three-Dimensional Speaker Localization: Audio-Refined Visual Scaling Factor EstimationabstractNeither a monocular RGB camera nor a small-size microphone array is capable of accurate three-dimensional (3D) speaker localization. By taking advantage of accurate visual object detection, and audio-visual complementary sensor fusion, we formulate the three-dimensional (3D) speaker localization problem as a visual scaling factor estimation problem. As a result, we effectively reduce the traditional audio-only 3D speaker localization from an exhaustive grid search to a one-dimensional (1D) optimization problem. We propose a multi-modal perception system with two optimization approaches. We show that the proposed methods are effective, accurate, and robust against interference and, as corroborated by indicative empirical results on real dataset, competitive to the conventional uni-modal and the state-of-the-art audio-visual speaker localization approaches. Xinyuan Qian 0001, Qi Liu 0005, Haizhou Li 0001 |
IEEE Signal Process. Lett. | 4 |
| 2021 | Exploiting Morphological and Phonological Features to Improve Prosodic Phrasing for Mongolian Speech SynthesisabstractProsodic phrasing is an important factor that affects naturalness and intelligibility in text-to-speech synthesis. Studies show that deep learning techniques improve prosodic phrasing when large text and speech corpus are available. However, for low-resource languages, such as Mongolian, prosodic phrasing remains a challenge for various reasons. First, the database suitable for system training is limited. Second, word composition knowledge that is prosody-informing has not been used in prosodic phrase modeling. To address these problems, in this article, we propose a feature augmentation method in conjunction with a self-attention neural classifier. We augment input text with morphological and phonological decompositions of words to enhance the text encoder. We study the use of self-attention classifier, that makes use of global context of a sentence, as a decoder for phrase break prediction. Both objective and subjective evaluations validate the effectiveness of the proposed phrase break prediction framework, that consistently improves voice quality in a Mongolian text-to-speech synthesis system. Rui Liu 0008, Berrak Sisman, Feilong Bao, Guanglai Gao, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2021 | Expressive TTS Training With Frame and Style Reconstruction LossabstractWe propose a novel training strategy for Tacotron-based text-to-speech (TTS) system that improves the speech styling at utterance level. One of the key challenges in prosody modeling is the lack of reference that makes explicit modeling difficult. The proposed technique doesn’t require prosody annotations from training data. It doesn’t attempt to model prosody explicitly either, but rather encodes the association between input text and its prosody styles using a Tacotron-based TTS framework. This study marks a departure from the style token paradigm where prosody is explicitly modeled by a bank of prosody embeddings. It adopts a combination of two objective functions: 1) frame level reconstruction loss, that is calculated between the synthesized and target spectral features; 2) utterance level style reconstruction loss, that is calculated between the deep style features of synthesized and target speech. The style reconstruction loss is formulated as a perceptual loss to ensure that utterance level speech style is taken into consideration during training. Experiments show that the proposed training strategy achieves remarkable performance and outperforms the state-of-the-art baseline in both naturalness and expressiveness. To our best knowledge, this is the first study to incorporate utterance level perceptual quality as a loss function into Tacotron training for improved expressiveness. Rui Liu 0008, Berrak Sisman, Guanglai Gao, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2021 | Multi-Tone Phase Coding of Interaural Time Difference for Sound Source Localization With Spiking Neural NetworksabstractMammals exhibit remarkable capability of detecting and localizing sound sources in complex acoustic environments by using binaural cues in the spiking manner. Emulating the auditory process for sound source localization (SSL) by mammals, we propose a computational model for accurate and robust SSL under the neuromorphic spiking neural network (SNN) framework. The center of this model is a Multi-Tone Phase Coding (MTPC) scheme, which encodes the interaural time difference (ITD) between binaural pure tones into discriminative spike patterns that can be directly classified by SNNs. As such, SSL can be implemented as an event-driven task on highly efficient, neuromorphic parallel processors. We evaluate the proposed computational model on a directional audio dataset recorded from a microphone array in a realistic acoustic environment with background noise, obstruction, reflection, and other interferences. We report superior localization capability with a mean absolute error (MAE) of 1.02°or 100% classification accuracy with an angle resolution of 5°, which surpasses other SNN-based biologically plausible neuromorphic approaches by a relatively large margin and on par with human performance in similar tasks. This study opens up many application opportunities in human-robot interaction where energy efficiency is crucial. As a case study, we successfully deploy the proposed SSL system in a robotic platform to track the speaker and orient the robot's attention. Zihan Pan, Malu Zhang, Jibin Wu, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2021 | An Overview of Voice Conversion and Its Challenges: From Statistical Modeling to Deep LearningabstractSpeaker identity is one of the important characteristics of human speech. In voice conversion, we change the speaker identity from one to another, while keeping the linguistic content unchanged. Voice conversion involves multiple speech processing techniques, such as speech analysis, spectral conversion, prosody conversion, speaker characterization, and vocoding. With the recent advances in theory and practice, we are now able to produce human-like voice quality with high speaker similarity. In this article, we provide a comprehensive overview of the state-of-the-art of voice conversion techniques and their performance evaluation methods from the statistical approaches to deep learning, and discuss their promise and limitations. We will also report the recent Voice Conversion Challenges (VCC), the performance of the current state of technology, and provide a summary of the available resources for voice conversion research. Berrak Sisman, Junichi Yamagishi, Simon King 0001, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2021 | Target Speaker Verification With Selective Auditory Attention for Single and Multi-Talker SpeechabstractSpeaker verification has been studied mostly under the single-talker condition. It is adversely affected in the presence of interference speakers. Inspired by the study on target speaker extraction, e.g., SpEx, we propose a unified speaker verification framework for both single- and multi-talker speech, that is able to pay selective auditory attention to the target speaker. This target speaker verification (tSV) framework jointly optimizes a speaker attention module and a speaker representation module via multi-task learning. We study four different target speaker embedding schemes under the tSV framework. The experimental results show that all four target speaker embedding schemes significantly outperform other competitive solutions for multi-talker speech. Notably, the best tSV speaker embedding scheme achieves 76.0% and 55.3% relative improvements over the baseline system on the WSJ0-2mix-extr and Libri2Mix corpora in terms of equal-error-rate for 2-talker speech, while the performance of tSV for single-talker speech is on par with that of traditional speaker verification system, that is trained and evaluated under the same single-talker condition. Chenglin Xu, Wei Rao 0002, Jibin Wu, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2021 | D-Score: Holistic Dialogue Evaluation Without ReferenceabstractIn artistic gymnastics, difficulty score or D-score is used for judging performance. Starting from zero, an athlete earns points from different aspects such as composition requirement, difficulty, and connection between moves. The final score is a composition of the quality of various performance indicators. Similarly, when evaluating dialogue responses, human judges generally follow a number of criteria, among which language fluency, context coherence, logical consistency, and semantic appropriateness are on top of the agenda. In this paper, we propose an automatic dialogue evaluation framework called D-score that resembles the way gymnastics is evaluated. Following the four human judging criteria above, we devise a range of evaluation tasks and model them under a multi-task learning framework. The proposed framework, without relying on any human-written reference, learns to appreciate the overall quality of human-human conversations through a representation that is shared by all tasks without over-fitting to individual task domain. We evaluate D-score by performing comprehensive correlation analyses with human judgement on three dialogue evaluation datasets, among which two are from past DSTC series, and benchmark against state-of-the-art baselines. D-score not only outperforms the best baseline by a large margin in terms of system-level Spearman correlation but also represents an important step towards explainable dialogue scoring. Chen Zhang 0055, Grandee Lee, Luis Fernando D'Haro, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2021 | Transfer Learning From Speech Synthesis to Voice Conversion With Non-Parallel Training DataabstractWe present a novel voice conversion (VC) framework by learning from a text-to-speech (TTS) synthesis system, that is called TTS-VC transfer learning or TTL-VC for short. We first develop a multi-speaker speech synthesis system with sequence-to-sequence encoder-decoder architecture, where the encoder extracts the linguistic representations of input text, while the decoder, conditioned on target speaker embedding, takes the context vectors and the attention recurrent network cell output to generate target acoustic features. We take advantage of the fact that TTS system maps input text to speaker independent context vectors, thus re-purpose such a mapping to supervise the training of the latent representations of an encoder-decoder voice conversion system. In the voice conversion system, the encoder takes speech instead of text as the input, while the decoder is functionally similar to the TTS decoder. As we condition the decoder on a speaker embedding, the system can be trained on non-parallel data for any-to-any voice conversion. During voice conversion training, we present both text and speech to speech synthesis and voice conversion networks respectively. At run-time, the voice conversion network uses its own encoder-decoder architecture without the need of text input. Experiments show that the proposed TTL-VC system outperforms two competitive voice conversion baselines consistently, namely phonetic posteriorgram and AutoVC methods, in terms of speech quality, naturalness, and speaker similarity. Mingyang Zhang 0003, Yi Zhou 0020, Li Zhao 0003, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2021 | Language Agnostic Speaker Embedding for Cross-Lingual Personalized Speech GenerationabstractCross-lingual personalized speech generation seeks to synthesize a target speakers voice from only a few training samples that are in a different language. One popular technique is to condition a speech synthesizer on a speaker embedding, that characterizes the target speaker. Unfortunately, such a speaker embedding is usually affected by the language being spoken, which compromises the speaker similarity in cross-lingual personalized speech generation. In this paper, we propose a novel speaker encoding mechanism that learns a language agnostic speaker embedding to characterize speaker individuality. Specifically, we adopt an encoder-decoder architecture to disentangle the language information from speaker embeddings via multi-task learning. We conduct experiments on both voice conversion and text-to-speech synthesis between English and Mandarin that involve cross-lingual speech generation. All objective and subjective evaluations consistently confirm that the proposed speaker embedding is language agnostic, thus improving cross-lingual personalized speech generation in terms of speaker similarity. Yi Zhou 0020, Xiaohai Tian, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2020 | Modeling Code-Switch Languages Using Bilingual Parallel CorpusabstractLanguage modeling is the technique to estimate the probability of a sequence of words.A bilingual language model is expected to model the sequential dependency for words across languages, which is difficult due to the inherent lack of suitable training data as well as diverse syntactic structure across languages.We propose a bilingual attention language model (BALM) that simultaneously performs language modeling objective with a quasi-translation objective to model both the monolingual as well as the cross-lingual sequential dependency.The attention mechanism learns the bilingual context from a parallel corpus.BALM achieves state-of-theart performance on the SEAME code-switch database by reducing the perplexity of 20.5% over the best-reported result.We also apply BALM in bilingual lexicon induction, and language normalization tasks to validate the idea. Grandee Lee, Haizhou Li 0001 |
ACL | 2 |
| 2020 | Robust Real-time Face Tracking for People Wearing Face MasksabstractDue to the outbreak of the novel coronavirus (or known as COVID-19), people are advised to wear masks when they stay outdoors in many countries. This could result in difficulty for some public safety surveillance systems involving face detection or tracking. Therefore, the development of face detection and tracking algorithms for people wearing face masks is particularly important. In this paper, a real-time tracking algorithm for people with or without face masks is proposed. This algorithm is trained on public face datasets with faces without masks. Although the training does not involve face images of people wearing face masks, we show that the proposed algorithm is robust as it is able to perform well in face tracking for people wearing face masks. We also discuss the possible scenarios where the algorithm could lose track of the target when experimenting in tracking masked faces. This can motivate future research in this area. Xinggan Peng, Huiping Zhuang, Guang-Bin Huang, Haizhou Li 0001, Zhiping Lin 0001 |
ICARCV | 4 |
| 2020 | Real-Time Multiple Object Tracking with Discriminative FeaturesabstractTracking-by-detection methods track multiple objects by detecting the objects of interest in each frame and associating the detected objects with the tracks. By allowing object detection and appearance embedding to be learned in a shared network, recent tracking-by-detection methods can implement the tracking task in real time with the power of deep neural networks. However, they just focus on the detection stage and do not take advantage of the embedding features well in the association stage. In this paper, we exploit the discriminative embedding features in the association stage to improve the tracking performance. By combing the embedding features with the bounding boxes to associate the detected objects with the tracks, the number of identity switches during tracking can be reduced. Further, after associating the detected objects with the tracks, the embedding feature of each track is not only updated according to the associated object, but also learned to distinguish the similar detected objects. The experiments show that our method can achieve competitive tracking performance in real time compared to the state-of-the-art tracking methods. Zhenyu Weng, Yuesheng Zhu, Zhiping Lin 0001, Haizhou Li 0001 |
ICARCV | 4 |
| 2020 | Teacher-Student Training For Robust Tacotron-Based TTSabstractWhile neural end-to-end text-to-speech (TTS) is superior to conventional statistical methods in many ways, the exposure bias problem in the autoregressive models remains an issue to be resolved. The exposure bias problem arises from the mismatch between the training and inference process, that results in unpredictable performance for out-of-domain test data at run-time. To overcome this, we propose a teacher-student training scheme for Tacotron-based TTS by introducing a distillation loss function in addition to the feature loss function. We first train a Tacotron2-based TTS model by always providing natural speech frames to the decoder, that serves as a teacher model. We then train another Tacotron2-based model as a student model, of which the decoder takes the predicted speech frames as input, similar to how the decoder works during run-time inference. With the distillation loss, the student model learns the output probabilities from the teacher model, that is called knowledge distillation. Experiments show that our proposed training scheme consistently improves the voice quality for out-of-domain test data both in Chinese and English systems. Rui Liu 0008, Berrak Sisman, Jingdong Li, Feilong Bao, Guanglai Gao, Haizhou Li 0001 |
ICASSP | 6 |
| 2020 | On the Importance of Vocal Tract Constriction for Speaker Characterization: The Whispered Speech StudyabstractCharacterizing speakers under stressed condition is a challenge because speakers deviate from the normal speech production process. Whispered speech is one among them that is produced by abducting the vocal folds to pass the air out of mouth. During this process, the airflow thus passed is influenced by the physiological structure of vocal tract, that can be described by the level of constriction in vocal tract during speech production. We believe such information may be helpful for capturing speaker traits in such scenarios. The vocal tract constriction evidence in combination with mel frequency cepstral coefficients is investigated for speaker recognition experiments. The studies are conducted on CHAINS corpus that includes both neutral and whispered speech. We find that considering the constriction evidence from vocal tract improves the speaker characterization for whispered speech and under different mismatched scenarios. Rohan Kumar Das, Haizhou Li 0001 |
ICASSP | 2 |
| 2020 | Assessing the Scope of Generalized Countermeasures for Anti-SpoofingabstractMost of the research on anti-spoofing countermeasures are specific to a type of spoofing attacks, where models are trained on data of a particular nature, either synthetic or replay. However, one does not have such leverage as there is no prior knowledge about the kind of spoofing attack in practice. Therefore, there is a requirement to assess the scope of generalized countermeasures for anti-spoofing. The ASVspoof 2019 challenge covers both synthetic as well as replay attacks, which makes the database suitable for such study. In this work, we consider widely popular constant-Q cepstral coefficient features along with two other promising front-ends that capture long-term signal characteristics to assess their scope as generalized countermeasures. Additionally, a comprehensive study is made across different editions of ASVspoof corpora to highlight the need of robust generalized countermeasures in unseen conditions. Rohan Kumar Das, Haizhou Li 0001 |
ICASSP | 3 |
| 2020 | Effective Wavenet Adaptation for Voice Conversion with Limited DataabstractWaveNet has shown its great potential as a direct conversion model in voice conversion. However, due to the model complexity, WaveNet always requires a large amount of training data, which has limited its applications in voice conversion, where training data is scarce. In this paper, we propose a WaveNet adaptation method that effectively reduces the need of adaptation data. We first train a speaker independent WaveNet conversion model with multi-speaker dataset. Adaptation is then applied with limited target speaker’s data. Specifically, singular value decomposition (SVD) is applied to dilated convolution layers of WaveNet to reduce the number of parameters, which makes adaptation more effective with limited data. Experiments conducted on CMU-ARCTIC and CSTR-VCTK corpus show that the proposed method outperforms baseline methods in terms of both quality and similarity. Hongqiang Du, Xiaohai Tian, Lei Xie 0001, Haizhou Li 0001 |
ICASSP | 4 |
| 2020 | Automatic Lyrics Alignment and Transcription in Polyphonic Music: Does Background Music Help?abstractAutomatic lyrics alignment and transcription in polyphonic music are challenging tasks because the singing vocals are corrupted by the background music. In this work, we propose to learn music genre-specific characteristics to train polyphonic acoustic models. We first compare several automatic speech recognition pipelines for the application of lyrics transcription. We then present the lyrics alignment and transcription performance of music-informed acoustic models for the best-performing pipeline, and systematically study the impact of music genre and language model on the performance. With such genre-based approach, we explicitly model the music without removing it during acoustic modeling. The proposed approach outperforms all competing systems in the lyrics alignment and transcription tasks on well-known polyphonic test datasets. Chitralekha Gupta, Emre Yilmaz 0001, Haizhou Li 0001 |
ICASSP | 3 |
| 2020 | Time-Domain Neural Network Approach for Speech Bandwidth ExtensionabstractIn this paper, we study the time-domain neural network approach for speech bandwidth extension. We propose a network architecture, named multi-scale fusion neural network (MfNet), that gradually restores the low-frequency signal and predicts the high-frequency signal through the exchange of information across different scale representations. We propose a training scheme to optimize the network with a combination of perceptual loss and time-domain adversarial loss. Experiments show the proposed multi-scale fusion network consistently outperforms the competing methods in terms of perceptual evaluation of speech quality (PESQ), signal to distortion rate (SDR), signal to noise ratio (SNR), log-spectral distance (LSD) and word error rate (WER). More promisingly, the multi-scale fusion network requires only 10% of the parameters of the time-domain reference baseline. Chenglin Xu, Nana Hou, Lei Xie 0001, Chng Eng Siong, Haizhou Li 0001 |
ICASSP | 6 |
| 2020 | Independent Language Modeling Architecture for End-To-End ASRabstractThe attention-based end-to-end (E2E) automatic speech recognition (ASR) architecture allows for joint optimization of acoustic and language models within a single network. However, in a vanilla E2E ASR architecture, the decoder sub-network (subnet), which incorporates the role of the language model (LM), is conditioned on the encoder output. This means that the acoustic encoder and the language model are entangled that doesn’t allow language model to be trained separately from external text data. To address this problem, in this work, we propose a new architecture that separates the decoder subnet from the encoder output. In this way, the decoupled subnet becomes an independently trainable LM subnet, which can easily be updated using the external text data. We study two strategies for updating the new architecture. Experimental results show that, 1) the independent LM architecture benefits from external text data, achieving 9.3% and 22.8% relative character and word error rate reduction on Mandarin HKUST and English NSC datasets respectively; 2) the proposed architecture works well with external LM and can be generalized to different amount of labelled data. Van Tung Pham, Haihua Xu 0001, Yerbolat Khassanov, Zhiping Zeng, Chng Eng Siong, Chongjia Ni, Bin Ma 0001, Haizhou Li 0001 |
ICASSP | 8 |
| 2020 | End-to-End Code-Switching TTS with Cross-Lingual Language ModelabstractCode-switching text-to-speech (TTS) aims to enable a system to speak two languages with a single voice and in the same utterance. In this paper, we propose to incorporate cross-lingual word embedding into an end-to-end TTS system, to improve the voice rendering. The cross-lingual word embedding, generated from a pre-trained cross-lingual language model, is able to encode words of two languages in the same embedding space, therefore, allows words across languages to share each other's contextual information, which is useful for the voice rendering of code-switching content. To investigate the effectiveness of this idea, we conduct studies on two multi-speaker monolingual corpora, namely, THCHS30 Mandarin and LibriTTS English database. The evaluation results show that our proposed framework outperforms the baseline systems when presented with code-switching text input, and achieves state-of-the-art performance. Xuehao Zhou, Xiaohai Tian, Grandee Lee, Rohan Kumar Das, Haizhou Li 0001 |
ICASSP | 5 |
| 2020 | Speaker-Utterance Dual Attention for Speaker and Utterance VerificationabstractIn this paper, we study a novel technique that exploits the interaction between speaker traits and linguistic content to improve both speaker verification and utterance verification performance. We implement an idea of speaker-utterance dual attention (SUDA) in a unified neural network. The dual attention refers to an attention mechanism for the two tasks of speaker and utterance verification. The proposed SUDA features an attention mask mechanism to learn the interaction between the speaker and utterance information streams. This helps to focus only on the required information for respective task by masking the irrelevant counterparts. The studies conducted on RSR2015 corpus confirm that the proposed SUDA outperforms the framework without attention mask as well as several competitive systems for both speaker and utterance verification. Tianchi Liu 0004, Rohan Kumar Das, Maulik C. Madhavi, Shengmei Shen, Haizhou Li 0001 |
INTERSPEECH | 5 |
| 2020 | Low Latency Auditory Attention Detection with Common Spatial Pattern Analysis of EEG Signals
Siqi Cai 0002, Enze Su, Yonghao Song, Longhan Xie, Haizhou Li 0001 |
INTERSPEECH | 5 |
| 2020 | The Attacker's Perspective on Automatic Speaker Verification: An OverviewabstractSecurity of automatic speaker verification (ASV) systems is compromised by various spoofing attacks.While many types of non-proactive attacks (and their defenses) have been studied in the past, attacker's perspective on ASV, represents a far less explored direction.It can potentially help to identify the weakest parts of ASV systems and be used to develop attackeraware systems.We present an overview on this emerging research area by focusing on potential threats of adversarial attacks on ASV, spoofing countermeasures, or both.We conclude the study with discussion on selected attacks and leveraging from such knowledge to improve defense mechanisms against adversarial attacks. Rohan Kumar Das, Xiaohai Tian, Tomi Kinnunen, Haizhou Li 0001 |
INTERSPEECH | 4 |
| 2020 | SpEx+: A Complete Time Domain Speaker Extraction NetworkabstractSpeaker extraction aims to extract the target speech signal from a multi-talker environment given a target speaker's reference speech.We recently proposed a time-domain solution, SpEx, that avoids the phase estimation in frequency-domain approaches.Unfortunately, SpEx is not fully a time-domain solution since it performs time-domain speech encoding for speaker extraction, while taking frequency-domain speaker embedding as the reference.The size of the analysis window for timedomain and the size for frequency-domain input are also different.Such mismatch has an adverse effect on the system performance.To eliminate such mismatch, we propose a complete time-domain speaker extraction solution, that is called SpEx+.Specifically, we tie the weights of two identical speech encoder networks, one for the encoder-extractor-decoder pipeline, another as part of the speaker encoder.Experiments show that the SpEx+ achieves 0.8dB and 2.1dB SDR improvement over the state-of-the-art SpEx baseline, under different and same gender conditions on WSJ0-2mix-extr database respectively. Meng Ge, Chenglin Xu, Longbiao Wang, Chng Eng Siong, Jianwu Dang 0001, Haizhou Li 0001 |
INTERSPEECH | 6 |
| 2020 | Speaker and Phoneme-Aware Speech Bandwidth Extension with Residual Dual-Path NetworkabstractSpeech bandwidth extension aims to generate a wideband signal from a narrowband (low-band) input by predicting the missing high-frequency components. It is believed that the general knowledge about the speaker and phonetic content strengthens the prediction. In this paper, we propose to augment the low-band acoustic features with i-vector and phonetic posteriorgram (PPG), which represent speaker and phonetic content of the speech, respectively. We also propose a residual dual-path network (RDPN) as the core module to process the augmented features, which fully utilizes the utterance-level temporal continuity information and avoids gradient vanishing. Experiments show that the proposed method achieves 20.2% and 7.0% relative improvements over the best baseline in terms of log-spectral distortion (LSD) and signal-to-noise ratio (SNR), respectively. Furthermore, our method is 16 times more compact than the best baseline in terms of the number of parameters. Nana Hou, Chenglin Xu, Van Tung Pham, Joey Tianyi Zhou, Chng Eng Siong, Haizhou Li 0001 |
INTERSPEECH | 6 |
| 2020 | Multi-Task Learning for End-to-End Noise-Robust Bandwidth ExtensionabstractBandwidth extension aims to reconstruct wideband speech signals from narrowband inputs to improve perceptual quality. Prior studies mostly perform bandwidth extension under the assumption that the narrowband signals are clean without noise. The use of such extension techniques is greatly limited in practice when signals are corrupted by noise. To alleviate such problem, we propose an end-to-end time-domain framework for noise-robust bandwidth extension, that jointly optimizes a mask-based speech enhancement and an ideal bandwidth extension module with multi-task learning. The proposed framework avoids decomposing the signals into magnitude and phase spectra, therefore, requires no phase estimation. Experimental results show that the proposed method achieves 14.3% and 15.8% relative improvements over the best baseline in terms of perceptual evaluation of speech quality (PESQ) and log-spectral distortion (LSD), respectively. Furthermore, our method is 3 times more compact than the best baseline in terms of the number of parameters. Nana Hou, Chenglin Xu, Joey Tianyi Zhou, Chng Eng Siong, Haizhou Li 0001 |
INTERSPEECH | 5 |
| 2020 | Multi-Modal Attention for Speech Emotion RecognitionabstractEmotion represents an essential aspect of human speech that is manifested in speech prosody.Speech, visual, and textual cues are complementary in human communication.In this paper, we study a hybrid fusion method, referred to as multi-modal attention network (MMAN) to make use of visual and textual cues in speech emotion recognition.We propose a novel multimodal attention mechanism, cLSTM-MMA, which facilitates the attention across three modalities and selectively fuse the information.cLSTM-MMA is fused with other uni-modal subnetworks in the late fusion.The experiments show that speech emotion recognition benefits significantly from visual and textual cues, and the proposed cLSTM-MMA alone is as competitive as other fusion methods in terms of accuracy, but with a much more compact network structure.The proposed hybrid network MMAN achieves state-of-the-art performance on IEMOCAP database for emotion recognition. Zexu Pan, Zhaojie Luo, Haizhou Li 0001 |
INTERSPEECH | 4 |
| 2020 | The INTERSPEECH 2020 Far-Field Speaker Verification ChallengeabstractThe INTERSPEECH 2020 Far-Field Speaker Verification Challenge (FFSVC 2020) addresses three different research problems under well-defined conditions: far-field text-dependent speaker verification from single microphone array, far-field textindependent speaker verification from single microphone array, and far-field text-dependent speaker verification from distributed microphone arrays.All three tasks pose a cross-channel challenge to the participants.To simulate the real-life scenario, the enrollment utterances are recorded from close-talk cellphone, while the test utterances are recorded from the far-field microphone arrays.In this paper, we describe the database, the challenge, and the baseline system, which is based on a ResNetbased deep speaker network with cosine similarity scoring.For a given utterance, the speaker embeddings of different channels are equally averaged as the final embedding.The baseline system achieves minDCFs of 0.62, 0.66, and 0.64 and EERs of 6.27%, 6.55%, and 7.18% for task 1, task 2, and task 3, respectively. Xiaoyi Qin, Ming Li 0026, Hui Bu, Wei Rao 0002, Rohan Kumar Das, Shri Narayanan, Haizhou Li 0001 |
INTERSPEECH | 7 |
| 2020 | Audio-Visual Speaker Recognition with a Cross-Modal Discriminative NetworkabstractAudio-visual speaker recognition is one of the tasks in the recent 2019 NIST speaker recognition evaluation (SRE).Studies in neuroscience and computer science all point to the fact that vision and auditory neural signals interact in the cognitive process.This motivated us to study a cross-modal network, namely voice-face discriminative network (VFNet) that establishes the general relation between human voice and face.Experiments show that VFNet provides additional speaker discriminative information.With VFNet, we achieve 16.54% equal error rate relative reduction over the score level fusion audio-visual baseline on evaluation set of 2019 NIST SRE. Ruijie Tao, Rohan Kumar Das, Haizhou Li 0001 |
INTERSPEECH | 3 |
| 2020 | Light Convolutional Neural Network with Feature Genuinization for Detection of Synthetic Speech AttacksabstractModern text-to-speech (TTS) and voice conversion (VC) systems produce natural sounding speech that questions the security of automatic speaker verification (ASV).This makes detection of such synthetic speech very important to safeguard ASV systems from unauthorized access.Most of the existing spoofing countermeasures perform well when the nature of the attacks is made known to the system during training.However, their performance degrades in face of unseen nature of attacks.In comparison to the synthetic speech created by a wide range of TTS and VC methods, genuine speech has a more consistent distribution.We believe that the difference between the distribution of synthetic and genuine speech is an important discriminative feature between the two classes.In this regard, we propose a novel method referred to as feature genuinization that learns a transformer with convolutional neural network (CNN) using the characteristics of only genuine speech.We then use this genuinization transformer with a light CNN classifier.The ASVspoof 2019 logical access corpus is used to evaluate the proposed method.The studies show that the proposed feature genuinization based LCNN system outperforms other state-ofthe-art spoofing countermeasures, depicting its effectiveness for detection of synthetic speech attacks. Zhenzong Wu, Rohan Kumar Das, Haizhou Li 0001 |
INTERSPEECH | 4 |
| 2020 | Deep Convolutional Spiking Neural Networks for Keyword Spotting
Emre Yilmaz 0001, Özgür Bora Gevrek, Jibin Wu, Xuanbo Meng, Haizhou Li 0001 |
INTERSPEECH | 6 |
| 2020 | Self-and-Mixed Attention Decoder with Deep Acoustic Structure for Transformer-Based LVCSRabstractTransformer has shown impressive performance in automatic speech recognition.It uses an encoder-decoder structure with self-attention to learn the relationship between high-level representation of source inputs and embedding of target outputs.In this paper, we propose a novel decoder structure that features a self-and-mixed attention decoder (SMAD) with a deep acoustic structure (DAS) to improve the acoustic representation of Transformer-based LVCSR.Specifically, we introduce a self-attention mechanism to learn a multi-layer deep acoustic structure for multiple levels of acoustic abstraction.We also design a mixed attention mechanism that learns the alignment between different levels of acoustic abstraction and its corresponding linguistic information simultaneously in a shared embedding space.The ASR experiments on Aishell-1 show that the proposed structure achieves CERs of 4.8% on the dev set and 5.1% on the test set, which are the best reported results on this task to the best of our knowledge. Xinyuan Zhou, Grandee Lee, Emre Yilmaz 0001, Yanhua Long, Jiaen Liang, Haizhou Li 0001 |
INTERSPEECH | 6 |
| 2020 | Converting Anyone's Emotion: Towards Speaker-Independent Emotional Voice ConversionabstractEmotional voice conversion aims to convert the emotion of speech from one state to another while preserving the linguistic content and speaker identity.The prior studies on emotional voice conversion are mostly carried out under the assumption that emotion is speaker-dependent.We consider that there is a common code between speakers for emotional expression in a spoken language, therefore, a speaker-independent mapping between emotional states is possible.In this paper, we propose a speaker-independent emotional voice conversion framework, that can convert anyone's emotion without the need for parallel data.We propose a VAW-GAN based encoderdecoder structure to learn the spectrum and prosody mapping.We perform prosody conversion by using continuous wavelet transform (CWT) to model the temporal dependencies.We also investigate the use of F0 as an additional input to the decoder to improve emotion conversion performance.Experiments show that the proposed speaker-independent framework achieves competitive results for both seen and unseen speakers. Kun Zhou 0003, Berrak Sisman, Mingyang Zhang 0003, Haizhou Li 0001 |
INTERSPEECH | 4 |
| 2020 | Multi-Encoder-Decoder Transformer for Code-Switching Speech RecognitionabstractCode-switching (CS) occurs when a speaker alternates words of two or more languages within a single sentence or across sentences.Automatic speech recognition (ASR) of CS speech has to deal with two or more languages at the same time.In this study, we propose a Transformer-based architecture with two symmetric language-specific encoders to capture the individual language attributes, that improve the acoustic representation of each language.These representations are combined using a language-specific multi-head attention mechanism in the decoder module.Each encoder and its corresponding attention module in the decoder are pre-trained using a large monolingual corpus aiming to alleviate the impact of limited CS training data.We call such a network a multi-encoder-decoder (MED) architecture.Experiments on the SEAME corpus show that the proposed MED architecture achieves 10.2% and 10.8% relative error rate reduction on the CS evaluation sets with Mandarin and English as the matrix language respectively. Xinyuan Zhou, Emre Yilmaz 0001, Yanhua Long, Yijie Li 0001, Haizhou Li 0001 |
INTERSPEECH | 5 |
| 2020 | Two decades into Speaker Recognition Evaluation - are we there yet?
Kong-Aik Lee, Seyed Omid Sadjadi, Haizhou Li 0001, Douglas A. Reynolds |
Comput. Speech Lang. | 3 |
| 2020 | Supervised learning in spiking neural networks with synaptic delay-weight plasticity
Malu Zhang, Jibin Wu, Ammar Belatreche, Zihan Pan, Xiurui Xie, Yansong Chua, Guoqi Li 0002, Hong Qu 0002, Haizhou Li 0001 |
Neurocomputing | 9 |
| 2020 | DeepConversion: Voice conversion with limited parallel training dataabstractA deep neural network approach to voice conversion usually depends on a large amount of parallel training data from source and target speakers. In this paper, we propose a novel conversion pipeline, DeepConversion, that leverages a large amount of non-parallel, multi-speaker data, but requires only a small amount of parallel training data. It is believed that we can represent the shared characteristics of speakers by training a speaker independent general model on a large amount of publicly available, non-parallel, multi-speaker speech data. Such general model can then be used to learn the mapping between source and target speaker more effectively from a limited amount of parallel training data. We also propose a strategy to make full use of the parallel data in all models along the pipeline. In particular, the parallel data is used to adapt the general model towards the source-target speaker pair to achieve a coarse grained conversion, and to develop a compact Error Reduction Network (ERN) for a fine-grained conversion. The parallel data is also used to adapt the WaveNet vocoder towards the source-target pair. The experiments show that DeepConversion that only uses a limited amount of parallel training data, consistently outperforms the traditional approaches that use a large amount of parallel training data, in both objective and subjective evaluations. Mingyang Zhang 0003, Berrak Sisman, Li Zhao 0003, Haizhou Li 0001 |
Speech Commun. | 4 |
| 2020 | Modeling Prosodic Phrasing With Multi-Task Learning in Tacotron-Based TTSabstractTacotron-based end-to-end speech synthesis has shown remarkable voice quality. However, the rendering of prosody in the synthesized speech remains to be improved, especially for long sentences, where prosodic phrasing errors can occur frequently. In this letter, we extend the Tacotron-based speech synthesis framework to explicitly model the prosodic phrase breaks. We propose a multi-task learning scheme for Tacotron training, that optimizes the system to predict both Mel spectrum and phrase breaks. To our best knowledge, this is the first implementation of multi-task learning for Tacotron based TTS with a prosodic phrasing model. Experiments show that our proposed training scheme consistently improves the voice quality for both Chinese and Mongolian systems. Rui Liu 0008, Berrak Sisman, Feilong Bao, Guanglai Gao, Haizhou Li 0001 |
IEEE Signal Process. Lett. | 5 |
| 2020 | Multi-Task WaveRNN With an Integrated Architecture for Cross-Lingual Voice ConversionabstractSpoken languages are similar phonetically because humans have a common vocal production system. However, each language has a unique phonetic repertoire and phonotactic rule. In cross-lingual voice conversion, source speaker and target speaker speak different languages. The challenge is how to project the speaker identity of the source speaker to that of the target across two different phonetic systems. A typical voice conversion system employs a generator-vocoder pipeline, where the generator is responsible for conversion, and the vocoder is for waveform reconstruction. We propose a novel Multi-Task WaveRNN with an integrated architecture for cross-lingual voice conversion. The WaveRNN is trained on two sets of monolingual data via a two-task learning. The integrated architecture takes linguistic features as input and outputs speech waveform directly. Voice conversion experiments are conducted between English and Mandarin, which confirm the effectiveness of the proposed method in terms of speech quality and speaker similarity. Yi Zhou 0020, Xiaohai Tian, Haizhou Li 0001 |
IEEE Signal Process. Lett. | 3 |
| 2020 | Automatic Leaderboard: Evaluation of Singing Quality Without a Standard ReferenceabstractAutomatic evaluation of singing quality can be done with the help of a reference singing or the digital sheet music of the song. However, such a standard reference is not always available. In this article, we propose a framework to rank a large pool of singers according to their singing quality without any standard reference. We define musically motivated absolute measures based on pitch histogram, and relative measures based on inter-singer statistics to evaluate the quality of singing attributes such as intonation, and rhythm. The absolute measures evaluate the goodness of pitch histogram specific to a singer, while the relative measures use the similarity between singers in terms of pitch, rhythm, and timbre as an indicator of singing quality. With the relative measures, we formulate the concept of veracity or truth-finding for the ranking of singing quality. We successfully validate a self-organizing approach to rank-ordering a large pool of singers. The fusion of absolute and relative measures results in an average Spearman's rank correlation of 0.71 with human judgments in a 10-fold cross-validation experiment, which is close to the inter-judge correlation. Chitralekha Gupta, Haizhou Li 0001, Ye Wang 0007 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | SpEx: Multi-Scale Time Domain Speaker Extraction NetworkabstractSpeaker extraction aims to mimic humans' selective auditory attention by extracting a target speaker's voice from a multi-talker environment. It is common to perform the extraction in frequency-domain, and reconstruct the time-domain signal from the extracted magnitude and estimated phase spectra. However, such an approach is adversely affected by the inherent difficulty of phase estimation. Inspired by Conv-TasNet, we propose a time-domain speaker extraction network (SpEx) that converts the mixture speech into multi-scale embedding coefficients instead of decomposing the speech signal into magnitude and phase spectra. In this way, we avoid phase estimation. The SpEx network consists of four network components, namely speaker encoder, speech encoder, speaker extractor, and speech decoder. Specifically, the speech encoder converts the mixture speech into multi-scale embedding coefficients, the speaker encoder learns to represent the target speaker with a speaker embedding. The speaker extractor takes the multi-scale embedding coefficients and target speaker embedding as input and estimates a receptive mask. Finally, the speech decoder reconstructs the target speaker's speech from the masked embedding coefficients. We also propose a multi-task learning framework and a multi-scale embedding implementation. Experimental results show that the proposed SpEx achieves 37.3%, 37.7% and 15.0% relative improvements over the best baseline in terms of signal-to-distortion ratio (SDR), scale-invariant SDR (SI-SDR), and perceptual evaluation of speech quality (PESQ) under an open evaluation condition. Chenglin Xu, Wei Rao 0002, Chng Eng Siong, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2020 | Significance of Subband Features for Synthetic Speech DetectionabstractIn text-to-speech or voice conversion based synthetic speech detection, it is a common practice that spectral information over the entire frequency band is used for feature representation. We propose a new method, referred to as subband transform, that characterizes the signals by subband. It is found that subband transform captures the artifacts in synthetic speech more effectively than full band transform. We propose equal subband transform, octave subband transform, and mel subband transform for three novel features, namely, constant-Q equal subband transform (CQ-EST), constant-Q octave subband transform (CQ-OST) and discrete Fourier mel subband transform (DF-MST). We evaluate the three features on the ASVspoof 2015, noisy ASVspoof 2015 and ASVspoof 2019 logical access corpora. The experiments show that the proposed CQ-EST feature achieves an average equal error rate of 0.056% on ASVspoof 2015 evaluation set. The study observes that the features based on subband transform outperform those based on full band transform under both clean and noisy conditions. In addition, the tandem detection cost function of CQ-OST can reach 0.188 on ASVspoof 2019 logical access evaluation set. Rohan Kumar Das, Haizhou Li 0001 |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2019 | MPD-AL: An Efficient Membrane Potential Driven Aggregate-Label Learning Algorithm for Spiking NeuronsabstractOne of the long-standing questions in biology and machine learning is how neural networks may learn important features from the input activities with a delayed feedback, commonly known as the temporal credit-assignment problem. The aggregate-label learning is proposed to resolve this problem by matching the spike count of a neuron with the magnitude of a feedback signal. However, the existing threshold-driven aggregate-label learning algorithms are computationally intensive, resulting in relatively low learning efficiency hence limiting their usability in practical applications. In order to address these limitations, we propose a novel membrane-potential driven aggregate-label learning algorithm, namely MPD-AL. With this algorithm, the easiest modifiable time instant is identified from membrane potential traces of the neuron, and guild the synaptic adaptation based on the presynaptic neurons’ contribution at this time instant. The experimental results demonstrate that the proposed algorithm enables the neurons to generate the desired number of spikes, and to detect useful clues embedded within unrelated spiking activities and background noise with a better learning efficiency over the state-of-the-art TDP1 and Multi-Spike Tempotron algorithms. Furthermore, we propose a data-driven dynamic decoding scheme for practical classification tasks, of which the aggregate labels are hard to define. This scheme effectively improves the classification accuracy of the aggregate-label learning algorithms as demonstrated on a speech recognition task. Malu Zhang, Jibin Wu, Yansong Chua, Xiaoling Luo 0001, Zihan Pan, Haizhou Li 0001 |
AAAI | 7 |
| 2019 | Long Range Acoustic and Deep Features Perspective on ASVspoof 2019abstractTo secure automatic speaker verification (ASV) systems from intruders, robust countermeasures for spoofing attack detection are required. The ASVspoof series of challenge provides a shared anti-spoofing task. The recent edition, ASVspoof 2019, focuses on attacks by both synthetic and replay speech that are referred to as logical and physical access attacks, respectively. In the ASVspoof 2019 submission, we considered novel countermeasures based on long range acoustic features, that are unique in many ways as they are derived using octave power spectrum and subbands, as opposed to the commonly used linear power spectrum. During the post-challenge study, we further investigate the use of deep features that enhances the discriminative ability between genuine and spoofed speech. In this paper, we summarize the findings from the perspective of long range acoustic and deep features for spoof detection. We make a comprehensive analysis on the nature of different kinds of spoofing attacks and system development. Rohan Kumar Das, Haizhou Li 0001 |
ASRU | 3 |
| 2019 | WaveNet Factorization with Singular Value Decomposition for Voice ConversionabstractWaveNet vocoder has seen its great advantage over traditional vocoders in voice quality. However, it usually requires a relatively large amount of speech data to train a speaker-dependent WaveNet vocoder. Therefore, it remains a challenge to build a high-quality WaveNet vocoder for low resource tasks, e.g. voice conversion, where speech samples are limited in real applications. We propose to use singular value decomposition (SVD) to reduce WaveNet parameters while maintaining its output voice quality. Specifically, we apply SVD on dilated convolution layers, and impose semi-orthogonal constraint to improve the performance. Experiments conducted on CMU-ARCTIC database show that as compared with the original WaveNet vocoder, the proposed method maintains similar performance, in terms of both quality and similarity, while using much less training data. Hongqiang Du, Xiaohai Tian, Lei Xie 0001, Haizhou Li 0001 |
ASRU | 4 |
| 2019 | On the Study of Generative Adversarial Networks for Cross-Lingual Voice ConversionabstractCross-lingual voice conversion (VC) aims to convert the source speaker's voice to sound like that of the target speaker, when the source and target speakers speak different languages. In this paper, we propose to use Generative Adversarial Networks (GANs) for cross-lingual voice-conversion. We further the studies on Variational Autoencoding Wasserstein GAN (VAW-GAN) and cycle-consistent adversarial network (CycleGAN), that are known to be effective for mono-lingual voice conversion. As cross-lingual voice conversion needs to converts the voice across different phonetic system, it is more challenging than mono-lingual voice conversion. By using VAW-GAN and CycleGAN, we successfully convert the speaker identity while carrying over the source speaker's linguistic content. The proposed idea is unique in the sense that it neither relies on bilingual data and their alignment, nor any external process, such as ASR. Moreover, it works with limited amount of training data of any two languages. To our best knowledge, this is the first comprehensive study of Generative Adversarial Networks in cross-lingual voice conversion. In the experiments, we achieve high-quality converted voice, that performs equally well or better than mono-lingual voice conversion. Berrak Sisman, Mingyang Zhang 0003, Minghui Dong, Haizhou Li 0001 |
ASRU | 4 |
| 2019 | Time-Domain Speaker Extraction NetworkabstractSpeaker extraction is to extract a target speaker's voice from multi-talker speech. It simulates humans' cocktail party effect or the selective listening ability. The prior work mostly performs speaker extraction in frequency domain, then reconstructs the signal with some phase approximation. The inaccuracy of phase estimation is inherent to the frequency domain processing, that affects the quality of signal reconstruction. In this paper, we propose a time-domain speaker extraction network (TseNet) that doesn't decompose the speech signal into magnitude and phase spectrums, therefore, doesn't require phase estimation. The TseNet consists of a stack of dilated depthwise separable convolutional networks, that capture the long-range dependency of the speech signal with a manageable number of parameters. It is also conditioned on a reference voice from the target speaker, that is characterized by speaker i-vector, to perform the selective listening to the target speaker. Experiments show that the proposed TseNet achieves 16.3% and 7.0% relative improvements over the baseline in terms of signal-to-distortion ratio (SDR) and perceptual evaluation of speech quality (PESQ) under open evaluation condition. Chenglin Xu, Wei Rao 0002, Chng Eng Siong, Haizhou Li 0001 |
ASRU | 4 |
| 2019 | End-to-End Code-Switching ASR for Low-Resourced Language PairsabstractDespite the significant progress in end-to-end (E2E) automatic speech recognition (ASR), E2E ASR for low resourced code-switching (CS) speech has not been well studied. In this work, we describe an E2E ASR pipeline for the recognition of CS speech in which a low-resourced language is mixed with a high resourced language. Low-resourcedness in acoustic data hinders the performance of E2E ASR systems more severely than the conventional ASR systems. To mitigate this problem in the transcription of archives with code-switching Frisian-Dutch speech, we integrate a designated decoding scheme and perform rescoring with neural network-based language models to enable better utilization of the available textual resources. We first incorporate a multi-graph decoding approach which creates parallel search spaces for each monolingual and mixed recognition tasks to maximize the utilization of the textual resources from each language. Further, language model rescoring is performed using a recurrent neural network pre-trained with cross-lingual embedding and further adapted with the limited amount of in-domain CS text. The ASR experiments demonstrate the effectiveness of the described techniques in improving the recognition performance of an E2E CS ASR system in a low-resourced scenario. Xianghu Yue, Grandee Lee, Emre Yilmaz 0001, Fang Deng, Haizhou Li 0001 |
ASRU | 5 |
| 2019 | A Modularized Neural Network with Language-Specific Output Layers for Cross-Lingual Voice ConversionabstractThis paper presents a cross-lingual voice conversion framework that adopts a modularized neural network. The modularized neural network has a common input structure that is shared for both languages, and two separate output modules, one for each language. The idea is motivated by the fact that phonetic systems of languages are similar because humans share a common vocal production system, but acoustic renderings, such as prosody and phonotactic, vary a lot from language to language. The modularized neural network is trained to map Phonetic PosteriorGram (PPG) to acoustic features for multiple speakers. It is conditioned on a speaker i-vector to generate the desired target voice. We validated the idea between English and Mandarin languages in objective and subjective tests. In addition, mixed-lingual PPG derived from a unified English-Mandarin acoustic model is proposed to capture the linguistic information from both languages. It is found that our proposed modularized neural network significantly outperforms the baseline approaches in terms of speech quality and speaker individuality, and mixed-lingual PPG representation further improves the conversion performance. Yi Zhou 0020, Xiaohai Tian, Emre Yilmaz 0001, Rohan Kumar Das, Haizhou Li 0001 |
ASRU | 5 |
| 2019 | Word and Class Common Space Embedding for Code-switch Language ModellingabstractCode-switch language modelling is challenging due to limited linguistic resources and less predictable word sequences. Many state-of-the-art systems rely on linguistic information such as Part-of-Speech (POS) or classes to generalize the lexicon. Such systems generally use multi-task learning or conditional network to improve over baseline RNN language model by providing a better word prediction. To overcome the data sparsity through continuous space modelling and back-off mechanism, we propose to constrain the word and class embedding in a common space by means of cross-lingual word embedding, and to make use of the predicted class embedding as a back-off scheme when word prediction model is weak. The proposed word and class Common Space embedding Language Model (CSLM) is able to model word prediction better and is more robust when only sparse training data are available. The CSLM outperforms the state-of-the-art language model by 9.7% on the code-switch SEAME corpus. Grandee Lee, Haizhou Li 0001 |
ICASSP | 2 |
| 2019 | Automatic Lyrics-to-audio Alignment on Polyphonic Music Using Singing-adapted Acoustic ModelsabstractLyrics-to-audio alignment is to automatically align the lyrical words with the mixed singing audio (singing voice+musical accompaniment). Such alignment can be achieved with an automatic speech recognition (ASR) system. We propose to adapt the acoustic model of a speech recognizer towards solo singing voice. This avoids the hurdles of annotating a large polyphonic music training dataset. Moreover, a lexicon-modification based duration modelling has been incorporated to account for the long duration vowels in singing. As practical application demand the alignment on polyphonic music, we study the effect of different singing vocal separation methods in the task of lyrics-to-audio alignment in polyphonic music. The extracted vocals are forced-aligned with the singing-adapted models. We demonstrate that the use of audio source separation method and effective end-pointing of the songs has a high impact on the alignment performance through the experiments. We report a mean average absolute error of 3.87 seconds, which is comparable with the state-of-the-art lyrics-to-audio alignment system that is trained on a large polyphonic music database. Bidisha Sharma, Chitralekha Gupta, Haizhou Li 0001, Ye Wang 0007 |
ICASSP | 3 |
| 2019 | Auditory Inspired Spatial Differentiation for Replay Spoofing Attack DetectionabstractThe security of Automatic Speaker Verification systems is greatly threatened by spoofing attacks of various kinds. Among them, replay attacks are noteworthy due to the ease with which they can be employed. Most countermeasures for replay attacks use subband features based on parallel filter banks. This paper explores the effect of `spatial differentiation' used in auditory system modelling to improve frequency selectivity and hence provide a more selective front-end for replay attack detection. Experiments were done using a parallel filter bank consisting of simple 2ndorder IIR bandpass filters following which, processing analogous to spatial differentiation was employed to obtain higher order stable IIR filters, in turn leading to highly selective filter banks. Two novel features based on spatially differentiated higher order filter bank have been proposed. Together they yield a relative improvement of 29.9% in replay speech detection over a constant Q transform based baseline system, when evaluated on the ASVspoof 2017 Version 2.0 database. Buddhi Wickramasinghe, Eliathamby Ambikairajah, Julien Epps, Vidhyasaharan Sethu, Haizhou Li 0001 |
ICASSP | 5 |
| 2019 | Optimization of Speaker Extraction Neural Network with Magnitude and Temporal Spectrum Approximation LossabstractThe SpeakerBeam-FE (SBF) method is proposed for speaker extraction. It attempts to overcome the problem of unknown number of speakers in an audio recording during source separation. The mask approximation loss of SBF is sub-optimal, which doesn't calculate direct signal reconstruction error and consider the speech context. To address these problems, this paper proposes a magnitude and temporal spectrum approximation loss to estimate a phase sensitive mask for the target speaker with the speaker characteristics. Moreover, this paper explores a concatenation framework instead of the context adaptive deep neural network in the SBF method to encode a speaker embedding into the mask estimation network. Experimental results under open evaluation condition show that the proposed method achieves 70.4% and 17.7% relative improvement over the SBF baseline on signal-to-distortion ratio (SDR) and perceptual evaluation of speech quality (PESQ), respectively. A further analysis demonstrates 69.1% and 72.3% relative SDR improvements obtained by the proposed method for different and same gender mixtures. Chenglin Xu, Wei Rao 0002, Chng Eng Siong, Haizhou Li 0001 |
ICASSP | 4 |
| 2019 | Cross-lingual Voice Conversion with Bilingual Phonetic Posteriorgram and Average ModelingabstractThis paper presents a cross-lingual voice conversion approach using bilingual Phonetic PosteriorGram (PPG) and average modeling. The proposed approach makes use of bilingual PPGs to represent speaker-independent features of speech signals from different languages in the same feature space. In particular, a bilingual PPG is formed by stacking two monolingual PPG vectors, which are extracted from two monolingual speech recognition systems. The conversion model is trained to learn the relationship between bilingual PPGs and the corresponding acoustic features. To leverage the linguistic and acoustic information from other speakers in different languages, an average model is trained with multiple speakers in both source and target languages. I-vector is utilized as an additional input feature of the average model for network adaptation. Experiments are performed for intralingual and cross-lingual voice conversion between English and Mandarin speakers. Both objective and subjective evaluations demonstrate the effectiveness of our proposed approach. Yi Zhou 0020, Xiaohai Tian, Haihua Xu 0001, Rohan Kumar Das, Haizhou Li 0001 |
ICASSP | 5 |
| 2019 | Neural Population Coding for Effective Temporal ClassificationabstractNeural encoding plays an important role in faithfully describing the temporally rich patterns, whose instances include human speech and environmental sounds. To classify such spatio-temporal patterns with the Spiking Neural Networks (SNNs), how these patterns are encoded has a direct impact on the complexity of the task. In this paper, we study several existing temporal and population coding schemes in speech and audio recognition. We show that, with population neural coding, the encoded patterns are linearly separable using the Support Vector Machine (SVM). We note that the population neural coding effectively project the temporal information into the spatial domain, thus improving linear separability of the patterns. We achieve an accuracy of 95% and 100% on TIDIGITS and RWCP datasets respectively with SVM classifier. We further implement the Tempotron as an SNN-based classifier on the same datasets and achieve similar results. The study suggests that an effective neural coding scheme is just as important as the classifier. Zihan Pan, Jibin Wu, Malu Zhang, Haizhou Li 0001, Yansong Chua |
IJCNN | 4 |
| 2019 | Deep Spiking Neural Network with Spike Count based Learning RuleabstractDeep spiking neural networks (SNNs) support asynchronous event-driven computation, massive parallelism and demonstrate great potential to improve the energy efficiency of its synchronous analog counterpart. However, insufficient attention has been paid to neural encoding when designing SNN learning rules. Remarkably, the temporal credit assignment has been performed on rate-coded spiking inputs, leading to poor learning efficiency. In this paper, we introduce a novel spike-based learning rule for rate-coded deep SNNs, whereby the spike count of each neuron is used as a surrogate for gradient backpropagation. We evaluate the proposed learning rule by training deep spiking multi-layer perceptron (MLP) and spiking convolutional neural network (CNN) on the UCI machine learning and MNIST handwritten digit datasets. We show that the proposed learning rule achieves state-of-the-art accuracies on all benchmark datasets. The proposed learning rule allows introducing latency, spike rate and hardware constraints into the SNN learning, which is superior to the indirect approach in which conventional artificial neural networks are first trained and then converted to SNNs. Hence, it allows direct deployment to the neuromorphic hardware and supports efficient inference. Notably, a test accuracy of 98.40% was achieved on the MNIST dataset in our experiments with only 10 simulation time steps, when the same latency constraint is imposed during training. Jibin Wu, Yansong Chua, Malu Zhang, Qu Yang, Guoqi Li 0002, Haizhou Li 0001 |
IJCNN | 6 |
| 2019 | Competitive STDP-based Feature Representation Learning for Sound Event ClassificationabstractHumans are good at discriminating environmental sounds and associating them with opportunities or dangers. While the deep learning approach to sound event classification (SEC) is achieving human parity, unsolved problems remain, instances include high computational cost, requirement of massive labeled training data, and question of biological plausibility. Motivated by the human auditory system, we propose a biologically plausible SEC system, which integrates the auditory front-end, population coding, competitive spike-timing-dependent plasticity (STDP) based feature representation learning and supervised temporal classification into a unified spiking neural network (SNN) system. The proposed SEC system achieves a classification accuracy on the RWCP database that is on par with other competitive baseline systems. Furthermore, the STDP-based feature representation learning shows low intra-class variability and high inter-class variability in our experiments, which is highly desirable for pattern classification tasks. Jibin Wu, Malu Zhang, Haizhou Li 0001, Yansong Chua |
IJCNN | 3 |
| 2019 | Joint Training Framework for Text-to-Speech and Voice Conversion Using Multi-Source Tacotron and WaveNetabstractWe investigated the training of a shared model for both textto-speech (TTS) and voice conversion (VC) tasks.We propose using an extended model architecture of Tacotron, that is a multi-source sequence-to-sequence model with a dual attention mechanism as the shared model for both the TTS and VC tasks.This model can accomplish these two different tasks respectively according to the type of input.An end-to-end speech synthesis task is conducted when the model is given text as the input while a sequence-to-sequence voice conversion task is conducted when it is given the speech of a source speaker as the input.Waveform signals are generated by using WaveNet, which is conditioned by using a predicted mel-spectrogram.We propose jointly training a shared model as a decoder for a target speaker that supports multiple sources.Listening experiments show that our proposed multi-source encoder-decoder model can efficiently achieve both the TTS and VC tasks. Mingyang Zhang 0003, Xin Wang 0037, Fuming Fang, Haizhou Li 0001, Junichi Yamagishi |
INTERSPEECH | 4 |
| 2019 | Instantaneous Phase and Long-Term Acoustic Cues for Orca Activity Detection
Rohan Kumar Das, Haizhou Li 0001 |
INTERSPEECH | 2 |
| 2019 | Long Range Acoustic Features for Spoofed Speech Detection
Rohan Kumar Das, Haizhou Li 0001 |
INTERSPEECH | 3 |
| 2019 | An Adaptive-Q Cochlear Model for Replay Spoofing Detection
Tharshini Gunendradasan, Eliathamby Ambikairajah, Julien Epps, Haizhou Li 0001 |
INTERSPEECH | 4 |
| 2019 | NUS Speak-to-Sing: A Web Platform for Personalized Speech-to-Singing Conversion
Chitralekha Gupta, Karthika Vijayan, Bidisha Sharma, Xiaoxue Gao, Haizhou Li 0001 |
INTERSPEECH | 5 |
| 2019 | Acoustic Modeling for Automatic Lyrics-to-Audio AlignmentabstractAutomatic lyrics to polyphonic audio alignment is a challenging task not only because the vocals are corrupted by background music, but also there is a lack of annotated polyphonic corpus for effective acoustic modeling. In this work, we propose (1) using additional speech and music-informed features and (2) adapting the acoustic models trained on a large amount of solo singing vocals towards polyphonic music using a small amount of in-domain data. Incorporating additional information such as voicing and auditory features together with conventional acoustic features aims to bring robustness against the increased spectro-temporal variations in singing vocals. By adapting the acoustic model using a small amount of polyphonic audio data, we reduce the domain mismatch between training and testing data. We perform several alignment experiments and present an in-depth alignment error analysis on acoustic features, and model adaptation techniques. The results demonstrate that the proposed strategy provides a significant error reduction of word boundary alignment over comparable existing systems, especially on more challenging polyphonic data with long-duration musical interludes. Chitralekha Gupta, Emre Yilmaz 0001, Haizhou Li 0001 |
INTERSPEECH | 3 |
| 2019 | I4U Submission to NIST SRE 2018: Leveraging from a Decade of Shared ExperiencesabstractThe I4U consortium was established to facilitate a joint entry to NIST speaker recognition evaluations (SRE). The latest edition of such joint submission was in SRE 2018, in which the I4U submission was among the best-performing systems. SRE'18 also marks the 10-year anniversary of I4U consortium into NIST SRE series of evaluation. The primary objective of the current paper is to summarize the results and lessons learned based on the twelve sub-systems and their fusion submitted to SRE'18. It is also our intention to present a shared view on the advancements, progresses, and major paradigm shifts that we have witnessed as an SRE participant in the past decade from SRE'08 to SRE'18. In this regard, we have seen, among others, a paradigm shift from supervector representation to deep speaker embedding, and a switch of research challenge from channel compensation to domain adaptation. Kong-Aik Lee, Ville Hautamäki, Tomi Kinnunen, Hitoshi Yamamoto, Koji Okabe, Ville Vestman, Jing Huang 0019, Guo-Hong Ding, Hanwu Sun, Anthony Larcher, Rohan Kumar Das, Haizhou Li 0001, Mickael Rouvier, Pierre-Michel Bousquet, Wei Rao 0002, Qing Wang 0039, Fahimeh Bahmaninezhad, Héctor Delgado, Massimiliano Todisco |
INTERSPEECH | 12 |
| 2019 | Linguistically Motivated Parallel Data Augmentation for Code-Switch Language Modeling
Grandee Lee, Xianghu Yue, Haizhou Li 0001 |
INTERSPEECH | 3 |
| 2019 | A Unified Framework for Speaker and Utterance Verification
Tianchi Liu 0004, Maulik C. Madhavi, Rohan Kumar Das, Haizhou Li 0001 |
INTERSPEECH | 4 |
| 2019 | Target Speaker Extraction for Multi-Talker Speaker Verification
Wei Rao 0002, Chenglin Xu, Chng Eng Siong, Haizhou Li 0001 |
INTERSPEECH | 4 |
| 2019 | A Combination of Model-Based and Feature-Based Strategy for Speech-to-Singing Alignment
Bidisha Sharma, Haizhou Li 0001 |
INTERSPEECH | 2 |
| 2019 | Multi-Level Adaptive Speech Activity Detector for Speech in Naturalistic Environments
Bidisha Sharma, Rohan Kumar Das, Haizhou Li 0001 |
INTERSPEECH | 3 |
| 2019 | On the Importance of Audio-Source Separation for Singer Identification in Polyphonic Music
Bidisha Sharma, Rohan Kumar Das, Haizhou Li 0001 |
INTERSPEECH | 3 |
| 2019 | A Speaker-Dependent WaveNet for Voice Conversion with Non-Parallel Data
Xiaohai Tian, Chng Eng Siong, Haizhou Li 0001 |
INTERSPEECH | 3 |
| 2019 | VQVAE Unsupervised Unit Discovery and Multi-Scale Code2Spec Inverter for Zerospeech Challenge 2019abstractWe describe our submitted system for the ZeroSpeech Challenge 2019.The current challenge theme addresses the difficulty of constructing a speech synthesizer without any text or phonetic labels and requires a system that can (1) discover subword units in an unsupervised way, and (2) synthesize the speech with a target speaker's voice.Moreover, the system should also balance the discrimination score ABX, the bit-rate compression rate, and the naturalness and the intelligibility of the constructed voice.To tackle these problems and achieve the best tradeoff, we utilize a vector quantized variational autoencoder (VQ-VAE) and a multi-scale codebook-tospectrogram (Code2Spec) inverter trained by mean square error and adversarial loss.The VQ-VAE extracts the speech to a latent space, forces itself to map it into the nearest codebook and produces compressed representation.Next, the inverter generates a magnitude spectrogram to the target voice, given the codebook vectors from VQ-VAE.In our experiments, we also investigated several other clustering algorithms, including K-Means and GMM, and compared them with the VQ-VAE result on ABX scores and bit rates.Our proposed approach significantly improved the intelligibility (in CER), the MOS, and discrimination ABX scores compared to the official ZeroSpeech 2019 baseline or even the topline. Andros Tjandra, Berrak Sisman, Mingyang Zhang 0003, Sakriani Sakti, Haizhou Li 0001, Satoshi Nakamura 0001 |
INTERSPEECH | 5 |
| 2019 | Code-Switching Detection Using ASR-Generated Language PosteriorsabstractCode-switching (CS) detection refers to the automatic detection of language switches in code-mixed utterances.This task can be achieved by using a CS automatic speech recognition (ASR) system that can handle such language switches.In our previous work, we have investigated the code-switching detection performance of the Frisian-Dutch CS ASR system by using the time alignment of the most likely hypothesis and found that this technique suffers from over-switching due to numerous very short spurious language switches.In this paper, we propose a novel method for CS detection aiming to remedy this shortcoming by using the language posteriors which are the sum of the framelevel posteriors of phones belonging to the same language.The CS ASR-generated language posteriors contain more complete language-specific information on frame level compared to the time alignment of the ASR output.Hence, it is expected to yield more accurate and robust CS detection.The CS detection experiments demonstrate that the proposed language posterior-based approach provides higher detection accuracy than the baseline system in terms of equal error rate.Moreover, a detailed CS detection error analysis reveals that using language posteriors reduces the false alarms and results in more robust CS detection. Qinyi Wang, Emre Yilmaz 0001, Adem Derinel, Haizhou Li 0001 |
INTERSPEECH | 4 |
| 2019 | Robust Sound Recognition: A Neuromorphic Approach
Jibin Wu, Zihan Pan, Malu Zhang, Rohan Kumar Das, Yansong Chua, Haizhou Li 0001 |
INTERSPEECH | 6 |
| 2019 | Multi-Graph Decoding for Code-Switching ASRabstract\n Contains fulltext :\n 214606.pdf (Publisher’s version ) (Open Access)\n Emre Yilmaz 0001, Samuel Cohen, Xianghu Yue, David A. van Leeuwen, Haizhou Li 0001 |
INTERSPEECH | 5 |
| 2019 | Large-Scale Speaker Diarization of Radio Broadcast ArchivesabstractThis paper describes our initial efforts to build a large-scale speaker diarization (SD) and identification system on a recently digitized radio broadcast archive from the Netherlands which has more than 6500 audio tapes with 3000 hours of Frisian-Dutch speech recorded between 1950-2016. The employed large-scale diarization scheme involves two stages: (1) tape-level speaker diarization providing pseudo-speaker identities and (2) speaker linking to relate pseudo-speakers appearing in multiple tapes. Having access to the speaker models of several frequently appearing speakers from the previously collected FAME! speech corpus, we further perform speaker identification by linking these known speakers to the pseudo-speakers identified at the first stage. In this work, we present a recently created longitudinal and multilingual SD corpus designed for large-scale SD research and evaluate the performance of a new speaker linking system using x-vectors with PLDA to quantify cross-tape speaker similarity on this corpus. The performance of this speaker linking system is evaluated on a small subset of the archive which is manually annotated with speaker information. The speaker linking performance reported on this subset (53 hours) and the whole archive (3000 hours) is compared to quantify the impact of scaling up in the amount of speech data. Emre Yilmaz 0001, Adem Derinel, Kun Zhou 0003, Henk van den Heuvel, Niko Brümmer, Haizhou Li 0001, David A. van Leeuwen |
INTERSPEECH | 6 |
| 2019 | On the End-to-End Solution to Mandarin-English Code-Switching Speech RecognitionabstractCode-switching (CS) refers to a linguistic phenomenon where a speaker uses different languages in an utterance or between alternating utterances.In this work, we study end-to-end (E2E) approaches to the Mandarin-English code-switching speech recognition task.We first examine the effectiveness of using data augmentation and byte-pair encoding (BPE) subword units.More importantly, we propose a multitask learning recipe, where a language identification task is explicitly learned in addition to the E2E speech recognition task.Furthermore, we introduce an efficient word vocabulary expansion method for language modeling to alleviate data sparsity issues under the code-switching scenario.Experimental results on the SEAME data, a Mandarin-English code-switching corpus, demonstrate the effectiveness of the proposed methods. Zhiping Zeng, Yerbolat Khassanov, Van Tung Pham, Haihua Xu 0001, Chng Eng Siong, Haizhou Li 0001 |
INTERSPEECH | 6 |
| 2019 | Automatic evaluation of end-to-end dialog systems with adequacy-fluency metrics
Luis Fernando D'Haro, Rafael E. Banchs, Chiori Hori, Haizhou Li 0001 |
Comput. Speech Lang. | 4 |
| 2019 | Group Sparse Representation With WaveNet Vocoder Adaptation for Spectrum and Prosody ConversionabstractThe statistical approach to voice conversion typically consists of a feature conversion module followed by a vocoder. So far, the feature conversion studies are mainly focused on the conversion of spectrum. However, speaker identity is also characterized by prosodic features, such as fundamental frequency F0 and energy contour among others. In this paper, we study the transformation of speaker characteristics both in terms of spectrum and prosody. We propose two novel techniques that effectively use a limited amount of source-target training data and leverage a large general speech corpus to improve the voice conversion quality. First, we study the phonetic sparse representation under the group sparsity mathematical formulation. We use phonetic posteriorgrams PPGs together with spectral and prosody features to form tandem feature in the phonetic dictionary. The tandem feature allow us to estimate an activation matrix that is less dependent on source speakers, thus providing a better voice conversion quality. Second, we study the use of WaveNet vocoder that can be trained on general speech corpus from multiple speakers and adapted on target speaker data to improve the vocoding quality. We benefit from the large general speech databases that are used to train the PPG generator, and the WaveNet vocoder. The experiments show that the proposed conversion framework outperforms the traditional spectrum and prosody conversion techniques in both objective and subjective evaluations. Berrak Sisman, Mingyang Zhang 0003, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2019 | Spike Timing or Rate? Neurons Learn to Make Decisions for Both Through Threshold-Driven PlasticityabstractSpikes play an essential role in information transmission in central nervous system, but how neurons learn from them remains a challenging question. Most algorithms studied how to train spiking neurons to process patterns encoded with a sole assumption of either a rate or a temporal code. Is there a general learning algorithm capable of processing both codes regardless of the intense debate on them within neuroscience community? In this paper, we propose several threshold-driven plasticity algorithms to address the above question. In addition to formulating the algorithms, we also provide proofs with respect to several properties, such as robustness and convergence. The experimental results illustrate that our algorithms are simple, effective and yet efficient for training neurons to learn spike patterns. Due to their simplicity and high efficiency, our algorithms would be potentially beneficial for both software and hardware implementations. Neurons with our algorithms can also detect and recognize embedded features from a background sensory activity. With the as-proposed algorithms, a single neuron can successfully perform multicategory classifications by making decisions based on its output spike number in response to each category. Spike patterns being processed can be encoded with both spike rates and precise timings. When afferent spike timings matter, neurons will automatically extract temporal features without being explicitly instructed as to which point to fire. Qiang Yu 0005, Haizhou Li 0001, Kay Chen Tan |
IEEE Trans. Cybern. | 2 |
| 2019 | A Cost-Sensitive Deep Belief Network for Imbalanced ClassificationabstractImbalanced data with a skewed class distribution are common in many real-world applications. Deep Belief Network (DBN) is a machine learning technique that is effective in classification tasks. However, conventional DBN does not work well for imbalanced data classification because it assumes equal costs for each class. To deal with this problem, cost-sensitive approaches assign different misclassification costs for different classes without disrupting the true data sample distributions. However, due to lack of prior knowledge, the misclassification costs are usually unknown and hard to choose in practice. Moreover, it has not been well studied as to how cost-sensitive learning could improve DBN performance on imbalanced data problems. This paper proposes an evolutionary cost-sensitive deep belief network (ECS-DBN) for imbalanced classification. ECS-DBN uses adaptive differential evolution to optimize the misclassification costs based on the training data that presents an effective approach to incorporating the evaluation measure (i.e., G-mean) into the objective function. We first optimize the misclassification costs, and then apply them to DBN. Adaptive differential evolution optimization is implemented as the optimization algorithm that automatically updates its corresponding parameters without the need of prior domain knowledge. The experiments have shown that the proposed approach consistently outperforms the state of the art on both benchmark data sets and real-world data set for fault diagnosis in tool condition monitoring. Chong Zhang 0003, Kay Chen Tan, Haizhou Li 0001, Geok Soon Hong |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2018 | End-to-End Hierarchical Language Identification SystemabstractRecently, hierarchical language identification systems have shown significant improvement over single level systems in both closed and open set language identification tasks. However, developing such a system requires the features and classifier selection at each node in the hierarchical structure to be hand crafted. Motivated by the superior ability of end-to-end deep neural network architecture to jointly optimize the feature extraction and classification process, we propose a novel approach developing an end-to-end hierarchical language identification system. The proposed approach also demonstrates the in -built ability of the end-to-end hierarchical structure training that enables an out-of-set language model, without using any additional out-of-set language training data. Experiments are conducted on the NIST LRE 2015 data set. The overall results show relative improvements of 18.6% and 27.3% in terms of Cavgin closed and open set tasks over the corresponding baseline systems. Saad Irtza, Vidhyasaharan Sethu, Eliathamby Ambikairajah, Haizhou Li 0001 |
ICASSP | 4 |
| 2018 | On the Importance of Analytic Phase of Speech Signals in Spoken Language RecognitionabstractIn this paper, we study the role of long-time analytic phase of speech signals in spoken language recognition (SLR) and employ a set of features termed as instantaneous frequency cepstral coefficients (IFCC). We extract IFCC from long-time analytic phase, in an effort to capture long range acoustic features from speech signals. These features are used in combination with the traditional shifted delta cepstral coefficients (SDCC) for SLR. As the SDCC are extracted from spectral magnitude and IFCC are from analytic phase, they characterize long-time information of speech in different ways. The experiments conducted with NIST LRE 2017 task reveals the complementary effects of IFCC features to SDCC and deep bottleneck (DBN) features. The fusion of IFCC with SDCC/DBN features delivered relative improvements of 23.23% and 16.78% in average equal error rate over the SDCC and DBN features, respectively, indicating the benefits of information from analytic phase in SLR. Karthika Vijayan, Haizhou Li 0001, Hanwu Sun, Kong-Aik Lee |
ICASSP | 2 |
| 2018 | Unsupervised Domain Adaptation via Domain Adversarial Training for Speaker RecognitionabstractThe i-vector approach to speaker recognition has achieved good performance when the domain of the evaluation dataset is similar to that of the training dataset. However, in realworld applications, there is always a mismatch between the training and evaluation datasets, that leads to performance degradation. To address this problem, this paper proposes to learn the domain-invariant and speaker-discriminative speech representations via domain adversarial training. Specifically, with domain adversarial training method, we use a gradient reversal layer to remove the domain variation and project the different domain data into the same subspace. Moreover, we compare the proposed method with other state-of-the-art unsupervised domain adaptation techniques for i-vector approach to speaker recognition (e.g. autoencoder based domain adaptation, inter dataset variability compensation, dataset-invariant covariance normalization, and so on). Experiments on 2013 domain adaptation challenge (DAC) dataset demonstrate that the proposed method is not only effective in solving the dataset mismatch problem, but also outperforms the compared unsupervised domain adaptation methods. Qing Wang 0039, Wei Rao 0002, Sining Sun, Lei Xie 0001, Chng Eng Siong, Haizhou Li 0001 |
ICASSP | 6 |
| 2018 | Single Channel Speech Separation with Constrained Utterance Level Permutation Invariant Training Using Grid LSTMabstractUtterance level permutation invariant training (uPIT) technique is a state-of-the-art deep learning architecture for speaker independent multi-talker separation. uPIT solves the label ambiguity problem by minimizing the mean square error (MSE) over all permutations between outputs and targets. However, uPIT may be sub-optimal at segmental level because the optimization is not calculated over the individual frames. In this paper, we propose a constrained uPIT (cuPIT) to solve this problem by computing a weighted MSE loss using dynamic information (i.e., delta and acceleration). The weighted loss ensures the temporal continuity of output frames with the same speaker. Inspired by the heuristics (i.e., vocal tract continuity) in computational auditory scene analysis, we then extend the model by adding a Grid LSTM layer, that we name it as cuPIT-Grid LSTM, to automatically learn both temporal and spectral patterns over the input magnitude spectrum simultaneously. The experimental results show 9.6% and 8.5% relative improvements on WSJ0-2mix dataset under both closed and open conditions comparing with the uPIT baseline. Chenglin Xu, Wei Rao 0002, Chng Eng Siong, Haizhou Li 0001 |
ICASSP | 5 |
| 2018 | An Event-Based Cochlear Filter Temporal Encoding Scheme for Speech SignalsabstractSpiking Neural Network (SNN), the third generation of neural networks, has been shown to perform well in pattern recognition tasks involving temporal information, such as speech recognition and motion detection. However, most neural networks, including the SNN, for speech recognition rely on short-time frequency analysis, such as the mel-frequency cepstral coefficients (MFCC), for low-level feature extraction. MFCC feature extraction works by analyzing a window of time signal in multiple frequency bands one window at a time, in a synchronous fashion. This is in contrast to the event-based principle of SNN, whereby electrical impulses are emitted and processed in an asynchronous fashion. Just as speech signals arrive at the human's cochlear filterbank concurrently, but spikes encoding the power in each frequency band are emitted asynchronously, we propose an event-based cochlear filter encoding scheme, whereby the power in each frequency band is directly extracted in the time domain and spikes encoded using the latency code are emitted asynchronously to represent the power of each frequency band. This replaces the traditional MFCC frontend used in most speech recognition models, and makes possible an end-to- end event-based SNN implementation for a speech recognition task. The proposed event-based neural encoding is not only biologically plausible, but also outperforms the MFCC as an encoding frontend for an SNN classifier in a speech recognition task, in terms of higher classification accuracy and lower latency. Such an end-to-end SNN model could be implemented on a neuromorphic chip to fully realize the advantages of event-based processing. Zihan Pan, Haizhou Li 0001, Jibin Wu, Yansong Chua |
IJCNN | 2 |
| 2018 | A Biologically Plausible Speech Recognition Framework Based on Spiking Neural NetworksabstractHumans perform remarkably well for speech recognition using sparse and asynchronous events carried by electrical impulses. Motivated by the observations that human brains primarily learn features from environmental stimuli in an unsupervised manner and consume extremely low power for complex cognitive tasks, we propose a biologically plausible speech recognition mechanism using unsupervised self-organizing map (SOM) for feature representation and event-driven spiking neural network (SNN) for spatiotemporal pattern classification. Moreover, we improve the biological realism of the proposed framework by using mel-scaled filter bank as the front-end, so as to mimic the human auditory system. Our experiments on the TIDIGITS dataset achieve speech recognition accuracy surpassing those of other bio-inspired systems. The proposed SOM-SNN framework can be implemented using the artificial silicon cochlear and neuromorphic processor, so as to fully exploit the potential of event-based speech recognition system. Jibin Wu, Yansong Chua, Haizhou Li 0001 |
IJCNN | 3 |
| 2018 | Automatic Pronunciation Evaluation of Singing
Chitralekha Gupta, Haizhou Li 0001, Ye Wang 0007 |
INTERSPEECH | 2 |
| 2018 | Wavelet Analysis of Speaker Dependent and Independent Prosody for Voice Conversion
Berrak Sisman, Haizhou Li 0001 |
INTERSPEECH | 2 |
| 2018 | A Voice Conversion Framework with Tandem Feature Sparse Representation and Speaker-Adapted WaveNet Vocoder
Berrak Sisman, Mingyang Zhang 0003, Haizhou Li 0001 |
INTERSPEECH | 3 |
| 2018 | Co-whitening of I-vectors for Short and Long Duration Speaker Verification
Longting Xu, Kong-Aik Lee, Haizhou Li 0001, Zhen Yang 0001 |
INTERSPEECH | 3 |
| 2018 | Mandarin-English Code-switching Speech Recognition
Haihua Xu 0001, Van Tung Pham, Kyaw Zin Tun, Zhi Hao Lim, Chng Eng Siong, Haizhou Li 0001 |
INTERSPEECH | 6 |
| 2018 | A Shifted Delta Coefficient Objective for Monaural Speech Separation Using Multi-task Learning
Chenglin Xu, Wei Rao 0002, Chng Eng Siong, Haizhou Li 0001 |
INTERSPEECH | 4 |
| 2018 | Learning Acoustic Word Embeddings with Temporal Context for Query-by-Example Speech SearchabstractWe propose to learn acoustic word embeddings with temporal context for query-by-example (QbE) speech search.The temporal context includes the leading and trailing word sequences of a word.We assume that there exist spoken word pairs in the training database.We pad the word pairs with their original temporal context to form fixed-length speech segment pairs.We obtain the acoustic word embeddings through a deep convolutional neural network (CNN) which is trained on the speech segment pairs with a triplet loss.By shifting a fixed-length analysis window through the search content, we obtain a running sequence of embeddings.In this way, searching for the spoken query is equivalent to the matching of acoustic word embeddings.The experiments show that our proposed acoustic word embeddings learned with temporal context are effective in QbE speech search.They outperform the state-of-the-art frame-level feature representations and reduce run-time computation since no dynamic time warping is required in QbE speech search.We also find that it is important to have sufficient speech segment pairs to train the deep CNN for effective acoustic word embeddings. Yougen Yuan, Cheung-Chi Leung, Lei Xie 0001, Hongjie Chen 0001, Bin Ma 0001, Haizhou Li 0001 |
INTERSPEECH | 6 |
| 2018 | Adaptive Wavenet Vocoder for Residual Compensation in GAN-Based Voice ConversionabstractIn this paper, we propose to use generative adversarial networks (GAN) together with a WaveNet vocoder to address the over-smoothing problem arising from the deep learning approaches to voice conversion, and to improve the vocoding quality over the traditional vocoders. As GAN aims to minimize the divergence between the natural and converted speech parameters, it effectively alleviates the over-smoothing problem in the converted speech. On the other hand, WaveNet vocoder allows us to leverage from the human speech of a large speaker population, thus improving the naturalness of the synthetic voice. Furthermore, for the first time, we study how to use WaveNet vocoder for residual compensation to improve the voice conversion performance. The experiments show that the proposed voice conversion framework consistently outperforms the baselines. Berrak Sisman, Mingyang Zhang 0003, Sakriani Sakti, Haizhou Li 0001, Satoshi Nakamura 0001 |
SLT | 4 |
| 2018 | Generative X-Vectors for Text-Independent Speaker VerificationabstractSpeaker verification (SV) systems using deep neural network embeddings, so-called the x-vector systems, are becoming popular due to its good performance superior to the i-vector systems. The fusion of these systems provides improved performance benefiting both from the discriminatively trained x-vectors and generative i-vectors capturing distinct speaker characteristics. In this paper, we propose a novel method to include the complementary information of i-vector and x-vector, that is called generative x-vector. The generative x-vector utilizes a transformation model learned from the i-vector and x-vector representations of the background data. Canonical correlation analysis is applied to derive this transformation model, which is later used to transform the standard x-vectors of the enrollment and test segments to the corresponding generative x-vectors. The SV experiments performed on the NIST SRE 2010 dataset demonstrate that the system using generative x-vectors provides considerably better performance than the baseline i-vector and x-vector systems. Furthermore, the generative x-vectors outperform the fusion of i-vector and x-vector systems for long-duration utterances, while yielding comparable results for short-duration utterances. Longting Xu, Rohan Kumar Das, Emre Yilmaz 0001, Haizhou Li 0001 |
SLT | 5 |
| 2018 | Using language cluster models in hierarchical language identification
Saad Irtza, Vidhyasaharan Sethu, Eliathamby Ambikairajah, Haizhou Li 0001 |
Speech Commun. | 4 |
| 2018 | Re-ranking spoken term detection with acoustic exemplars of keywords
Van Tung Pham, Haihua Xu 0001, Nancy F. Chen, Chng Eng Siong, Haizhou Li 0001 |
Speech Commun. | 6 |
| 2018 | Farewell Editorial
Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2018 | Generalizing I-Vector Estimation for Rapid Speaker RecognitionabstractAn i-vector is a compact representation that captures both the speaker and session variabilities rendered in a spoken utterance. Over the past years, it has prevailed over other techniques and is now the de facto representation for text-independent speaker recognition. Standard i-vector extraction requires intense computation at run-time. Reducing the computation will allow effective use of i-vector in more applications. Such intense computation arises from the posterior covariance matrix, when estimating the i-vector. There have been studies on how to simplify the computation of posterior covariance matrix with modest success. In this paper, we propose a novel approach to i-vector extraction without the need to evaluate the full posterior covariance thereby speeding up the run-time extraction process. This is achieved by generalizing the i-vector estimation in two ways. First, we introduce the use of occupancy reweighting in conjunction with whitening over the Baum-Welch statistics as part of the preprocessing step. Second, we introduce the so-called subspace-orthogonalizing prior (SOP) to replace the standard Gaussian prior in i-vector formulation. Experiments conducted on the extended-core task of NIST SRE'10 show that the proposed rapid SOP approach achieves considerable speed-up over the standard i-vector with comparable equal error rates. Longting Xu, Kong-Aik Lee, Haizhou Li 0001, Zhen Yang 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2017 | Multilingual bottle-neck feature learning from untranscribed speechabstractWe propose to learn a low-dimensional feature representation for multiple languages without access to their manual transcription. The multilingual features are extracted from a shared bottleneck layer of a multi-task learning deep neural network which is trained using un-supervised phoneme-like labels. The unsupervised phoneme-like labels are obtained from language-dependent Dirichlet process Gaussian mixture models (DPGMMs). Vocal tract length normalization (VTLN) is applied to mel-frequency cepstral coefficients to reduce talker variation when DPGMMs are trained. The proposed features are evaluated using the ABX phoneme discriminability test in the Zero Resource Speech Challenge 2017. In the experiments, we show that the proposed features perform well across different languages, and they consistently outperform our previously proposed DPGMM posteriorgrams which topped the performance in the same challenge in 2015. Hongjie Chen 0001, Cheung-Chi Leung, Lei Xie 0001, Bin Ma 0001, Haizhou Li 0001 |
ASRU | 5 |
| 2017 | Sparse representation of phonetic features for voice conversion with and without parallel dataabstractThis paper presents a voice conversion framework that uses phonetic information in an exemplar-based voice conversion approach. The proposed idea is motivated by the fact that phone-dependent exemplars lead to better estimation of activation matrix, therefore, possibly better conversion. We propose to use the phone segmentation results from automatic speech recognition (ASR) to construct a sub-dictionary for each phone. The proposed framework can work with or without parallel training data. With parallel training data, we found that phonetic sub-dictionary outperforms the state-of-the-art baseline in objective and subjective evaluations. Without parallel training data, we use Phonetic PosteriorGrams (PPGs) as the speaker-independent exemplars in the phonetic sub-dictionary to serve as a bridge between speakers. We report that such technique achieves a competitive performance without the need of parallel training data. Berrak Sisman, Haizhou Li 0001, Kay Chen Tan |
ASRU | 2 |
| 2017 | Statistical parametric speech synthesis using generative adversarial networks under a multi-task learning frameworkabstractIn this paper, we aim at improving the performance of synthesized speech in statistical parametric speech synthesis (SPSS) based on a generative adversarial network (GAN). In particular, we propose a novel architecture combining the traditional acoustic loss function and the GAN's discriminative loss under a multi-task learning (MTL) framework. The mean squared error (MSE) is usually used to estimate the parameters of deep neural networks, which only considers the numerical difference between the raw audio and the synthesized one. To mitigate this problem, we introduce the GAN as a second task to determine if the input is a natural speech with specific conditions. In this MTL framework, the MSE optimization improves the stability of GAN, and at the same time GAN produces samples with a distribution closer to natural speech. Listening tests show that the multi-task architecture can generate more natural speech that satisfies human perception than the conventional methods. Shan Yang 0001, Lei Xie 0001, Xiaoyan Lou, Dong-Yan Huang, Haizhou Li 0001 |
ASRU | 7 |
| 2017 | Extracting bottleneck features and word-like pairs from untranscribed speech for feature representationabstractWe propose a framework to learn a frame-level speech representation in a scenario where no manual transcription is available. Our framework is based on pairwise learning using bottleneck features (BNFs). Initial frame-level features are extracted from a bottleneck-shaped multilingual deep neural network (DNN) which is trained with unsupervised phoneme-like labels. Word-like pairs are discovered in the untranscribed speech using the initial features, and frame alignment is performed on each word-like speech pair. The matching frame pairs are used as input-output to train another DNN with the mean square error (MSE) loss function. The final frame-level features are extracted from an internal hidden layer of MSE-based DNN. Our pairwise learned feature representation is evaluated on the ZeroSpeech 2017 challenge. The experiments show that pairwise learning improves phoneme discrimination in 10s and 120s test conditions. We find that it is important to use BNFs as initial features when pairwise learning is performed. With more word pairs obtained from the Switchboard corpus and its manual transcription, the phoneme discrimination of three languages in the evaluation data can further be improved despite data mismatch. Yougen Yuan, Cheung-Chi Leung, Lei Xie 0001, Hongjie Chen 0001, Bin Ma 0001, Haizhou Li 0001 |
ASRU | 6 |
| 2017 | A data-driven prognostics framework for tool remaining useful life estimation in tool condition monitoringabstractTool Condition Monitoring (TCM) is an important topic in manufacturing industry, which improves product quality, production efficiency, reduces costs and downtime. This paper develops a new data-driven framework for estimating tool remaining useful life (RUL) in TCM. The framework includes the following modular components: data preprocessing with a proposed adaptive Baysian change point detection (ABCPD) for automatic data alignment, time window process, feature extraction, feature selection and a multi-layer neural network as the main machine learning algorithm. The proposed framework is evaluated on a real-world gun drilling experimental dataset with multiple sensor measurements (i.e. thrust force, torque, 12 vibration signals). Different model selection, sensor selection, feature selection methods have been investigated in this paper. The simulation performance of the proposed framework is studied with the gun drilling dataset and it has been shown that the proposed framework has good performance. Chong Zhang 0003, Geok Soon Hong, Kay Chen Tan, Junhong Zhou, Hian-Leng Chan, Haizhou Li 0001 |
ETFA | 7 |
| 2017 | Adaptation of PLDA for multi-source text-independent speaker verificationabstractProbabilistic linear discriminant analysis (PLDA) is widely described as an effective model for text-independent speaker verification in the i-vector space. The PLDA scoring function is typically formulated as the likelihood ratio between the speaker-adapted and the universal PLDAs. In this case, the adaptation of PLDA was performed through the speaker factors. In this paper, we show that the channel factors of the PLDA could be equivalently exploited to deal with the multi-source conditions. In speaker verification, with the proposed method, a PLDAmodel trained on conversational telephone speech could be adequately adapted for interview-style microphone recordings. Experimental results on NIST SRE'08 and SRE'10 datasets confirm that the proposed method is effective, especially for the case whereby enrollment and test utterances were captured from different sources. Kong-Aik Lee, Bin Ma 0001, Haizhou Li 0001, Li-Rong Dai 0001 |
ICASSP | 5 |
| 2017 | On time-frequency mask estimation for MVDR beamforming with application in robust speech recognitionabstractAcoustic beamforming has played a key role in the robust automatic speech recognition (ASR) applications. Accurate estimates of the speech and noise spatial covariance matrices (SCM) are crucial for successfully applying the minimum variance distortionless response (MVDR) beamforming. Reliable estimation of time-frequency (TF) masks can improve the estimation of the SCMs and significantly improve the performance of the MVDR beamforming in ASR tasks. In this paper, we focus on the TF mask estimation using recurrent neural networks (RNN). Specifically, our methods include training the RNN to estimate the speech and noise masks independently, training the RNN to minimize the ASR cost function directly, and performing multiple passes to iteratively improve the mask estimation. The proposed methods are evaluated individually and overally on the CHiME-4 challenge. The results show that the proposed methods improve the ASR performance individually and also work complementarily. The overall performance achieves a word error rate of 8.9% with 6-microphone configuration, which is much better than 12.0% achieved with the state-of-the-art MVDR implementation. Shengkui Zhao, Douglas L. Jones, Chng Eng Siong, Haizhou Li 0001 |
ICASSP | 5 |
| 2017 | Pairwise learning using multi-lingual bottleneck features for low-resource query-by-example spoken term detectionabstractWe propose to use a feature representation obtained by pairwise learning in a low-resource language for query-by-example spoken term detection (QbE-STD). We assume that word pairs identified by humans are available in the low-resource target language. The word pairs are parameterized by a multi-lingual bottleneck feature (BNF) extractor that is trained using transcribed data in high-resource languages. The multi-lingual BNFs of the word pairs are used as an initial feature representation to train an autoencoder (AE). We extract features from an internal hidden layer of the pairwise trained AE to perform acoustic pattern matching for QbE-STD. Our experiments on the TIMIT and Switchboard corpora show that the pairwise learning brings 7.61% and 8.75% relative improvements in mean average precision (MAP) respectively over the initial feature representation. Yougen Yuan, Cheung-Chi Leung, Lei Xie 0001, Hongjie Chen 0001, Bin Ma 0001, Haizhou Li 0001 |
ICASSP | 6 |
| 2017 | Multimodal Prediction of Affective Dimensions via Fusing Multiple Regression Techniques
Dong-Yan Huang, Wan Ding, Huaiping Ming, Minghui Dong, Xinguo Yu, Haizhou Li 0001 |
INTERSPEECH | 7 |
| 2017 | Investigating Scalability in Hierarchical Language Identification System
Saad Irtza, Vidhyasaharan Sethu, Eliathamby Ambikairajah, Haizhou Li 0001 |
INTERSPEECH | 4 |
| 2017 | The I4U Mega Fusion and Collaboration for NIST Speaker Recognition Evaluation 2016abstract18th Annual Conference of the International Speech Communication Association, INTERSPEECH 2017, Stockholm, Sweden, 20-24 August 2017 Kong-Aik Lee, Ville Hautamäki, Tomi Kinnunen, Anthony Larcher, Andreas Nautsch, Themos Stafylakis, Gang Liu 0001, Mickael Rouvier, Wei Rao 0002, Federico Alegre, Man-Wai Mak, Achintya Kumar Sarkar, Héctor Delgado, Rahim Saeidi, Hagai Aronowitz, Aleksandr Sizov, Hanwu Sun, Trung Hieu Nguyen 0001, Guangsen Wang, Bin Ma 0001, Ville Vestman, Md. Sahidullah, M. Halonen, Anssi Kanervisto, Gaël Le Lan, Fahimeh Bahmaninezhad, Sergey Isadskiy, Christian Rathgeb, Christoph Busch 0001, Georgios Tzimiropoulos, Q. Qian, Q. Zhao, J. Xue, R. Jin, T. Zhao, Pierre-Michel Bousquet, Moez Ajili, Waad Ben Kheder, Driss Matrouf, Zhi Hao Lim, Chenglin Xu, Haihua Xu 0001, Chng Eng Siong, Benoit G. B. Fauve, Kaavya Sriskandaraja, Vidhyasaharan Sethu, W. W. Lin, Dennis Alexander Lehmann Thomsen, Zheng-Hua Tan, Massimiliano Todisco, Nicholas W. D. Evans, Haizhou Li 0001, John H. L. Hansen, Jean-François Bonastre, Eliathamby Ambikairajah |
INTERSPEECH | 59 |
| 2017 | Gain Compensation for Fast i-Vector Extraction Over Short Duration
Kong-Aik Lee, Haizhou Li 0001 |
INTERSPEECH | 2 |
| 2017 | ISCA Medal for Scientific Achievement
Haizhou Li 0001 |
INTERSPEECH | 1 |
| 2017 | Denoising Recurrent Neural Network for Deep Bidirectional LSTM Based Voice Conversion
Jie Wu 0017, Dong-Yan Huang, Lei Xie 0001, Haizhou Li 0001 |
INTERSPEECH | 4 |
| 2017 | Weighted Spatial Covariance Matrix Estimation for MUSIC Based TDOA Estimation of Speech Source
Chenglin Xu, Sining Sun, Wei Rao 0002, Chng Eng Siong, Haizhou Li 0001 |
INTERSPEECH | 6 |
| 2017 | Modeling Latent Topics and Temporal Distance for Story Segmentation of Broadcast NewsabstractThis paper studies a strategy to model latent topics and temporal distance of text blocks for story segmentation, that we call graph regularization in topic modeling or GRTM. We propose two novel approaches that consider both temporal distance and lexical similarity of text blocks, collectively referred to as data proximity, in learning latent topic representation, where a graph regularizer is involved to derive the latent topic representation while preserving data proximity. In the first approach, we extend the idea of Laplacian probabilistic latent semantic analysis (LapPLSA) by introducing a distance penalty function in the affinity matrix of a graph for latent topic estimation. The estimated latent topic distributions are used to replace the traditional term-frequency vectors as the data representation of the text blocks and to measure the cohesive strength between them. In the second approach, we perform Laplacian eigenmaps, which makes use of the graph regularizer for dimensionality reduction, on latent topic distributions estimated by conventional topic modeling. We conduct the experiments on the automatic speech recognition transcripts of the TDT2 English broadcast news corpus. The experiments show the proposed strategy outperforms the conventional techniques. LapPLSA performs the best with the highest F1-measure of 0.816. The effects of the penalty constant in the distance penalty function, the number of latent topics, and the size of training data on the segmentation performances are also studied. Hongjie Chen 0001, Lei Xie 0001, Cheung-Chi Leung, Xiaoming Lu, Bin Ma 0001, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2017 | An Exemplar-Based Approach to Frequency Warping for Voice ConversionabstractThe voice conversion's task is to modify a source speaker's voice to sound like that of a target speaker. A conversion method is considered successful when the produced speech sounds natural and similar to the target speaker. This paper presents a new voice conversion framework in which we combine frequency warping and exemplar-based method for voice conversion. Our method maintains high-resolution details during conversion by directly applying frequency warping on the high-resolution spectrum to represent the target. The warping function is generated by a sparse interpolation from a dictionary of exemplar warping functions. As the generated warping function is dependent only on a very small set of exemplars, we do away with the statistical averaging effects inherited from Gaussian mixture models. To compensate for the conversion error, we also apply residual exemplars into the conversion process. Both objective and subjective evaluations on the VOICES database validated the effectiveness of the proposed voice conversion framework. We observed a significant improvement in speech quality over the state-of-the-art parametric methods. Xiaohai Tian, Siu Wa Lee, Zhizheng Wu 0001, Chng Eng Siong, Haizhou Li 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2016 | Exploring Convolutional and Recurrent Neural Networks in Sequential Labelling for Dialogue Topic Tracking
Seokhwan Kim, Rafael E. Banchs, Haizhou Li 0001 |
ACL (1) | 3 |
| 2016 | Content-aware local variability vector for speaker verification with short utteranceabstractI-vector has shown to be very effective in speaker verification with long-duration speech utterances. But when test utterances are of short duration, content mismatch between the enrollment and test utterances limit the performance of i-vector system. This paper proposes to extract local session variability vectors on different phonetic classes from the utterances instead of estimating the session variability across the whole utterance as i-vector does. Using the posteriors given by a deep neural network (DNN) trained for phone state classification, the local vectors represent the session variability contained in specific phonetic content. Our experiments show that the content-aware local vectors are better at coping with the content mismatch between training and test utterances of short durations for text-independent, text-constrained and text-dependent tasks. Kong-Aik Lee, Chng Eng Siong, Bin Ma 0001, Haizhou Li 0001, Li-Rong Dai 0001 |
ICASSP | 5 |
| 2016 | Exemplar-inspired strategies for low-resource spoken keyword search in SwahiliabstractWe present exemplar-inspired low-resource spoken keyword search strategies for acoustic modeling, keyword verification, and system combination. This state-of-the-art system was developed by the SINGA team in the context of the 2015 NIST Open Keyword Search Evaluation (OpenKWS15) using conversational Swahili provided by the IARPA Babel program. In this work, we elaborate on the following: (1) exploiting exemplar training samples to construct a non-parametric acoustic model using kernel density estimation at test time; (2) rescoring hypothesized keyword detections through quantifying their acoustic similarity with exemplar training samples; (3 ) extending our previously proposed system combination approach to incorporate prosody features of exemplar keyword samples. Nancy F. Chen, Van Tung Pham, Haihua Xu 0001, Van Hai Do, Chongjia Ni, I-Fan Chen, Sunil Sivadas, Chin-Hui Lee 0001, Chng Eng Siong, Bin Ma 0001, Haizhou Li 0001 |
ICASSP | 12 |
| 2016 | Combining multiple kernel models for automatic intelligibility detection of pathological speechabstractAutomatic detection of pathological voice is a challenging task in speech processing. Appropriate acoustic cues of voice can be used to differentiate between normal voices and pathological voices. We propose a method to represent each speech utterance using three types of speech signal representations (i.e., cross-correlation matrix, Gaussian distribution and linear subspace) respectively. Various kernels were applied to these representations for measuring resemblance and difference. Four classifiers, i.e., KNN, kernel partial least squares, kernel SVM, and logistic regression, are studied for comparing their performance of classification. Finally, a simple fusion of learning classifiers from different acoustic representations was carried out at the score decision level for enhancing the performance. The different classifiers were evaluated on the Interspeech 2012 challenge development data set and test data set. Their effects in a fusion scheme are studied. The accuracy of the fusion system attained 78.0 % on test set, with an improved gain of 9.1 % over the challenge baseline 68.9 %. Dong-Yan Huang, Minghui Dong, Haizhou Li 0001 |
ICASSP | 3 |
| 2016 | A hierarchical framework for language identificationabstractMost current language recognition systems model different levels of information such as acoustic, prosodic, phonotactic, etc. independently and combine the model likelihoods in order to make a decision. However, these are single level systems that treat all languages identically and hence incapable of exploiting any similarities that may exist within groups of languages. In this paper, a hierarchical language identification (HLID) framework is proposed that involves a series of classification decisions at multiple levels involving language clusters of decreasing sizes with individual languages identified only at the final level. The performance of proposed hierarchical framework is compared with a state-of-the-art LID system on the NIST 2007 database and the results indicate that the proposed approach outperforms state-of-the-art systems. Saad Irtza, Vidhyasaharan Sethu, Haris Bavattichalil, Eliathamby Ambikairajah, Haizhou Li 0001 |
ICASSP | 5 |
| 2016 | Exemplar-based sparse representation of timbre and prosody for voice conversionabstractVoice conversion (VC) aims to make one speaker (source) to sound like spoken by another speaker (target) without changing the language content. Most of the state-of-the-art voice conversion systems focus only on timbre conversion. However, the speaker identity is characterized by the source-related cues such as fundamental frequency and energy as well. In this work, we propose an exemplarbased sparse representation of timbre and prosody for voice conversion that does not necessitate separately timbre conversion and prosody conversions. The experiment results show that, in addition to the conversion of spectral features, the proper conversion of prosody features will improve the quality and speaker identity of the converted speech. Huaiping Ming, Dong-Yan Huang, Lei Xie 0001, Shaofei Zhang, Minghui Dong, Haizhou Li 0001 |
ICASSP | 6 |
| 2016 | Cross-lingual deep neural network based submodular unbiased data selection for low-resource keyword searchabstractIn this paper, we propose a cross-lingual deep neural network (DNN) based submodular unbiased data selection approach for low-resource keyword search (KWS). A small amount (e.g. one hour) of transcribed data is used to conduct cross-lingual transfer. The frame-level senone sequence activated by the cross-lingual DNN is used to represent each untranscribed speech utterance. The proposed submodular function considers utterance length normalization and the feature distribution matched to a development set. Experiments are conducted by selecting 9 hours of Tamil speech for the 2014 NIST Open Keyword Search Evaluation (OpenKWS14). The proposed data selection approach provides 35.8% relative actual term weighted value (ATWV) improvement over random selection on the OpenKWS14 Evalpartl data set. Further analysis of the experimental results shows that both utterance length normalization and the feature distribution estimated from a development set deployed in the submodular function can suppress the preference to select long utterances. The selected utterances can cover a more diverse range of tri-phones, words, and acoustic variations from a wider set of utterances. Moreover, the wider coverage of words also benefits the acquired linguistic knowledge, which also contributes to improving KWS performance. Chongjia Ni, Cheung-Chi Leung, Lei Wang 0020, Feng Rao, Nancy F. Chen, Bin Ma 0001, Haizhou Li 0001 |
ICASSP | 9 |
| 2016 | Keyword search using query expansion for graph-based rescoring of hypothesized detectionsabstractIn this work, we propose a novel framework for rescoring keyword search (KWS) detections using acoustic samples extracted from the training data. We view the keyword rescoring task as an information retrieval task and adopt the idea of query expansion. We expand a textual keyword with multiple speech keyword samples extracted from the training data. In this way, the hypothesized detections are compared with the multiple keywords using non-parametric approaches such as dynamic time warping (DTW). The obtained similarity scores are used in a graph based method to re-rank the original confidence scores estimated by the automatic speech recognition (ASR) systems. Experimental results on the NIST OpenKWS15 Evaluation show that our rescoring method is effective, especially for the subword system. For subword experiments, the graph-based rescoring with training samples obtains 5.1% and 1.5% absolute improvement over two baseline systems. One is a standard parametric ASR system, while the other is the graph-based rescoring without training samples. Van Tung Pham, Haihua Xu 0001, Nancy F. Chen, Chng Eng Siong, Haizhou Li 0001 |
ICASSP | 6 |
| 2016 | Spoofing detection from a feature representation perspectiveabstractSpoofing detection, which discriminates the spoofed speech from the natural speech, has gained much attention recently. Low-dimensional features that are used in speaker recognition/verification are also used in spoofing detection. Unfortunately, they don't capture sufficient information required for spoofing detection. In this work, we investigate the use of high-dimensional features for spoofing detection, that maybe more sensitive to the artifacts in the spoofed speech. Six types of high-dimensional feature are employed. For each kind of feature, four different representations are extracted, i.e. the original high-dimensional feature, corresponding low-dimensional feature, the low- and the high-frequency regions of the original high-dimensional feature. Dynamic features are also calculated to assess the effectiveness of the temporal information to detect the artifacts across frames. A neural network-based classifier is adopted to handle the high-dimensional features. Experimental results on the standard ASVspoof 2015 corpus suggest that high-dimensional features and dynamic features are useful for spoofing attack detection. A fusion of them has been shown to achieve 0.0% the equal error rates for nine of ten attack types. Xiaohai Tian, Zhizheng Wu 0001, Chng Eng Siong, Haizhou Li 0001 |
ICASSP | 5 |
| 2016 | An expectation-maximization eigenvector clustering approach to direction of arrival estimation of multiple speech sourcesabstractThis paper presents an eigenvector clustering approach for estimating the direction of arrival (DOA) of multiple speech signals using a microphone array. Existing clustering approaches usually only use low frequencies to avoid spatial aliasing. In this study, we propose a probabilistic eigenvector clustering approach to use all frequencies. In our work, time-frequency (TF) bins dominated by only one source are first detected using a combination of noise-floor tracking, onset detection and coherence test. For each selected TF bin, the largest eigenvector of its spatial covariance matrix is extracted for clustering. A mixture density model is introduced to model the distribution of the eigenvectors, where each component distribution corresponds to one source and is parameterized by the source DOA. To use eigenvectors of all frequencies, the steering vectors of all frequencies of the sources are used in the distribution function. The DOAs of the sources can be estimated by maximizing the likelihood of the eigenvectors using an expectation-maximization (EM) algorithm. Simulation and experimental results show that the proposed approach significantly improves the root-mean-square error (RMSE) for DOA estimation of multiple speech sources compared to the MUSIC algorithm implemented on the single-source dominated TF bins and our previous clustering approach. Shengkui Zhao, Thi Ngoc Tho Nguyen, Douglas L. Jones, Chng Eng Siong, Haizhou Li 0001 |
ICASSP | 6 |
| 2016 | Approximate search of audio queries by using DTW with phone time boundary and data augmentationabstractDynamic Time Warping (DTW) is widely used in language independent query-by-example (QbE) spoken term detection (STD) tasks due to its high performance. However, there are two limitations of DTW based template matching, 1) it is not straightforward to perform approximate match of audio queries; 2) DTW is sensitive to the mismatch of signal conditions between the query and the speech search data. To allow approximate search, we propose a partial template matching strategy using phone time boundary information generated by a phone recognizer. To have more invariant representation of audio signals, we use bottleneck features (BNF) as the input of DTW. The BNF network is trained from augmented data, which is generated by adding reverberation and additive noises to the clean training data. Experimental results on QUESST 2015 task shows the effectiveness of the proposed methods for QbE-STD when the queries and search data are both distorted by reverberation and noises. Haihua Xu 0001, Jingyong Hou, Van Tung Pham, Cheung-Chi Leung, Lei Wang 0020, Van Hai Do, Hang Lv 0001, Lei Xie 0001, Bin Ma 0001, Chng Eng Siong, Haizhou Li 0001 |
ICASSP | 12 |
| 2016 | Audio and face video emotion recognition in the wild using deep neural networks and small datasetsabstractThis paper presents the techniques used in our contribution to Emotion Recognition in the Wild 2016’s video based sub-challenge. The purpose of the sub-challenge is to classify the six basic emotions (angry, sad, happy, surprise, fear & disgust) and neutral. Compared to earlier years’ movie based datasets, this year’s test dataset introduced reality TV videos containing more spontaneous emotion. Our proposed solution is the fusion of facial expression recognition and audio emotion recognition subsystems at score level. For facial emotion recognition, starting from a network pre-trained on ImageNet training data, a deep Convolutional Neural Network is fine-tuned on FER2013 training data for feature extraction. The classifiers, i.e., kernel SVM, logistic regression and partial least squares are studied for comparison. An optimal fusion of classifiers learned from different kernels is carried out at the score level to improve system performance. For audio emotion recognition, a deep Long Short-Term Memory Recurrent Neural Network (LSTM-RNN) is trained directly using the challenge dataset. Experimental results show that both subsystems individually and as a whole can achieve state-of-the art performance. The overall accuracy of the proposed approach on the challenge test dataset is 53.9%, which is better than the challenge baseline of 40.47% . Wan Ding, Dong-Yan Huang, Weisi Lin, Minghui Dong, Xinguo Yu, Haizhou Li 0001 |
ICMI | 7 |
| 2016 | SERAPHIM: A Wavetable Synthesis System with 3D Lip Animation for Real-Time Speech and Singing Applications on Mobile Platforms
Paul Y. Chan, Minghui Dong, Grace Xue Hui Ho, Haizhou Li 0001 |
INTERSPEECH | 4 |
| 2016 | SERAPHIM Live! - Singing Synthesis for the Performer, the Composer, and the 3D Game Developer
Paul Y. Chan, Minghui Dong, Grace Xue Hui Ho, Haizhou Li 0001 |
INTERSPEECH | 4 |
| 2016 | Unsupervised Bottleneck Features for Low-Resource Query-by-Example Spoken Term Detection
Hongjie Chen 0001, Cheung-Chi Leung, Lei Xie 0001, Bin Ma 0001, Haizhou Li 0001 |
INTERSPEECH | 5 |
| 2016 | SingaKids-Mandarin: Speech Corpus of Singaporean Children Speaking Mandarin Chinese
Nancy F. Chen, Rong Tong, Darren Wee, Pei Xuan Lee, Bin Ma 0001, Haizhou Li 0001 |
INTERSPEECH | 6 |
| 2016 | Out of Set Language Modelling in Hierarchical Language Identification
Saad Irtza, Vidhyasaharan Sethu, Sarith Fernando, Eliathamby Ambikairajah, Haizhou Li 0001 |
INTERSPEECH | 5 |
| 2016 | The 2015 NIST Language Recognition Evaluation: The Shared View of I2R, Fantastic4 and SingaMSabstractTechnical report for NIST LRE 2015 Workshop Kong-Aik Lee, Haizhou Li 0001, Li Deng 0001, Ville Hautamäki, Wei Rao 0002, Anthony Larcher, Hanwu Sun, Trung Hieu Nguyen 0001, Guangsen Wang, Aleksandr Sizov, Jianshu Chen, Ivan Kukanov, Amir Hossein Poorjam, Trung Ngo Trong, Chenglin Xu, Haihua Xu 0001, Bin Ma 0001, Chng Eng Siong, Sylvain Meignier |
INTERSPEECH | 2 |