EDBT 2026 Demo / reviewers in the wild / expert
Jianwu Dang 0001
dblp:07/6174-1 · also Jiangwu Dang 0001
· DBLP profile ↗
210ranked-venue papers
9as first author
112since 2021 · last 2026
0000-0002-9237-4821ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 147 · 8 first-author · 81 since 2021Artificial intelligence and machine learning · 134 · 7 first-author · 64 since 2021Databases, data management, data science and information retrieval · 7 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-authorSystems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | UniSonate: A Unified Model for Speech, Music, and Sound Effect Generation with Text InstructionsabstractChunyu Qiang, Xiaopeng Wang, Kang Yin, Yuzhe Liang, Yuxin Guo, Teng Ma, Ziyu Zhang, Tianrui Wang, Cheng Gong, Yushen Chen, Ruibo Fu, Longbiao Wang, Jianwu Dang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Chunyu Qiang, Yuzhe Liang, Tianrui Wang, Yushen Chen, Ruibo Fu, Longbiao Wang, Jianwu Dang 0001 |
ACL (1) | 13 |
| 2026 | Evaluating the Expressive Appropriateness of Speech in Rich ContextsabstractTianrui Wang, Ziyang Ma, Yizhou Peng, Haoyu Wang, Zhikang Niu, Zikang Huang, Yihao Wu, Yi-Wen Chao, Yu Jiang, Yuheng Lu, Guanrou Yang, Xuanchen Li, Hexin Liu, Chunyu Qiang, Cheng Gong, Yifan Yang, Tianchi Liu, Junyu Wang, Nana Hou, Meng Ge, Fuming You, Yang Wei, Zhongqian Sun, Hu Haifeng, Xiaobao Wang, Eng Siong Chng, Xie Chen, Longbiao Wang, Jianwu Dang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Tianrui Wang, Ziyang Ma 0001, Yizhou Peng, Zhikang Niu, Zikang Huang, Yi-Wen Chao, Yuheng Lu, Guanrou Yang, Xuanchen Li, Hexin Liu, Chunyu Qiang, Yifan Yang 0005, Tianchi Liu 0004, Nana Hou, Meng Ge, Fuming You, Zhongqian Sun, Haifeng Hu 0009, Xiaobao Wang, Chng Eng Siong, Xie Chen 0001, Longbiao Wang, Jianwu Dang 0001 |
ACL (1) | 29 |
| 2025 | Enriching Multimodal Sentiment Analysis Through Textual Emotional Descriptions of Visual-Audio ContentabstractMultimodal Sentiment Analysis (MSA) stands as a critical research frontier, seeking to comprehensively unravel human emotions by amalgamating text, audio, and visual data. Yet, discerning subtle emotional nuances within audio and video expressions poses a formidable challenge, particularly when emotional polarities across various segments appear similar. In this paper, our objective is to spotlight emotion-relevant attributes of audio and visual modalities to facilitate multimodal fusion in the context of nuanced emotional shifts in visual-audio scenarios. To this end, we introduce DEVA, a progressive fusion framework founded on textual sentiment descriptions aimed at accentuating emotional features of visual-audio content. DEVA employs an Emotional Description Generator (EDG) to transmute raw audio and visual data into textualized sentiment descriptions, thereby amplifying their emotional characteristics. These descriptions are then integrated with the source data to yield richer, enhanced features. Furthermore, DEVA incorporates the Text-guided Progressive Fusion Module (TPF), leveraging varying levels of text as a core modality guide. This module progressively fuses visual-audio minor modalities to alleviate disparities between text and visual-audio modalities. Experimental results on widely used sentiment analysis benchmark datasets, including MOSI, MOSEI, and CH-SIMS, underscore significant enhancements compared to state-of-the-art models. Moreover, fine-grained emotion experiments corroborate the robust sensitivity of DEVA to subtle emotional variations. Dongxiao He, Xiaobao Wang, Longbiao Wang, Jianwu Dang 0001 |
AAAI | 5 |
| 2025 | HeterGP: Bridging Heterogeneity in Graph Neural Networks with Multi-View PromptingabstractThe challenges tied to unstructured graph data are manifold, primarily falling into node, edge, and graph-level problem categories. Graph Neural Networks (GNNs) serve as effective tools to tackle these issues. However, individual tasks often demand distinct model architectures, and training these models typically requires abundant labeled data, a luxury often unavailable in practical settings. Recently, various "prompt tuning" methodologies have emerged to empower GNNs to adapt to multi-task learning with limited labels. The crux of these methods lies in bridging the gap between pre-training tasks and downstream objectives. Nonetheless, a prevalent oversight in existing studies is the homophily-centric nature of prompt tuning frameworks, disregarding scenarios characterized by high heterogeneity. To remedy this oversight, we introduce a novel prompting strategy named HeterGP tailored for highly heterophilic scenarios. Specifically, we present a dual-view approach to capture both homophilic and heterophilic information, along with a prompt graph design that encompasses token initialization and insertion patterns. Through extensive experiments conducted in a few-shot context encompassing node and graph classification tasks, our method showcases superior performance in highly heterophilic environments compared to state-of-the-art prompt tuning techniques. Fengyu Yan, Xiaobao Wang, Dongxiao He, Longbiao Wang, Jianwu Dang 0001, Di Jin 0001 |
AAAI | 5 |
| 2025 | Reducing the Gap Between Pretrained Speech Enhancement and Recognition Models Using a Real Speech-Trained Bridging ModuleabstractThe information loss or distortion caused by single-channel speech enhancement (SE) harms the performance of automatic speech recognition (ASR). Observation addition (OA) is an effective post-processing method to improve ASR performance by balancing noisy and enhanced speech. Determining the OA coefficient is crucial. However, the currently supervised OA coefficient module, called the bridging module, only utilizes simulated noisy speech for training, which has a severe mismatch with real noisy speech. In this paper, we propose training strategies to train the bridging module with real noisy speech. First, DNSMOS is selected to evaluate the perceptual quality of real noisy speech with no need for the corresponding clean label to train the bridging module. Additional constraints during training are introduced to enhance the robustness of the bridging module further. Each utterance is evaluated by the ASR back-end using various OA coefficients to obtain the word error rates (WERs). The WERs are used to construct a multidimensional vector. This vector is introduced into the bridging module with multi-task learning and is used to determine the optimal OA coefficients. The experimental results on the CHiME-4 dataset show that the proposed methods all had significant improvement compared with the simulated data trained bridging module, especially under real evaluation sets. Zhongjian Cui, Chenrui Cui, Tianrui Wang, Mengnan He, Meng Ge, Caixia Gong, Longbiao Wang, Jianwu Dang 0001 |
ICASSP | 9 |
| 2025 | Augmenting Short Enrollment Speech via Synthesis for Target Speaker ExtractionabstractA high-quality enrollment speech is crucial to target speaker extraction (TSE), since it provides essential cues for identifying the target speaker in the mixture. However, real applications usually only permit a short enrollment speech, e.g. a wakeup word for a mobile device, that provides limited cues. To address this issue, we propose an enrollment augmentation strategy that allows us to enrich the limited enrollment speech with massive text data through speech synthesis. By doing so, the extended enrollment speech contains enhanced speaker timbre and phonetic content which leads to better extraction quality. Furthermore, we propose a training data augmentation strategy to improve the model’s robustness and generalization in short enrollment speech scenarios. Experiments on Libri2Mix demonstrate that our proposed strategies bring a significant improvement in extreme scenarios where only 0.5s and 1-word enrollment speech is provided. We also release our code at https://github.com/HuangZikang-TJU/Aug4TSE. Zikang Huang, Jingru Lin, Meng Ge, Xiaobao Wang, Longbiao Wang, Jianwu Dang 0001 |
ICASSP | 7 |
| 2025 | Discrete Unit-based Low-latency Multi-lingual Speech Synthesis for LIMMITS'25 ChallengeabstractIn this paper, we present the system developed by our team, CCATTS, for the LIMMITS’25 challenge, focusing on few-shot and zero-shot TTS. We adopt a two-stage TTS strategy. In track 1, we fine-tune the pre-trained ZMM-TTS model and successfully achieve multilingual low-latency TTS. In track 2, we propose a token-based framework by modifying the first-stage model of ZMM-TTS, disentangling speech features into four types of discrete tokens—Content, Acoustic, Emotion, and Speaker—and integrating it with our designed token2wav module. This module consists of a HiFi-GAN-style decoder, an acoustic refiner, and U-Net flow matching, to generate high-quality speech. The official competition results demonstrate that our method achieves strong performance in both tracks. Tianrui Wang, Chunyu Qiang, Qiuyu Liu, Yuheng Lu, Xiaobao Wang, Longbiao Wang, Jianwu Dang 0001 |
ICASSP | 11 |
| 2025 | A Chinese Expressive Long-dialogue Speech Dataset with ScriptsabstractWith the advancement of large-scale models, the demand for emotionally rich, long-context, and highly natural communication in human-computer interaction increases. However, the exploration of long-context or script-level speech conversation tasks remains limited due to the lack of specific supervised data. To address this, we introduce a three-stage data processing pipeline for creating a Chinese expressive long-dialogue speech dataset with scripts (CELSDS). We collect videos from TV series, manually annotate speaker information for each character, apply Optical Character Recognition (OCR) to extract speech content, annotate episode summaries, and use a large language model (LLM) to generate sentence-level scenario descriptions. To our knowledge, this is the first Chinese long-context dialogue dataset that incorporates speaker and content annotations, script-level episode summaries, and sentence-level scenario details. Using this dataset, we develop a baseline model for both speech-to-script and script-to-speech generation tasks. The annotations and data production code are open-sourced at: https://github.com/lijin0120/CELSDS. Tianrui Wang, Meng Ge, Chenrui Cui, Jianrong Wang, Longbiao Wang, Jianwu Dang 0001 |
ICASSP | 8 |
| 2025 | Mamba-SEUNet: Mamba UNet for Monaural Speech EnhancementabstractIn recent speech enhancement (SE) research, transformer and its variants have emerged as the predominant methodologies. However, the quadratic complexity of the self-attention mechanism imposes certain limitations on practical deployment. Mamba, as a novel state-space model (SSM), has gained widespread application in natural language processing and computer vision due to its strong capabilities in modeling long sequences and relatively low computational complexity. In this work, we introduce Mamba-SEUNet, an innovative architecture that integrates Mamba with U-Net for SE tasks. By leveraging bidirectional Mamba to model forward and backward dependencies of speech signals at different resolutions, and incorporating skip connections to capture multi-scale information, our approach achieves state-of-the-art (SOTA) performance. Experimental results on the VCTK+DEMAND dataset indicate that Mamba-SEUNet attains a PESQ score of 3.59, while maintaining low computational complexity. When combined with the Perceptual Contrast Stretching technique, Mamba-SEUNet further improves the PESQ score to 3.73. Zizhen Lin, Tianrui Wang, Meng Ge, Longbiao Wang, Jianwu Dang 0001 |
ICASSP | 6 |
| 2025 | A Prompt Learning Framework with Large Language Model Augmentation for Few-shot Multi-label Intent DetectionabstractIntent detection (ID) is essential in spoken language understanding, especially in multi-label settings where intent labels are interdependent and diverse. Existing methods like SE-MLP and QA-FT struggle in few-shot settings, due to limited data availability and efficiency concerns. To address this, we introduce a Prompt Learning framework with large language Model Augmentation (PLMA) for few-shot multi-label ID. In this study, we make three contributions. First, PLMA integrates large language models (LLMs) with small language models (SLMs), using LLMs to enhance SLMs, with prompt learning as the core framework. Second, it leverages LLMs to improve the model’s understanding of both queries and labels by intent spans extraction and answer space expansion. Third, PLMA refines the QA-FT. A novel QA format template is designed to allow single-step training and inference for each input, improving efficiency. Experiments on NLU++ show that PLMA significantly outperforms other baselines in both in-domain and cross-domain settings, demonstrating its effectiveness in combining the strengths of both small and large models. Ning Zhuang, Junlei Li, Xiaobao Wang, Longbiao Wang, Jianwu Dang 0001 |
ICASSP | 7 |
| 2025 | A Progressive Generation Framework with Speech Pre-trained Model for Expressive Voice ConversionabstractExpressive voice conversion (EVC) aims to modify the speaker identity and emotional style of speech while preserving its content. Existing approaches often focus on disentangling speaker, emotion, and content information but overlook the progressive generation mechanisms in human speech production. To address this, we propose a three-stage framework that includes a speech disentanglement module, a progressive generator, and an acoustic refiner. This framework enables speech pre-trained models to parse linguistic content, emotional style, and speaker identity, which are then progressively integrated into the speech reconstruction branch to generate high-quality speech with replaceable emotional style and speaker identity. Experiments with six different pre-trained models show that our framework activates their disentanglement capabilities, surpassing baseline performance in EVC, and supports speaker and emotion control from different target samples. This framework also provides a valuable reference for evaluating the disentanglement capabilities of speech pre-training models. Tianrui Wang, Meng Ge, Zhikang Niu, Chunyu Qiang, Zikang Huang, Ziyang Ma 0001, Xiaobao Wang, Xie Chen 0001, Longbiao Wang, Jianwu Dang 0001 |
ICME | 12 |
| 2025 | Rethinking Contrastive Learning in Graph Anomaly Detection: A Clean-View PerspectiveabstractGraph anomaly detection aims to identify unusual patterns in graph-based data, with wide applications in fields such as web security and financial fraud detection. Existing methods typically rely on contrastive learning, assuming that a lower similarity between a node and its local subgraph indicates abnormality. However, these approaches overlook a crucial limitation: the presence of interfering edges invalidates this assumption, since it introduces disruptive noise that compromises the contrastive learning process. Consequently, this limitation impairs the ability to effectively learn meaningful representations of normal patterns, leading to suboptimal detection performance. To address this issue, we propose a Clean-View Enhanced Graph Anomaly Detection framework (CVGAD), which includes a multi-scale anomaly awareness module to identify key sources of interference in the contrastive learning process. Moreover, to mitigate bias from the one-step edge removal process, we introduce a novel progressive purification module. This module incrementally refines the graph by iteratively identifying and removing interfering edges, thereby enhancing model performance. Extensive experiments on five benchmark datasets validate the effectiveness of our approach. Di Jin 0001, Jingyi Cao, Xiaobao Wang, Bingdao Feng, Dongxiao He, Longbiao Wang, Jianwu Dang 0001 |
IJCAI | 7 |
| 2025 | Integration of Old and New Knowledge for Generalized Intent Discovery: A Consistency-driven Prototype-Prompting FrameworkabstractIntent detection aims to identify user intents from natural language inputs, where supervised methods rely heavily on labeled in-domain (IND) data and struggle with out-of-domain (OOD) intents, limiting their practical applicability. Generalized Intent Discovery (GID) addresses this by leveraging unlabeled OOD data to discover new intents without additional annotation. However, existing methods focus solely on clustering unsupervised data while neglecting domain adaptation. Therefore, we propose a consistency-driven prototype-prompting framework for GID from the perspective of integrating old and new knowledge, which includes a prototype-prompting framework for transferring old knowledge from external sources, and a hierarchical consistency constraint for learning new knowledge from target domains. We conducted extensive experiments and the results show that our method significantly outperforms all baseline methods, achieving state-of-the-art results, which strongly demonstrates the effectiveness and generalization of our methods. Our source code is publicly available at https://github.com/smileix/cpp. Xiaobao Wang, Ning Zhuang, Longbiao Wang, Jianwu Dang 0001 |
IJCAI | 6 |
| 2025 | A Three-Stage Beamforming with Harmonic Guidance for Multi-Channel Speech Enhancement
Nurali Alip, Tianrui Wang, Meng Ge, Jingru Lin, Longbiao Wang, Jianwu Dang 0001 |
INTERSPEECH | 7 |
| 2025 | ASDA: Audio Spectrogram Differential Attention Mechanism for Self-Supervised Representation Learning
Tianrui Wang, Meng Ge, Longbiao Wang, Jianwu Dang 0001 |
INTERSPEECH | 5 |
| 2025 | Word-Level Emotional Expression Control in Zero-Shot Text-to-Speech SynthesisabstractWhile emotional text-to-speech (TTS) has made significant progress, most existing research remains limited to utterance-level emotional expression and fails to support word-level control. Achieving word-level expressive control poses fundamental challenges, primarily due to the complexity of modeling multi-emotion transitions and the scarcity of annotated datasets that capture intra-sentence emotional and prosodic variation. In this paper, we propose WeSCon, the first self-training framework that enables word-level control of both emotion and speaking rate in a pretrained zero-shot TTS model, without relying on datasets containing intra-sentence emotion or speed transitions.
Our method introduces a transition-smoothing strategy and a dynamic speed control mechanism to guide the pretrained TTS model in performing word-level expressive synthesis through a multi-round inference process. To further simplify the inference, we incorporate a dynamic emotional attention bias mechanism and fine-tune the model via self-training, thereby activating its ability for word-level expressive control in an end-to-end manner. Experimental results show that WeSCon effectively overcomes data scarcity, achieving state-of-the-art performance in word-level emotional expression control while preserving the strong zero-shot synthesis capabilities of the original TTS model. Tianrui Wang, Meng Ge, Chunyu Qiang, Ziyang Ma 0001, Zikang Huang, Guanrou Yang, Xiaobao Wang, Chng Eng Siong, Xie Chen 0001, Longbiao Wang, Jianwu Dang 0001 |
NeurIPS | 13 |
| 2025 | Dual-stream Noise and Speech Information Perception based Speech Enhancement
Longbiao Wang, Qiquan Zhang, Jianwu Dang 0001 |
Expert Syst. Appl. | 4 |
| 2025 | APIN: Amplitude- and phase-aware interaction network for speech emotion recognition
Lili Guo 0001, Jie Li 0069, Shifei Ding, Jianwu Dang 0001 |
Speech Commun. | 4 |
| 2025 | HC-APNet: Harmonic Compensation Auditory Perception Network for low-complexity speech enhancement
Meng Ge, Longbiao Wang, Yang-Hao Zhou, Jianwu Dang 0001 |
Speech Commun. | 5 |
| 2025 | LORT: Locally refined convolution and Taylor transformer for monaural speech enhancement
Zizhen Lin, Tianrui Wang, Meng Ge, Longbiao Wang, Jianwu Dang 0001 |
Speech Commun. | 6 |
| 2025 | Heterogeneous Graph Neural Networks using Self-supervised Reciprocally Contrastive LearningabstractHeterogeneous graph neural network (HGNN) is a popular technique for modeling and analyzing heterogeneous graphs. Most existing HGNN-based approaches are supervised or semi-supervised learning methods requiring graphs to be annotated, which is costly and time-consuming. Self-supervised contrastive learning has been proposed to address the problem of requiring annotated data by mining intrinsic properties in the given data. However, the existing contrastive learning methods are not suitable for heterogeneous graphs because they construct contrastive views only based on data perturbation or pre-defined structural properties (e.g., meta-path) in graph data while ignoring noises in node attributes and graph topologies. We develop a robust heterogeneous graph contrastive learning approach, namely HGCL, which introduces two views on respective guidances of node attributes and graph topologies and integrates and enhances them by a reciprocally contrastive mechanism to better model heterogeneous graphs. In this new approach, we adopt distinct but suitable attribute and topology fusion mechanisms in the two views, which are conducive to mining relevant information in attributes and topologies separately. We further use both attribute similarity and topological correlation to construct high-quality contrastive samples. Extensive experiments on four large real-world heterogeneous graphs demonstrate the superiority and robustness of HGCL over several state-of-the-art methods. Cuiying Huo, Dongxiao He, Yawen Li 0001, Di Jin 0001, Jianwu Dang 0001, Witold Pedrycz, Lingfei Wu 0001, Weixiong Zhang |
ACM Trans. Intell. Syst. Technol. | 5 |
| 2025 | Elevating Knowledge-Enhanced Entity and Relationship Understanding for Sarcasm DetectionabstractSarcasm thrives on popular social media platforms such as Twitter and Reddit, where users frequently employ it to convey emotions in an ironic or satirical manner. The ability to detect sarcasm plays a pivotal role in comprehending individuals’ true sentiments. To achieve a comprehensive grasp of sentence semantics, it is crucial to integrate external knowledge that can aid in deciphering entities and their intricate relationships within a sentence. Although some efforts have been made in this regard, their use of external knowledge is still relatively superficial. Specifically, Knowledge-enhanced entity and relationship understanding still face significant challenges. In this paper, we propose the Knowledge Enhanced Sentiment Dependency Graph Convolutional Network (KSDGCN) framework, which constructs a commonsense-augmented sentiment graph and a commonsense-replaced dependency graph for each text to explicitly capture the role of external knowledge for sarcasm detection. Furthermore, we validate the irrational relationships between co-occurring entity pairs within sentences and background knowledge by a signed attention mechanism. We conduct experiments on four benchmark datasets, and the results show that KSDGCN outperforms existing state-of-the-art methods and is highly interpretable. Xiaobao Wang, Yujing Wang 0003, Dongxiao He, Yawen Li 0001, Longbiao Wang, Jianwu Dang 0001, Di Jin 0001 |
IEEE Trans. Knowl. Data Eng. | 7 |
| 2025 | DSTCNet: Deep Spectro-Temporal-Channel Attention Network for Speech Emotion RecognitionabstractSpeech emotion recognition (SER) plays an important role in human-computer interaction, which can provide better interactivity to enhance user experiences. Existing approaches tend to directly apply deep learning networks to distinguish emotions. Among them, the convolutional neural network (CNN) is the most commonly used method to learn emotional representations from spectrograms. However, CNN does not explicitly model features' associations in the spectral-, temporal-, and channel-wise axes or their relative relevance, which will limit the representation learning. In this article, we propose a deep spectro-temporal-channel network (DSTCNet) to improve the representational ability for speech emotion. The proposed DSTCNet integrates several spectro-temporal-channel (STC) attention modules into a general CNN. Specifically, we propose the STC module that infers a 3-D attention map along the dimensions of time, frequency, and channel. The STC attention can focus more on the regions of crucial time frames, frequency ranges, and feature channels. Finally, experiments were conducted on the Berlin emotional database (EmoDB) and interactive emotional dyadic motion capture (IEMOCAP) databases. The results reveal that our DSTCNet can outperform the traditional CNN-based and several state-of-the-art methods. Lili Guo 0001, Shifei Ding, Longbiao Wang, Jianwu Dang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | Minimally-Supervised Speech Synthesis with Conditional Diffusion Model and Language Model: A Comparative Study of Semantic CodingabstractRecently, there has been a growing interest in text-to-speech (TTS) methods that can be trained with minimal supervision by combining two types of discrete speech representations and using two sequence-to-sequence tasks to decouple TTS. However, existing methods suffer from three problems: the high-frequency waveform distortion of discrete speech representations, the prosodic averaging problem caused by the duration prediction model in non-autoregressive frameworks, and difficulty in prediction due to the information redundancy and dimension explosion of existing semantic coding methods. To address these problems, three progressive methods are proposed. First, we propose Diff-LM-Speech, an autoregressive structure consisting of a language model and diffusion models, which models the semantic embedding into the mel-spectrogram based on a diffusion model to achieve higher audio quality. We also introduce a prompt encoder structure based on a variational autoencoder and a prosody bottleneck to improve prompt representation ability. Second, we propose Tetra-Diff-Speech, a non-autoregressive structure consisting of four diffusion model-based modules that design a duration diffusion model to achieve diverse prosodic expressions. Finally, we propose Tri-Diff-Speech, a non-autoregressive structure consisting of three diffusion model-based modules that verify the non-necessity of existing semantic coding models and achieve the best results. Experimental results show that our proposed methods outperform baseline methods. We provide a website with audio samples.1 Chunyu Qiang, Hao Li 0078, He Qu, Ruibo Fu, Tao Wang 0074, Longbiao Wang, Jianwu Dang 0001 |
ICASSP | 8 |
| 2024 | Learning Speech Representation from Contrastive Token-Acoustic PretrainingabstractFor fine-grained generation and recognition tasks such as minimally-supervised text-to-speech (TTS), voice conversion (VC), and automatic speech recognition (ASR), the intermediate representations extracted from speech should serve as a "bridge" between text and acoustic information, containing information from both modalities. The semantic content is emphasized, while the paralinguistic information such as speaker identity and acoustic details should be de-emphasized. However, existing methods for extracting fine-grained intermediate representations from speech suffer from issues of excessive redundancy and dimension explosion. Contrastive learning is a good method for modeling intermediate representations from two modalities. However, existing contrastive learning methods in the audio field focus on extracting global descriptive information for downstream audio classification tasks, making them unsuitable for TTS, VC, and ASR tasks. To address these issues, we propose a method named "Contrastive Token-Acoustic Pretraining (CTAP)", which uses two encoders to bring phoneme and speech into a joint multimodal space, learning how to connect phoneme and speech at the frame level. The CTAP model is trained on 210k speech and phoneme pairs, achieving minimally-supervised TTS, VC, and ASR. The proposed CTAP method offers a promising solution for fine-grained generation and recognition downstream tasks in speech processing. We provide a website with audio samples.1 Chunyu Qiang, Hao Li 0078, Yixin Tian, Ruibo Fu, Tao Wang 0074, Longbiao Wang, Jianwu Dang 0001 |
ICASSP | 7 |
| 2024 | High-Fidelity Speech Synthesis with Minimal Supervision: All Using Diffusion ModelsabstractText-to-speech (TTS) methods have shown promising results in voice cloning, but they require a large number of labeled text-speech pairs. Minimally-supervised speech synthesis decouples TTS by combining two types of discrete speech representations(semantic & acoustic) and using two sequence-to-sequence tasks to enable training with minimal supervision. However, existing methods suffer from information redundancy and dimension explosion in semantic representation, and high-frequency waveform distortion in discrete acoustic representation. Autoregressive frameworks exhibit typical instability and uncontrollability issues. And non-autoregressive frameworks suffer from prosodic averaging caused by duration prediction models. To address these issues, we propose a minimally-supervised high-fidelity speech synthesis method, where all modules are constructed based on the diffusion models. The non-autoregressive framework enhances controllability, and the duration diffusion model enables diversified prosodic expression. Contrastive Token-Acoustic Pretraining (CTAP) is used as an intermediate semantic representation to solve the problems of information redundancy and dimension explosion in existing semantic coding methods. Mel-spectrogram is used as the acoustic representation. Both semantic and acoustic representations are predicted by continuous variable regression tasks to solve the problem of high-frequency fine-grained waveform distortion. Experimental results show that our proposed method outperforms the baseline method. We provide audio samples on our website.1 Chunyu Qiang, Hao Li 0078, Yixin Tian, Yi Zhao 0006, Longbiao Wang, Jianwu Dang 0001 |
ICASSP | 7 |
| 2024 | VoiCor: A Residual Iterative Voice Correction Framework for Monaural Speech Enhancement
Tianrui Wang, Meng Ge, Andong Li, Longbiao Wang, Jianwu Dang 0001, Yungang Jia |
INTERSPEECH | 6 |
| 2024 | An Initial Investigation of Language Adaptation for TTS Systems under Low-resource ScenariosabstractSelf-supervised learning (SSL) representations from massively multilingual models offer a promising solution for low-resource language speech tasks. Despite advancements, language adaptation in TTS systems remains an open problem. This paper explores the language adaptation capability of ZMM-TTS, a recent SSL-based multilingual TTS system proposed in our previous work. We conducted experiments on 12 languages using limited data with various fine-tuning configurations. We demonstrate that the similarity in phonetics between the pretraining and target languages, as well as the language category, affects the target language’s adaptation performance. Additionally, we find that the fine-tuning dataset size and number of speakers influence adaptability. Surprisingly, we also observed that using paired data for fine-tuning is not always optimal compared to audio-only data. Beyond speech intelligibility, our analysis covers speaker similarity, language identification, and predicted MOS. Erica Cooper, Xin Wang 0037, Chunyu Qiang, Mengzhe Geng, Dan Wells, Longbiao Wang, Jianwu Dang 0001, Marc Tessier, Aidan Pine, Korin Richmond, Junichi Yamagishi |
INTERSPEECH | 8 |
| 2024 | Exploring Pre-trained Speech Model for Articulatory Feature Extraction in Dysarthric Speech Using ASR
Yuqin Lin, Longbiao Wang, Jianwu Dang 0001, Nobuaki Minematsu |
INTERSPEECH | 3 |
| 2024 | Error Correction by Paying Attention to Both Acoustic and Confidence References for Automatic Speech Recognition
Yuchun Shu, Longbiao Wang, Jianwu Dang 0001 |
INTERSPEECH | 6 |
| 2024 | Language-Emphasized Cross-Lingual In-Context Learning for Multilingual LLM
Junlei Li, Xiaobao Wang, Ning Zhuang, Longbiao Wang, Jianwu Dang 0001 |
NLPCC (3) | 6 |
| 2024 | MBCFNet: A Multimodal Brain-Computer Fusion Network for human intention recognition
Gaoyan Zhang, Shogo Okada, Longbiao Wang, Jianwu Dang 0001 |
Knowl. Based Syst. | 6 |
| 2024 | Robust voice activity detection using an auditory-inspired masked modulation encoder based convolutional attention network
Longbiao Wang, Meng Ge, Masashi Unoki, Sheng Li 0010, Jianwu Dang 0001 |
Speech Commun. | 6 |
| 2024 | Adversarial Domain Generalized Transformer for Cross-Corpus Speech Emotion RecognitionabstractSpeech emotion recognition (SER) promotes the development of intelligent devices, which enable natural and friendly human-computer interactions. However, the recognition performance of existing approaches is significantly reduced on unseen datasets, and the lack of sufficient training data limits the generalizability of deep learning models. In this work, we analyze the impact of the domain generalization method on cross-corpus SER and propose an adversarial domain generalized transformer (ADoGT), which is aimed at learning a shared feature distribution for the source and target domains. Specifically, we investigate the effect of domain adversarial learning by eliminating nonaffective information. We also combine the center loss with the softmax function as joint supervision to learn discriminative features. Moreover, we introduce unsupervised transfer learning to extract additional features, and incorporate a gated fusion model to learn the complementary information of the features learned by the supervised feature extractor and pretrained model. The proposed transformer based domain generalization method is evaluated using four emotional datasets. We also provide an ablation study of different domain adversarial model structures and feature fusion models. The results of comparative experiments demonstrate the effectiveness of the proposed ADoGT. Yuan Gao 0040, Longbiao Wang, Jiaxing Liu 0001, Jianwu Dang 0001, Shogo Okada |
IEEE Trans. Affect. Comput. | 4 |
| 2024 | ZMM-TTS: Zero-Shot Multilingual and Multispeaker Speech Synthesis Conditioned on Self-Supervised Discrete Speech RepresentationsabstractNeural text-to-speech (TTS) has achieved human-like synthetic speech for single-speaker, single-language synthesis. Multilingual TTS systems are limited to resource-rich languages due to the lack of large paired text and studio-quality audio data. TTS systems are typically built using a single speaker's voice, but there is growing interest in developing systems that can synthesize voices for new speakers using only a few seconds of their speech. This paper presents ZMM-TTS, a multilingual and multispeaker framework utilizing quantized latent speech representations from a large-scale, pre-trained, self-supervised model. Our paper combines text-based and speech-based self-supervised learning models for multilingual speech synthesis. Our proposed model has zero-shot generalization ability not only for unseen speakers but also for unseen languages. We have conducted comprehensive subjective and objective evaluations through a series of experiments. Our model has proven effective in terms of speech naturalness and similarity for both seen and unseen speakers in six high-resource languages. We also tested the efficiency of our method on two hypothetically low-resource languages. The results are promising, indicating that our proposed approach can synthesize audio that is intelligible and has a high degree of similarity to the target speaker's voice, even without any training data for the new, unseen language. Xin Wang 0037, Erica Cooper, Dan Wells, Longbiao Wang, Jianwu Dang 0001, Korin Richmond, Junichi Yamagishi |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2024 | A Prompt-Based Hierarchical Pipeline for Cross-Domain Slot FillingabstractIn task-oriented dialogue systems, slot filling aims to identify the semantic slot types of each token in user utterances. Due to the lack of sufficient supervised data in many scenarios, it is necessary to transfer relevant knowledge by using cross-domain slot filling. Previous studies rely on manually additional meta-information to build the relationships among similar slots across domains, yet not fully utilizing the knowledge learned by language models in the pre-training stage. In this study, we propose a prompt-based hierarchical pipeline (PHP) with three innovations. First, we design a hierarchical pipeline to separately model domain-independent syntactic structures and domain-specific semantic structures, i.e., span detection and slot prediction. Second, we improve the prompt paradigm with discriminative structure to fully utilize pre-trained language models, which reformulates downstream tasks into pre-trained tasks. Finally, we polish our template and verbalizer to effectively utilize task-specific prior knowledge, adding some meta-information and updating their additional trainable parameters. We conducted extensive experiments on three datasets to evaluate our method, and experimental results show that our method significantly outperforms the previous state-of-the-art results. Yuke Si, Longbiao Wang, Xiaobao Wang, Jianwu Dang 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 6 |
| 2023 | Augmenting Affective Dependency Graph via Iterative Incongruity Graph Learning for Sarcasm DetectionabstractRecently, progress has been made towards improving automatic sarcasm detection in computer science. Among existing models, manually constructing static graphs for texts and then using graph neural networks (GNNs) is one of the most effective approaches for drawing long-range incongruity patterns. However, the manually constructed graph structure might be prone to errors (e.g., noisy or incomplete) and not optimal for the sarcasm detection task. Errors produced during the graph construction step cannot be remedied and may accrue to the following stages, resulting in poor performance. To surmount the above limitations, we explore a novel Iterative Augmenting Affective Graph and Dependency Graph (IAAD) framework to jointly and iteratively learn the incongruity graph structure. IAAD can alternatively update the incongruity graph structure and node representation until the learning graph structure is optimal for the metrics of sarcasm detection. More concretely, we begin with deriving an affective and a dependency graph for each instance, then an iterative incongruity graph learning module is employed to augment affective and dependency graphs for obtaining the optimal inconsistent semantic graph with the goal of optimizing the graph for the sarcasm detection task. Extensive experiments on three datasets demonstrate that the proposed model outperforms state-of-the-art baselines for sarcasm detection with significant margins. Xiaobao Wang, Yiqi Dong, Di Jin 0001, Yawen Li 0001, Longbiao Wang, Jianwu Dang 0001 |
AAAI | 6 |
| 2023 | Self-Supervised Audio-Visual Speaker Representation with Co-Meta LearningabstractIn self-supervised speaker verification, the quality of pseudo labels determines the upper bound of its performance and it is not uncommon to end up with massive amount of unreliable pseudo labels. We observe that the complementary information in different modalities ensures a robust supervisory signal for audio and visual representation learning. This motivates us to propose an audio-visual self-supervised learning framework named Co-Meta Learning. Inspired by the Coteaching+, we design a strategy that allows the information of two modalities to be coordinated through the Update by Disagreement. Moreover, we use the idea of modelagnostic meta learning (MAML) to update the network parameters, which makes the hard samples of two modalities to be better resolved by the other modality through gradient regularization. Compared to the baseline, our proposed method achieves a 29.8%, 11.7% and 12.9% relative improvement on Vox-O, Vox-E and Vox-H trials of Voxceleb1 evaluation dataset respectively. Hanyi Zhang, Longbiao Wang, Kong-Aik Lee, Meng Liu 0017, Jianwu Dang 0001 |
ICASSP | 6 |
| 2023 | Brain Network Features Differentiate Intentions from Different Emotional Expressions of the Same TextabstractIntent differentiation in speech communication relies not only on linguistic information but also on paralinguistic information. The same textual content, when pronounced with different prosodies and emotions, may express totally different intentions. The true intentions in this condition can be easily grasped by our brain. Therefore, combining text, speech, and electroencephalography (EEG) for intent discrimination on the same text may be an effective approach. Before fusing speech and text modalities, the current study focused on exploring effective EEG-based features for Chinese intent recognition as no previous research has utilized EEG signals for this purpose. To tackle this issue, we first created a Chinese multimodal spoken language intention understanding (CMSLIU) dataset, in which the same texts were pronounced with varying prosodies to express different intents. To identify effective brain features that were most relevant to intent recognition improvement, we compared the event-related spectral perturbation and effective brain connectivity patterns on two intent conditions (praise vs. irony). It was found that the praise expression tended to elicit stronger high-frequency brain activities while the irony expression involved a more suppressive network connection in the right hemisphere. These features were trained on the CMSLIU dataset and achieved an intention classification accuracy of 78.66%, which indicated a great potential of the EEG features in intent discrimination on the same text. Gaoyan Zhang, Jianwu Dang 0001 |
ICASSP | 4 |
| 2023 | VF-Taco2: Towards Fast and Lightweight Synthesis for Autoregressive Models with Variation Autoencoder and Feature DistillationabstractWith the development of deep learning, end-to-end neural text-to-speech (TTS) systems have achieved significant improvements in high-quality speech synthesis. However, most of these systems are attention-based autoregressive models, resulting in slow synthesis speed and large model parameter sizes. In this paper, we propose a new fast and lightweight TTS framework named VF-Taco2, which can quickly synthesize speech without GPUs. We first profiled the complexity of decoder process in the current autoregressive model and designed a novel multiple frames prediction module based on variational autoencoder (VAE) to alleviate quality degradation when a larger "reduction factor" is applied. Besides, feature distillation is leveraged to compress a relatively large proposed model to its small version with a minor loss of speech quality. Compared to the original Tacotron 2, our VF-Taco2 achieves a 3.6x-4.4x Mel-spectrum generation acceleration on different performance CPUs, and the parameters are compressed by 1.5x with speech quality maintained. Longbiao Wang, Xixin Wu, Qiuyu Liu, Jianwu Dang 0001 |
ICASSP | 6 |
| 2023 | Leveraging Positional-Related Local-Global Dependency for Synthetic Speech DetectionabstractAutomatic speaker verification (ASV) systems are vulnerable to spoofing attacks. As synthetic speech exhibits local and global artifacts compared to natural speech, incorporating local-global dependency would lead to better anti-spoofing performance. To this end, we propose the Rawformer that leverages positional-related local-global dependency for synthetic speech detection. The two-dimensional convolution and Transformer are used in our method to capture local and global dependency, respectively. Specifically, we design a novel positional aggregator that integrates local-global dependency by adding positional information and flattening strategy with less information loss. Furthermore, we propose the squeeze-and-excitation Rawformer (SE-Rawformer), which introduces squeeze-and-excitation operation to acquire local dependency better. The results demonstrate that our proposed SE-Rawformer leads to 37% relative improvement compared to the single state-of-the-art system on ASVspoof 2019 LA and generalizes well on ASVspoof 2021 LA. Especially, using the positional aggregator in the SE-Rawformer brings a 43% improvement on average. Meng Liu 0017, Longbiao Wang, Kong-Aik Lee, Hanyi Zhang, Jianwu Dang 0001 |
ICASSP | 6 |
| 2023 | Cross-Modal Audio-Visual Co-Learning for Text-Independent Speaker VerificationabstractVisual speech (i.e., lip motion) is highly related to auditory speech due to the co-occurrence and synchronization in speech production. This paper investigates this correlation and proposes a cross-modal speech co-learning paradigm. The primary motivation of our cross-modal co-learning method is modeling one modality aided by exploiting knowledge from another modality. Specifically, two cross-modal boosters are introduced based on an audio-visual pseudo-siamese structure to learn the modality-transformed correlation. Inside each booster, a max-feature-map embedded Transformer variant is proposed for modality alignment and enhanced feature generation. The network is co-learned both from scratch and with pretrained models. Experimental results on the test scenarios demonstrate that our proposed method achieves around 60% and 20% average relative performance improvement over baseline unimodal and fusion systems, respectively. Meng Liu 0017, Kong-Aik Lee, Longbiao Wang, Hanyi Zhang, Chang Zeng, Jianwu Dang 0001 |
ICASSP | 6 |
| 2023 | Speech and Noise Dual-Stream Spectrogram Refine Network With Speech Distortion Loss For Robust Speech RecognitionabstractIn recent years, the joint training of speech enhancement front-end and automatic speech recognition (ASR) back-end has been widely used to improve the robustness of ASR systems. Traditional joint training methods only use enhanced speech as input for the backend. However, it is difficult for speech enhancement systems to directly separate speech from input due to the diverse types of noise with different intensities. Furthermore, speech distortion and residual noise are often observed in enhanced speech, and the distortion of speech and noise is different. Most existing methods focus on fusing enhanced and noisy features to address this issue. In this paper, we propose a dual-stream spectrogram refine network to simultaneously refine the speech and noise and decouple the noise from the noisy input. Our proposed method can achieve better performance with a relative 8.6% CER reduction. Haoyu Lu, Tongtong Song, Longbiao Wang, Jianwu Dang 0001, Xiaobao Wang, Shiliang Zhang |
ICASSP | 5 |
| 2023 | Time-Domain Speech Enhancement Assisted by Multi-Resolution Frequency Encoder and DecoderabstractTime-domain speech enhancement (SE) has recently been intensively investigated. Among recent works, DEMUCS [1] introduces multi-resolution STFT loss to enhance performance. However, some resolutions used for STFT contain non-stationary signals, and it is challenging to learn multi-resolution frequency losses simultaneously with only one output. For better use of multi-resolution frequency information, we supplement multiple spectrograms in different frame lengths into the time-domain encoders. They extract stationary frequency information in both narrowband and wideband. We also adopt multiple decoder outputs, each of which computes its corresponding resolution frequency loss. Experimental results show that (1) it is more effective to fuse stationary frequency features than non-stationary features in the encoder, and (2) the multiple outputs consistent with the frequency loss improve performance. Experiments on the Voice-Bank dataset show that the proposed method obtained a 0.14 PESQ improvement. Masato Mimura, Longbiao Wang, Jianwu Dang 0001, Tatsuya Kawahara |
ICASSP | 4 |
| 2023 | Noise-Disentanglement Metric Learning for Robust Speaker VerificationabstractAutomatic speaker verification (ASV) suffers from performance degradation in noisy environments. To solve this problem, we propose the noise-disentanglement metric learning to reduce the speaker-irrelevant noisy components and build a noise-invariant embedding space. Specifically, the disentanglement module, including the speaker encoder and re-construction module, is dedicated to decoupling speech signals. The speaker encoder is used to disentangle speaker-related components, and the reconstruction module increases the model’s ability to constrain the noise information by re-constructing the signal. In addition, distribution optimization is introduced to supervise the spatial structure of speaker embeddings under noisy environments. Experiments on Vox-Celeb1 indicate that the proposed method improves the performance of the speaker verification system in both clean and noisy conditions. Hanyi Zhang, Longbiao Wang, Kong-Aik Lee, Meng Liu 0017, Jianwu Dang 0001 |
ICASSP | 6 |
| 2023 | Multi-Modal Sarcasm Detection Based on Cross-Modal Composition of Inscribed Entity RelationsabstractSarcasm, a linguistic technique employed to express emotions opposite to their literal meaning, has garnered significant attention from researchers due to the rise of social media. Detecting sarcasm in a multi-modal context has become a focal point in recent studies. However, existing research primarily relies on identifying inconsistencies between text semantics and image semantics, often lacking a deep understanding of images. Consequently, capturing inconsistencies between images and texts poses a challenge in many cases. In this paper, we propose the Entity-Relational Graph Convolutional Network (ERGCN) as a solution to detect sarcasm by examining the relationship between entities within images. Our approach involves extracting entities and text descriptions from each image, which provides valuable entity information. Subsequently, we employ external knowledge to construct a cross-modal graph for each text and image pair, emphasizing the presence of internal contradictory information. Finally, we utilize the graph convolutional network to identify inconsistent information across modalities and successfully detect sarcasm. Experimental results demonstrate that our model achieves state-of-the-art performance on a widely used multimodal Twitter dataset. Lingshan Li, Di Jin 0001, Xiaobao Wang, Fengyu Guo, Longbiao Wang, Jianwu Dang 0001 |
ICTAI | 6 |
| 2023 | Commonsense Knowledge Enhanced Sentiment Dependency Graph for Sarcasm DetectionabstractSarcasm is widely utilized on social media platforms such as Twitter and Reddit. Sarcasm detection is required for analyzing people's true feelings since sarcasm is commonly used to portray a reversed emotion opposing the literal meaning. The syntactic structure is the key to make better use of commonsense when detecting sarcasm. However, it is extremely challenging to effectively and explicitly explore the information implied in syntactic structure and commonsense simultaneously. In this paper, we apply the pre-trained COMET model to generate relevant commonsense knowledge, and explore a novel scenario of constructing a commonsense-augmented sentiment graph and a commonsense-replaced dependency graph for each text. Based on this, a Commonsense Sentiment Dependency Graph Convolutional Network (CSDGCN) framework is proposed to explicitly depict the role of external commonsense and inconsistent expressions over the context for sarcasm detection by interactively modeling the sentiment and dependency information. Experimental results on several benchmark datasets reveal that our proposed method beats the state-of-the-art methods in sarcasm detection, and has a stronger interpretability. Di Jin 0001, Xiaobao Wang, Yawen Li 0001, Longbiao Wang, Jianwu Dang 0001 |
IJCAI | 6 |
| 2023 | Local and Global Context Modeling with Relation Matching Task for Dialog Act RecognitionabstractIn dialog act recognition (DAR) of an utterance in a conversation, the prior studies have focused either on the global context using the whole utterances in the dialog, or the local context using the neighbouring utterance flow in the dialog. However, their methods attempt to deal with all types of dialogs indiscriminately. In this study, we propose a model to extract the local context information by an inter-utterance relation matching task (RMT), and a DAR framework to incorporate the local context information into a hierarchical network to fulfil both local and global context modeling. Extensive evaluations were conducted on a Mandarin dialog corpus and two benchmark English corpora. It is found that the different dialog types possess different window lengths for RMT, which is related to the length of subtopics in a given type of dialog. According to ablation experiments, the global information contributed more to the DAR in the hierarchical framework, while the contribution ratio of the local to the global context information was larger than 0.1. The results demonstrated that the proposed RMT and DAR framework significantly improved the DAR performance. Yuke Si, Yan Zhang 0004, Xiaobao Wang, Longbiao Wang, Jianwu Dang 0001, Chng Eng Siong, Haizhou Li 0001 |
IJCNN | 6 |
| 2023 | Locate and Beamform: Two-dimensional Locating All-neural Beamformer for Multi-channel Speech Separation
Yanjie Fu, Meng Ge, Honglong Wang, Longbiao Wang, Gaoyan Zhang, Jianwu Dang 0001, Chengyun Deng |
INTERSPEECH | 8 |
| 2023 | Rethinking the Visual Cues in Audio-Visual Speaker Extraction
Meng Ge, Zexu Pan, Longbiao Wang, Jianwu Dang 0001, Shiliang Zhang |
INTERSPEECH | 6 |
| 2023 | Improving Zero-shot Cross-domain Slot Filling via Transformer-based Slot Semantics Fusion
Yuke Si, Longbiao Wang, Xiaobao Wang, Jianwu Dang 0001 |
INTERSPEECH | 6 |
| 2023 | Discrimination of the Different Intents Carried by the Same Text Through Integrating Multimodal Information
Gaoyan Zhang, Longbiao Wang, Jianwu Dang 0001 |
INTERSPEECH | 4 |
| 2023 | SDNet: Stream-attention and Dual-feature Learning Network for Ad-hoc Array Speech Separation
Honglong Wang, Chengyun Deng, Yanjie Fu, Meng Ge, Longbiao Wang, Gaoyan Zhang, Jianwu Dang 0001 |
INTERSPEECH | 7 |
| 2023 | TMS: Temporal multi-scale in time-delay neural network for speaker verification
Ruiteng Zhang, Jianguo Wei, Xugang Lu, Wenhuan Lu, Di Jin 0001, Lin Zhang 0054, Junhai Xu, Jianwu Dang 0001 |
Appl. Intell. | 8 |
| 2023 | Disordered speech recognition considering low resources and abnormal articulation
Yuqin Lin, Jianwu Dang 0001, Longbiao Wang, Sheng Li 0010, Chenchen Ding |
Speech Commun. | 2 |
| 2023 | A CIF-Based Speech Segmentation Method for Streaming E2E ASRabstractLong utterances segmentation is crucial in end-to-end (E2E) streaming automatic speech recognition (ASR). However, commonly used voice activity detection(VAD)-based and fixed-length segmentation methods may lead to long segments and semantic incompleteness, affecting the user experience and ASR performance. In this paper, we propose a speech segmentation method for streaming E2E ASR to solve the above issues. Both the decoder's dependence on acoustic information and the human average breath frequency are used for judging segment boundaries. Frame-level decoder's dependence information is provided by the Continuous Integrate-and-Fire (CIF) predictor, which optimizes jointly with ASR to guarantee a more suitable segmentation for ASR. Besides, the proposed method does not increase the model parameters and real-time factor (RTF). The experimental results show that our method can accurately detect the pauses in speech, and the segment usually contains relatively complete semantic information. Compared with VAD-based segmentation, 53.5% latency reduction and 3.7% CER reduction relatively are achieved. Yuchun Shu, Haoneng Luo, Shiliang Zhang, Longbiao Wang, Jianwu Dang 0001 |
IEEE Signal Process. Lett. | 5 |
| 2023 | CFDRN: A Cognition-Inspired Feature Decomposition and Recombination Network for Dysarthric Speech RecognitionabstractAs an essential technology in human–computer interactions, automatic speech recognition (ASR) ensures a convenient life for healthy people; however, people with speech disorders, who truly need support from such a technology, have experienced difficulties in the use of ASR. Disordered ASR is challenging because of the large variabilities in disordered speech. Humans tend to separately process different spectro-temporal features of speech in the left and right hemispheres of their brain, showing significantly better ability in speech perception than machines, especially in disordered speech perception. Inspired by human speech processing, this paper proposes a cognition-inspired feature decomposition and recombination network (CFDRN) for dysarthric ASR. In the CFDRN, slow- and rapid-varying temporal processors are designed to decompose features into stable and changeable features, respectively. A gated fusion module was developed to selectively recombine the decomposed features. Moreover, this study utilised an adaptation approach based on unsupervised pre-training techniques to alleviate data scarcity issues in dysarthric ASR. The CFDRNs were added to the layers of the pre-trained model, and the entire model is adapted from normal speech to disordered speech. The effectiveness of the proposed method was validated on the widely used TORGO and UASpeech dysarthria datasets under three popular unsupervised pre-training techniques, wav2vec 2.0, HuBERT, and data2vec. When compared to the baseline methods, the proposed CFDRN with the three pre-training techniques achieved 13.73%$\sim$16.23% and 4.50%$\sim$13.20% word error rate reductions on the TORGO and UASpeech datasets, respectively. Furthermore, this study clarified several major factors affecting dysarthric ASR performance. Yuqin Lin, Longbiao Wang, Yanbing Yang 0003, Jianwu Dang 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2023 | Meta-Generalization for Domain-Invariant Speaker VerificationabstractAutomatic speaker verification (ASV) exhibits unsatisfactory performance under domain mismatch conditions owing to intrinsic and extrinsic factors, such as variations in speaking styles and recording devices encountered in real-world applications. To ensure robust performance under unseen conditions, domain generalization has been explored. However, an inherent contradiction exists between model discrimination and domain generalization, in which the discrimination ability may be reduced while learning to generalize. In this paper, to extract discriminative yet domain-invariant representations, we propose the meta-generalized speaker verification (MGSV) via meta-learning. Specifically, we propose a metric-based distribution optimization and a gradient-based meta-optimization to simultaneously supervise the spatial relationship between embeddings and improve the generalization ability of the model on unseen domains. In addition, we design multiple-single (MS) and simulated speaker verification (SSV) sampling strategies based on single-domain (SD) and single-single (SS) strategies to simulate the train/test domain mismatch more relevantly, thereby mining transferable speaker-related knowledge. SSV is chosen as the most effective method, as it substantially improves the domain generalization by ensuring that the model has learned to discriminate efficiently. Additionally, to intuitively reflect the model performance on the unseen domains, the proposed method is validated on cross-genre, cross-device, and cross-dataset tasks. The experimental results demonstrate that our proposed method achieves remarkable performance in handling domain mismatch issues in speaker verification. Hanyi Zhang, Longbiao Wang, Kong-Aik Lee, Meng Liu 0017, Jianwu Dang 0001, Helen M. Meng |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2022 | Domain-Invariant Feature Learning for Cross Corpus Speech Emotion RecognitionabstractTo deal with speech emotion recognition (SER) in real-life applications, researchers have to focus on cross corpus SER, where the feature distribution of source and target datasets are different. In this paper, we propose an efficient domain adversarial training method to cope with the non-affective information during feature extraction. Through the proposed domain-adversarial learning, we can reduce the domain divergence between train and test data. Furthermore, we incorporate center loss with the emotion classifier to reduce the intra-class variation of features learned from the same emotion. We conduct experiments on four emotional benchmark datasets to verify the performance of the proposed method. The experimental results demonstrate that our proposed model outperform the baseline system in both cross-corpus and multi-corpus evaluation. Yuan Gao 0040, Shogo Okada, Longbiao Wang, Jiaxing Liu 0001, Jianwu Dang 0001 |
ICASSP | 5 |
| 2022 | L-SpEx: Localized Target Speaker ExtractionabstractSpeaker extraction aims to extract the target speaker’s voice from a multi-talker speech mixture given an auxiliary reference utterance. Recent studies show that speaker extraction benefits from the location or direction of the target speaker. However, these studies assume that the target speaker’s location is known in advance or detected by an extra visual cue, e.g., face image or video. In this paper, we propose an end-to-end localized target speaker extraction on pure speech cues, that is called L-SpEx. Specifically, we design a speaker localizer driven by the target speaker’s embedding to extract the spatial features, including direction-of-arrival (DOA) of the target speaker and beamforming output. Then, the spatial cues and target speaker’s embedding are both used to form a top-down auditory attention to the target speaker. Experiments on the multi-channel reverberant dataset called MCLibri2Mix show that our L-SpEx approach significantly outperforms the baseline system. Meng Ge, Chenglin Xu, Longbiao Wang, Chng Eng Siong, Jianwu Dang 0001, Haizhou Li 0001 |
ICASSP | 5 |
| 2022 | Using Multiple Reference Audios and Style Embedding Constraints for Speech SynthesisabstractThe end-to-end speech synthesis model can directly take an utterance as reference audio, and generate speech from the text with prosody and speaker characteristics similar to the reference audio. However, an appropriate acoustic embedding must be manually selected during inference. Due to the fact that only the matched text and speech are used in the training process, using unmatched text and speech for inference would cause the model to synthesize speech with low content quality. In this study, we propose to mitigate these two problems by using multiple reference audios and style embedding constraints rather than using only the target audio. Multiple reference audios are automatically selected using the sentence similarity determined by Bidirectional Encoder Representations from Transformers (BERT). In addition, we use "target" style embedding from a pre-trained encoder as a constraint by considering the mutual information between the predicted and "target" style embedding. The experimental results show that the proposed model can improve the speech naturalness and content quality with multiple reference audios and can also outperform the baseline model in ABX preference tests of style similarity. Longbiao Wang, Zhen-Hua Ling, Ju Zhang 0001, Jianwu Dang 0001 |
ICASSP | 5 |
| 2022 | Compressing Transformer-Based ASR Model by Task-Driven Loss and Attention-Based Multi-Level Feature DistillationabstractThe current popular knowledge distillation (KD) methods effectively compress the transformer-based end-to-end speech recognition model. However, existing methods fail to utilize complete information of the teacher model, and they distill only a limited number of blocks of the teacher model. In this study, we first integrate a task-driven loss function into the decoder’s intermediate blocks to generate task-related feature representations. Then, we propose an attention-based multi-level feature distillation to automatically learn the feature representation summarized by all blocks of the teacher model. Under the 1.1M parameters model, the experimental results on the Wall Street Journal dataset reveal that our approach achieves a 12.1% WER reduction compared with the baseline system. Yongjie Lv, Longbiao Wang, Meng Ge, Sheng Li 0010, Chenchen Ding, Lixin Pan, Yuguang Wang 0003, Jianwu Dang 0001, Kiyoshi Honda |
ICASSP | 8 |
| 2022 | Cache: Modeling Contribution-Aware Context Hierarchically for Long-Range Dialogue State TrackingabstractRecently, many studies on dialogue state tracking (DST) based on the copy-augmented encoder-decoder framework have been proposed and have achieved encouraging performance. However, these studies commonly lose earlier information during encoding the long dialogues with RNNs, and have difficulty for the decoder to focus on specific dialogue turns from lengthy context, which causes decreased performance as the dialogue gets longer. In this work, we propose a novel method to model Contribution-Aware Context HiErarchically (CACHE) with a hierarchical encoder and a slot-turn attention module. The hierarchical encoder is designed to prevent information loss by reducing the length of the sequence sent to each encoder. The slot-turn attention module is explored to help the decoder focus on the slotrelated dialogue turn information. To evaluate models more appropriately, we introduce a new metric continued joint accuracy considering the prediction accuracy of both current and historical dialogue turns. Experiments on MultiWOZ 2.0 show that CACHE is an effective model for tracking states especially in long context. Jianshu Qi, Yuke Si, Longbiao Wang, Jianwu Dang 0001 |
ICASSP | 4 |
| 2022 | Multi-Stage Graph Representation Learning for Dialogue-Level Speech Emotion RecognitionabstractWith the development of speech emotion recognition (SER), most of current research is utterance-level and cannot fit the need of actual scenarios. In this paper, we propose a novel strategy that focuses on capturing dialogue-level contextual information. On the basis of utterance-level representation learned by convolutional neural network (CNN) which is followed by the bidirectional long short-term memory network (BLSTM), the proposed dialogue-level method consists of two modules. The first module is Dialogue Multi-stage Graph Representation Learning Algorithm (DialogMSG). The multi-stage graph that modeling from different dialogue scope is introduced to capture more effective information. The other one is a double-constrained module. This module includes not only an utterance-level classifier but also a dialogue-level graph classifier which is named as Atmosphere. The results of extensive experiments show that the proposed method outperforms the current state of the art on the IEMOCAP benchmark dataset. Yaodong Song, Jiaxing Liu 0001, Longbiao Wang, Jianwu Dang 0001 |
ICASSP | 5 |
| 2022 | Learning Domain-Invariant Transformation for Speaker VerificationabstractAutomatic speaker verification (ASV) faces domain shift caused by the mismatch of intrinsic and extrinsic factors such as recording device and speaking style in real-world applications, which leads to unsatisfactory performance. To this end, we propose the meta generalized transformation via meta-learning to build a domain-invariant embedding space. Specifically, the transformation module is motivated to learn the domain generalization knowledge by executing meta-optimization on the meta-train and meta-test sets which are designed to simulate domain shift. Furthermore, distribution optimization is incorporated to supervise the metric structure of embeddings. In terms of the transformation module, we investigate various instantiations and observe the multilayer perceptron with gating (gMLP) is the most effective given its extrapolation capability. The experimental results on cross-genre and cross-dataset settings demonstrate that the meta generalized transformation dramatically improves the robustness of ASV systems to domain shift, while outperforms the state-of-the-art methods. Hanyi Zhang, Longbiao Wang, Kong-Aik Lee, Meng Liu 0017, Jianwu Dang 0001 |
ICASSP | 5 |
| 2022 | Improving Dialogue Generation via Proactively Querying Grounded KnowledgeabstractRecent advances in pre-trained language models have significantly improved neural response generation. Further, an intelligent dialogue system should be able to give accurate responses that meet the needs of users. However, using appropriate knowledge has so far been proved difficult, because faced with a mass of relevant knowledge, the model not only needs to accurately retrieve the target information, but also integrate the information into dialogue response. In this paper, we propose a novel knowledge-based dialogue system which integrates the strength of a transformer-based generator and a knowledge retriever capable of proactively constructing queries for accurate information. Specifically, generator are used to give a response skeleton and knowledge retriever constructs a query based on response skeleton to obtain specific knowledge, which avoids excessive noise interference when generating responses. Experiments show that conversational systems that leverage knowledge integrator could generate more informative and human-like responses than strong baseline systems. Longbiao Wang, Jianwu Dang 0001 |
ICASSP | 3 |
| 2022 | An Improved Stimulus Reconstruction Method for EEG-Based Short-Time Auditory Attention Detection
Gaoyan Zhang, Masashi Unoki, Jianwu Dang 0001, Longbiao Wang |
ICONIP (5) | 5 |
| 2022 | Dual-stream Speech Dereverberation Network Using Long-term and Short-term CuesabstractFor reverberation, the current speech is usually influenced by the previous frames. Traditional neural network-based speech dereverberation (SD) methods directly map the current speech frame that only has short-term cues to clean speech or learn a mask, which can not utilize long-term information to remove late reverberation and further limit SD's ability. To address this issue, we propose a dual-stream speech dereverberation network (DualSDNet) using long-term and short-term cues. First, we analyze the effectiveness of using a finite impulse response (FIR) based on long-term information recorded filter by reverberation generation progress. Second, to make full use of both long-term and short-term information, we further design a dual-stream network, it can map both long and short speech to high-dimensional representation and pay more attention to a more helpful time index. The results of the REVERB Challenge data show that our DualSDNet consistently outperforms the state-of-the-art SD baselines. Meng Ge, Longbiao Wang, Jianwu Dang 0001 |
IJCNN | 4 |
| 2022 | Iterative Sound Source Localization for Unknown Number of SourcesabstractSound source localization aims to seek the direction of arrival (DOA) of all sound sources from the observed multichannel audio.For the practical problem of unknown number of sources, existing localization algorithms attempt to predict a likelihood-based coding (i.e., spatial spectrum) and employ a pre-determined threshold to detect the source number and corresponding DOA value.However, these threshold-based algorithms are not stable since they are limited by the careful choice of threshold.To address this problem, we propose an iterative sound source localization approach called ISSL, which can iteratively extract each source's DOA without threshold until the termination criterion is met.Unlike threshold-based algorithms, ISSL designs an active source detector network based on binary classifier to accept residual spatial spectrum and decide whether to stop the iteration.By doing so, our ISSL can deal with an arbitrary number of sources, even more than the number of sources seen during the training stage.The experimental results show that our ISSL achieves significant performance improvements in both DOA estimation and source number detection compared with the existing threshold-based algorithms. Yanjie Fu, Meng Ge, Xinyuan Qian 0001, Longbiao Wang, Gaoyan Zhang, Jianwu Dang 0001 |
INTERSPEECH | 7 |
| 2022 | Improve emotional speech synthesis quality by learning explicit and implicit representations with semi-supervised training
Jiaxu He, Longbiao Wang, Di Jin 0001, Xiaobao Wang, Junhai Xu, Jianwu Dang 0001 |
INTERSPEECH | 7 |
| 2022 | Data Augmentation Using McAdams-Coefficient-Based Speaker Anonymization for Fake Audio DetectionabstractFake audio detection (FAD) is a technique to distinguish synthetic speech from natural speech.In most FAD systems, removing irrelevant features from acoustic speech while keeping only robust discriminative features is essential.Intuitively, speaker information entangled in acoustic speech should be suppressed for the FAD task.Particularly in a deep neural network (DNN)-based FAD system, the learning system may learn speaker information from a training dataset and cannot generalize well on a testing dataset.In this paper, we propose to use the speaker anonymization (SA) technique to suppress speaker information from acoustic speech before inputting it into a DNN-based FAD system.We adopted the McAdamscoefficient-based SA (MC-SA) algorithm, and this is expected that the entangled speaker information will not be involved in the DNN-based FAD learning.Based on this idea, we implemented a light convolutional neural network bidirectional long short-term memory (LCNN-BLSTM)-based FAD system and conducted experiments on the Audio Deep Synthesis Detection Challenge (ADD2022) datasets.The results showed that removing the speaker information from acoustic speech improved the relative performance in the first track of ADD2022 by 17.66%. Kai Li 0018, Sheng Li 0010, Xugang Lu, Masato Akagi, Meng Liu 0017, Lin Zhang 0054, Chang Zeng, Longbiao Wang, Jianwu Dang 0001, Masashi Unoki |
INTERSPEECH | 9 |
| 2022 | VCSE: Time-Domain Visual-Contextual Speaker Extraction NetworkabstractSpeaker extraction seeks to extract the target speech in a multitalker scenario given an auxiliary reference.Such reference can be auditory, i.e., a pre-recorded speech, visual, i.e., lip movements, or contextual, i.e., phonetic sequence.References in different modalities provide distinct and complementary information that could be fused to form top-down attention on the target speaker.Previous studies have introduced visual and contextual modalities in a single model.In this paper, we propose a two-stage time-domain visual-contextual speaker extraction network named VCSE, which incorporates visual and selfenrolled contextual cues stage by stage to take full advantage of every modality.In the first stage, we pre-extract a target speech with visual cues and estimate the underlying phonetic sequence.In the second stage, we refine the pre-extracted target speech with the self-enrolled contextual cues.Experimental results on the real-world Lip Reading Sentences 3 (LRS3) database demonstrate that our proposed VCSE network consistently outperforms other state-of-the-art baselines. Meng Ge, Zexu Pan, Longbiao Wang, Jianwu Dang 0001 |
INTERSPEECH | 5 |
| 2022 | Global Signal-to-noise Ratio Estimation Based on Multi-subband Processing Using Convolutional Neural Network
Meng Ge, Longbiao Wang, Masashi Unoki, Sheng Li 0010, Jianwu Dang 0001 |
INTERSPEECH | 6 |
| 2022 | Finer-grained Modeling units-based Meta-Learning for Low-resource Tibetan Speech Recognition
Siqing Qin, Longbiao Wang, Sheng Li 0010, Yuqin Lin, Jianwu Dang 0001 |
INTERSPEECH | 5 |
| 2022 | Monaural Speech Enhancement Based on Spectrogram Decomposition for Convolutional Neural Network-sensitive Feature Extraction
Longbiao Wang, Sheng Li 0010, Jianwu Dang 0001, Tatsuya Kawahara |
INTERSPEECH | 4 |
| 2022 | Language-specific Characteristic Assistance for Code-switching Speech Recognition
Tongtong Song, Meng Ge, Longbiao Wang, Yongjie Lv, Yuqin Lin, Jianwu Dang 0001 |
INTERSPEECH | 8 |
| 2022 | TopicKS: Topic-driven Knowledge Selection for Knowledge-grounded Dialogue Generation
Shiquan Wang, Yuke Si, Longbiao Wang, Zhiqiang Zhuang, Xiaowang Zhang, Jianwu Dang 0001 |
INTERSPEECH | 7 |
| 2022 | Hierarchical Tagger with Multi-task Learning for Cross-domain Slot Filling
Yuke Si, Shiquan Wang, Longbiao Wang, Jianwu Dang 0001 |
INTERSPEECH | 5 |
| 2022 | Self-Distillation Based on High-level Information Supervision for Compressing End-to-End ASR Model
Tongtong Song, Longbiao Wang, Yuqin Lin, Yongjie Lv, Meng Ge, Qiang Yu 0005, Jianwu Dang 0001 |
INTERSPEECH | 9 |
| 2022 | MIMO-DoAnet: Multi-channel Input and Multiple Outputs DoA Network with Unknown Number of Sound SourcesabstractRecent neural network based Direction of Arrival (DoA) estimation algorithms have performed well on unknown number of sound sources scenarios.These algorithms are usually achieved by mapping the multi-channel audio input to the single output (i.e.overall spatial pseudo-spectrum (SPS) of all sources), that is called MISO.However, such MISO algorithms strongly depend on empirical threshold setting and the angle assumption that the angles between the sound sources are greater than a fixed angle.To address these limitations, we propose a novel multi-channel input and multiple outputs DoA network called MIMO-DoAnet.Unlike the general MISO algorithms, MIMO-DoAnet predicts the SPS coding of each sound source with the help of the informative spatial covariance matrix.By doing so, the threshold task of detecting the number of sound sources becomes an easier task of detecting whether there is a sound source in each output, and the serious interaction between sound sources disappears during inference stage.Experimental results show that MIMO-DoAnet achieves relative 18.6% and absolute 13.3%, relative 34.4% and absolute 20.2% F1 score improvement compared with the MISO baseline system in 3, 4 sources scenes.The results also demonstrate MIMO-DoAnet alleviates the threshold setting problem and solves the angle assumption problem effectively. Meng Ge, Yanjie Fu, Gaoyan Zhang, Longbiao Wang, Jianwu Dang 0001 |
INTERSPEECH | 8 |
| 2022 | Learning affective representations based on magnitude and dynamic relative phase information for speech emotion recognition
Lili Guo 0001, Longbiao Wang, Jianwu Dang 0001, Chng Eng Siong, Seiichi Nakagawa |
Speech Commun. | 3 |
| 2022 | One-shot emotional voice conversion based on feature separation
Wenhuan Lu, Xinyue Zhao, Jianguo Wei, Jianhua Tao 0001, Jianwu Dang 0001 |
Speech Commun. | 7 |
| 2022 | Toward Efficient Processing and Learning With Spikes: New Approaches for Multispike LearningabstractSpikes are the currency in central nervous systems for information transmission and processing. They are also believed to play an essential role in low-power consumption of the biological systems, whose efficiency attracts increasing attentions to the field of neuromorphic computing. However, efficient processing and learning of discrete spikes still remain a challenging problem. In this article, we make our contributions toward this direction. A simplified spiking neuron model is first introduced with the effects of both synaptic input and firing output on the membrane potential being modeled with an impulse function. An event-driven scheme is then presented to further improve the processing efficiency. Based on the neuron model, we propose two new multispike learning rules which demonstrate better performance over other baselines on various tasks, including association, classification, and feature detection. In addition to efficiency, our learning rules demonstrate high robustness against the strong noise of different types. They can also be generalized to different spike coding schemes for the classification task, and notably, the single neuron is capable of solving multicategory classifications with our learning rules. In the feature detection task, we re-examine the ability of unsupervised spike-timing-dependent plasticity with its limitations being presented, and find a new phenomenon of losing selectivity. In contrast, our proposed learning rules can reliably solve the task over a wide range of conditions without specific constraints being applied. Moreover, our rules cannot only detect features but also discriminate them. The improved performance of our methods would contribute to neuromorphic computing as a preferable choice. Qiang Yu 0005, Shenglan Li, Huajin Tang, Longbiao Wang, Jianwu Dang 0001, Kay Chen Tan |
IEEE Trans. Cybern. | 5 |
| 2022 | Constructing Accurate and Efficient Deep Spiking Neural Networks With Double-Threshold and Augmented SchemesabstractSpiking neural networks (SNNs) are considered as a potential candidate to overcome current challenges, such as the high-power consumption encountered by artificial neural networks (ANNs); however, there is still a gap between them with respect to the recognition accuracy on various tasks. A conversion strategy was, thus, introduced recently to bridge this gap by mapping a trained ANN to an SNN. However, it is still unclear that to what extent this obtained SNN can benefit both the accuracy advantage from ANN and high efficiency from the spike-based paradigm of computation. In this article, we propose two new conversion methods, namely TerMapping and AugMapping. The TerMapping is a straightforward extension of a typical threshold-balancing method with a double-threshold scheme, while the AugMapping additionally incorporates a new scheme of augmented spike that employs a spike coefficient to carry the number of typical all-or-nothing spikes occurring at a time step. We examine the performance of our methods based on the MNIST, Fashion-MNIST, and CIFAR10 data sets. The results show that the proposed double-threshold scheme can effectively improve the accuracies of the converted SNNs. More importantly, the proposed AugMapping is more advantageous for constructing accurate, fast, and efficient deep SNNs compared with other state-of-the-art approaches. Our study, therefore, provides new approaches for further integration of advanced techniques in ANNs to improve the performance of SNNs, which could be of great merit to applied developments with spike-based neuromorphic computing. Qiang Yu 0005, Chenxiang Ma, Shiming Song 0001, Gaoyan Zhang, Jianwu Dang 0001, Kay Chen Tan |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2021 | Learning Language and Speaker Information for Code-Switch Speech Synthesis with Limited DataabstractEnd-to-end speech synthesis demonstrates remarkable performance in monolingual speech, whereas code-switching (CS) speech synthesis remains a challenge owing to the sparsity of data and diverse syntactic structures across languages. Previous studies show that large mixed-lingual corpora are essential for effective learning text/language representations and target speaker information. In this study, we propose a method using three independent encoders (text, language, and speaker), which requires only a small amount of mixed-lingual data to realize the CS speech synthesis of Mandarin and English. Additionally, to distinguish between Mandarin and English, we investigate two text-representation methods: (1) the implicit method, which uses Pinyin and the CMU11http://www.speech.cs.cmu.edu/cgi-bin/cmudict dictionary to represent both languages; and (2) the explicit method, which uses language markers i.e., masks, to differentiate the languages. Through our proposed method, we can improve synthesized speech in terms of quality and speaker similarity using a small amount of mixed-lingual data. In addition, the experimental results demonstrate that the proposed method achieves performance improvement of 0.06 in terms of the mean opinion score and absolute improvement of 0.64% in terms of the character error rate compared to the baseline method. Mengxin Chai, Shaotong Guo, Longbiao Wang, Jianwu Dang 0001, Ju Zhang 0001 |
ASRU | 5 |
| 2021 | DeepLip: A Benchmark for Deep Learning-Based Audio-Visual Lip BiometricsabstractAudio-visual lip biometrics (AV-LB) has been an emerging biometrics technology that straddles auditory and visual speech processing. Previous works mainly focused on the front-end lip-based feature engineering combined with a shallow statistical back-end model. Over the past decade, convolutional neural network (CNN, or ConvNet) has been widely used and achieved good performance in computer vision and speech processing tasks. However, the lack of a sizeable public AV-LB database led to a stagnation in deep-learning exploration on AV-LB tasks. In addition to the dual audio-visual streams, one essential requirement on the video stream is the region of interest (ROI) around the lips has to be of sufficient resolution. To this end, we compile a moderate-size database using existing public databases. Using this database, we present a deep learning-based AV-LB benchmark, dubbed DeepLip11https://github.com/DanielMengLiu/DeepLip, realized with convolutional video and audio unimodal modules, and a multimodal fusion module. Our experiments show that DeepLip outperforms the traditional lip-biometrics system in context modeling and achieves over 50% relative improvements compared with its unimodal system, with an equal error rate of 0.75% and 1.11% on the test datasets, respectively. Meng Liu 0017, Longbiao Wang, Kong-Aik Lee, Hanyi Zhang, Chang Zeng, Jianwu Dang 0001 |
ASRU | 6 |
| 2021 | Domain-Adversarial Autoencoder with Attention Based Feature Level Fusion for Speech Emotion RecognitionabstractOver the past two decades, although speech emotion recognition (SER) has garnered considerable attention, the problem of insufficient training data has been unresolved. A potential solution for this problem is to pre-train a model and transfer knowledge from large amounts of audio data. However, the data used for pre-training and testing originate from different domains, resulting in the latent representations to contain non-affective information. In this paper, we propose a domain-adversarial autoencoder to extract discriminative representations for SER. Through domain-adversarial learning, we can reduce the mismatch between domains while retaining discriminative information for emotion recognition. We also introduce multi-head attention to capture emotion information from different subspaces of input utterances. Experiments on IEMOCAP show that the proposed model outperforms the state-of-the-art systems by improving the unweighted accuracy by 4.15%, thereby demonstrating the effectiveness of the proposed model. Yuan Gao 0040, Jiaxing Liu 0001, Longbiao Wang, Jianwu Dang 0001 |
ICASSP | 4 |
| 2021 | Multi-Stage Speaker Extraction with Utterance and Frame-Level Reference SignalsabstractSpeaker extraction requires a sample speech from the target speaker as the reference. However, enrolling a speaker with a long speech is not practical. We propose a speaker extraction technique, that performs in multiple stages to take full advantage of short reference speech sample. The extracted speech in early stages is used as the reference speech for late stages. For the first time, we use frame-level sequential speech embedding as the reference for target speaker. This is a departure from the traditional utterance-based speaker embedding reference. In addition, a signal fusion scheme is proposed to combine the decoded signals in multiple scales with automatically learned weights. Experiments on WSJ0-2mix and its noisy versions (WHAM! and WHAMR!) show that SpEx++ consistently outperforms other state-of-the-art baselines. Meng Ge, Chenglin Xu, Longbiao Wang, Chng Eng Siong, Jianwu Dang 0001, Haizhou Li 0001 |
ICASSP | 5 |
| 2021 | Improving Naturalness and Controllability of Sequence-to-Sequence Speech Synthesis by Learning Local Prosody RepresentationsabstractState-of-the-art neural text-to-speech (TTS) networks are trained with a large amount of speech data, which significantly improves the quality of synthetic speech compared with traditional approaches. However, the prosody and controllability of the generated speech is still insufficient, especially in tonal languages. Moreover, the generated prosody is solely defined by the input text, which does not allow for different styles for the same sentence or words. In this study, we extended Tacotron2 with a pitch prediction task to capture discrete pitch-related representations. Specifically, the learned pitch-related suprasegmental information is fed simultaneously with traditional character features into the decoder to generate final Mel spectrogram. Experiments show that the proposed method can improve the quality of the generated speech (mean opinion score of 4.37 vs. 4.22). Moreover, we demonstrated that we can easily achieve word-level pitch control during generation by changing local pitch-related representations before passing them to the decoder network. Longbiao Wang, Zhen-Hua Ling, Shaotong Guo, Ju Zhang 0001, Jianwu Dang 0001 |
ICASSP | 6 |
| 2021 | Representation Learning with Spectro-Temporal-Channel Attention for Speech Emotion RecognitionabstractConvolutional neural network (CNN) is found to be effective in learning representation for speech emotion recognition. CNNs do not explicitly model the associations or relative importance of features in the spectral/temporal/channel-wise axes. In this paper, we propose an attention module, named spectro-temporal-channel (STC) attention module that is integrated with CNN to improve representation learning ability. Our module infers an attention map along the three dimensions, namely time, frequency, and CNN channel. Experiments are conducted on the IEMOCAP database to evaluate the effectiveness of the proposed representation learning method. The results demonstrate that the proposed method outperforms the traditional CNN method by an absolute increase of 3.13% in terms of F1 score. Lili Guo 0001, Longbiao Wang, Chenglin Xu, Jianwu Dang 0001, Chng Eng Siong, Haizhou Li 0001 |
ICASSP | 4 |
| 2021 | Robust Voice Activity Detection Using a Masked Auditory Encoder Based Convolutional Neural NetworkabstractVoice activity detection (VAD) based on deep learning has achieved remarkable success. However, when the traditional features (e.g., raw waveforms and MFCCs) are directly fed to the deep neural network model, the performance decreases because of noise interference. Here, we propose a robust VAD approach using a masked auditory encoder based convolutional neural network (M-AECNN). First, we analyze the effectiveness of using auditory features as deep learning encoder. These features can roughly simulate the transmission of sound to human inner-ear hair cells; thus, they are more robust than the raw waveform and frequency domain features designed as encoders. Second, similar to the human ear’s masking effect for different speech frequencies, the proposed auditory encoder can further improve the robustness of VAD by increasing the gain for cleaner speech frequencies. Extensive experimental results demonstrate that this approach achieves about 10.5% absolute improvement in the area under the curve on the AURORA-2J dataset compared with a VAD method based on a CNN and MFCCs. Longbiao Wang, Masashi Unoki, Sheng Li 0010, Rui Wang 0102, Meng Ge, Jianwu Dang 0001 |
ICASSP | 7 |
| 2021 | Multimodal Emotion Recognition with Capsule Graph Convolutional Based Representation FusionabstractDue to the more robust characteristics compared to unimodal, audio-video multimodal emotion recognition (MER) has attracted a lot of attention. The efficiency of representation fusion algorithm often determines the performance of MER. Although there are many fusion algorithms, information redundancy and information complementarity are usually ignored. In this paper, we propose a novel representation fusion method, Capsule Graph Convolutional Network (CapsGCN). Firstly, after unimodal representation learning, the extracted audio and video representations are distilled by capsule network and encapsulated into multimodal capsules respectively. Multimodal capsules can effectively reduce data redundancy by the dynamic routing algorithm. Secondly, the multimodal capsules with their inter-relations and intra-relations are treated as a graph structure. The graph structure is learned by Graph Convolutional Network (GCN) to get hidden representation which is a good supplement for information complementarity. Finally, the multimodal capsules and hidden relational representation learned by CapsGCN are fed to multihead self-attention to balance the contributions of source representation and relational representation. To verify the performance, visualization of representation, the results of commonly used fusion methods, and ablation studies of the proposed CapsGCN are provided. Our proposed fusion method achieves 80.83% accuracy and 80.23% F1 score on eNTERFACE05’. Jiaxing Liu 0001, Longbiao Wang, Zhilei Liu, Yahui Fu 0001, Lili Guo 0001, Jianwu Dang 0001 |
ICASSP | 7 |
| 2021 | Replay-Attack Detection Using Features With Adaptive Spectro-Temporal ResolutionabstractVariable-resolution processing aims to improve the feature representation ability by enlarging the local discriminative details. In previous anti-spoofing studies, different phones and frequency regions were both proven to have various levels of sensitivity to replay distortion. In this paper, an adaptive spectro-temporal resolution is proposed to obtain the optimal scale in the feature space: the frequency resolution is adaptive to frequency discrimination, while the temporal resolution is adaptive to continuous phones. In the process, phone-frequency F-ratio analysis is applied to investigate the sensitivity divergences to replay distortion among phones and frequencies. Then, attentive filters are designed to automatically adapt to the phone-frequency discrimination. Validation experiments for the proposed method are conducted on two well-acknowledged magnitude and phase features. A comparative analysis on the ASVspoof 2017 V2.0 database demonstrates that our proposed adaptive spectro-temporal resolution method attains considerably higher error reduction rates than the approaches involving the corresponding original resolution features. Meng Liu 0017, Longbiao Wang, Kong-Aik Lee, Xuanda Chen, Jianwu Dang 0001 |
ICASSP | 5 |
| 2021 | Meta-Learning for Cross-Channel Speaker VerificationabstractAutomatic speaker verification (ASV) has been successfully deployed for identity recognition. With increasing use of ASV technology in real-world applications, channel mismatch caused by the recording devices and environments severely degrade its performance, especially in the case of unseen channels. To this end, we propose a meta speaker embedding network (MSEN) via meta-learning to generate channel-invariant utterance embeddings. Specifically, we optimize the differences between the embeddings of a support set and a query set in order to learn a channel-invariant embedding space for utterances. Furthermore, we incorporate distribution optimization (DO) to stabilize the performance of MSEN. To quantitatively measure the effect of MSEN on unseen channels, we specially design the generalized cross-channel (GCC) evaluation. The experimental results on the HI-MIA corpus demonstrate that the proposed MSEN reduce considerably the impact of channel mismatch, while significantly outperforms other state-of-the-art methods. Hanyi Zhang, Longbiao Wang, Kong-Aik Lee, Meng Liu 0017, Jianwu Dang 0001 |
ICASSP | 5 |
| 2021 | CONSK-GCN: Conversational Semantic- and Knowledge-Oriented Graph Convolutional Network for Multimodal Emotion RecognitionabstractEmotion recognition in conversations (ERC) has received significant attention in recent years due to its widespread applications in diverse areas, such as social media, health care, and artificial intelligence interactions. However, different from nonconversational text, it is particularly challenging to model the effective context-aware dependence for the task of ERC. To address this problem, we propose a new Conversational Semantic- and Knowledge-oriented Graph Convolutional Network (ConSK-GCN) approach that leverages both semantic dependence and commonsense knowledge. First, we construct the contextual inter-interaction and intradependence of the interlocutors via a conversational graph-based convolutional network based on multimodal representations. Second, we incorporate commonsense knowledge to guide ConSK-GCN to model the semantic-sensitive and knowledge-sensitive contextual dependence. The results of extensive experiments show that the proposed method outperforms the current state of the art on the IEMOCAP dataset. Yahui Fu 0001, Shogo Okada, Longbiao Wang, Lili Guo 0001, Yaodong Song, Jiaxing Liu 0001, Jianwu Dang 0001 |
ICME | 7 |
| 2021 | Exploring Effective Speech Representation via ASR for High-Quality End-to-End Multispeaker TTS
Longbiao Wang, Sheng Li 0010, Chenchen Ding, Ju Zhang 0001, Jianwu Dang 0001 |
ICONIP (6) | 7 |
| 2021 | Speech Dereverberation Based on Scale-Aware Mean Square Error Loss
Luya Qiang, Meng Ge, Longbiao Wang, Sheng Li 0010, Jianwu Dang 0001 |
ICONIP (5) | 8 |
| 2021 | Simultaneous Progressive Filtering-Based Monaural Speech Enhancement
Longbiao Wang, Luya Qiang, Sheng Li 0010, Meng Ge, Gaoyan Zhang, Jianwu Dang 0001 |
ICONIP (5) | 8 |
| 2021 | Multi-Modal Emotion Recognition Based On deep Learning Of EEG And Audio SignalsabstractAutomatic recognition of human emotional states has attracted many researchers' attention in Human-Computer Interactions and emotional brain-computer interface recently. However, the accuracy of emotion recognition is not satisfying. Considering the advantage of information supplement based on deep learning of multi-modal signals related to emotion, this study proposed a novel emotion recognition architecture to fuse emotional features from brain electroencephalography (EEG) signal and the corresponding audio signal in emotion recognition on DEAP dataset. We used convolutional neural network (CNN) to extract EEG features and bidirectional long short term memory (BiLSTM) neural networks to extract audio features. After that, we combine the multi-modal features into a deep learning architecture to recognize arousal and valence levels. Results showed an improved accuracy compared with previous studies that merely used the EEG signals in both arousal level and valence level, which suggests the effectiveness of our proposed multi-modal fused emotion recognition model. In future work, multi-modal data from nature interaction scenes will be collected and inputted into this architecture to further validate the effectiveness of the method. Gaoyan Zhang, Jianwu Dang 0001, Longbiao Wang, Jianguo Wei |
IJCNN | 3 |
| 2021 | Metric Learning Based Feature Representation with Gated Fusion Model for Speech Emotion Recognition
Yuan Gao 0040, Jiaxing Liu 0001, Longbiao Wang, Jianwu Dang 0001 |
Interspeech | 4 |
| 2021 | TacoLPCNet: Fast and Stable TTS by Conditioning LPCNet on Mel Spectrogram Predictions
Longbiao Wang, Ju Zhang 0001, Shaotong Guo, Yuguang Wang 0003, Jianwu Dang 0001 |
Interspeech | 6 |
| 2021 | Time-Frequency Representation Learning with Graph Convolutional Network for Dialogue-Level Speech Emotion Recognition
Jiaxing Liu 0001, Yaodong Song, Longbiao Wang, Jianwu Dang 0001 |
Interspeech | 4 |
| 2021 | Domain-Specific Multi-Agent Dialog Policy Learning in Multi-Domain Task-Oriented Scenarios
Yuke Si, Longbiao Wang, Jianwu Dang 0001 |
Interspeech | 4 |
| 2021 | Joint Feature Enhancement and Speaker Recognition with Multi-Objective Task-Oriented Network
Longbiao Wang, Kong-Aik Lee, Meng Liu 0017, Jianwu Dang 0001 |
Interspeech | 5 |
| 2021 | A Sentiment Similarity-Oriented Attention Model with Multi-task Learning for Text-Based Emotion Recognition
Yahui Fu 0001, Lili Guo 0001, Longbiao Wang, Zhilei Liu, Jiaxing Liu 0001, Jianwu Dang 0001 |
MMM (1) | 6 |
| 2021 | Exploiting Explicit and Inferred Implicit Personas for Multi-turn Dialogue Generation
Ruifang He, Longbiao Wang, Yuke Si, Jianwu Dang 0001 |
NLPCC (1) | 7 |
| 2021 | Replay attack detection using variable-frequency resolution phase and magnitude features
Meng Liu 0017, Longbiao Wang, Jianwu Dang 0001, Kong-Aik Lee, Seiichi Nakagawa |
Comput. Speech Lang. | 3 |
| 2021 | Multi-resolution modulation-filtered cochleagram feature for LSTM-based dimensional emotion recognition from speech
Zhichao Peng, Jianwu Dang 0001, Masashi Unoki, Masato Akagi |
Neural Networks | 2 |
| 2021 | Efficient learning with augmented spikes: A case study with image classification
Shiming Song 0001, Chenxiang Ma, Junhai Xu, Jianwu Dang 0001, Qiang Yu 0005 |
Neural Networks | 5 |
| 2021 | A Tibetan Language Model That Considers the Relationship Between Suffixes and Functional WordsabstractThe complete semantic representation of a Tibetan sentence is mainly determined by the addition of a specific functional word. The choice of Tibetan functional words is mainly influenced (both explicitly and implicitly) by the sequence of Tibetan suffixes. In this article, we propose an RNN-based Tibetan radical suffix unit (TRSU) to consider this relationship. Specifically, for the Tibetan radical suffix unit-explicit (TRSU-E) method, the fixed suffix in Tibetan is used to determine the virtual functional words. For the Tibetan radical suffix unit-implicit (TRSU-I) method, the decision is assisted by adding a specific suffix. To test the method, we design a standard Tibetan corpus, which consists of different genres. Our experimental results show that the complexity of our method is reduced by up to 22.2% relative to the best baseline. Furthermore, with the hidden semantic information and implicit suffix, TRSU-I outperforms TRSU-E by reducing the perplexity (PPL) by 3%. Moreover, good results are achieved on the English Penn Treebank data set. Kuntharrgyal Khysru, Di Jin 0001, Jianwu Dang 0001 |
IEEE Signal Process. Lett. | 5 |
| 2021 | Robust Detection of Link Communities With Summary Description in Social NetworksabstractCommunity detection has been extensively studied for various applications. Recent research has started to explore node contents to identify semantically meaningful communities. However, links in real networks typically have semantic descriptions and communities of links can better characterize community behaviors than communities of nodes. The second issue in community finding is that the most existing methods assume network topologies and descriptive contents carry the same or compatible information of node group membership, restricting them to one topic per community, which is generally violated in real networks. The third issue is that the existing methods use top ranked words or phrases to label topics when interpreting communities, which is often inadequate for comprehension. To address these issues altogether, we propose a new Bayesian probabilistic approach for modeling real networks and developing an efficient variational algorithm for model inference. Our new method explores the intrinsic correlation between communities and topics to discover link communities and extract semantically meaningful community summaries at the same time. If desired, it is able to derive more than one topical summary per community to provide rich explanations. We present experimental results to show the effectiveness of our new approach and evaluate the method by a case study. Di Jin 0001, Xiaobao Wang, Dongxiao He, Jianwu Dang 0001, Weixiong Zhang |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2021 | Robust Environmental Sound Recognition With Sparse Key-Point Encoding and Efficient Multispike LearningabstractThe capability for environmental sound recognition (ESR) can determine the fitness of individuals in a way to avoid dangers or pursue opportunities when critical sound events occur. It still remains mysterious about the fundamental principles of biological systems that result in such a remarkable ability. Additionally, the practical importance of ESR has attracted an increasing amount of research attention, but the chaotic and nonstationary difficulties continue to make it a challenging task. In this article, we propose a spike-based framework from a more brain-like perspective for the ESR task. Our framework is a unifying system with consistent integration of three major functional parts which are sparse encoding, efficient learning, and robust readout. We first introduce a simple sparse encoding, where key points are used for feature representation, and demonstrate its generalization to both spike- and nonspike-based systems. Then, we evaluate the learning properties of different learning rules in detail with our contributions being added for improvements. Our results highlight the advantages of multispike learning, providing a selection reference for various spike-based developments. Finally, we combine the multispike readout with the other parts to form a system for ESR. Experimental results show that our framework performs the best as compared to other baseline approaches. In addition, we show that our spike-based framework has several advantageous characteristics including early decision making, small dataset acquiring, and ongoing dynamic processing. Our framework is the first attempt to apply the multispike characteristic of nervous neurons to ESR. The outstanding performance of our approach would potentially contribute to draw more research efforts to push the boundaries of spike-based paradigm to a new horizon. Qiang Yu 0005, Yanli Yao, Longbiao Wang, Huajin Tang, Jianwu Dang 0001, Kay Chen Tan |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2020 | Working Memory-Driven Neural Networks with a Novel Knowledge Enhancement Paradigm for Implicit Discourse Relation RecognitionabstractRecognizing implicit discourse relation is a challenging task in discourse analysis, which aims to understand and infer the latent relations between two discourse arguments, such as temporal, comparison. Most of the present models largely focus on learning-based methods that utilize only intra-sentence textual information to identify discourse relations, ignoring the wider contexts beyond the discourse. Moreover, people comprehend the meanings and the relations of discourses, heavily relying on their interconnected working memories (e.g., instant memory, long-term memory). Inspired by this, we propose a Knowledge-Enhanced Attentive Neural Network (KANN) framework to address these issues. Specifically, it establishes a mutual attention matrix to capture the reciprocal information between two arguments, as instant memory. While implicitly stated knowledge in the arguments is retrieved from external knowledge source and encoded as inter-words semantic connection embeddings to further construct knowledge matrix, as long-term memory. We devise a novel paradigm with two ways by the collaboration of the memories to enrich the argument representation: 1) integrating the knowledge matrix into the mutual attention matrix, which implicitly maps knowledge into the process of capturing asymmetric interactions between two discourse arguments; 2) directly concatenating the argument representations and the semantic connection embeddings, which explicitly supplements knowledge to help discourse understanding. The experimental results on the PDTB also show that our KANN model is effective. Fengyu Guo, Ruifang He, Jianwu Dang 0001 |
AAAI | 3 |
| 2020 | Topic Enhanced Sentiment Spreading Model in Social Networks Considering User InterestabstractEmotion is a complex emotional state, which can affect our physiology and psychology and lead to behavior changes. The spreading process of emotions in the text-based social networks is referred to as sentiment spreading. In this paper, we study an interesting problem of sentiment spreading in social networks. In particular, by employing a text-based social network (Twitter) , we try to unveil the correlation between users' sentimental statuses and topic distributions embedded in the tweets, then to automatically learn the influence strength between linked users. Furthermore, we introduce user interest to refine the influence strength. We develop a unified probabilistic framework to formalize the problem into a topic-enhanced sentiment spreading model. The model can predict users' sentimental statuses based on their historical emotional status, topic distributions in tweets and social structures. Experiments on the Twitter dataset show that the proposed model significantly outperforms several alternative methods in predicting users' sentimental status. We also discover an intriguing phenomenon that positive and negative sentiment is more relevant to user interest than neutral ones. Our method offers a new opportunity to understand the underlying mechanism of sentimental spreading in online social networks. Xiaobao Wang, Di Jin 0001, Katarzyna Musial, Jianwu Dang 0001 |
AAAI | 4 |
| 2020 | End-to-End Articulatory Modeling for Dysarthric Articulatory Attribute DetectionabstractIn this study, we focus on detecting articulatory attribute errors for dysarthric patients with cerebral palsy (CP) or amyotrophic lateral sclerosis (ALS). There are two major challenges for this task. The pronunciation of dysarthric patients is unclear and inaccurate, which results in poor performances of traditional automatic speech recognition (ASR) systems and traditional automatic speech attribute transcription (ASAT). In addition, the data is limited because of the difficulty of recording. This study proposes an end-to-end automatic speech attribute transcription (E2E-ASAT) method for detecting articulatory attribute errors more precisely. To use the limited data more effectively, the parameters of the acoustic model are refactored into two layers and only one layer is retrained. Our proposed method showed good performances in both ASR and articulatory attribute detection. Our system has a potential as a rehabilitation tool. Yuqin Lin, Longbiao Wang, Jianwu Dang 0001, Sheng Li 0010, Chenchen Ding |
ICASSP | 3 |
| 2020 | Speech Emotion Recognition with Local-Global Aware Deep Representation LearningabstractConvolutional neural network (CNN) based deep representation learning methods for speech emotion recognition (SER) have demonstrated great success. The basic design of CNN restricts the ability to model only local information well. Capsule network (CapsNet) can overcome the shortages of CNNs to capture the shallow global features from the spectrogram, although CapsNet cannot learn the local and deep global information. In this paper, we propose a local-global aware deep representation learning system that mainly includes two modules. One module contains a multi-scale CNN, time- frequency CNN (TFCNN) to learn the local representation. In the other module, we introduce a structure with dense connections of multiple blocks to learn shallow and deep global information. Every block in this structure is a complete CapsNet improved by a new routing algorithm. The local and global representations are fed to the classifier and achieve an absolute increase of at least 4.25% than benchmarks on IEMOCAP. Jiaxing Liu 0001, Zhilei Liu, Longbiao Wang, Lili Guo 0001, Jianwu Dang 0001 |
ICASSP | 5 |
| 2020 | Spectrograms Fusion with Minimum Difference Masks Estimation for Monaural Speech DereverberationabstractSpectrograms fusion is an effective method for incorporating complementary speech dereverberation systems. Previous linear spectrograms fusion by averaging multiple spectrograms shows outstanding performance. However, various systems with different features cannot apply this simple method. In this study, we design the minimum difference masks (MDMs) to classify the time-frequency (T-F) bins in spectrograms according to the nearest distances from labels. Then, we propose a two-stage nonlinear spectrograms fusion system for speech dereverberation. First, we conduct a multitarget learning-based speech dereverberation front-end model to get spectrograms simultaneously. Then, MDMs are estimated to take the best parts of different spectrograms. We are using spectrograms in the first stage and MDMs in the second stage to recombine T-F bins. The experiments on the REVERB challenge show that a strong feature complementarity between spectrograms and MDMs. Moreover, the proposed framework can consistently and significantly improve PESQ and SRMR, both real and simulated data, e.g., an average PESQ gain of 0.1 in all simulated data and an average SRMR gain of 1.22 in all real data. Longbiao Wang, Meng Ge, Sheng Li 0010, Jianwu Dang 0001 |
ICASSP | 5 |
| 2020 | A Hierarchical Model for Dialog Act Recognition Considering Acoustic and Lexical Context InformationabstractDialog act recognition (DAR) is important to capture speakers' intention in a dialog system. Traditional methods commonly use the lexical information from transcripts, acoustic information from speech, and dialog context information to do DAR. However, in these methods, textual context information may be considered, whereas acoustic context information is ignored, which leads to ambiguity in certain DAs especially in Mandarin. To solve the problem, we propose a hierarchical model for DAR considering context information of both lexical and acoustic prosody. The experimental results on a Mandarin dialog corpus demonstrate that the contextual-acoustic information is helpful for recognizing DAs. The contextually specific prosodies involved in the utterances such as the echo question and open-end question are beneficial to identify the users' intention. We also investigate the effect of the context length on the DAR. The proper context length is approximately equal to the length of the entire subtopics. Yuke Si, Longbiao Wang, Jianwu Dang 0001, Mengfei Wu |
ICASSP | 3 |
| 2020 | Integrating Group Homophily and Individual Personality of Topics Can Better Model Network CommunitiesabstractCommunity detection is an important research field in the understanding of networks. The definition of network communities focuses on denser intracommunity links and sparpser intercommunity links. It cannot explain the fundamental generation mechanisms of the two types of links, which is challenging to reveal. Unfortunately, none of existing works can solve this challenge which is important for accurately modeling community structures. This paper investigates a typical category of networks which possess contents on links. Based on analyses of real networks, we get an observation that nodes with distinctive personality regarding content topics are more active across communities, while nodes without it are more active inside a community, behaving in a similar way known as homophily. This observation provides clues to the generation of intracommunity and intercommunity links. Based on above observation, this paper proposes a novel generative community detection model called GHIPT (Group Homophily and Individual Personality of Topics) by integrating group homophily and individual personality of topics. Besides deriving more precise community results by accurately modeling intracommunity and intercommunity links, GHIPT is able to identify those nodes with distinctive personality who are more willing to interact with others from different communities. It further validates that they change their community memberships more frequently. GHIPT is evaluated on two real networks, i.e., Reddit and DBLP. Experimental results show that it outperforms all the state-of-the-art baselines. In addition to case studies on above two datasets, a case study on COVID-19 dataset provides new insights to support the ongoing fight against COVID-19 pandemic. Yingkui Wang, Di Jin 0001, Carl Yang 0001, Jianwu Dang 0001 |
ICDM | 4 |
| 2020 | Investigation of Effectively Synthesizing Code-Switched Speech Using Highly Imbalanced Mix-Lingual Data
Shaotong Guo, Longbiao Wang, Sheng Li 0010, Ju Zhang 0001, Yuguang Wang 0003, Jianwu Dang 0001, Kiyoshi Honda |
ICONIP (1) | 7 |
| 2020 | Adversarial Shared-Private Attention Network for Joint Slot Filling and Intent Detection
Mengfei Wu, Longbiao Wang, Yuke Si, Jianwu Dang 0001 |
ICONIP (4) | 4 |
| 2020 | Hierarchical Interactive Matching Network for Multi-turn Response Selection in Retrieval-Based Chatbots
Ruifang He, Longbiao Wang, Jianwu Dang 0001 |
ICONIP (1) | 5 |
| 2020 | Deep Discriminative Embedding with Ranked Weight for Speaker Verification
Dao Zhou, Longbiao Wang, Kong-Aik Lee, Meng Liu 0017, Jianwu Dang 0001 |
ICONIP (5) | 5 |
| 2020 | SpEx+: A Complete Time Domain Speaker Extraction NetworkabstractSpeaker extraction aims to extract the target speech signal from a multi-talker environment given a target speaker's reference speech.We recently proposed a time-domain solution, SpEx, that avoids the phase estimation in frequency-domain approaches.Unfortunately, SpEx is not fully a time-domain solution since it performs time-domain speech encoding for speaker extraction, while taking frequency-domain speaker embedding as the reference.The size of the analysis window for timedomain and the size for frequency-domain input are also different.Such mismatch has an adverse effect on the system performance.To eliminate such mismatch, we propose a complete time-domain speaker extraction solution, that is called SpEx+.Specifically, we tie the weights of two identical speech encoder networks, one for the encoder-extractor-decoder pipeline, another as part of the speaker encoder.Experiments show that the SpEx+ achieves 0.8dB and 2.1dB SDR improvement over the state-of-the-art SpEx baseline, under different and same gender conditions on WSJ0-2mix-extr database respectively. Meng Ge, Chenglin Xu, Longbiao Wang, Chng Eng Siong, Jianwu Dang 0001, Haizhou Li 0001 |
INTERSPEECH | 5 |
| 2020 | Segment-Level Effects of Gender, Nationality and Emotion Information on Text-Independent Speaker VerificationabstractSpeaker embeddings extracted from neural network (NN) achieve excellent performance on general speaker verification (SV) missions. Most current SV systems use only speaker labels. Therefore, the interaction between different types of domain information decrease the prediction accuracy of SV. To overcome this weakness and improve SV performance, four effective SV systems were proposed by using gender, nationality, and emotion information to add more constraints in the NN training stage. More specifically, multitask learning-based systems which including multitask gender (MTG), multitask nationality (MTN) and multitask gender and nationality (MTGN) were used to enhance gender and nationality information learning. Domain adversarial training-based system which including emotion domain adversarial training (EDAT) was used to suppress different emotions information learning. Experimental results indicate that encouraging gender and nationality information and suppressing emotion information learning improve the performance of SV. In the end, our proposed systems achieved 16.4 and 22.9% relative improvements in the equal error rate for MTL- and DAT-based systems, respectively. Kai Li 0018, Masato Akagi, Jianwu Dang 0001 |
INTERSPEECH | 4 |
| 2020 | Staged Knowledge Distillation for End-to-End Dysarthric Speech Recognition and Speech Attribute Transcription
Yuqin Lin, Longbiao Wang, Sheng Li 0010, Jianwu Dang 0001, Chenchen Ding |
INTERSPEECH | 4 |
| 2020 | Temporal Attention Convolutional Network for Speech Emotion Recognition with Latent Representation
Jiaxing Liu 0001, Zhilei Liu, Longbiao Wang, Yuan Gao 0040, Lili Guo 0001, Jianwu Dang 0001 |
INTERSPEECH | 6 |
| 2020 | Dimensional Emotion Prediction Based on Interactive Context in Conversation
Sixia Li, Jianwu Dang 0001 |
INTERSPEECH | 3 |
| 2020 | Singing Voice Extraction with Attention-Based Spectrograms Fusion
Longbiao Wang, Sheng Li 0010, Chenchen Ding, Meng Ge, Jianwu Dang 0001, Hiroshi Seki |
INTERSPEECH | 7 |
| 2020 | EEG-Based Short-Time Auditory Attention Detection Using Multi-Task Deep Learning
Gaoyan Zhang, Jianwu Dang 0001, Di Zhou 0008, Longbiao Wang |
INTERSPEECH | 3 |
| 2020 | Cortical Oscillatory Hierarchy for Natural Sentence Processing
Jianwu Dang 0001, Gaoyan Zhang, Masashi Unoki |
INTERSPEECH | 2 |
| 2020 | Dynamic Margin Softmax Loss for Speaker Verification
Dao Zhou, Longbiao Wang, Kong-Aik Lee, Meng Liu 0017, Jianwu Dang 0001, Jianguo Wei |
INTERSPEECH | 6 |
| 2020 | Neural Entrainment to Natural Speech Envelope Based on Subject Aligned EEG Signals
Di Zhou 0008, Gaoyan Zhang, Jianwu Dang 0001 |
INTERSPEECH | 3 |
| 2020 | Speaker-Aware Speech Emotion Recognition by Fusing Amplitude and Phase Information
Lili Guo 0001, Longbiao Wang, Jianwu Dang 0001, Zhilei Liu, Haotian Guan |
MMM (1) | 3 |
| 2020 | Relation Modeling with Graph Convolutional Networks for Facial Action Unit Detection
Zhilei Liu, Jiahui Dong, Cuicui Zhang, Longbiao Wang, Jianwu Dang 0001 |
MMM (2) | 5 |
| 2019 | Community Detection in Social Networks Considering Topic CorrelationsabstractNetwork contents including node contents and edge contents can be utilized for community detection in social networks. Thus, the topic of each community can be extracted as its semantic information. A plethora of models integrating topic model and network topologies have been proposed. However, a key problem has not been resolved that is the semantic division of a community. Since the definition of community is based on topology, a community might involve several topics. To ach Yingkui Wang, Di Jin 0001, Katarzyna Musial, Jianwu Dang 0001 |
AAAI | 4 |
| 2019 | Emotional Contagion-Based Social Sentiment Mining in Social Networks by Introducing Network CommunitiesabstractThe rapid development of social media services has facilitated the communication of opinions through online news, blogs, microblogs, instant-messages, and so on. This article concentrates on the mining of readers' social sentiments evoked by social media materials. Existing methods are only applicable to a minority of social media like news portals with emotional voting information, while ignore the emotional contagion between writers and readers. However, incorporating such factors is challenging since the learned hidden variables would be very fuzzy (because of the short and noisy text in social networks). In this paper, we try to solve this problem by introducing a high-order network structure, i.e. communities. We first propose a new generative model called Community-Enhanced Social Sentiment Mining (CESSM), which 1) considers the emotional contagion between writers and readers to capture precise social sentiment, and 2) incorporates network communities to capture coherent topics. We then derive an inference algorithm based on Gibbs sampling. Empirical results show that, CESSM achieves significantly superior performance against the state-of-the-art techniques for text sentiment classification and interestingness in social sentiment mining. Xiaobao Wang, Di Jin 0001, Mengquan Liu, Dongxiao He, Katarzyna Musial, Jianwu Dang 0001 |
CIKM | 6 |
| 2019 | Robust Sound Event Classification with Local Time-Frequency Information and Convolutional Neural Networks
Yanli Yao, Qiang Yu 0005, Longbiao Wang, Jianwu Dang 0001 |
ICANN (4) | 4 |
| 2019 | A Multi-spike Approach for Robust Sound RecognitionabstractThe extraordinary performance of the brain on various cognitive tasks motivates the design of a biologically plausible system for the challenging task of environmental sound recognition. In this paper, we propose a novel approach based on multi-spike learning and key-point encoding. Our encoding extracts local temporal and spectral information from the sound and converts it into spatiotemporal spike pattern, which is further learned by the following spiking neural networks. Our experiments demonstrate the robustness and effectiveness of our approach across a variety of noise conditions, outperforming other conventional baseline methods in both mismatched and multi-condition scenarios. Qiang Yu 0005, Yanli Yao, Longbiao Wang, Huajin Tang, Jianwu Dang 0001 |
ICASSP | 5 |
| 2019 | Replay Attack Detection Using Magnitude and Phase Information with Attention-based Adaptive FiltersabstractAutomatic Speech Verification (ASV) systems are highly vulnerable to spoofing attacks, and replay attack poses the greatest threat among various spoofing attacks. In this paper, we propose a novel multi-channel feature extraction method with attention-based adaptive filters (AAF). Original phase information, discarded by conventional feature extraction techniques after Fast Fourier Transform (FFT), is promising in distinguishing genuine from replay spoofed speech. Accordingly, phase and magnitude information are respectively extracted as phase channel and magnitude channel complementary features in our system. First, we make discriminative ability analysis on full frequency bands with F-ratio methods. Then attention-based adaptive filters are implemented to maximize capturing of high discriminative information on frequency bands, and the results on ASVspoof 2017 challenge indicate that our proposed approach achieved relative error reduction rates of 78.7% and 59.8% on development and evaluation dataset than the baseline method. Meng Liu 0017, Longbiao Wang, Jianwu Dang 0001, Seiichi Nakagawa, Haotian Guan, Xiangang Li |
ICASSP | 3 |
| 2019 | NVSRN: A Neural Variational Scaling Reasoning Network for Initiative Response GenerationabstractOpen-domain multi-turn dialogue systems are booming in human-machine interactions, which encourage to chat actively and freely in an intelligent natural way. Previous generative conversational models usually employ a single and deterministic encoder-decoder framework to model the semantic consistency between the context and corresponding response. However, they neglect the various dialog patterns (we denote the regularity of topic shifting as dialog pattern) in the conversations, leading to uninformative, non-initiative yet plausible responses. Although the existing variational methods have improved the response diversity to some extent by introducing a global variability into the generative process, they fail to simulate the transfer between topics with directional information due to the weak interpretability of the Gaussian-distributed latent variables. In this paper, we propose a novel Neural Variational Scaling Reasoning Network (NVSRN) for initiative response generation. To this end, our approach has two core ingredients: neural dialog pattern reasoner (reasoner) and topic scaling mechanism. Specifically, inspired by the advantage of von Mises-Fisher (vMF) distribution modeling the directional data (e.g., the topic transfer state), we employ it as the latent space of the reasoner to explore the regularity of topic shifting, which is then used to reason the topic of response. Based on this, a topic scaling mechanism is designed to control the transfer degree of topic in the response generator. The experimental results on two large dialog datasets demonstrate that the proposed model outperforms state-of-the-art baselines. The human evaluation shows the proposed model can produce more informative and initiative responses actively. Jinxin Chang, Ruifang He, Haiyang Xu 0001, Longbiao Wang, Xiangang Li, Jianwu Dang 0001 |
ICDM | 7 |
| 2019 | Interaction Process Label Recognition in Group DiscussionabstractIn qualifying and analyzing the performance of group interaction, interaction processing analysis (IPA) defined by Bale is considered a useful approach. IPA is a system for labeling a total of 12 interaction categories for the interaction process. Automatic IPA can manually encompass the gap in spending manpower and can efficiently qualify group performance. In this paper, we present computational interaction processing analysis by developing a model to recognize categories of IPA. We extract both verbal features and nonverbal features for IPA category recognition modeling with SVM, RF, DNN and LSTM machine learning algorithms and analyze the contribution of multimodal features and unimodal features for the total data and each label. We also investigate the effect of context information by training sequences with different lengths with an LSTM and evaluating them. The results show that multimodal features achieve the best performance with an F1 score of 0.601 for the recognition of 12 IPA categories using the total data. Multimodal features are better than the unimodal features for the total data and most labels. The results of investigating context information show that a suitable length of sequence enables a longer sequence to achieve the best F1 score of 0.602 and a better performance for recognition. Sixia Li, Shogo Okada, Jianwu Dang 0001 |
ICMI | 3 |
| 2019 | A Fast Convolutional Self-attention Based Speech Dereverberation Method for Robust Speech Recognition
Meng Ge, Longbiao Wang, Jianwu Dang 0001 |
ICONIP (3) | 4 |
| 2019 | Time-Frequency Deep Representation Learning for Speech Emotion Recognition Integrating Self-attention
Jiaxing Liu 0001, Zhilei Liu, Longbiao Wang, Lili Guo 0001, Jianwu Dang 0001 |
ICONIP (4) | 5 |
| 2019 | A Spiking Neural Network with Distributed Keypoint Encoding for Robust Sound RecognitionabstractCompared to traditional artificial neural networks, spiking neural networks (SNNs) operate on an additional dimension of time which makes them more suitable for processing sound signals. However, two of the major challenges in sound recognition with SNNs are neural encoding and learning which demand more research efforts. In this paper, we propose a novel method by combining an improved local time-frequency encoding using key-points detection and biologically plausible tempotron spike learning for robust sound recognition. In the neural encoding part, local energy peaks, called key-points, are firstly extracted from local temporal and spectral regions in the spectrogram. The extracted key-points in each frequency channel are then distributed to multiple sub-channels according to their energy amplitudes with their temporal positions being retained. The resulted spatio-temporal spike patterns are then used as the inputs for spiking neural networks to learn and classify patterns of different categories. We use the RWCP database to evaluate the performance of our proposed system in mismatched environments. Our experimental results highlight that our proposed system, namely DKP-SNN, is effective and reliable for robust sound recognition, resulting in an improved recognition performance as compared to baseline methods. Yanli Yao, Qiang Yu 0005, Longbiao Wang, Jianwu Dang 0001 |
IJCNN | 4 |
| 2019 | Acoustic and Articulatory Study of Ewe Vowels: A Comparative Study of Male and Female
Kowovi Comivi Alowonou, Jianguo Wei, Wenhuan Lu, Kiyoshi Honda, Jianwu Dang 0001 |
INTERSPEECH | 6 |
| 2019 | Environment-Dependent Attention-Driven Recurrent Convolutional Neural Network for Robust Speech Enhancement
Meng Ge, Longbiao Wang, Jianwu Dang 0001, Xiangang Li |
INTERSPEECH | 5 |
| 2019 | CNN-BLSTM Based Question Detection from Dialogs Considering Phase and Context Information
Yuke Si, Longbiao Wang, Jianwu Dang 0001, Mengfei Wu |
INTERSPEECH | 3 |
| 2019 | Combination of links and node contents for community discovery using a graph regularization approach
Jinxin Cao, Hongcui Wang, Di Jin 0001, Jianwu Dang 0001 |
Future Gener. Comput. Syst. | 4 |
| 2019 | Story co-segmentation of Chinese broadcast news using weakly-supervised semantic similarity
Wei Feng 0005, Xuecheng Nie, Yujun Zhang 0002, Jianwu Dang 0001 |
Neurocomputing | 5 |
| 2018 | Robust Detection of Link Communities in Large Social Networks by Exploiting Link SemanticsabstractCommunity detection has been extensively studied for various applications, focusing primarily on network topologies. Recent research has started to explore node contents to identify semantically meaningful communities and interpret their structures using selected words. However, links in real networks typically have semantic descriptions, e.g., comments and emails in social media, supporting the notion of communities of links. Indeed, communities of links can better describe multiple roles that nodes may play and provide a richer characterization of community behaviors than communities of nodes. The second issue in community finding is that most existing methods assume network topologies and descriptive contents to be consistent and to carry the compatible information of node group membership, which is generally violated in real networks. These methods are also restricted to interpret one community with one topic. The third problem is that the existing methods have used top ranked words or phrases to label topics when interpreting communities. However, it is often difficult to comprehend the derived topics using words or phrases, which may be irrelevant. To address these issues altogether, we propose a new unified probabilistic model that can be learned by a dual nested expectation-maximization algorithm. Our new method explores the intrinsic correlation between communities and topics to discover link communities robustly and extract adequate community summaries in sentences instead of words for topic labeling at the same time. It is able to derive more than one topical summary per community to provide rich explanations. We present experimental results to show the effectiveness of our new approach, and evaluate the quality of the results by a case study. Di Jin 0001, Xiaobao Wang, Ruifang He, Dongxiao He, Jianwu Dang 0001, Weixiong Zhang |
AAAI | 5 |
| 2018 | Implicit Discourse Relation Recognition using Neural Tensor Network with Interactive Attention and Sparse LearningabstractImplicit discourse relation recognition aims to understand and annotate the latent relations between two discourse arguments, such as temporal, comparison, etc. Most previous methods encode two discourse arguments separately, the ones considering pair specific clues ignore the bidirectional interactions between two arguments and the sparsity of pair patterns. In this paper, we propose a novel neural Tensor network framework with Interactive Attention and Sparse Learning (TIASL) for implicit discourse relation recognition. (1) We mine the most correlated word pairs from two discourse arguments to model pair specific clues, and integrate them as interactive attention into argument representations produced by the bidirectional long short-term memory network. Meanwhile, (2) the neural tensor network with sparse constraint is proposed to explore the deeper and the more important pair patterns so as to fully recognize discourse relations. The experimental results on PDTB show that our proposed TIASL framework is effective. Fengyu Guo, Ruifang He, Di Jin 0001, Jianwu Dang 0001, Longbiao Wang, Xiangang Li |
COLING | 4 |
| 2018 | Interaction-Aware Topic Model for Microblog Conversations through Network Embedding and User AttentionabstractTraditional topic models are insufficient for topic extraction in social media. The existing methods only consider text information or simultaneously model the posts and the static characteristics of social media. They ignore that one discusses diverse topics when dynamically interacting with different people. Moreover, people who talk about the same topic have different effects on the topic. In this paper, we propose an Interaction-Aware Topic Model (IATM) for microblog conversations by integrating network embedding and user attention. A conversation network linking users based on reposting and replying relationship is constructed to mine the dynamic user behaviours. We model dynamic interactions and user attention so as to learn interaction-aware edge embeddings with social context. Then they are incorporated into neural variational inference for generating the more consistent topics. The experiments on three real-world datasets show that our proposed model is effective. Ruifang He, Di Jin 0001, Longbiao Wang, Jianwu Dang 0001, Xiangang Li |
COLING | 5 |
| 2018 | Gender-Aware CNN-BLSTM for Speech Emotion Recognition
Linjuan Zhang, Longbiao Wang, Jianwu Dang 0001, Lili Guo 0001, Qiang Yu 0005 |
ICANN (1) | 3 |
| 2018 | A Feature Fusion Method Based on Extreme Learning Machine for Speech Emotion RecognitionabstractSpeech emotion recognition is important to understand users' intention in human-computer interaction. However, it is a challenging task partly because we cannot clearly know which feature and model are effective to distinguish emotions. Previous studies utilize convolutional neural network (CNN) directly on spectrograms to extract features, and bidirectional long short term memory (BLSTM) is the state-of-the-art model. However, there are two problems of CNN-BLSTM. Firstly, it doesn't utilize heuristic features based on priori knowledge. Secondly, BLSTM has a complex structure and high complexity in training. To address the first problem, we propose a feature fusion method that combines CNN-based features and heuristic-based discriminative features which are extracted from heuristic features using deep neural network (DNN). In addition, we utilize extreme learning machine (ELM) instead of BLSTM to solve the second problem. The experiments conducted on EmoDB and our method leads to 40% relative error reduction in Fl-score compared to CNN-BLSTM. Lili Guo 0001, Longbiao Wang, Jianwu Dang 0001, Linjuan Zhang, Haotian Guan |
ICASSP | 3 |
| 2018 | Auditory-Inspired End-to-End Speech Emotion Recognition Using 3D Convolutional Recurrent Neural Networks Based on Spectral-Temporal RepresentationabstractThe human auditory system has far superior emotion recognition abilities compared with recent speech emotion recognition systems, so research has focused on designing emotion recognition systems by mimicking the human auditory system. Psychoacoustic and physiological studies indicate that the human auditory system decomposes speech signals into acoustic and modulation frequency components, and further extracts temporal modulation cues. Speech emotional states are perceived from temporal modulation cues using the spectral and temporal receptive field of the neuron. This paper proposes an emotion recognition system in an end-to-end manner using three-dimensional convolutional recurrent neural networks (3D-CRNNs) based on temporal modulation cues. Temporal modulation cues contain four-dimensional spectral-temporal (ST) integration representations directly as the input of 3D-CRNNs. The convolutional layer is used to extract high-level multiscale ST representations, and the recurrent layer is used to extract long-term dependency for emotion recognition. The proposed method was verified on the IEMOCAP database. The results show that our proposed method can exceed the recognition accuracy compared to that of the state-of-the-art systems. Zhichao Peng, Masashi Unoki, Jianwu Dang 0001, Masato Akagi |
ICME | 4 |
| 2018 | Efficient Multi-spike Learning with Tempotron-Like LTP and PSD-Like LTD
Qiang Yu 0005, Longbiao Wang, Jianwu Dang 0001 |
ICONIP (1) | 3 |
| 2018 | Convolutional Neural Network with Spectrogram and Perceptual Features for Speech Emotion Recognition
Linjuan Zhang, Longbiao Wang, Jianwu Dang 0001, Lili Guo 0001, Haotian Guan |
ICONIP (4) | 3 |
| 2018 | Speech Emotion Recognition by Combining Amplitude and Phase Information Using Convolutional Neural Network
Lili Guo 0001, Longbiao Wang, Jianwu Dang 0001, Linjuan Zhang, Haotian Guan, Xiangang Li |
INTERSPEECH | 3 |
| 2018 | Multiple Phase Information Combination for Replay Attacks Detection
Dongbo Li, Longbiao Wang, Jianwu Dang 0001, Meng Liu 0017, Zeyan Oo, Seiichi Nakagawa, Haotian Guan, Xiangang Li |
INTERSPEECH | 3 |
| 2018 | Revealing Spatiotemporal Brain Dynamics of Speech Production Based on EEG and Eye Movement
Gaoyan Zhang, Jianwu Dang 0001, Minbo Chen, YingjianFu, Longbiao Wang |
INTERSPEECH | 4 |
| 2018 | Autoencoder Based Community Detection with Adaptive Integration of Network Topology and Node Contents
Jinxin Cao, Di Jin 0001, Jianwu Dang 0001 |
KSEM (2) | 3 |
| 2018 | Incorporating network structure with node contents for community detection on large networks using deep learning
Jinxin Cao, Di Jin 0001, Liang Yang 0002, Jianwu Dang 0001 |
Neurocomputing | 4 |
| 2018 | Unsupervised measure of Chinese lexical semantic similarity using correlated graph model for news story segmentation
Wei Feng 0005, Xuecheng Nie, Yujun Zhang 0002, Lei Xie 0001, Jianwu Dang 0001 |
Neurocomputing | 5 |
| 2018 | Phase and reverberation aware DNN for distant-talking speech enhancement
Zeyan Oo, Longbiao Wang, Khomdet Phapatanaburi, Masahiro Iwahashi, Seiichi Nakagawa, Jianwu Dang 0001 |
Multim. Tools Appl. | 6 |
| 2017 | Phase aware deep neural network for noise robust voice activity detectionabstractPhase information is ignored for almost all voice activity detection (VAD). To exploit full information in the original signal, this paper proposes a deep neural network (DNN) using magnitude and phase information (that is, phase aware DNN) to achieve better VAD performance. Mel-frequency cepstral coefficient (MFCC), power-normalized cepstral coefficients (PNCC), instantaneous frequency derivative (IF), baseband phase difference (BPD) and modified group delay cepstral coefficient (MGDCC) are used as magnitude and phase information. The proposed methods were evaluated using CENSREC-1-C database under noise condition. The results show that the phase aware DNN significantly outperforms the DNN using only magnitude information. For DNN-based classifier, the equal error rate (EER) was reduced from 23.70% of MFCC, to 20.43% of joint dual magnitude and single phase features (augmenting PNCC, MGDCC and IF), to 19.92% of joint dual phase and single magnitude feature features (augmenting PNCC, MGDCC and BPD). By combining joint dual magnitude and single phase features with joint dual phase and single magnitude features, the EER was reduced to 19.44%. Longbiao Wang, Khomdet Phapatanaburi, Zeyan Oo, Seiichi Nakagawa, Masahiro Iwahashi, Jianwu Dang 0001 |
ICME | 6 |
| 2017 | Stochastic Sequential Minimal Optimization for Large-Scale Linear SVM
Shili Peng, Qinghua Hu, Jianwu Dang 0001, Zhichao Peng |
ICONIP (1) | 3 |
| 2017 | Exploiting the Tibetan Radicals in Recurrent Neural Network for Low-Resource Language Models
Tongtong Shen, Longbiao Wang, Xie Chen 0001, Kuntharrgyal Khysru, Jianwu Dang 0001 |
ICONIP (2) | 5 |
| 2017 | Neuronal Classifier for both Rate and Timing-Based Spike Patterns
Qiang Yu 0005, Longbiao Wang, Jianwu Dang 0001 |
ICONIP (6) | 3 |
| 2017 | Phonemic Restoration Based on the Movement Continuity of Articulation
Cenxi Zhao, Longbiao Wang, Jianwu Dang 0001 |
ICONIP (6) | 3 |
| 2017 | Identification of Generalized Communities with Semantics in Networks with ContentabstractDiscovery of communities in networks is a fundamental data analysis task. Recently, researchers have tried to improve its performance by exploiting node contents, and further interpret the communities using the derived semantics. However, the existing methods typically assume that the communities are assortative (i.e. members of each group are mostly connected to other members of the same group), and are unable to find the generalized community structure, e.g. structures with either assortative or disassortative communities (i.e. vertices of the same group have most of their connections outside their group), or a combination. In addition, these methods often assume that the network topology and node contents share the same group memberships, and thus cannot perform well when the contents mismatch with network structure. Also, they are limited to using only one topic to interpret each community. To address these two issues, we propose a new generative probabilistic model which is learned by using a nested expectation-maximization algorithm. It describes the generalized communities (based on network) and the content clusters (based on contents) separately, and further explores and models their correlation to improve as much as possible each of the communities and clusters based on the other. By depicting and utilizing this correlation, our model is not only robust with respect to the above problems, but is also able to interpret each community using more than one topic, which provides richer explanations. We validate the robustness of this proposed new approach on an artificial benchmark, and test its interpretability using a case study analysis. We finally show its definite superiority for community detection by comparing with seven state-of-the-art algorithms on eight real networks. Di Jin 0001, Xiaobao Wang, Dongxiao He, Wenhuan Lu, Françoise Fogelman-Soulié, Jianwu Dang 0001 |
ICTAI | 6 |
| 2017 | A Neuro-Experimental Evidence for the Motor Theory of Speech Perception
Jianwu Dang 0001, Gaoyan Zhang |
INTERSPEECH | 2 |
| 2017 | Twitter summarization with social-temporal context
Ruifang He, Yang Liu 0124, Guangchuan Yu, Jiliang Tang, Qinghua Hu, Jianwu Dang 0001 |
World Wide Web | 6 |
| 2016 | Detect Overlapping Communities via Ranking Node PopularitiesabstractDetection of overlapping communities has drawn much attention lately as they are essential properties of real complex networks. Despite its influence and popularity, the well studied and widely adopted stochastic model has not been made effective for finding overlapping communities. Here we extend the stochastic model method to detection of overlapping communities with the virtue of autonomous determination of the number of communities. Our approach hinges upon the idea of ranking node popularities within communities and using a Bayesian method to shrink communities to optimize an objective function based on the stochastic generative model. We evaluated the novel approach, showing its superior performance over five state-of-the-art methods, on large real networks and synthetic networks with ground-truths of overlapping communities. Di Jin 0001, Hongcui Wang, Jianwu Dang 0001, Dongxiao He, Weixiong Zhang |
AAAI | 3 |
| 2016 | Investigations into vowel and consonant structures in articulatory and auditory spaces using Laplacian eigenmapsabstractMany studies have investigated the relationship between the articulatory and auditory features for isolated speech sound and vowels. For fully understanding the mechanisms of speech production and perception, it is necessary to investigate the consonants in the same way. For this reason, in this study, we investigate the manifolds of vowels and consonants out of Japanese reading speech using Laplacian eigenmaps. We constructed uniform articulatory and auditory spaces based on the vowels and consonants to investigate their manifolds. It is found that the distribution of consonants in articulatory space could be classified into labial and lingual groups which reflected their articulatory properties, while in auditory space their distribution was clustered according to voiced and unvoiced, plosive and fricative properties. In vowel-consonant acoustic space, the consonants distributed as a hoe-like shape, with voiced consonants located on the blade of the hoe and fused with vowels. We defined average correlation coefficients to measure the similarity of manifold between three speakers. The results indicated that the vowel/consonant structures had high consistency among the three speakers. Jianwu Dang 0001, Shengbei Wang, Masashi Unoki |
ICASSP | 1 |
| 2016 | Effects of Subglottal-Coupling and Interdental-Space on Formant Trajectories During Front-to-Back Vowel Transitions in Chinese
Shuanglin Fan, Kiyoshi Honda, Jianwu Dang 0001 |
INTERSPEECH | 3 |
| 2016 | A New Model for Acoustic Wave Propagation and Scattering in the Vocal Tract
Jianguo Wei, Wendan Guan, Darcy Qingzhi Hou, Dingyi Pan, Wenhuan Lu, Jianwu Dang 0001 |
INTERSPEECH | 6 |
| 2016 | Audio-visual speech recognition integrating 3D lip information obtained from the Kinect
Jianrong Wang, Ju Zhang 0001, Kiyoshi Honda, Jianguo Wei, Jianwu Dang 0001 |
Multim. Syst. | 5 |
| 2016 | Sketch4Image: a novel framework for sketch-based image retrieval based on product quantization with coding residuals
Yahong Han, Jianwu Dang 0001 |
Multim. Tools Appl. | 3 |
| 2016 | Mapping ultrasound-based articulatory images and vowel sounds with a deep neural network framework
Jianguo Wei, Qiang Fang 0003, Xinyuan Zheng, Wenhuan Lu, Jianwu Dang 0001 |
Multim. Tools Appl. | 6 |
| 2016 | Multi-modal recording and modeling of vocal tract movements
Jianguo Wei, Song Wang 0005, Wenhuan Lu, Darcy Qingzhi Hou, Qiang Fang 0003, Jianwu Dang 0001 |
Multim. Tools Appl. | 6 |
| 2015 | Vocal responses to frequency modulated composite sinewaves via auditory and vibrotactile pathwaysabstractFeedback control mechanisms for speaking have been examined using the transformed auditory feedback (TAF) technique. Previous studies have shown that speakers demonstrate fundamental frequency (F0) changes when they monitor their voice with artificial alterations of F0. However, those studies underestimate the role of vibrotactile information involved in feedback F0 control. This pilot study aims at exploring whether and how vibrotactile information from the larynx influences vowel F0. Participants in our experiment were asked to sustain vowel with their F0 adjusted to composite sinewave stimuli, which were given via auditory and vibrotactile channels using a headset on the ears or a bone-conduction transducer on the larynx. Results revealed the greater compensatory responses to combined vibrotactile-auditory stimuli than to the responses to auditory-only stimuli. The effect of vibrotactile stimuli on feedback F0 adjustment was also observed with the shorter latency of the responses. Kiyoshi Honda, Jianwu Dang 0001, Jianguo Wei |
ICASSP | 3 |
| 2015 | Perception of Mandarin tones by native tibetan speakers
Wenfu Bao, Jianwu Dang 0001, Zhilei Liu |
INTERSPEECH | 3 |
| 2015 | Combined cine- and tagged-MRI for tracking landmarks on the tongue surface
Honghao Bao, Wenhuan Lu, Kiyoshi Honda, Jianguo Wei, Qiang Fang 0003, Jianwu Dang 0001 |
INTERSPEECH | 6 |
| 2015 | Measuring oral and nasal airflow in production of Chinese plosive
Yujie Chi, Kiyoshi Honda, Jianguo Wei, Jianwu Dang 0001 |
INTERSPEECH | 5 |
| 2015 | Community detection with manifold learning on speaker i-vector space for Chinese
Hongcui Wang, Di Jin 0001, Lantian Li, Jianwu Dang 0001 |
INTERSPEECH | 4 |
| 2014 | Image decomposing for inpainting using compressed sensing in DCT domain
Yahong Han, Jianwu Dang 0001 |
Frontiers Comput. Sci. | 3 |
| 2014 | Detection of speaker individual information using a phoneme effect suppression method
Songgun Hyon, Jianwu Dang 0001, Hongcui Wang, Kiyoshi Honda |
Speech Commun. | 2 |
| 2013 | An anisotropic diffusion filter based on multidirectional separability
Jianguo Wei, Xin Wang 0037, Wenhuan Lu, Qiang Fang 0003, Jianwu Dang 0001 |
INTERSPEECH | 6 |
| 2013 | An MRI-based acoustic study of Mandarin vowels
Yuguang Wang 0003, Jianwu Dang 0001, Jianguo Wei, Hongcui Wang, Kiyoshi Honda |
INTERSPEECH | 2 |
| 2012 | Noise estimation using a constrained sequential HMM IN log-spectral domainabstractHow to utilize the time correlation of speech/nonspeech presence is a crucial problem faced by noise estimators. The popular technique of exploiting such correlation is to smooth noisy spectra by using a temporal recursive filter with a time-varying smoothing factor. But this technique cannot warrant the statistical optimality. In theory, hidden Markov model (HMM) is more desirable than this technique. It can give an elaborate description of speech/nonspeech transition. Moreover, some theoretical frameworks, such as maximum likelihood (ML), are available for optimal estimation. This paper presents a constrained sequential HMM to model the time correlation of speech/nonspeech presence of an individual log-power sequence. Its parameter set is on-line adapted to varying signals based on a ML framework. We compared its performance with that of well-established algorithms by speech enhancement experiments. The results confirmed its promising performance. Dongwen Ying, Xugang Lu, Yonghong Yan 0002, Jianwu Dang 0001, Frank K. Soong |
ICASSP | 5 |
| 2012 | A method of speaker identification based on phoneme mean F-ratio contribution
Songgun Hyon, Hongcui Wang, Jianguo Wei, Jianwu Dang 0001 |
INTERSPEECH | 5 |
| 2011 | Voice Activity Detection Based on an Unsupervised Learning FrameworkabstractHow to construct models for speech/nonspeech discrimination is a crucial point for voice activity detectors (VADs). Semi-supervised learning is the most popular way for model construction in conventional VADs. In this correspondence, we propose an unsupervised learning framework to construct statistical models for VAD. This framework is realized by a sequential Gaussian mixture model. It comprises an initialization process and an updating process. At each subband, the GMM is firstly initialized using EM algorithm, and then sequentially updated frame by frame. From the GMM, a self-regulatory threshold for discrimination is derived at each subband. Some constraints are introduced to this GMM for the sake of reliability. For the reason of unsupervised learning, the proposed VAD does not rely on an assumption that the first several frames of an utterance are nonspeech, which is widely used in most VADs. Moreover, the speech presence probability in the time-frequency domain is a byproduct of this VAD. We tested it on speech from TIMIT database and noise from NOISEX-92 database. The evaluations effectively showed its promising performance in comparison with VADs such as ITU G.729B, GSM AMR, and a typical semi-supervised VAD. Dongwen Ying, Yonghong Yan 0002, Jianwu Dang 0001, Frank K. Soong |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2010 | Morphological normalization of vocal tract shapeabstractThe articulatory databases are not utilized so widely as acoustic databases. One of the reasons is the difficulty of reducing morphological variations among subjects. To reduce morphological differences in speech organs among speakers and remain their speech dynamics, this study proposed a framework of normalizing vocal tract by using a Thin-plate spline method. Electromagnetic Midsagittal Articulographic data for three subjects have been used in this research. The template for normalization was obtained by averaging all three subjects' palates and tongue shapes. The landmarks of the template and subjects have been defined according to a gridline system of the vocal tract. The results show that the variances among subjects were reduced 0.8 mm in horizontal and 2.4 mm in vertical direction. The similar vowel structure of pre/post-normalization data indicates that speaker specific characteristics can be maintained by this framework. The effects of the normalization in acoustic space are also investigated by using a physiological articulatory model. Results show that the variations have also been reduced in acoustic space. Jianguo Wei, Jianwu Dang 0001 |
ICASSP | 2 |
| 2010 | Vowel Production Manifold: Intrinsic Factor Analysis of Vowel ArticulationabstractThe manner in which the organization of vowels in a compact space reflects the relationship between their production and perception remains to be clarified in the field of speech science, although it is believed that vowels exist in a compact space with a regular structure rather than as an unstructured blob. In the articulatory domain, some traditional representations such as those based on the results of tongue position analysis and linear factor analysis are used. However, the former partially encodes information on vowel articulation, and the latter only reflects the linear degrees of freedom of vowel articulation. Since nonlinear degrees of freedom exist during the production of vowels, the traditional linear factor analysis is not suitable. In this paper, we proposed the use of Laplacian eigenmaps for analyzing the intrinsic factors affecting vowel articulation and obtain a compact manifold representation of vowels. On the manifold, vowels have distinct cluster positions depending on the similarities in their production and perception. On the basis of vowel articulation, we state that the first dimension of the manifold structure is related to the tongue height position and the second dimension to mouth opening. The third dimension is related to the articulation location of vowels along the vocal tract which is curved along the manifold. A similar topological manifold structure is explored in the vowel acoustic space. Quantitatively, on the basis of the conditional entropy criterion and vowel identification experiments, we confirmed that the analyzed compact manifold structure encodes more information on vowels than do traditional representations. Xugang Lu, Jianwu Dang 0001 |
IEEE Trans. Speech Audio Process. | 2 |
| 2009 | Feedforward control of a 3d physiological articulatory model for vowel production
Qiang Fang 0003, Akikazu Nishikido, Jianwu Dang 0001 |
INTERSPEECH | 3 |
| 2008 | A model based investigation of activation patterns of the tongue muscles for vowel production
Qiang Fang 0003, Satoru Fujita, Xugang Lu, Jianwu Dang 0001 |
INTERSPEECH | 4 |
| 2008 | An investigation of dependencies between frequency components and speaker characteristics for text-independent speaker identification
Xugang Lu, Jianwu Dang 0001 |
Speech Commun. | 2 |
| 2007 | Physiological Feature Extraction for Text Independent Speaker Identification using Non-Uniform Subband ProcessingabstractThe features used for speech recognition should emphasize linguistic information while suppressing speaker differences. For speaker recognition, features should have more speaker individual information while attenuating the linguistic information. In most studies, however, the identical acoustic features are used for the different missions of speaker and speech recognitions. In this paper, we propose a new physiological feature extraction method which emphasizes individual information for speaker identification. For the purpose, physiological features of speakers were analyzed from the point of view of speech production. It is found that the speaker individual information is encoded in different frequency regions of speech sound. The speaker discriminative information was quantified using Fisher's F-ratio in each frequency region. Based on the F-ratio, we proposed a non-uniform sub-band processing strategy to extract new feature which can emphasize or refine the physiological aspects involved in speech production. We combined the new feature with GMM for speaker identification task and applied on NTT-VR speaker recognition database. Compared with MFCC feature, by using the proposed feature, the identification error rate was reduced 20.1%. Xugang Lu, Jianwu Dang 0001 |
ICASSP (4) | 2 |
| 2007 | Dimension reduction for speaker identification based on mutual informationabstractAbstract Dimension reduction is a necessary step for speech feature extraction in a speaker identification system. Discrete Cosine Transform (DCT) or Principal Component Analysis (PCA) is widely used for dimension reduction. By choosing basis vectors from basis vector pool of DCT or PCA which contribute more to data distribution variance or reconstruction accuracy of speech data set, we can transform the data set by projecting them on to the selected basis vectors. However, keeping the maximum distribution variance or high reconstruction accuracy does not guarantee the optimal keeping of high speaker discriminative information. In this paper, we proposed a basis vector selection method based on mutual information concept which guarantees the keeping of high speaker discriminative information. The mutual information is used to measure the dependency between the features extracted using basis vectors and speaker class labels. The high mutual information related basis vectors are chosen for feature extraction. Considering one speaker feature may be encoded in more than one basis vectors, we proposed to use joint mutual information concept which takes the dependency between feature variables into consideration. Based on the selected basis vectors from DCT or PCA basis vector pool, we extracted features for speaker identification experiments. Experimental results showed that the speaker identification error rate using proposed feature was reduced 11% and 8% on average for DCT and PCA based features respectively. Xugang Lu, Jianwu Dang 0001 |
INTERSPEECH | 2 |
| 2006 | A simulation based parameter optimization for a coarticulation modelabstractA coarticulation model, namely ‘carrier model’, has been proposed previously by Dang et al. to improve the performance of a physiological articulatory model based speech synthesizer. The carrier model offers a good framework to account for coarticulation in the planning stage, while its parameters need to be refined for improving the performance of the model. This study is to refine the parameters of the carrier model and estimate typical phonetic targets by minimizing the differences between model simulations and observations. A simulation based optimization framework is proposed for this purpose. The framework consists of two layers: obtaining planned targets in a low layer; estimating phonetic targets and optimizing the parameters in a high layer. A direct search method was applied to the low layer due to the non-analytic nature of the articulation model, while the high layer adopts bilevel optimization strategy to decompose the complicated problem into a set of subproblems. A general evaluation was conducted by combining the refined carrier model and the learned phonetic targets together using the physiological articulatory model and the average error between observations and simulations was 0.15 cm over 103 VCV combinations on the jaw, tongue tip and tongue dorsum. Index Terms: speech production, coarticulation, optimization Jianguo Wei, Xugang Lu, Jianwu Dang 0001 |
INTERSPEECH | 3 |
| 2006 | Communication Between Speech Production and Perception Within the Brain-Observation and Simulation
Jianwu Dang 0001, Masato Akagi, Kiyoshi Honda |
J. Comput. Sci. Technol. | 1 |
| 2005 | Investigation and modeling of coarticulation during speech
Jianwu Dang 0001, Jianguo Wei, Takeharu Suzuki, Pascal Perrier |
INTERSPEECH | 1 |
| 2003 | Consideration of muscle co-contraction in a physiological articulatory modelabstractPhysiological models of the speech organs must consider cocontraction of the muscles, a common phenomenon taking place during articulation. This study investigated cocontraction of the tongue muscles using the physiological articulatory model that replicates midsagittal regions of the speech organs to simulate articulatory movements during speech [1,2]. The relation between the muscle force and tongue movement obtained by the model simulation indicated that each muscle drives the tongue towards an equilibrium position (EP) corresponding to the magnitude of the activation forces. Contributions of the muscles to the tongue movement were evaluated by the distance between the equilibrium positions. Based on the EPs and the muscle contributions, an invariant mapping (the EP map) was established to function the connection of a spatial location to a muscle force. Cocontractions between agonist and antagonist muscles were simulated using the EP maps. The simulations demonstrated that coarticulation with multiple targets could be compatibly realized using the co-contraction mechanism. The implementation of the co-contraction mechanism enables relatively independent control over the tongue tip and body. Jianwu Dang 0001, Kiyoshi Honda |
INTERSPEECH | 1 |
| 2002 | Investigation of coarticulation based on electromagnetic articulographic data
Jianwu Dang 0001, Masaaki Honda, Kiyoshi Honda |
INTERSPEECH | 1 |
| 2000 | Improvement of a physiological articulatory model for synthesis of vowel sequences
Jianwu Dang 0001, Kiyoshi Honda |
INTERSPEECH | 1 |
| 1998 | Speech production of vowel sequences using a physiological articulatory modelabstractThis report describes the development of a physiologically-based articulatory model, which consists of the tongue, mandible, hyoid bone and vocal tract wall. These organs are represented in a quasi-3D shape to replicate a midsagittal layer with a thickness of 2 cm for tongue tissue and 3 cm for tract wall. The geometry of these organs and muscles are extracted from volumetric MR images of a male speaker. Both the soft and rigid structures are represented by mass-points and viscoelastic springs for connective tissue, where the springs for bony organs are set to extremely large stiffness. This design is suitable to compute soft tissue deformations and rigid organ displacements simultaneously using a single algorithm, and thus reduces computational complexities of the simulation. A novel control method is developed to produce dynamic actions of the vocal Jianwu Dang 0001, Kiyoshi Honda |
ICSLP | 1 |
| 1996 | An improved vocal tract model of vowel production implementing piriform resonance and transvelar nasal coupling
Jianwu Dang 0001, Kiyoshi Honda |
ICSLP | 1 |
| 1994 | Investigation of the acoustic characteristics of the velum for vowels
Jianwu Dang 0001, Kiyoshi Honda |
ICSLP | 1 |
| 1994 | A physiological model of speech production and the implication of tongue-larynx interaction
Kiyoshi Honda, Hiroyuki Hirai, Jianwu Dang 0001 |
ICSLP | 3 |