Xin Wang 0037

dblp:10/5630-37 · DBLP profile ↗
← Back
109ranked-venue papers
27as first author
72since 2021 · last 2026
0000-0001-8246-0606ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 72 · 17 first-author · 41 since 2021Artificial intelligence and machine learning · 58 · 12 first-author · 36 since 2021Security and privacy · 8 · 2 first-author · 8 since 2021Computer networks · 5 · 1 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 5 · 5 since 2021Systems, architecture and hardware · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 The third VoicePrivacy challenge: Preserving emotional expressiveness and linguistic content in voice anonymization
Natalia A. Tomashenko, Xiaoxiao Miao, Pierre Champion, Sarina Meyer, Michele Panariello, Xin Wang 0037, Nicholas W. D. Evans, Emmanuel Vincent 0001, Junichi Yamagishi, Massimiliano Todisco
Comput. Speech Lang.6
2026 ASVspoof 5: Design, collection and validation of resources for spoofing, deepfake, and adversarial attack detection using crowdsourced speech
abstract
ASVspoof 5 is the fifth edition in a series of challenges which promote the study of speech spoofing and deepfake attacks as well as the design of detection solutions. We introduce the ASVspoof 5 database which is generated in a crowdsourced fashion from data collected in diverse acoustic conditions (cf. studio-quality data for earlier ASVspoof databases) and from ∼ 2,000 speakers (cf. ∼ 100 earlier). The database contains attacks generated with 32 different algorithms, also crowdsourced, and optimised to varying degrees using new surrogate detection models. Among them are attacks generated with a mix of legacy and contemporary text-to-speech synthesis and voice conversion models, in addition to adversarial attacks which are incorporated for the first time. ASVspoof 5 protocols comprise seven speaker-disjoint partitions. They include two distinct partitions for the training of different sets of attack models, two more for the development and evaluation of surrogate detection models, and then three additional partitions which comprise the ASVspoof 5 training, development and evaluation sets. An auxiliary set of data collected from an additional 30k speakers can also be used to train speaker encoders for the implementation of attack algorithms. Also described herein is an experimental validation of the new ASVspoof 5 database using a set of automatic speaker verification and spoof/deepfake baseline detectors. With the exception of protocols and tools for the generation of spoofed/deepfake speech, the resources described in this paper, already used by participants of the ASVspoof 5 challenge in 2024, are now all freely available to the community.
Xin Wang 0037, Héctor Delgado, Hemlata Tak, Jee-Weon Jung, Hye-Jin Shim, Massimiliano Todisco, Ivan Kukanov, Xuechen Liu 0001, Md. Sahidullah, Tomi Kinnunen, Nicholas W. D. Evans, Kong-Aik Lee, Junichi Yamagishi, Myeonghun Jeong, Yongyi Zang, Soumi Maiti, Florian Lux, Nicolas Müller, Wangyou Zhang, Chengzhe Sun 0001, Shuwei Hou, Siwei Lyu, Sébastien Le Maguer, Hanjie Guo, Vishwanath Pratap Singh
Comput. Speech Lang.1
2026 Search Me in the Dark: Access Pattern-Hidden Range Query Over Encrypted Spatial Data
abstract
With the widespread use of encrypted spatial data, many range query schemes emerge to address potential security risks caused by access pattern leakage. However, most existing schemes rely on a dual-server model to hide access patterns and often involve complex spatial relation judgments during range comparisons, leading to low query efficiency. To address these issues, we propose a novel Fast and Access Hidden Range Query (FAHRQ) scheme. First, we introduce an efficient range membership verification technique based on Bloom filters and Lagrange interpolation function, combine homomorphic encryption to ensure the confidentiality of spatial data and the computational flexibility of related operations, and realize the access pattern hidden under single server. Then, we construct an index using R-tree and employ Bloom filters and prefix 0-1 encoding to accelerate the minimum bounding rectangle intersection judgment, enabling secure and efficient range queries over encrypted spatial data while maintaining retrieval accuracy. Finally, we give a formal security analysis to show that our scheme achieves access pattern hidden while protecting data security, and conduct extensive experiments to demonstrate that our scheme improves query efficiency by 5 – 7× compared to existing schemes.
Yinbin Miao, Xin Wang 0037, Kaifa Zheng, Xinghua Li 0001, Zhiquan Liu 0001, Robert H. Deng
IEEE Trans. Inf. Forensics Secur.2
2026 Robust Identity-Based Signcryption Scheme for Vehicular Ad Hoc Networks
abstract
Vehicular Ad Hoc Networks (VANETs) are the cornerstone of intelligent transportation systems and autonomous driving. Vehicle-to-road communication, as one of the core services, faces increasing risks of privacy breaches. Signcryption technology effectively ensures secure information transmission. However, existing signcryption schemes still have deficiencies in terms of transmission robustness and identity privacy protection. To solve these issues, this paper proposes a Robust Identity-based Signcryption scheme (RIBSC) for VANETs. In RIBSC, we first design an area session key distribution mechanism based on Chinese Residual Theorem (CRT), which can dynamically revoke the decryption ability of malicious Roadside Units (RSUs) in real time. Only RSUs approved by Trusted Detection Center (TDC) can obtain a valid session private key by conducting one modular operation. We then utilize the traceable pseudonym mechanism to protect the identity privacy of vehicles and RSUs, which can track their true identities when illegal activities occur. We finally provide a rigorous security proof under the random oracle model, and demonstrate the performance advantages of RIBSC through extensive experiments. More attractively, the session information is fixed at only 148 bytes, regardless of the number of RSUs.
Xin Wang 0037, Yinbin Miao, Xinghua Li 0001, Hongwei Li 0001, Robert H. Deng
IEEE Trans. Inf. Forensics Secur.1
2026 Efficient Heterogeneous Signcryption With Forward Privacy for Vehicular Platoon Communication
Xin Wang 0037, Yinbin Miao, Xinghua Li 0001, Zhiquan Liu 0001, Jun Feng 0007, Robert H. Deng
IEEE Trans. Inf. Forensics Secur.1
2025 Post-training for Deepfake Speech Detection
abstract
We introduce a post-training approach that adapts self-supervised learning (SSL) models for deepfake speech detection by bridging the gap between general pre-training and domain-specific fine-tuning. We present AntiDeepfake models, a series of post-trained models developed using a large-scale multilingual speech dataset containing over $\mathbf{5 6, 0 0 0}$ hours of genuine speech and $\mathbf{1 8, 0 0 0}$ hours of speech with various artifacts in over one hundred languages. Experimental results show that the post-trained models already exhibit strong robustness and generalization to unseen deepfake speech. When they are further fine-tuned on the Deepfake-Eval-2024 dataset, these models consistently surpass existing state-of-the-art detectors that do not leverage post-training. Model checkpoint1and source code2are available online.1Zenodo: https://doi.org/10.5281/zenodo.15580542 Hugging Face: https://huggingface.co/nii-yamagishilab2GitHub: https://github.com/nii-yamagishilab/AntiDeepfake
Wanying Ge, Xin Wang 0037, Xuechen Liu 0001, Junichi Yamagishi
ASRU2
2025 SEF-MK: Speaker-Embedding-Free Voice Anonymization through Multi-k-means Quantization
abstract
Voice anonymization protects speaker privacy by concealing identity while preserving linguistic and paralinguistic content. Self-supervised learning (SSL) representations encode linguistic features but preserve speaker traits. We propose a novel speaker-embedding-free framework called SEF-MK. Instead of using a single k-means model trained on the entire dataset, SEF-MK anonymizes SSL representations for each utterance by randomly selecting one of multiple k-means models, each trained on a different subset of speakers. We explore this approach from both attacker and user perspectives. Extensive experiments show that, compared to a single k -means model, SEF-MK with multiple $\mathbf{k}$-means models better preserves linguistic and emotional content from the user’s viewpoint. However, from the attacker’s perspective, utilizing multiple $\mathbf{k}$-means models boosts the effectiveness of privacy attacks. These insights can aid users in designing voice anonymization systems to mitigate attacker threats.11Code and audio samples can be found at https://github.com/Beilong-Tang/sef-mk
Beilong Tang, Xiaoxiao Miao, Xin Wang 0037, Ming Li 0026
ASRU3
2025 DNPS: A Robust Aggregation Method for Heterogeneous Distributed Learning Based on Gradient Direction and Norm Probability Screening
abstract
Heterogeneous distributed machine learning systems are vulnerable to Byzantine attacks, where malicious worker nodes disrupt global model convergence by submitting incorrect model updates, leading to degraded system performance. Current aggregation methods that filter based on gradient norm or direction often underperform in heterogeneous environments and struggle to defend effectively against Byzantine attacks. To address this issue, we propose the Direction-Norm Probability Screening (DNPS) algorithm-a novel approach that harmonizes gradient norm and direction to blunt these attacks. DNPS integrates gradient norm and directional information to establish a probability screening mechanism for identifying and filtering Byzantine worker nodes, thereby fortifying the system's defen-sive capabilities and enhancing its robustness against Byzantine attacks. Experimental results show the DNPS algorithm provides more effective protection under both non-attack and Byzantine attack scenarios than existing aggregation methods.
Ming Yang 0023, Caiyun Li, Yunpeng He, Xin Wang 0037
CSCWD4
2025 Audio Codec Augmentation for Robust Collaborative Watermarking of Speech Synthesis
abstract
Automatic detection of synthetic speech is becoming increasingly important as current synthesis methods are both near indistinguishable from human speech and widely accessible to the public. Audio watermarking and other active disclosure methods of are attracting research activity, as they can complement traditional deepfake defenses based on passive detection. In both active and passive detection, robustness is of major interest. Traditional audio watermarks are particularly susceptible to removal attacks by audio codec application. Most generated speech and audio content released into the wild passes through an audio codec purely as a distribution method. We recently proposed collaborative watermarking as method for making generated speech more easily detectable over a noisy but differentiable transmission channel. This paper extends the channel augmentation to work with non-differentiable traditional audio codecs and neural audio codecs and evaluates transferability and effect of codec bitrate over various configurations. The results show that collaborative watermarking can be reliably augmented by black-box audio codecs using a waveform-domain straight-through-estimator for gradient approximation. Furthermore, that results show that channel augmentation with a neural audio codec transfers well to traditional codecs. Listening tests demonstrate collaborative watermarking incurs negligible perceptual degradation with high bitrate codecs or DAC at 8kbps.
Lauri Juvela, Xin Wang 0037
ICASSP2
2025 Towards An Integrated Approach for Expressive Piano Performance Synthesis from Music Scores
abstract
This paper presents an integrated system that transforms symbolic music scores into expressive piano performance audio. By combining a Transformer-based Expressive Performance Rendering (EPR) model with a fine-tuned neural MIDI synthesiser, our approach directly generates expressive audio performances from score inputs. To the best of our knowledge, this is the first system to offer a streamlined method for converting score MIDI files lacking expression control into rich, expressive piano performances. We conducted experiments using subsets of the ATEPP dataset, evaluating the system with both objective metrics and subjective listening tests. Our system not only accurately reconstructs human-like expressiveness, but also captures the acoustic ambience of environments such as concert halls and recording studios. Additionally, the proposed system demonstrates its ability to achieve musical expressiveness while ensuring good audio quality in its outputs.
Jingjing Tang 0002, Erica Cooper, Xin Wang 0037, Junichi Yamagishi, György Fazekas
ICASSP3
2025 SecureSpeech: Prompt-based Speaker and Content Protection
abstract
Given the increasing privacy concerns from identity theft and the re-identification of speakers through content in the speech field, this paper proposes a prompt-based speech generation pipeline that ensures dual anonymization of both speaker identity and spoken content. This is addressed through 1) generating a speaker identity un-linkable to the source speaker, controlled by descriptors, and 2) replacing sensitive content within the original text using a name entity recognition model and a large language model. The pipeline utilizes the anonymized speaker identity and text to generate high-fidelity, privacy-friendly speech via a text-to-speech synthesis model. Experimental results demonstrate an achievement of significant privacy protection while maintaining a decent level of content retention and audio quality. This paper also investigates the impact of varying speaker descriptions on the utility and privacy of generated speech to determine potential biases.
Belinda Soh Hui Hui, Xiaoxiao Miao, Xin Wang 0037
IJCB3
2025 LENS-DF: Deepfake Detection and Temporal Localization for Long-Form Noisy Speech
abstract
This study introduces LENS-DF, a novel and comprehensive recipe for training and evaluating audio deepfake detection and temporal localization under complicated and realistic audio conditions. The generation part of the recipe outputs audios from the input dataset with several critical characteristics, such as longer duration, noisy conditions, and containing multiple speakers, in a controllable fashion. The corresponding detection and localization protocol uses models. We conduct experiments based on self-supervised learning front-end and simple back-end. The results indicate that models trained using data generated with LENS-DF consistently outperform those trained via conventional recipes, demonstrating the effectiveness and usefulness of LENS-DF for robust audio deepfake detection and localization. We also conduct ablation studies on the variations introduced, investigating their impact on and relevance to realistic challenges in the field1.
Xuechen Liu 0001, Wanying Ge, Xin Wang 0037, Junichi Yamagishi
IJCB3
2025 FedSaaS: Class-Consistency Federated Semantic Segmentation via Global Prototype Supervision and Local Adversarial Harmonization
abstract
Federated semantic segmentation enables pixel-level classification in images through collaborative learning while maintaining data privacy. However, existing research commonly overlooks the fine-grained class relationships within the semantic space when addressing heterogeneous problems, particularly domain shift. This oversight results in ambiguities between class representation. To overcome this challenge, we propose a novel federated segmentation framework that strikes class consistency, termed FedSaaS. Specifically, we introduce class exemplars as a criterion for both local- and global-level class representations. On the server side, the uploaded class exemplars are leveraged to model class prototypes, which supervise global branch of clients, ensuring alignment with global-level representation. On the client side, we incorporate an adversarial mechanism to harmonize contributions of global and local branches, leading to consistent output. Moreover, multilevel contrastive losses are employed on both sides to enforce consistency between two-level representations in the same semantic space. Extensive experiments on five driving scene segmentation datasets demonstrate that our framework outperforms state-of-the-art methods, significantly improving average segmentation accuracy and effectively addressing the class-consistency representation problem.
Xin Wang 0037, Dongrun Li, Ming Yang 0023, Peng Cheng 0001
IJCAI3
2025 Dyn-D^2P: Dynamic Differentially Private Decentralized Learning with Provable Utility Guarantee
abstract
Most existing decentralized learning methods with differential privacy (DP) guarantee rely on constant gradient clipping bounds and fixed-level DP Gaussian noises for each node throughout the training process, leading to a significant accuracy degradation compared to non-private counterparts. In this paper, we propose a new Dynamic Differentially Private Decentralized learning approach (termed Dyn-D^2P) tailored for general time-varying directed networks. Leveraging the Gaussian DP (GDP) framework for privacy accounting, Dyn-D^2P dynamically adjusts gradient clipping bounds and noise levels based on gradient convergence. This proposed dynamic noise strategy enables us to enhance model accuracy while preserving the total privacy budget. Extensive experiments on benchmark datasets demonstrate the superiority of Dyn-D^2P over its counterparts employing fixed-level noises, especially under strong privacy guarantees. Furthermore, we provide a provable utility bound for Dyn-D^2P that establishes an explicit dependency on network-related parameters, with a scaling factor of 1/sqrt{n} in terms of the number of nodes n up to a bias error term induced by gradient clipping. To our knowledge, this is the first model utility analysis for differentially private decentralized non-convex optimization with dynamic gradient clipping bounds and noise levels.
Zehan Zhu, Yan Huang 0036, Xin Wang 0037, Shouling Ji, Jinming Xu 0002
IJCAI3
2025 Bridging Privacy Preservation and Optimization in Heterogeneous Decentralized Learning: Regularization Tuning and Model Pruning
abstract
Federated learning (FL) is an efficient distributed optimization algorithm but faces significant challenges related to the risk of privacy leakage during training. Many existing FL methods rely on centralized communication network topologies, which have inherent limitations in practical applications. These limitations include vulnerability to single points of failure, susceptibility to communication bottlenecks, and an inability to effectively adapt to dynamic and decentralized environments. To address these challenges, this paper proposes the PODL-RM method, which offers an optimized framework for decentralized learning under time-varying directed communication topologies while ensuring personalized differential privacy (DP) protection for each client. To mitigate the negative effects of data heterogeneity and DP noise perturbation on model performance, PODL-RM combines regularization tuning with model pruning techniques. These components work synergistically to enhance both convergence efficiency and model accuracy. We conduct a rigorous theoretical analysis of the proposed method, formally establishing its convergence in the context of non-convex optimization problems. Experimental results demonstrate that the proposed method significantly outperforms state-of-the-art algorithms in time-varying directed communication topologies, yielding superior convergence performance and reduced communication costs.
Xin Wang 0037, Ming Yang 0023
IJCNN3
2025 From Sharpness to Better Generalization for Speech Deepfake Detection
Wen Huang 0004, Xuechen Liu 0001, Xin Wang 0037, Junichi Yamagishi, Yanmin Qian
INTERSPEECH3
2025 The Text-to-speech in the Wild (TITW) Database
Jee-Weon Jung, Wangyou Zhang, Soumi Maiti, Yihan Wu 0008, Xin Wang 0037, Yuta Matsunaga, Seyun Um, Jinchuan Tian, Hye-Jin Shim, Nicholas W. D. Evans, Joon Son Chung, Shinnosuke Takamichi, Shinji Watanabe 0001
INTERSPEECH5
2025 A Comparative Study on Proactive and Passive Detection of Deepfake Speech
Chia-Hua Wu, Wanying Ge, Xin Wang 0037, Junichi Yamagishi, Yu Tsao 0001, Hsin-Min Wang
INTERSPEECH3
2025 Mitigating Language Mismatch in SSL-Based Speaker Anonymization
Wen-Chin Huang, Xin Wang 0037, Xiaoxiao Miao, Junichi Yamagishi
INTERSPEECH3
2025 DFed-LaMA: Differentially Private Federated Learning via Adaptive Layer-Wise Model Aggregation
abstract
Personalized federated learning (PFL) is a distributed learning paradigm designed to address data heterogeneity across clients. While PFL enhances model adaptability through local personalization, achieving a balance between robust privacy protection and effective personalized learning remains a critical challenge—particularly for sensitive data. To address this issue, we propose DFed-LaMA, a novel differentially private PFL framework that incorporates adaptive layer-wise model aggregation, optimizing the trade-off between personalization and privacy. Specifically, clients dynamically identify and personalize the most relevant model layers—determined via Kullback-Leibler divergence between local and global models—before applying differential privacy (DP) perturbation and uploading them to the server. Furthermore, we introduce a model-product integration strategy to align local updates with global objectives, mitigating performance degradation induced by DP noise. Extensive experiments on multiple benchmark datasets demonstrate that DFed-LaMA outperforms state-of-the-art methods in classification accuracy, training stability, and privacy guarantees.
Xin Wang 0037, Heng Zhang 0001, Ming Yang 0023
MASS3
2025 Zero Trust-Based Dynamic and Continuous Access Control for Mobile Devices
abstract
Communication among mobile devices in the Inter-net of Things (IoT) typically relies on fixed security boundaries and centralized trust models, resulting in weak trust mechanisms, unauthorized access, and limited adaptability to dynamic environments. To address this, this paper proposes a zero-trust-based secure communication access control method that follows the principles of zero trust and least privilege, establishing a dynamic trusted framework encompassing identity authentication, real-time trust evaluation, multi-layer access control, continuous behavior analysis, and trajectory visualization. A dynamic trust evaluation model combining static attributes and historical interaction data is designed to ensure that only users who pass trust assessments are granted access. An interaction trust computation method incorporating direct trust, indirect trust, and a time-decay factor is introduced, along with a sliding window algorithm for dynamic threshold updates, enhancing responsiveness to changes in device behavior. A machine-learning-based continuous behavior monitoring mechanism is implemented to perform real-time modeling and detection of performance, traffic, and access patterns, improving anomaly identification and rapid response. The prototype system was validated through multi-device collaborative interaction simulations, demonstrating significant advantages in traffic anomaly detection and communication security over existing approaches.
Xiaoya Cao, Zhenya Chen, Ming Yang 0023, Xin Wang 0037
TrustCom4
2025 Adapting general disentanglement-based speaker anonymization for enhanced emotion preservation
abstract
A general disentanglement-based speaker anonymization system typically separates speech into content, speaker, and prosody features using individual encoders. This paper explores how to adapt such a system when a new speech attribute, for example, emotion, needs to be preserved to a greater extent. Two strategies for this are examined. First, we show that integrating emotion embeddings from a pre-trained emotion encoder can help preserve emotional cues, even though this approach slightly compromises privacy protection. Alternatively, we propose an emotion compensation strategy as a post-processing step applied to anonymized speaker embeddings. This conceals the original speaker’s identity and reintroduces the emotional traits lost during speaker embedding anonymization. Specifically, we model the emotion attribute using support vector machines to learn separate boundaries for each emotion. During inference, the original speaker embedding is processed in two ways: one, by an emotion indicator to predict emotion and select the emotion-matched SVM accurately; and two, by a speaker anonymizer to conceal speaker characteristics. The anonymized speaker embedding is then modified along the corresponding SVM boundary towards an enhanced emotional direction to save the emotional cues. The proposed strategies are also expected to be useful for adapting a general disentanglement-based speaker anonymization system to preserve other target paralinguistic attributes, with potential for a range of downstream tasks 2 .
Xiaoxiao Miao, Xin Wang 0037, Natalia A. Tomashenko, Cheng Lock Donny Soh, Ian McLoughlin 0001
Comput. Speech Lang.3
2025 Efficient Encrypted Trajectory Similarity Query Over Mobile E-Health Cloud
abstract
Mobile electronic health systems collect a large amount of people’s trajectory data through smart devices (e.g., sensors). Generally, to provide data confidentiality, trajectories are encrypted before being uploaded to cloud servers. Trajectory similarity query has gained widespread attention as a mean to control the spread of infectious diseases. Nonetheless, existing solutions often suffer from low query efficiency and access pattern exposure. To solve these issues, we propose an efficient encrypted trajectory similarity query$(\textsf {TraSQ})$for mobile e-health systems. First, we utilize XZ* index and bijective function to generate the unique trajectory encoding value, thereby reducing storage and query costs over large-scale trajectory datasets. Then, we employ a dual cloud model, secure computing protocols and obfuscation technique to conceal the true trajectory similarity from cloud servers, protecting access pattern. Finally, we formally prove that our scheme achieves privacy protection against chosen plaintext attack (CPA), and conduct extensive experiments to demonstrate that our scheme improves the query efficiency by at least 25.4% when compared with existing solutions.
Xin Wang 0037, Yinbin Miao, Qingming Li, Kefeng Ding, Shouling Ji
IEEE Internet Things J.1
2025 A Benchmark for Multi-Speaker Anonymization
abstract
Privacy-preserving voice protection approaches primarily suppress privacy-related information derived from paralinguistic attributes while preserving the linguistic content. Existing solutions focus particularly on single-speaker scenarios. However, they lack practicality for real-world applications, i.e., multi-speaker scenarios. In this paper, we present an initial attempt to provide a multi-speaker anonymization benchmark by defining the task and evaluation protocol, proposing benchmarking solutions, and discussing the privacy leakage of overlapping conversations. The proposed benchmark solutions are based on a cascaded system that integrates spectral-clustering-based speaker diarization and disentanglement-based speaker anonymization using a selection-based anonymizer. To improve utility, the benchmark solutions are further enhanced by two conversation-level speaker vector anonymization methods. The first method minimizes the differential similarity across speaker pairs in the original and anonymized conversations, which maintains original speaker relationships in the anonymized version. The other minimizes the aggregated similarity across anonymized speakers, which achieves better differentiation between speakers. Experiments conducted on both non-overlap simulated and real-world datasets demonstrate the effectiveness of the multi-speaker anonymization system with the proposed speaker anonymizers. Additionally, we analyzed overlapping speech regarding privacy leakage and provided potential solutions (Code and audio samples are available athttps://github.com/xiaoxiaomiao323/MSA), evaluation datasets can be download fromhttps://zenodo.org/records/14249171
Xiaoxiao Miao, Ruijie Tao, Chang Zeng, Xin Wang 0037
IEEE Trans. Inf. Forensics Secur.4
2025 Trace Your Footprint: Efficient Spatial Keyword Query Over Encrypted Trajectory Data
abstract
With the popularity of mobile devices, spatial-textual trajectory query has been deployed in applications such as trajectory-based navigation and travel route recommendation. Massive trajectory data have been outsourced to cloud servers for storage and sharing such as spatial keyword search. However, existing solutions only support similarity queries in the spatial dimension and still incur high storage and query costs, which cannot scale well in large-scale trajectory data scenarios. To solve the above issues, we first achieve an Efficient Range Query over Encrypted Trajectory Data (ERT) using Douglas-Peucker trajectory compression algorithm, random matrix multiplication, filtering-verification mechanism and polynomial fitting technology. Then, we further propose an enhanced Efficient Spatial Keyword Query over Encrypted Trajectory Data (ESKT) by constructing a unified spatial-textual index structure, which can find relevant trajectories that are within some arbitrary geometric range and contain all query keywords. Finally, we formally prove that our schemes are secure against chosen-plaintext-attack, and conduct extensive experiments to demonstrate that our schemes improve the query efficiency by almost 100× when compared with state-of-the-art solutions.
Yinbin Miao, Xin Wang 0037, Xinghua Li 0001, Shujiang Xu, Zhiquan Liu 0001, Kim-Kwang Raymond Choo, Robert H. Deng
IEEE Trans. Inf. Forensics Secur.2
2024 Efficient Secure Inference Scheme for Large Neural Networks
abstract
Secure inference is the main technology to avoid privacy leakage when reasoning with machine learning models. The existing secure inference schemes have been widely explored in the academic field, but when the neural network model is large, existing solutions based on Trusted Execution Environments (TEE) must be configured with a large number of trusted devices due to insufficient hardware memory, which incurs high deployment costs. In addition, existing homomorphic encryption schemes have a high computational cost for ciphertexts, which incurs high time cost when the model is large. To address these issues, we propose a low time overhead secure inference scheme by training a generative adversarial network, which can reduce the calculations for ciphertext. We also use a trusted execution environment to compute nonlinear functions in neural networks, which solves the problem of high computational cost for nonlinear functions in HE while suppressing the growth of noise in HE. Security analysis proves that our scheme achieves semi-honest security, and extensive experiments demonstrate that our scheme has lower time overhead in large neural networks when compare with previous solutions.
Lin Chen 0033, Yiwei Yang 0003, Xin Wang 0037, Yinbin Miao, Chao Hong
HPCC3
2024 Can Large-Scale Vocoded Spoofed Data Improve Speech Spoofing Countermeasure with a Self-Supervised Front End?
abstract
A speech spoofing countermeasure (CM) that discriminates between unseen spoofed and bona fide data requires diverse training data. While many datasets use spoofed data generated by speech synthesis systems, it was recently found that data vocoded by neural vocoders were also effective as the spoofed training data. Since many neural vocoders are fast in building and generation, this study used multiple neural vocoders and created more than 9,000 hours of vocoded data on the basis of the VoxCeleb2 corpus. This study investigates how this large-scale vocoded data can improve spoofing countermeasures that use data-hungry self-supervised learning (SSL) models. Experiments demonstrated that the overall CM performance on multiple test sets improved when using features extracted by an SSL model continually trained on the vocoded data. Further improvement was observed when using a new SSL distilled from the two SSLs before and after the continual training. The CM with the distilled SSL outperformed the previous best model on challenging unseen test sets, including the ASVspoof 2019 logical access, WaveFake, and In-the-Wild.
Xin Wang 0037, Junichi Yamagishi
ICASSP1
2024 Spoofing Attack Augmentation: Can Differently-Trained Attack Models Improve Generalisation?
abstract
A reliable deepfake detector or spoofing countermeasure (CM) should be robust in the face of unpredictable spoofing attacks. To encourage the learning of more generaliseable artefacts, rather than those specific only to known attacks, CMs are usually exposed to a broad variety of different attacks during training. Even so, the performance of deeplearning-based CM solutions are known to vary, sometimes substantially, when they are retrained with different initialisations, hyper-parameters or training data partitions. We show in this paper that the potency of spoofing attacks, also deep-learning-based, can similarly vary according to training conditions, sometimes resulting in substantial degradations to detection performance. Nevertheless, while a RawNet2 CM model is vulnerable when only modest adjustments are made to the attack algorithm, those based upon graph attention networks and self-supervised learning are reassuringly robust. The focus upon training data generated with different attack algorithms might not be sufficient on its own to ensure generaliability; some form of spoofing attack augmentation at the algorithm level can be complementary.
Wanying Ge, Xin Wang 0037, Junichi Yamagishi, Massimiliano Todisco, Nicholas W. D. Evans
ICASSP2
2024 Collaborative Watermarking for Adversarial Speech Synthesis
abstract
Advances in neural speech synthesis have brought us technology that is not only close to human naturalness, but is also capable of instant voice cloning with little data, and is highly accessible with pre-trained models available. Naturally, the potential flood of generated content raises the need for synthetic speech detection and watermarking. Recently, considerable research effort in synthetic speech detection has been related to the Automatic Speaker Verification and Spoofing Countermeasure Challenge (ASVspoof), which focuses on passive countermeasures. This paper takes a complementary view to generated speech detection: a synthesis system should make an active effort to watermark the generated speech in a way that aids detection by another machine, but remains transparent to a human listener. We propose a collaborative training scheme for synthetic speech watermarking and show that a HiFi-GAN neural vocoder collaborating with the ASVspoof 2021 baseline countermeasure models consistently improves detection performance over conventional classifier training. Furthermore, we demonstrate how collaborative training can be paired with augmentation strategies for added robustness against noise and time-stretching. Finally, listening tests demonstrate that collaborative training has little adverse effect on perceptual quality of vocoded speech.
Lauri Juvela, Xin Wang 0037
ICASSP2
2024 Synvox2: Towards A Privacy-Friendly Voxceleb2 Dataset
abstract
The success of deep learning in speaker recognition relies heavily on the use of large datasets. However, the data-hungry nature of deep learning methods has already being questioned on account the ethical, privacy, and legal concerns that arise when using large-scale datasets of natural speech collected from real human speakers. For example, the widely-used VoxCeleb2 dataset for speaker recognition is no longer accessible from the official website. To mitigate these concerns, this work presents an initiative to generate a privacyfriendly synthetic VoxCeleb2 dataset that ensures the quality of the generated speech in terms of privacy, utility, and fairness. We also discuss the challenges of using synthetic data for the downstream task of speaker verification.
Xiaoxiao Miao, Xin Wang 0037, Erica Cooper, Junichi Yamagishi, Nicholas W. D. Evans, Massimiliano Todisco, Jean-François Bonastre, Mickael Rouvier
ICASSP2
2024 PrivSGP-VR: Differentially Private Variance-Reduced Stochastic Gradient Push with Tight Utility Bounds
Zehan Zhu, Yan Huang 0036, Xin Wang 0037, Jinming Xu 0002
IJCAI3
2024 Revisiting and Improving Scoring Fusion for Spoofing-aware Speaker Verification Using Compositional Data Analysis
abstract
Interspeech 2024, 1-5 September 2024, Kos, Greece
Xin Wang 0037, Tomi Kinnunen, Kong-Aik Lee, Paul-Gauthier Noé, Junichi Yamagishi
INTERSPEECH1
2024 An Initial Investigation of Language Adaptation for TTS Systems under Low-resource Scenarios
abstract
Self-supervised learning (SSL) representations from massively multilingual models offer a promising solution for low-resource language speech tasks. Despite advancements, language adaptation in TTS systems remains an open problem. This paper explores the language adaptation capability of ZMM-TTS, a recent SSL-based multilingual TTS system proposed in our previous work. We conducted experiments on 12 languages using limited data with various fine-tuning configurations. We demonstrate that the similarity in phonetics between the pretraining and target languages, as well as the language category, affects the target language’s adaptation performance. Additionally, we find that the fine-tuning dataset size and number of speakers influence adaptability. Surprisingly, we also observed that using paired data for fine-tuning is not always optimal compared to audio-only data. Beyond speech intelligibility, our analysis covers speaker similarity, language identification, and predicted MOS.
Erica Cooper, Xin Wang 0037, Chunyu Qiang, Mengzhe Geng, Dan Wells, Longbiao Wang, Jianwu Dang 0001, Marc Tessier, Aidan Pine, Korin Richmond, Junichi Yamagishi
INTERSPEECH3
2024 To what extent can ASV systems naturally defend against spoofing attacks?
Jee-Weon Jung, Xin Wang 0037, Nicholas W. D. Evans, Shinji Watanabe 0001, Hye-Jin Shim, Hemlata Tak, Siddhant Arora, Junichi Yamagishi, Joon Son Chung
INTERSPEECH2
2024 Speaker Detection by the Individual Listener and the Crowd: Parametric Models Applicable to Bonafide and Deepfake Speech
Tomi Kinnunen, Rosa González Hautamäki, Xin Wang 0037, Junichi Yamagishi
INTERSPEECH3
2024 Spoof Diarization: "What Spoofed When" in Partially Spoofed Audio
Lin Zhang 0054, Xin Wang 0037, Erica Cooper, Mireia Díez, Federico Landini, Nicholas W. D. Evans, Junichi Yamagishi
INTERSPEECH2
2024 Poster Abstract: Intrusion Detection for In-vehicle Networks Based on Parc-net Architecture
abstract
The Controller Area Network (CAN) serves as a pivotal communication protocol for Electronic Control Units (ECUs) in modern automotive systems. However, the increasing interconnectivity and sophistication of these ECUs introduce significant vulnerabilities, rendering in-vehicle networks susceptible to a variety of cyber threats. This work presents a multi-class classification model for intrusion detection within vehicle CAN networks, utilizing Bidirectional Long Short-Term Memory (BILSTM) and Parc-Net architectures. We also introduce a novel feature extraction module that computes the rate of ID changes, significantly enhancing the model's detection capabilities. The proposed method was evaluated on the CAR-HACKING dataset, demonstrating remarkable performance. The proposed Intrusion Detection System (IDS) greatly contributes to the security of vehicle CAN networks by facilitating real-time detection and localization of potential intrusions.
Mingming Tan, Heng Zhang 0001, Xin Wang 0037, Ming Li 0026, Meng Huang 0003, Jian Zhang 0082
MSN3
2024 Spoofing-Aware Speaker Verification Robust Against Domain and Channel Mismatches
abstract
In real-world applications, it is challenging to build a speaker verification system that is simultaneously robust against common threats, including spoofing attacks, channel mismatch, and domain mismatch. Traditional automatic speaker verification (ASV) systems often tackle these issues separately, leading to suboptimal performance when faced with simultaneous challenges. In this paper, we propose an integrated framework that incorporates pair-wise learning and spoofing attack simulation into the meta-learning paradigm to enhance robustness against these multifaceted threats. This novel approach employs an asymmetric dual-path model and a multi-task learning strategy to handle ASV, anti-spoofing, and spoofing-aware ASV tasks concurrently. A new testing dataset, CNComplex, is introduced to evaluate system performance under these combined threats. Experimental results demonstrate that our integrated model significantly improves performance over traditional ASV systems across various scenarios, showcasing its potential for real-world deployment. Additionally, the proposed framework’s ability to generalize across different conditions highlights its robustness and reliability, making it a promising solution for practical ASV applications.
Chang Zeng, Xiaoxiao Miao, Xin Wang 0037, Erica Cooper, Junichi Yamagishi
SLT3
2024 FedDue: Optimizing Personalized Federated Learning Through Dynamic Update Classifier
Dongrun Li, Xin Wang 0037, Yanhan Wang, Ming Yang 0023
WASA (1)2
2024 Joint speaker encoder and neural back-end model for fully end-to-end automatic speaker verification with multiple enrollment utterances
abstract
Conventional automatic speaker verification systems can usually be decomposed into a front-end model such as time delay neural network (TDNN) for extracting speaker embeddings and a back-end model such as statistics-based probabilistic linear discriminant analysis (PLDA) or neural network-based neural PLDA (NPLDA) for similarity scoring. However, the sequential optimization of the front-end and back-end models may lead to a local minimum, which theoretically prevents the whole system from achieving the best optimization. Although some methods have been proposed for jointly optimizing the two models, such as the generalized end-to-end (GE2E) model and NPLDA E2E model, most of these methods have not fully investigated how to model the intra-relationship between multiple enrollment utterances. In this paper, we propose a new E2E joint method for speaker verification especially designed for the practical scenario of multiple enrollment utterances. To leverage the intra-relationship among multiple enrollment utterances, our model comes equipped with frame-level and utterance-level attention mechanisms. Additionally, focal loss is utilized to balance the importance of positive and negative samples within a mini-batch and focus on the difficult samples during the training process. We also utilize several data augmentation techniques, including conventional noise augmentation using MUSAN and RIRs datasets and a unique speaker embedding-level mixup strategy for better optimization.
Chang Zeng, Xiaoxiao Miao, Xin Wang 0037, Erica Cooper, Junichi Yamagishi
Comput. Speech Lang.3
2024 Application of Prompt Learning Models in Identifying the Collaborative Problem Solving Skills in an Online Task
abstract
Collaborative problem solving (CPS) competence is considered one of the essential 21st-century skills. To facilitate the assessment and learning of CPS competence, researchers have proposed a series of frameworks to conceptualize CPS and explored ways to make sense of the complex processes involved in collaborative problem solving. However, encoding explicit behaviors into subskills within the frameworks of CPS skills is still a challenging task. Traditional studies have relied on manual coding to decipher behavioral data for CPS, but such coding methods can be very time-consuming and cannot support real-time analyses. Scholars have begun to explore approaches for constructing automatic coding models. Nevertheless, the existing models built using machine learning or deep learning techniques depend on a large amount of training data and have relatively low accuracy. To address these problems, this paper proposes a prompt-based learning pre-trained model. The model can achieve high performance even with limited training data. In this study, three experiments were conducted, and the results showed that our model not only produced the highest accuracy, macro F1 score, and kappa values on large training sets, but also performed the best on small training sets of the CPS behavioral data. The application of the proposed prompt-based learning pre-trained model contributes to the CPS skills coding task and can also be used for other CSCW coding tasks to replace manual coding.
Mengxiao Zhu 0001, Xin Wang 0037, Xiantao Wang, Wei Huang 0002
Proc. ACM Hum. Comput. Interact.2
2024 ZMM-TTS: Zero-Shot Multilingual and Multispeaker Speech Synthesis Conditioned on Self-Supervised Discrete Speech Representations
abstract
Neural text-to-speech (TTS) has achieved human-like synthetic speech for single-speaker, single-language synthesis. Multilingual TTS systems are limited to resource-rich languages due to the lack of large paired text and studio-quality audio data. TTS systems are typically built using a single speaker's voice, but there is growing interest in developing systems that can synthesize voices for new speakers using only a few seconds of their speech. This paper presents ZMM-TTS, a multilingual and multispeaker framework utilizing quantized latent speech representations from a large-scale, pre-trained, self-supervised model. Our paper combines text-based and speech-based self-supervised learning models for multilingual speech synthesis. Our proposed model has zero-shot generalization ability not only for unseen speakers but also for unseen languages. We have conducted comprehensive subjective and objective evaluations through a series of experiments. Our model has proven effective in terms of speech naturalness and similarity for both seen and unseen speakers in six high-resource languages. We also tested the efficiency of our method on two hypothetically low-resource languages. The results are promising, indicating that our proposed approach can synthesize audio that is intelligible and has a high degree of similarity to the target speaker's voice, even without any training data for the new, unseen language.
Xin Wang 0037, Erica Cooper, Dan Wells, Longbiao Wang, Jianwu Dang 0001, Korin Richmond, Junichi Yamagishi
IEEE ACM Trans. Audio Speech Lang. Process.2
2024 The VoicePrivacy 2022 Challenge: Progress and Perspectives in Voice Anonymisation
abstract
The VoicePrivacy Challenge promotes the development of voice anonymisation solutions for speech technology. In this paper we present a systematic overview and analysis of the second edition held in 2022. We describe the voice anonymisation task and datasets used for system development and evaluation, present the different attack models used for evaluation, and the associated objective and subjective metrics. We describe three anonymisation baselines, provide a summary description of the anonymisation systems developed by challenge participants, and report objective and subjective evaluation results for all. In addition, we describe post-evaluation analyses and a summary of related work reported in the open literature. Results show that solutions based on voice conversion better preserve utility, that an alternative which combines automatic speech recognition with synthesis achieves greater privacy, and that a privacy-utility trade-off remains inherent to current anonymisation solutions. Finally, we present our ideas and priorities for future VoicePrivacy Challenge editions.
Michele Panariello, Natalia A. Tomashenko, Xin Wang 0037, Xiaoxiao Miao, Pierre Champion, Hubert Nourtel, Massimiliano Todisco, Nicholas W. D. Evans, Emmanuel Vincent 0001, Junichi Yamagishi
IEEE ACM Trans. Audio Speech Lang. Process.3
2024 Blockchain-Based Proxy-Oriented Data Integrity Checking Mechanism in Cloud-Assisted Intelligent Transportation Systems
abstract
Cloud-assisted intelligent transportation systems depend on cloud computing to provide powerful computing capabilities and big data storage services. As precise intelligent traffic control and dispatch policies are heavily based on real-time traffic information (e.g., unmanned driving test information), any altered data may cause severe consequences. The integrity of outsourced critical traffic control data has been the most concerning security issue. To this end, a lightweight proxy-oriented data integrity checking mechanism has been devised, without incurring substantial certificates management. The mechanism enables a data manager in traffic information control center to delegate the proxy to produce the signatures of encrypted data and outsource them to the cloud server, dramatically alleviating the work intensity of the data manager. By integrating blockchain into the mechanism, it gives assistance to the data manager for validating malicious integrity checking behaviors. The comprehensive security analysis and performance evaluation demonstrate the feasibility of the mechanism in the deployment of cloud-assisted intelligent transportation systems.
Yuan Zhang 0006, Xin Wang 0037
IEEE Trans. Intell. Transp. Syst.4
2023 Modelling Attention Levels with Ocular Responses in a Speech-in-Noise Recall Task
abstract
We applied state-space modelling technique to estimate the cognitive workload of a speech-in-noise (SIN) recall task, based on participants’ oculo-motor responses to speech signals. We estimated common latent attention levels in 15 time bins and observed temporal changes between pupillary dilations and saccade frequencies, given that the both conditions were independent. We also compared two speech type factors (natural vs. synthetic) and three levels of signal-to-noise (-1dB, -3dB, and -5dB) using the estimated parameter distribution. The comparison of experimental factors provided us with insights into differences in participants’ processing of spoken information during a SIN recall task.
Mateusz Dubiel, Minoru Nakayama, Xin Wang 0037
ETRA3
2023 Hiding Speaker's Sex in Speech Using Zero-Evidence Speaker Representation in an Analysis/Synthesis Pipeline
abstract
The use of modern vocoders in an analysis/synthesis pipeline allows us to investigate high-quality voice conversion that can be used for privacy purposes. Here, we propose to transform the speaker embedding and the pitch in order to hide the sex of the speaker. ECAPA-TDNN-based speaker representation fed into a HiFiGAN vocoder is protected using a neural-discriminant analysis approach, which is consistent with the zero-evidence concept of privacy. This approach significantly reduces the information in speech related to the speaker’s sex while preserving speech content and some consistency in the resulting protected voices.
Paul-Gauthier Noé, Xiaoxiao Miao, Xin Wang 0037, Junichi Yamagishi, Jean-François Bonastre, Driss Matrouf
ICASSP3
2023 Can Knowledge of End-to-End Text-to-Speech Models Improve Neural Midi-to-Audio Synthesis Systems?
abstract
With the similarity between music and speech synthesis from symbolic input and the rapid development of text-to-speech (TTS) techniques, it is worthwhile to explore ways to improve the MIDI-to-audio performance by borrowing from TTS techniques. In this study, we analyze the shortcomings of a TTS-based MIDI-to-audio system and improve it in terms of feature computation, model selection, and training strategy, aiming to synthesize highly natural-sounding audio. Moreover, we conducted an extensive model evaluation through listening tests, pitch measurement, and spectrogram analysis. This work demonstrates not only synthesis of highly natural music but offers a thorough analytical approach and useful outcomes for the community. Our code, pre-trained models, supplementary materials, and audio samples are open sourced at https://github.com/nii-yamagishilab/midi-to-audio.
Xuan Shi, Erica Cooper, Xin Wang 0037, Junichi Yamagishi, Shri Narayanan
ICASSP3
2023 Spoofed Training Data for Speech Spoofing Countermeasure Can Be Efficiently Created Using Neural Vocoders
abstract
A good training set for speech spoofing countermeasures requires diverse TTS and VC spoofing attacks, but generating TTS and VC spoofed trials for a target speaker may be technically demanding. Instead of using full-fledged TTS and VC systems, this study uses neural-network-based vocoders to do copy-synthesis on bona fide utterances. The output data can be used as spoofed data. To make better use of pairs of bona fide and spoofed data, this study introduces a contrastive feature loss that can be plugged into the standard training criterion. On the basis of the bona fide trials from the ASVspoof 2019 logical access training set, this study empirically compared a few training sets created in the proposed manner using a few neural non-autoregressive vocoders. Results on multiple test sets suggest good practices such as fine-tuning neural vocoders using bona fide data from the target domain. The results also demonstrated the effectiveness of the contrastive feature loss. Combining the best practices, the trained CM achieved overall competitive performance. Its EERs on the ASVspoof 2021 hidden subsets also outperformed the top-1 challenge submission.
Xin Wang 0037, Junichi Yamagishi
ICASSP1
2023 Neural Network-Based Safety Optimization Control for Constrained Discrete-Time Systems
abstract
This paper proposes a constraint-aware safety control approach via adaptive dynamic programming (ADP) to address the control optimization issues for discrete-time systems subjected to state constraints. First, the constrained control framework is established via the primal-dual approach with the modified Lagrangian based on the relaxed barrier function. Herein, the Lagrangian multiplier is designed to achieve the trade-off between optimization performance and state constraints. In addition, the sub optimality of the dual method is built by proving that the dual gap can be arbitrarily small. And the value iteration algorithm is utilized to implement the constraint-aware ADP controller. Furthermore, the weight estimation error is proved to be bounded when the learning rate satisfies a given sufficient condition. Finally, numerical simulation proves the effectiveness and superiority of the proposed method.
Shangwei Zhao, Ming Yang 0023, Xin Wang 0037
IECON5
2023 Towards Single Integrated Spoofing-aware Speaker Verification Embeddings
Sung Hwan Mun, Hye-Jin Shim, Hemlata Tak, Xin Wang 0037, Xuechen Liu 0001, Md. Sahidullah, Myeonghun Jeong, Min Hyun Han, Massimiliano Todisco, Kong-Aik Lee, Junichi Yamagishi, Nicholas W. D. Evans, Tomi Kinnunen, Nam Soo Kim, Jee-Weon Jung
INTERSPEECH4
2023 Improving Generalization Ability of Countermeasures for New Mismatch Scenario by Combining Multiple Advanced Regularization Terms
Chang Zeng, Xin Wang 0037, Xiaoxiao Miao, Erica Cooper, Junichi Yamagishi
INTERSPEECH2
2023 Range-Based Equal Error Rate for Spoof Localization
Lin Zhang 0054, Xin Wang 0037, Erica Cooper, Nicholas W. D. Evans, Junichi Yamagishi
INTERSPEECH2
2023 Data-driven control for dynamic quantized nonlinear systems with state constraints based on barrier functions
Shangwei Zhao, Xin Wang 0037, Ming Yang 0023
Inf. Sci.3
2023 Using iterative adaptation and dynamic mask for child speech extraction under real-world multilingual conditions
Shi Cheng 0001, Jun Du 0002, Shutong Niu, Alejandrina Cristià, Xin Wang 0037, Qing Wang 0008, Chin-Hui Lee 0001
Speech Commun.5
2023 ASVspoof 2021: Towards Spoofed and Deepfake Speech Detection in the Wild
abstract
Benchmarking initiatives support the meaningful comparison of competing solutions to prominent problems in speech and language processing. Successive benchmarking evaluations typically reflect a progressive evolution from ideal lab conditions towards to those encountered in the wild. ASVspoof, the spoofing and deepfake detection initiative and challenge series, has followed the same trend. This article provides a summary of the ASVspoof 2021 challenge and the results of 54 participating teams that submitted to the evaluation phase. For the logical access (LA) task, results indicate that countermeasures are robust to newly introduced encoding and transmission effects. Results for the physical access (PA) task indicate the potential to detect replay attacks in real, as opposed to simulated physical spaces, but a lack of robustness to variations between simulated and real acoustic environments. The Deepfake (DF) task, new to the 2021 edition, targets solutions to the detection of manipulated, compressed speech data posted online. While detection solutions offer some resilience to compression effects, they lack generalization across different source datasets. In addition to a summary of the top-performing systems for each task, new analyses of influential data factors and results for hidden data subsets, the article includes a review of post-challenge results, an outline of the principal challenge limitations and a road-map for the future of ASVspoof.
Xuechen Liu 0001, Xin Wang 0037, Md. Sahidullah, Jose Patino 0001, Héctor Delgado, Tomi Kinnunen, Massimiliano Todisco, Junichi Yamagishi, Nicholas W. D. Evans, Andreas Nautsch, Kong-Aik Lee
IEEE ACM Trans. Audio Speech Lang. Process.2
2023 Speaker Anonymization Using Orthogonal Householder Neural Network
abstract
Speaker anonymization aims to conceal a speaker's identity while preserving content information in speech. Current mainstream neural-network speaker anonymization systems disentangle speech into prosody-related, content, and speaker representations. The speaker representation is then anonymized by a selection-based speaker anonymizer that uses a mean vector over a set of randomly selected speaker vectors from an external pool of English speakers. However, the resulting anonymized vectors are subject to severe privacy leakage against powerful attackers, reduction in speaker diversity, and language mismatch problems for unseen-language speaker anonymization. To generate diverse, language-neutral speaker vectors, this paper proposes an anonymizer based on an orthogonal Householder neural network (OHNN). Specifically, the OHNN acts like a rotation to transform the original speaker vectors into anonymized speaker vectors, which are constrained to follow the distribution over the original speaker vector space. A basic classification loss is introduced to ensure that anonymized speaker vectors from different speakers have unique speaker identities. To further protect speaker identities, an improved classification loss and similarity loss are used to push original-anonymized sample pairs away from each other. Experiments on VoicePrivacy Challenge datasets in English and theAISHELL-3dataset in Mandarin demonstrate the proposed anonymizer's effectiveness.
Xiaoxiao Miao, Xin Wang 0037, Erica Cooper, Junichi Yamagishi, Natalia A. Tomashenko
IEEE ACM Trans. Audio Speech Lang. Process.2
2023 The PartialSpoof Database and Countermeasures for the Detection of Short Fake Speech Segments Embedded in an Utterance
abstract
Automatic speaker verification is susceptible to various manipulations and spoofing, such as text-to-speech synthesis, voice conversion, replay, tampering, adversarial attacks, and so on. We consider a new spoofing scenario called "Partial Spoof" (PS) in which synthesized or transformed speech segments are embedded into a bona fide utterance. While existing countermeasures (CMs) can detect fully spoofed utterances, there is a need for their adaptation or extension to the PS scenario. We propose various improvements to construct a significantly more accurate CM that can detect and locate short-generated spoofed speech segments at finer temporal resolutions. First, we introduce newly developed self-supervised pre-trained models as enhanced feature extractors. Second, we extend our PartialSpoof database by adding segment labels for various temporal resolutions. Since the short spoofed speech segments to be embedded by attackers are of variable length, six different temporal resolutions are considered, ranging from as short as 20 ms to as large as 640 ms. Third, we propose a new CM that enables the simultaneous use of the segment-level labels at different temporal resolutions as well as utterance-level labels to execute utterance- and segment-level detection at the same time. We also show that the proposed CM is capable of detecting spoofing at the utterance level with low error rates in the PS scenario as well as in a related logical access (LA) scenario. The equal error rates of utterance-level detection on the PartialSpoof database and ASVspoof 2019 LA database were 0.77 and 0.90%, respectively.
Lin Zhang 0054, Xin Wang 0037, Erica Cooper, Nicholas W. D. Evans, Junichi Yamagishi
IEEE ACM Trans. Audio Speech Lang. Process.2
2023 Enabling Anonymous Authorized Auditing Over Keyword-Based Searchable Ciphertexts in Cloud Storage Systems
abstract
Cloud storage provides Data Owners (DOs) with flexible data storage and management services, but simultaneously poses various security concerns, of which the integrity of outsourced data, as the most important security issue, determines the widespread use of cloud storage services. In this article, we present an$\mathbf {\overline{A}}$nonymous$\mathbf {\overline{A}}$uthorized$\mathbf {\overline{A}}$uditing scheme over$\mathbf {\overline{K}}$eyword-based$\mathbf {\overline{S}}$earchable$\mathbf {\overline{C}}$iphertexts (AAA-KSC) in cloud storage systems. In AAA-KSC, we devise an anonymous authorization mechanism, only the designated Third Party Auditor (TPA) could decrypt DO's identity and execute auditing tasks. In particular, by utilizing the verification metadata, without breaking the privacy of DO's concerned keyword, TPA could check the integrity of outsourced ciphertexts containing the specific keyword. More attractively, we design an identity-based homomorphic signature algorithm that makes use of the polynomial encapsulation mechanism to tackle with large scale data files, and thus AAA-KSC enables TPA to fulfil integrity verification with nearly constant computational costs, which are independent of the number of encrypted files that contain the same keyword. We provide the correctness, and prove that AAA-KSC is adequately secure. The performance evaluation experimentally demonstrates the high efficiency and feasibility of AAA-KSC in the real applications of cloud storage systems.
Xin Wang 0037, Yinbin Miao, Jingting Xue
IEEE Trans. Serv. Comput.1
2022 Estimating the Confidence of Speech Spoofing Countermeasure
abstract
Conventional speech spoofing countermeasures (CMs) are designed to make a binary decision on an input trial. However, a CM trained on a closed-set database is theoretically not guaranteed to perform well on unknown spoofing attacks. In some scenarios, an alternative strategy is to let the CM defer a decision when it is not confident. The question is then how to estimate a CM’s confidence regarding an input trial. We investigated a few confidence estimators that can be easily plugged into a neural-network-based CM. On the ASVspoof2019 logical access database, the results demonstrate that an energy-based estimator and a neural-network-based one achieved acceptable performance in identifying unknown attacks in the test set. On a test set with additional unknown attacks and bona fide trials from other databases, the confidence estimators performed moderately well, and the CMs better discriminated bona fide and spoofed trials that had a high confidence score. Additional results also revealed the difficulty in enhancing a confidence estimator by adding unknown attacks to the training set.
Xin Wang 0037, Junichi Yamagishi
ICASSP1
2022 Attention Back-End for Automatic Speaker Verification with Multiple Enrollment Utterances
abstract
Probabilistic linear discriminant analysis (PLDA) or cosine similarity have been widely used in traditional speaker verification systems as back-end techniques to measure pairwise similarities. To make better use of multiple enrollment utterances, we propose a novel attention back-end model that can be used for both textindependent (TI) and text-dependent (TD) speaker verification, and we use scaled-dot self-attention and feed-forward self-attention networks as architectures that learn the intra-relationships of enrollment utterances. To verify the proposed model, we conduct a series of experiments on the CNCeleb and VoxCeleb datasets by combining it with several state-of-the-art speaker encoders including TDNN and ResNet. Experimental results obtained using multiple enrollment utterances on CNCeleb show that the proposed attention back-end model leads to lower EER and minDCF scores than its PLDA and cosine similarity counterparts for each speaker encoder, and an experiment on VoxCeleb demonstrates that our model can be used even for a single enrollment case.
Chang Zeng, Xin Wang 0037, Erica Cooper, Xiaoxiao Miao, Junichi Yamagishi
ICASSP2
2022 Analyzing Language-Independent Speaker Anonymization Framework under Unseen Conditions
abstract
In our previous work, we proposed a language-independent speaker anonymization system based on self-supervised learning models.Although the system can anonymize speech data of any language, the anonymization was imperfect, and the speech content of the anonymized speech was distorted.This limitation is more severe when the input speech is from a domain unseen in the training data.This study analyzed the bottleneck of the anonymization system under unseen conditions.It was found that the domain (e.g., language and channel) mismatch between the training and test data affected the neural waveform vocoder and anonymized speaker vectors, which limited the performance of the whole system.Increasing the training data diversity for the vocoder was found to be helpful to reduce its implicit language and channel dependency.Furthermore, a simple correlation-alignment-based domain adaption strategy was found to be significantly effective to alleviate the mismatch on the anonymized speaker vectors.Audio samples 1 and source code 2 are available online.
Xiaoxiao Miao, Xin Wang 0037, Erica Cooper, Junichi Yamagishi, Natalia A. Tomashenko
INTERSPEECH2
2022 Investigating Active-Learning-Based Training Data Selection for Speech Spoofing Countermeasure
abstract
Training a speech spoofing countermeasure (CM) that generalizes to various unseen test data is challenging. Methods such as data augmentation and self-supervised learning can help, but the imperfect CM performance still calls for additional strategies. This paper investigates CM training using active learning (AL) to select useful training data from a large pool set, which is an unexplored area for speech anti-spoofing. Existing AL methods are compared to select useful data from a large pool set. A new AL method is also proposed that actively removes useless data from a pool. Experiments demonstrate that an energy-score-based AL method and the proposed data-removing method outperformed our strong baseline, and the relative reduction in detection error rates was higher than 40% on multiple test sets. Furthermore, compared with a top-line method that blindly used the whole pool set for training, the two AL-based CMs used less training data and achieved better or similar performance.
Xin Wang 0037, Junichi Yamagishi
SLT1
2022 The VoicePrivacy 2020 Challenge: Results and findings
Natalia A. Tomashenko, Xin Wang 0037, Emmanuel Vincent 0001, Jose Patino 0001, Brij Mohan Lal Srivastava, Paul-Gauthier Noé, Andreas Nautsch, Nicholas W. D. Evans, Junichi Yamagishi, Benjamin O'Brien, Anaïs Chanclu, Jean-François Bonastre, Massimiliano Todisco, Mohamed Maouche
Comput. Speech Lang.2
2022 Privacy and Utility of X-Vector Based Speaker Anonymization
abstract
We study the scenario where individuals (speakers) contribute to the publication of an anonymized speech corpus. Datausersleverage this public corpus for downstream tasks, e.g., training an automatic speech recognition (ASR) system, whileattackersmay attempt to de-anonymize it using auxiliary knowledge. Motivated by this scenario, speaker anonymization aims to conceal speaker identity while preserving the quality and usefulness of speech data. In this article, we study x-vector based speaker anonymization, the leading approach in the VoicePrivacy Challenge, which converts the speaker’s voice into that of a random pseudo-speaker. We show that the strength of anonymization varies significantly depending on how the pseudo-speaker is chosen. We explore four design choices for this step: the distance metric between speakers, the region of speaker space where the pseudo-speaker is picked, its gender, and whether to assign it to one or all utterances of the original speaker. We assess the quality of anonymization from the perspective of the three actors involved in our threat model, namely the speaker, the user and the attacker. To measure privacy and utility, we use respectively the linkability score achieved by the attackers and the decoding word error rate achieved by an ASR model trained on the anonymized data. Experiments on LibriSpeech show that the best combination of design choices yields state-of-the-art performance in terms of both privacy and utility. Experiments on Mozilla Common Voice further show that it guarantees the same anonymization level against re-identification attacks among 50 speakers as original speech among 20,000 speakers.
Brij Mohan Lal Srivastava, Mohamed Maouche, Md. Sahidullah, Emmanuel Vincent 0001, Aurélien Bellet, Marc Tommasi, Natalia A. Tomashenko, Xin Wang 0037, Junichi Yamagishi
IEEE ACM Trans. Audio Speech Lang. Process.8
2022 Conditional Anonymous Certificateless Public Auditing Scheme Supporting Data Dynamics for Cloud Storage Systems
abstract
Cloud computing provides users with convenient data storage services, which simultaneously poses various security concerns, the integrity of outsourced data has been termed as one of the most concerning security issues. Certificateless public auditing not only enables a third-party auditor (TPA) to check data integrity, but also avoids complex certificates management and inherent drawbacks of key escrow. Up to date, a few certificateless public auditing schemes have been proposed, without providing fast data dynamics or protecting the identity privacy of users. In this paper, we propose a lightweight conditional anonymous certificateless public auditing (CACPA) scheme, supporting much faster data dynamics in cloud storage systems. Based on a homomorphic hash function, we design a certificateless signature and integrate it into the construction of CACPA, reducing the computational costs of TPA substantially. CACPA achieves conditional identity privacy preservation, anyone cannot infer the real identity of a user based on outsourced data, only the private key generator (PKG) can revoke the users when some misbehaviors occur. We provide security analysis of CACPA, and conduct performance evaluation demonstrating the lightweight advantages of CACPA, and therefore it is suitable for auditors with resource-constrained mobile devices.
Xin Wang 0037, Dawu Gu, Jingting Xue
IEEE Trans. Netw. Serv. Manag.2
2021 How Similar or Different is Rakugo Speech Synthesizer to Professional Performers?
abstract
We have been working on speech synthesis for rakugo (a traditional Japanese form of verbal entertainment similar to one-person stand-up comedy) toward speech synthesis that authentically entertains audiences. In this paper, we propose a novel evaluation methodology using synthesized rakugo speech and real rakugo speech uttered by professional performers of three different ranks. The naturalness of the synthesized speech was comparable to that of the human speech, but the synthesized speech entertained listeners less than the performers of any rank. However, we obtained some interesting insights into challenges to be solved in order to achieve a truly entertaining rakugo synthesizer. For example, naturalness was not the most important factor, even though it has generally been emphasized as the most important point to be evaluated in the conventional speech synthesis field. More important factors were the understandability of the content and distinguishability of the characters in the rakugo story, both of which the synthesized rakugo speech was relatively inferior at as compared with the professional performers. We also found that fundamental frequency (fo) modeling should at least be further improved to better entertain audiences. These results show important steps to reaching authentically entertaining speech synthesis.
Shuhei Kato, Yusuke Yasuda, Xin Wang 0037, Erica Cooper, Junichi Yamagishi
ICASSP3
2021 End-to-End Text-to-Speech Using Latent Duration Based on VQ-VAE
abstract
Explicit duration modeling is a key to achieving robust and efficient alignment in text-to-speech synthesis (TTS). We propose a new TTS framework using explicit duration modeling that incorporates duration as a discrete latent variable to TTS and enables joint optimization of whole modules from scratch. We formulate our method based on conditional VQ-VAE to handle discrete duration in a variational autoencoder and provide a theoretical explanation to justify our method. In our framework, a connectionist temporal classification (CTC) -based force aligner acts as the approximate posterior, and text-to-duration works as the prior in the variational autoencoder. We evaluated our proposed method with a listening test and compared it with other TTS methods based on soft-attention or explicit duration modeling. The results showed that our systems rated between soft-attention-based methods (Transformer-TTS, Tacotron2) and explicit duration modeling-based methods (Fastspeech).
Yusuke Yasuda, Xin Wang 0037, Junichi Yamagishi
ICASSP2
2021 Visualizing Classifier Adjacency Relations: A Case Study in Speaker Verification and Voice Anti-Spoofing
abstract
Whether it be for results summarization, or the analysis of classifier fusion, some means to compare different classifiers can often provide illuminating insight into their behaviour, (dis)similarity or complementarity. We propose a simple method to derive 2D representation from detection scores produced by an arbitrary set of binary classifiers in response to a common dataset. Based upon rank correlations, our method facilitates a visual comparison of classifiers with arbitrary scores and with close relation to receiver operating characteristic (ROC) and detection error trade-off (DET) analyses. While the approach is fully versatile and can be applied to any detection task, we demonstrate the method using scores produced by automatic speaker verification and voice anti-spoofing systems. The former are produced by a Gaussian mixture model system trained with VoxCeleb data whereas the latter stem from submissions to the ASVspoof 2019 challenge.
Tomi Kinnunen, Andreas Nautsch, Md. Sahidullah, Nicholas W. D. Evans, Xin Wang 0037, Massimiliano Todisco, Héctor Delgado, Junichi Yamagishi, Kong-Aik Lee
Interspeech5
2021 A Comparative Study on Recent Neural Spoofing Countermeasures for Synthetic Speech Detection
abstract
A great deal of recent research effort on speech spoofing countermeasures has been invested into back-end neural networks and training criteria.We contribute to this effort with a comparative perspective in this study.Our comparison of countermeasure models on the ASVspoof 2019 logical access task takes into account recently proposed margin-based training criteria, widely used front ends, and common strategies to deal with varied-length input trials.We also measured intramodel differences through multiple training-evaluation rounds with random initialization.Our statistical analysis demonstrates that the performance of the same model may be significantly different when just changing the random initial seed.Thus, we recommend similar analysis or multiple training-evaluation rounds for further research on the database.Despite the intramodel differences, we observed a few promising techniques such as the average pooling to process varied-length inputs and a new hyper-parameter-free loss function.The two techniques led to the best single model in our experiment, which achieved an equal error rate of 1.92% and was significantly different in statistical sense from most of the other experimental models.
Xin Wang 0037, Junichi Yamagishi
Interspeech1
2021 An Initial Investigation for Detecting Partially Spoofed Audio
abstract
International audience
Lin Zhang 0054, Xin Wang 0037, Erica Cooper, Junichi Yamagishi, Jose Patino 0001, Nicholas W. D. Evans
Interspeech2
2021 Denoising-and-Dereverberation Hierarchical Neural Vocoder for Robust Waveform Generation
abstract
This paper presents a denoising and dereverberation hierarchical neural vocoder (DNR-HiNet) to convert noisy and reverberant acoustic features into a clean speech waveform. We implement it mainly by modifying the amplitude spectrum predictor (ASP) in the original HiNet vocoder. This modified denoising and dereverberation ASP (DNR-ASP) can predict clean log amplitude spectra (LAS) from input degraded acoustic features. To achieve this, the DNR-ASP first predicts the noisy and reverberant LAS, noise LAS related to the noise information, and room impulse response related to the reverberation information then performs initial denoising and dereverberation. The initial processed LAS are then enhanced by another neural network as the final clean LAS. To further improve the quality of the generated clean LAS, we also introduce a bandwidth extension model and frequency resolution extension model in the DNR-ASP. The experimental results indicate that the DNR-HiNet vocoder was able to generate a denoised and dereverberated waveform given noisy and reverberant acoustic features and outperformed the original HiNet vocoder and a few other neural vocoders. We also applied the DNR-HiNet vocoder to speech enhancement tasks, and its performance was competitive with several advanced speech enhancement methods.
Yang Ai, Xin Wang 0037, Junichi Yamagishi, Zhen-Hua Ling
SLT3
2021 Investigation of learning abilities on linguistic features in sequence-to-sequence text-to-speech synthesis
Yusuke Yasuda, Xin Wang 0037, Junichi Yamagishi
Comput. Speech Lang.2
2020 Zero-Shot Multi-Speaker Text-To-Speech with State-Of-The-Art Neural Speaker Embeddings
abstract
While speaker adaptation for end-to-end speech synthesis using speaker embeddings can produce good speaker similarity for speakers seen during training, there remains a gap for zero-shot adaptation to unseen speakers. We investigate multi-speaker modeling for end-to-end text-to-speech synthesis and study the effects of different types of state-of-the-art neural speaker embeddings on speaker similarity for unseen speakers. Learnable dictionary encoding-based speaker embeddings with angular softmax loss can improve equal error rates over x-vectors in a speaker verification task; these embeddings also improve speaker similarity and naturalness for unseen speakers when used for zero-shot adaptation to new speakers in end-to-end speech synthesis.
Erica Cooper, Cheng-I Lai, Yusuke Yasuda, Fuming Fang, Xin Wang 0037, Nanxin Chen, Junichi Yamagishi
ICASSP5
2020 A Study of Child Speech Extraction Using Joint Speech Enhancement and Separation in Realistic Conditions
abstract
In this paper, we design a novel joint framework of speech enhancement and speech separation for child speech extraction in realistic conditions, targeting the problem of extracting child speech from daily conversations in BabyTrain mega corpus. To the best of our knowledge, it is the first discussion of a feasible method for child speech extraction in realistic conditions. First, we make detailed analysis of the BabyTrain mega corpus, which is recorded in adverse environments. We observe problems of background noises, reverberations and child speech that is partially obscured by adult speech (for instance due to speaker overlap but also imitation by the adult). Motivated by this, we conduct a joint framework of speech enhancement and speech separation for child speech extraction. To measure the extraction results in realistic conditions, we propose several objective measurements to evaluate the performance of the our system, which is different from those commonly used for simulation data. Compared with the unprocessed approach and classification approach, our proposed approach can yield the best performance among all subsets of BabyTrain.
Xin Wang 0037, Jun Du 0002, Alejandrina Cristià, Lei Sun 0010, Chin-Hui Lee 0001
ICASSP1
2020 Effect of Choice of Probability Distribution, Randomness, and Search Methods for Alignment Modeling in Sequence-to-Sequence Text-to-Speech Synthesis Using Hard Alignment
abstract
Sequence-to-sequence text-to-speech (TTS) is dominated by soft-attention-based methods. Recently, hard-attention-based methods have been proposed to prevent fatal alignment errors, but their sampling method of discrete alignment is poorly investigated. This research investigates various combinations of sampling methods and probability distributions for alignment transition modeling in a hard-alignment-based sequence-to-sequence TTS method called SSNT-TTS. We clarify the common sampling methods of discrete variables including greedy search, beam search, and random sampling from a Bernoulli distribution in a more general way. Furthermore, we introduce the binary Concrete distribution to model discrete variables more properly. The results of a listening test shows that deterministic search is more preferable than stochastic search, and the binary Concrete distribution is robust with stochastic search for natural alignment transition.
Yusuke Yasuda, Xin Wang 0037, Junichi Yamagishi
ICASSP2
2020 Transferring Neural Speech Waveform Synthesizers to Musical Instrument Sounds Generation
abstract
Recent neural waveform synthesizers such as WaveNet, WaveG-low, and the neural-source-filter (NSF) model have shown good performance in speech synthesis despite their different methods of waveform generation. The similarity between speech and music audio synthesis techniques suggests interesting avenues to explore in terms of the best way to apply speech synthesizers in the music domain. This work compares three neural synthesizers used for musical instrument sounds generation under three scenarios: training from scratch on music data, zero-shot learning from the speech domain, and fine-tuning-based adaptation from the speech to the music domain. The results of a large-scale perceptual test demonstrated that the performance of three synthesizers improved when they were pre-trained on speech data and fine-tuned on music data, which indicates the usefulness of knowledge from speech data for music audio generation. Among the synthesizers, WaveGlow showed the best potential in zero-shot learning while NSF performed best in the other scenarios and could generate samples that were perceptually close to natural audio.
Yi Zhao 0006, Xin Wang 0037, Lauri Juvela, Junichi Yamagishi
ICASSP2
2020 Reverberation Modeling for Source-Filter-Based Neural Vocoder
abstract
This paper presents a reverberation module for source-filter-based neural vocoders that improves the performance of reverberant effect modeling. This module uses the output waveform of neural vocoders as an input and produces a reverberant waveform by convolving the input with a room impulse response (RIR). We propose two approaches to parameterizing and estimating the RIR. The first approach assumes a global time-invariant (GTI) RIR and directly learns the values of the RIR on a training dataset. The second approach assumes an utterance-level time-variant (UTV) RIR, which is invariant within one utterance but varies across utterances, and uses another neural network to predict the RIR values. We add the proposed reverberation module to the phase spectrum predictor (PSP) of a HiNet vocoder and jointly train the model. Experimental results demonstrate that the proposed module was helpful for modeling the reverberation effect and improving the perceived quality of generated reverberant speech. The UTV-RIR was shown to be more robust than the GTI-RIR to unknown reverberation conditions and achieved a perceptually better reverberation effect.
Yang Ai, Xin Wang 0037, Junichi Yamagishi, Zhen-Hua Ling
INTERSPEECH2
2020 Design Choices for X-Vector Based Speaker Anonymization
abstract
International audience
Brij Mohan Lal Srivastava, Natalia A. Tomashenko, Xin Wang 0037, Emmanuel Vincent 0001, Junichi Yamagishi, Mohamed Maouche, Aurélien Bellet, Marc Tommasi
INTERSPEECH3
2020 Introducing the VoicePrivacy Initiative
abstract
The VoicePrivacy initiative aims to promote the development of privacy preservation tools for speech technology by gathering a new community to define the tasks of interest and the evaluation methodology, and benchmarking solutions through a series of challenges. In this paper, we formulate the voice anonymization task selected for the VoicePrivacy 2020 Challenge and describe the datasets used for system development and evaluation. We also present the attack models and the associated objective and subjective evaluation metrics. We introduce two anonymization baselines and report objective evaluation results.
Natalia A. Tomashenko, Brij Mohan Lal Srivastava, Xin Wang 0037, Emmanuel Vincent 0001, Andreas Nautsch, Junichi Yamagishi, Nicholas W. D. Evans, Jose Patino 0001, Jean-François Bonastre, Paul-Gauthier Noé, Massimiliano Todisco
INTERSPEECH3
2020 Using Cyclic Noise as the Source Signal for Neural Source-Filter-Based Speech Waveform Model
abstract
Neural source-filter (NSF) waveform models generate speech waveforms by morphing sine-based source signals through dilated convolution in the time domain.Although the sinebased source signals help the NSF models to produce voiced sounds with specified pitch, the sine shape may constrain the generated waveform when the target voiced sounds are less periodic.In this paper, we propose a more flexible source signal called cyclic noise, a quasi-periodic noise sequence given by the convolution of a pulse train and a static random noise with a trainable decaying rate that controls the signal shape.We further propose a masked spectral loss to guide the NSF models to produce periodic voiced sounds from the cyclic noise-based source signal.Results from a large-scale listening test demonstrated the effectiveness of the cyclic noise and the masked spectral loss on speaker-independent NSF models in copy-synthesis experiments on the CMU ARCTIC database.
Xin Wang 0037, Junichi Yamagishi
INTERSPEECH1
2020 Fine-Grained Similarity Measurement between Educational Videos and Exercises
abstract
In online learning systems, measuring the similarity between educational videos and exercises is a fundamental task with great application potentials. In this paper, we explore to measure the fine-grained similarity by leveraging multimodal information. The problem remains pretty much open due to several domain-specific characteristics. First, unlike general videos, educational videos contain not only graphics but also text and formulas, which have a fixed reading order. Both spatial and temporal information embedded in the frames should be modeled. Second, there are semantic associations between adjacent video segments. The semantic associations will affect the similarity and different exercises usually focus on the related context of different ranges. Third, the fine-grained labeled data for training the model is scarce and costly. To tackle the aforementioned challenges, we propose VENet to measure the similarity at both video-level and segment-level by just exploiting the video-level labeled data. Extensive experimental results on real-world data demonstrate the effectiveness of VENet.
Xin Wang 0037, Wei Huang 0002, Qi Liu 0003, Yu Yin 0002, Zhenya Huang, Le Wu 0001, Jianhui Ma 0001
ACM Multimedia1
2020 ASVspoof 2019: A large-scale public database of synthesized, converted and replayed speech
Xin Wang 0037, Junichi Yamagishi, Massimiliano Todisco, Héctor Delgado, Andreas Nautsch, Nicholas W. D. Evans, Md. Sahidullah, Ville Vestman, Tomi Kinnunen, Kong-Aik Lee, Lauri Juvela, Paavo Alku, Yu-Huai Peng, Hsin-Te Hwang, Yu Tsao 0001, Hsin-Min Wang, Sébastien Le Maguer, Zhen-Hua Ling
Comput. Speech Lang.1
2020 Tandem Assessment of Spoofing Countermeasures and Automatic Speaker Verification: Fundamentals
abstract
Recent years have seen growing efforts to develop spoofing countermeasures (CMs) to protect automatic speaker verification (ASV) systems from being deceived by manipulated or artificial inputs. The reliability of spoofing CMs is typically gauged using the equal error rate (EER) metric. The primitive EER fails to reflect application requirements and the impact of spoofing and CMs upon ASV and its use as a primary metric in traditional ASV research has long been abandoned in favour of risk-based approaches to assessment. This paper presents several new extensions to the tandem detection cost function (t-DCF), a recent risk-based approach to assess the reliability of spoofing CMs deployed in tandem with an ASV system. Extensions include a simplified version of the t-DCF with fewer parameters, an analysis of a special case for a fixed ASV system, simulations which give original insights into its interpretation and new analyses using the ASVspoof 2019 database. It is hoped that adoption of the t-DCF for the CM assessment will help to foster closer collaboration between the anti-spoofing and ASV research communities.
Tomi Kinnunen, Héctor Delgado, Nicholas W. D. Evans, Kong-Aik Lee, Ville Vestman, Andreas Nautsch, Massimiliano Todisco, Xin Wang 0037, Md. Sahidullah, Junichi Yamagishi, Douglas A. Reynolds
IEEE ACM Trans. Audio Speech Lang. Process.8
2020 Neural Source-Filter Waveform Models for Statistical Parametric Speech Synthesis
abstract
Neural waveform models have demonstrated better performance than conventional vocoders for statistical parametric speech synthesis. One of the best models, called WaveNet, uses an autoregressive (AR) approach to model the distribution of waveform sampling points, but it has to generate a waveform in a time-consuming sequential manner. Some new models that use inverse-autoregressive flow (IAF) can generate a whole waveform in a one-shot manner but require either a larger amount of training time or a complicated model architecture plus a blend of training criteria. As an alternative to AR and IAF-based frameworks, we propose a neural source-filter (NSF) waveform modeling framework that is straightforward to train and fast to generate waveforms. This framework requires three components to generate waveforms: a source module that generates a sine-based signal as excitation, a non-AR dilated-convolution-based filter module that transforms the excitation into a waveform, and a conditional module that pre-processes the input acoustic features for the source and filter modules. This framework minimizes spectral-amplitude distances for model training, which can be efficiently implemented using short-time Fourier transform routines. As an initial NSF study, we designed three NSF models under the proposed framework and compared them with WaveNet using our deep learning toolkit. It was demonstrated that the NSF models generated waveforms at least 100 times faster than our WaveNet-vocoder, and the quality of the synthetic speech from the best NSF model was comparable to that from WaveNet on a large single-speaker Japanese speech corpus.
Xin Wang 0037, Shinji Takaki, Junichi Yamagishi
IEEE ACM Trans. Audio Speech Lang. Process.1
2020 A Vector Quantized Variational Autoencoder (VQ-VAE) Autoregressive Neural F0 Model for Statistical Parametric Speech Synthesis
abstract
Recurrent neural networks (RNNs) can predict fundamental frequency (F0) for statistical parametric speech synthesis systems, given linguistic features as input. However, these models assume conditional independence between consecutive F0values, given the RNN state. In a previous study, we proposed autoregressive (AR) neural F0models to capture the causal dependency of successive F0values. In subjective evaluations, a deep AR model (DAR) outperformed an RNN. Here, we propose a Vector Quantized Variational Autoencoder (VQ-VAE) neural F0model that is both more efficient and more interpretable than the DAR. This model has two stages: one uses the VQ-VAE framework to learn a latent code for the F0contour of each linguistic unit, and other learns to map from linguistic features to latent codes. In contrast to the DAR and RNN, which process the input linguistic features frame-by-frame, the new model converts one linguistic feature vector into one latent code for each linguistic unit. The new model achieves better objective scores than the DAR, has a smaller memory footprint and is computationally faster. Visualization of the latent codes for phones and moras reveals that each latent code represents an F0shape for a linguistic unit.
Xin Wang 0037, Shinji Takaki, Junichi Yamagishi, Simon King 0001, Keiichi Tokuda
IEEE ACM Trans. Audio Speech Lang. Process.1
2019 Audiovisual Speaker Conversion: Jointly and Simultaneously Transforming Facial Expression and Acoustic Characteristics
abstract
An audiovisual speaker conversion method is presented for simultaneously transforming the facial expressions and voice of a source speaker into those of a target speaker. Transforming the facial and acoustic features together makes it possible for the converted voice and facial expressions to be highly correlated and for the generated target speaker to appear and sound natural. It uses three neural networks: a conversion network that fuses and transforms the facial and acoustic features, a waveform generation network that produces the waveform from both the converted facial and acoustic features, and an image reconstruction network that outputs an RGB facial image also based on both the converted features. The results of experiments using an emotional audiovisual database showed that the proposed method achieved significantly higher naturalness compared with one that separately transformed acoustic and facial features.
Fuming Fang, Xin Wang 0037, Junichi Yamagishi, Isao Echizen
ICASSP2
2019 STFT Spectral Loss for Training a Neural Speech Waveform Model
abstract
This paper proposes a new loss using short-time Fourier transform (STFT) spectra for the aim of training a high-performance neural speech waveform model that predicts raw continuous speech waveform samples directly. Not only amplitude spectra but also phase spectra obtained from generated speech waveforms are used to calculate the proposed loss. We also mathematically show that training of the waveform model on the basis of the proposed loss can be interpreted as maximum likelihood training that assumes the amplitude and phase spectra of generated speech waveforms following Gaussian and von Mises distributions, respectively. Furthermore, this paper presents a simple network architecture as the speech waveform model, which is composed of uni-directional long short-term memories (LSTMs) and an auto-regressive structure. Experimental results showed that the proposed neural model synthesized high-quality speech waveforms.
Shinji Takaki, Toru Nakashika, Xin Wang 0037, Junichi Yamagishi
ICASSP3
2019 Neural Source-filter-based Waveform Model for Statistical Parametric Speech Synthesis
abstract
Neural waveform models such as the WaveNet are used in many recent text-to-speech systems, but the original WaveNet is quite slow in waveform generation because of its autoregressive (AR) structure. Although faster non-AR models were recently reported, they may be prohibitively complicated due to the use of a distilling training method and the blend of other disparate training criteria. This study proposes a non-AR neural source-filter waveform model that can be directly trained using spectrum-based training criteria and the stochastic gradient descent method. Given the input acoustic features, the proposed model first uses a source module to generate a sine-based excitation signal and then uses a filter module to transform the excitation signal into the output speech waveform. Our experiments demonstrated that the proposed model generated waveforms at least 100 times faster than the AR WaveNet and the quality of its synthetic speech is close to that of speech generated by the AR WaveNet. Ablation test results showed that both the sine-wave excitation signal and the spectrum-based training criteria were essential to the performance of the proposed model.
Xin Wang 0037, Shinji Takaki, Junichi Yamagishi
ICASSP1
2019 Investigation of Enhanced Tacotron Text-to-speech Synthesis Systems with Self-attention for Pitch Accent Language
abstract
End-to-end speech synthesis is a promising approach that directly converts raw text to speech. Although it was shown that Tacotron2 outperforms classical pipeline systems with regards to naturalness in English, its applicability to other languages is still unknown. Japanese could be one of the most difficult languages for which to achieve end-to-end speech synthesis, largely due to its character diversity and pitch accents. Therefore, state-of-the-art systems are still based on a traditional pipeline framework that requires a separate text analyzer and duration model. Towards end-to-end Japanese speech synthesis, we extend Tacotron to systems with self-attention to capture long-term dependencies related to pitch accents and compare their audio quality with classical pipeline systems under various conditions to show their pros and cons. In a large-scale listening test, we investigated the impacts of the presence of accentual-type labels, the use of force or predicted alignments, and acoustic features used as local condition parameters of the Wavenet vocoder. Our results reveal that although the proposed systems still do not match the quality of a top-line pipeline system for Japanese, we show important stepping stones towards end-to-end Japanese speech synthesis.
Yusuke Yasuda, Xin Wang 0037, Shinji Takaki, Junichi Yamagishi
ICASSP2
2019 Joint Training Framework for Text-to-Speech and Voice Conversion Using Multi-Source Tacotron and WaveNet
abstract
We investigated the training of a shared model for both textto-speech (TTS) and voice conversion (VC) tasks.We propose using an extended model architecture of Tacotron, that is a multi-source sequence-to-sequence model with a dual attention mechanism as the shared model for both the TTS and VC tasks.This model can accomplish these two different tasks respectively according to the type of input.An end-to-end speech synthesis task is conducted when the model is given text as the input while a sequence-to-sequence voice conversion task is conducted when it is given the speech of a source speaker as the input.Waveform signals are generated by using WaveNet, which is conditioned by using a predicted mel-spectrogram.We propose jointly training a shared model as a decoder for a target speaker that supports multiple sources.Listening experiments show that our proposed multi-source encoder-decoder model can efficiently achieve both the TTS and VC tasks.
Mingyang Zhang 0003, Xin Wang 0037, Fuming Fang, Haizhou Li 0001, Junichi Yamagishi
INTERSPEECH2
2019 MOSNet: Deep Learning-Based Objective Assessment for Voice Conversion
abstract
Existing objective evaluation metrics for voice conversion (VC) are not always correlated with human perception. Therefore, training VC models with such criteria may not effectively improve naturalness and similarity of converted speech. In this paper, we propose deep learning-based assessment models to predict human ratings of converted speech. We adopt the convolutional and recurrent neural network models to build a mean opinion score (MOS) predictor, termed as MOSNet. The proposed models are tested on large-scale listening test results of the Voice Conversion Challenge (VCC) 2018. Experimental results show that the predicted scores of the proposed MOSNet are highly correlated with human MOS ratings at the system level while being fairly correlated with human MOS ratings at the utterance level. Meanwhile, we have modified MOSNet to predict the similarity scores, and the preliminary results show that the predicted scores are also fairly correlated with human ratings. These results confirm that the proposed models could be used as a computational evaluator to measure the MOS of VC systems to reduce the need for expensive human rating.
Chen-Chou Lo, Szu-Wei Fu, Wen-Chin Huang, Xin Wang 0037, Junichi Yamagishi, Yu Tsao 0001, Hsin-Min Wang
INTERSPEECH4
2019 Training Multi-Speaker Neural Text-to-Speech Systems Using Speaker-Imbalanced Speech Corpora
abstract
When the available data of a target speaker is insufficient to train a high quality speaker-dependent neural text-to-speech (TTS) system, we can combine data from multiple speakers and train a multi-speaker TTS model instead. Many studies have shown that neural multi-speaker TTS model trained with a small amount data from multiple speakers combined can generate synthetic speech with better quality and stability than a speaker-dependent one. However when the amount of data from each speaker is highly unbalanced, the best approach to make use of the excessive data remains unknown. Our experiments showed that simply combining all available data from every speaker to train a multi-speaker model produces better than or at least similar performance to its speaker-dependent counterpart. Moreover by using an ensemble multi-speaker model, in which each subsystem is trained on a subset of available data, we can further improve the quality of the synthetic speech especially for underrepresented speakers whose training data is limited.
Hieu-Thi Luong, Xin Wang 0037, Junichi Yamagishi, Nobuyuki Nishizawa
INTERSPEECH2
2019 ASVspoof 2019: Future Horizons in Spoofed and Fake Audio Detection
abstract
ASVspoof, now in its third edition, is a series of community-led challenges which promote the development of countermeasures to protect automatic speaker verification (ASV) from the threat of spoofing. Advances in the 2019 edition include: (i) a consideration of both logical access (LA) and physical access (PA) scenarios and the three major forms of spoofing attack, namely synthetic, converted and replayed speech; (ii) spoofing attacks generated with state-of-the-art neural acoustic and waveform models; (iii) an improved, controlled simulation of replay attacks; (iv) use of the tandem detection cost function (t-DCF) that reflects the impact of both spoofing and countermeasures upon ASV reliability. Even if ASV remains the core focus, in retaining the equal error rate (EER) as a secondary metric, ASVspoof also embraces the growing importance of fake audio detection. ASVspoof 2019 attracted the participation of 63 research teams, with more than half of these reporting systems that improve upon the performance of two baseline spoofing countermeasures. This paper describes the 2019 database, protocols and challenge results. It also outlines major findings which demonstrate the real progress made in protecting against the threat of spoofing and fake audio.
Massimiliano Todisco, Xin Wang 0037, Ville Vestman, Md. Sahidullah, Héctor Delgado, Andreas Nautsch, Junichi Yamagishi, Nicholas W. D. Evans, Tomi Kinnunen, Kong-Aik Lee
INTERSPEECH2
2018 Cyborg Speech: Deep Multilingual Speech Synthesis for Generating Segmental Foreign Accent with Natural Prosody
abstract
We describe a new application of deep-learning-based speech synthesis, namely multilingual speech synthesis for generating controllable foreign accent. Specifically, we train a DBLSTM-based acoustic model on non-accented multilingual speech recordings from a speaker native in several languages. By copying durations and pitch contours from a pre-recorded utterance of the desired prompt, natural prosody is achieved. We call this paradigm “cyborg speech” as it combines human and machine speech parameters. Segmentally accented speech is produced by interpolating specific quinphone linguistic features towards phones from the other language that represent non-native mispronunciations. Experiments on synthetic American-English-accented Japanese speech show that subjective synthesis quality matches monolingual synthesis, that natural pitch is maintained, and that naturalistic phone substitutions generate output that is perceived as having an American foreign accent, even though only non-accented training data was used.
Gustav Eje Henter, Jaime Lorenzo-Trueba, Xin Wang 0037, Mariko Kondo, Junichi Yamagishi
ICASSP3
2018 Speech Waveform Synthesis from MFCC Sequences with Generative Adversarial Networks
abstract
This paper proposes a method for generating speech from filterbank mel frequency cepstral coefficients (MFCC), which are widely used in speech applications, such as ASR, but are generally considered unusable for speech synthesis. First, we predict fundamental frequency and voicing information from MFCCs with an autoregressive recurrent neural net. Second, the spectral envelope information contained in MFCCs is converted to all-pole filters, and a pitch-synchronous excitation model matched to these filters is trained. Finally, we introduce a generative adversarial network -based noise model to add a realistic high-frequency stochastic component to the modeled excitation signal. The results show that high quality speech reconstruction can be obtained, given only MFCC information at test time.
Lauri Juvela, Bajibabu Bollepalli, Xin Wang 0037, Hirokazu Kameoka, Manu Airaksinen, Junichi Yamagishi, Paavo Alku
ICASSP3
2018 A Comparison of Recent Waveform Generation and Acoustic Modeling Methods for Neural-Network-Based Speech Synthesis
abstract
Recent advances in speech synthesis suggest that limitations such as the lossy nature of the amplitude spectrum with minimum phase approximation and the over-smoothing effect in acoustic modeling can be overcome by using advanced machine learning approaches. In this paper, we build a framework in which we can fairly compare new vocoding and acoustic modeling techniques with conventional approaches by means of a large scale crowdsourced evaluation. Results on acoustic models showed that generative adversarial networks and an autoregressive (AR) model performed better than a normal recurrent network and the AR model performed best. Evaluation on vocoders by using the same AR acoustic model demonstrated that a Wavenet vocoder outperformed classical source-filter-based vocoders. Particularly, generated speech waveforms from the combination of AR acoustic model and Wavenet vocoder achieved a similar score of speech quality to vocoded speech.
Xin Wang 0037, Jaime Lorenzo-Trueba, Shinji Takaki, Lauri Juvela, Junichi Yamagishi
ICASSP1
2018 Investigating Accuracy of Pitch-accent Annotations in Neural Network-based Speech Synthesis and Denoising Effects
abstract
We investigated the impact of noisy linguistic features on the performance of a Japanese speech synthesis system based on neural network that uses WaveNet vocoder. We compared an ideal system that uses manually corrected linguistic features including phoneme and prosodic information in training and test sets against a few other systems that use corrupted linguistic features. Both subjective and objective results demonstrate that corrupted linguistic features, especially those in the test set, affected the ideal system’s performance significantly in a statistical sense due to a mismatched condition between the training and test sets. Interestingly, while an utterance-level Turing test showed that listeners had a difficult time differentiating synthetic speech from natural speech, it further indicated that adding noise to the linguistic features in the training set can partially reduce the effect of the mismatch, regularize the model, and help the system perform better when linguistic features of the test set are noisy. Index Terms: speech synthesis, deep neural network, Japanese prosody, WaveNet
Hieu-Thi Luong, Xin Wang 0037, Junichi Yamagishi, Nobuyuki Nishizawa
INTERSPEECH2
2018 Investigating very deep highway networks for parametric speech synthesis
Xin Wang 0037, Shinji Takaki, Junichi Yamagishi
Speech Commun.1
2018 Autoregressive Neural F0 Model for Statistical Parametric Speech Synthesis
abstract
Recurrent neural networks (RNNs) have been successfully used as fundamental frequency (F0) models for text-to-speech synthesis. However, this paper showed that a normal RNN may not take into account the statistical dependency of the F0 data across frames and consequently only generate noisy F0 contours when F0 values are sampled from the model. A better model may take into account the causal dependency of the current F0 datum on the previous frames' F0 data. One such model is the shallow autoregressive (AR) recurrent mixture density network (SAR) that we recently proposed. However, as this study showed, an SAR is equivalent to the combination of trainable linear filters and a conventional RNN. It is still weak for F0 modeling. To better model the temporal dependency in F0 contours, we propose a deep AR model (DAR). On the basis of an RNN, this DAR propagates the previous frame's F0 value through the RNN, which allows nonlinear AR dependency to be achieved. We also propose F0 quantization and data dropout strategies for the DAR. Experiments on a Japanese corpus demonstrated that this DAR can generate appropriate F0 contours by using the random-sampling-based generation method, which is impossible for the baseline RNN and SAR. When a conventional mean-based generation method was used in the proposed DAR and other experimental models, the DAR generated accurate and less oversmoothed F0 contours and achieved a better mean-opinion-score in a subjective evaluation test.
Xin Wang 0037, Shinji Takaki, Junichi Yamagishi
IEEE ACM Trans. Audio Speech Lang. Process.1
2017 An autoregressive recurrent mixture density network for parametric speech synthesis
abstract
Neural-network-based generative models, such as mixture density networks, are potential solutions for speech synthesis. In this paper we follow this path and propose a recurrent mixture density network that incorporates a trainable autoregressive model. An advantage of incorporating an autoregressive model is that the time dependency within acoustic feature trajectories can be modeled without using the conventional dynamic features. More interestingly, experiments show that this autoregressive model learns to be a filter that emphasizes the high frequency components of the target acoustic feature trajectories in the training stage. In the synthesis stage, it boosts the low frequency components of the generated feature trajectories and hence increases their global variance. Experimental results show that the proposed model achieved higher likelihood on the training data and generated speech with better quality than other models when dynamic features were not utilized in any model.
Xin Wang 0037, Shinji Takaki, Junichi Yamagishi
ICASSP1
2017 Principles for Learning Controllable TTS from Annotated and Latent Variation
abstract
For building flexible and appealing high-quality speech synthesisers, it is desirable to be able to accommodate and reproduce fine variations in vocal expression present in natural speech. Synthesisers can enable control over such output properties by adding adjustable control parameters in parallel to their text input. If not annotated in training data, the values of these control inputs can be optimised jointly with the model parameters. We describe how this established method can be seen as approximate maximum likelihood and MAP inference in a latent variable model. This puts previous ideas of (learned) synthesiser inputs such as sentence-level control vectors on a more solid theoretical footing. We furthermore extend the method by restricting the latent variables to orthogonal subspaces via a sparse prior. This enables us to learn dimensions of variation present also within classes in coarsely annotated speech. As an example, we train an LSTM-based TTS system to learn nuances in emotional expression from a speech database annotated with seven different acted emotions. Listening tests show that our proposal successfully can synthesise speech with discernible differences in expression within each emotion, without compromising the recognisability of synthesised emotions compared to an identical system without learned nuances.
Gustav Eje Henter, Jaime Lorenzo-Trueba, Xin Wang 0037, Junichi Yamagishi
INTERSPEECH3
2017 An RNN-Based Quantized F0 Model with Multi-Tier Feedback Links for Text-to-Speech Synthesis
abstract
A recurrent-neural-network-based F0 model for text-to-speech (TTS) synthesis that generates F0 contours given textual features is proposed. In contrast to related F0 models, the proposed one is designed to learn the temporal correlation of F0 contours at multiple levels. The frame-level correlation is covered by feeding back the F0 output of the previous frame as the additional input of the current frame; meanwhile, the correlation over long-time spans is similarly modeled but by using F0 features aggregated over the phoneme and syllable. Another difference is that the output of the proposed model is not the interpolated continuous-valued F0 contour but rather a sequence of discrete symbols, including quantized F0 levels and a symbol for the unvoiced condition. By using the discrete F0 symbols, the proposed model avoids the influence of artificially interpolated F0 curves. Experiments demonstrated that the proposed F0 model, which was trained using a dropout strategy, generated smooth F0 contours with relatively better perceived quality than those from baseline RNN models.
Xin Wang 0037, Shinji Takaki, Junichi Yamagishi
INTERSPEECH1
2016 A full training framework of cross-stream dependence modelling for HMM-based singing voice synthesis
abstract
A cross-stream dependence modelling (CSDM) method has been proposed to model the dependence of spectral distributions on F0 observations for hidden Markov model (HMM) based speech synthesis. However, this method incorporates CSDM only for the embedded training of HMM estimation while ignoring CSDM in the clustering of context-dependent HMMs. This paper applies CSDM to HMM-based singing voice synthesis and presents a decision-tree-based model clustering method with explicit CSDM. This method, in conjunction with the previous CSDM method, forms a full CSDM training framework. Experimental results demonstrate that this full CSDM training framework achieves better performance than the previous CSDM method and the baseline without CSDM in a singing voice synthesis task.
Xin Wang 0037, Minghui Dong, Zhen-Hua Ling
ICASSP1
2016 Using Text and Acoustic Features in Predicting Glottal Excitation Waveforms for Parametric Speech Synthesis with Recurrent Neural Networks
abstract
This work studies the use of deep learning methods to directly model glottal excitation waveforms from context dependent text features in a text-to-speech synthesis system. Glottal vocoding is integrated into a deep neural network-based text-to-speech framework where text and acoustic features can be flexibly used as both network inputs or outputs. Long short-term memory recurrent neural networks are utilised in two stages: first, in mapping text features to acoustic features and second, in predicting glottal waveforms from the text and/or acoustic features. Results show that using the text features directly yields similar quality to the prediction of the excitation from acoustic features, both outperforming a baseline system based on using a fixed glottal pulse for excitation generation.
Lauri Juvela, Xin Wang 0037, Shinji Takaki, Manu Airaksinen, Junichi Yamagishi, Paavo Alku
INTERSPEECH2
2016 Speech Enhancement for a Noise-Robust Text-to-Speech Synthesis System Using Deep Recurrent Neural Networks
abstract
Quality of text-to-speech voices built from noisy recordings is diminished. In order to improve it we propose the use of a recurrent neural network to enhance acoustic parameters prior to training. We trained a deep recurrent neural network using a parallel database of noisy and clean acoustics parameters as input and output of the network. The database consisted of multiple speakers and diverse noise conditions. We investigated using text-derived features as an additional input of the network. We processed a noisy database of two other speakers using this network and used its output to train an HMM acoustic text-to-synthesis model for each voice. Listening experiment results showed that the voice built with enhanced parameters was ranked significantly higher than the ones trained with noisy speech and speech that has been enhanced using a conventional enhancement system. The text-derived features improved results only for the female voice, where it was ranked as highly as a voice trained with clean speech.
Cassia Valentini-Botinhao, Xin Wang 0037, Shinji Takaki, Junichi Yamagishi
INTERSPEECH2
2016 Enhance the Word Vector with Prosodic Information for the Recurrent Neural Network Based TTS System
abstract
Word embedding, which is a dense and low-dimensional vector representation of word, is recently used to replace of the conventional prosodic context as an input feature to the acoustic model of a TTS system. However, these word vectors trained from text data may encode insufficient information related to speech. This paper presents a post-filtering approach to enhance the raw word vectors with prosodic information for the TTS task. Based on a publicly available speech corpus with manual prosodic annotation, a post-filter can be trained to transform the raw word vectors. Experiment shows that using the enhanced word vectors as an input to the neural network-based acoustic model improves the accuracy of the predicted F0 trajectory. Besides, we also show that the enhanced vectors provide better initial values than the raw vectors for error back-propagation of the network, which results in further improvement.
Xin Wang 0037, Shinji Takaki, Junichi Yamagishi
INTERSPEECH1
2016 Concept-to-Speech generation with knowledge sharing for acoustic modelling and utterance filtering
Xin Wang 0037, Zhen-Hua Ling, Li-Rong Dai 0001
Comput. Speech Lang.1
2014 Concept-to-speech generation by integrating syntagmatic features into HMM-based speech synthesis
Xin Wang 0037, Zhen-Hua Ling, Li-Rong Dai 0001
INTERSPEECH1
2013 An anisotropic diffusion filter based on multidirectional separability
Jianguo Wei, Xin Wang 0037, Wenhuan Lu, Qiang Fang 0003, Jianwu Dang 0001
INTERSPEECH3