VLDB 2026 Research / reviewers in the wild / expert
Tao Wang 0074
dblp:12/5838-74
· DBLP profile ↗
34ranked-venue papers
10as first author
28since 2021 · last 2026
0000-0003-1490-6973ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 25 · 6 first-author · 19 since 2021Artificial intelligence and machine learning · 20 · 8 first-author · 15 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SpeechPalette: A Comprehensive Speech Editing Method for Text-Based Speech Editing, One-Shot TTS and Attributes EditingabstractSpeech editing has garnered more and more attention due to its diverse applications. However, existing systems often require substantial manual effort or have limited capabilities in attribute editing, imposing significant constraints. In this work, we present SpeechPalette, a comprehensive high-quality speech editing method that allows users to easily modify various attributes of the selected speech segment according to their preferences. Specifically, the proposed model approaches speech editing from a decoupling perspective, disentangling critical information such as text, pitch, duration and more from the input speech. Then, reconstruction is achieved through a mask and prediction mechanism. Furthermore, we leverage a diffusion model to predict the residuals between the real and predicted speech, further enhancing synthesis quality. The proposed method not only excels at text-based speech editing but also handles tasks involving pitch and speed rate adjustments. Moreover, it also demonstrates remarkable performance in one-shot text-to-speech scenarios. While recent large-scale models achieve impressive synthesis quality through massive computational resources, SpeechPalette offers a balanced approach with explicit fine-grained control over speech attributes, practical deployment requirements, and competitive performance relative to similarly-sized systems. Experimental results across a range of tasks consistently demonstrate the superior performance of our method compared to baseline systems. Additionally, comprehensive ablation studies validate the effectiveness of our proposed approach. Tao Wang 0074, Jiangyan Yi, Ruibo Fu, Chunyu Qiang, Dading Chong, Dongyang Dai, Zhengqi Wen, Jianhua Tao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2025 | Mel-Refine: A Plug-and-Play Approach to Refine Mel-Spectrogram in Audio GenerationabstractMainstream Text-to-Audio (TTA) models that rely on Mel-spectrograms often struggle to generate audio with rich content, leading to blurred or incoherent outputs. This stems from an inability to model intricate spectral details and textures. We investigate the role of U-Net components in generation and find that high-frequency components in skip-connections and the backbone are crucial for texture, while low-frequency backbone components are vital for the denoising process. Based on this, we propose “Mel-Refine,” a plug-and-play approach that enhances Mel-spectrogram quality by adjusting component weights during inference. Our method requires no additional training or fine-tuning and is fully compatible with any diffusion-based TTA architecture. Experiments show that Mel-Refine boosts the performance of the latest TTA model, Tango2, by $25 \%$, demonstrating its effectiveness. Hongming Guo, Ruibo Fu, Yizhong Geng, Shuchen Shi, Tao Wang 0074, Chunyu Qiang, Ya Li 0001, Zhengqi Wen, Xuefei Liu, Chenxing Li |
ASRU | 5 |
| 2025 | DPI-TTS: Directional Patch Interaction for Fast-Converging and Style Temporal Modeling in Text-to-SpeechabstractIn recent years, speech diffusion models have advanced rapidly. Alongside the widely used U-Net architecture, transformer-based models such as the Diffusion Transformer (DiT) have also gained attention. However, current DiT speech models treat Mel spectrograms as general images, which overlooks the specific acoustic properties of speech. To address these limitations, we propose a method called Directional Patch Interaction for Text-to-Speech (DPI-TTS), which builds on DiT and achieves fast training without compromising accuracy. Notably, DPI-TTS employs a low-to-high frequency, frame-by-frame progressive inference approach that aligns more closely with acoustic properties, enhancing the naturalness of the generated speech. Additionally, we introduce a fine-grained style temporal modeling method that further improves speaker style similarity. Experimental results demonstrate that our method increases the training speed by nearly 2 times and significantly outperforms the baseline models. Ruibo Fu, Zhengqi Wen, Tao Wang 0074, Chunyu Qiang, Jianhua Tao 0001, Chenxing Li, Shuchen Shi, Yuankun Xie, Xuefei Liu, Guanjun Li |
ICASSP | 4 |
| 2025 | WMCodec: End-to-End Neural Speech Codec with Deep Watermarking for Authenticity VerificationabstractRecent advances in speech spoofing necessitate stronger verification mechanisms in neural speech codecs to ensure authenticity. Current methods embed numerical watermarks before compression and extract them from reconstructed speech for verification, but face limitations such as separate training processes for the watermark and codec, and insufficient cross-modal information integration, leading to reduced watermark imperceptibility, extraction accuracy, and capacity. To address these issues, we propose WMCodec, the first neural speech codec to jointly train compression-reconstruction and watermark embedding-extraction in an end-to-end manner, optimizing both imperceptibility and extractability of the watermark. Furthermore, We design an iterative Attention Imprint Unit (AIU) for deeper feature integration of watermark and speech, reducing the impact of quantization noise on the watermark. Experimental results show WMCodec outperforms AudioSeal with Encodec in most quality metrics for watermark imperceptibility and consistently exceeds both AudioSeal with Encodec and reinforced TraceableSpeech in extraction accuracy of watermark. At bandwidth of 6 kbps with a watermark capacity of 16 bps, WMCodec maintains over 99% extraction accuracy under common attacks, demonstrating strong robustness. Junzuo Zhou, Jiangyan Yi, Yong Ren 0006, Jianhua Tao 0001, Tao Wang 0074, Chuyuan Zhang |
ICASSP | 5 |
| 2025 | HeRo: A State Machine-Based, Fault-Tolerant Framework for Heterogeneous Multi-Robot CollaborationabstractHeterogeneous robots can work together to accomplish a variety of complex tasks and have shown great potential in many fields. There are many efforts to make robot task orchestration more efficient. However, current methods still have some limitations, including the lack of a high-level abstraction for programming method and fault handling mechanism. In this paper, we design a state machine-based, fault-tolerant framework for heterogeneous multi-robot collaboration named HeRo, to effectively support the development of heterogeneous multi-robot systems. HeRo has three key techniques: (1) a state machine-based programming language to flexibly model robot behaviors and tasks; (2) a state synchronization mechanism to achieve information exchange and maintain the consistency among heterogeneous robots in distributed environments; (3) a fault detection and recovery mechanism to monitor the system's runtime states and use Large Language Model (LLM) combined with Planning Domain Definition Language (PDDL) to enable automated recovery. We evaluate the effectiveness and fault recovery capability of the framework by setting up manufacturing task and fault scenarios with varying difficulty in the ARIAC simulation environment, achieving a 100% task completion rate, with low system overhead and flexible scalability. Guoquan Wu, Tao Wang 0074, Wei Chen 0018, Jun Wei 0001 |
ICRA | 3 |
| 2024 | Minimally-Supervised Speech Synthesis with Conditional Diffusion Model and Language Model: A Comparative Study of Semantic CodingabstractRecently, there has been a growing interest in text-to-speech (TTS) methods that can be trained with minimal supervision by combining two types of discrete speech representations and using two sequence-to-sequence tasks to decouple TTS. However, existing methods suffer from three problems: the high-frequency waveform distortion of discrete speech representations, the prosodic averaging problem caused by the duration prediction model in non-autoregressive frameworks, and difficulty in prediction due to the information redundancy and dimension explosion of existing semantic coding methods. To address these problems, three progressive methods are proposed. First, we propose Diff-LM-Speech, an autoregressive structure consisting of a language model and diffusion models, which models the semantic embedding into the mel-spectrogram based on a diffusion model to achieve higher audio quality. We also introduce a prompt encoder structure based on a variational autoencoder and a prosody bottleneck to improve prompt representation ability. Second, we propose Tetra-Diff-Speech, a non-autoregressive structure consisting of four diffusion model-based modules that design a duration diffusion model to achieve diverse prosodic expressions. Finally, we propose Tri-Diff-Speech, a non-autoregressive structure consisting of three diffusion model-based modules that verify the non-necessity of existing semantic coding models and achieve the best results. Experimental results show that our proposed methods outperform baseline methods. We provide a website with audio samples.1 Chunyu Qiang, Hao Li 0078, He Qu, Ruibo Fu, Tao Wang 0074, Longbiao Wang, Jianwu Dang 0001 |
ICASSP | 6 |
| 2024 | Learning Speech Representation from Contrastive Token-Acoustic PretrainingabstractFor fine-grained generation and recognition tasks such as minimally-supervised text-to-speech (TTS), voice conversion (VC), and automatic speech recognition (ASR), the intermediate representations extracted from speech should serve as a "bridge" between text and acoustic information, containing information from both modalities. The semantic content is emphasized, while the paralinguistic information such as speaker identity and acoustic details should be de-emphasized. However, existing methods for extracting fine-grained intermediate representations from speech suffer from issues of excessive redundancy and dimension explosion. Contrastive learning is a good method for modeling intermediate representations from two modalities. However, existing contrastive learning methods in the audio field focus on extracting global descriptive information for downstream audio classification tasks, making them unsuitable for TTS, VC, and ASR tasks. To address these issues, we propose a method named "Contrastive Token-Acoustic Pretraining (CTAP)", which uses two encoders to bring phoneme and speech into a joint multimodal space, learning how to connect phoneme and speech at the frame level. The CTAP model is trained on 210k speech and phoneme pairs, achieving minimally-supervised TTS, VC, and ASR. The proposed CTAP method offers a promising solution for fine-grained generation and recognition downstream tasks in speech processing. We provide a website with audio samples.1 Chunyu Qiang, Hao Li 0078, Yixin Tian, Ruibo Fu, Tao Wang 0074, Longbiao Wang, Jianwu Dang 0001 |
ICASSP | 5 |
| 2024 | Fewer-Token Neural Speech Codec with Time-Invariant CodesabstractLanguage model based text-to-speech (TTS) models, like VALL-E, have gained attention for their outstanding in-context learning capability in zero-shot scenarios. Neural speech codec is a critical component of these models, which can convert speech into discrete token representations. However, excessive token sequences from the codec may negatively affect prediction accuracy and restrict the progression of Language model based TTS models. To address this issue, this paper proposes a novel neural speech codec with time-invariant codes named TiCodec. By encoding and quantizing time-invariant information into a separate code, TiCodec can reduce the amount of frame-level information that needs encoding, effectively decreasing the number of tokens as codes of speech. Furthermore, this paper introduces a time-invariant encoding consistency loss to enhance the consistency of time-invariant code within an utterance, which can benefit the zero-shot TTS task. Experimental results demonstrate that TiCodec can not only enhance the quality of reconstruction speech with fewer tokens but also increase the similarity and naturalness, as well as reduce the word error rate of the synthesized speech by the TTS model. The code is publicly available at https://github.com/y-ren16/TiCodec. Yong Ren 0006, Tao Wang 0074, Jiangyan Yi, Jianhua Tao 0001, Chu Yuan Zhang, Junzuo Zhou |
ICASSP | 2 |
| 2024 | Multi-modal Adversarial Training for Zero-Shot Voice Cloning
John Janiczek, Dading Chong, Dongyang Dai, Arlo Faria, Tao Wang 0074, Yuzong Liu |
INTERSPEECH | 6 |
| 2024 | PPPR: Portable Plug-in Prompt Refiner for Text to Audio Generation
Shuchen Shi, Ruibo Fu, Zhengqi Wen, Jianhua Tao 0001, Tao Wang 0074, Chunyu Qiang, Xuefei Liu |
INTERSPEECH | 5 |
| 2024 | Residual Speaker Representation for One-Shot Voice ConversionabstractInternational audience Jiangyan Yi, Tao Wang 0074, Yong Ren 0006, Rongxiu Zhong, Zhengqi Wen, Jianhua Tao 0001 |
INTERSPEECH | 3 |
| 2024 | TraceableSpeech: Towards Proactively Traceable Text-to-Speech with Watermarking
Junzuo Zhou, Jiangyan Yi, Tao Wang 0074, Jianhua Tao 0001, Ye Bai 0001, Chu Yuan Zhang, Yong Ren 0006, Zhengqi Wen |
INTERSPEECH | 3 |
| 2024 | Emotion selectable end-to-end text-based speech editing
Tao Wang 0074, Jiangyan Yi, Ruibo Fu, Jianhua Tao 0001, Zhengqi Wen, Chu Yuan Zhang |
Artif. Intell. | 1 |
| 2024 | Assessing growth potential of careers with occupational mobility network and ensemble framework
Tao Wang 0074, Witold Pedrycz, Yanjie Song 0001 |
Eng. Appl. Artif. Intell. | 2 |
| 2024 | CFAD: A Chinese dataset for fake audio detection
Haoxin Ma, Jiangyan Yi, Chenglong Wang 0001, Xinrui Yan, Jianhua Tao 0001, Tao Wang 0074, Ruibo Fu |
Speech Commun. | 6 |
| 2023 | Slow-Fast Time Parameter Aggregation Network for Class-Incremental Lip ReadingabstractClass incremental learning has yet to be explored in the field of lip-reading, which can circumvent data privacy issues and avoid the high training costs associated with joint training. In this paper, we introduce a benchmark for Class-Incremental Lip-Reading (CILR). To simultaneously improve the plasticity for new classes and stability for old classes in incremental learning, we propose a Slow-Fast Time Parameter Aggregation Network (TPAN) that decouples representation learning of new and old knowledge, taking into account the task characteristics of lip-reading. The TPAN comprises two dynamically evolving branches: one that uses fast gradient descent and the other employs slow momentum updates to retain old knowledge while adapting to new knowledge. Additionally, to achieve efficient knowledge transfer of the incremental model, we design a Hybrid Sequence-Distribution Distillation (HSDD) strategy to transfer knowledge in temporal feature view and classification probability view. We present a comprehensive comparison of the proposed method and previous state-of-the-art class incremental learning methods on the most commonly used lip-reading datasets LRW and LRW1000. The experimental result show that the proposed method can reduce the effect of catastrophic forgetting and improve the incremental accuracy. Xueyi Zhang 0001, Tao Wang 0074, Jun Tang 0001, Songyang Lao, Haizhou Li 0001 |
ACM Multimedia | 3 |
| 2023 | Adversarial Multi-Task Learning for Mandarin Prosodic Boundary Prediction With Multi-Modal EmbeddingsabstractProsodic boundaries are still crucial to the naturalness of end-to-end speech synthesis systems. This article proposes to use adversarial multi-task learning to predict prosodic boundaries. Adversarial multi-task learning is utilized to transfer knowledge from an auxiliary POS tagging task to a prosodic boundary prediction task. Furthermore, multi-modal embeddings are composed of contextual word and speech embedding features obtained from the pre-trained bidirectional encoder representations from transformers (BERT) model and Speech2Vec. We can utilize linguistic and acoustic information from large amounts of external text and speech data without prosodic boundary labels. At the inference stage, the prosodic boundary predicting model can use the syntactic features learnt from the POS tagging task without any extra computation cost due to only employing the prosodic boundary predicting task to decode. We conducted experiments on Mandarin datasets. The results show that the models using multi-modal embeddings from the pre-trained BERT and Speech2Vec outperform the models trained with single modal embedding. Furthermore, the models trained with adversarial training obtain further performance gains by up to 2.95% in$F_{1}$score. Jiangyan Yi, Jianhua Tao 0001, Ruibo Fu, Tao Wang 0074, Chu Yuan Zhang, Chenglong Wang 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2023 | Amer: A New Attribute-Missing Network Embedding ApproachabstractNetwork embedding which aims to learn a low dimensional representation of nodes is a powerful technique for network analysis. While network embedding for networks with complete attributes has been widely investigated, in many real-world applications the attributes of partial nodes are unobserved (i.e., missing) due to privacy concern or resource limit. Very recently, several network embedding methods have been proposed for attribute-missing networks. They first complete the missing attributes and then use the complemented network to learn network embedding. The parameters of these two processes cannot be adjusted by each other, resulting in compromised results. To address this problem, we propose a unified model in which the process of completing missing attributes and the process of learning embedding are not separated but closely intertwined. Being specific, completing missing attributes is under the guidance of learning network representation via mutual information maximization, and the complemented attributes directly enter network representation module which will generate further feedback for completing missing attributes. We further impose attribute-structure relationship constraint for completing missing attributes by designing a new generative adversarial networks (GANs) model. To the best of our knowledge, this is the first unified model for attribute-missing network embedding. Empirical results on real-world datasets show the superiority of our new method over other state-of-the-art methods on four network analysis tasks, including node classification, node clustering, link prediction, and network visualization. Di Jin 0001, Rui Wang 0102, Tao Wang 0074, Dongxiao He, Weiping Ding 0001, Longbiao Wang, Witold Pedrycz |
IEEE Trans. Cybern. | 3 |
| 2023 | Adversarial Representation Mechanism Learning for Network EmbeddingabstractNetwork embedding which is to learn a low dimensional representation of nodes in a network has been used in many network analysis tasks. Some network embedding methods, including those based on Generative Adversarial Networks (GAN) (a promising deep learning model), have been proposed recently. Existing GAN-based methods typically use GAN to learn a Gaussian distribution as a prior for network embedding, which makes it difficult to distinguish the node representation from Gaussian distribution. It did not apply the adversarial learning strategy on the representation mechanism but just on representation results. Thus, it does not make full use of the essential advantage of GAN, and leads to compromised performance of the method. To address this problem, we propose a novel adversarial learning framework consisting of three players for network embedding, which applies the adversarial learning strategy on the representation mechanism, called Adversarial representation mechanism GAN (ArmGAN). Specifically, the first two players, named encoder and competitor, aim to learn two different representation mechanisms (i.e., two ways projecting data onto latent space). They compete with each other to improve their representation mechanisms. The third player is the discriminator, which discriminate the representation mechanism of the encoder from that of the competitor. In addition, we design a perturbation strategy to produce fake networks from the original network, and feed the fake networks to the competitor to obtain a “fake” representation mechanism. We evaluated ArmGAN on a variety of tasks including node clustering, node classification, link prediction and visualization. Moreover, we compared ArmGAN with 10 state-of-the-art methods (including DGI, which is well-known for its high accuracy) on 7 real-world networks. The experimental results show the significant superiority of ArmGAN over the existing methods. Dongxiao He, Tao Wang 0074, Lu Zhai, Di Jin 0001, Liang Yang 0002, Zhiyong Feng 0002, Philip S. Yu |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2022 | Powerful Graph Convolutional Networks with Adaptive Propagation Mechanism for Homophily and HeterophilyabstractGraph Convolutional Networks (GCNs) have been widely applied in various fields due to their significant power on processing graph-structured data. Typical GCN and its variants work under a homophily assumption (i.e., nodes with same class are prone to connect to each other), while ignoring the heterophily which exists in many real-world networks (i.e., nodes with different classes tend to form edges). Existing methods deal with heterophily by mainly aggregating higher-order neighborhoods or combing the immediate representations, which leads to noise and irrelevant information in the result. But these methods did not change the propagation mechanism which works under homophily assumption (that is a fundamental part of GCNs). This makes it difficult to distinguish the representation of nodes from different classes. To address this problem, in this paper we design a novel propagation mechanism, which can automatically change the propagation and aggregation process according to homophily or heterophily between node pairs. To adaptively learn the propagation process, we introduce two measurements of homophily degree between node pairs, which is learned based on topological and attribute information, respectively. Then we incorporate the learnable homophily degree into the graph convolution framework, which is trained in an end-to-end schema, enabling it to go beyond the assumption of homophily. More importantly, we theoretically prove that our model can constrain the similarity of representations between nodes according to their homophily degree. Experiments on seven real-world datasets demonstrate that this new approach outperforms the state-of-the-art methods under heterophily or low homophily, and gains competitive performance under homophily. Tao Wang 0074, Di Jin 0001, Rui Wang 0102, Dongxiao He |
AAAI | 1 |
| 2022 | Context-Aware Mask Prediction Network for End-to-End Text-Based Speech EditingabstractThe text-based speech editor allows the editing of speech through intuitive cutting, copying, and pasting operations to speed up the process of editing speech. However, the major drawback of current systems is that edited speech often sounds unnatural and it is not obvious how to synthesize records according to a new word not appearing in the transcript. This paper proposes a novel end-to-end text-based speech editing method called context-aware mask prediction network (CampNet), which avoids the unnatural phenomenon caused by cut-copy-paste operation in the traditional method and can synthesize a new word not appearing in the transcript. Besides, three text-based speech editing operations based on CampNet are designed: deletion, replacement, and insertion. These operations can comprehensively cover different kinds of situations that text-based speech editing can face. The subjective and objective experiments on VCTK and LibriTTS data sets show that the speech editing results based on CampNet are better than TTS technology, manual editing, and VoCo method (the combination of speech synthesis and speech conversion). We also conducted detailed ablation experiments to explore the effect of the CampNet structure on its performance. Examples of generated speech can be found at https://hairuo55.github.io/CampNet-demo. Tao Wang 0074, Jiangyan Yi, Liqun Deng, Ruibo Fu, Jianhua Tao 0001, Zhengqi Wen |
ICASSP | 1 |
| 2022 | ADD 2022: the first Audio Deep Synthesis Detection ChallengeabstractAudio deepfake detection is an emerging topic, which was included in the ASVspoof 2021. However, the recent shared tasks have not covered many real-life and challenging scenarios. The first Audio Deep synthesis Detection challenge (ADD) was motivated to fill in the gap. The ADD 2022 includes three tracks: low-quality fake audio detection (LF), partially fake audio detection (PF) and audio fake game (FG). The LF track focuses on dealing with bona fide and fully fake utterances with various real-world noises etc. The PF track aims to distinguish the partially fake audio from the real. The FG track is a rivalry game, which includes two tasks: an audio generation task and an audio fake detection task. In this paper, we describe the datasets, evaluation metrics, and protocols. We also report major findings that reflect the recent advances in audio deepfake detection tasks. Jiangyan Yi, Ruibo Fu, Jianhua Tao 0001, Shuai Nie 0001, Haoxin Ma, Chenglong Wang 0001, Tao Wang 0074, Zhengkun Tian, Ye Bai 0001, Cunhang Fan, Shan Liang 0007, Shuai Zhang 0014, Xinrui Yan, Zhengqi Wen, Haizhou Li 0001 |
ICASSP | 7 |
| 2022 | NeuralDPS: Neural Deterministic Plus Stochastic Model With Multiband Excitation for Noise-Controllable Waveform GenerationabstractThe traditional vocoders have the advantages of high synthesis efficiency, strong interpretability, and speech editability, while the neural vocoders have the advantage of high synthesis quality. To combine the advantages of two vocoders, inspired by the traditional deterministic plus stochastic model, this paper proposes a novel neural vocoder named NeuralDPS which can retain high speech quality and acquire high synthesis efficiency and noise controllability. Firstly, this framework contains four modules: a deterministic source module, a stochastic source module, a neural V/UV decision module and a neural filter module. The input required by the vocoder is just the spectral parameter, which avoids the error caused by estimating additional parameters, such as F0. Secondly, to solve the problem that different frequency bands may have different proportions of deterministic components and stochastic components, a multiband excitation strategy is used to generate a more accurate excitation signal and reduce the neural filter’s burden. Thirdly, a method to control noise components of speech is proposed. In this way, the signal-to-noise ratio (SNR) of speech can be adjusted easily. Objective and subjective experimental results show that our proposed NeuralDPS vocoder can obtain similar performance with the WaveNet and it generates waveforms at least 280 times faster than the WaveNet vocoder. It is also 28% faster than WaveGAN’s synthesis efficiency on a single CPU core. We have also verified through experiments that this method can effectively control the noise components in the predicted speech and adjust the SNR of speech. Tao Wang 0074, Ruibo Fu, Jiangyan Yi, Jianhua Tao 0001, Zhengqi Wen |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2022 | CampNet: Context-Aware Mask Prediction for End-to-End Text-Based Speech EditingabstractThe text-based speech editor allows the editing of speech through intuitive cutting, copying, and pasting operations to speed up the process of editing speech. However, the major drawback of current systems is that edited speech often sounds unnatural due to cut-copy-paste operation. In addition, it is not obvious how to synthesize records according to a new word not appearing in the transcript, which often needs the help of text-to-speech (TTS) and voice conversion (VC) technology at the same time. This paper first proposes a novel end-to-end text-based speech editing method called context-aware mask prediction network (CampNet). The model can simulate the text-based speech editing process by randomly masking part of speech and then predicting the masked region by sensing the speech context. It can solve unnatural prosody in the edited region and synthesize the speech corresponding to the unseen words in the transcript. Secondly, for the possible operation of text-based speech editing, we design three text-based operations based on CampNet: deletion, insertion, and replacement. These operations can cover various situations of speech editing. Thirdly, to synthesize the speech corresponding to long text in insertion and replacement operations, a word-level autoregressive generation method is proposed, which can synthesize the speech of arbitrary length text. Fourthly, we propose a speaker adaptation method using only one sentence for CampNet and explore the ability of few-shot learning based on CampNet, which provides a new idea for speech forgery tasks. The subjective and objective experiments11Examples of generated speech can be found athttps://hairuo55.github.io/CampNet.on VCTK and LibriTTS datasets show that the speech editing results based on CampNet are better than TTS technology, manual editing, and VoCo method (the combination of TTS and VC). We also conduct detailed ablation experiments to explore the effect of the CampNet structure on its performance. Finally, the experiment shows that speaker adaptation with only one sentence can further improve the naturalness of speech editing for one-shot learning. Tao Wang 0074, Jiangyan Yi, Ruibo Fu, Jianhua Tao 0001, Zhengqi Wen |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2022 | Under-Display Camera Image Enhancement via Cascaded Curve EstimationabstractThe new trend of full-screen devices encourages manufacturers to position a camera behind a screen, i.e., the newly-defined Under-Display Camera (UDC). Therefore, UDC image restoration has been a new realistic single image enhancement problem. In this work, we propose a curve estimation network operating on the hue (H) and saturation (S) channels to perform adaptive enhancement for degraded images captured by UDCs. The proposed network aims to match the complicated relationship between the images captured by under-display and display-free cameras. To extract effective features, we cascade the proposed curve estimation network with sharing weights, and we introduce a spatial and channel attention module in each curve estimation network to exploit attention-aware features. In addition, we learn the curve estimation network in a semi-supervised manner to alleviate the restriction of the requirement for amounts of labeled images and improve the generalization ability for unseen degraded images in various realistic scenes. The semi-supervised network consists of a supervised branch trained on labeled data and an unsupervised branch trained on unlabeled data. To train the proposed model, we build a new dataset comprised of real-world labeled and unlabeled images. Extensive experiments demonstrate that our proposed algorithm performs favorably against state-of-the-art image enhancement methods for UDC images in terms of accuracy and speed, especially on ultra-high-definition (UHD) images. Jun Luo 0012, Wenqi Ren, Tao Wang 0074, Chongyi Li, Xiaochun Cao |
IEEE Trans. Image Process. | 3 |
| 2021 | Bi-Level Style and Prosody Decoupling Modeling for Personalized End-to-End Speech SynthesisabstractEnd-to-end framework can generate high-quality and high-similarity speech in the personalized speech synthesis task. However, the generalization of out-of-domain texts is still a challenging task. Limited target data leads to unacceptable errors and poor prosody and similarity performance of the synthetic speech. In this paper, we present a bi-level function decoupling framework to realise separate modeling and controlling for solving above problems. Firstly, on the style representation modeling level, compared with the conventional methods that use single embedding to model all the text dependent discrepancies, it is proposed that the speaker embedding and prosody embedding are modeled separately based on the reference audio and phonetic posteriorgram (PPG) by a multi-head attention mechanism. Secondly, on the model structure level, the decoder model structure is factored into average-net and adaptation-net, where the duration prosody controlling and speaker timbre imitation are mainly designed in relatively separate areas. Experimental results on Mandarin dataset show that the proposed methods lead to an improvement on both robustness, naturalness and similarity. Ruibo Fu, Jianhua Tao 0001, Zhengqi Wen, Jiangyan Yi, Tao Wang 0074, Chunyu Qiang |
ICASSP | 5 |
| 2021 | Prosody and Voice Factorization for Few-Shot Speaker Adaptation in the Challenge M2voc 2021abstractThe paper describes the CASIA speech synthesis system entry for challenge M2VoC 2021. The low similarity and naturalness of synthesized speech remains a challenging problem for speaker adaptation with few resources. Since the end-to-end acoustic model is too complex to interpret, overfitting will occur when training with few data. To prevent the model from overfitting, this paper proposes a novel speaker adaptation framework that decomposes the prosody and voice characteristics in the end-to-end model. A prosody control attention is proposed to control the phonemes’ duration of different speakers. To make the attention controlled by the prosody information, a set of phoneme-level transition tokens is auto-learned from the prosody encoder in our framework and these transition tokens can determine the duration of phonemes in the attention mechanism. Secondly, when we need to use small data set for speaker adaptation, we just need to adapt the speaker related prosody model and decoder, which can prevent the model from overfitting. Further, we use a data puring model to automatically optimize the quality of datasets. Experiments demonstrate the effectiveness of speaker adaptation based on our method, and we (team identifier is T03) get the top three results in competition M2VoC by using this framework. Tao Wang 0074, Ruibo Fu, Jiangyan Yi, Jianhua Tao 0001, Zhengqi Wen, Chunyu Qiang |
ICASSP | 1 |
| 2021 | Half-Truth: A Partially Fake Audio Detection DatasetabstractDiverse promising datasets have been designed to further the development of fake audio detection, such as ASVspoof databases.However, previous datasets ignore an attacking situation, in which the hacker hides some small fake clips in real speech audio.This poses a serious threat since that it is difficult to distinguish the small fake clip from the whole speech utterance.Therefore, this paper develops such a dataset for half-truth audio detection (HAD).Partially fake audio in the HAD dataset involves only changing a few words in an utterance.The audio of the words is generated with the very latest state-of-the-art speech synthesis technology.We can not only detect fake uttrances but also localize manipulated regions in a speech using this dataset.Some benchmark results are presented on this dataset.The results show that partially fake audio presents much more challenging than fully fake audio for fake audio detection.The HAD dataset is publicly available 1 . Jiangyan Yi, Ye Bai 0001, Jianhua Tao 0001, Haoxin Ma, Zhengkun Tian, Chenglong Wang 0001, Tao Wang 0074, Ruibo Fu |
Interspeech | 7 |
| 2020 | Focusing on Attention: Prosody Transfer and Adaptative Optimization Strategy for Multi-Speaker End-to-End Speech SynthesisabstractEnd-to-end speech synthesis can generate high-quality synthetic speech and achieve high similarity scores with low-resource adaptation data. However, the generalization of out-domain texts is still a challenging task. The limited adaptation data leads to unacceptable errors and the poor prosody performance of the synthetic speech. In this paper, we present two novel methods to handle the above problems by focusing on the attention. Firstly, compared with the conventional methods that extract prosody embeddings for conditioning input, a duration controller with feedback mechanism is proposed, which can control the states transition in the sequence-to-sequence model more directly and precisely. Secondly, to alleviate the unmatching text-audio pairs' impact on model, an adaptative optimization strategy which would consider the matching degree of the training sample is also proposed. Experimental results on Mandarin dataset show that proposed methods lead to an improvement on both robustness and overall naturalness. Ruibo Fu, Jianhua Tao 0001, Zhengqi Wen, Jiangyan Yi, Tao Wang 0074 |
ICASSP | 5 |
| 2020 | Dynamic Soft Windowing and Language Dependent Style Token for Code-Switching End-to-End Speech Synthesis
Ruibo Fu, Jianhua Tao 0001, Zhengqi Wen, Jiangyan Yi, Chunyu Qiang, Tao Wang 0074 |
INTERSPEECH | 6 |
| 2020 | Dynamic Speaker Representations Adjustment and Decoder Factorization for Speaker Adaptation in End-to-End Speech Synthesis
Ruibo Fu, Jianhua Tao 0001, Zhengqi Wen, Jiangyan Yi, Tao Wang 0074, Chunyu Qiang |
INTERSPEECH | 5 |
| 2020 | Non-Autoregressive End-to-End TTS with Coarse-to-Fine Decoding
Tao Wang 0074, Xuefei Liu, Jianhua Tao 0001, Jiangyan Yi, Ruibo Fu, Zhengqi Wen |
INTERSPEECH | 1 |
| 2020 | Bi-Level Speaker Supervision for One-Shot Speech Synthesis
Tao Wang 0074, Jianhua Tao 0001, Ruibo Fu, Jiangyan Yi, Zhengqi Wen, Chunyu Qiang |
INTERSPEECH | 1 |
| 2020 | Spoken Content and Voice Factorization for Few-Shot Speaker Adaptation
Tao Wang 0074, Jianhua Tao 0001, Ruibo Fu, Jiangyan Yi, Zhengqi Wen, Rongxiu Zhong |
INTERSPEECH | 1 |