Yujia Xiao

dblp:212/6415 · DBLP profile ↗
← Back
11ranked-venue papers
7as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 6 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 3 first-author · 1 since 2021Computer networks · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 TrafficAudio: Audio Representation for Lightweight Encrypted Traffic Classification in IoT
abstract
Encrypted traffic classification has become a crucial task for network management and security with the widespread adoption of encrypted protocols across the Internet and the Internet of Things. However, existing methods often rely on discrete representations and complex models, which leads to incomplete feature extraction, limited fine-grained classification accuracy, and high computational costs. To this end, we propose TrafficAudio, a novel encrypted traffic classification method based on audio representation. TrafficAudio comprises three modules: audio representation generation (ARG), audio feature extraction (AFE), and spatiotemporal traffic classification (STC). Specifically, the ARG module first represents raw network traffic as audio to preserve temporal continuity of traffic. Then, the audio is processed by the AFE module to compute low-dimensional Mel-frequency cepstral coefficients (MFCC), encoding both temporal and spectral characteristics. Finally, spatiotemporal features are extracted from MFCC through a parallel architecture of one-dimensional convolutional neural network and bidirectional gated recurrent unit layers, enabling fine-grained traffic classification. Experiments on five public datasets across six classification tasks demonstrate that TrafficAudio consistently outperforms ten state-of-the-art baselines, achieving accuracies of 99.74%, 98.40%, 99.76%, 99.25%, 99.77%, and 99.74%. Furthermore, TrafficAudio significantly reduces computational complexity, achieving reductions of 86.88% in floating-point operations and 43.15% of model parameters over the best-performing baseline.
Yilu Chen, Ye Wang 0015, Yujia Xiao, Lichen Liu, Yan Jia 0001, Zhaoquan Gu
IEEE Trans. Netw. Serv. Manag.4
2025 ZSVC: Zero-shot Style Voice Conversion with Disentangled Latent Diffusion Models and Adversarial Training
abstract
Style voice conversion aims to transform the speaking style of source speech into a desired style while keeping the original speaker’s identity. However, previous style voice conversion approaches primarily focus on well-defined domains such as emotional aspects, limiting their practical applications. In this study, we present ZSVC, a novel Zero-shot Style Voice Conversion approach that utilizes a speech codec and a latent diffusion model with speech prompting mechanism to facilitate in-context learning for speaking style conversion. To disentangle speaking style and speaker timbre, we introduce information bottleneck to filter speaking style in the source speech and employ Uncertainty Modeling Adaptive Instance Normalization (UMAdaIN) to perturb the speaker timbre in the style prompt. Moreover, we propose a novel adversarial training strategy to enhance in-context learning and improve style similarity. Experiments conducted on 44,000 hours of speech data demonstrate the superior performance of ZSVC in generating speech with diverse speaking styles in zero-shot scenarios.
Xinfa Zhu, Lei He 0005, Yujia Xiao, Xi Wang 0016, Xu Tan 0003, Sheng Zhao 0002, Lei Xie 0001
ICASSP3
2025 IoT-SCNet: Semi-Supervised Contrastive Network Traffic Images Learning for IoT Device Identification
abstract
The widespread deployment of Internet of Things (IoT) devices has made them vulnerable targets for cyber attacks, highlighting the great significance of IoT device identification for network security management. Existing studies have primarily focused on either manual extraction of excessive network traffic features or heavy reliance on labeled data. To address these limitations, we propose a novel IoT device identification approach (named IoT-SCNet) via semi-supervised contrastive learning of network traffic visual representations. Specifically, IoT-SCNet converts two packet-level features and raw traffic into network traffic images of each device, and applies three augmentation strategies designed for network traffic to construct semantically meaningful positive/negative image pairs. By deep neural networks automatically capturing discriminative patterns and a contrastive task, IoT-SCNet achieves effective identification of diverse IoT device types. Comprehensive evaluations across three benchmark datasets demonstrate the superior performance of IoT-SCNet, achieving remarkable identification accuracies of 99.83% on UNSW, 97.50% on Yourthings, and 99.34% on CIC IoT datasets.
Yujia Xiao, Yilu Chen, Lichen Liu, Ye Wang 0015, Zhaoquan Gu, Jie Liu 0001
IEEE Internet Things J.1
2024 Contrastive Context-Speech Pretraining for Expressive Text-to-Speech Synthesis
abstract
The latest Text-to-Speech (TTS) systems can produce speech with voice quality and naturalness comparable to human speech. Yet the demand for large amount of high-quality data from target speakers remains a significant challenge. Particularly for long-form expressive reading, target speaker's training speech that covers rich contextual information are needed. In this paper a novel design of context-aware speech pre-trained model is developed for expressive TTS based on contrastive learning. The model can be trained with abundant speech data without explicitly labelled speaker identities. It captures the intricate relationship between the speech expression of a spoken sentence and the contextual text information. By incorporating cross-modal text and speech features into the TTS model, it enables the generation of coherent and expressive speech, which is especially beneficial when there is a scarcity of target speaker data. The pre-trained model is evaluated first in the task of Context-Speech retrieval and then as the integral part of a zero-shot TTS system. Experimental results demonstrate that the pretraining framework effectively learns Context-Speech representations and significantly enhances the expressiveness of synthesized speech. Audio demos are available at: https://ccsp2024.github.io/demo/.
Yujia Xiao, Xi Wang 0016, Xu Tan 0003, Lei He 0005, Xinfa Zhu, Sheng Zhao 0002, Tan Lee
ACM Multimedia1
2024 UniStyle: Unified Style Modeling for Speaking Style Captioning and Stylistic Speech Synthesis
abstract
Understanding the speaking style, such as the emotion of the interlocutor's speech, and responding with speech in an appropriate style is a natural occurrence in human conversations. However, technically, existing research on speech synthesis and speaking style captioning typically proceeds independently. In this work, an innovative framework, referred to as UniStyle, is proposed to incorporate both the capabilities of speaking style captioning and style-controllable speech synthesizing. Specifically, UniStyle consists of a UniConnector and a style prompt-based speech generator. The role of the UniConnector is to bridge the gap between different modalities, namely speech audio and text descriptions. It enables the generation of text descriptions with speech as input and the creation of style representations from text descriptions for speech synthesis with the speech generator. Besides, to overcome the issue of data scarcity, we propose a two-stage and semi-supervised training strategy, which reduces data requirements while boosting performance. Extensive experiments conducted on open-source corpora demonstrate that UniStyle achieves state-of-the-art performance in speaking style captioning and synthesizes expressive speech with various speaker timbres and speaking styles in a zero-shot manner.
Xinfa Zhu, Lei He 0005, Yujia Xiao, Xi Wang 0016, Xu Tan 0003, Sheng Zhao 0002, Lei Xie 0001
ACM Multimedia5
2023 ContextSpeech: Expressive and Efficient Text-to-Speech for Paragraph Reading
abstract
While state-of-the-art Text-to-Speech systems can generate natural speech of very high quality at sentence level, they still meet great challenges in speech generation for paragraph / long-form reading.Such deficiencies are due to i) ignorance of cross-sentence contextual information, and ii) high computation and memory cost for long-form synthesis.To address these issues, this work develops a lightweight yet effective TTS system, ContextSpeech.Specifically, we first design a memory-cached recurrence mechanism to incorporate global text and speech context into sentence encoding.Then we construct hierarchically-structured textual semantics to broaden the scope for global context enhancement.Additionally, we integrate linearized self-attention to improve model efficiency.Experiments show that ContextSpeech significantly improves the voice quality and prosody expressiveness in paragraph reading with competitive model efficiency.
Yujia Xiao, Shaofei Zhang, Xi Wang 0016, Xu Tan 0003, Lei He 0005, Sheng Zhao 0002, Frank K. Soong, Tan Lee
INTERSPEECH1
2022 Improving Fastspeech TTS with Efficient Self-Attention and Compact Feed-Forward Network
abstract
FastSpeech, as a feed-forward transformer based TTS, can avoid the slow serial, autoregressive inference to generate the target mel-spectrogram in a parallel way. As a non-autoregressive TTS, the latency and computation load in inference is shifted from vocoder to transformer where the efficiency is limited by the quadratic time and memory complexity in the self-attention mechanism, particularly for a long text sequence. To tackle this challenges, We propose two models, ProbSparseFS and LinearizedFS, which have efficient self-attention arrangements to improve the inference speed and memory complexity. LinearizedFS has achieved 3.4x memory savings and 2.1x inference speedup, compared with the those of the baseline FastSpeech. A further optimized LinearizedFS with a lightweight FFN can accelerate the inference speed by 3.6x more. We do subjective voice quality evaluations in MOS and CMOS of News report and Audiobook applications, for multi-speaker and multi-style scenarios. Test results verified that the proposed models yield a TTS quality which is on-par with that of the baseline system but with much better memory efficiency and inference speed.
Yujia Xiao, Xi Wang 0016, Lei He 0005, Frank K. Soong
ICASSP1
2022 Prosodyspeech: Towards Advanced Prosody Model for Neural Text-to-Speech
abstract
This paper proposes ProsodySpeech, a novel prosody model to enhance encoder-decoder neural Text-To-Speech (TTS), to generate high expressive and personalized speech even with very limited training data. First, we use a Prosody Extractor built from a large speech corpus with various speakers to generate a set of prosody exemplars from multiple reference speeches, in which Mutual Information based Style content separation (MIST) is adopted to alleviate "content leakage" problem. Second, we use a Prosody Distributor to make a soft selection of appropriate prosody exemplars in phone-level with the help of an attention mechanism. The resulting prosody feature is then aggregated into the output of text encoder, together with additional phone-level pitch feature to enrich the prosody. We apply this method into two tasks: highly expressive multi style/emotion TTS and few-shot personalized TTS. The experiments show the proposed model outperforms baseline FastSpeech 2 + GST with significant improvements in terms of similarity and style expression.
Yuanhao Yi, Lei He 0005, Shifeng Pan, Xi Wang 0016, Yujia Xiao
ICASSP5
2020 Improving Prosody with Linguistic and Bert Derived Features in Multi-Speaker Based Mandarin Chinese Neural TTS
abstract
Recent advances of neural TTS have made "human parity" synthesized speech possible when a large amount of studio-quality training data from a voice talent is available. However, with only limited, casual recordings from an ordinary speaker, human-like TTS is still a big challenge, in addition to other artifacts like incomplete sentences, repetition of words, etc. Chinese, a language, of which the text is different from that of other roman-letter based languages like English, has no blank space between adjacent words, hence word segmentation errors can cause serious semantic confusions and unnatural prosody. In this study, with a multi-speaker TTS to accommodate the insufficient training data of a target speaker, we investigate linguistic features and Bert-derived information to improve the prosody of our Mandarin Chinese TTS. Three factors are studied: phone-related and prosody-related linguistic features; better predicted breaks with a refined Bert-CRF model; augmented phoneme sequence with character embedding derived from a Bert model. Subjective tests on in- and out-domain tasks of News, Chat and Audiobook, have shown that all factors are effective for improving prosody of our Mandarin TTS. The model with additional character embeddings from Bert is the best one, which outperforms the baseline by 0.17 MOS gain.
Yujia Xiao, Lei He 0005, Huaiping Ming, Frank K. Soong
ICASSP1
2018 Paired Phone-Posteriors Approach to ESL Pronunciation Quality Assessment
Yujia Xiao, Frank K. Soong, Wenping Hu
INTERSPEECH1
2017 Proficiency Assessment of ESL Learner's Sentence Prosody with TTS Synthesized Voice as Reference
Yujia Xiao, Frank K. Soong
INTERSPEECH1