Xiaowei Yi

dblp:120/6771 · DBLP profile ↗
← Back
32ranked-venue papers
3as first author
17since 2021 · last 2025
0000-0002-8250-1698ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 15 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 2 first-author · 10 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2025 Towards High-Capacity Provably Secure Steganography via Cascade Sampling
Meiyang Lv, Haocheng Fu, Xiaowei Yi, Hongxian Huang, Yun Cao 0001
ICICS (3)3
2025 DisentTalk: Cross-lingual Talking Face Generation via Semantic Disentangled Diffusion Model
abstract
Recent advances in talking face generation have significantly improved facial animation synthesis. However, existing approaches face fundamental limitations: 3DMM-based methods maintain temporal consistency but lack fine-grained regional control, while Stable Diffusion-based methods enable spatial manipulation but suffer from temporal inconsistencies. The integration of these approaches is hindered by incompatible control mechanisms and semantic entanglement of facial representations. This paper presents DisentTalk, introducing a data-driven semantic disentanglement framework that decomposes 3DMM expression parameters into meaningful subspaces for fine-grained facial control. Building upon this disentangled representation, we develop a hierarchical latent diffusion architecture that operates in 3DMM parameter space, integrating region-aware attention mechanisms to ensure both spatial precision and temporal coherence. To address the scarcity of high-quality Chinese training data, we introduce CHDTF, a Chinese high-definition talking face dataset. Extensive experiments show superior performance over existing methods across multiple metrics, including lip synchronization, expression quality, and temporal consistency. Project Page: https://kangweiiliu.github.io/DisentTalk.
Kangwei Liu 0003, Junwu Liu, Yun Cao 0001, Jinlin Guo, Xiaowei Yi
ICME5
2025 Controllable Expressive 3D Facial Animation via Diffusion in a Unified Multimodal Space
abstract
Audio-driven emotional 3D facial animation encounters two significant challenges: (1) reliance on single-modal control signals (videos, text, or emotion labels) without leveraging their complementary strengths for comprehensive emotion manipulation, and (2) deterministic regression-based mapping that constrains the stochastic nature of emotional expressions and non-verbal behaviors, limiting the expressiveness of synthesized animations. To address these challenges, we present a diffusion-based framework for controllable expressive 3D facial animation. Our approach introduces two key innovations: (1) a FLAME-centered multimodal emotion binding strategy that aligns diverse modalities (text, audio, and emotion labels) through contrastive learning, enabling flexible emotion control from multiple signal sources, and (2) an attention-based latent diffusion model with content-aware attention and emotion-guided layers, which enriches motion diversity while maintaining temporal coherence and natural facial dynamics. Extensive experiments demonstrate that our method outperforms existing approaches across most metrics, achieving a 21.6% improvement in emotion similarity while preserving physiologically plausible facial dynamics. Project Page: https://kangweiiliu.github.io/Control_3D_Animation.
Kangwei Liu 0003, Junwu Liu, Xiaowei Yi, Jinlin Guo, Yun Cao 0001
ICME3
2025 Triple-Stage Robust Audio Steganography Framework with AAC Encoding for Lossy Social Media Channels
abstract
Robust audio steganography has significant application value for secure communication, especially with the rise of social media platforms.However, the complexity of audio encoding and the distortions introduced by lossy channels have hindered the research in this field.This paper systematically analyzes the origins of this challenge, and evaluates the limitations of previous methods.Building on this foundation, we propose a triple-stage Robust Audio Steganography Framework (RASF), specifically designed for AAC encoding process.RASF consists three essential stages: Psy-Window Control to synchronize psychoacoustic model parameters, Robust Embedding Domain Construction to establish a robust embedding domain using stable quantized coefficients, and Error Correction to ensure reliable data recovery.Experiments demonstrate that the proposed framework achieves high capacity and strong robustness against compression.Notably, tests conducted on social media platforms reveal a very low bit error rate, enabling zero-bit-error transmission when combined with error-correcting codes.RASF addresses critical gaps in robust audio steganography, offering a practical solution for covert communication over lossy social media channels.
Ziping Zhang, Jiamin Zeng, Xiaowei Yi, Yun Cao 0001
IH&MMSec4
2025 ExpFormer: Cross-lingual One-shot Talking Head Generation via Enhanced 3D Expression Modeling
abstract
One-shot audio-driven talking head generation suffers from significant limitations, primarily due to the complex relationship between speech and facial dynamics. Existing methods exhibit two major issues: (1) unrealistic expressions due to noisy 3D estimations in the lip region, and (2) inconsistent speaking styles, particularly evident in cross-lingual scenarios. To address these challenges, we propose ExpFormer, a novel 3DMM-based framework with two key innovations: (1) a lip movement enhancement strategy that reinforces lip region saliency during training, effectively mitigating 3D estimation errors in the lip region while preserving identity consistency, and (2) a transformer-based architecture with periodic positional encoding that captures both fine-grained lip synchronization and long-term speaking patterns, enabling natural facial animations across languages. To systematically evaluate cross-lingual generalization, we introduce a Mandarin Chinese dataset. Extensive experiments demonstrate that EXPFORMER significantly outperforms existing methods in visual quality, identity preservation, and lip synchronization across five different languages while achieving real-time performance.
Kangwei Liu 0003, Xiaowei Yi, Junwu Liu, Yun Cao 0001
IJCNN2
2025 DiffEmotionVC: A Dual-Granularity Disentangled Diffusion Framework for Any-to-Any Emotional Voice Conversion
Xiaosu Su, Xiaowei Yi, Yun Cao 0001
INTERSPEECH3
2024 RLVC: Robust and Lightweight Voice Conversion Using Cross-Adaptive Instance Normalization
abstract
Voice conversion refers to transforming the speaker of a voice into a target speaker while keeping the content unchanged. Current solutions either rely on feature representations from large-scale pre-trained models or require complex model designs and intensive training, lacking exploration of intrinsic speech features. This restricts the exploration of lightweight and robust methods. In this study, we remove pre-trained models and depart from complex mutual information minimization for feature decoupling. Instead, we revisit decoupling methods based on instance normalization. To address it, we introduce a novel feature coupling module named cross-adaptive instance normalization (CAIN), which extends the adaptive instance normalization (AdaIN). Beyond offering style injection capabilities, CAIN is designed to maintain content consistency by reconstructing frame-level statistics in mel-spectrograms. The results indicate that CAIN, serving as a lightweight plugin, significantly improves conventional instance normalization-driven approaches. Building upon this, we propose RLVC, which achieves robust performance with a mere 5.29M parameters.
Yewei Gu, Xianfeng Zhao, Xiaowei Yi
ICME3
2024 ProDub: Progressive Growing of Facial Dubbing Networks for Enhanced Lip Sync and Fidelity
abstract
Facial dubbing has attracted growing research interests due to its creative and practical applications. An ideal facial dubbing video should exhibit accurate lip-sync and high visual quality. However, prior methods fall short in fully exploring the relationship between two pairs of critical elements: distinguishing mouth shape and texture for accurate lip-sync performance; aligning the driving audio and high-frequency details for better visual quality. To address these challenges, we propose a progressive framework ProDub for this task. Specifically, we propose an audio-supervised contrastive approach to disentangle the mouth shape and texture, along with a novel lip-shape-aware loss as a constraint for producing accurate lip-sync. For high-quality visual output, a lip-aware temporal-enhanced network is designed to improve the lip details while ensuring temporal coherency based on a learned prior. Extensive experiments demonstrate that our ProDub improves lip-sync by 10.7% and visual quality by 28.5% compared to state-of-the-art methods .1
Kangwei Liu 0003, Xiaowei Yi, Xianfeng Zhao
ICME2
2023 Robust Feature Decoupling in Voice Conversion by Using Locality-Based Instance Normalization
Yewei Gu, Xianfeng Zhao, Xiaowei Yi
INTERSPEECH3
2023 A Compressed Synthetic Speech Detection Method with Compression Feature Embedding
Jinghong Zhang, Xiaowei Yi, Xianfeng Zhao
INTERSPEECH2
2023 Inversion Image Pairs for Anti-forensics in the Frequency Domain
Houchen Pu, Xiaowei Yi, Xianfeng Zhao
IWDW2
2022 Improving Robustness of Speech Anti-Spoofing System Using Resnext with Neighbor Filters
abstract
Since recent advances in speech synthesis techniques, it is important to develop robust speech anti-spoofing systems against all major spoofing attacks. In this paper, we propose a novel spoofing speech detection model by jointing ResNeXt with neighbor filters (NF-ResNeXt) to improve the robustness of speech anti-spoofing models. Inspired by higher-order cepstral coefficients are more difficult to be maintained during the speech synthesis procedure, we present a novel neighbor filter module for extracting the residual features to enhance the robustness of cepstral features. Then, we introduce a neural network architecture based on ResNeXt model for processing the residual features and calculating the scores of speech clips being spoofed. The NF-ResNeXt model is trained on the training set of ASVspoof 2019 logical access (LA) dataset and achieves an equal error rate (EER) of 5.13% on the evaluation dataset, which outperforms the existing state-of-the-art speech anti-spoofing models.
Xianfeng Zhao, Xiaowei Yi
ICME3
2022 Voice Conversion Using Learnable Similarity-Guided Masked Autoencoder
Yewei Gu, Xianfeng Zhao, Xiaowei Yi, Junchao Xiao
IWDW3
2022 A fast and secure MP3 steganographic scheme with multi-domain
Yunzhao Yang, Xiaowei Yi, Xianfeng Zhao, Jinghong Zhang
Signal Process.2
2022 Content-Aware Robust JPEG Steganography for Lossy Channels Using LPCNet
abstract
Most robust steganographic methods pursue insignificantly zero-bit error rate of hidden messages for ensuring the reliability of communication. Cover JPEG images are required to recompress many times for enhancing the robustness, that reduces the security. In this letter, we propose a content-aware steganographic scheme for hiding speech signals into JPEG images by utilizing the redundancy of speech signals. Firstly, we design a steganographic communication model that combines speech coding with embedding process by using embedded redundancy of speech messages. It is suitable for all lossy channels. Secondly, for improving the robustness of speech signals under a given embedding rate, we propose a content-aware protection method by exploiting different effects of speech coding parameters on speech quality after transcoding. Finally, an optimized linear prediction net (LPCNet) model is implemented to improve the end-to-end quality of speech signals. Compared with existing algorithms, experimental results show that the end-to-end quality of embedding speech is improved by 95% and the transmission efficiency is raised by 2.16 times. Meanwhile, our scheme can resist the steganalysis attack based on the JPEG recompression feature.
Xiaowei Yi, Xianfeng Zhao, Yunzhao Yang
IEEE Signal Process. Lett.2
2021 Fake Speech Detection Using Residual Network with Transformer Encoder
abstract
Fake speech detection aims to distinguish fake speech from natural speech. This paper presents an effective fake speech detection scheme based on residual network with transformer encoder (TE-ResNet) for improving the performance of fake speech detection. Firstly, considering inter-frame correlation of the speech signal, we utilize transformer encoder to extract contextual representations of the acoustic features. Then, a residual network is used to process deep features and calculate score that the speech is fake. Besides, to increase the quantity of training data, we apply five speech data augmentation techniques on the training dataset. Finally, we fuse the different fake speech detection models on score-level by logistic regression for compensating the shortcomings of each single model. The proposed scheme is evaluated on two public speech datasets. Our experiments demonstrate that the proposed TE-ResNet outperforms the existing state-of-the-art methods both on development and evaluation datasets. In addition, the proposed fused model achieves improved performance for detection of unseen fake speech technology, which can obtain equal error rates (EERs) of 3.99% and 5.89% on evaluation set of FoR-normal dataset and ASVspoof 2019 LA dataset respectively.
Xiaowei Yi, Xianfeng Zhao
IH&MMSec2
2021 FMFCC-A: A Challenging Mandarin Dataset for Synthetic Speech Detection
Yewei Gu, Xiaowei Yi, Xianfeng Zhao
IWDW3
2020 Deepfake Video Detection Using Audio-Visual Consistency
Yewei Gu, Xianfeng Zhao, Xiaowei Yi
IWDW4
2020 MP3 steganalysis based on joint point-wise and block-wise correlations
Yuntao Wang 0003, Xiaowei Yi, Xianfeng Zhao
Inf. Sci.2
2020 An AAC steganography scheme for adaptive embedding with distortion minimization model
Xiaowei Yi, Xianfeng Zhao
Multim. Tools Appl.2
2020 An Adaptive Double-Layered Embedding Scheme for MP3 Steganography
abstract
In this letter, we devise an adaptive double-layered embedding scheme that is suitable for MP3 steganography. According to the encoding characteristics of MP3, the linbits are used as the embedding domain. The proposed scheme is divided into two layers and the messages are embedded into each layer with binary STCs. The cost function used in the first layer is designed by employing the masking effect to achieve optimal imperceptibility. In order to reduce the modification of coefficients, the cost function for the second layer is revised according to the embedding results of the first layer. Experiments demonstrate that our scheme is able to achieve better acoustic concealment and higher embedding modifying efficiency indeed. In addition, our scheme can well resist the attack from the statistical handcrafted analysis methods.
Yunzhao Yang, Xianfeng Zhao, Xiaowei Yi
IEEE Signal Process. Lett.4
2019 RHFCN: : Fully CNN-based Steganalysis of MP3 with Rich High-pass Filtering
abstract
Recent studies have shown that convolutional neural networks (CNNs) can boost the performance of audio steganalysis. In this paper, we propose a well-designed fully CNN architecture for MP3 steganalysis based on rich high-pass filtering (HPF). On the one hand, multi-type HPFs are employed for "residual" extraction to enlarge the traces of the signal in view of the truth that signal introduced by secret messages can be seen as high-pass frequency noise. On the other hand, to utilize the spatial characteristics of feature maps better, fully connected (Fc) layers are replaced with convolutional layers. Moreover, this fully CNN architecture can be applied to the steganalysis of MP3 with size mismatch. The proposed network is evaluated on various MP3 steganographic algorithms, bitrates and relative payloads, and the experimental results demonstrate that our proposed network performs better than state-of-the-art methods.
Yuntao Wang 0003, Xiaowei Yi, Xianfeng Zhao, Ante Su
ICASSP2
2019 Recurrent Convolutional Neural Networks for AMR Steganalysis Based on Pulse Position
abstract
With the rapid development of stream multimedia, the adaptive multi-rate (AMR) audio steganography are emerging recently. However, the traditional steganalysis methods face great challenges in detecting short time speech at low embedding rates. To address this problem, we propose a steganalytic scheme by combining Recurrent Neural Network (RNN) and Convolutional Neural Network (CNN), SRCNet. AMR fixed codebook (FCB) steganography embed messages by modifying the pulse positions, which would destroy the FCB correlation. Firstly we analyzed the FCB correlations at different distances, and summarized these correlations into four categories. Furthermore, we utilizes RNN to extract higher level contextual representations of FCBs and CNN to fuse spatial-temporal features for the steganalysis. The proposed approach was evaluated on a public data-set. The experiment results validate that the proposed framework greatly outperforms the existing state-of-the-art methods. The correct detection rate of SRCNet has been improved above at least 10% when the sample is as short as 100ms at the 20% embedding rate. In particular, the network achieves the significant improvements for detecting the STCs based adaptive AMR steganography.
Xiaowei Yi, Xianfeng Zhao
IH&MMSec2
2019 Defining Joint Embedding Distortion for Adaptive MP3 Steganography
abstract
In this paper, a universal joint embedding distortion function (JED) is proposed to improve the undetectability and imperceptibility of MP3 steganography, which can be applied to Huffman codeword mapping (HCM) and sign bit flipping (SBF). Content-aware and statistical distortions are synthetically modeled to formulate the atom modification of the quantified modified discrete cosine transform (QMDCT) coefficients. On the one hand, to retain the hearing imperceptibility, the absolute threshold of hearing is employed to measure the auditory sensitivity of each QMDCT coefficient. On the other hand, considering most of the existing universal MP3 steganalysis features are designed based on correlations, the forward and backward transition probability are utilized to characterize the correlations between adjacent QMDCT coefficients. What's more, we present an implementation of JED in sign bits domain. Experimental results demonstrate that our method is able to achieve higher embedding capacity and better imperceptibility. The detection accuracy of the proposed scheme is about 75% with the bitrate of 320kbps and embedding rate of 11kbit/s, which is respectively decreased by 9.54% ~ 16.94% than existing MP3 steganographic methods.
Yunzhao Yang, Yuntao Wang 0003, Xiaowei Yi, Xianfeng Zhao
IH&MMSec3
2019 Improving Audio Steganalysis Using Deep Residual Networks
Xiaowei Yi, Xianfeng Zhao
IWDW2
2019 Light Multiscale Conventional Neural Network for MP3 Steganalysis
Jinghong Zhang, Xiaowei Yi, Xianfeng Zhao, Yun Cao 0001
IWDW2
2019 AHCM: Adaptive Huffman Code Mapping for Audio Steganography Based on Psychoacoustic Model
abstract
Most current audio steganographic methods are content non-adaptive which have poor security and low embedding capacity. This paper proposes a generalized adaptive Huffman code mapping (AHCM) framework for obtaining higher secure payload. To avoid the frame-offset effect of audio codec, we first establish a distortion-limited suppressible code space, which realizes data embedding by using equal-length entropy codes. Furthermore, a stego key is used to dynamically build Huffman code mapping of each frame for improving acoustic imperceptibility and statistical undetectability. We then consider integrating psychoacoustic model (PAM) of intra-frame with frame-level perceptual distortion of inter-frame to obtain minimized total distortion. Finally, we present an implementation of the proposed AHCM framework on MP3 audios. A distortion function based on the PAM and an optimal steganographic frame path are, respectively, devised for adaptively embedding via employing syndrome-trellis codes. Experimental results demonstrate that our approach is, indeed, able to achieve higher secure steganographic capacity and better acoustic concealment. The detection accuracy of 320-kbps-mp3 datasets is lower than 65% when the embedding payload reaches 11 kbps, which is decreased by 11.8%-13.4% than the state-of-the-art steganographic methods.
Xiaowei Yi, Xianfeng Zhao, Yuntao Wang 0003
IEEE Trans. Inf. Forensics Secur.1
2018 CNN-based Steganalysis of MP3 Steganography in the Entropy Code Domain
abstract
This paper presents an effective steganalytic scheme based on CNN for detecting MP3 steganography in the entropy code domain. These steganographic methods hide secret messages into the compressed audio stream through Huffman code substitution, which usually achieve high capacity, good security and low computational complexity. First, unlike most previous CNN based steganalytic methods, the quantified modified DCT (QMDCT) coefficients matrix is selected as the input data of the proposed network. Second, a high pass filter is used to extract the residual signal, and suppress the content itself, so that the network is more sensitive to the subtle alteration introduced by the data hiding methods. Third, the $ 1 \times 1 $ convolutional kernel and the batch normalization layer are applied to decrease the danger of overfitting and accelerate the convergence of the back-propagation. In addition, the performance of the network is optimized via fine-tuning the architecture. The experiments demonstrate that the proposed CNN performs far better than the traditional handcrafted features. In particular, the network has a good performance for the detection of an adaptive MP3 steganography algorithm, equal length entropy codes substitution (EECS) algorithm which is hard to detect through conventional handcrafted features. The network can be applied to various bitrates and relative payloads seamlessly. Last but not the least, a sliding window method is proposed to steganalyze audios of arbitrary size.
Yuntao Wang 0003, Xiaowei Yi, Xianfeng Zhao, Zhoujun Xu
IH&MMSec3
2018 Pitch Delay Based Adaptive Steganography for AMR Speech Stream
Xiaowei Yi, Xianfeng Zhao
IWDW2
2017 Adaptive MP3 Steganography Using Equal Length Entropy Codes Substitution
Xiaowei Yi, Xianfeng Zhao, Linna Zhou
IWDW2
2014 Efficient authentication of scalable media streams over wireless networks
Xiaowei Yi, Hengtai Ma, Changwen Zheng
Multim. Tools Appl.1
2013 Joint FEC codes and hash chains for optimizing authentication of JPEG2000 image streaming
abstract
This paper shows an optimizing stream-level authentication approach for secure and robust multimedia delivery over wireless networks. The proposed approach can achieve the optimum end-to-end authentic quality with a lower overhead. The received packet of codestreams is really effective to decrease the distortion, when it is both decodable and verifiable. Therefore, according to the coding and verification dependencies, an authentication optimization model (AOM) is designed to minimize the distortion and overhead. An implementation on the JPEG2000 codestream is realized via using the proposed AOM. By utilizing forward error correction (FEC) codes and hash chains, our approach can ensure that the verification dependencies are consistent with the coding dependencies. In other words, the proposed scheme does not cause quality degradations and it also reduce overhead redundancy. Experimental results demonstrate that our scheme achieves more optimizing end-to-end rate-distortion (R-D) performances at any packet-loss rate.
Xiaowei Yi, Hengtai Ma, Changwen Zheng
ICME1