Yingqiang Qiu

dblp:173/7130 · DBLP profile ↗
← Back
10ranked-venue papers
6as first author
7since 2021 · last 2026
0000-0001-5352-1318ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 6 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Leveraging Gaussian Attention for Adapter-Based Fine-Tuning of WavLM in Audio Deepfake Detection
abstract
The advancement of deepfake technology has enabled deepfake audio to achieve unprecedented realism, posing significant challenges to automatic speaker verification (ASV) systems and threatening information security in critical sectors such as telecommunications and banking. As a countermeasure, Audio Deepfake Detection (ADD) has received growing attention in recent years. Existing ADD techniques derived from deep‐learning framework often encounter the problem of insufficient generalizability on unseen deepfake methods and insufficient robustness with changes to the codec and noise condition. To address these challenges, we propose an efficient method that adapts self‐supervised learning (SSL) model WavLM for deepfake detection. Instead of costly full‐model retraining, we introduce a lightweight “adapter” module, which acts as a small, trainable component to efficiently teach WavLM this new task. This adapter incorporates a Gaussian attention mechanism, which guides the model to consistently focus on subtle vocal artifacts and ignore irrelevant noise or compression distortions, thereby ensuring robust performance across different real‐world conditions. Extensive evaluations show that our framework not only achieves competitive performance on the ASVspoof 2019 LA and DF in‐domain benchmarks but, more importantly, excels in generalization. Our method achieves EERs of 6.11% on the challenging ‘in‐the‐wild’ dataset and 4.90% on the WaveFake dataset, outperforming existing methods against previously unseen deepfake attacks. This superior generalization is further explained through SHAP analysis, which reveals that our method enables the model to focus on discriminative information within voiced segments, effectively ignoring fragile ‘shortcut’ cues from silent portions that often hinder generalization in other systems.
Xiaodan Lin, Yingqiang Qiu, Gewei Tan
Int. J. Intell. Syst.3
2025 Separable and high-capacity reversible data hiding for encrypted 3D mesh models based on dual multi-MSB predictions
Jiacheng Ge, Yingqiang Qiu, Kaimeng Chen, Xiaodan Lin, Yufeng Dai
Signal Process.2
2025 High-Capacity Reversible Data Hiding in Encrypted Images With Edge-Directed Prediction and Adaptive Entropy Coding
abstract
As cloud services rapidly evolve and the demand for privacy protection grows, reversible data hiding in encrypted images (RDHEI) has gained significant attention. To enhance the data embedding capacity of RDHEI, this paper proposes an adaptive prediction-error entropy encoding framework that dynamically allocates edge-directed prediction (EDP) errors to either separate or grouped entropy coding modes, thereby optimizing net payload capacity. The image owner first predicts the pixel values of the cover image using EDP algorithm and calculates the prediction errors. After computing the prediction errors, the optimal thresholds are adaptively determined by minimizing the expected codeword length through entropy coding theory. Using these optimized thresholds, the prediction errors are classified into separate or grouped encoding categories, and losslessly compressed via arithmetic encoding. Through the processes of image encryption and self-embedding, an encrypted image with embedding room is generated and subsequently uploaded to the cloud server. The data hider can easily locate the data embedding room in the encrypted domain of the image and embed encrypted additional data to obtain the marked encrypted image. The authorized recipient can correctly extract the embedded data, restore the original plaintext image without any loss, or do both. The experimental results demonstrate the effectiveness of the proposed approach, surpassing many state-of-the-art RDHEI techniques.
Yingqiang Qiu, Kaimeng Chen, Xiaodan Lin, Guogang Li, Huanqiang Zeng
IEEE Signal Process. Lett.1
2024 DS-TDNN: Dual-Stream Time-Delay Neural Network With Global-Aware Filter for Speaker Verification
abstract
Conventional time-delay neural networks (TDNNs) struggle to handle long-range context, their ability to represent speaker information is therefore limited for long utterances. Existing solutions either depend on increasing model complexity or try to strike a balance between local features and global context to address this issue. To effectively leverage the long-term dependencies of audio signals and constrain model complexity, we introduce a novel module called Global-aware Filter layer (GF layer) in this work, which employs a set of learnable transform-domain filters between a 1D discrete Fourier transform and its inverse transform to capture global context. Additionally, we develop a dynamic filtering strategy and a sparse regularization method to enhance the performance of the GF layer and prevent overfitting. Based on the GF layer, we present a dual-stream TDNN architecture called DS-TDNN for automatic speaker verification (ASV), which utilizes two unique branches to extract both local and global features in parallel and employs an efficient strategy to fuse different-scale information. Experiments on the Voxceleb and SITW databases demonstrate that the DS-TDNN achieves a relative improvement of 10% together with a relative decline of 20% in computational cost over the ECAPA-TDNN in the speaker verification task. This improvement becomes more evident as the utterance's duration grows. Furthermore, the DS-TDNN also beats popular deep residual models and attention-based systems on utterances of arbitrary length.
Yangfu Li, Jiapan Gan, Xiaodan Lin, Yingqiang Qiu, Hongjian Zhan, Hui Tian 0002
IEEE ACM Trans. Audio Speech Lang. Process.4
2022 A Cony-Attention Network for Detecting the Presence of ENF Signal in Short-Duration Audio
abstract
Detecting the presence of the electric network frequency (ENF) signal in audio recordings is a prerequisite of applying the ENF criterion that plays an essential role in numerous forensic applications. However, existing detection methods are powerless to handle short-duration audio recordings that have attracted considerable attention due to the popularity of voice messaging apps. This paper proposes a novel deep learning-based approach for ENF detection in short audio recordings, reducing the minimum operating range of audio duration to 1/10 of the state-of-the-art methods. Meanwhile, a convolutional attention network termed Conv-AttNet is proposed to improve the detection performance of convolutional neural networks (CNN) through the attention mechanism. Experiments on both synthetic and real-world audio recordings reveal that Conv-AttNet is able to detect the ENF signal buried in only 2 seconds of audio recordings, surpassing both matched filtering and typical CNN like ResNet50. In addition, the detection accuracy can be further increased by utilizing audio recordings of longer duration.
Yangfu Li, Xiaodan Lin, Yingqiang Qiu, Huanqiang Zeng
MMSP3
2022 High-Capacity Framework for Reversible Data Hiding in Encrypted Image Using Pixel Prediction and Entropy Encoding
abstract
While the existing reserving room before encryption (RRBE) based reversible data hiding in encrypted image (RDHEI) schemes can achieve decent embedding capacity, the capacity of the existing vacating room by encryption (VRBE) based schemes is relatively low. To address this issue, this paper proposes a generalized framework for high-capacity RDHEI for both the RRBE and VRBE cases. First, an efficient embedding room generation algorithm (ERGA) is designed to produce large embedding room using pixel prediction and entropy encoding. Then, we propose two RDHEI schemes, one for RRBE, another for VRBE. In the RRBE scenario, the image owner generates the embedding room with ERGA and encrypts the preprocessed image using stream cipher with two encryption keys. Then, the data hider locates the embedding room and embeds the additional encrypted data. In the VRBE scenario, the cover image is encrypted by an improved block modulation and permutation encryption algorithm, where the spatial redundancy in the plain-text image is greatly preserved. Then, the data hider applies ERGA on the encrypted image to generate the embedding room and conducts data embedding. For both schemes, receivers with different authentication keys can conduct either error-free data extraction or error-free image recovery. The experimental results show that the two proposed schemes outperform many state-of-the-art RDHEI schemes. Besides, they can ensure high security level, where the original image can be hardly discovered from the encrypted version before or after data hiding by unauthorized users.
Yingqiang Qiu, Qichao Ying, Yuyan Yang, Huanqiang Zeng, Sheng Li 0006, Zhenxing Qian
IEEE Trans. Circuits Syst. Video Technol.1
2021 Optimized Lossless Data Hiding in JPEG Bitstream and Relay Transfer-Based Extension
abstract
This paper proposes a new framework of lossless data hiding (LDH) in JPEG images. The proposed framework contains two algorithms, i.e., the optimized basic LDH and the relay transfer based extension. In the basic algorithm, we aim to preserve the filesize after data embedding. The data hiding process is optimized by variable-length-code (VLC) mapping, combination and permutation. In the extended algorithm, we focus on embedding more bits into the bitstream with a condition that the filesize increment is allowed. To decrease the filesize increment, we propose a relay transfer based algorithm to preprocess the JPEG bitstream. Subsequently, we embed data into the processed bitstream using the basic LDH. Both algorithms provide better performances than previous arts. After lossless data hiding, the marked JPEG bitstream is compliant to common JPEG decoders. Since all operations are implemented on VLCs and the Huffman codes, no distortion is generated on the image pixels. Experimental results demonstrate that the proposed approach outperforms previous methods.
Yingqiang Qiu, Zhenxing Qian, Han He, Hui Tian 0002, Xinpeng Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2020 Reversible data hiding in encrypted images using adaptive reversible integer transformation
Yingqiang Qiu, Zhenxing Qian, Huanqiang Zeng, Xiaodan Lin, Xinpeng Zhang 0001
Signal Process.1
2018 Lossless data hiding in JPEG bitstream using alternative embedding
Yingqiang Qiu, Han He, Zhenxing Qian, Sheng Li 0006, Xinpeng Zhang 0001
J. Vis. Commun. Image Represent.1
2016 Adaptive Reversible Data Hiding by Extending the Generalized Integer Transformation
abstract
This letter proposes a novel reversible data hiding (RDH) method with an adaptive embedding capability by extending the generalized integer transformation (GIT). We modify the GIT algorithm to a further generalized form, and accordingly we propose an adaptive embedding algorithm. When hiding data into the original image, parameters of the extended GIT can be identified according to the content of each block. Along with a multi-level location map, a better embedding efficiency of the net payload can be achieved. On the decoding side, the original image can be losslessly recovered after extracting the hidden message. Experimental results show that the proposed method outperforms state-of-the-art integer-transformation based RDH methods.
Yingqiang Qiu, Zhenxing Qian, Lun Yu
IEEE Signal Process. Lett.1