VLDB 2026 Research / reviewers in the wild / expert
Rui Yang 0006
dblp:92/1942-6
· DBLP profile ↗
22ranked-venue papers
2as first author
9since 2021 · last 2026
0009-0008-7446-7216ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 14 · 1 first-author · 6 since 2021Security and privacy · 8 · 1 first-author · 3 since 2021Computer networks · 3 · 1 first-authorArtificial intelligence and machine learning · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Context-Aware TFL: A Universal Context-Aware Contrastive Learning Framework for Temporal Forgery LocalizationabstractMost research efforts in the multimedia forensics domain have focused on detecting forgery audio-visual content and reached sound achievements. However, these works only consider deepfake detection as a classification task and ignore the case where partial segments of the video are tampered with. Temporal forgery localization (TFL) of small fake audio-visual clips embedded in real videos is still challenging and more in line with realistic application scenarios. To resolve this issue, we propose a universal context-aware contrastive learning framework (UniCa-CLF) for TFL. Our approach leverages supervised contrastive learning to discover and identify forged instants by means of anomaly detection, allowing for the precise localization of temporal forged segments. To this end, we propose a specialized context-aware perception layer that utilizes a heterogeneous activation operation and an adaptive context updater to construct a context-aware contrastive objective, which enhances the discriminability of forged instant features by contrasting them with genuine instant features in terms of their distances to the global context. An efficient context-aware contrastive coding is introduced to further push the limit of instant feature distinguishability between genuine and forged instants in a supervised sample-by-sample manner, suppressing the cross-sample influence to improve temporal forgery localization performance. Extensive experimental results over five public datasets demonstrate that our proposed UniCaCLF significantly outperforms the state-of-the-art competing algorithms. The source code and pre-trained models of our proposed UniCaCLF are made publicly available at GitHub repository.TimeTimeTimeTime Qilin Yin, Wei Lu 0001, Xiangyang Luo 0001, Rui Yang 0006, Xiaochun Cao |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2025 | Weakly-supervised Audio Temporal Forgery Localization via Progressive Audio-language Co-learning NetworkabstractAudio temporal forgery localization (ATFL) aims to find the precise forgery regions of the partial spoof audio that is purposefully modified. Existing ATFL methods rely on training efficient networks using fine-grained annotations, which are obtained costly and challenging in real-world scenarios. To meet this challenge, in this paper, we propose a progressive audio-language co-learning network (LOCO) that adopts co-learning and self-supervision manners to prompt localization performance under weak supervision scenarios. Specifically, an audio-language co-learning module is first designed to capture forgery consensus features by aligning semantics from temporal and global perspectives. In this module, forgery-aware prompts are constructed by using utterance-level annotations together with learnable prompts, which can incorporate semantic priors into temporal content features dynamically. In addition, a forgery localization module is applied to produce forgery proposals based on fused forgery-class activation sequences. Finally, a progressive refinement strategy is introduced to generate pseudo frame-level labels and leverage supervised semantic contrastive learning to amplify the semantic distinction between real and fake content, thereby continuously optimizing forgery-aware features. Extensive experiments show that the proposed LOCO achieves SOTA performance on three public benchmarks. Junyan Wu, Wei Lu 0001, Xiangyang Luo 0001, Rui Yang 0006, Shize Guo |
IJCAI | 5 |
| 2025 | Robust watermarking against arbitrary scaling and cropping attacks
Shaowu Wu, Wei Lu 0001, Xiaolin Yin, Rui Yang 0006 |
Signal Process. | 4 |
| 2025 | Robust Image Watermarking With Synchronization Using Template Enhanced-Extracted NetworkabstractAn efficient robust watermarking method should be resistant to various distortions, including distortions from image processing and geometric attacks. Geometric attacks are significant challenges for watermarking methods because they destroy the synchronization of the watermark between the embedding side and extracting side. It is a considerable challenge to accomplish watermark synchronization for watermarking methods. To address this challenge, a novel robust watermarking method with synchronization is proposed. At the embedding side, the watermark and the template are embedded to generate the watermarked image. If the watermarked image is attacked, the watermark and template are also distorted. At the extracting side, a template enhanced-extracted network is proposed to achieve watermark synchronization. The template enhanced-extracted network effectively extracts the distorted template from the distorted image. The template-enhanced subnet can indirectly enhance the strength of the distorted template in the distorted image and improve the accuracy of the template-extracted subnet. The visual quality of the watermarked image is guaranteed because there is no need to embed the template with high strength. Then, the attack factor is predicted based on the distorted template. By leveraging this prediction, correct watermark extraction with synchronization is achieved. The experimental results demonstrate that the proposed watermarking method with synchronization yields excellent robustness under image processing, geometric attacks and combined attacks. Shaowu Wu, Xiaolin Yin, Wei Lu 0001, Xiangyang Luo 0001, Rui Yang 0006 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | Accurate and Efficient Privacy-Preserving Feature Extraction on Encrypted ImagesabstractIn cloud computing, it is necessary to outsource image processing algorithms securely without exposing private image content. The scale-invariant feature transform (SIFT) is a famous local descriptor widely used in computer vision. There are already some privacy-preserving schemes for computing SIFT on encrypted images. However, the state-of-the-art works have to convert fixed-point numbers into their binary representations, which reduces efficiency and accuracy. In this paper, we propose a novel privacy-preserving SIFT scheme built from secure protocols designed explicitly for fixed-point numbers to solve this problem. Specifically, using RLWE-based homomorphic encryption, we propose word-wise protocols to perform secure division, square root operation, comparison, derivation, and matrix inversion in a single-instruction multiple-data manner. These protocols allow direct processing of fixed-point numbers without converting them to binary numbers, thus achieving high computational efficiency. We have also realized critical SIFT steps missing from previous works, including Euclidean gradient amplitude computation, histogram peak interpolation, and precise interval localization, leading to improved accuracy of SIFT features in the encrypted domain. We conduct security analysis and perform extensive experiments to evaluate the execution efficiency and accuracy. The experimental results show that the proposed scheme outperforms the state-of-the-art works in terms of computational efficiency and accuracy. Peijia Zheng, Xiongjie Fang, Rui Yang 0006, Wei Lu 0001, Xiaochun Cao, Jiwu Huang |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2024 | Coarse-to-Fine Proposal Refinement Framework for Audio Temporal Forgery Detection and LocalizationabstractRecently, a novel form of audio partial forgery has posed challenges to its forensics, requiring advanced countermeasures to detect subtle forgery manipulations within long-duration audio. However, existing countermeasures still serve a classification purpose and fail to perform meaningful analysis of the start and end timestamps of partial forgery segments. To address this challenge, we introduce a novel coarse-to-fine proposal refinement framework (CFPRF) that incorporates a frame-level detection network (FDN) and a proposal refinement network (PRN) for audio temporal forgery detection and localization. Specifically, the FDN aims to mine informative inconsistency cues between real and fake frames to obtain discriminative features that are beneficial for roughly indicating forgery regions. The PRN is responsible for predicting confidence scores and regression offsets to refine the coarse-grained proposals derived from the FDN. To learn robust discriminative features, we devise a difference-aware feature learning (DAFL) module guided by contrastive representation learning to enlarge the sensitive differences between different frames induced by minor manipulations. We further design a boundary-aware feature enhancement (BAFE) module to capture the contextual information of multiple transition boundaries and guide the interaction between boundary information and temporal features via a cross-attention mechanism. Extensive experiments show that our CFPRF achieves state-of-the-art performance on various datasets, including LAV-DF, ASVS2019PS, and HAD. Junyan Wu, Wei Lu 0001, Xiangyang Luo 0001, Rui Yang 0006, Qian Wang 0002, Xiaochun Cao |
ACM Multimedia | 4 |
| 2023 | On Physically Occluded Fake Identity Document DetectionabstractMany online applications require the users to upload their identity documents for authentication. The fake identity document is one of the main threats which compromises the security and reliability of such online applications. Existing techniques focus on the detection of digitally forged identity documents, which neglect the impact of physical forgeries. In this paper, we look into the problem of detecting physically occluded fake identity documents, which can be easily generated without any image processing knowledge. We observe that the physical occlusions inevitably produce occluded boundaries on the document. To take the advantage, we propose an Occluded Boundary Representation Learning (OBRL) module to progressively learn the occluded boundary features. These are then fed into an Occluded Boundary Message Passing (OBMP) module to effectively diffuse the physical occlusion traces to enhance the backbone features for robust detection. We newly construct a Physically Occluded Fake ID Card image dataset (POID) for evaluation. Various experiments are conducted on the POID, where our scheme is able to achieve 99.6% of accuracy in detecting physically occluded fake ID card images with a mAP of over 85% to localize the occlusion regions. Sheng Li 0006, Silu Cao, Rui Yang 0006, Jishen Zeng, Zhenxing Qian, Xinpeng Zhang 0001 |
ACM Multimedia | 4 |
| 2023 | ReLoc: A Restoration-Assisted Framework for Robust Image Tampering LocalizationabstractWith the spread of tampered images, locating the tampered regions in digital images has drawn increasing attention. The existing tampering localization methods, however, suffer from severe performance degradation when the images are subjected to some post-processing, as the tampering traces would be distorted by the post-processing operations. The poor robustness against post-processing has become a bottleneck for the practical applications of image tampering localization techniques. In order to address this issue, this paper proposes a novelrestoration-assisted framework for image tamperinglocalization (ReLoc). The ReLoc framework mainly consists of an image restoration module and a tampering localization module. The key idea of ReLoc is to use the restoration module to recover a high-quality counterpart from the distorted tampered image, such that the distorted tampering traces can be re-enhanced, facilitating the tampering localization module to identify the tampered regions. To achieve this, the restoration module is optimized not only with the conventional constraints on image visual quality, but also with a forensics-oriented objective function. Furthermore, the restoration module and the localization module are trained alternately, which can stabilize the training process and is beneficial for improving the performance. The robustness of ReLoc has been evaluated by using several common post-processing operations, including lossy compressions, online social network transmission, and image resizing. Extensive experimental results show that ReLoc can significantly improve the localization performance compared to using a restoration-free model. In addition, we have shown that the restoration module in a well-trained ReLoc model is transferable for different localization modules and across different datasets. Peiyu Zhuang, Haodong Li 0001, Rui Yang 0006, Jiwu Huang |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2022 | Decision-Based Attack to Speaker Recognition System via Local Low-Frequency PerturbationabstractDespite neural network-based speaker recognition systems (SRS) have enjoyed significant success, they are proved to be quite vulnerable to adversarial examples. In practice, the SRS model parameters are not always available. Attackers have to probe the model only via querying, and such decision-based attacking merely relies on the output label is quite challenging. This letter proposes a two-step query-efficient decision-based attack based on local low-frequency perturbation. Specifically, instead of imposing perturbation on the entire audio sample, a local attacking region is firstly sought, confining the perturbed distortion to a local region. Second, considering that the majority of energy concentrates on the low-frequency bands, the proposed method suggests performing perturbation generation in the low-frequency domain. Experimental results demonstrate that, compared with the recent methods, our method could implement target attacking to SRS with a higher attacking success rate, at the cost of much lower queries and adversarial perturbation. Jiacheng Deng 0001, Li Dong 0006, Rangding Wang, Rui Yang 0006, Diqun Yan |
IEEE Signal Process. Lett. | 4 |
| 2019 | Robust Copy-Move Detection of Speech Recording Using Similarities of Pitch and FormantabstractCopy-move forgery on very short speech segments, followed by post-processing operations to eliminate traces of the forgery, presents a great challenge to forensic detection. In this paper, we propose a robust method for detecting and locating a speech copy-move forgery. We found that pitch and formant can be used as the features representing a voiced speech segment, and these two features are very robust against commonly used post-processing operations. In the proposed algorithm, we first divide the speech recording into voiced speech segments and unvoiced speech segments. We then extract the pitch sequence and the first two formant sequences as the feature set of each voiced speech segment. Dynamic time warping is applied to compute the similarities of each feature set. By comparing the similarities with a threshold, we can detect and locate copy-move forgeries in speech recording. The extensive experiments show that the proposed method is very effective in detecting and locating copy-move forgeries, even on a forged speech segment as short as one voiced speech segment. The proposed method is also robust against several kinds of commonly used post-processing operations and background noise, which highlights the promising potential of the proposed method as a speech copy-move forgery localization tool in practical forensics applications. Qi Yan 0004, Rui Yang 0006, Jiwu Huang |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2019 | Detection of Speech Smoothing on Very Short ClipsabstractAudio editing software can easily be used to manipulate digital speech for forgery. Smoothing on the tampered boundary is usually performed to eliminate the obvious traces of forgery after tampering. This presents a considerable challenge for the forensic detection of tampered speech because the smoothing model is unknown and the smoothing operation often modifies only several tens of samples with the editing software. In this paper, we propose to apply six filtering models to approximate the smoothing in audio editing software for training the classifier. We analyze the impact of filtering operations on speech signals, especially on differential signals. On the basis of the local variance of the differential signal, we design a simple and yet efficient feature set. Theoretical analysis and extensive experiments show that the proposed features are very effective in detecting several common filtering operations on very short speech clips. The experimental results also show that the proposed method can detect unknown smoothing performed by commonly used audio editing software, such as Cooledit and Adobe Audition. This highlights the promising potential of the proposed method for use as a forgery localization tool of digital speech signals in practical forensic applications. The proposed method is capable of detecting smoothing on very short speech clips containing only several tens of samples and practical forgery used in audio editing software. Qi Yan 0004, Rui Yang 0006, Jiwu Huang |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2018 | Improved Audio Steganalytic Feature and Its Applications in Audio ForensicsabstractDigital multimedia steganalysis has attracted wide attention over the past decade. Currently, there are many algorithms for detecting image steganography. However, little research has been devoted to audio steganalysis. Since the statistical properties of image and audio files are quite different, features that are effective in image steganalysis may not be effective for audio. In this article, we design an improved audio steganalytic feature set derived from both the time and Mel-frequency domains for detecting some typical steganography in the time domain, including LSB matching, Hide4PGP, and Steghide. The experiment results, evaluated on different audio sources, including various music and speech clips of different complexity, have shown that the proposed features significantly outperform the existing ones. Moreover, we use the proposed features to detect and further identify some typical audio operations that would probably be used in audio tampering. The extensive experiment results have shown that the proposed features also outperform the related forensic methods, especially when the length of the audio clip is small, such as audio clips with 800 samples. This is very important in real forensic situations. Weiqi Luo 0001, Haodong Li 0001, Qi Yan 0004, Rui Yang 0006, Jiwu Huang |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2017 | Detection of Double Compressed AMR Audio Using Stacked AutoencoderabstractThe adaptive multi-rate (AMR) audio codec adopted by many portable recording devices is widely used in speech compression. The use of AMR speech recordings as evidence in court is growing. Nowadays, it is easy to tamper with digital speech recordings, which makes audio forensics increasingly important. The detection of double compressed audio is one of the key issues in audio forensics. In this paper, we propose a framework for detecting double compressed AMR audio based on the stacked autoencoder (SAE) network and the universal background model-Gaussian mixture model (UBM-GMM). Instead of hand-crafted features, we used the SAE to learn the optimal features automatically from the audio waveforms. Audio frames are used as network input and the last hidden layer's output constitutes the features of a single frame. For an audio clip with many frames, the features of all the frames are aggregated and classified by UBM-GMM. Experimental results show that our method is effective in distinguishing single/double compressed AMR audio and outperforms the existing methods by achieving a detection accuracy of 98% on the TIMIT database. Exhaustive experiments demonstrate the effectiveness and robustness of the proposed method. Rui Yang 0006, Bin Li 0011, Jiwu Huang |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2015 | Copy-move detection of audio recording with pitch similarityabstractThe widespread availability of audio editing software has made it very easy to create forgeries without perceptual trace. Copy-move is one of popular audio forgeries. It is very important to identify audio recording with duplicated segments. However, copy-move detection in digital audio with sample by sample comparison is invalid due to post-processing after forgeries. In this paper we present a method based on pitch similarity to detect copy-move forgeries. We use a robust pitch tracking method to extract the pitch of every syllable and calculate the similarities of these pitch sequences. Then we can use the similarities to detect copy-move forgeries of digital audio recording. Experimental result shows that our method is feasible and efficient. Qi Yan 0004, Rui Yang 0006, Jiwu Huang |
ICASSP | 2 |
| 2014 | Detecting double compressed AMR audio using deep learningabstractThe Adaptive Multi-Rate (AMR) audio codec is a widely used audio data compression scheme optimized for speech and adopted by many devices. With the audio editing software, it is easy to perform tampering on digital speech recording, which makes the audio forensics become an important and urgent issue. Usually, the tampered AMR audio is double compressed AMR audio. In this paper, we proposed a method to detect the double compressed AMR audio. Such technique may be served as a tool for authenticating the originality of audio recordings and detecting the forgery positions. Our proposed method is based on deep learning algorithm and a majority voting strategy is designed for decision. The experimental results show that our method is effective to detect the double compressed AMR audio. Besides, the potential application of this technique is also discussed. Rui Yang 0006, Jiwu Huang |
ICASSP | 2 |
| 2014 | Anti-forensics of JPEG Detectors via Adaptive Quantization Table ReplacementabstractDue to the popularity of JPEG compression standard, JPEG images have been widely used in various applications. Nowadays, detection of JPEG forgeries becomes an important issue in digital image forensics, and lots of related works have been reported. However, most existing works mainly rely on a pre-trained classifier according to the quantization table shown in the file header of the suspicious JPEG image, and they assume that such a table is authentic. This assumption leaves a potential flaw for those wise forgers to confuse or even invalidate the current JPEG forensic detectors. Based on our analysis and experiments, we found that the generalization ability of most current JPEG forensic detectors is not very good. If the quantization table changes, their performances would decrease significantly. Based on this observation, we propose a universal anti-forensic scheme via replacing the quantization table adaptively. The extensive experimental results evaluated on 10,000 natural images have shown the effectiveness of the proposed scheme for confusing four typical JPEG forensic works. Haodong Li 0001, Weiqi Luo 0001, Rui Yang 0006, Jiwu Huang |
ICPR | 4 |
| 2014 | Identifying Compression History of Wave Audio and Its ApplicationsabstractAudio signal is sometimes stored and/or processed in WAV (waveform) format without any knowledge of its previous compression operations. To perform some subsequent processing, such as digital audio forensics, audio enhancement and blind audio quality assessment, it is necessary to identify its compression history. In this article, we will investigate how to identify a decompressed wave audio that went through one of three popular compression schemes, including MP3, WMA (windows media audio) and AAC (advanced audio coding). By analyzing the corresponding frequency coefficients, including modified discrete cosine transform (MDCT) and Mel-frequency cepstral coefficients (MFCCs), of those original audio clips and their decompressed versions with different compression schemes and bit rates, we propose several statistics to identify the compression scheme as well as the corresponding bit rate previously used for a given WAV signal. The experimental results evaluated on 8,800 audio clips with various contents have shown the effectiveness of the proposed method. In addition, some potential applications of the proposed method are discussed. Weiqi Luo 0001, Rui Yang 0006, Jiwu Huang |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2012 | Compression history identification for digital audio signalabstractCompression history identification plays a very important role in digital multimedia forensics. However, most existing literatures mainly focus on digital image forensics, and just a few works consider digital audio. In this paper, we investigate two popular compression schemes in digital audio, that is, MP3 and WMA, and try to reveal the compression history for a questionable audio signal in the original uncompressed WAV format via analyzing some statistical characteristics of the modified discrete cosine transform coefficients of the audio. The extensive experimental results have shown that the proposed method can effectively identify whether the given audio has been previously compressed with MP3 and/or WMA, and can further estimate the hidden compression rates, even the compression rate is as high as 128 K bps (bits per second). Weiqi Luo 0001, Rui Yang 0006, Jiwu Huang |
ICASSP | 3 |
| 2012 | Exposing MP3 audio forgeries using frame offsetsabstractAudio recordings should be authenticated before they are used as evidence. Although audio watermarking and signature are widely applied for authentication, these two techniques require accessing the original audio before it is published. Passive authentication is necessary for digital audio, especially for the most popular audio format: MP3. In this article, we propose a passive approach to detect forgeries of MP3 audio. During the process of MP3 encoding the audio samples are divided into frames, and thus each frame has its own frame offset after encoding. Forgeries lead to the breaking of framing grids. So the frame offset is a good indication for locating forgeries, and it can be retrieved by the identification of the quantization characteristic. In this way, the doctored positions can be automatically located. Experimental results demonstrate that the proposed approach is effective in detecting some common forgeries, such as deletion, insertion, substitution, and splicing. Even when the bit rate is as low as 32 kbps, the detection rate is above 99%. Rui Yang 0006, Zhenhua Qu, Jiwu Huang |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2011 | Geometric Invariant Audio Watermarking Based on an LCM FeatureabstractThe development of a geometric invariant audio watermarking scheme without degrading acoustical quality is challenging work. This paper proposes a multi-bit spread-spectrum audio watermarking scheme based on a geometric invariant log coordinate mapping (LCM) feature. The LCM feature is very robust to audio geometric distortions. The watermark is embedded in the LCM feature, but it is actually embedded in the Fourier coefficients which are mapped to the feature via LCM, so the embedding is actually performed in the DFT domain without interpolation, thus eliminating completely the severe distortion resulted from the non-uniform interpolation mapping. The watermarked audio achieves high auditory quality in both objective and subjective quality assessments. A mixed correlation between the LCM feature and a key-generated PN tracking sequence is proposed to align the log-coordinate mapping, thus synchronizing the watermark efficiently with only one FFT and one IFFT. Both the theoretical analysis and experimental results show that the proposed audio watermarking scheme is not only resilient against common signal processing operations, including low-pass filtering, MP3 recompression, echo addition, volume change, normalization, test functions in the Stirmark benchmark, and DA/AD conversion, but also has conquered the challenging audio geometric distortion and achieves the best robustness against simultaneous geometric distortions, such as pitch invariant time-scale modification (TSM) by ±20%, tempo invariant pitch shifting by 20%, resample TSM with scaling factors between 75% and 140%, and random cropping by 95%. This is mainly contributed by the proposed geometric invariant LCM feature. To our best knowledge, audio watermarking based on LCM has not been reported before. Xiangui Kang, Rui Yang 0006, Jiwu Huang |
IEEE Trans. Multim. | 2 |
| 2008 | Robust Audio Watermarking Based on Log-Polar Frequency Index
Rui Yang 0006, Xiangui Kang, Jiwu Huang |
IWDW | 1 |
| 2006 | Robust Audio Watermarking Based on Low-Order Zernike Moments
Shijun Xiang, Jiwu Huang, Rui Yang 0006, Chuntao Wang, Hongmei Liu 0001 |
IWDW | 3 |