Guang Hua 0001

dblp:120/6837 · DBLP profile ↗
← Back
48ranked-venue papers
18as first author
27since 2021 · last 2026
0000-0001-5184-4359ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 22 · 9 first-author · 10 since 2021Artificial intelligence and machine learning · 16 · 3 first-author · 13 since 2021Security and privacy · 10 · 6 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Removing Box-Free Watermarks for Image-to-Image Models via Query-Based Reverse Engineering
abstract
The intellectual property of deep generative networks (GNets) can be protected using a cascaded hiding network (HNet) which embeds watermarks (or marks) into GNet outputs, known as box-free watermarking. Although both GNet and HNet are encapsulated in a black box (called operation network, or ONet), with only the generated and marked outputs from HNet being released to end users and deemed secure, in this paper, we reveal an overlooked vulnerability in such systems. Specifically, we show that the hidden GNet outputs can still be reliably estimated via query-based reverse engineering, leaking the generated and unmarked images, despite the attacker's limited knowledge of the system. Our first attempt is to reverse-engineer an inverse model for HNet under the stringent black-box condition, for which we propose to exploit the query process with specially curated input images. While effective, this method yields unsatisfactory image quality. To improve this, we subsequently propose an alternative method leveraging the equivalent additive property of box-free model watermarking and reverse-engineering a forward surrogate model of HNet, with better image quality preservation. Extensive experimental results on image processing and image generation tasks demonstrate that both attacks achieve impressive watermark removal success rates (100%) while also maintaining excellent image quality (reaching the highest PSNR of 34.69 dB), substantially outperforming existing attacks, highlighting the urgent need for robust defensive strategies to mitigate the identified vulnerability in box-free model watermarking.
Haonan An 0001, Guang Hua 0001, Hangcheng Cao, Zhengru Fang, Guowen Xu, Susanto Rahardja, Yuguang Fang
AAAI2
2026 Fragment-energy audio watermarking resilient to de-synchronization attacks
Juan Zhao 0007, Tianrui Zong, Iynkaran Natgunanathan, Yong Xiang 0001, Guang Hua 0001, Longxiang Gao, Wanlei Zhou 0001
Expert Syst. Appl.6
2026 Box-Free Model Watermarks are Prone to Black-Box Removal Attacks
abstract
Box-free model watermarking is an emerging technique to safeguard the intellectual property of deep learning models, particularly those for low-level image processing tasks. Existing works have verified and improved its effectiveness in several aspects. However, in this paper, we systematically investigate the vulnerability and demonstrate that box-free model watermarking is prone to removal attacks, even under the real-world threat model such that the protected model and the watermark extractor are in black boxes. Under this setting, we carry out three studies. 1) We develop an extractor-gradient-guided (EGG) remover and show its effectiveness when the extractor uses ReLU activation only. 2) More generally, for an unknown extractor, we leverage adversarial attacks and design the EGG remover based on the estimated gradients. 3) Under the most stringent condition that the extractor is inaccessible, we design a transferable remover based on a set of private proxy models. In all cases, the proposed removers can successfully remove embedded watermarks while preserving the quality of the processed images, and we also demonstrate that the EGG remover can even replace the watermarks. Extensive experimental results verify the effectiveness and generalizability of the proposed attacks, revealing the vulnerabilities of the existing box-free methods and calling for further research.
Haonan An 0001, Guang Hua 0001, Zhiping Lin 0001, Yuguang Fang
IEEE Trans. Pattern Anal. Mach. Intell.2
2026 Nonintrusive Watermarking for CycleGAN
abstract
Generative adversarial networks (GANs) are a set of powerful generative models, among which CycleGAN, featuring the unique cycle-consistency loss, has gained special popularity. However, this unique structure and the cycle-consistency loss make watermarking CycleGAN particularly challenging, rendering existing deep neural network (DNN) watermarking methods, whether model-agnostic or GAN-specific, inapplicable. Meanwhile, existing DNN watermarking methods are intrusive in nature, requiring direct or indirect modification of model parameters for watermark embedding, which raises fidelity concerns. To solve the above problems, we propose the first nonintrusive and robust watermarking method for CycleGAN. We empirically show that without modifying the CycleGAN model, a user-defined watermark image can still be extracted from model outputs using a dedicated watermark decoder. Extensive experimental results verify that while achieving the so-called absolute fidelity, the proposed method is robust to various attacks, from image post-processing to model stealing.
Yebin Zheng, Haonan An 0001, Guang Hua 0001, Yongming Chen, Zhiping Lin 0001
IEEE Signal Process. Lett.3
2026 Decoder Gradient Shields: A Family of Provable and High-Fidelity Methods Against Gradient-Based Box-Free Watermark Removal
abstract
Box-free model watermarking has gained significant attention in deep neural network (DNN) intellectual property protection due to its model-agnostic nature and its ability to flexibly manage high-entropy image outputs from generative models. Typically operating in a black-box manner, it employs an encoder-decoder framework for watermark embedding and extraction. While existing research has focused primarily on the encoders for the robustness to resist various attacks, the decoders have been largely overlooked, leading to attacks against the watermark. In this paper, we identify one such attack against the decoder, where query responses are utilized to obtain backpropagated gradients to train a watermark remover. To address this issue, we propose Decoder Gradient Shields (DGSs), a family of defense mechanisms, including DGS at the output (DGS-O), at the input (DGS-I), and in the layers (DGS-L) of the decoder, with a closed-form solution for DGS-O and provable performance for all DGS. Leveraging the joint design of reorienting and rescaling of the gradients from watermark channel gradient leaking queries, the proposed DGSs effectively prevent the watermark remover from achieving training convergence to the desired low-loss value, while preserving image quality of the decoder output. We demonstrate the effectiveness of our proposed DGSs in diverse application scenarios. Our experimental results on deraining and image generation tasks with the state-of-the-art box-free watermarking show that our DGSs achieve a defense success rate of 100% under all settings.
Haonan An 0001, Guang Hua 0001, Hangcheng Cao, Yihang Tao, Guowen Xu, Susanto Rahardja, Yuguang Fang
IEEE Trans. Dependable Secur. Comput.2
2025 Decoder Gradient Shield: Provable and High-Fidelity Prevention of Gradient-Based Box-Free Watermark Removal
abstract
The intellectual property of deep image-to-image models can be protected by the so-called box-free watermarking. It uses an encoder and a decoder, respectively, to embed into and extract from the model’s output images invisible copyright marks. Prior works have improved watermark robustness, focusing on the design of better watermark encoders. In this paper, we reveal an overlooked vulnerability of the unprotected watermark decoder which is jointly trained with the encoder and can be exploited to train a watermark removal network. To defend against such an attack, we propose the decoder gradient shield (DGS) as a protection layer in the decoder API to prevent gradient-based watermark removal with a closed-form solution. The fundamental idea is inspired by the classical adversarial attack, but is utilized for the first time as a defensive mechanism in the box-free model watermarking. We then demonstrate that DGS can reorient and rescale the gradient directions of watermarked queries and stop the watermark remover’s training loss from converging to the level without DGS, while retaining decoder output image quality. Experimental results verify the effectiveness of the proposed method. Code of paper is available at https://github.com/haonanAN309/CVPR-2025-Official-Implementation-Decoder-Gradient-Shield.
Haonan An 0001, Guang Hua 0001, Zhengru Fang, Guowen Xu, Susanto Rahardja, Yuguang Fang
CVPR2
2025 A Key-Driven Framework for Identity-Preserving Face Anonymization
Guang Hua 0001, Sheng Li 0006, Guorui Feng
NDSS2
2025 Detecting Every Object From Events
abstract
Object detection is critical in autonomous driving, and it is more practical yet challenging to localize objects of unknown categories: an endeavour known as Class-Agnostic Object Detection (CAOD). Existing studies on CAOD predominantly rely on RGB cameras, but these frame-based sensors usually have high latency and limited dynamic range, leading to safety risks under extreme conditions like fast-moving objects, overexposure, and darkness. In this study, we turn to the event-based vision, featured by its sub-millisecond latency and high dynamic range, for robust CAOD. We propose Detecting Every Object in Events (DEOE), an approach aimed at achieving high-speed, class-agnostic object detection in event-based vision. Built upon the fast event-based backbone: recurrent vision transformer, we jointly consider the spatial and temporal consistencies to identify potential objects. The discovered potential objects are assimilated as soft positive samples to avoid being suppressed as backgrounds. Moreover, we introduce a disentangled objectness head to separate the foreground-background classification and novel object discovery tasks, enhancing the model's generalization in localizing novel objects while maintaining a strong ability to filter out the background. Extensive experiments confirm the superiority of our proposed DEOE in both open-set and closed-set settings, outperforming strong baseline methods.
Haitian Zhang, Chang Xu 0027, Xinya Wang, Bingde Liu, Guang Hua 0001, Lei Yu 0006, Wen Yang 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2025 Robust ENF Estimation in Contaminated Audio
abstract
Electric network frequency (ENF) is an important criterion in audio forensic analysis. However, environmental uncertainties often introduce various types of noises, diminishing the number of useful ENF samples in audio recordings. This issue is even more challenging in short-duration recordings. To address this issue, we propose an adaptive-window-based harmonic recombination (AWHR) method, which can accurately estimate ENF from noisy audio. Initially, we identify noisy samples and use a metric called noise ratio (NR) to determine the optimal harmonic. Adaptive windows are selectively applied to the noisy samples of the optimal harmonic to mitigate frequency spikes, prevent distortions, and preserve signal quality. This step also reduces computational complexity by minimizing the number of samples requiring enhancement. Finally, via a proposed harmonic recombination mechanism, we improve the number of useful ENF samples, which reduces the NR. Given the lack of ENF datasets designed to evaluate considerably contaminated audio, we have also built an ENF noisy audio harmonic (ENF-NAH) dataset. Experiments on public ENF-WHU and our ENF-NAH datasets show that the proposed AWHR method is effective in handling varying levels of contamination and is applicable to both long and short audio recordings.
Shiyu Zuo, Lexuan Xu, Sijin Wu, Guang Hua 0001
IEEE Trans. Inf. Forensics Secur.5
2024 Poisoning-Free Defense Against Black-Box Model Extraction
abstract
Recent research has shown that an adversary can use a surrogate model to steal the functionality of a target deep learning model even under the black-box condition and without data curation, while the existing defense mainly relies on API poisoning to disturb the surrogate training. Unfortunately, due to poisoning, the defense is achieved at the price of fidelity loss, sacrificing the interests of honest users. To solve this problem, we propose an Adversarial Fine-Tuning (AdvFT) framework, incorporating the generative adversarial network (GAN) structure that disturbs the feature representations of out-of-distribution (OOD) queries while preserving those of in-distribution (ID) ones, circumventing the need for OOD sample collection and API poisoning. Extensive experiments verify the effectiveness of the proposed framework. Code is available at github.com/Hatins/AdvFT.
Haitian Zhang, Guang Hua 0001, Wen Yang 0001
ICASSP2
2024 "Seeing" ENF From Neuromorphic Events: Modeling and Robust Estimation
abstract
Most artificial lights exhibit subtle fluctuations in intensity and frequency in response to the influence of the grid's alternating current, providing the potential to estimate the Electric Network Frequency (ENF) from conventional frame-based videos. Nevertheless, the performance of Video-based ENF (V-ENF) estimation largely relies on the imaging quality and thus may suffer from significant interference caused by non-ideal sampling, scene diversity, motion interference, and extreme lighting conditions. In this paper, we show that the ENF can be extracted without the above limitations from a new modality provided by the so-called event camera, a neuromorphic sensor that encodes the light intensity variations and asynchronously emits events with extremely high temporal resolution and high dynamic range. Specifically, we formulate and validate the physical mechanism for the ENF captured in events and then propose a simple yet robust Event-based ENF (E-ENF) estimation method through mode filtering and harmonic enhancement. To validate the effectiveness, we build the first Event-Video ENF Dataset (EV-ENFD) and its extension EV-ENFD+ with diverse scenarios, including static, dynamic, and extreme lighting scenes. Comprehensive experiments have been conducted on our proposed datasets, showcasing that our proposed E-ENF significantly outperforms the V-ENF in extracting accurate ENF traces, especially in challenging environments.
Lexuan Xu, Guang Hua 0001, Lei Yu 0006
IEEE Trans. Pattern Anal. Mach. Intell.2
2024 Multimodal Proxy-Free Face Anti-Spoofing Exploiting Local Patch Features
abstract
Face anti-spoofing (FAS) is vital to ensure the security of the face recognition systems, for which the essential task is to capture the unique spoof face features. Most of the existing methods extract spoof features from the whole faces, overlooking clues in local face patches. Meanwhile, researchers usually use intermediate parameters as a proxy in face classification, but this requires the design of additional loss functions. To solve these problems, we propose a multimodal proxy-free FAS model which uses contrastive language image pre-training (CLIP) as the backbone. Specifically, we use patches cropped from the original face to augment the data, forcing the network to learn local spoof features, such as the edges of printing attacks. At the same time, we introduce dynamic central difference convolutional (DCDC) adapter to extract fine-grained features in patches. Furthermore, we propose to adopt a proxy-free pairwise similarity learning (PSL) loss to achieve the goal that the maximum intra-class distance is smaller than the minimum inter-class distance. Experiments on several benchmark datasets show that the proposed method achieves state-of-the-art performance.
Xiangyu Yu, Xinghua Huang, Xiaohui Ye, Guang Hua 0001
IEEE Signal Process. Lett.5
2024 FAWA: Fast Adversarial Watermark Attack
abstract
Recently, adversarial attacks have shown to lead the state-of-the-art deep neural networks (DNNs) to misclassification. However, most adversarial attacks are generated according to whether they are perceptual to human visual system, measured by geometric metrics such as the$\ell _2$-norm, which ignores the common watermarks in cyber-physical systems. In this article, we propose a fast adversarial watermark attack (FAWA) method based on fast differential evolution technique, which optimally superimposes a watermark on an image to fool DNNs. We also attempt to explain the reason why the attack is successful and propose two hypotheses on the vulnerability of DNN classifiers and the influence of the watermark attack on higher-layer features extraction respectively. In addition, we propose two countermeasure methods against FAWA based on random rotation and median filtering respectively. Experimental results show that our method achieves 41.3 percent success rate in fooling VGG-16 and have good transferability. Our approach is also shown to be effective in deceiving deep learning as a service (DLaaS) systems as well as the physical world. The proposed FAWA, hypotheses, and the countermeasure methods, provide a timely help for DNN designers to gain some knowledge of model vulnerability while designing DNN classifiers and related DLaaS applications.
Hao Jiang 0010, Jintao Yang, Guang Hua 0001, Lixia Li, Shenghui Tu, Song Xia
IEEE Trans. Computers3
2024 Unambiguous and High-Fidelity Backdoor Watermarking for Deep Neural Networks
abstract
The unprecedented success of deep learning could not be achieved without the synergy of big data, computing power, and human knowledge, among which none is free. This calls for the copyright protection of deep neural networks (DNNs), which has been tackled via DNN watermarking. Due to the special structure of DNNs, backdoor watermarks have been one of the popular solutions. In this article, we first present a big picture of DNN watermarking scenarios with rigorous definitions unifying the black- and white-box concepts across watermark embedding, attack, and verification phases. Then, from the perspective of data diversity, especially adversarial and open set examples overlooked in the existing works, we rigorously reveal the vulnerability of backdoor watermarks against black-box ambiguity attacks. To solve this problem, we propose an unambiguous backdoor watermarking scheme via the design of deterministically dependent trigger samples and labels, showing that the cost of ambiguity attacks will increase from the existing linear complexity to exponential complexity. Furthermore, noting that the existing definition of backdoor fidelity is solely concerned with classification accuracy, we propose to more rigorously evaluate fidelity via examining training data feature distributions and decision boundaries before and after backdoor embedding. Incorporating the proposed prototype guided regularizer (PGR) and fine-tune all layers (FTAL) strategy, we show that backdoor fidelity can be substantially improved. Experimental results using two versions of the basic ResNet18, advanced wide residual network (WRN28_10) and EfficientNet-B0, on MNIST, CIFAR-10, CIFAR-100, and FOOD-101 classification tasks, respectively, illustrate the advantages of the proposed method.
Guang Hua 0001, Andrew Beng Jin Teoh, Yong Xiang 0001, Hao Jiang 0010
IEEE Trans. Neural Networks Learn. Syst.1
2023 "Seeing" Electric Network Frequency from Events
abstract
Most of the artificial lights fluctuate in response to the grid's alternating current and exhibit subtle variations in terms of both intensity and spectrum, providing the potential to estimate the Electric Network Frequency (ENF)from conventional frame-based videos. Nevertheless, the performance of Video-based ENF (V-ENF) estimation largely re-lies on the imaging quality and thus may suffer from significant interference caused by non-ideal sampling, motion, and extreme lighting conditions. In this paper, we show that the ENF can be extracted without the above limitations from a new modality provided by the so-called event camera, a neuromorphic sensor that encodes the light intensity variations and asynchronously emits events with extremely high temporal resolution and high dynamic range. Specifically, we first formulate and validate the physical mechanism for the ENF captured in events, and then propose a simple yet robust Event-based ENF (E-ENF) estimation method through mode filtering and harmonic enhancement. Furthermore, we build an Event-Video ENF Dataset (EV-ENFD) that records both events and videos in diverse scenes. Extensive experiments on EV-ENFD demonstrate that our proposed E-ENF method can extract more accurate ENF traces, outperforming the conventional V-ENF by a large margin, especially in challenging environments with object motions and extreme lighting conditions. The code and dataset are available at https://github.com/x1x-creater/E-ENF.
Lexuan Xu, Guang Hua 0001, Lei Yu 0006
CVPR2
2023 Deep fidelity in DNN watermarking: A study of backdoor watermarking for classification models
Guang Hua 0001, Andrew Beng Jin Teoh
Pattern Recognit.1
2023 SSVS-SSVD Based Desynchronization Attacks Resilient Watermarking Method for Stereo Signals
abstract
Most of the audio signals in real-world applications are stereo signals. However, the previous desynchronization attacks resilient watermarking methods cannot preserve perceptual quality or achieve robustness when constrained by high embedding rates and stereo host. In this paper, based on two novel features segmental singular values summation (SSVS) and segmental singular values difference (SSVD) that are generated using discrete cosine transform (DCT) and singular value decomposition (SVD), we present a robust watermarking method for stereo signals that not only is robust to desynchronization attacks and common signal processing attacks but also has a larger embedding rate compared with the previous methods. In the proposed method, we first apply DCT and SVD on each segment of the host signal to extract the SSVS feature and the SSVD feature. Then we generate the adaptive embedding parameters and embed watermark bits via optimized embedding strategies based on these features. Due to the use of the adaptive embedding parameters and the optimized embedding strategies, the proposed method significantly increases the embedding rate without compromising the robustness and perceptual quality. Analysis results show our proposed method outperforms the state-of-the-art methods by a large margin, where the perceptual quality improvement is over 14%, and the robustness against desynchronization attacks is improved by more than 49% when the embedding rate is 70 bps.
Juan Zhao 0007, Tianrui Zong, Yong Xiang 0001, Longxiang Gao, Guang Hua 0001, Keshav Sood, Yushu Zhang 0001
IEEE ACM Trans. Audio Speech Lang. Process.5
2023 Frequency Spectrum Modification Process-Based Anti-Collusion Mechanism for Audio Signals
abstract
The collusion attack combines multiple multimedia files into one new file to erase the user identity information. The traditional anti-collusion methods (which aim to trace the traitors) can defend the collusion attack, but they cannot well defend some hybrid collusion attacks (e.g., a collusion attack combined with desynchronization attacks). To address this issue, we propose a frequency spectrum modification process (FSMP) to defend the collusion attack by significantly downgrading the perceptual quality of the colluded file. The severe perceptual quality degradation can demotivate the attackers from launching the collusion attack. Because FSMP is orthogonal to the existing traitor-trace-based methods, it can be combined with the existing methods to provide a double-layer protection against different attacks. In FSMP, after several signal processing procedures (e.g., uneven framing and smoothing), multiple signals (called FSMP signals) can be generated from the host signal. Launching collusion attack using the generated FSMP signals would lead to the energy disturbance and attenuation effect (EDAE) over the colluded signals. Due to the EDAE, FSMP can significantly degrade the perceptual quality of the colluded audio file, thereby thwarting the collusion attack. In addition, FSMP can well defend different hybrid collusion attacks. Theoretical analysis and experimental results confirm the validity of the proposed method.
Juan Zhao 0007, Tianrui Zong, Yong Xiang 0001, Guang Hua 0001, Longxiang Gao, Gleb Beliakov
IEEE Trans. Cybern.4
2023 Categorical Inference Poisoning: Verifiable Defense Against Black-Box DNN Model Stealing Without Constraining Surrogate Data and Query Times
abstract
Deep Neural Network (DNN) models have offered powerful solutions for a wide range of tasks, but the cost to develop such models is nontrivial, which calls for effective model protection. Although black-box distribution can mitigate some threats, model functionality can still be stolen via black-box surrogate attacks. Recent studies have shown that surrogate attacks can be launched in several ways, while the existing defense methods commonly assume attackers with insufficient in-distribution (ID) data and restricted attacking strategies. In this paper, we relax these constraints and assume a practical threat model in which the adversary not only has sufficient ID data and query times but also can adjust the surrogate training data labeled by the victim model. Then, we propose a two-step categorical inference poisoning (CIP) framework, featuring both poisoning for performance degradation (PPD) and poisoning for backdooring (PBD). In the first poisoning step, incoming queries are classified into ID and (out-of-distribution) OOD ones using an energy score (ES) based OOD detector, and the latter are further classified into high ES and low ES ones, which are subsequently passed to a strong and a weak PPD process, respectively. In the second poisoning step, difficult ID queries are detected by a proposed reliability score (RS) measurement and are passed to PBD. In doing so, the first step OOD poisoning leads to substantial performance degradation in surrogate models, the second step ID poisoning further embeds backdoors in them, while both can preserve model fidelity. Extensive experiments confirm that CIP can not only achieve promising performance against state-of-the-art black-box surrogate attacks like KnockoffNets and data-free model extraction (DFME) but also work well against stronger attacks with sufficient ID and deceptive data, better than the existing dynamic adversarial watermarking (DAWN) and deceptive perturbation defense methods. PyTorch code is available athttps://github.com/Hatins/CIP_master.git.
Haitian Zhang, Guang Hua 0001, Xinya Wang, Hao Jiang 0010, Wen Yang 0001
IEEE Trans. Inf. Forensics Secur.2
2022 Mask-guided cycle-GAN for specular highlight removal
Yuanfeng Zheng, Haoran Yan, Guang Hua 0001
Pattern Recognit. Lett.4
2022 A Data-Driven High-Resolution Time-Frequency Distribution
abstract
The design of high-resolution and cross-term (CT) free time-frequency distributions (TFDs) has been an open problem. Classical kernel based methods are limited by the trade-off between resolution and CT suppression, even under optimally derived parameters. To break the current limitation, we propose a data-driven model directly based on Wigner-Ville distribution (WVD). The proposed data-driven high-resolution TFD (DH-TFD) includes several stacked multi-channel convolutional kernels. Specifically, convolutional layers with skipping operators are utilized to learn coarse features, while a weighted block is employed to refine these features independently in both channel and spatial dimensions. By doing so, CTs can be effectively eliminated while maintaining a high resolution. Numerical experiments on both synthetic and real-world data confirm the superiority of the proposed DH-TFD in simultaneously extracting and representing a target signal over state-of-the-art methods.
Lei Yu 0006, Guang Hua 0001
IEEE Signal Process. Lett.4
2021 Towards End-to-End Synthetic Speech Detection
abstract
The constant Q transform (CQT) has been shown to be one of the most effective speech signal pre-transforms to facilitate synthetic speech detection, followed by either hand-crafted (subband) constant Q cepstral coefficient (CQCC) feature extraction and a back-end binary classifier, or a deep neural network (DNN) directly for further feature extraction and classification. Despite the rich literature on such a pipeline, we show in this paper that the pre-transform and hand-crafted features could simply be replaced by end-to-end DNNs. Specifically, we experimentally verify that by only using standard components, a light-weight neural network could outperform the state-of-the-art methods for the ASVspoof2019 challenge. The proposed model is termed Time-domain Synthetic Speech Detection Net (TSSDNet), having ResNet- or Inception-style structures. We further demonstrate that the proposed models also have attractive generalization capability. Trained on ASVspoof2019, they could achieve promising detection performance when tested on disjoint ASVspoof2015, significantly better than the existing cross-dataset results. This paper reveals the great potential of end-to-end DNNs for synthetic speech detection, without hand-crafted features.
Guang Hua 0001, Andrew Beng Jin Teoh
IEEE Signal Process. Lett.1
2021 ENF Detection in Audio Recordings via Multi-Harmonic Combining
abstract
The detection of the electric network frequency (ENF) in digital recordings is an essential step before the subsequent ENF extraction and forensic analysis. In this letter, we extend the state-of-the-art single-tone time-frequency (TF) domain ENF detector to the multi-tone scenario and propose a multi-harmonic combining (MHC) method, exploiting ENF harmonic components for improved detection performance. To exclude the corrupted components interfering rather than contributing to ENF detection, the proposed detector first performs a pre-screening based on the estimated average subband signal-to-noise ratios (SNRs) to exclude interfering components. Then, with the selected harmonic candidates, a second screening process is applied based on the TF test statistics (TSs), i.e., the variances of observed subband traces. After that, the multi-harmonic components are combined to form the final TS, whose sign determines the final decision. The advantages of the proposed method are illustrated via both synthetic analysis and real-world experimental results using the ENF-WHU dataset.
Han Liao, Guang Hua 0001
IEEE Signal Process. Lett.2
2021 Segmental DCT Coefficient Reversal Based Anti-Collusion Audio Fingerprinting Mechanism
abstract
Collusion attacks are challenging to tackle in audio fingerprinting. A new direction to resist collusion attacks is to degrade the perceptual quality of the colluded files so that these files cannot be reused. The existing method in this direction has low embedding capacity and limited anti-collusion performance when the number of colluders is odd. In this letter, we present an anti-collusion mechanism that has a higher embedding capacity and can significantly degrade the perceptual quality of the colluded files regardless of the number of colluders. In the proposed mechanism, we first segment the host audio file into frames and perform the discrete cosine transform (DCT) on each frame. Then multiple fingerprint bits are embedded into each frame by reversing the DCT coefficients of the corresponding frequency band. As a result, when a collusion attack occurs, our proposed embedding mechanism can introduce perceptibly annoying differences between frames in the colluded file, which leads to severe perceptual quality degradation. Theoretical analysis and experimental results validate the superiority of the proposed anti-collusion mechanism.
Juan Zhao 0007, Tianrui Zong, Yong Xiang 0001, Longxiang Gao, Guang Hua 0001
IEEE Signal Process. Lett.5
2021 Non-Linear-Echo Based Anti-Collusion Mechanism for Audio Signals
abstract
Collusion attacks are considered to be challenging attacks in audio copyright protection. The traditional watermarking algorithms cannot identify the traitors when other attacks, such as desynchronization attacks, are applied with a collusion attack. Instead of tracing the traitors, in this paper we aim to tackle collusion attacks by removing the commercial value from the colluded copy, which will demotivate the attackers from launching collusion attacks. Since the commercial value of an audio signal is directly reflected by its perceptual quality, we propose a novel non-linear-echo generation (NLEG) based algorithm to significantly degrade the perceptual quality of the colluded copy by embedding a time delay sequence into the host signal. The proposed NLEG is also designed to be resilient to common signal processing attacks and desynchronization attacks. Furthermore, the proposed NLEG can be combined with other digital watermarking techniques to enhance its performance on protecting the copyright information. Experimental results show the validity of the proposed NLEG.
Tianrui Zong, Yong Xiang 0001, Iynkaran Natgunanathan, Longxiang Gao, Guang Hua 0001, Wanlei Zhou 0001
IEEE ACM Trans. Audio Speech Lang. Process.5
2021 Detection of Electric Network Frequency in Audio Recordings-From Theory to Practical Detectors
abstract
Recently, it has been discovered that the electric network frequency (ENF) could be captured by digital audio, video, or even image files, and could further be exploited in forensic investigations. However, the existence of the ENF in multimedia content is not a sure thing, and if the ENF is not present, ENF-based forensic analysis would become useless or even misleading. In this paper, we address the problem of ENF detection in digital audio recordings, which is modeled as the detection of a weak (ENF) signal contaminated by unknown colored wide-sense stationary (WSS) Gaussian noise, while the signal also contains multiple unknown random parameters. We first derive three Neyman-Pearson (NP) detectors, i.e., general matched filter (GMF), matched filter (MF)-like detector, and the asymptotic approximation of the GMF, and choose the MF-like detector as the clairvoyant detector. For practical detectors, we show that the generalized likelihood ratio test (GLRT) could not be efficiently obtained due to the unknown noise and large matrix inversion. Alternatively, we propose two least-squares (LS)-based time domain detectors termed as LS-likelihood ratio test (LRT) and naive-LRT. Further, we propose a time-frequency (TF) domain detector, termed as TF detector, which exploits the a priori knowledge of the ENF. The performances of the derived detectors are extensively analyzed in terms of test statistic distributions, threshold selection, and computational complexity. The naive-LRT detector is found to be only effective for very short recordings. As the data recording length increases, both LS-LRT and TF detectors yield effective detection results, while the latter is approximately a constant false alarm rate (CFAR) detector. Practical experiments using real audio recordings justify the effectiveness of the proposed detectors and our analysis.
Guang Hua 0001, Han Liao, Dengpan Ye
IEEE Trans. Inf. Forensics Secur.1
2021 Robust ENF Estimation Based on Harmonic Enhancement and Maximum Weight Clique
abstract
The electric network frequency (ENF) is an important and extensively researched forensic criterion to authenticate digital recordings, but currently it is still challenging to extract reliable ENF traces from recordings in uncontrollable environments. In this paper, we present a framework for robust ENF extraction from real-world audio recordings, featuring multi-tone harmonic ENF enhancement and graph-based harmonic selection. We first extend the recently developed single-tone robust filtering algorithm (RFA) to the multi-tone scenario and propose a harmonic robust filtering algorithm (HRFA). It can enhance each harmonic component without cross-component interference, thus alleviating the effects of unwanted noise and audio content. In addition, considering the fact that some harmonic components could still be severely corrupted after the HRFA, interfering rather than facilitating ENF estimation, we propose a graph-based harmonic selection algorithm (GHSA), which finds a subset of harmonic components having the overall highest mutual cross-correlation. Noticeably, the harmonic selection problem is found to be equivalent to the maximum weight clique problem in graph theory, and the Bron-Kerbosch algorithm is adopted in the GHSA. With the enhanced and carefully selected harmonic components, both the existing maximum likelihood estimator (MLE) and weighted MLE are incorporated to yield the final ENF estimation results. The proposed framework is evaluated using both synthetic signals and the ENF-WHU dataset consisting of 130 real-world audio recordings, demonstrating its advantages over both the existing single- and multi-tone competitors. This work further improves the applicability of the ENF as a forensic criterion in real-world situations.
Guang Hua 0001, Han Liao, Dengpan Ye, Jiayi Ma 0001
IEEE Trans. Inf. Forensics Secur.1
2020 Efficient Estimation of Mixing Matrix Using a Two-sensor Array
abstract
Blind source separation involves the estimation of mixing matrix given only observed mixtures. Considering the computational cost in practice, this paper proposes an efficient mixing matrix estimation (MME) method using an easily configured two-sensor array. Time-frequency (TF) analysis, which has been an important field of research for MME, is invoked herein to precisely detect single-source TF points (SSPs) in combination with the principle of matching pursuit. We theoretically verify that the residual TF vectors at SSPs derived by mixture TF vectors tend to have zero-valued elements, based on which a criterion is designed to identify a set of exact SSPs for accurate MME. Furthermore, the number of sources can be simultaneously estimated by adopting mean-shift clustering on SSPs. Numerical simulations are carried out on both real-valued and complex-valued mixing matrices to provide corroborating evidence for the theoretical claims.
Qinmengying Yan, Guang Hua 0001
ICASSP4
2020 Over-Complete-Dictionary-Based Improved Spread Spectrum Watermarking Security
abstract
This letter presents a theoretical analysis of the security of two over-complete-dictionary-based improved spread spectrum (ISS) watermarking systems, i.e., subspace-ISS (sub-ISS) and random matching pursuit (RMP)-ISS, in comparison with the conventional orthogonal-transform-based ISS system. Considering Kerckhoffs' principle and the watermarked-only attack (WOA) scenario, we derive the associated decoding regions under brute force attack in which the adversary randomly draws probing keys until getting access to the watermarking channel. We first reveal that the sub-ISS system is as equally secure as the conventional ISS system in terms of message decoding, but the sensed key could provide computational security of the spreading key. More importantly, the RMP-ISS system is shown to be securer than the conventional ISS system for it being able to survive brute force attack. Based on the findings, we further propose an empirical design mechanism that optimizes the security of the RMP-ISS scheme.
Guang Hua 0001
IEEE Signal Process. Lett.1
2020 Informed Histogram-Based Watermarking
abstract
The existing works on histogram-based watermark embedding share the common notions of non-informed random host sample selection and uniform embedding, resulting in limited and yet similar performances. This letter develops two methods to improve imperceptibility and robustness of histogram-based embedding. We first propose a content-aware method which performs uniform embedding in ranked energy-significant regions of the host signal. We show that the embedded ternary watermark signal is more likely to be masked by the strong host signal component, thus improving imperceptibility without compromising robustness. Further, a non-uniform embedding method is proposed, which modifies host samples according to their relationships with the center of the target histogram bin. It ensures minimum host sample modifications without reducing the payload size. Meanwhile, watermark robustness could also be improved since it becomes the most difficult to move samples at bin centers into other bins via attacks. The proposed methods are supported by extensive experimental results.
Guang Hua 0001, Yong Xiang 0001, Leo Yu Zhang
IEEE Signal Process. Lett.1
2020 ENF Signal Enhancement in Audio Recordings
abstract
In electric network frequency (ENF) based audio forensics, the ENF signal captured in a questioned audio recording is estimated and analyzed for authentication purposes. However, the captured ENF signal is usually contaminated by very strong noise and interference. In this paper, we propose a robust filtering algorithm (RFA) for ENF signal enhancement in audio recordings, which could effectively suppress the additive noise and facilitate subsequent ENF estimation, especially in practical low signal-to-noise ratio (SNR) situations. The proposed algorithm encodes the time domain expression of the preprocessed audio signal (ENF signal plus noise) as the instantaneous frequencies (IFs) of an analytical sinusoidal frequency modulated (SFM) signal. Then, a kernel function is utilized to generate a sinusoidal time-frequency distribution (STFD) whose peaks correspond to the IFs of the analytical signal, i.e., the denoised ENF signal. It is then proven that finding the STFD peaks is equivalent to finding the averaged phases of the kernel function if the additive noise is a zero mean wide sense stationary (WSS) process. The RFA serves as a noise reduction mechanism yielding improved SNR at the filter output. Combined with generic frequency estimation methods, ENF extraction accuracy could be substantially improved with the use of the proposed RFA than without using it. Both synthetic and experimental results are provided to illustrate the effectiveness of our proposal. Reliable ENF extraction could be achieved under noise level down to -20 dB SNR, and the RFA is suitable for a wide range of ENF-based forensic applications.
Guang Hua 0001
IEEE Trans. Inf. Forensics Secur.1
2019 Random Matching Pursuit for Image Watermarking
abstract
The classical solution to an underdetermined system of linear equations mainly has two opposite directions, which lead to either a large ℓ2-norm sparse solution or a non-sparse minimum ℓ2-norm solution. In this paper, we systematically show that by modifying the well-known basic matching pursuit algorithm originally proposed to identify the sparse solution, an alternative solution between the two classical ones could be obtained. The modified algorithm, termed as random matching pursuit (RMP), is then used to create a novel image watermarking framework. Compared to conventional systems, the security is substantially improved by the use of random over-complete dictionaries and the order parameter of RMP. Capacity can also be increased thanks to the transform with over-complete dictionaries that could expand signal dimension. Meanwhile, imperceptibility and robustness properties of the proposed design framework are not compromised. The classical spread spectrum and improved spread spectrum techniques are applied to the proposed framework for practical implementations. The novelty and effectiveness of the proposed systems are supported by rigorous performance analysis and experimental results using an image data set. This paper reveals the potential of using over-complete dictionaries in multimedia watermarking systems, which theoretically leads to the exploration of alternative candidates among the infinite solutions to underdetermined linear systems other than minimum ℓ2-norm and sparse ones.
Guang Hua 0001, Lifan Zhao, Guoan Bi, Yong Xiang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2019 Compressed Sensing Based Selective Encryption With Data Hiding Capability
abstract
This paper proposes a joint selective encryption and data hiding scheme based on compressed sensing (CS), with a focus to its application in secure imaging. Specifically, working with a semantic-secure stream cipher, we suggest to selectively encrypt the sign bits of the CS measurements during its quantization stage and insert the authentication information using a nonseparable histogram-shifting based data hiding scheme. The rationale behind the sign encryption is that CS measurements, when measured by random subspace projection, is random in nature and thus, from both theoretical and experimental points of view, the mean squared errors associated with authorized users and attackers are significant. Due to the indistinguishability of the output ciphertext and the nonlinearity of the CS decoder, it is robust, when comparing with the existing selective encryption system of multimedia data, against known error concealment attacks. When applied in imaging, we demonstrate that it could effectively degrade the visual quality level while saving the computation load by at least 90%. Moreover, we further show that a state-of-the-art data hiding system can be seamlessly incorporated into the sign encryption, thus allowing soft data authentication without heavy computation. The proposed scheme is expected to strengthen the security of applications in the field where both energy and privacy are the concerns, such as sensitive information protection for multimedia data in wireless sensor networks.
Jia Wang 0008, Leo Yu Zhang, Junxin Chen 0001, Guang Hua 0001, Yushu Zhang 0001, Yong Xiang 0001
IEEE Trans. Ind. Informatics4
2019 Improving the Visual Quality of Size-Invariant Visual Cryptography for Grayscale Images: An Analysis-by-Synthesis (AbS) Approach
abstract
In visual cryptography (VC) for grayscale image, size reduction leads to bad perceptual quality to the reconstructed secret image. To improve the quality, the current efforts are limited to the design of VC algorithm for binary image, and measuring the quality with metrics that are not directly related to how the human visual system (HVS) perceives halftone images. We propose an analysis-by-synthesis (AbS) framework to integrate the halftoning process and the VC encoding: the secret pixel/block is reconstructed from the shares in the encoder and the error between the reconstructed secret and the original secret images is fed back and compensated concurrently by the error diffusion process. In doing so, the error between the reconstructed secret and original secret is pushed to high frequency band, thus producing visually pleasing reconstructed secret image. This framework is simple and flexible in that it can be combined with many existing size-invariant VC algorithms, including probabilistic VC, random grid VC and vector/block VC. More importantly, it is proved that this AbS framework is as secure as the traditional VC algorithms. Experimental results demonstrate the effectiveness of the proposed AbS framework.
Bin Yan 0001, Yong Xiang 0001, Guang Hua 0001
IEEE Trans. Image Process.3
2018 Face Spoofing Video Detection Using Spatio-Temporal Statistical Binary Pattern
abstract
Face has been used as a popular biometric trait to identify a person. However, attacking such face recognition system is not challenging today by using for example a fake face photo or video in front of a camera. In this paper, we present a novel feature, namely, SBP-TOP, to effectively identify such spoofings. SBP-TOP presents the texture information for the face region from both spatial and temporal perspectives. We have tested the proposed feature on two well-known face spoofing datasets: new Michigan State University mobile face spoofing database (MSU MFSD) and CASIA Face Anti-Spoofing Database (CASIA). The results indicate an accuracy over 95% on both datasets and there is an improvement over the state-of-the-art feature by around 10% and 3.2%, respectively.
Ying Zhang 0047, Rohit Kumar Dubey, Guang Hua 0001, Vrizlynn L. L. Thing
TENCON3
2018 An overview of protection of privacy in multibiometrics
Iynkaran Natgunanathan, Abid Mehmood, Yong Xiang 0001, Guang Hua 0001, Gang Li 0009, Shaun Bangay
Multim. Tools Appl.4
2018 Spread Spectrum Audio Watermarking Using Multiple Orthogonal PN Sequences and Variable Embedding Strengths and Polarities
abstract
Copyright protection of audio data is a serious problem and spread spectrum (SS) based audio watermarking is a promising technology to tackle this problem. Although a number of SS-based audio watermarking methods have been reported in the literature, they cannot achieve high robustness and embedding capacity at the same time. In this paper, we propose a novel SS-based audio watermarking method that can embed a large number of watermark bits into an audio signal without compromising the robustness against common attacks. Compared with the existing audio watermarking methods, the proposed one is especially robust against severe noise addition and compression attacks, while achieving high embedding capacity. Moreover, the new audio watermarking method is computationally efficient. The validity of the proposed SS-based audio watermarking method is demonstrated by simulation results.
Yong Xiang 0001, Iynkaran Natgunanathan, Dezhong Peng, Guang Hua 0001, Bo Liu 0001
IEEE ACM Trans. Audio Speech Lang. Process.4
2017 Underdetermined blind separation of overlapped speech mixtures in time-frequency domain with estimated number of sources
Guang Hua 0001, Lei Yu 0006, Yunlong Cai, Guoan Bi
Speech Commun.2
2017 Patchwork-Based Multilayer Audio Watermarking
abstract
A multilayer watermarking system is a system that is able to embed watermarks to a host media signal repeatedly in an overlaying manner, without incurring troubles in extracting the watermarks in each layer. In this paper, we present a novel patchwork-based audio watermarking algorithm that can embed and extract watermark bits successfully in such a multilayer framework. In the proposed method, a new watermark embedding algorithm is designed to ensure that the embedded watermarks in a certain layer do not affect the detection of watermarks in other layers. Adding multiple layers of watermark bits inevitably reduces the perceptual quality. However, to minimize the perceptual quality degradation in multilayer watermarking, the audio fragments for watermark embedding are selected from a set of specially arranged discrete cosine transform coefficients of the host audio signal. Watermark embedding is achieved by modifying the mean values of selected sample fragments. With the use of an embedding error buffer, the proposed system can withstand a wide range of common attacks. To maintain the balance between the perceptual quality and robustness, watermark embedding strength is adjusted according to the specific layer used. The proposed multilayer scheme ensures the independence of the processing in different layers. The effectiveness of the proposed system is demonstrated and verified by extensive simulation results.
Iynkaran Natgunanathan, Yong Xiang 0001, Guang Hua 0001, Gleb Beliakov, John Yearwood
IEEE ACM Trans. Audio Speech Lang. Process.3
2016 Twenty years of digital audio watermarking - a comprehensive review
abstract
Digital audio watermarking is an important technique to secure and authenticate audio media. This paper provides a comprehensive review of the twenty years’ research and development works for digital audio watermarking, based on an exhaustive literature survey and careful selections of representative solutions. We generally classify the existing designs into time domain and transform domain methods , and relate all the reviewed works using two generic watermark embedding equations in the two domains. The most important designing criteria, i.e., imperceptibility and robustness, are thoroughly reviewed. For imperceptibility , the existing measurement and control approaches are classified into heuristic and analytical types, followed by intensive analysis and discussions. Then, we investigate the robustness of the existing solutions against a wide range of critical attacks categorized into basic, desynchronization, and replacement attacks, respectively. This reveals current challenges in developing a global solution robust against all the attacks considered in this paper. Some remaining problems as well as research potentials for better system designs are also discussed. In addition, audio watermarking applications in terms of US patents and commercialized solutions are reviewed. This paper serves as a comprehensive tutorial for interested readers to gain a historical, technical, and also commercial view of digital audio watermarking.
Guang Hua 0001, Jiwu Huang, Yun Q. Shi 0001, Jonathan Goh, Vrizlynn L. L. Thing
Signal Process.1
2016 When Compressive Sensing Meets Data Hiding
abstract
We present a novel framework of performing multimedia data hiding using an over-complete dictionary, which brings compressive sensing to the application of data hiding. Unlike the conventional orthonormal full-space dictionary, the over-complete dictionary produces an underdetermined system with infinite transform results. We first discuss the minimum norm formulation (ℓ2-norm) which yields a closed-form solution and the concept of watermark projection, so that higher embedding capacity and an additional privacy preserving feature can be obtained. Furthermore, we study the sparse formulation (ℓ2-norm) and illustrate that as long as the ℓ0-norm of the sparse representation of the host signal is less than the signal's dimension in the original domain, an informed sparse domain data hiding system can be established by modifying the coefficients of the atoms that have not participated in representing the host signal. A single support modification-based data hiding system is then proposed and analyzed as an example. Several potential research directions are discussed for further studies. More generally, apart from the ℓ2- and ℓ0-norm constraints, other conditions for reliable detection performance are worth of future investigation.
Guang Hua 0001, Yong Xiang 0001, Guoan Bi
IEEE Signal Process. Lett.1
2016 Audio Authentication by Exploring the Absolute-Error-Map of ENF Signals
abstract
Recently, the electric network frequency (ENF), a natural signature embedded in many audio recordings, has been utilized as a criterion to examine the authenticity of audio recordings. ENF-based audio authentication system involves extraction of the ENF signal from a questioned audio recording, and matching it with the reference signal stored in an ENF database. This establishes a popular application of audio timestamp verification. In this paper, we explore another important application, i.e., ENF-based audio tampering detection, which has received less research attention. Specifically, we introduce the absolute-error-map (AEM) between the ENF signals obtained from the testing audio recording and the database. The AEM serves as an ensemble of the raw data associated with the ENF matching process. Through intensive analysis of the AEM, we propose two algorithms to jointly deal with timestamp verification and tampering detection, including insertion, deletion, and splicing attacks, respectively. The first algorithm is based on exhaustive point search and measurement, while the second algorithm leverages the image erosion technique to achieve fast detection of tampering type and tampered region, thus the second algorithm sacrifices some accuracy for speed. The authentication mechanism is that the system first determines if the testing data have been tampered with, and then outputs the timestamp information if no tampering is detected. Otherwise, it outputs the tampering type and tampered region. We demonstrate the effectiveness of the proposed solution via both synthetic and practical examples from our practically deployed audio authentication system.
Guang Hua 0001, Ying Zhang 0047, Jonathan Goh, Vrizlynn L. L. Thing
IEEE Trans. Inf. Forensics Secur.1
2015 Robust transmit beampattern design for uniform linear arrays using correlated LFM waveforms
abstract
This paper presents a robust design of the transmit beampattern for uniform linear antenna arrays. Existing designs are usually completed at the stage of achieving an optimal transmit covariance matrix from identifying a weighting matrix with the assumption of ideally orthogonal waveforms. However, we propose a compensation technique to achieve the optimal covariance matrix without the requirement of orthogonality. The corresponding solutions identify a set of weighting matrices that are robust against the imperfection of the waveforms. As a result, a set of easy-to-generate partially correlated linear frequency modulated (LFM) waveforms can be used to achieve identical transmit beampatterns which could be synthesized by ideally orthogonal multiple-input multiple-output (MIMO) radar waveforms. The proposed robust design is evaluated via numerical examples.
Guang Hua 0001, Saman S. Abeysekera
ICASSP1
2015 Time-Spread Echo-Based Audio Watermarking With Optimized Imperceptibility and Robustness
abstract
We present a time-spread echo-based audio watermarking scheme with optimized imperceptibility and robustness. Specifically, convex optimization based finite-impulse-response (FIR) filter design is utilized to obtain the optimal echo filter coefficients. The desired power spectrum of the echo filter is shaped by the proposed maximum power spectral margin (MPSM) and the absolute threshold of hearing (ATH) of human auditory system (HAS) to ensure the optimal imperceptibility. Meanwhile, the auto-correlation function of the echo filter coefficients is specified as the constraint in the problem formulation, which controls the robustness in terms of watermark detection. In this way, a joint optimization of imperceptibility and robustness can be quantitatively performed. As a result, the proposed watermarking scheme is superior to existing solutions such as the ones based on pseudo noise (PN) sequence or modified pseudo noise (MPN) sequence. Note that the designed echo kernel is also highly secure in that only with the same filter coefficients can one successfully detect the watermark. Experimental results are provided to evaluate the imperceptibility and robustness of the proposed watermarking scheme.
Guang Hua 0001, Jonathan Goh, Vrizlynn L. L. Thing
IEEE ACM Trans. Audio Speech Lang. Process.1
2015 Cepstral Analysis for the Application of Echo-Based Audio Watermark Detection
abstract
Cepstral analysis is an important signal processing procedure for audio watermark detection in echo-based audio watermarking systems. However, with the use of two common versions, i.e., complex and real cepstra, this procedure is usually treated as a very standard routine. This paper starts from noting inappropriate cepstral analysis from existing works, and provides rigorous derivations to reveal the advantages of using real cepstrum than complex cepstrum in echo-based audio watermark detection. Furthermore, we introduce two alternatives, termed as real part and imaginary part cepstrum, respectively, based on which a joint detection scheme is proposed. This is achieved by noting that both real part and imaginary part cepstra contain a full version of the echo kernel coefficients, which can be appropriately combined to obtain a composite cepstrum to further suppress the interferences. The advantages of the joint detection scheme over conventional approach using real cepstrum are illustrated via both performance analysis and experimental results. The accuracies of the mathematical approximations for each version of cepstrum are evaluated by normalized misalignment. The detection robustness is evaluated using the peak-to-average power ratio. The relationships among echo length, echo delays, and scaling factor, during watermark detection phase, are also discussed. Experimental results of watermark detection rate are provided to compare the performance of complex, real, and composite cepstra, respectively.
Guang Hua 0001, Jonathan Goh, Vrizlynn L. L. Thing
IEEE Trans. Inf. Forensics Secur.1
2014 A Dynamic Matching Algorithm for Audio Timestamp Identification Using the ENF Criterion
abstract
The electric network frequency (ENF) criterion is a recently developed technique for audio timestamp identification, which involves the matching between extracted ENF signal and reference data. For nearly a decade, conventional matching criterion has been based on the minimum mean squared error (MMSE) or maximum correlation coefficient. However, the corresponding performance is highly limited by low signal-to-noise ratio, short recording durations, frequency resolution problems, and so on. This paper presents a threshold-based dynamic matching algorithm (DMA), which is capable of autocorrecting the noise affected frequency estimates. The threshold is chosen according to the frequency resolution determined by the short-time Fourier transform (STFT) window size. A penalty coefficient is introduced to monitor the autocorrection process and finally determine the estimated timestamp. It is then shown that the DMA generalizes the conventional MMSE method. By considering the mainlobe width in the STFT caused by limited frequency resolution, the DMA achieves improved identification accuracy and robustness against higher levels of noise and the offset problem. Synthetic performance analysis and practical experimental results are provided to illustrate the advantages of the DMA.
Guang Hua 0001, Jonathan Goh, Vrizlynn L. L. Thing
IEEE Trans. Inf. Forensics Secur.1
2013 On the transmit beampattern design using MIMO and phased-array radar
abstract
The design of the transmit beampattern of a radar system aims to focus the energy of the transmitted signals to the desired spatial section(s) in order to enhance the direction of arrival (DOA) estimation at the receiver. In this paper, we consider this problem for both multiple-input-multiple-output (MIMO) and phased-array radar from the perspective of finite impulse response (FIR) filter design. For MIMO radar, we formulate the design as a feasibility problem (FP) and show its advantages over the existing methods. For phased-array radar, we use the spectral factorization technique which has been considered to be better than conventional filter design methods. It is shown that the two systems can achieve similar transmit beampatterns for identical number of transmit antennas and waveform energy. This is contradictory to the existing argument that MIMO radar can achieve more flexible transmit beampattern design over its phased-array counterpart. Simulation results are provided and discussed.
Guang Hua 0001, Saman S. Abeysekera
ICASSP1
2012 Colocated MIMO radar transmit beamforming using orthogonal waveforms
abstract
Multiple-input-multiple-output (MIMO) radar transmit beamforming mainly relies on designing transmitted signals with an appropriate covariance matrix. The signals can be designed using a two-step method which first optimizes the covariance matrix and then searches for the signals accordingly. A more efficient way is to synthesize transmitted signals by designing a weight matrix given a set of orthogonal waveforms, which makes use of both MIMO waveform diversity and phased-array transmit gain. In this paper, we propose a method to design the transmit beampattern by solving a semidefinite programming (SDP) problem. Then the eigen-decomposition of the optimal covariance matrix yields the weight matrix. Therefore it is called the SDP-EIG method. As a result, the overall transmitted waveforms are obtained more simply and efficiently, and the number of orthogonal signals required to form a desired beam reaches its minimum.
Guang Hua 0001, Saman S. Abeysekera
ICASSP1