VLDB 2026 Research / reviewers in the wild / expert
Bin Ma 0003
dblp:70/6176-3
· DBLP profile ↗
113ranked-venue papers
21as first author
97since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 54 · 7 first-author · 50 since 2021Artificial intelligence and machine learning · 21 · 2 first-author · 21 since 2021Security and privacy · 21 · 8 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 11 · 3 first-author · 6 since 2021Computer networks · 8 · 1 first-author · 8 since 2021Databases, data management, data science and information retrieval · 6 · 6 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Robust Facial Feature Watermarking Method Based on Learnable Quantum Controlled Encryption for Deepfake Detection
Shihui Xue, Yongjin Xian, Bin Ma 0003 |
ICIC (11) | 3 |
| 2026 | A survey on JPEG image forensics: Exploring key advances and persistent challenges in compression and quantization analysis
Hao Wang 0060, Xin Cheng 0018, Jiawei Zhang 0011, Hao Wu 0078, Xue Xie, Xiangyang Luo 0001, Bin Ma 0003 |
Comput. Secur. | 7 |
| 2026 | Anti-forensic for quantization steps estimation based on direct and preemptive adversarial attacks
Jiawei Zhang 0011, Hao Wu 0078, Xin Cheng 0018, Xiangyang Luo 0001, Bin Ma 0003, Hao Wang 0060 |
Eng. Appl. Artif. Intell. | 6 |
| 2026 | Sureillance camera authentication system based on PRNU
Jian Li 0034, Lisheng Yan, Bin Ma 0003, Xiaolong Li 0001, Zhenxing Qian |
Expert Syst. Appl. | 3 |
| 2026 | PRNU-adapted deep fingerprint learning and reparameterized correlation for camera source identification
Jian Li 0034, Bin Ma 0003, Xiaolong Li 0001, Zhenxing Qian |
Expert Syst. Appl. | 3 |
| 2026 | An Arbitrary Oblivious Selection-Enabled Privacy-Preserving Cross-Chain Regulatory Market
Bin Ma 0003, Zijie Gao, Yongjin Xian, Xiangyang Luo 0001 |
IEEE Internet Things J. | 1 |
| 2026 | Cross-chain identity privacy protection scheme based on oblivious transfer protocol and key agreement
Yuli Wang, Zhichao Cai, Bin Ma 0003 |
Inf. Sci. | 3 |
| 2026 | Hybrid RDH Method for JPEG Images Based on Quantization-Table-Modification and STCabstractQuantization-table-modification (QTM) is widely utilized in current studies of reversible data hiding (RDH) for JPEG images. However, the performance of existing QTM-based methods is far from optimal due to the uniform quantization step division and imperfect distortion modeling. In this letter, by incorporating Syndrome-Trellis-Code (STC) into QTM, a novel hybrid JPEG images RDH method is proposed. Firstly, instead of the uniformly dividing strategy conducted in previous works, by adaptively dividing the quantization steps, a hybrid embedding mechanism combining binary and ternary embedding is proposed. Then, the corresponding capacity-distortion model is established, by which the spatial domain distortion is estimated. Finally, based on the derived capacity-distortion model, for performance optimization, STC is utilized to minimize the cover modification. In this way, JPEG images RDH can be effectively conducted so that the visual quality of the marked image is well maintained. Experimental results demonstrate that the proposed method significantly outperforms some state-of-the-art works in terms of visual quality. Jiuchao Ban, Mengyao Xiao, Xiaolong Li 0001, Bin Ma 0003, Yao Zhao 0001 |
IEEE Signal Process. Lett. | 4 |
| 2026 | Re-Cropping Framework: A Grid Recovery Method for Quantization Step Estimation in Non-Aligned Recompressed ImagesabstractThe manipulation history of Joint Photographic Experts Group (JPEG) compression plays an important role in JPEG image forensics and information hiding. For non-aligned recompressed images, different cropping methods produce non-aligned outputs with varying feature distributions. One such important factor is the shifts of the discrete cosine transform (DCT) grid (i.e., the misalignment parameters) between two compression processes. Although many methods have been proposed to estimate the misalignment parameters, the limited amount of useful information available in small-sized images leads to low accuracy of these methods. To enhance the accuracy of misalignment parameter estimation for small-sized non-aligned images, we propose a novel two-branch network structure that accounts for the unique horizontal and vertical characteristics of non-aligned images. This structure employs convolution to simulate second-order difference (SOD) and incorporates it throughout the training process to optimize the difference parameters dynamically. Based on the insight that cropping operations leave traces in all color channels, we derive the Cg channel through a color space transformation. This approach expands the input dimensionality to four channels (Y, Cb, Cr, and Cg), thereby compensating for the information scarcity in small-sized images. The experimental results show that our method outperforms existing methods on different image sizes, regardless of the known or unknown quality factor (QF) of the first compression. Finally, we propose a re-cropping framework based on the estimated misalignment parameters. The influence of the first cropping is counteracted by a re-cropping operation, which improves the accuracy of existing methods in estimating the first quantization step for non-aligned recompressed images. Xin Cheng 0018, Hao Wang 0060, Xiangyang Luo 0001, Qingxiao Guan, Bin Ma 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | A Meta-Learning-Based Active Defense Scheme Against Deep Facial Forgery AttacksabstractDeepfake technology poses a serious threat to society by synthesizing a victims facial features and attributes to carry out deception. Traditional active defense methods against deepfake attacks are typically designed for specific models, and protected images often lose their anti-forgery capability after compression or reconstruction, severely limiting their practical applicability. This paper proposes a Meta-Learning-based active defense Scheme against deep facial forgery attacks (MLPDS), which effectively safeguards facial images against diverse deepfake attacks in real-world scenarios. Our approach adopts a general paradigminjecting noise into the original image to construct a cross-model defense algorithm against deepfake attacks. Specifically, by leveraging a meta-learning strategy, we integrate perturbations generated by multiple deepfake models, enabling robust protection against a variety of forgery models. Furthermore, to maintain the high fidelity of the images, we propose a symmetric gradient quantization strategy based on the arctan function to minimize the perceptual discrepancy between the perturbed and original images. Finally, an end-to-end optimization network is employed to generate universal perturbations tailored to specific images, supported by a pixel-level error metric that constrains deviations from the original content. Since no retraining is required to protect newly encountered images, this approach significantly improves the efficiency and practicality of real-time anti-deepfake defense. Experiments show that the proposed MLPDS algorithm can effectively resist attacks from multiple forgery models, outperforming state-of-the-art defense methods and significantly reducing image distortion with an average PSNR gain of approximately 7 dB, which fully meets the practical desire for efficient and reliable deepfake defense. Bin Ma 0003, Meihong Yang, Jian Xu 0025, Yongjin Xian, Xiaolong Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | An End-to-End Framework for Joint Makeup Style Transfer and Image SteganographyabstractExisting image steganography schemes always introduce obvious modification traces to the cover image, resulting in the risk of secret information leakage. To address this issue, an end-to-end framework for joint makeup style transfer and image steganography is proposed in this paper to achieve imperceptible higher-capacity data hiding. In the scheme, a Parsing-guided Semantic Feature Alignment (PSFA) module is designed to transfer the style of a makeup image to an object non-makeup image, thereby generating a content-style integrated feature matrix. Meanwhile, a Multi-Scale Feature Fusion and Data Embedding (MFFDE) module was devised to encode the secret image into its latent features and fuse them with the generated content-style integrated feature matrix, as well as the non-makeup image features across multiple scales, to achieve the makeup-stego image. As a result, the style of the makeup image is well transformed and the secret image is imperceptibly embedded simultaneously without directly modifying the pixels of the original non-makeup image. Additionally, a Residual-aware Information Compensation Network (RICN) is developed to compensate the loss of the secret image arising from the multilevel data embedding, thereby further enhancing the quality of the reconstructed secret image. Experimental results show that the proposed scheme achieves superior steganalysis resistance capability and visual quality in both makeup-stego images and recovered secret images, compared with other state-of-the-art schemes. Meihong Yang, Bin Ma 0003, Jian Xu 0025, Yongjin Xian, Linna Zhou |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | Toward Robust Proactive Deepfake Detection via Orthogonal Moment Watermarking
Chunpeng Wang 0001, Xianqiu Xu, Shanshan Zhang 0001, Bin Ma 0003, Qi Li 0029, Yunan Liu 0001 |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2026 | Rethinking Cross-Table Quantization Step Estimation: From Global and Local PerspectivesabstractThe quantization step is a crucial parameter in the JPEG compression process, and provides prior knowledge for JPEG image steganography and forensics. Existing neural network-based methods typically estimate the quantization steps for all discrete cosine transform (DCT) subbands jointly, by treating the entire quantization table as a unified input and leveraging the inter-subband relationships. However, subband relationships vary across different quantization tables, leading to poor generalization for methods that rely heavily on such relationships. To address the above issues, we depart from the strategy that relies on inter-subband relationships and instead train the model on a specific single subband. To compensate for the possible decrease in accuracy due to the lack of relationships between subbands, we extract the ranking features and histogram features from the DCT coefficient histograms of the subbands. Ranking features capture local patterns in DCT histograms by modeling the relative relationships between neighboring coefficients, thereby compensating for the absence of local detail. On the other hand, histogram features represent the overall distribution pattern of the DCT coefficient histograms and capture the global trends and statistical properties in the subbands. We subsequently employ convolutional groups and multilayer perceptron (MLP) structures to extract compression artifacts from these two features. Finally, we introduce a comprehensive evaluation metric, called GenAQt, to quantify the algorithm’s generalization ability across quantization tables. The experimental results demonstrate that our method maintains high accuracy across quantization tables, with RelGenAQt (relative accuracy decrease) exceeding 81% and AbsGenAQt (absolute accuracy decrease) being less than 0.38. Xin Cheng 0018, Hao Wang 0060, Xiangyang Luo 0001, Bin Ma 0003, Baowei Wang, Bin Li 0011 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2026 | Security Enhancement for Person Re-Identification Through Diffusion Driven Semantic Attacks
Kaixin Du, Bin Ma 0003, Meihong Yang, Jian Xu 0025, Xiaolong Li 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2026 | HENet: A Heterogeneous Encoding Network for General and Robust Adversarial Example GenerationabstractGenerator-based adversarial attack methods aim to fool deep neural networks (DNNs) by training a generator for crafting adversarial examples (AEs). However, as DNNs evolve from Convolutional Neural Networks (CNNs) to Transformers, the existing generator-based methods can hardly achieve satisfactory attack performance against different target model architectures in semi-whitebox attack scenarios. In addition, the generated AEs are susceptible to various distortions (especially for JPEG compression with low quality factors), which deteriorate the attack ability and increase the unreliability. To address these issues, we propose a dual-branch guided generative model called Heterogeneous Encoding Network (HENet) to form a robust generator-based adversarial attack framework. Specifically, our HENet introduces an Adaptive Feature Fusion Module (AFFM) to solve the dimensions and representativeness contradictions between CNNs and Transformers, which steers the perturbation generation based on a richer latent space and achieves better general attack ability. To further improve the robustness against JPEG compression, we design and integrate a Dynamic Differentiable JPEG Simulator (DDJS), which introduces an adaptive quantization mask to determine the flow of the gradient backpropagation in each frequency position. Extensive experiments prove the proposed method achieves a better attack success rate, lower perturbation magnitude, and higher robustness for various target network architectures under compressed, distorted, and lossless scenarios. Our codes will be made publicly available. Jiawei Zhang 0011, Hao Wang 0060, Hao Wu 0078, Bin Li 0011, Xiangyang Luo 0001, Bin Ma 0003 |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2026 | Boosting Targeted Adversarial Transferability with Feature Contrastive OptimizationabstractTransferable adversarial examples (AEs) have attracted considerable attention due to their ability to expose vulnerabilities in black-box deep neural networks (DNNs). However, achieving superior transferability for targeted attacks remains a challenge. In this article, inspired by the observation that AEs with smaller intra-class distances and larger inter-class distances tend to exhibit higher transferability, we propose a novel targeted attack based on Feature Contrastive Optimization (FCO). This attack enhances adversarial transferability by minimizing intra-class distances and maximizing inter-class distances. Specifically, we first define positive samples (belonging to the target class) and negative samples (belonging to non-target classes) that correspond to targeted AEs. Subsequently, leveraging these defined positive and negative samples, we propose two metrics—Intra-class Compactness (IC) and Inter-class Separability (IS)—to construct a novel Feature Contrastive (FC) loss. By integrating this plug-and-play FC loss into standard adversarial objectives, the generated AEs are encouraged to better align with the target class distribution while diverging from those of non-target classes. Extensive experiments on the ImageNet-compatible dataset demonstrate that our approach consistently improves targeted transferability across a broad range of DNN architectures. Jingtian Wang, Xiaolong Li 0001, Jian Li 0034, Bin Ma 0003, Yao Zhao 0001, Jinhua Zeng |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2025 | Speed Master: Quick or Slow Play to Attack Speaker RecognitionabstractBackdoor attacks pose a significant threat during the model's training phase. Attackers craft pre-defined triggers to break deep neural networks, ensuring the model accurately classifies clean samples during inference yet erroneously classifies samples added with these triggers. Recent studies have shown that speaker recognition systems trained on large-scale data are susceptible to backdoor attacks. Existing attackers employ unnoticed ambient sounds as triggers. However, these sounds are not inherently part of the training samples themselves. In essence, triggers can be designed to maintain an intrinsic connection with the original speech to enhance stealthiness. Our paper presents a novel attack methodology named Speed Master, which undermines deep neural networks by manipulating the speed of speech samples. Specifically, we execute poison-only backdoor attacks using speed or tempo adjustment. Changes in speech rate have become a common occurrence, as seen on platforms that allow users to adjust playback speed. In real-world scenarios, people naturally adjust their speaking rate depending on the context. As a result, changes in a speaker’s speech rate are typically perceived as normal and are unlikely to raise suspicion. Furthermore, detecting such subtle adjustments becomes challenging for users without reference speech. Our comprehensive experiments demonstrate that Speed Master can achieve an ASR over 99% in the digital domain, with only a 0.6% poisoning rate. Additionally, we validate the feasibility of Speed Master in the real world and its resistance to typical defensive measures. Zhe Ye 0001, Ying Ren, Xiangui Kang, Diqun Yan, Bin Ma 0003, Shiqi Wang 0001 |
AAAI | 6 |
| 2025 | SML: A Backdoor Defense for Non-Intrusive Speech Quality Assessment via Semi-Supervised and Multi-Task LearningabstractNon-intrusive speech quality assessment (NISQA) is widely used in speech downstream tasks due to its ability to predict the quality of speech without a reference speech. However, few researchers have focused on the backdoor security of NISQA. Despite the backdoor defenses have been extensively studied to mitigate the threat of maliciously modifications in deep neural networks. In particular, semi-supervised based backdoor defenses have excellent defensive performance by depriving backdoor attacks of their most essential need. But these defense methods rely on data-augmentation consistency and thus cannot be applied to NISQA. In this work, we propose a backdoor defense based on semi-supervised and multi-task learning (SML). Semi-supervised learning is based on the simple assumption that the same input should be as consistent as possible in two similar models. Multi-task learning further improves the prediction performance of mean opinion score (MOS) by learning the tasks of perceptual evaluation of speech quality (PESQ), short-time objective intelligibility (STOI) and speech distortion index (SDI). Extensive experiments involving five backdoor defenses against five backdoor attacks on two benchmark datasets demonstrate the superiority of our SML approach. Ying Ren, Jiahong Ye, Diqun Yan, Bin Ma 0003 |
ICASSP | 6 |
| 2025 | Pixel2Feature Attack (P2FA): Rethinking the Perturbed Space to Enhance Adversarial TransferabilityabstractAdversarial examples have been shown to deceive Deep Neural Networks (DNNs), raising widespread concerns about this security threat. More seriously, as different DNN models share critical features, feature-level attacks can generate transferable adversarial examples, thereby deceiving black-box models in real-world scenarios. Nevertheless, we have theoretically discovered the principle behind the limited transferability of existing feature-level attacks: Their attack effectiveness is essentially equivalent to perturbing features in one step along the direction of feature importance in the feature space, despite performing multiple perturbations in the pixel space. This finding indicates that existing feature-level attacks are inefficient in disrupting features through multiple pixel-space perturbations. To address this problem, we propose a P2FA that efficiently perturbs features multiple times. Specifically, we directly shift the perturbed space from pixel to feature space. Then, we perturb the features multiple times rather than just once in the feature space with the guidance of feature importance to enhance the efficiency of disrupting critical shared features. Finally, we invert the perturbed features to the pixels to generate more transferable adversarial examples. Numerous experimental results strongly demonstrate the superior transferability of P2FA over State-Of-The-Art (SOTA) attacks. Renpu Liu, Hao Wu 0078, Jiawei Zhang 0011, Xin Cheng 0018, Xiangyang Luo 0001, Bin Ma 0003 |
ICML | 6 |
| 2025 | Active Defense Against Deepfakes: An Integrated Framework of Adversarial Data Embedding and Blockchain Authentication
Yuli Wang, Yiyan Liang, Bin Ma 0003 |
ICONIP (3) | 3 |
| 2025 | Toward Robust Deepfake Detection: A Proactive Method Based on Watermarking and Knowledge DistillationabstractFace deepfake detection is a critical technology for verifying the authenticity of facial media content and has long been a focal point in multimedia forensics. However, existing methods face significant challenges, primarily due to their limited ability to generalize across domains. Consequently, the growing variety of forgery techniques, combined with the degradation of visual quality in forged images, makes reliable detection even more difficult. To address these challenges, we propose WKD, a proactive deepfake detection framework based on Watermarking and Knowledge Distillation. The key insights of WKD are twofold: First, we embed watermark information into the Fractional-order Quaternion Radial Harmonic Fourier Moments (FrQRHFMs) space of the host image, achieving a robust balance between imperceptibility and robustness. Second, we design a dual-task learning framework consisting of a watermark extractor and a forgery discriminator, where learnable Low-Rank Adaptation (LoRA) layers are used to transfer knowledge from the extractor to the discriminator, thereby providing additional clues for deepfake detection. Specifically, the integrity of the watermark is compromised only when the host image undergoes a deepfake forgery, while it remains unaffected by conventional attacks. Experimental results on benchmark datasets demonstrate that WKD achieves state-of-the-art performance in both intra-domain and cross-domain deepfake detection, particularly when images are subjected to various conventional attacks. Chunpeng Wang 0001, Qi Li 0029, Bin Ma 0003, Yunan Liu 0001 |
ACM Multimedia | 6 |
| 2025 | DVW: Diffusion Visible WatermarkabstractWith the rapid development of the diffusion models, numerous exquisitely generated images have significantly increased the risk of image misuse and abuse. Despite various AI parties and companies having devoted themselves to embedding watermarks into the generated images to curb the potential detriments, the isolated embedding from the generation process makes the watermarks vulnerable to watermark removal networks. To address this issue, we propose a novel generative image watermark scheme, dubbed Diffusion Visible Watermark (DVW), which can generate watermarked images in one step without additional training or fine-tuning of the diffusion models. Specifically, DVW introduces a masked distribution alignment strategy to fuse the watermark distribution with a Gaussian noise distribution. By iterative denoising the fused aligned distribution with the pretraining diffusion models, the watermarked images with coordinated and unified distribution can be generated with natural robustness against removal. In addition, we design and integrate a dynamic transparency module to adaptively control the watermark coverage degree for better visual quality. Comprehensive experiments and analysis are conducted on two representative kinds of diffusion models, GLIDE and StableDiffusion, to prove the superior and generic robustness of our DVW against watermark removal without sacrificing the generation ability of the diffusion models. Jiawei Zhang 0011, Xiaoli Jiang, Hao Wang 0060, Lin Yuan 0002, Xiangyang Luo 0001, Bin Ma 0003 |
ACM Multimedia | 6 |
| 2025 | FALU: A Proactive Deepfake Detection Scheme Based on Average Hashing and Mamba-Like Linear Attention U-NetabstractThe widespread emergence of Deepfake content has made it increasingly important to distinguish real and fake faces. Although many methods focus on detecting Deepfake content, only a few address the protection of real faces against forgery. Therefore, this paper proposes a proactive Deepfake detection scheme named FALU, which combines the uniqueness of facial identity features with the robustness of average hashing. The method first divides the input image into facial and non-facial regions, extracts identity-related features from the facial region, encodes them using average hashing, and embeds the result as a watermark into the non-facial region. During detection, the watermark is extracted from the non-facial region and compared with a newly generated hash code from the facial region. High correlation indicates authenticity, while low correlation suggests Deepfake forgery. To facilitate efficient and reliable watermark embedding, FALU integrates the symmetric sampling structure of U-Net with Mamba-like linear attention mechanism, proposing a lightweight encoder network. This scheme ensures the persistent presence of secret information before and after manipulation, thereby enhancing face source detection and tampering identification. Experimental results demonstrate that the proposed scheme effectively counters traditional Deepfake techniques and shows significant potential for preserving personal privacy. Jian Li 0034, Bin Ma 0003, Xiaolong Li 0001, Zhenxing Qian |
MMAsia | 3 |
| 2025 | LDSGAN: Unsupervised Image-to-Image Translation With Long-Domain Search GAN for Generating High-Quality Anime ImagesabstractImage‐to‐image ( I2I ) translation has emerged as a valuable tool for privacy protection in the digital age, offering effective ways to safeguard portrait rights in cyberspace. In addition, I2I translation is applied in real‐world tasks such as image synthesis, super‐resolution, virtual fitting, and virtual live streaming. Traditional I2I translation models demonstrate strong performance when handling similar datasets. However, when the domain distance between two datasets is large, translation quality may degrade significantly due to notable differences in image shape and edges. To address this issue, we propose Long‐Domain Search GAN ( LDSGAN ), an unsupervised I2I translation network that employs a GAN structure as its backbone, incorporating a novel Real‐Time Routing Search ( RTRS ) module and Sketch Loss. Specifically, RTRS aids in expanding the search space within the target domain, aligning feature projection with images closest to the optimization target. Additionally, Sketch Loss retains human visual similarity during long‐domain distance translation. Experimental results indicate that LDSGAN surpasses existing I2I translation models in both image quality and semantic similarity between input and generated images, as reflected by its mean FID and LPIPS scores of 31.509 and 0.581, respectively. Hao Wang 0060, Chenbin Wang, Xin Cheng 0018, Hao Wu 0078, Jiawei Zhang 0011, Xiangyang Luo 0001, Bin Ma 0003 |
Int. J. Intell. Syst. | 8 |
| 2025 | Privacy-Preserving IoT Image Transmission: Multistage SVD Data Embedding and Heatmap AlignmentabstractThe images transmitted by IoT devices, particularly those used for surveillance or sensor data, are vulnerable to malicious screenshots and unauthorized access, leading to potential privacy breaches. To address this, we propose a multi-stage Singular Value Decomposition (SVD)-based robust data-hiding scheme for JPEG images aimed at mitigating screenshot attacks. The method exploits the decorrelation properties of the Discrete Cosine Transform (DCT) to preprocess the carrier image, facilitating the selection of specific frequency coefficients. These coefficients undergo a dual-stage SVD transformation, where dimensionality reduction reduces the impact of noise from non-critical image regions. Additionally, we optimize Grad-CAM heatmap generation to better align with human visual perception, enabling the identification of stable and reliable feature regions for embedding secret information. This approach ensures that the visual integrity of the carrier image is maintained while preserving the legibility of the embedded information, even under attack.Our method enhances both the visual fidelity of the carrier image and the robustness of the embedded information. Experimental results demonstrate that the proposed scheme outperforms existing methods, achieving at least a 5% improvement in confidential information extraction accuracy and a data extraction rate exceeding 95% across screenshot angles ranging from -40∘ to 40∘. Extensive evaluations confirm the superior performance, efficiency, and security of our method in mitigating screenshot attacks, showcasing its broad applicability to various image formats and resilience to distortions. Kaixin Du, Bin Ma 0003, Meihong Yang, Xiaoyu Wang 0011, Xiaolong Li 0001 |
IEEE Internet Things J. | 2 |
| 2025 | An IoT-Oriented Image Retrieval Scheme Based on Multifeature Fusion for Cloud-Edge EnvironmentsabstractWith the rapid advancement of the Internet of Things (IoT), massive volumes of multimedia data are continuously generated by distributed sensing devices and edge nodes. Efficient and accurate image retrieval from such data has become a key component in enabling advanced IoT applications. However, the constraints of edge computing—including limited bandwidth, low power budgets, and heterogeneous hardware—pose significant challenges to conventional image retrieval schemes. To address these issues, this paper proposes a lightweight and effective Content-Based Image Retrieval (CBIR) framework optimized for cloud-enabled IoT environments. Specifically, this paper introduces a new multi-feature construction scheme that integrates the Color Granular Descriptor (CGD) for fine-grained color characterization, the Double-Radius Local Binary Pattern (DR-LBP) for enhanced local texture extraction, and the Lower-Order Polar Harmonic Fourier Moments (LPHFMs) for capturing global shape features with strong rotational and scale invariance. The proposed scheme achieves high retrieval precision with low computational cost, making it well-suited for deployment in resource-constrained IoT environments. Extensive evaluations conducted on widely used benchmark datasets—including Corel-1K, Corel-5K, Corel-10K, Oxford105K, and GHIM-10K—demonstrate the superior performance and robustness of the proposed method, validating its practical applicability to IoT scenarios. Zhongquan Tao, Bin Ma 0003, Jian Xu 0025, Xiaolong Li 0001 |
IEEE Internet Things J. | 2 |
| 2025 | A High-Performance Region Recognition Network-Enhanced Deep CNN for Image Content Perceptual HashingabstractPerceptual image hashing has emerged as a crucial forensic tool within the Internet of Things (IoT) ecosystem. Traditional perceptual hashing algorithms predominantly rely on global image features to generate hash codes, which limit their ability to represent key features of images effectively. This paper introduces a Perceptual Region Recognition Network (PRRN) to accurately identify key feature regions in images based on their texture distribution characteristics, thereby generating image perceptual hashing codes that reflect the key content of the images. At the same time, a perceptual hashing feature extraction module, which integrates a Residual Network (ResNet) and a Weighted Feature Fusion Network (WFFN), is built to extract deep semantic features of the object image. Where, ResNet is leveraged to extract high-level semantic features, while WFFN ensures the preservation of low-level local features. Furthermore, skip connections are employed to achieve content enhancements for intricate details of critical image regions. Additionally, the Mean Squared Error (MSE) loss is incorporated to enhance the accuracy of key region localization, further improving the sensitivity of image perceptual hash codes and accelerating the network’s convergence speed. Extensive experimental evaluations demonstrate that the proposed PRRN-based perceptual image hashing scheme significantly outperforms other state-of-the-art methods in terms of image feature representation capability. Specifically, it achieves an average improvement of over 1.2 in attack-resistant capability for images compared with other counterparts, making it a promising candidate for practical applications in the IoT environment. Meihong Yang, Baolin Qi, Bin Ma 0003, Jian Xu 0025, Yongjin Xian, Xiaolong Li 0001 |
IEEE Internet Things J. | 3 |
| 2025 | A GAN-based anti-forensics method by modifying the quantization table in JPEG header file
Hao Wang 0060, Xin Cheng 0018, Hao Wu 0078, Xiangyang Luo 0001, Bin Ma 0003, Hui Zong, Jiawei Zhang 0011 |
J. Vis. Commun. Image Represent. | 5 |
| 2025 | Highly applicable and imperceptible watermark attack network
Chunpeng Wang 0001, Qi Li 0029, Jian Li 0034, Ziqi Wei 0001, Ting Luo 0001, Bin Ma 0003 |
Signal Process. | 8 |
| 2025 | One-class network leveraging spectro-temporal features for generalized synthetic speech detection
Jiahong Ye, Diqun Yan, Songyin Fu, Bin Ma 0003, Zhihua Xia |
Speech Commun. | 4 |
| 2025 | High Precision CNN Predictor of Color Images for Reversible Data Hiding
Hongtao Duan 0005, Zhongquan Tao, Bin Ma 0003, Jian Xu 0025, Yongjin Xian |
IEEE Signal Process. Lett. | 3 |
| 2025 | Cross-Domain Robust Image Steganography via Dual-Domain Enhancement NetworkabstractCurrent steganographic techniques predominantly focus on single-domain security designs, while neglecting the fact that cross-domain conversions between spatial and frequency domains may compromise embedded features, introducing detectable noise and artifacts that render stego images vulnerable to steganalyzers. This letter proposes a high-performance image steganography by dual-domain adversarial training to enhance both the security and image quality in the spatial and JPEG domains. The proposed method employs a dual-domain adversarial training strategy, integrating spatial and JPEG-domain steganalyzers to guide the generator toward producing compression-resilient stego images. In addition, a dual-objective loss function is introduced, consisting of a spatial fidelity loss to ensure visual imperceptibility and a frequency-domain consistency loss to mitigate compression-induced distortions. This design enables the model to effectively learn domain-aware embedding strategies, thereby achieving enhanced cross-domain robustness and security. Extensive experiments demonstrate that the proposed method outperforms other advanced image steganographic methods in terms of security and robustness. Kun Li 0010, Bin Ma 0003, Weike You, Linna Zhou |
IEEE Signal Process. Lett. | 2 |
| 2025 | High-Performance Optimization Framework for Reversible Data Hiding PredictorabstractExisting deep learning-based reversible data hiding (RDH) predictors are affected by the difference of pixel complexity, which leads to the reduction of prediction accuracy. Therefore, this letter proposes an optimization framework tailored for RDH predictors, which integrates the local complexity of pixels into the predictor's regression optimization process. By analyzing the image's texture features, the framework adaptively determines the optimal prediction coefficients, thereby improving prediction accuracy. Notably, this optimization framework is versatile and can be applied to optimize other deep learning-based RDH predictors. Additionally, recognizing the critical role of interpolation strategies in RDH pixel prediction, we introduce a multi-scale fusion-enhanced interpolation network specifically designed for RDH, which integrates features across different scales to provide accurate reference pixels for subsequent predictions. Finally, experimental results demonstrate that the proposed method outperforms several advanced RDH predictors in terms of both prediction accuracy and embedding performance. Bin Ma 0003, Hongtao Duan 0005, Ruihe Ma, Yongjin Xian, Xiaolong Li 0001 |
IEEE Signal Process. Lett. | 1 |
| 2025 | Cross-Domain Deepfake Detection Based on Latent Domain Knowledge DistillationabstractThe rapid development of deepfake technology poses challenges to face-centered data security. Existing methods primarily focus on how to transfer deepfake detectors from the source domain to the target domain to handle diverse deepfake techniques. In practical application scenarios, it is usually difficult to access the true and false labels of the source domain. In this letter, we introduce a new adaptation framework called Latent Domain Knowledge Distillation (LDKD) for cross-domain deepfake detection. In the proposed framework, we construct a knowledge distillation structure that includes a student network and a teacher network, which are jointly optimized in a coupled manner to facilitate the model's adaptation to the target domain. Furthermore, to improve the quality of pseudo-labels generated by the teacher network, we propose a Fourier Latent Domain Generation Module (FLGM) and a Stochastic Complementary Mask Module (SCMM). The former is used to generate latent domains to bridge domain differences at the image level, while the latter is employed to mine richer contextual cues for the model. Extensive cross-domain experimental results demonstrate that our method achieves state-of-the-art performance, and the model analysis proves the effectiveness of our key components. Chunpeng Wang 0001, Lingshan Meng, Na Ren, Bin Ma 0003 |
IEEE Signal Process. Lett. | 5 |
| 2025 | HashShield: A Robust DeepFake Forensic Framework With Separable Perceptual HashingabstractThe proliferation of DeepFakes has heightened the necessity to distinguish between authentic and counterfeit faces. While numerous methods concentrate on detecting DeepFakes, only a few address safeguarding genuine faces from manipulation. This letter proposes a novel active forensics system for DeepFake forensics utilizing separable perceptual hash enhancement algorithm. A separable perceptual hash code specifically designed for face deep forgery is introduced, achieving robustness while maintaining sensitivity and imperceptibility when embedded within the original image. Additionally, a multi-scale perceptual smoothing loss function is employed to optimize perceptual similarity, structural smoothness, and embedding stability. As a result, this system ensures the consistence of confidential information both before and after manipulation, thereby enhancing the capability of face source detection and DeepFake identification. Experimental results demonstrate that the proposed scheme can effectively counter traditional deep forgery techniques while exhibiting significant potential in preserving personal privacy. Meihong Yang, Baolin Qi, Ruihe Ma, Yongjin Xian, Bin Ma 0003 |
IEEE Signal Process. Lett. | 5 |
| 2025 | Paradoxical Role of Adversarial Attacks: Enabling Crosslinguistic Attacks and Information Hiding in Multilingual Speech RecognitionabstractWith the rise of automatic speech recognition (ASR) research and practical applications, enabling adversarial attacks on ASR systems via subtle perturbations has become a priority. Most prior research has focused on single-language, single-model ASR systems. However, multilingual ASR systems hold opportunities for crosslinguistic attacks and covert message transmission. This letter introduces a new approach for crosslinguistic adversarial attacks in multilingual ASR, focusing on information hiding. For example, in military settings, adversarial examples applied to eavesdropping devices can encode messages detectable only by friendly devices, leaving adversaries, even with identical methods, unable to access them. This letter examines multilingual ASR system properties and introduces a crosslinguistic adversarial example with minimal perturbation, allowing friendly classifiers to extract hidden information while being undetectable by hostile classifiers. The experimental results on 5 models and 5 datasets show that the proposed method achieves a success rate of over 90% and an SNR close to 40 dB. Zhihua Xia, Bin Ma 0003, Diqun Yan |
IEEE Signal Process. Lett. | 3 |
| 2025 | Robust Image Steganography via Color ConversionabstractIn this paper, we propose a robust image steganography method utilizing color conversion, leveraging de-colorization and colorization models to achieve covert transmission of secret information. The motivation is to use color conversion of the stego image to conceal steganographic behavior. For the sender, secret information is embedded into the color cover image using a robust embedding algorithm based on quaternion exponent moments. The stego images are then de-colorized to obtain grayscale images, which can be transmitted over public channels. For the receiver, a corresponding colorization network is designed to reconstruct the stego image and extract the secret information. Additionally, an attack module using Gaussian noise is implemented to enhance the robustness of the proposed steganography. Given a color image, its grayscale version can be chosen from various options, making it difficult for attackers to detect steganographic activity as long as the generated grayscale image appears normal and meaningful. Extensive simulation results demonstrate the feasibility and scalability of the proposed steganography method. Qi Li 0029, Bin Ma 0003, Xianping Fu, Xiaoyu Wang 0011, Chunpeng Wang 0001, Xiaolong Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Color Image High-Capacity Differential Steganography Algorithm Based on Multiple Adversarial NetworksabstractAiming to mitigate image distortion caused by steganography algorithms at high-capacity information embedding and enhance the steganalysis resistance capability of generated stego images, this paper proposes a high-capacity differential steganography algorithm for color images based on multiple adversarial networks. Instead of directly modifying the pixels of the cover image, the algorithm embeds the secret information into the differential plane generated by the two most similar channels of the cover image. Consequently, the distortion of the stego image is minimized while embedding a secret image of the same size. At the same time, the fidelity of the stego and extracted secret images is continually improved through adversarial training between the generator and discriminator in the proposed steganography network. Furthermore, multiple steganalysis networks are parallelly utilized to enhance the steganalysis resistance capability of stego images. In addition, the Lion optimizer is utilized for the first time to improve the convergence speed of the proposed steganographic network. Experimental results show that the comprehensive performance of the proposed algorithm outperforms other state-of-the-art steganography algorithms significantly. Bin Ma 0003, Jian Xu 0025, Xiaoyu Wang 0011, Xiaolong Li 0001, Jian Li 0034 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Light-Field Image Multiple Reversible Robust Watermarking Against Geometric AttacksabstractLight-field (LF) images contain rich visual information and have broader application scenarios than traditional images. However, their complex structure also makes their copyright protection more challenging. Currently, there are few watermarking schemes suitable for LF images, and most of them fail to restore the original image after embedding the watermark. In addition, geometric attacks remain a difficult problem in the field of LF image watermarking. In this study, we propose a multiple reversible robust LF image watermarking scheme based on code division multiplexing (CDM) and quaternion polar harmonic Fourier moments (QPHFMs). This scheme embeds multiple identical watermarks into the LF macro-pixel image, and the compensation information for information loss caused by watermark embedding is reversibly embedded into the LF sub-aperture images. The watermark can be extracted and the original LF image can be fully recovered if the image has not been attacked. The watermark can be extracted to verify the copyright ownership of the LF image even when the image has been attacked. Experimental results demonstrate that the proposed watermarking scheme is resistant to various attacks and exhibits strong robustness. Chunpeng Wang 0001, Xiaoyu Wang 0011, Linna Zhou, Qi Li 0029, Bin Ma 0003, Yun Q. Shi 0001 |
IEEE Trans. Dependable Secur. Comput. | 7 |
| 2025 | Boosting Transferability of Adversarial Examples with Spatio-Temporal ContextabstractTransferable adversarial examples have received increasing attention for their utility in spoofing multiple models, but existing attacks still perform poorly in terms of transferability. In light of this, a novel attack method called Spatio-Temporal Context-Based Enhanced Momentum Iteration (STCEMI) is proposed for transferability enhancement. First, two spatially and temporally oriented context exploitation strategies are devised, respectively. On the one hand, the blended image is obtained by summing a randomly scrambled version of the original image with itself, and the correction of spatial context momentum to the gradient of the current position is achieved by utilizing the blended image to optimize the perturbation. On the other hand, with the short-time context obtained from single-step iteration along the backward and forward gradient directions, the gradient of the current iteration can be corrected by the temporal context momentum. Second, considering the complementarity of spatial and temporal contexts, two strategies are naturally integrated to construct the spatio-temporal context-based attack, STCEMI, with the objective of achieving stronger transferability. The results of extensive experiments demonstrate that the adversarial images generated by STCEMI achieve the highest cross-model attack success rate across multiple mainstream normally trained and adversarially trained models. Jingtian Wang, Xiaolong Li 0001, Bin Ma 0003, Yao Zhao 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2025 | Invisible Adversarial Watermarking: A Novel Security Mechanism for Enhancing Copyright ProtectionabstractInvisible watermarking can be used as an important tool for copyright certification in the Metaverse. However, with the advent of deep learning, Deep Neural Networks (DNNs) have posed new threats to this technique. For example, artificially trained DNNs can perform unauthorized content analysis and achieve illegal access to protected images. Furthermore, some specially crafted DNNs may even erase invisible watermarks embedded within the protected images, which eventually leads to the collapse of this protection and certification mechanism. To address these issues, inspired by the adversarial attack, we introduce Invisible Adversarial Watermarking (IAW), a novel security mechanism to enhance the copyright protection efficacy of watermarks. Specifically, we design an Adversarial Watermarking Fusion Model (AWFM) to efficiently generate Invisible Adversarial Watermark Images (IAWIs). By modeling the embedding of watermarks and adversarial perturbations as a unified task, the generated IAWIs can effectively defend against unauthorized identification, access, and erase via DNNs and identify the ownership by extracting the embedded watermark. Experimental results show that the proposed IAW presents superior extraction accuracy, attack ability, and robustness on different DNNs, and the protected images maintain good visual quality, which ensures its effectiveness as an image protection mechanism. Jiawei Zhang 0011, Hao Wu 0078, Xiangyang Luo 0001, Bin Ma 0003 |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2024 | A Pixel Distribution Complexity Classification Enhanced Convolutional Neural Network Predictor for Reversible Data Hiding
Bin Ma 0003, Hongtao Duan 0005, Ruihe Ma, Chunpeng Wang 0001, Xiaolong Li 0001 |
ICIC (9) | 1 |
| 2024 | LCRPS: Large-Capacity Residual Plane Steganography Based on Multiple Adversarial Networks
Bin Ma 0003, Ruihe Ma, Yongjin Xian, Chunpeng Wang 0001 |
ICONIP (7) | 1 |
| 2024 | Robust Video Watermarking Network Based on Channel Spatial AttentionabstractRobust video watermarking refers to the ability to extract the originally embedded watermark information from a video even after malicious modifications and attacks. Currently, traditional watermarking methods have the drawback of lacking robustness against multiple watermark attacks simultaneously. Neural network-based approaches have not fully considered the multi-scale features of videos and tend to lose information during the fusion of scale features. Therefore, we propose a video watermarking scheme based on Channel Spatial Attention. Our model can extract feature information at different scales, allowing the watermark to adapt to features of different scales in the video. Through a series of comparative experiments, our method has shown significant improvements over traditional video watermarking methods and deep learning-based video watermarking models. Jian Li 0034, Bin Ma 0003, Chunpeng Wang 0001, Huanhuan Zhao, Zhengzhong Zhao |
IJCNN | 3 |
| 2024 | Advancing Quantization Steps Estimation: A Two-Stream Network Approach for Enhancing Robustness
Xin Cheng 0018, Hao Wang 0060, Xiangyang Luo 0001, Bin Ma 0003 |
ACM Multimedia | 5 |
| 2024 | Dual-Task Cascaded for Proactive Deepfake Detection Using QPCET Watermarking
Chunpeng Wang 0001, Chaoyi Shi, Yunan Liu 0001, Jian Li 0034, Yongjin Xian, Bin Ma 0003 |
PRCV (2) | 7 |
| 2024 | A Reversible Data Hiding in Encryption Domain for JPEG Image Based on Controllable Ciphertext Range of Paillier Homomorphic Encryption Algorithm
Bin Ma 0003, Chunxin Zhao, Ruihe Ma, Yongjin Xian, Chunpeng Wang 0001 |
PRICAI (3) | 1 |
| 2024 | PWPH: Proactive Deepfake Detection Method Based on Watermarking and Perceptual HashingabstractThe popularity of Deepfake technology has raised the challenge of recognizing real and fake faces. While detection methods already exist, most of them are passive forensics and face challenges of generalizability and migration. Currently, some research attempts to protect the original image by priorly inserting invisible information. However, there are still shortcomings in terms of image quality and information robustness due to information embedding, i.e., watermarking. Therefore, we employ the robustness of perceptual hash coding and combine it with information hiding techniques to propose a proactive Deepfake detection solution, referred to as PWPH in this paper. Our approach is simple and efficient: first, the image containing a face is divided into two parts: FA (face area), and NFA (non-face area). A perceptual hash code is generated from the non-face area (NFA). Then, the hash codes are embedded as watermarks into the FA. At the extraction stage, we use the same method as the encoder to retrieve the embedded watermark from FA. The watermark is then compared with the hash code generated from the NFA of the detected image. The extracted watermark is sensitive to distortion and may vanish during Deepfake processing. Experimental results validate that our method, requiring just one encoder and decoder, enables active detection and source tracking. Furthermore, its efficacy in typical Deepfake scenarios such as face swapping and expression reconstruction is confirmed through comparison with prior arts. Jian Li 0034, Shuanshuan Li, Bin Ma 0003, Chunpeng Wang 0001, Linna Zhou, Yule Wang |
SMC | 3 |
| 2024 | HIWANet: A high imperceptibility watermarking attack network
Chunpeng Wang 0001, Qi Li 0029, Hao Zhang 0061, Jian Li 0034, Bin Ma 0003 |
Eng. Appl. Artif. Intell. | 8 |
| 2024 | Improving the transferability of adversarial examples through black-box feature attacks
Maoyuan Wang, Bin Ma 0003, Xiangyang Luo 0001 |
Neurocomputing | 3 |
| 2024 | Adversarial watermark: A robust and reliable watermark against removal
Wanyun Huang, Jiawei Zhang 0011, Xiangyang Luo 0001, Bin Ma 0003 |
J. Inf. Secur. Appl. | 5 |
| 2024 | Trustworthy adaptive adversarial perturbations in social networks
Jiawei Zhang 0011, Hao Wang 0060, Xiangyang Luo 0001, Bin Ma 0003 |
J. Inf. Secur. Appl. | 5 |
| 2024 | Reversible data hiding algorithm based on adaptive prediction and code division multiplexing
Xiaoyu Wang 0011, Xingyuan Wang 0001, Bin Ma 0003, Qi Li 0029, Chunpeng Wang 0001, Yongjin Xian |
Multim. Tools Appl. | 3 |
| 2024 | A High-Performance Image Steganography Scheme Based on Dual-Adversarial NetworksabstractThis letter proposes a high-performance image steganography scheme based on dual-adversarial networks to enhance the performance of secret message hiding. According to the characteristics of generative adversarial networks, a dual-adversarial steganography scheme is devised to improve both the visual quality and the steganalysis resistance capability of the stego image. In the first adversarial block, the U-net structure is employed to reconstruct the original image as the generated image, and the adversarial noise is imperceptibly embedded into the generated image to produce the adversarial image that is most suitable for data hiding. In the second adversarial block, secret messages are undetectably embedded into the adversarial image under the confrontation of multiple steganalysis networks. Moreover, a multiple-channel attention module is introduced to enhance the performance of the adversarial image and accelerate the convergence speed of the proposed dual-adversarial networks. Additionally, the MSE loss is employed to minimize the divergence between the original and the adversarial image. Experimental results indicate that the average PSNR of the adversarial images reach 41 dB, and the detection probability is 2.79% lower than that of other advanced schemes. The proposed scheme outperforms its counterparts in terms of performance. Bin Ma 0003, Kun Li 0010, Jian Xu 0025, Chunpeng Wang 0001, Xiaolong Li 0001 |
IEEE Signal Process. Lett. | 1 |
| 2024 | Dual-Task Mutual Learning With QPHFM Watermarking for Deepfake DetectionabstractDeepfake technology has rapidly evolved and emerged in recent years, posing significant threats to individuals' reputations and security. Although passive detection methods can achieve reasonable accuracy, they still lack proactive defense mechanisms. To address this issue, this letter proposes a proactive detection framework that combines Quaternion Polar Harmonic Fourier Moments (QPHFMs) with Dual-Task Mutual Learning (DTML) framework. Firstly, watermark information is embedded into QPHFMs, ensuring high imperceptibility while enhancing robustness against common attacks. Secondly, DTML is introduced, where the knowledge distilled from watermark detection can facilitate more accurate deepfake detection. Experimental results on benchmark datasets demonstrate that our method surpasses state-of-the-art techniques, delivering exceptional performance in watermark robustness and imperceptibility while simultaneously accomplishing accurate deepfake detection. Chunpeng Wang 0001, Chaoyi Shi, Simiao Wang, Bin Ma 0003 |
IEEE Signal Process. Lett. | 5 |
| 2024 | Quantization Step Estimation of Color Images Based on Res2Net-C With Frequency Clustering Prior KnowledgeabstractThe quantization step is a crucial parameter in JPEG compression, that can reveal the compression history of a JPEG image. Estimating the quantization steps for single compressed and recompressed images is attracting considerable interest in the field of image forensics and steganalysis. Several effective methods have been proposed, but the performance of these methods still needs to be improved on small-sized and low-quality images. To solve the above problems, feature enrichment is performed on images in the frequency domain, resulting in clustering discrete cosine transform (DCT) coefficients of the same frequency. Then, we construct a hierarchical connection within the residual blocks of the network to represent multi-scale features, enabling the network to learn deep features of the image. At the same time, we use multiple small-sized convolution kernels instead of one large-sized convolution kernel to minimize the impact of block artifacts. Based on the above two ideas, we construct a network model, Res2Net-C, to discover information about the quantization steps in the frequency domain. The integration of multi-channel information of color images is achieved by multi-channel convolution, and the quantization steps of the chrominance and luminance channels of the color images are estimated. The experimental results show that the accuracy of the proposed method for estimating the quantization steps is 29.97% better than that of the existing algorithm with a single compressed dataset and 4.87% better than that of the existing algorithm with a recompressed image dataset. In addition, the method has good performance with mixed datasets that contain both single compressed and recompressed images. Xin Cheng 0018, Hao Wang 0060, Xiangyang Luo 0001, Bin Ma 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | A High-Performance Robust Reversible Data Hiding Algorithm Based on Polar Harmonic Fourier MomentsabstractAiming at the problem of most Robust Reversible Data Hiding (RRDH) schemes failing to anti geometric deformation attacks, a new RRDH algorithm based on Polar Harmonic Fourier Moments (PHFMs) is presented in this paper, thereby enhancing both the robustness of the embedded data and perceptual quality of the data-embedded image. Firstly, by leveraging the anti-geometric transformation and high-fidelity features of PHFMs, the image is transformed into its frequency domain for RRDH. Then, a quantitation index modulation (QIM) algorithm is designed to embed secret data into the integer part of PHFMs coefficients. By minimizing the differences between the secret-data-embedded image and the original image, the amount of compensation data is reduced. Meanwhile, a two-dimensional RDH scheme is further adopted to embed the compensation data, thus reducing the distortion of the full data-embedded image. Finally, the robustness of the embedded data and the fidelity of the full data-embedded image are both improved. The combination of PHFMs transformation and two-dimensional RDH enables the proposed RRDH algorithm to achieve high visual quality and strong resistance capability against geometric transformation attacks. Extensive experimental results demonstrate that the proposed RRDH algorithm outperforms other state-of-the-art techniques. Bin Ma 0003, Zhongquan Tao, Ruihe Ma, Chunpeng Wang 0001, Jian Li 0034, Xiaolong Li 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | General Forensics for Aligned Double JPEG Compression Based on the Quantization InterferenceabstractDetection of aligned double Joint Photographic Experts Group (JPEG) compressed images is a crucial area of research within the field of digital image forensics. The detection tasks for aligned double JPEG compression can be categorized into two sub-tasks, namely detecting double JPEG images with the same quantization matrix (DJSQM) or double JPEG images with different quantization matrices (DJDQM). Existing methods for one of these sub-tasks may not be effective for the other. To address this issue, a novel approach is proposed by recompressing both DJDQM and DJSQM using modified quantization coefficients. The perturbation in the recompression process results in a perturbed error image, which is valid for both DJDQM and DJSQM. Subsequently, the relative change rate is used to combine the perturbed error image, the original error image, and the quantization error to derive the interference error and the interference quantization error. The interference error and interference quantization error further expand the difference between single and double compressed images by preserving the general validity of the original image information. Furthermore, the recompression process of DJDQM and DJSQM results in the conversion of truncation and rounding errors at the pixel level, which can be represented by the pixel state map. The pixel state map characterizes the differing transformation relationships between single and double compressed images and provides additional valid features, thereby enhancing the performance of the proposed method. The empirical results demonstrate that the proposed method outperforms existing methods on detecting aligned double JPEG compressed images. Hao Wang 0060, Jiawei Zhang 0011, Xiangyang Luo 0001, Bin Ma 0003, Bin Li 0011, Jinsheng Sun |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Improving the Transferability of Adversarial Attacks through Experienced Precise Nesterov MomentumabstractDeep neural networks are vulnerable to adversarial examples. Although the adversarial example has superior white-box attack success rate, its transferability is poor under the black-box setting. Momentum is often integrated into attacks so as to prevent adversarial examples from overfitting the source model and improve the transferability of adversarial examples. How-ever, conventional momentum merely accumulates few gradients during the early iterations, resulting in the early adversarial examples already overfitting the source model. Therefore, we propose Experienced Momentum (EM), which is trained on a set of models derived by Random Channels Swapping (RCS). Since EM takes the direction of loss increasing for multiple models into account, assigning EM to the initial value of momentum to makes adversarial examples transferable across models during the early iterations. Moreover, conventional Nesterov momentum only take the previous gradients into consideration but ignore the gradient of the current data point during the whole pre-update, making the estimate of the next position imprecise. It prompts us to propose Precise Nesterov momentum (PN), which not only retains the looking-ahead property but also adopts the gradient of the current data point during the pre-update. To further improve transferability, we combine EM and PN as Experienced Precise Nesterov momentum (EPN). Extensive experiments on the ImageNet dataset against normally trained and defense models demonstrate that the proposed EPN is more effective than conventional momentum for improving transferability. Hao Wu 0078, Jiawei Zhang 0011, Bin Ma 0003, Xiangyang Luo 0001 |
IJCNN | 5 |
| 2023 | High-Quality PRNU Anonymous Algorithm for JPEG Images
Jian Li 0034, Huanhuan Zhao, Bin Ma 0003, Chunpeng Wang 0001, Zhengzhong Zhao |
IWDW | 3 |
| 2023 | Cross-channel Image Steganography Based on Generative Adversarial Network
Bin Ma 0003, Yongjin Xian, Chunpeng Wang 0001, Guanxu Zhao |
IWDW | 1 |
| 2023 | CAWNet: A Channel Attention Watermarking Attack Network Based on CWABlock
Chunpeng Wang 0001, Ziqi Wei 0001, Qi Li 0029, Bin Ma 0003 |
PRCV (9) | 6 |
| 2023 | Improving Transferability of Adversarial Attacks with Gaussian Gradient Enhance Momentum
Maoyuan Wang, Hao Wu 0078, Bin Ma 0003, Xiangyang Luo 0001 |
PRCV (9) | 4 |
| 2023 | FPA-WAN: Feature Pyramid Attention Based Watermarking Attack NetworkabstractDigital watermarking technology is a method of embedding specific information in digital images and videos, often used for copyright protection and identity verification. However, unscrupulous users may also use this technique to falsify and tamper with data. Therefore, a reliable watermark attack network is needed to detect and remove watermarks embedded by unscrupulous users. In this paper, we propose a watermarking attack network based on feature pyramid attention, which can effectively remove watermarks embedded in digital images and greatly guarantee the image quality of carrier images. The network consists of two main modules: the feature extraction module and the watermark attack module. In the feature extraction module, we use a convolutional neural network and a residual block to extract the features of the image. Then, in the watermarking attack module, we use the pyramid attention mechanism to focus on the important regions in the feature map and apply the attention weights to the watermarking attack operation. To validate the effectiveness of this network, we conducted experiments using a variety of standard data sets. Experimental results show that the network can effectively attack the watermark information embedded in digital images while guaranteeing the quality of the images after the attack. Overall, the watermarking attack network based on feature pyramid attention proposed in this paper is an effective attack with high imperceptibility that can be applied in the field of digital media protection and security in practical scenarios. Chunpeng Wang 0001, Qi Li 0029, Ziqi Wei 0001, Bin Ma 0003 |
SMC | 6 |
| 2023 | Reversible PRNU anonymity for device privacy protection based on data hiding
Jian Li 0034, Bin Ma 0003, Chuan Qin 0001, Chunpeng Wang 0001 |
Expert Syst. Appl. | 3 |
| 2023 | Multi-dimensional hypercomplex continuous orthogonal moments for light-field images
Chunpeng Wang 0001, Linna Zhou, Ziqi Wei 0001, Hao Zhang 0061, Bin Ma 0003 |
Expert Syst. Appl. | 7 |
| 2023 | PRNU Anonymous Algorithm Used for Privacy Protection in Biometric Authentication SystemsabstractThe photo response non-uniformity (PRNU) is used to connect an image to its source sensor. In this paper, researchers propose a PRNU anonymity method based on image segmentation to cut the relationship between the image and its source camera. According to the distribution rule of PRNU in the high and low frequency band of the image, the high and low frequency information of the part is also processed differently, which ensures the quality of the output image to a large extent. Experiments on the datasets show that the proposed method can preserve the biometric characteristics of the device while maintaining the anonymity of the device. Comparing with prior art, peak signal to noise ratio (PSNR) and cosine similarity are improved by 1.9 dB and 0.02 points, respectively. Jian Li 0034, Bin Ma 0003, Meihong Yang, Chunpeng Wang 0001, Xinan Cui |
Int. J. Semantic Web Inf. Syst. | 3 |
| 2023 | A screen-shooting resilient data-hiding algorithm based on two-level singular value decomposition
Bin Ma 0003, Kaixin Du, Jian Xu 0025, Chunpeng Wang 0001, Jian Li 0034, Linna Zhou |
J. Inf. Secur. Appl. | 1 |
| 2023 | Wavelet-FCWAN: Fast and Covert Watermarking Attack Network in Wavelet Domain
Chunpeng Wang 0001, Fanran Sun, Qi Li 0029, Jian Li 0034, Bin Ma 0003 |
J. Vis. Commun. Image Represent. | 7 |
| 2023 | Reversible data hiding based on prediction-error value ordering and multiple-embedding
Wenfa Qi, Tong Zhang 0024, Xiaolong Li 0001, Bin Ma 0003, Zongming Guo |
Signal Process. | 4 |
| 2023 | High-performance reversible data hiding based on ridge regression prediction algorithm
Xiaoyu Wang 0011, Xingyuan Wang 0001, Bin Ma 0003, Qi Li 0029, Chunpeng Wang 0001, Yun Q. Shi 0001 |
Signal Process. | 3 |
| 2023 | Sedenion polar harmonic Fourier moments and their application in multi-view color image watermarking
Chunpeng Wang 0001, Bin Ma 0003, Jian Li 0034, Hao Zhang 0061, Qi Li 0029 |
Signal Process. | 3 |
| 2023 | A Novel Reversible Data Hiding Scheme Based on Pixel-Residual HistogramabstractPrediction-error expansion (PEE) is the most popular reversible data hiding (RDH) technique due to its efficient capacity-distortion tradeoff. With the generated prediction-error histogram (PEH) and adaptively selected expansion bins, the image redundancy is well exploited by PEE. However, for the most widely used rhombus predictor, the rounding operation which groups different prediction-errors into one value is completely unnecessary. The embedding can be extended to a general case by removing the rounding operation, and more histogram bins can be derived for expansion with a new mapping mechanism. Therefore, in this article, instead of pixel prediction-error, we propose to compute the pixel residuals without the rounding operation, and a new embedding mechanism based on pixel-residual histogram (PRH) modification is devised. In PRH, four bins correspond to one bin in PEH. Then, different from the one-to-one mapping between the prediction-error and pixel modification, a four-to-one mapping between the pixel-residual and pixel modification is established, and the performance is optimized by adaptively selecting four expansion bin pairs for embedding. Since more modification selections are considered, better performance can be obtained. Moreover, the proposed scheme is extended to the two-dimensional (2D) histogram and multiple histograms based embedding, and the performance is further enhanced. The superiority of the proposed method is experimentally verified by comparing it with some state-of-the-art works. Mengyao Xiao, Xiaolong Li 0001, Yao Zhao 0001, Bin Ma 0003, Guodong Guo |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2022 | Improving the Transferability of Adversarial Attacks Through Both Front and Rear Vector Method
Hao Wu 0078, Jiawei Zhang 0011, Xiangyang Luo 0001, Bin Ma 0003 |
IWDW | 5 |
| 2022 | A robust zero-watermarking algorithm for lossless copyright protection of medical images
Xingyuan Wang 0001, Chunpeng Wang 0001, Changxu Wang, Bin Ma 0003, Qi Li 0029 |
Appl. Intell. | 5 |
| 2022 | Light-field image watermarking based on geranion polar harmonic Fourier moments
Chunpeng Wang 0001, Bin Ma 0003, Jian Li 0034, Ting Luo 0001, Qi Li 0029 |
Eng. Appl. Artif. Intell. | 3 |
| 2022 | A high-performance insulators location scheme based on YOLOv4 deep learning network with GDIoU loss functionabstractAbstract This paper proposes a Gaussian Distance Intersection over Union (GDIoU) loss function‐based YOLOv4 deep learning network to solve the problem of slow speed and low accuracy insulator location in power facilities health inspection. In the scheme, A GDIoU loss function is designed to accelerate the convergence speed of the YOLOv4 deep learning network; at the same time, the GDIoU loss is added as one part of the network propagation loss, and the insulator's location accuracy is accordingly improved. Moreover, a re‐location scheme for tilt insulators correction is proposed to enhance the location accuracy of the insulators in different spatial angle states. Large amounts of field insulator images were gathered as training and testing samples to evaluate the performance of the proposed scheme. The experimental results have demonstrated that the GDIoU‐based YOLOv4 deep learning network combined with the tilt correction scheme can improve the insulator location speed by three times compared with the peer schemes, and the average precision is increased by 7.37% compared with the naive YOLOv4 network. The performance of the proposed scheme meets the requirement of online insulator location adequately. Bin Ma 0003, Yongkang Fu, Chunpeng Wang 0001, Jian Li 0034, Yuli Wang |
IET Image Process. | 1 |
| 2022 | ESGAN for generating high quality enhanced samples
Xiangyang Luo 0001, Bin Ma 0003 |
Multim. Syst. | 5 |
| 2022 | GAN-generated fake face detection via two-stream CNN with PRNU in the wild
Kehui Zeng, Bin Ma 0003, Xiangyang Luo 0001, Qilin Yin, Guangjie Liu 0001, Sunil Kr. Jha |
Multim. Tools Appl. | 3 |
| 2022 | Fast Expansion-Bins-Determination for Multiple Histograms Modification Based Reversible Data HidingabstractReversible data hiding (RDH) is a research hotspot nowadays. By RDH, after data extraction, the cover image can be restored without information loss. Among numerous existing RDH techniques, multiple histograms modification (MHM) is a general reversible embedding framework, and it is experimentally verified better than the traditional single histogram based methods. However, the expansion-bins-determination process for MHM is conducted through naive exhaustive search, which is time consuming. Based on this consideration, a fast expansion-bins-determination method for MHM is proposed in this paper. Specifically, to determine the optimal expansion bins, instead of solving the optimization problem of discrete variables, we consider a general form of this problem with differentiable objective function and real variables, so that advanced analysis tools such as Lagrange multiplier can be utilized. By the proposed approach, compared with the original MHM, the expansion bins can be determined quickly with only a tiny performance loss, and thus the practicality of MHM is improved. Shi-Mei Ma, Xiaolong Li 0001, Mengyao Xiao, Bin Ma 0003, Yao Zhao 0001 |
IEEE Signal Process. Lett. | 4 |
| 2022 | Concealed Attack for Robust Watermarking Based on Generative Model and Perceptual LossabstractWhile existing watermarking attack methods can disturb the correct extraction of watermark information, the visual quality of watermarked images will be greatly damaged. Therefore, a concealed attack based on generative adversarial network and perceptual losses for robust watermarking is proposed. First, the watermarked image is utilized as the input of generative networks, and its generating target (i.e. attacked watermarked image) is the original image. Inspired by the U-Net network, the generative networks consist of encoder-decoder architecture with skip connection, which can combine the low-level and high-level information to ensure the imperceptibility of the generated image. Next, to further improve the imperceptibility of the generated image, instead of the loss function based on MSE, a perceptual loss based on feature extraction is introduced. In addition, a discriminative network is also introduced to make the appearance and distribution of generated image similar to those of the original image. The addition of the discriminative network can remove watermark information effectively. Extensive experiments are conducted to verify the feasibility of the proposed concealed attack method. Experimental and analysis results demonstrate that the proposed concealed attack method has better imperceptibility and attack ability in comparison to the existing watermarking attack methods. Qi Li 0029, Xingyuan Wang 0001, Bin Ma 0003, Xiaoyu Wang 0011, Chunpeng Wang 0001, Suo Gao, Yun Q. Shi 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | RD-IWAN: Residual Dense Based Imperceptible Watermark Attack NetworkabstractDigital watermarking technology and watermark attack methods are mutually reinforcing and complementary. Currently, traditional watermark attack methods are relatively mature, but these traditional attack methods will inevitably damage the visual quality of original images (OIs). Therefore, this paper proposes a covert attack method called residual dense based imperceptible watermark attack network (RD-IWAN). First, this paper designs a watermark attack residual dense network (WARDN) based on the residual dense network (RDN), which can effectively remove the watermark information in the middle and low frequency features of the watermarked image (WMI). Second, to improve the attack ability of the network, this paper innovatively proposes a progressive preprocessing method based on the information enhancement preprocessing method. Concurrently, to ensure the imperceptibility of this watermark attack method, a comprehensive loss function that combines the perceptual loss and mean square error loss (MSE) of OI and attacked watermarked image (AWMI) is designed in this study. Finally, attack experiments are designed and performed on watermarks with different embedding strengths and sizes. Experimental results show that, compared to traditional attack methods, the watermark attack method proposed in this paper exhibits stronger attack ability and higher imperceptibility. Chunpeng Wang 0001, Qixian Hao, Shujiang Xu, Bin Ma 0003, Qi Li 0029, Jian Li 0034, Yun Q. Shi 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Stereoscopic Image Description With Trinion Fractional-Order Continuous Orthogonal MomentsabstractSome research progress has been made on fractional-order continuous orthogonal moments (FrCOMs) in the past two years. Compared with integer-order continuous orthogonal moments (InCOMs), FrCOMs increase the number of affine invariants and effectively improve numerical stability. However, the existing types of FrCOMs are still very limited, of which all are planar image oriented. No report on stereoscopic images is available yet. To this end, in this paper, FrCOMs corresponding to various types of InCOMs are first deduced, and then, they are combined with trinion theory to construct trinion FrCOMs (TFrCOMs) applicable to stereoscopic images. Furthermore, the reconstruction performance and geometric invariance of TFrCOMs are analyzed theoretically and experimentally. Finally, an application in the stereoscopic image zero-watermarking algorithm is investigated to verify the superior performance of TFrCOMs. Chunpeng Wang 0001, Bin Ma 0003, Jian Li 0034, Qi Li 0029, Yun Q. Shi 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Detecting Aligned Double JPEG Compressed Color Image With Same Quantization Matrix Based on the Stability of ImageabstractJoint photographic experts group (JPEG) compression is widely used in image processing and computer vision. Detecting double compressed JPEG images is a common problem in forensics and detecting compressed images with the same quantization matrix remains a challenging task. However, most existing methods were designed for detection in grayscale images and cannot fully use the unique characteristics of color images (such as the relationship between channels and color information). In addition, the performance of existing methods is unsatisfactory for low JPEG quality factors and in cross detection experiments. To solve these problems, we analyze the stability of a color image to obtain the convergence error and transposition error. According to the convergence characteristics of color JPEG images, the continuous compression by the same quantization matrix can make the JPEG image tend to be stable. The final stable state and the convergence process are determined by the number of compressions of the original image. Thus, continuously compressed JPEG images can be regarded as a continuous frame to obtain the convergence error. As the color image converges, its ability to resist interference decreases. To reflect the changes in anti-interference ability, the transposition operation is used to disturb the color JPEG image to obtain the transposition error. In addition, quaternion mapping is used to retain the relationship between continuously compressed JPEG images and enlarge the influence caused by transposition operation. In our experiments on several image databases, the proposed method outperforms existing methods in different settings. Hao Wang 0060, Xiangyang Luo 0001, Yuhui Zheng, Bin Ma 0003, Jinsheng Sun, Sunil Kr. Jha |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2021 | ExNN-SMOTE: Extended Natural Neighbors Based SMOTE to Deal with Imbalanced DataabstractMany practical applications suffer from the problem of imbalanced classification. The minority class has poor classification performance; on the other hand, its misclassification cost is high. One reason for classification difficulty is the intrinsic complicated distribution characteristics (CDCs) in imbalanced data itself. Classical oversampling method SMOTE generates synthetic minority class examples between neighbors, which is parameter dependent. Furthermore, due to blindness of neighbor selection, SMOTE suffers from overgeneralization in the minority class. To solve such problems, we propose an oversampling method, called extended natural neighbors based SMOTE (ExNN-SMOTE). In ExNN-SMOTE, neighbors are determined adaptively by capturing data distribution characteristics. Extensive experiments over synthetic and real datasets demonstrate the effectiveness of ExNN-SMOTE dealing with CDCs and the superiority of ExNN-SMOTE over other SMOTE-related methods. Hongjiao Guan, Bin Ma 0003, Yingtao Zhang, Xianglong Tang |
ACML | 2 |
| 2021 | SURFR: A Real-Time Platform for Non-Coding RNA Fragmentation Analysis Using WaveletsabstractIt is well known that microRNAs (miRNAs or miRs) are small (~18-25 nt) yet highly potent non-coding RNA-derived RNAs (ndRNAs), originating from pre-miRNA fragmentation, that have been shown to alter the post-transcriptional functionality of many messenger RNAs (mRNAs). Biologically, the identification and study of miRNAs is very critical due to their increasing significance as biomarkers for many types of cancers and other genetic diseases. While empirical evidence supporting the existence of several novel ndRNAs excised from other longer non coding RNAs (ncRNAs) is growing, recent evidence suggests the full extent of their prevalence is likely underappreciated. Although some computational methods have been designed to help domain experts identify and understand miRNAs by analyzing Next Generation Sequencing (NGS) datasets, there are some crucial challenges, such as efficiency, effectiveness, and generalizability, in the state-of-the-art in-silico methods. To address such problems, our group proposed a new algorithm to mine ndRNAs by applying wavelet-based signal processing techniques as opposed to the current string-based NGS sequence alignment/analysis. However, due to novelty of the approach, our initial version of the algorithm was focused specifically on mining miRNAs, snoRNA-derived RNAs (sdRNAs) and transfer RNA (tRNA) fragments (tRFs) because of their importance in the literature plus the availability of experimentally validated databases to confirm our findings. Despite the computational issues, we still lack a basic understanding of the existence and the range of ndRNA functionalities from a) ndRNAs other than miRs, sdRNAs & tRFs in humans, and b) all ndRNAs in millions of organisms other than humans. Hence, there is an urgent requirement to automate the extraction and experimentation of ndRNAs, especially considering the rate at which NGS data is being produced. Therefore, in the current article, we extended our algorithm to be applicable to ~500 organisms—including eukaryotes, plants, bacteria, fungi, and protists—along with all their ncRNAs available in the current NCBI annotation. We also constructed a real-time user-friendly platform, SURFR, available at salts.soc.southalabama.edu/surfr, to aid domain experts and the aspiring biomedical scientists to perform RNA-Seq experiments to study ndRNAs. Not only our platform is extremely efficient, but we are also capable of allowing the users to identify, analyze, visualize, and compare ndRNAs from up to 30 NGS files to perform rigorous experimentation. Moreover, access to NGS files from public databases like SRA, and ndRNAs from private databases like TCGA are made readily available to the users to further validate their novel findings. Finally, we provide theoretical validation to examine our platform’s effectiveness. Mohan Vamsi Kasukurthi, Dominika Houserova, Dongqi Li, Jingwei Lin, Guanhuan Yang, Shaobo Tan, David M. Bourrie, Bin Ma 0003, Glen M. Borchert, Jingshan Huang |
BIBM | 10 |
| 2021 | A New Classification Algorithm and a New Oversampling Method of Mapping Common Data Elements to the BRIDG ModelabstractThe Common Data Elements (CDEs) standard of the International organization for Standardization (ISO) 11179 is commonly used in the field of clinical data processing. The Biomedical Research Integrated Domain Group (BRIDG) model is the framework for biomedical and clinical research. Mapping CDEs to BRIDG (also known as CDE classification) would help with interoperability and data analysis in the field of clinical research. That said, manually mapping CDEs to their corresponding BRIDG class is highly time-consuming and labor-intensive. In this paper we present a new classification algorithm along with a new oversampling method. Our algorithm uses the Term Frequency-Inverse Document Frequency (TF-IDF) as the feature representation method. By assigning different weights to various attributes, we enable more important attributes to perform more important roles during the mapping process. In addition, the oversampling method generates every new attribute in the minor class by picking the length and setting the word of the new attribute according to the existing training set. Our research outcomes demonstrate significant contributions to the field in the following ways: (1) Generation of a new CDE classification algorithm that outperforms existing algorithms in the literature, including the Random Forest Classifier, Linear Support Vector Classification (SVC), Multinomial Naive Bayes (NB), Logistic Regression, and Long Short-Term Memory (LSTM) networks, in terms of accuracy, precision, recall, and F-1 score measures. (2) Generation of a new oversampling method able to improve CDE classification accuracy for Random Forest and Multinomial NB. (3) Our classification algorithm employs two novel attributes, namely “Data Element Preferred Definition” and “Document,” which are more efficient at classifying CDEs than the six attributes traditionally selected by domain experts. Mohan Vamsi Kasukurthi, Jiajie Yang, Dongqi Li, Guanhuan Yang, Jingwei Lin, Shaobo Tan, David M. Bourrie, Bin Ma 0003, Glen M. Borchert, Jingshan Huang |
BIBM | 10 |
| 2021 | A Generalized Optimization Embedded Framework of Undersampling Ensembles for Imbalanced ClassificationabstractImbalanced classification exists commonly in practical applications, and it has always been a challenging issue. Traditional classification methods have poor performance on imbalanced data, especially, on the minority class. However, the minority class is usually of our interest, and its misclassification cost is higher. The critical factor is the intrinsic complicated distribution characteristics in imbalanced data itself. Resampling ensemble learning achieves promising results and is a research focus recently. However, some resampling ensembles do not consider complicated distribution characteristics, thus limiting the performance improvement. In this paper, a generalized optimization embedded framework (GOEF) is proposed based on undersampling bagging. The GOEF aims to pay more attention to the learning of local regions to handle the complicated distribution characteristics. Specifically, the GOEF utilizes out-of-bag data to explore heterogeneous local areas and chooses misclassified examples to optimize base classifiers. The optimization can focus on a single class or both classes. Extensive experiments over synthetic and real datasets demonstrate that GOEF with the minority class optimization performs the best in terms of AUC, G-mean, and sensitivity, compared with five resampling ensemble methods. Hongjiao Guan, Yingtao Zhang, Bin Ma 0003, Jian Li 0034, Chunpeng Wang 0001 |
DSAA | 3 |
| 2021 | Modify the Quantization Table in the JPEG Header File for Forensics and Anti-forensics
Hao Wang 0060, Xiangyang Luo 0001, Qilin Yin, Bin Ma 0003, Jinsheng Sun |
IWDW | 5 |
| 2021 | Octonion continuous orthogonal moments and their applications in color stereoscopic image reconstruction and zero-watermarking
Chunpeng Wang 0001, Qixian Hao, Bin Ma 0003, Jian Li 0034, Hongling Gao |
Eng. Appl. Artif. Intell. | 3 |
| 2021 | An encrypted coverless information hiding method based on generative models
Qi Li 0029, Xingyuan Wang 0001, Xiaoyu Wang 0011, Bin Ma 0003, Chunpeng Wang 0001, Yun Q. Shi 0001 |
Inf. Sci. | 4 |
| 2021 | Local quaternion polar harmonic Fourier moments-based multiple zero-watermarking scheme for color medical images
Xingyuan Wang 0001, Chunpeng Wang 0001, Bin Ma 0003, Yun Q. Shi 0001 |
Knowl. Based Syst. | 4 |
| 2021 | A reversible data hiding algorithm for audio files based on code division multiplexing
Bin Ma 0003, Jin-Cheng Hou, Chun-Peng Wang, Yun Q. Shi 0001 |
Multim. Tools Appl. | 1 |
| 2021 | Medical image super-resolution via deep residual neural network in the shearlet domain
Chunpeng Wang 0001, Simiao Wang, Qi Li 0029, Bin Ma 0003, Jian Li 0034, Meihong Yang, Yun Q. Shi 0001 |
Multim. Tools Appl. | 5 |
| 2021 | Medical Image Key Area Protection Scheme Based on QR Code and Reversible Data HidingabstractMedical image data, like most patient information, has high requirements for privacy and confidentiality. To improve the security of medical image transmission within the open network, we proposed a medical image key area protection algorithm based on reversible data hiding. First, the coefficient of variation is used to identify the key area, that is, the lesion area of the image. Then, the other regions are divided into blocks to analyze the texture complexity. Next, we propose a new reversible data hiding algorithm, which embeds the content of the key area into the high-texture regions. On this basis, a quick response (QR) code is generated using the ciphertext of the basic image information to replace the original lesion area. Experimental results show that this method can not only safely transmit sensitive patient information by hiding the content of the lesion, it can also store copyright information through QR code and achieve accurate image retrieval. Jian Xu 0025, Bin Ma 0003, Chunpeng Wang 0001, Jian Li 0034, Yuli Wang |
Secur. Commun. Networks | 3 |
| 2021 | High Precision Error Prediction Algorithm Based on Ridge Regression Predictor for Reversible Data HidingabstractAn efficient predictor is crucial for high embedding capacity and low image distortion. In this letter, a ridge regression-based high precision error prediction algorithm for reversible data hiding is proposed. The ridge regression is a penalized least-square algorithm, which solves the overfitting problem of the least-square method. Reversible data hiding based on ridge regression predictor minimizes the residual sum of squares between predicted and target pixels subject to the constraint expressed in terms of the L2-norm. Compared to a least-square-based predictor, the ridge regression-based predictor can obtain more small prediction errors, proving that the proposed method has a higher accuracy. In addition, the eight neighbor pixels of the target pixels and their two different combinations are selected as training and support sets, respectively. This selection scheme further improves the prediction accuracy. Experimental results show that the proposed method outperforms state-of-the-art adaptive reversible data hiding in terms of prediction accuracy and embedding performance. Xiaoyu Wang 0011, Xingyuan Wang 0001, Bin Ma 0003, Qi Li 0029, Yun Q. Shi 0001 |
IEEE Signal Process. Lett. | 3 |
| 2021 | Efficient Reversible Data Hiding for JPEG Images With Multiple Histograms ModificationabstractMost current reversible data hiding (RDH) techniques are designed for uncompressed images. However, JPEG images are more commonly used in our daily lives. Up to now, several RDH methods for JPEG images have been proposed, yet few of them investigated the adaptive data embedding as the lack of accurate measurement for the embedding distortion. To realize adaptive embedding and optimize the embedding performance, in this article, a novel RDH scheme for JPEG images based on multiple histogram modification (MHM) and rate-distortion optimization is proposed. Firstly, with selected coefficients, the RDH for JPEG images is generalized into a MHM embedding framework. Then, by estimating the embedding distortion, the rate-distortion model is formulated, so that the expansion bins can be adaptively determined for different histograms and images. Finally, to optimize the embedding performance in real time, a greedy algorithm with low computation complexity is proposed to derive the nearly optimal embedding efficiently. Experiments show that the proposed method can yield better embedding performance compared with state-of-the-art methods in terms of both visual quality and file size preservation. Mengyao Xiao, Xiaolong Li 0001, Bin Ma 0003, Xinpeng Zhang 0001, Yao Zhao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | Accurate Computation of Fractional-Order Exponential MomentsabstractExponential moments (EMs) are important radial orthogonal moments, which have good image description ability and have less information redundancy compared with other orthogonal moments. Therefore, it has been used in various fields of image processing in recent years. However, EMs can only take integer order, which limits their reconstruction and antinoising attack performances. The promotion of fractional-order exponential moments (FrEMs) effectively alleviates the numerical instability problem of EMs; however, the numerical integration errors generated by the traditional calculation methods of FrEMs still affect the accuracy of FrEMs. Therefore, the Gaussian numerical integration (GNI) is used in this paper to propose an accurate calculation method of FrEMs, which effectively alleviates the numerical integration error. Extensive experiments are carried out in this paper to prove that the GNI method can significantly improve the performance of FrEMs in many aspects. Shujiang Xu, Qixian Hao, Bin Ma 0003, Chunpeng Wang 0001, Jian Li 0034 |
Secur. Commun. Networks | 3 |
| 2020 | Robust image watermarking using invariant accurate polar harmonic Fourier moments and chaotic mapping
Bin Ma 0003, Lili Chang, Chunpeng Wang 0001, Jian Li 0034, Xingyuan Wang 0001, Yun Q. Shi 0001 |
Signal Process. | 1 |
| 2020 | Image Description With Polar Harmonic Fourier MomentsabstractDue to their good rotational invariance and stability, image continuous orthogonal moments are intensively applied in rotationally invariant recognition and image processing. However, most moments produce numerical instability, which impacts the image reconstruction and recognition performance. In this paper, a new set of invariant continuous orthogonal moments, polar harmonic Fourier moments (PHFMs), free of numerical instability is designed. The radial basis functions (RBFs) of the PHFMs are much simpler than those of the Chebyshev-Fourier moments (CHFMs), orthogonal Fourier-Mellin moments (OFMMs), Zernike moments (ZMs), and pseudo-Zernike moments (PZMs). For the same degree, the RBFs of the PHFMs have more zeros and are more evenly distributed than those of the ZMs and PZMs. Therefore, PHFMs do not suffer from information suppression problem; hence, the image description ability of the PHFMs is superior to that of the ZMs and PZMs. Moreover, the RBFs of the PHFMs are always less than or equal to 1.0 near the unit disk center, whereas those of the OFMMs, PZMs, CHFMs, and radial harmonic Fourier moments (RHFMs) are infinite (implying numerical instability). This indicates that PHFMs can outperform these moments in image reconstruction tasks. We theoretically and experimentally demonstrate that PHFMs outperform the above moments in reconstructing images and recognizing rotationally invariant objects considering noise and various attacks. This paper also details the significance of the PHFM phase in image reconstruction, angle estimation using PHFMs, and the accurate moment selection of the PHFMs. Chunpeng Wang 0001, Xingyuan Wang 0001, Bin Ma 0003, Yun Q. Shi 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2019 | SURFr: Algorithm for identification and analysis of ncRNA-derived RNAsabstractNoncoding RNAs (ncRNAs) regulate gene expression in essential cellular processes and play key roles in many human diseases. Small nucleolar RNAs (snoRNAs) are a relatively large group of ncRNAs. Many studies have identified significant alterations of different snoRNAs in prostate, breast, and lung malignancies and that many of these snoRNAs have been shown to be processed into microRNA-like molecules known as snoRNA-derived RNAs (sdRNAs). Similarly, several small RNAs have also recently been found to be excised from well characterized tRNAs to function in both normal cellular metabolism and disease. In this paper, we present a new computational methodology for the identification, analysis, and visualization of ncRNA-derived RNAs (ndRNAs) with linear time and space complexity (named SURFr for Short Uncharacterized R NA Finder). We provide the results of our algorithm's running time by analyzing several publicly available datasets and finally discuss future research directions to our work. Mohan Vamsi Kasukurthi, Glen M. Borchert, Jingshan Huang, Dihua Zhang, Mika Housevera, Shaobo Tan, Bin Ma 0003, Dongqi Li, Ryan G. Benton, Jingwei Lin |
BIBM | 9 |
| 2019 | Reversible Data Hiding Based Key Region Protection Method in Medical ImagesabstractThe transmission of medical image data in an open network environment is subject to privacy issues including patient privacy and data leakage. In the past, image encryption and information-hiding technology have been used to solve such security problems. But these methodologies, in general, suffered from difficulties in retrieving original images. We present in this paper an algorithm to protect key regions in medical images. First, coefficient of variation is used to locate the key regions, a.k.a. the lesion areas, of an image; other areas are then processed in blocks and analyzed for texture complexity. Next, our reversible data-hiding algorithm is used to embed the contents from the lesion areas into a high-texture area, and the Arnold transformation is performed to protect the original lesion information. In addition to this, we use the ciphertext of the basic information about the image and the decryption parameter to generate the Quick Response (QR) Code to replace the original key regions. Consequently, only authorized customers can obtain the encryption key to extract information from encrypted images. Experimental results show that our algorithm can not only restore the original image without information loss, but also safely transfer the medical image copyright and patient-sensitive information. Jian Li 0034, Shaobo Tan, Bin Ma 0003, Meihong Yang, Jingshan Huang, Ryan G. Benton, Mohan Vamsi Kasukurthi, Dongqi Li, Jingwei Lin, Glen M. Borchert |
BIBM | 3 |
| 2019 | A SVM-Based Algorithm to Diagnose Sleep ApneaabstractObstructive sleep apnea syndrome (OSAS) is a breathing disorder presenting during sleep. Although polysomnography (PSG) is the gold standard to diagnose OSAS, it is an expensive method that is quite complicated to use. Worse, it takes a long time between testing and getting a diagnosis from PSG. Thus, we have designed an algorithm aimed at diagnosing OSAS in a more efficient manner. First, blood oxygen saturation (SpO2) data are processed to obtain statistical features, which are then trained to establish a classification model based on a support vector machine (SVM) strategy; the resulting SVM model performs the diagnosis of OSAS. Furthermore, in order to allow remote diagnosis, we combine our algorithm with a monitoring system. To achieve this, physiological data are collected from a smart phone and then uploaded to the SVM model in the cloud. Once processed, a diagnosis report is returned to the smart phone. A preliminary evaluation of our algorithm based on real-world data is extremely promising as we find its accuracy, sensitivity, and specificity to be 90.2%, 87.6%, and 94.1%, respectively. Bin Ma 0003, Shaobo Tan, Meihong Yang, Jingshan Huang, Zhaolong Wu, Ryan G. Benton, Dongqi Li, Mohan Vamsi Kasukurthi, Jingwei Lin, Glen M. Borchert |
BIBM | 1 |
| 2019 | Use CPET data to predict the intervention effect of aerobic exercise on young hypertensive patientsabstractThe incidence of hypertension has recently shown a significant increase in young people, with aerobic exercise intervention being recognized as an effective approach to decrease blood pressure (BP). However, BP response to aerobic exercise can be highly individualized, and no research has been conducted on predicting the effect of aerobic exercise intervention for reducing BP in young hypertensive patients. In this work, we use the data generated from a cardiopulmonary exercise test (CPET) in young hypertensive patients (before aerobic exercise intervention) to derive information from multiple cardiopulmonary metabolic indices. The data, presented as time series, are then analyzed by a machine learning method to predict the effect of aerobic exercise intervention in lowering BP. This study provides several novel insights for making personalized aerobic exercise intervention programs for young adults with stage I hypertension. Guanyi Yang, Ryan G. Benton, Glen M. Borchert, Bin Ma 0003, Jingshan Huang, Xiuyu Leng, Fangwan Huang, Mohan Vamsi Kasukurthi, Dongqi Li, Jingwei Lin, Shaobo Tan, Guiying Lu |
BIBM | 4 |
| 2019 | Transform Domain Based Medical Image Super-resolution via Deep Multi-scale NetworkabstractThis paper proposes a new medical image super-resolution (SR) network, namely deep multi-scale network (DMSN), in the uniform discrete curvelet transform (UDCT) domain. DMSN is made up of a set of cascaded multi-scale fushion (MSF) blocks. In each MSF block, we use convolution kernels of different sizes to adaptively detect the local multi-scale feature, and then local residual learning (LRL) is used to learn effective feature from preceding MSF block and current multi-scale features. After obtaining multi-scale features of different MSF block, we use global feature fusion (GFF) to jointly and adaptively learn global hierarchical features in a holistic manner. Finally, compared with other prediction methods in spatial domain, we applied DMSN in UDCT domain, which enables a better representation of global topological structure and local texture detail of HR images. DM-SN shows superior performance over other state-of-the-art medical image SR methods. Chunpeng Wang 0001, Simiao Wang, Bin Ma 0003, Jian Li 0034, Xiangjun Dong 0001 |
ICASSP | 3 |
| 2019 | Code Division Multiplexing and Machine Learning Based Reversible Data Hiding Scheme for Medical ImageabstractIn this paper, a new reversible data hiding (RDH) scheme based on Code Division Multiplexing (CDM) and machine learning algorithms for medical image is proposed. The original medical image is firstly converted into frequency domain with integer-to-integer wavelet transform (IWT) algorithm, and then the secret data are embedded into the medium frequency subbands of medical image robustly with CDM and machine learning algorithms. According to the orthogonality of different spreading sequences employed in CDM algorithm, the secret data are embedded repeatedly, most of the elements of spreading sequences are mutually canceled, and the proposed method obtained high data embedding capacity at low image distortion. Simultaneously, the to-be-embedded secret data are represented by different spreading sequences, and only the receiver who has the spreading sequences the same as the sender can extract the secret data and original image completely, by which the security of the RDH is improved effectively. Experimental results show the feasibility of the proposed scheme for data embedding in medical image comparing with other state-of-the-art methods. Bin Ma 0003, Bing Li 0014, Xiaoyu Wang 0011, Chunpeng Wang 0001, Jian Li 0034, Yun Q. Shi 0001 |
Secur. Commun. Networks | 1 |
| 2018 | A PWM-Based Muscle Fatigue Detection and Recovery System
Bin Ma 0003, Chunxiao Li 0005, Zhaolong Wu, Ada Chaeli van der Zijp-Tan, Shaobo Tan, Dongqi Li, Ada Fong, Chandan Basetty, Glen M. Borchert, Jingshan Huang |
BIBM | 1 |
| 2018 | A Multiple Linear Regression Based High-Accuracy Error Prediction Algorithm for Reversible Data Hiding
Bin Ma 0003, Xiaoyu Wang 0011, Bing Li 0014, Yun Q. Shi 0001 |
IWDW | 1 |
| 2018 | A Multiple Linear Regression Based High-Performance Error Prediction Method for Reversible Data Hiding
Bin Ma 0003, Xiaoyu Wang 0011, Bing Li 0014, Yun Q. Shi 0001 |
SecureComm (2) | 1 |
| 2018 | Extraction of PRNU noise from partly decoded video
Jian Li 0034, Bin Ma 0003, Chunpeng Wang 0001 |
J. Vis. Commun. Image Represent. | 2 |
| 2016 | A Novel CDMA Based High Performance Reversible Data Hiding SchemeabstractIn this paper, based on the principle of Code Division Multiple Access (CDMA), a novel reversible data hiding scheme is presented. The to-be-embedded data are represented by different orthogonal spreading sequences and embedded into a cover image while degrading the image quality slightly. According to the feature of orthogonality, different spreading sequences are repeatedly embedded into the image without disturbing each other, and most elements of different spreading sequences are mutually cancelled in the process of multilevel data embedding. Thus, it keeps the distortion of the embedded image at a relatively low level even with a high embedding capacity. Moreover, the location-map of the proposed scheme can be highly compressed and thus the size is quite small; it further helps to obtain high net embedding capacity. Experimental results have demonstrated that the CDMA based reversible data hiding scheme can achieve higher image quality at the moderate-to-high embedding capacity than other state-of-the-art reversible data hiding works. Bin Ma 0003, Jian Xu 0025, Yun Q. Shi 0001 |
IH&MMSec | 1 |
| 2016 | A Reversible Data Hiding Scheme Based on Code Division MultiplexingabstractIn this paper, a novel code division multiplexing (CDM) algorithm-based reversible data hiding (RDH) scheme is presented. The covert data are denoted by different orthogonal spreading sequences and embedded into the cover image. The original image can be completely recovered after the data have been extracted exactly. The Walsh Hadamard matrix is employed to generate orthogonal spreading sequences, by which the data can be overlappingly embedded without interfering each other, and multilevel data embedding can be utilized to enlarge the embedding capacity. Furthermore, most elements of different spreading sequences are mutually cancelled when they are overlappingly embedded, which maintains the image in good quality even with a high embedding payload. A location-map free method is presented in this paper to save more space for data embedding, and the overflow/underflow problem is solved by shrinking the distribution of the image histogram on both the ends. This would further improve the embedding performance. Experimental results have demonstrated that the CDM-based RDH scheme can achieve the best performance at the moderate-to-high embedding capacity compared with other state-of-the-art schemes. Bin Ma 0003, Yun Q. Shi 0001 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2014 | A Reversible Image Watermarking Scheme Based on Modified Integer-to-Integer Discrete Wavelet Transform and CDMA Algorithm
Bin Ma 0003, Yun Q. Shi 0001 |
IWDW | 1 |