EDBT 2026 Demo / reviewers in the wild / expert
Zehua Ma
dblp:240/2623
· DBLP profile ↗
35ranked-venue papers
4as first author
32since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 26 · 3 first-author · 23 since 2021Artificial intelligence and machine learning · 12 · 1 first-author · 11 since 2021Security and privacy · 4 · 4 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LAMIC: Layout-Aware Multi-Image Composition via Scalability of Multimodal Diffusion TransformerabstractIn controllable image synthesis, generating coherent and consistent images from multiple references with spatial layout awareness remains an open challenge. We propose LAMIC, a Layout-Aware Multi-Image Composition framework that, for the first time, extends single-reference diffusion models to multi-reference scenarios in a training-free manner. Built upon the MMDiT model, LAMIC introduces two plug-and-play attention mechanisms: 1) Group Isolation Attention (GIA) to enhance entity disentanglement; and 2) Region-Modulated Attention (RMA) to enable layout-aware generation. To comprehensively evaluate model capabilities, we further introduce three metrics: 1) Inclusion Ratio (IN-R) and Fill Ratio (FI-R) for assessing layout control; and 2) Background Similarity (BG-S) for measuring background consistency. Extensive experiments show that LAMIC achieves state-of-the-art performance across most major metrics: it consistently outperforms existing multi-reference baselines in ID-S, BG-S, IN-R and AVG scores across all settings, and achieves the best DPG in complex composition tasks. These results demonstrate LAMIC's superior abilities in identity keeping, background preservation, layout control, and prompt-following, all achieved without any training or fine-tuning, showcasing strong zero-shot generalization ability. By inheriting the strengths of advanced single-reference models and enabling seamless extension to multi-image scenarios, LAMIC establishes a new training-free paradigm for controllable multi-image composition. As foundation models continue to evolve, LAMIC's performance is expected to scale accordingly. Yuzhuo Chen, Zehua Ma, Weiming Zhang 0001 |
AAAI | 2 |
| 2026 | AuthSig: Safeguarding Scanned Signatures Against Unauthorized Reuse in Paperless WorkflowsabstractWith the deepening trend of paperless workflows, signatures as a means of identity authentication are gradually shifting from traditional ink-on-paper to electronic formats. Despite the availability of dynamic pressure-sensitive and PKI-based digital signatures, static scanned signatures remain prevalent in practice due to their convenience. However, these static images, having almost lost their authentication attributes, cannot be reliably verified and are vulnerable to malicious copying and reuse. To address these issues, we propose AuthSig, a novel static electronic signature framework based on generative models and watermark, which binds authentication information to the signature image. Leveraging the human visual system’s insensitivity to subtle style variations, AuthSig finely modulates style embeddings during generation to implicitly encode watermark bits-enforcing a One Signature, One Use policy. To overcome the scarcity of handwritten signature data and the limitations of traditional augmentation methods, we introduce a keypoint-driven data augmentation strategy that effectively enhances style diversity to support robust watermark embedding. Experimental results show that AuthSig achieves over 98% extraction accuracy under both digital-domain distortions and signature-specific degradations, and remains effective even in print-scan scenarios. Ruiqiang Zhang, Zehua Ma, Guanjie Wang, Chang Liu 0089, Hengyi Wang, Weiming Zhang 0001 |
AAAI | 2 |
| 2025 | RoPaSS: Robust Watermarking for Partial Screen-Shooting ScenariosabstractScreen-shooting robust watermarking is an effective means of preventing screen content leakage from unauthorized camera shooting, as it can trace the leaked source through the watermark extraction thereby providing an effective deterrent. However, current screen-shooting resilient watermarking schemes rely on the image's contours to synchronize and then extract the watermark. While in practical applications, it's common for only a portion of the image to be captured, resulting in a limited performance of the previous watermarking schemes. To address this problem, we propose the RoPaSS: a robust watermarking scheme for partial screen-shooting scenarios, which effectively constructs symmetric characteristics on the embedding watermark to handle the sticky re-synchronization issue. Specifically, RoPaSS consists of a watermark encoder, a decoder, and three estimators, which are trained in two stages. In the first training stage, RoPaSS integrates the flipping operation into the watermark encoder and decoder training to increase the redundancy of watermark messages and artificially guide the generation of symmetric watermarks. In the second stage, estimators utilize the watermark symmetry as an additional reference to estimate the restoration parameters to resynchronize the partially captured watermarked image. Experiments have demonstrated the excellent performance of RoPaSS in partial screen-shooting traceability, with extraction accuracy of above 93% in frontal shooting and above 86% in 30° shooting even if only 50% of the image content is captured. Zehua Ma, Han Fang 0004, Kejiang Chen, Weiming Zhang 0001 |
AAAI | 1 |
| 2025 | TAG-WM: Tamper-Aware Generative Image Watermarking via Diffusion Inversion SensitivityabstractAI-generated content (AIGC) enables efficient visual creation but raises copyright and authenticity risks. As a common technique for integrity verification and source tracing, digital image watermarking is regarded as a potential solution to above issues. However, the widespread adoption and advancing capabilities of generative image editing tools have amplified malicious tampering risks, while simultaneously posing new challenges to passive tampering detection and watermark robustness. To address these challenges, this paper proposes a Tamper-Aware Generative image WaterMarking method named TAG-WM. The proposed method comprises four key modules: a dual-mark joint sampling (DMJS) algorithm for embedding copyright and localization watermarks into the latent space while preserving generative quality, the watermark latent reconstruction (WLR) utilizing reversed DMJS, a dense variation region detector (DVRD) leveraging diffusion inversion sensitivity to identify tampered areas via statistical deviation analysis, and the tamper-aware decoding (TAD) guided by localization results. The experimental results demonstrate that TAG-WM achieves state-of-the-art performance in both tampering robustness and localization capability even under distortion, while preserving lossless generation quality and maintaining a watermark capacity of 256 bits. The code is available at: https://github.com/Suchenl/TAG-WM. Yuzhuo Chen, Zehua Ma, Han Fang 0004, Weiming Zhang 0001, Nenghai Yu |
ICCV | 2 |
| 2025 | SynTag: Enhancing the Geometric Robustness of Inversion-Based Generative Image Watermarking
Han Fang 0004, Kejiang Chen, Zehua Ma, Jiajun Deng, Yicong Li 0004, Weiming Zhang 0001, Ee-Chien Chang |
ICCV | 3 |
| 2025 | A Watermark Updating Framework for Multi-stage Image Content DistributionabstractDeep image watermarking embeds identification data into images to facilitate source tracking. However, existing schemes are primarily designed for single-stage transmission scenarios, and in practical multi-stage distribution requirements, current methods degrade image quality and reduce watermark extraction accuracy. In this paper, we introduces WaterUp, a deep watermark updating framework. WaterUp automatically updates watermark information as the image is transmitted, preserving image quality while accurately recording the transmission path for traceability. The core of WaterUp is a flow-based encoder-decoder (FED), which utilizes a forward and backward network to enable efficient watermark updating with minimal computational and storage demands. Experimental results show that WaterUp outperforms state-of-the-art methods, maintaining high visual quality with a PSNR exceeding 38 dB across multiple transmissions. Bin Liu 0016, Jie Zhang 0073, Xiang Zhang 0011, Zehua Ma, Nenghai Yu |
ICME | 5 |
| 2025 | CamLopa: A Hidden Wireless Camera Localization Framework via Signal Propagation Path AnalysisabstractHidden wireless cameras pose significant privacy threats, necessitating effective detection and localization methods. However, existing localization solutions often require impractical activity spaces, expensive specialized devices, or pre-collected training data, limiting their practical deployment. To address these limitations, we introduce CamLopa, a training-free wireless camera localization framework that operates with minimal activity space constraints using low-cost, commercial-off-the-shelf (COTS) devices. CamLopa can achieve detection and localization in just 45 seconds of user activities with a Raspberry Pi board. During this short period, it analyzes the causal relationship between wireless traffic and user movement to detect the presence of a hidden camera. Upon detection, CamLopa utilizes a novel azimuth localization model based on wireless signal propagation path analysis for localization. This model leverages the time ratio of user paths crossing the First Fresnel Zone (FFZ) to determine the camera's azimuth angle. Subsequently, CamLopa refines the localization by identifying the camera's quadrant. We evaluate CamLopa across various devices and environments, demonstrating its effectiveness with a 95.37% detection accuracy for snooping cameras and an average localization error of 17.23°, under the significantly reduced activity space requirements and without the need for training. Our code and demo are available at https://github.com/CamLoPA/CamLoPA-Code. Xiang Zhang 0011, Jie Zhang 0073, Zehua Ma, Jinyang Huang, Meng Li 0006, Huan Yan 0004, Peng Zhao 0024, Zijian Zhang 0001, Bin Liu 0016, Qing Guo 0005, Tianwei Zhang 0004, Nenghai Yu |
SP | 3 |
| 2025 | DiffLoc: WiFi Hidden Camera Localization Based on Electromagnetic Diffraction
Xiang Zhang 0011, Jie Zhang 0073, Huan Yan 0004, Jinyang Huang, Zehua Ma, Bin Liu 0016, Meng Li 0006, Kejiang Chen, Qing Guo 0005, Tianwei Zhang 0004, Zhi Liu 0002 |
USENIX Security Symposium | 5 |
| 2025 | C³shartMark: A Chart Watermarking Scheme With Consecutive-Encoding and Concurrent-DecodingabstractChart images are widely employed as the intuitive form to express information, which renders them highly valuable. Consequently, there is an urgent demand to develop a watermarking algorithm for copyright protection and leakage prevention of chart images. Nevertheless, existing chart watermarking methods fail to thoroughly consider the chart image’s special characteristics and simply rely on the previous natural image-based watermarking framework. Compared to natural images, the chart image generally exhibits relatively simple layouts and textures, containing fewer complex texture regions that watermarks are typically embedded in. Therefore, the embedding locations of watermarks for different distortions can be relatively dispersed in natural images, while for chart images, watermark embedding regions under various distortion conditions tend to be relatively concentrated and share more overlaps. Inspired by the above special characteristics of chart images, to sufficiently leverage them and design a better framework, this paper proposes C3hartMark, a chart watermarking scheme with consecutive-encoding and concurrent-decoding. Instead of using the combined noise layer as existing methods to ensure multiple robustness, a novel consecutive training framework is introduced in this paper, which efficiently utilizes the overlapping of embedded watermark features in chart images, and simultaneously, mitigates the poor convergence brought by the combined noise layer. During the extraction stage, multiple concurrent decoders are introduced to extract the potential embedded watermarks for different distortions independently. Moreover, we also incorporate two special noise layers, namely Captioning and Fusion, to address the corresponding realistic distortions in chart images, and an agnostic noise layer to accommodate potential channel transmission distortions unknown during training. Through extensive experiments, we demonstrate that with the better visual quality, C3hartMark simultaneously outperforms existing state-of-the-art (SOTA) watermarking methods in terms of robustness, achieving 99.57% extraction accuracy under JPEG compression (QF=60). Linfeng Ma, Han Fang 0004, Zehua Ma, Zhaoyang Jia, Weiming Zhang 0001, Nenghai Yu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Synthesizing Glyph Vectors for Practical Information Hiding in DocumentsabstractDocuments are ubiquitous vehicles for information transmission. Beyond the visible content meant for reading, there is a growing interest in hiding additional information in documents. Recent studies have focused on the utilization of glyphs, which are stored in vector format within computer systems. Specifically, glyph variants are manually designed to substitute the original ones in documents, thereby representing information. However, such strategies are costly, only effective for specific font types, and fragile to physical distortions. To address these limitations, this paper presents AutoStegaFont+, a two-stage and dual-modality learning framework designed to synthesize glyph vectors capable of conveying hidden information under real-world distortions. In the first stage, we jointly train an encoder and a decoder with a specialized distortion layer to achieve robust information encoding and decoding of glyph images. Then, the second stage employs a differentiable rasterizer to transfer the information from encoded glyph images to corresponding vectors, enabling the automatic generation of encoded vectors. Extensive experiments demonstrate the robust performance of AutoStegaFont+ across a variety of real-world scenarios, including screenshots, print-camera shooting, and screen-camera shooting, while maintaining compatibility with diverse font types. Additionally, we investigate the information-carrying capacity of individual glyphs, exploring their impact on robustness and visual quality. Jie Zhang 0073, Chang Liu 0089, Han Fang 0004, Zehua Ma, Kejiang Chen, Weiming Zhang 0001, Nenghai Yu |
IEEE Trans. Dependable Secur. Comput. | 5 |
| 2024 | MuST: Robust Image Watermarking for Multi-Source TracingabstractIn recent years, with the popularity of social media applications, massive digital images are available online, which brings great convenience to image recreation. However, the use of unauthorized image materials in multi-source composite images is still inadequately regulated, which may cause significant loss and discouragement to the copyright owners of the source image materials. Ideally, deep watermarking techniques could provide a solution for protecting these copyrights based on their encoder-noise-decoder training strategy. Yet existing image watermarking schemes, which are mostly designed for single images, cannot well address the copyright protection requirements in this scenario, since the multi-source image composing process commonly includes distortions that are not well investigated in previous methods, e.g., the extreme downsizing. To meet such demands, we propose MuST, a multi-source tracing robust watermarking scheme, whose architecture includes a multi-source image detector and minimum external rectangle operation for multiple watermark resynchronization and extraction. Furthermore, we constructed an image material dataset covering common image categories and designed the simulation model of the multi-source image composing process as the noise layer. Experiments demonstrate the excellent performance of MuST in tracing sources of image materials from the composite images compared with SOTA watermarking methods, which could maintain the extraction accuracy above 98% to trace the sources of at least 3 different image materials while keeping the average PSNR of watermarked image materials higher than 42.51 dB. We released our code on https://github.com/MrCrims/MuST Guanjie Wang, Zehua Ma, Chang Liu 0089, Han Fang 0004, Weiming Zhang 0001, Nenghai Yu |
AAAI | 2 |
| 2024 | A Geometric Distortion Immunized Deep Watermarking Framework with Robustness Generalizability
Linfeng Ma, Han Fang 0004, Tianyi Wei, Zijin Yang, Zehua Ma, Weiming Zhang 0001, Nenghai Yu |
ECCV (67) | 5 |
| 2024 | DERO: Diffusion-Model-Erasure Robust WatermarkingabstractThe effective denoising demonstrated by the latent diffusion model poses a new threat to image watermarking, as attackers can erase the watermark by performing a forward diffusion, followed by backward denoising. While such denoising might introduce large distortion in the pixel domain, the image semantics remain similar. Unfortunately, most existing robust watermarking methods fail to tackle such an erasure attack since they are primarily designed for traditional channel distortions. To address such issue, this paper proposed DERO, a diffusion-model-erasure robust watermarking framework. Based on the frequency domain analysis of the diffusion model's denoising process, we designed a destruction and compensation noise layer (DCNL) to approximate the distortion effects caused by latent diffusion model erasure (LDE). In detail, DCNL consists of a multi-scale low-pass filtering and a white noise compensation process, where the high-frequency components of the image are first obliterated, and then full-frequency components are enriched with white noise. Such a process broadly simulates the LDE distortions. Besides, on the extraction side, we cascaded a pre-trained variational autoencoder before the decoder to extract the watermark in the latent domain, which closely adapts to the operation domain of the LDE process. Meanwhile, to improve the robustness of the decoder, we also design a latent feature augmentation (LFA) operation on the latent feature. Throughout the end-to-end training with the DCNL and LFA, DERO can successfully achieve robustness against LDE. Our experimental results demonstrate the effectiveness and the generalizability of the proposed framework. The LDE robustness is significantly improved from 75% with SOTA methods to an impressive 96% with DERO. Han Fang 0004, Kejiang Chen, Yupeng Qiu, Zehua Ma, Weiming Zhang 0001, Ee-Chien Chang |
ACM Multimedia | 4 |
| 2024 | DreamVTON: Customizing 3D Virtual Try-on with Personalized Diffusion Models
Zhenyu Xie, Haoye Dong, Zehua Ma, Xiaodan Liang |
ACM Multimedia | 4 |
| 2024 | Hidden WiFi Camera Localization via Signal Propagation Path AnalysisabstractHidden WiFi cameras pose significant privacy threats, necessitating effective localization methods. In this work, we introduce CamLoPA, a system designed for the detection and localization of WiFi cameras. CamLoPA achieves this in just 45 seconds of user walking. It begins by analyzing the causal relationship between WiFi traffic and user movement to identify the presence of a snooping camera. Upon detection, CamLoPA utilizes a novel azimuth location model based on WiFi signal propagation path analysis to localize the hidden camera. Comprehensive evaluations demonstrate that CamLoPA can accurately and swiftly detect and localize snooping WiFi cameras with minimal constraints. Xiang Zhang 0011, Zehua Ma, Jinyang Huang, Huan Yan 0004, Meng Li 0006, Zhi Liu 0002, Bin Liu 0016 |
MobiCom | 2 |
| 2024 | Robust Model Watermarking for Image Processing Networks via Structure ConsistencyabstractThe intellectual property of deep networks can be easily "stolen" by surrogate model attack. There has been significant progress in protecting the model IP in classification tasks. However, little attention has been devoted to the protection of image processing models. By utilizing consistent invisible spatial watermarks, the work (Zhang et al. 2020) first considered model watermarking for deep image processing networks and demonstrated its efficacy in many downstream tasks. Its success depends on the hypothesis that if a consistent watermark exists in all prediction outputs, that watermark will be learned into the attacker's surrogate model. However, when the attacker uses common data augmentation attacks (e.g., rotate, crop, and resize) during surrogate model training, it will fail because the underlying watermark consistency is destroyed. To mitigate this issue, we propose a new watermarking methodology, "structure consistency", based on which a new deep structure-aligned model watermarking algorithm is designed. Specifically, the embedded watermarks are designed to be aligned with physically consistent image structures, such as edges or semantic regions. Experiments demonstrate that our method is more robust than the baseline in resisting data augmentation attacks. Besides that, we test the generalization ability and robustness of our method to a broader range of adaptive attacks. Jie Zhang 0073, Dongdong Chen 0001, Jing Liao 0001, Zehua Ma, Han Fang 0004, Weiming Zhang 0001, Huamin Feng, Gang Hua 0001, Nenghai Yu |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | A Robust Database Watermarking Scheme That Preserves Statistical CharacteristicsabstractDatabase watermarking can be used for copyright verification and leakage traceability, effectively protecting the security of the database. However, the existing watermarking schemes commonly embed watermarks by modifying the original data, which changes the statistical characteristics and affects the statistical analysis of the database. Therefore, this paper proposes SCPW, aStatisticalCharacteristicsPreserving robust databaseWatermarking framework. First, we perform a theoretical analysis and propose a data modification scheme maintaining the statistical characteristics unchanged. Then, we establish the correspondence between the data and the watermarks that need to be embedded in it by grouping. Finally, the watermark message is embedded into the database through data verification and modification. Specifically, for data that needs to be watermarked, we first verify whether the potential watermark bits extracted from the data are the same as bits that need to be embedded. If they are the same, we regard this original data, usually a floating point number, as a “good number” and do not modify it. Otherwise, we modify the data until it becomes a “good number” using a data modification scheme that preserves the statistical characteristics proposed by the theoretical analysis. In addition, we also use the genetic algorithm to optimize the grouping results and increase the proportion of “good number”, thereby reducing the proportion of data that needs to be modified and further reducing distortion. To our best knowledge, SCPW is the first watermarking scheme that ensures the preservation of statistical characteristics, and the experimental results also prove its excellent ability to preserve statistical characteristics compared to existing schemes. Moreover, experiments also illustrate that our method is robust against a wide range of attacks. When under deletion attack (deletion rate = 90%), the bit error rate of watermark extraction is only 0.8%, which is more than 12% lower than the current best method. Zhiwen Ren, Han Fang 0004, Jie Zhang 0073, Zehua Ma, Ronghao Lin, Weiming Zhang 0001, Nenghai Yu |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2023 | DeAR: A Deep-Learning-Based Audio Re-recording Resilient WatermarkingabstractAudio watermarking is widely used for leaking source tracing. The robustness of the watermark determines the traceability of the algorithm. With the development of digital technology, audio re-recording (AR) has become an efficient and covert means to steal secrets. AR process could drastically destroy the watermark signal while preserving the original information. This puts forward a new requirement for audio watermarking at this stage, that is, to be robust to AR distortions. Unfortunately, none of the existing algorithms can effectively resist AR attacks due to the complexity of the AR process. To address this limitation, this paper proposes DeAR, a deep-learning-based audio re-recording resistant watermarking. Inspired by DNN-based image watermarking, we pioneer a deep learning framework for audio carriers, based on which the watermark signal can be effectively embedded and extracted. Meanwhile, in order to resist the AR attack, we delicately analyze the distortions that occurred in the AR process and design the corresponding distortion layer to cooperate with the proposed watermarking framework. Extensive experiments show that the proposed algorithm can resist not only common electronic channel distortions but also AR distortions. Under the premise of high-quality embedding (SNR=25.86dB), in the case of a common re-recording distance (20cm), the algorithm can effectively achieve an average bit recovery accuracy of 98.55%. Chang Liu 0089, Jie Zhang 0073, Han Fang 0004, Zehua Ma, Weiming Zhang 0001, Nenghai Yu |
AAAI | 4 |
| 2023 | AutoStegaFont: Synthesizing Vector Fonts for Hiding Information in DocumentsabstractHiding information in text documents has been a hot topic recently, with the most typical schemes of utilizing fonts. By constructing several fonts with similar appearances, information can be effectively represented and embedded in documents. However, due to the unstructured characteristic, font vectors are more difficult to synthesize than font images. Existing methods mainly use handcrafted features to design the fonts manually, which is time-consuming and labor-intensive. Moreover, due to the diversity of fonts, handcrafted features are not generalizable to different fonts. Besides, in practice, since documents might be distorted through transmission, ensuring extractability under distortions is also an important requirement. Therefore, three requirements are imposed on vector font generation in this domain: automaticity, generalizability, and robustness. However, none of the existing methods can satisfy these requirements well and simultaneously. To satisfy the above requirements, we propose AutoStegaFont, an automatic vector font synthesis scheme for hiding information in documents. Specifically, we design a two-stage and dual-modality learning framework. In the first stage, we jointly train an encoder and a decoder to invisibly encode the font images with different information. To ensure robustness, we target designing a noise layer to work with the encoder and decoder during training. In the second stage, we employ a differentiable rasterizer to establish a connection between the image and the vector modality. Then, we design an optimization algorithm to convey the information from the encoded image to the corresponding vector. Thus the encoded font vectors can be automatically generated. Extensive experiments demonstrate the superior performance of our scheme in automatically synthesizing vector fonts for hiding information in documents, with robustness to distortions caused by low-resolution screenshots, printing, and photography. Besides, the proposed framework has better generalizability to fonts with diverse styles and languages. Jie Zhang 0073, Han Fang 0004, Chang Liu 0089, Zehua Ma, Weiming Zhang 0001, Nenghai Yu |
AAAI | 5 |
| 2023 | AnisoTag: 3D Printed Tag on 2D Surface via Reflection AnisotropyabstractIn the past few years, the widespread use of 3D printing technology enables the growth of the market of 3D printed products. On Esty, a website focused on handmade items, hundreds of individual entrepreneurs are selling their 3D printed products. Inspired by the positive effects of machine-readable tags, like barcodes, on daily product marketing, we propose AnisoTag, a novel tagging method to encode data on the 2D surface of 3D printed objects based on reflection anisotropy. AnisoTag has an unobtrusive appearance and much lower extraction computational complexity, contributing to a lightweight low-cost tagging system for individual entrepreneurs. On AnisoTag, data are encoded by the proposed tool as reflective anisotropic microstructures, which would reflect distinct illumination patterns when irradiating by collimated laser. Based on it, we implement a real-time detection prototype with inexpensive hardware to determine the reflected illumination pattern and decode data according to their mapping. We evaluate AnisoTag with various 3D printer brands, filaments, and printing parameters, demonstrating its superior usability, accessibility, and reliability for practical usage. Zehua Ma, Hang Zhou 0007, Weiming Zhang 0001 |
CHI | 1 |
| 2023 | ProTegO: Protect Text Content against OCR Extraction AttackabstractOnline documents greatly improve the efficiency of information interaction but also cause potential security hazards, such as the ability to copy and reuse text content without authorization readily. To address copyright concerns, recent works have proposed converting reproducible text content into non-reproducible formats, making digital text content observable but not duplicable. However, as the Optical Character Recognition (OCR) technology develops, adversaries can still take screenshots of the target text region and use OCR to extract the text content. None of the existing methods can be well adapted to this kind of OCR extraction attack. In this paper, we propose "ProTegO'', a novel text content protection method against the OCR extraction attack, which generates adversarial underpaintings that do not affect human reading but can interfere with OCR after taking screenshots. Specifically, we design a text-style universal adversarial underpaintings generation framework, which can mislead both text recognition models and commercial OCR services. For invisibility, we take full advantage of the fusion property of human eyes and create complementary underpaintings to display alternatively on the screen. Experimental results demonstrate that ProTegO is a one-size-fits-all method that can ensure good visual quality while simultaneously achieving a high protection success rate on text recognition models with different architectures, outperforming the state-of-the-art methods. Furthermore, we validate the feasibility of ProTegO on a wide range of popular commercial OCR services, including Microsoft, Tencent, Alibaba, Huawei, Baidu, Apple, and Xiaomi. Codes will be available at https://github.com/Ruby-He/ProTegO. Yanru He, Kejiang Chen, Zehua Ma, Jie Zhang 0073, Huanyu Bian, Han Fang 0004, Weiming Zhang 0001, Nenghai Yu |
ACM Multimedia | 4 |
| 2023 | SPSW: Database Watermarking Based on Fake Tuples and Sparse Priority StrategyabstractDatabases play a crucial role in storing and managing vast amounts of data in various organizations and industries. Yet the risk of database leakage poses a significant threat to data privacy and security. To trace the source of database leakage, researchers have proposed many database watermarking schemes. Among them, fake-tuples-based database watermarking shows great potential as it does not modify the original data of the database, ensuring the seamless usability of the watermarked database. However, the existing fake-tuple-based database watermarking schemes need to insert a large number of fake tuples for the embedding of each watermark bit, resulting in low watermark transparency. Therefore, we propose a novel database watermarking scheme based on fake tuples and sparse priority strategy, named SPSW, which achieves the same watermark capacity with a lower number of inserted fake tuples compared to the existing embedding strategy. Specifically, for a database about to be watermarked, we prioritize embedding the sparsest watermark sequence, i.e., the sequence containing the most ‘0’ bits among the currently available watermark sequences. For each bit in the sparse watermark sequence, when it is set to ‘1’, SPSW will embed the corresponding set of fake tuples into the database. Otherwise, no modifications will be made to the database. Through theoretical analysis, the proposed sparse priority strategy not only improves transparency but also enhances the robustness of the watermark. The comparative experimental results with other database watermarking schemes further validate the superior performance of the proposed SPSW, aligning with the theoretical analysis. Zhiwen Ren, Zehua Ma, Weiming Zhang 0001, Nenghai Yu |
TrustCom | 2 |
| 2023 | Language universal font watermarking with multiple cross-media robustness
Weiming Zhang 0001, Han Fang 0004, Zehua Ma, Nenghai Yu |
Signal Process. | 4 |
| 2023 | Encoded Feature Enhancement in Watermarking Network for Distortion in Real ScenesabstractDeep-learning based watermarking framework has been extensively studied recently. The main structure of such framework is an encoder, a noise layer and a decoder. By training with different distortion sets in the noise layer, the whole network can realize different robustness. However, such framework has a huge drawback that the noise layer must be differentiable, otherwise it cannot be trained end-to-end. But for practical use, much distortions are non-differentiable, so such framework cannot be applied. To address such limitations, this paper propose a triple-phase watermarking framework for practical distortions. The proposed framework consists of three phases including a noise-free initial phase, a mask-guided frequency enhancement phase and an adversarial-training phase. Phase 1 aims to initialize an encoder to embed watermark with high visual quality and a decoder to extract the watermark. In order to generate high quality watermarked image, we design the just noticeable difference (JND)-mask image loss in phase 1 to guide the encoder. At phase 2, based on the investigation of the encoded features and distortions, we propose a mask-guided frequency enhancement algorithm to enhance the encoded feature which ensures the survival of such features after distortion, so that there will be enough features to be learned in phase 3. And phase 3 aims to train a stronger decoder to extract the watermark from the image after practical distortions. The combination of these 3 phases can well handle the non-differentiable problems and make the whole network trainable. Various experiments indicate the superior performance of the proposed scheme in the view of traditional differentiable image processing distortion robustness and practical non-differentiable distortion robustness. Han Fang 0004, Zhaoyang Jia, Hang Zhou 0007, Zehua Ma, Weiming Zhang 0001 |
IEEE Trans. Multim. | 4 |
| 2023 | OAcode: Overall Aesthetic 2D Barcode on ScreenabstractNowadays, two-dimensional (2D) barcodes have been widely used in various domains. And a series of aesthetic 2D barcode schemes have been proposed to improve the visual quality and readability of 2D barcodes for better integration with marketing materials. Yet we believe that the existing aesthetic 2D barcode schemes arepartiallyaesthetic because they only beautify the data area but retain the position detection patterns with the blackwhite appearance of traditional 2D barcode schemes. Thus, in this paper, we propose the firstoverallaesthetic 2D barcode scheme, called OAcode, in which the position detection pattern is canceled. Its detection process is based on the pre-designed symmetrical data area of OAcode, whose symmetry could be used as the calibration signal to restore the perspective transformation in the barcode scanning process. Moreover, an enhanced demodulation method is proposed to resist the lens distortion common in the camera-shooting process. The experimental results illustrate that when 5×5cmOAcode is captured with a resolution of 720×1280 pixels, at the screen-camera distance of 10cmand the angle less or equal to 25°, OAcode has 100% detection rate and 99.5% demodulation accuracy. For 10×10cmOAcode, it could be extracted by consumer-grade mobile phones at a distance of 90cmwith around 90% accuracy. Zehua Ma, Han Fang 0004, Weiming Zhang 0001, Nenghai Yu |
IEEE Trans. Multim. | 1 |
| 2022 | Tracing Text Provenance via Context-Aware Lexical SubstitutionabstractText content created by humans or language models is often stolen or misused by adversaries. Tracing text provenance can help claim the ownership of text content or identify the malicious users who distribute misleading content like machine-generated fake news. There have been some attempts to achieve this, mainly based on watermarking techniques. Specifically, traditional text watermarking methods embed watermarks by slightly altering text format like line spacing and font, which, however, are fragile to cross-media transmissions like OCR. Considering this, natural language watermarking methods represent watermarks by replacing words in original sentences with synonyms from handcrafted lexical resources (e.g., WordNet), but they do not consider the substitution’s impact on the overall sentence's meaning. Recently, a transformer-based network was proposed to embed watermarks by modifying the unobtrusive words (e.g., function words), which also impair the sentence's logical and semantic coherence. Besides, one well-trained network fails on other different types of text content. To address the limitations mentioned above, we propose a natural language watermarking scheme based on context-aware lexical substitution (LS). Specifically, we employ BERT to suggest LS candidates by inferring the semantic relatedness between the candidates and the original sentence. Based on this, a selection strategy in terms of synchronicity and substitutability is further designed to test whether a word is exactly suitable for carrying the watermark signal. Extensive experiments demonstrate that, under both objective and subjective metrics, our watermarking scheme can well preserve the semantic integrity of original sentences and has a better transferability than existing methods. Besides, the proposed LS approach outperforms the state-of-the-art approach on the Stanford Word Substitution Benchmark. Jie Zhang 0073, Kejiang Chen, Weiming Zhang 0001, Zehua Ma, Nenghai Yu |
AAAI | 5 |
| 2022 | Font Watermarking Network for Text ImagesabstractWith the popularization of online services, a lot of text watermarking algorithm have been proposed to protect the digital documents. However, most of them require extra manual design or text samantic modification. This paper proposes an end-to-end font watermarking network which is capable of automatic watermark embedding and extraction without changing the text content. In our scheme, a watermark is generated by changing the font attributes slightly and then embedded as a tiny perturbation. And we extract the watermark by detecting the attribute values of the font image. Experimental results highlight the superiority of the proposed watermarking scheme in terms of imperceptibility and speed comparing to the existing work. Weiming Zhang 0001, Han Fang 0004, Zehua Ma, Nenghai Yu |
ICIP | 5 |
| 2022 | PIMoG: An Effective Screen-shooting Noise-Layer Simulation for Deep-Learning-Based Watermarking NetworkabstractWith the omnipresence of camera phone and digital display, capturing digitally displayed image with camera phone are getting widely practiced. In the context of watermarking, this brings forth the issue of screen-shooting robustness. The key to acquiring screen-shooting robustness is designing a good noise layer that could represent screen-shooting distortions in a deep-learning-based watermarking framework. However, it is very difficult to quantitatively formulate the screen-shooting distortion since the screen-shooting process is too complex. In order to design an effective noise layer for screen-shooting robustness, we propose new insight in this paper, that is, it is not necessary to quantitatively simulate the overall procedure in the screen-shooting noise layer, only including the most influenced distortions is enough to generate an effective noise layer with strong robustness. To verify this insight, we propose a screen-shooting noise layer dubbed PIMoG. Specifically, we summarize the most influenced distortions of screen-shooting process into three parts (p erspective distortion, i llumination distortion and mo iré distortion) and further simulate them in a differentiable way. For the rest distortion, we utilize the G aussian noise to approximate the main part of them. As a result, the whole network can be trained end-to-end with such noise layer. Extensive experiments illustrate the superior performance of the proposed PIMoG noise layer. In addition to the noise layer design, we also propose a gradient mask-guided image loss and an edge mask-guided image loss to further improve the robustness and invisibility of the whole network respectively. Based on the proposed loss and PIMoG noise layer, the whole framework outperforms the SOTA watermarking method with at least 5% in extraction accuracy and achieves more than 97% accuracy in different screen-shooting conditions. Han Fang 0004, Zhaoyang Jia, Zehua Ma, Ee-Chien Chang, Weiming Zhang 0001 |
ACM Multimedia | 3 |
| 2022 | TERA: Screen-to-Camera Image Code With Transparency, Efficiency, Robustness and AdaptabilityabstractWith the rapid development of digital devices, the issue of how to transmit information among different devices with multimedia carriers has drawn much attention from the research community. This paper focuses on the important user scenario of “screen-to-camera information transmission”. Along this direction, image coding-based techniques have been shown to be the most popular and effective methods in the past decades. However, after careful study, we find that none of the existing methods can satisfy the four important properties simultaneously, i.e.,high transparency,high embedding efficiency,strong transmission robustnessandhigh adaptability to device types. This is mainly because these properties are contradictory with each other. In this paper, we thus propose a screen-to-camera image code dubbed “TERA” (transparency,efficiency,robustness andadaptability), which makes it possible to circumvent the contradiction among the above four properties for the first time. Generally, TERA adopts the color decomposition principle to ensure the visual quality and the superposition-based scheme to ensure embedding efficiency. BCH-coding-based information arrangement and a powerful attention-guided information decoding network are further designed to guarantee the robustness and adaptability. Through extensive experiments, the superiority and broad applications of our method are demonstrated. Han Fang 0004, Dongdong Chen 0001, Zehua Ma, Honggu Liu, Wenbo Zhou 0004, Weiming Zhang 0001, Nenghai Yu |
IEEE Trans. Multim. | 4 |
| 2021 | Deep Template-Based WatermarkingabstractTraditional watermarking algorithms have been extensively studied. As an important type of watermarking schemes, template-based approaches maintain a very high embedding rate. In such scheme, the message is often represented by some dedicatedly designed templates, and then the message embedding process is carried out by additive operation with the templates and the host image. To resist potential distortions, these templates often need to contain some special statistical features so that they can be successfully recovered at the extracting side. But in existing methods, most of these features are handcrafted and too simple, thus making them not robust enough to resist serious distortions unless very strong and obvious templates are used. Inspired by the powerful feature learning capacity of deep neural network, we propose the first deep template-based watermarking algorithm in this paper. Specifically, at the embedding side, we first design two new templates for message embedding and locating, which is achieved by leveraging the special properties of human visual system, i.e., insensitivity to specific chrominance components, the proximity principle and the oblique effect. At the extracting side, we propose a novel two-stage deep neural network, which consists of an auxiliary enhancing sub-network and a classification sub-network. Thanks to the power of deep neural networks, our method achieves both digital editing resilience and camera shooting resilience based on typical application scenarios. Through extensive experiments, we demonstrate that the proposed method can achieve much better robustness than existing methods while guaranteeing the original visual quality. Han Fang 0004, Dongdong Chen 0001, Qidong Huang, Jie Zhang 0073, Zehua Ma, Weiming Zhang 0001, Nenghai Yu |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2021 | Local Geometric Distortions Resilient Watermarking Scheme Based on SymmetryabstractAs an efficient watermark attack method, geometric distortions destroy the synchronization between the watermark encoder and decoder. Local geometric distortion is a considerable challenge in the watermarking field. Although many geometric distortion resilient watermarking schemes have been proposed, few perform well against local geometric distortions, such as random bending attacks (RBAs). To address this problem, this paper proposes a novel watermark synchronization process and a corresponding watermarking scheme. In our scheme, the watermark bits are represented by random patterns. The message is encoded to obtain a watermark unit, and the watermark unit is flipped to generate a symmetrical watermark. Then, the symmetrical watermark is additively embedded into the spatial domain of the host image. In watermark extraction, we first obtain the theoretical mean-square error minimized estimation of the watermark. Then, an autoconvolution function is applied to this estimation to detect the symmetry and obtain a watermark unit map. According to this map, the watermark can be accurately synchronized, and then extraction can be performed. Experimental results demonstrate the excellent robustness of the proposed watermarking scheme to local geometric distortions, global geometric distortions, common image processing operations, and some kinds of combined attacks. Zehua Ma, Weiming Zhang 0001, Han Fang 0004, Xiaoyi Dong, Linfeng Geng, Nenghai Yu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | A survey of energy-saving technologies in cloud data centers
Huiwen Cheng, Bo Liu 0045, Weiwei Lin 0001, Zehua Ma, Keqin Li 0001, Ching-Hsien Hsu |
J. Supercomput. | 4 |
| 2020 | Robust Superpixel-Guided Attentional Adversarial AttackabstractDeep Neural Networks are vulnerable to adversarial samples, which can fool classifiers by adding small perturbations onto the original image. Since the pioneering optimization-based adversarial attack method, many following methods have been proposed in the past several years. However most of these methods add perturbations in a "pixel-wise" and "global" way. Firstly, because of the contradiction between the local smoothness of natural images and the noisy property of these adversarial perturbations, this "pixel-wise" way makes these methods not robust to image processing based defense methods and steganalysis based detection methods. Secondly, we find adding perturbations to the background is less useful than to the salient object, thus the "global" way is also not optimal. Based on these two considerations, we propose the first robust superpixel-guided attentional adversarial attack method. Specifically, the adversarial perturbations are only added to the salient regions and guaranteed to be same within each superpixel. Through extensive experiments, we demonstrate our method can preserve the attack ability even in this highly constrained modification space. More importantly, compared to existing methods, it is significantly more robust to image processing based defense and steganalysis based detection. Xiaoyi Dong, Jiangfan Han, Dongdong Chen 0001, Huanyu Bian, Zehua Ma, Hongsheng Li 0001, Xiaogang Wang 0001, Weiming Zhang 0001, Nenghai Yu |
CVPR | 6 |
| 2020 | A Camera Shooting Resilient Watermarking Scheme for Underpainting DocumentsabstractThis paper designs a novel underpainting based camera shooting resilient (CSR) document watermarking algorithm for dealing with the leak source tracking problem. By applying such algorithm, we can extract the authentication watermark information from the candid photographs. The watermarked underpainting contains three significant properties. 1) Inconspicuousness. The watermarked underpainting is inconspicuous and it will not easily be maliciously attacked. 2) Robustness. We propose DCT-based watermark embedding algorithm and distortion compensation based extracting algorithm, which make the watermark robust to camera shooting process. 3) Autocorrelation. We design the flip-based method to arrange the watermarked underpainting. So that a complete watermark region can be accurately located even if part of the document is recorded. Compared with previous watermarking algorithms, the proposed scheme guaranteed content independent embedding as well as the robustness to the camera shooting process. Besides, the proposed scheme satisfies the accuracy of extraction even when the captured document is incomplete. Han Fang 0004, Weiming Zhang 0001, Zehua Ma, Hang Zhou 0007, Hao Cui 0004, Nenghai Yu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2019 | A robust image watermarking scheme in DCT domain based on adaptive texture direction quantization
Han Fang 0004, Hang Zhou 0007, Zehua Ma, Weiming Zhang 0001, Nenghai Yu |
Multim. Tools Appl. | 3 |