Jizhe Zhou 0001

dblp:172/4712-1 · also Ji-Zhe Zhou 0001 · DBLP profile ↗
← Back
33ranked-venue papers
4as first author
31since 2021 · last 2026
0000-0002-2447-1806ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 1 first-author · 18 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 3 first-author · 16 since 2021Security and privacy · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 FGM-HD: Boosting Generation Diversity of Fractal Generative Models through Hausdorff Dimension Induction
abstract
Improving the diversity of generated results while maintaining high visual quality remains a significant challenge in image generation tasks. Fractal Generative Models (FGMs) are efficient in generating high-quality images, but their inherent self-similarity limits the diversity of output images. To address this issue, we propose a novel approach based on the Hausdorff Dimension (HD), a widely recognized concept in fractal geometry used to quantify structural complexity, which aids in enhancing the diversity of generated outputs. To incorporate HD into FGM, we propose a learnable HD estimation method that predicts HD directly from image embeddings, addressing computational cost concerns. However, simply introducing HD into a hybrid loss is insufficient to enhance diversity in FGMs due to: 1) degradation of image quality, and 2) limited improvement in generation diversity. To this end, during training, we adopt an HD-based loss with a monotonic momentum-driven scheduling strategy to progressively optimize the hyperparameters, obtaining optimal diversity without sacrificing visual quality. Moreover, during inference, we employ HD-guided rejection sampling to select geometrically richer outputs. Extensive experiments on the ImageNet dataset demonstrate that our FGM-HD framework yields a 39% improvement in output diversity compared to vanilla FGMs, while preserving comparable image quality. To our knowledge, this is the very first work introducing HD into FGM. Our method effectively enhances the diversity of generated outputs while offering a principled theoretical contribution to FGM development.
Yuanpei Zhao, Jizhe Zhou 0001, Mao Li 0001
AAAI3
2026 IO-RAE: Information-Obfuscation Reversible Adversarial Example for Audio Privacy Protection
abstract
The rapid advancements in artificial intelligence have significantly accelerated the adoption of speech recognition technology, leading to its widespread integration across various applications. However, this surge in usage also highlights a critical issue: audio data is highly vulnerable to unauthorized exposure and analysis, posing significant privacy risks for businesses and individuals. This paper introduces an Information-Obfuscation Reversible Adversarial Example (IO-RAE) framework, the pioneering method designed to safeguard audio privacy using reversible adversarial examples. IO-RAE leverages large language models to generate misleading yet contextually coherent content, effectively preventing unauthorized eavesdropping by humans and Automatic Speech Recognition (ASR) systems. Additionally, we propose the Cumulative Signal Attack technique, which mitigates high-frequency noise and enhances attack efficacy by targeting low-frequency signals. Our approach ensures the protection of audio data without degrading its quality or usability. Experimental evaluations demonstrate the superiority of our method, achieving a targeted misguidance rate of 96.5% and a remarkable 100% untargeted misguidance rate in obfuscating target keywords across multiple ASR models, including a commercial black-box system from Google. Furthermore, the quality of the recovered audio, measured by the Perceptual Evaluation of Speech Quality score, reached 4.45, comparable to high-quality original recordings. Notably, the recovered audio processed by ASR systems exhibited an error rate of 0%, indicating nearly lossless recovery. These results highlight the practical applicability and effectiveness of our IO-RAE framework in protecting sensitive audio privacy.
Xia Du, Jizhe Zhou 0001, Qizhen Xu, Zheng Lin 0001, Chi-Man Pun
AAAI4
2026 Towards Generalized Image Manipulation Localization via Score-based Model
abstract
With the rapid evolution of synthetic media, Image Manipulation Localization (IML) has emerged as a critical component in multimedia forensics for ensuring the integrity of digital content. However, generalization remains a core challenge, as existing discriminative methods typically learn a fixed decision boundary that tends to overfit to specific training artifacts and fails to adapt to unseen manipulation types. To address this, we propose DiffIML, a novel framework that introduces score-based generative modeling to IML. Diverging from the direct estimation of hard boundaries, DiffIML approximates the score function, the gradient of the log-likelihood, to capture the intrinsic geometric topology of mask distributions. This paradigm leverages structural priors to iteratively recover coherent masks from noise, thereby circumventing the brittleness associated with discriminative models. Under this formulation, diffusion models serve as an effective numerical solver for the learned score function. To ensure practicality, we respectively resolve the efficiency and stability bottlenecks of standard diffusion by: (1) utilizing a Lightweight Mask-Specific VAE for fast latent-space process and a decoupled architecture with a lightweight denoising UNet, (2) edge supervision and error prior to mitigate error accumulation during sampling. Extensive experiments of two distinct protocols on eight non-generative and three generative benchmarks demonstrate that DiffIML consistently outperforms state-of-the-art methods, yielding remarkable generalization improvements on diverse unseen datasets. The code will be publicly available.
Yunfei Wang 0001, Tianxin Xu, Jizhe Zhou 0001
ICMR7
2026 An effective UNet using feature interaction and fusion for organ segmentation in medical image
Xiaolin Gou, Chuanlin Liao, Jizhe Zhou 0001, Fengshuo Ye, Yi Lin 0006
Eng. Appl. Artif. Intell.3
2026 Bones to identity: Generative contrastive fusion for cross-modality medical person identification from skeletal data
Chaoqun Niu, Dongdong Chen 0004, Jizhe Zhou 0001, Jian Wang 0124, Quanhui Liu, Caiyang Yu, Wei Ju 0001, Jiancheng Lv 0001
Pattern Recognit.3
2026 Defensive Adversarial CAPTCHA: A Semantics-Driven Framework for Natural Adversarial Example Generation
abstract
Traditional CAPTCHA (Completely Automated Public Turing Test to Tell Computers and Humans Apart) schemes are increasingly vulnerable to automated attacks powered by deep neural networks (DNNs). Existing adversarial attack methods often rely on the original image characteristics, resulting in distortions that hinder human interpretation and limit their applicability in scenarios where no initial input images are available. To address these challenges, we propose the Unsourced Adversarial CAPTCHA (DAC), a novel framework that generates high-fidelity adversarial examples guided by attacker-specified semantics information. Leveraging a Large Language Model (LLM), DAC enhances CAPTCHA diversity and enriches the semantic information. To address various application scenarios, we examine the white-box targeted attack scenario and the black-box untargeted attack scenario. For target attacks, we introduce two latent noise variables that are alternately guided in the diffusion step to achieve robust inversion. The synergy between gradient guidance and latent variable optimization achieved in this way ensures that the generated adversarial examples not only accurately align with the target conditions but also achieve optimal performance in terms of distributional consistency and attack effectiveness. In untargeted attacks, especially for black-box scenarios, we introduce bi-path unsourced adversarial CAPTCHA (BP-DAC), a two-step optimization strategy employing multimodal gradients and bi-path optimization for efficient misclassification. Experiments show that the defensive adversarial CAPTCHA generated by BP-DAC is able to defend against most of the unknown models, and the generated CAPTCHA is indistinguishable to both humans and DNNs.
Xia Du, Jizhe Zhou 0001, Zheng Lin 0001, Chi-Man Pun, Cong Wu 0003, Tao Li 0001, Zhe Chen 0015, Wei Ni 0001, Jun Luo 0001
IEEE Trans. Dependable Secur. Comput.3
2026 PASK: Sparse Framework for Crafting Natural Adversarial Example
abstract
As audio adversarial attacks continue to evolve, Automatic Speech Recognition (ASR) models have emerged as a significant target. Traditional audio attack methods often focus on minimizing perturbation magnitude and frequency, overlooking the importance of perturbation location. However, certain audio regions hold lower importance for ASR models, making attacks on these regions less effective and more perceptible as noise. Additionally, the human ear perceives noise differently depending on its placement within the audio sequence, with noise in silent segments being more noticeable. To address these challenges, this paper proposes Pitch Sparse Audio Attack (PASK), an innovative framework designed to enhance adversarial imperceptibility through sparse perturbations. PASK introduces two key techniques: Pitch Mapping, which provides a strategic starting point for perturbation, and an adaptive grouped selective mask that achieves targeted sparsity, focusing perturbations on high-impact audio regions. Experimental results demonstrate that PASK outperforms existing methods in both effectiveness and imperceptibility. Furthermore, a human study confirms that silent-segment perturbations are more easily detected, underscoring the perceptual advantages of our approach.
Xia Du, Jizhe Zhou 0001, Qizhen Xu, Chi-Man Pun
IEEE Trans. Multim.4
2025 Can We Get Rid of Handcrafted Feature Extractors? SparseViT: Nonsemantics-Centered, Parameter-Efficient Image Manipulation Localization Through Spare-Coding Transformer
abstract
Non-semantic features or semantic-agnostic features, which are irrelevant to image context but sensitive to image manipulations, are recognized as evidential to Image Manipulation Localization (IML). Since manual labels are impossible, existing works rely on handcrafted methods to extract non-semantic features. Handcrafted non-semantic features jeopardize IML model's generalization ability in unseen or complex scenarios. Therefore, for IML, the elephant in the room is: How to adaptively extract non-semantic features? Non-semantic features are context-irrelevant and manipulation-sensitive. That is, within an image, they are consistent across patches unless manipulation occurs. Then, spare and discrete interactions among image patches are sufficient for extracting non-semantic features. However, image semantics vary drastically on different patches, requiring dense and continuous interactions among image patches for learning semantic representations. Hence, in this paper, we propose a Sparse Vision Transformer (SparseViT), which reformulates the dense, global self-attention in ViT into a sparse, discrete manner. Such sparse self-attention breaks image semantics and forces SparseViT to adaptively extract non-semantic features for images. Besides, compared with existing IML models, the sparse self-attention mechanism largely reduced the model size (max 80% in FLOPs), achieving stunning parameter efficiency and computation reduction. Extensive experiments demonstrate that, without any handcrafted feature extractors, SparseViT is superior in both generalization and efficiency across benchmark datasets.
Xiaochen Ma 0001, Xuekang Zhu, Chaoqun Niu, Zeyu Lei, Jizhe Zhou 0001
AAAI6
2025 Mesoscopic Insights: Orchestrating Multi-Scale & Hybrid Architecture for Image Manipulation Localization
abstract
The mesoscopic level serves as a bridge between the macroscopic and microscopic worlds, addressing gaps overlooked by both. Image manipulation localization (IML), a crucial technique to pursue truth from fake images, has long relied on low-level (microscopic-level) traces. However, in practice, most tampering aims to deceive the audience by altering image semantics. As a result, manipulation commonly occurs at the object level (macroscopic level), which is equally important as microscopic traces. Therefore, integrating these two levels into the mesoscopic level presents a new perspective for IML research. Inspired by this, our paper explores how to simultaneously construct mesoscopic representations of micro and macro information for IML and introduces the Mesorch architecture to orchestrate both. Specifically, this architecture i) combines Transformers and CNNs in parallel, with Transformers extracting macro information and CNNs capturing micro details, and ii) explores across different scales, assessing micro and macro information seamlessly. Additionally, based on the Mesorch architecture, the paper introduces two baseline models aimed at solving IML tasks through mesoscopic representation. Extensive experiments across four datasets have demonstrated that our models surpass the current state-of-the-art in terms of performance, computational complexity, and robustness.
Xuekang Zhu, Xiaochen Ma 0001, Zhuohang Jiang, Xiwen Wang 0002, Zeyu Lei, Wentao Feng, Chi-Man Pun, Jizhe Zhou 0001
AAAI10
2025 M3: Manipulation Mask Manufacturer for Arbitrary-Scale Super-Resolution Mask
Xiaochen Ma 0001, Xuekang Zhu, Bingkui Tong, Zeyu Lei, Jizhe Zhou 0001
CVM (1)8
2025 Enhancing Diffusion Model Stability for Image Restoration via Gradient Management
abstract
Diffusion models have shown remarkable promise for image restoration by leveraging powerful priors. Prominent methods typically frame the restoration problem within a Bayesian inference framework, which iteratively combines a denoising step with a likelihood guidance step. However, the interactions between these two components in the generation process remain underexplored. In this paper, we analyze the underlying gradient dynamics of these components and identify significant instabilities. Specifically, we demonstrate conflicts between the prior and likelihood gradient directions, alongside temporal fluctuations in the likelihood gradient itself. We show that these instabilities disrupt the generative process and compromise restoration performance. To address these issues, we propose Stabilized Progressive Gradient Diffusion (SPGD), a novel gradient management technique. SPGD integrates two synergistic components: (1) a progressive likelihood warm-up strategy to mitigate gradient conflicts; and (2) adaptive directional momentum (ADM) smoothing to reduce fluctuations in the likelihood gradient. Extensive experiments across diverse restoration tasks demonstrate that SPGD significantly enhances generation stability, leading to state-of-the-art performance in quantitative metrics and visually superior results. Code is available at https://github.com/74587887/SPGD.
Hongjie Wu, Mingqin Zhang, Linchao He, Jizhe Zhou 0001, Jiancheng Lv 0001
ACM Multimedia4
2025 Saliency Guided Optimization of Diffusion Latents
Xiwen Wang 0002, Jizhe Zhou 0001, Xuekang Zhu, Mao Li 0001
MMM (3)2
2025 ForensicHub: A Unified Benchmark & Codebase for All-Domain Fake Image Detection and Localization
abstract
The field of Fake Image Detection and Localization (FIDL) is highly fragmented, encompassing four domains: deepfake detection (Deepfake), image manipulation detection and localization (IMDL), artificial intelligence-generated image detection (AIGC), and document image manipulation localization (Doc). Although individual benchmarks exist in some domains, a unified benchmark for all domains in FIDL remains blank. The absence of a unified benchmark results in significant domain silos, where each domain independently constructs its datasets, models, and evaluation protocols without interoperability, preventing cross-domain comparisons and hindering the development of the entire FIDL field. To close the domain silo barrier, we propose ForensicHub, the first unified benchmark & codebase for all-domain fake image detection and localization. Considering drastic variations on dataset, model, and evaluation configurations across all domains, as well as the scarcity of open-sourced baseline models and the lack of individual benchmarks in some domains, ForensicHub: i) proposes a modular and configuration-driven architecture that decomposes forensic pipelines into interchangeable components across datasets, transforms, models, and evaluators, allowing flexible composition across all domains; ii) fully implements 10 baseline models (3 of which are reproduced from scratch), 6 backbones, 2 new benchmarks for AIGC and Doc, and integrates 2 existing benchmarks of DeepfakeBench and IMDLBenCo through an adapter-based design; iii) establishes an image forensic fusion protocol evaluation mechanism that supports unified training and testing of diverse forensic models across tasks; iv) conducts indepth analysis based on the ForensicHub, offering 8 key actionable insights into FIDL model architecture, dataset characteristics, and evaluation standards. Specifically, ForensicHub includes 4 forensic tasks, 23 datasets, 42 baseline models, 6 backbones, 11 GPU-accelerated pixel- and image-level evaluation metrics, and realizes 16 kinds of cross-domain evaluations. ForensicHub represents a significant leap forward in breaking the domain silos in the FIDL field and inspiring future breakthroughs. Code is available at: https://github.com/scu-zjz/ForensicHub.
Xuekang Zhu, Xiaochen Ma 0001, Chenfan Qu, Kaiwen Feng, Chi-Man Pun, Jizhe Zhou 0001
NeurIPS9
2025 FAcupoint: The first dense facial acupoint localization dataset and baselines
Jizhe Zhou 0001, Hongyu Yang 0002, Yi Lin 0006
Expert Syst. Appl.3
2025 DP-TRAE: A Dual-Phase Merging Transferable Reversible Adversarial Example for Image Privacy Protection
abstract
In the field of digital security, Reversible Adversarial Examples (RAE) combine adversarial attacks with reversible data hiding techniques to effectively protect sensitive data and prevent unauthorized analysis by malicious Deep Neural Networks (DNNs). However, existing RAE techniques primarily focus on white-box attacks, lacking a comprehensive evaluation of their effectiveness in black-box scenarios. This limitation impedes their broader deployment in complex, dynamic environments. Furthermore, traditional black-box attacks are often characterized by poor transferability and high query costs, significantly limiting their practical applicability. To address these challenges, we propose the Dual-Phase Merging Transferable Reversible Attack method, which generates highly transferable initial adversarial perturbations in a white-box model and employs a memory-augmented black-box strategy to effectively mislead target models. Experimental results demonstrate the superiority of our approach, achieving a 99.0% attack success rate and 100% recovery rate in black-box scenarios with the DN-121 target model and 1000 attack iterations, highlighting its robustness in privacy protection. Moreover, we successfully implemented a black-box attack on a commercial model, further substantiating the potential of this approach for practical use.
Xia Du, Jizhe Zhou 0001, Chi-Man Pun, Zheng Lin 0001, Cong Wu 0003, Zhe Chen 0015, Jun Luo 0001
IEEE Trans. Dependable Secur. Comput.3
2024 Neural Boneprint: Person Identification from Bones Using Generative Contrastive Deep Learning
abstract
Forensic person identification is of paramount importance in accidents and criminal investigations. Existing methods based on soft tissue or DNA can be unavailable if the body is badly decomposed, white-ossified, or charred. However, bones last a long time. This raises a natural question: can we learn to identify a person using bone data? We present a novel feature of bones called Neural Boneprint for personal identification. In particular, we exploit the thoracic skeletal data including chest radiographs (CXRs) and computed tomography (CT) images enhanced by the volume rendering technique (VRT) as an example to explore the availability of the neural boneprint. We then represent the neural boneprint as a joint latent embedding of VRT images and CXRs through a bidirectional cross-modality translation and contrastive learning. Preliminary experimental results on real skeletal data demonstrate the effectiveness of the Neural Boneprint for identification. We hope that this approach will provide a promising alternative for challenging forensic cases where conventional methods are limited. The code is available at https://github.com/CheltonNiu/Neural-Boneprint.git.
Chaoqun Niu, Dongdong Chen 0004, Jizhe Zhou 0001, Jian Wang 0124, Quanhui Liu, Jiancheng Lv 0001
ACM Multimedia3
2024 Diffusion Posterior Proximal Sampling for Image Restoration
abstract
Diffusion models have demonstrated remarkable efficacy in generating high-quality samples. Existing diffusion-based image restoration algorithms exploit pre-trained diffusion models to leverage data priors, yet they still preserve elements inherited from the unconditional generation paradigm. These strategies initiate the denoising process with pure white noise and incorporate random noise at each generative step, leading to over-smoothed results. In this paper, we present a refined paradigm for diffusion-based image restoration. Specifically, we opt for a sample consistent with the measurement identity at each generative step, exploiting the sampling selection as an avenue for output stability and enhancement. The number of candidate samples used for selection is adaptively determined based on the signal-to-noise ratio of the timestep. Additionally, we start the restoration process with an initialization combined with the measurement signal, providing supplementary information to better align the generative process. Extensive experimental results and analyses validate that our proposed method significantly enhances image restoration performance while consuming negligible additional computational resources.
Hongjie Wu, Linchao He, Mingqin Zhang, Dongdong Chen 0004, Kunming Luo, Mengting Luo, Jizhe Zhou 0001, Hu Chen 0002, Jiancheng Lv 0001
ACM Multimedia7
2024 DP-RAE: A Dual-Phase Merging Reversible Adversarial Example for Image Privacy Protection
abstract
In digital security, Reversible Adversarial Examples (RAE) blend adversarial attacks with Reversible Data Hiding (RDH) within images to thwart unauthorized access. Traditional RAE methods, however, compromise attack efficiency for the sake of perturbation concealment, diminishing the protective capacity of valuable perturbations and limiting applications to white-box scenarios. This paper proposes a novel Dual-Phase merging Reversible Adversarial Example (DP-RAE) generation framework, combining a heuristic black-box attack and RDH with Grayscale Invariance (RDH-GI) technology. This dual strategy not only evaluates and harnesses the adversarial potential of past perturbations more effectively but also guarantees flawless embedding of perturbation information and complete recovery of the original image. Experimental validation reveals our method's superiority, secured an impressive 96.9% success rate and 100% recovery rate in compromising black-box models. In particular, it achieved a 90% misdirection rate against commercial models under a constrained number of queries. This marks the first successful attempt at targeted black-box reversible adversarial attacks for commercial recognition models. This achievement highlights our framework's capability to enhance security measures without sacrificing attack performance. Moreover, our attack framework is flexible, allowing the interchangeable use of different attack and RDH modules to meet advanced technological requirements.
Xia Du, Jizhe Zhou 0001, Chi-Man Pun, Qizhen Xu
ACM Multimedia3
2024 IMDL-BenCo: A Comprehensive Benchmark and Codebase for Image Manipulation Detection & Localization
abstract
A comprehensive benchmark is yet to be established in the Image Manipulation Detection & Localization (IMDL) field. The absence of such a benchmark leads to insufficient and misleading model evaluations, severely undermining the development of this field. However, the scarcity of open-sourced baseline models and inconsistent training and evaluation protocols make conducting rigorous experiments and faithful comparisons among IMDL models challenging. To address these challenges, we introduce IMDL-BenCo, the first comprehensive IMDL benchmark and modular codebase. IMDL-BenCo: i) decomposes the IMDL framework into standardized, reusable components and revises the model construction pipeline, improving coding efficiency and customization flexibility; ii) fully implements or incorporates training code for state-of-the-art models to establish a comprehensive IMDL benchmark; and iii) conducts deep analysis based on the established benchmark and codebase, offering new insights into IMDL model architecture, dataset characteristics, and evaluation standards.Specifically, IMDL-BenCo includes common processing algorithms, 8 state-of-the-art IMDL models (1 of which are reproduced from scratch), 2 sets of standard training and evaluation protocols, 15 GPU-accelerated evaluation metrics, and 3 kinds of robustness evaluation. This benchmark and codebase represent a significant leap forward in calibrating the current progress in the IMDL field and inspiring future breakthroughs.Code is available at: https://github.com/scu-zjz/IMDLBenCo
Xiaochen Ma 0001, Xuekang Zhu, Zhuohang Jiang, Bingkui Tong, Zeyu Lei, Chi-Man Pun, Jiancheng Lv 0001, Jizhe Zhou 0001
NeurIPS11
2024 Efficient physical image attacks using adversarial fast autoaugmentation methods
Xia Du, Chi-Man Pun, Jizhe Zhou 0001
Knowl. Based Syst.3
2024 Shunting at Arbitrary Feature Levels via Spatial Disentanglement: Toward Selective Image Translation
abstract
The past few years have witnessed considerable efforts devoted to translating images from one domain to another, mainly aiming at editing global style. Here, we focus on a more general case, selective image translation (SLIT), under an unsupervised setting. SLIT essentially operates through a shunt mechanism that involves learning gates to manipulate only the contents of interest (CoIs), which can be either local or global, while leaving the irrelevant parts unchanged. Existing methods typically rely on a flawed implicit assumption that CoIs are separable at arbitrary levels, ignoring the entangled nature of DNN representations. This leads to unwanted changes and learning inefficiency. In this work, we revisit SLIT from an information-theoretical perspective and introduce a novel framework, which equips two opposite forces to disentangle the visual features. One force encourages independence between spatial locations on the features, while the other force unites multiple locations to form a "block" that jointly characterizes an instance or attribute that a single location may not independently characterize. Importantly, this disentanglement paradigm can be applied to visual features of any layer, enabling shunting at arbitrary feature levels, which is a significant advantage not explored in existing works. Our approach has undergone extensive evaluation and analysis, confirming its effectiveness in significantly outperforming the state-of-the-art baselines.
Jian Wang 0124, Jizhe Zhou 0001, Jiancheng Lv 0001
IEEE Trans. Neural Networks Learn. Syst.4
2023 Pre-training-free Image Manipulation Localization through Non-Mutually Exclusive Contrastive Learning
abstract
Deep Image Manipulation Localization (IML) models suffer from training data insufficiency and thus heavily rely on pre-training. We argue that contrastive learning is more suitable to tackle the data insufficiency problem for IML. Crafting mutually exclusive positives and negatives is the prerequisite for contrastive learning. However, when adopting contrastive learning in IML, we encounter three categories of image patches: tampered, authentic, and contour patches. Tampered and authentic patches are naturally mutually exclusive, but contour patches containing both tampered and authentic pixels are non-mutually exclusive to them. Simply abnegating these contour patches results in a drastic performance loss since contour patches are decisive to the learning outcomes. Hence, we propose the Nonmutually exclusive Contrastive Learning (NCL) framework to rescue conventional contrastive learning from the above dilemma. In NCL, to cope with the non-mutually exclusivity, we first establish a pivot structure with dual branches to constantly switch the role of contour patches between positives and negatives while training. Then, we devise a pivot-consistent loss to avoid spatial corruption caused by the role-switching process. In this manner, NCL both inherits the self-supervised merits to address the data insufficiency and retains a high manipulation localization accuracy. Extensive experiments verify that our NCL achieves state-of-the-art performance on all five benchmarks without any pre-training and is more robust on unseen real-life samples. https://github.com/Knightzjz/NCL-IML.
Jizhe Zhou 0001, Xiaochen Ma 0001, Xia Du, Ahmed Y. Al Hammadi, Wentao Feng
ICCV1
2023 M2ATS: A Real-world Multimodal Air Traffic Situation Benchmark Dataset and Beyond
abstract
Air Traffic Control (ATC) is a complicated, time-evolving, and real-time procedure to direct flight operations in a safer and ordered manner. Although enormous data storages are available during air traffic operations for over 40 years, data-driven intelligent application in aviation is still an emerging task due to the safety-critical issue. With the prevalence of the Next Generation ATC system, artificial intelligence (AI) -empowered research topics are attracting increasing attention from both industrial and academic domains and a high-quality dataset naturally becomes the prerequisite for such practices. However, almost all ATC-related datasets are only unimodal for certain tasks, which fails to comprehensively illustrate the traffic situation to further support real-world studies. To address this gap, a multimodal air traffic situation (M2ATS) dataset is constructed to advance AI-related research in the ATC domain, including airspace information, flight plan, trajectory, and speech. M2ATS covers 10362 flights ATC situation data, involving 110000+ utterances (104 hours) with diversity golden text annotations, 16 intents, and 51 slots. Considering the real-world ATC requirements, a total of 10 multimedia-related tasks (24 baselines) are designed to validate the proposed dataset, covering automatic speech recognition, natural language processing, and spatial-temporal data processing. New ATC-related metrics corresponding to ATC applications are proposed in addition to the common metrics to evaluate task performance. Extensive experiment results demonstrate that the selective baselines can achieve designed tasks on this new dataset, and further investigations are also required to address task and data specificities. It is believed that the proposed new dataset is a new practice to advance AI applications to an industrial scene, which not only promotes ATC-related applications but also provides diverse research topics in the common multimedia community.
Dongyue Guo, Yi Lin 0006, Xuehang You, Zhongping Yang, Jizhe Zhou 0001, Bo Yang 0063, Jianwei Zhang 0013, Shasha Hu
ACM Multimedia5
2023 Multi-view Adaptive Bone Activation from Chest X-Ray with Conditional Adversarial Nets
Chaoqun Niu, Jian Wang 0124, Jizhe Zhou 0001, Tu Xiong, Huili Guo, Weibo Liang, Jiancheng Lv 0001
MMM (2)4
2023 Exploring the first-move balance point of Go-Moku based on reinforcement learning and Monte Carlo tree search
Pengsen Liu, Jizhe Zhou 0001, Jiancheng Lv 0001
Knowl. Based Syst.2
2023 Revisiting the transferability of adversarial examples via source-agnostic adversarial feature inducing method
Yatie Xiao, Jizhe Zhou 0001, Kongyang Chen, Zhenbang Liu
Pattern Recognit.2
2023 Towards Recognition for Radio-Echo Speech in Air Traffic Control: Dataset and a Contrastive Learning Approach
abstract
In the air traffic control (ATC) domain, automatic speech recognition (ASR) suffers from radio speech echo, which cannot be addressed by existing echo cancellation due to auditory-oriented optimization and poor generalization ability caused by volatile radio transmission. In this work, a contrastive learning-based framework is proposed to tackle the radio-echo speech for the ASR task based on convolution networks with multiple paths and recurrent neural networks. 1) By analyzing the communication mechanism of the ATC speech, a novel transmission method is designed to collect clean and noisy speech samples (with the same texts) via a bypass device in a real-world ATC environment. 2) To enhance the model capacity, a temporal and frequency attention block is innovatively designed to guide the model to focus on informative frames and frequencies, aiming at learning shared representations between the clean and noisy speech signals with the same texts. 3) By incorporating contrastive loss, the proposed approach is implemented by a multi-objective optimization, in which the loss weights are dynamically determined to enhance the ASR performance in a learnable manner. With the proposed transmission method, a real-world dataset is collected and annotated to validate the proposed approach. Experimental results demonstrate that the proposed approach outperforms other comparative baselines with different technical frameworks, achieving a 6.76% character error rate on the test dataset. Most importantly, all the proposed improvements are confirmed by designed experiments, in which contrastive learning with learnable multi-objective loss weights contributes to the primary performance improvement.
Yi Lin 0006, Xincheng Yu, Zichen Zhang 0020, Dongyue Guo, Jizhe Zhou 0001
IEEE ACM Trans. Audio Speech Lang. Process.6
2023 DHI-GAN: Improving Dental-Based Human Identification Using Generative Adversarial Networks
abstract
In this work, a novel semisupervised framework is proposed to tackle the small-sample problem of dental-based human identification (DHI), achieving enhanced performance via a "classifying while generating" paradigm. A generative adversarial network (GAN), called the DHI-GAN, is presented to implement this idea, in which an extra classifier is also dedicatedly proposed to achieve an efficient training procedure. Considering the complex specificities of this problem, except for the noise input of the generator, an identity embedding-guided architecture is proposed to retain informative features for each individual. A parallel spatial and channel fusion attention block is innovatively designed to encourage the model to learn discriminative and informative features by focusing on different regional details and abstract concepts. The attention block is also widely applied to the overall classifier to learn identity-dependent information. A loss combination of the ArcFace and focal loss is utilized to address the small-sample problem. Two parameters are proposed to control the generated samples that are fed into the classifier during the optimization procedure. The proposed DHI-GAN framework is finally validated on a real-world dataset, and the experimental results demonstrate that it outperforms other baselines, achieving a 92.5% top-one accuracy rate. Most importantly, the proposed GAN-based semisupervised training strategy is able to reduce the required number of training samples (individuals) and can also be incorporated into other classification models. Our code will be available at https://github.com/sculyi/MedicalImages/.
Yi Lin 0006, Jianwei Zhang 0013, Jizhe Zhou 0001, Peixi Liao, Hu Chen 0002, Zhenhua Deng, Yi Zhang 0018
IEEE Trans. Neural Networks Learn. Syst.4
2022 Abstract Rule Learning for Paraphrase Generation
abstract
In early years, paraphrase generation typically adopts rule-based methods, which are interpretable and able to make global transformations to the original sentence. But they struggle to produce fluent paraphrases. Recently, deep neural networks have shown impressive performances in generating paraphrases. However, the current neural models are black boxes and are prone to make local modifications to the inputs. In this work, we combine these two approaches into RULER, a novel approach that performs abstract rule learning for paraphrasing. The key idea is to explicitly learn generalizable rules that could enhance the paraphrase generation process of neural networks. In RULER, we first propose a rule generalizability metric to guide the model to generate rules underlying the paraphrasing. Then, we leverage neural networks to generate paraphrases by refining the sentences transformed by the learned rules. Extensive experimental results demonstrate the superiority of RULER over previous state-of-the-art methods in terms of paraphrase quality, generalization ability and interpretability.
Xianggen Liu, Wenqiang Lei, Jiancheng Lv 0001, Jizhe Zhou 0001
IJCAI4
2022 Separation Inference: A Unified Framework for Word Segmentation in East Asian Languages
abstract
Existing methods consider Word Segmentation (WS) as sequence tagging. Each tag indicates the position of the current character in a segment. The exactness of the position for any non-boundaries character is unnecessary. Any incorrect inner prediction reduces model performance. The position information restricts tag-to-tag transition. Thereby, extra context information and the Conditional Random Field (CRF) network are desired to control unreasonable tag transition. To steer away from the implicit restriction, we propose the Separation(Sp)-Adhesion(Ad), which targets straight on the essential character-to-character connections, to tackle the WS task directly. Merely bigram that is specially tailored for “Sp-Ad” is required and considered as the processing unit to identify the connection states of every two adjacent characters. The elimination of the position restriction makes the model independent of the CRF layer which is widely adopted to revise unreasonable tags. Therefore, CRF can then be substituted with a classification network. We construct the Separation Inference (SpIn) framework based on the bigram features and softmax classification network to tackle the WS task. SpIn significantly reduces the inference complexity, dispels extra context information, and boosts the accuracy of the WS task. Besides its effectiveness in Chinese Word Segmentation, performance boosts on Japanese and Korean Word Segmentation further prove SpIn is universal for East Asian Languages. Moreover, our extensive experiments also verify the cross-domain effectiveness of SpIn by attaining state-of-the-art performances in the benchmark tests of in-domain and cross-domain Chinese Word Segmentation.
Yu Tong 0003, Jingzhi Guo, Jizhe Zhou 0001
IEEE ACM Trans. Audio Speech Lang. Process.3
2021 Personal Privacy Protection via Irrelevant Faces Tracking and Pixelation in Video Live Streaming
abstract
To date, the privacy-protection intended pixelation tasks are still labor-intensive and yet to be studied. With the prevailing of video live streaming, establishing an online face pixelation mechanism during streaming is an urgency. In this paper, we develop a new method called Face Pixelation in Video Live Streaming (FPVLS) to generate automatic personal privacy filtering during unconstrained streaming activities. Simply applying multi-face trackers will encounter problems in target drifting, computing efficiency, and over-pixelation. Therefore, for fast and accurate pixelation of irrelevant people's faces, FPVLS is organized in a frame-to-video structure of two core stages. On individual frames, FPVLS utilizes image-based face detection and embedding networks to yield face vectors. In the raw trajectories generation stage, the proposed Positioned Incremental Affinity Propagation (PIAP) clustering algorithm leverages face vectors and positioned information to quickly associate the same person's faces across frames. Such frame-wise accumulated raw trajectories are likely to be intermittent and unreliable on video level. Hence, we further introduce the trajectory refinement stage that merges a proposal network with the two-sample test based on the Empirical Likelihood Ratio (ELR) statistic to refine the raw trajectories. A Gaussian filter is laid on the refined trajectories for final pixelation. On the video live streaming dataset we collected, FPVLS obtains satisfying accuracy, real-time efficiency, and contains the over-pixelation problems.
Jizhe Zhou 0001, Chi-Man Pun
IEEE Trans. Inf. Forensics Secur.1
2020 Privacy-sensitive Objects Pixelation for Live Video Streaming
abstract
With the prevailing of live video streaming, establishing an online pixelation method for privacy-sensitive objects is an urgency. Caused by the inaccurate detection of privacy-sensitive objects, simply migrating the tracking-by-detection structure applied in offline pixelation into the online form will incur problems in target initialization, drifting, and over-pixelation. To cope with the inevitable but impacting detection issue, we propose a novel Privacy-sensitive Objects Pixelation (PsOP) framework for automatic personal privacy filtering during live video streaming. Leveraging pre-trained detection networks, our PsOP is extendable to any potential privacy-sensitive objects pixelation. Employing the embedding networks and the proposed Positioned Incremental Affinity Propagation (PIAP) clustering algorithm as the backbone, our PsOP unifies the pixelation of discriminating and indiscriminating pixelation objects through trajectories generation. In addition to the pixelation accuracy boosting, experiment results on the streaming video data we built show that the proposed PsOP can significantly reduce the over-pixelation ratio in privacy-sensitive object pixelation.
Jizhe Zhou 0001, Chi-Man Pun, Yu Tong 0003
ACM Multimedia1
2020 News Image Steganography: A Novel Architecture Facilitates the Fake News Identification
abstract
A larger portion of fake news quotes untampered images from other sources with ulterior motives rather than conducting image forgery. Such elaborate engraftments keep the inconsistency between images and text reports stealthy, thereby, palm off the spurious for the genuine. This paper proposes an architecture named News Image Steganography (NIS) to reveal the aforementioned inconsistency through image steganography based on GAN. Extractive summarization about a news image is generated based on its source texts, and a learned steganographic algorithm encodes and decodes the summarization of the image in a manner that approaches perceptual invisibility. Once an encoded image is quoted, its source summarization can be decoded and further presented as the ground truth to verify the quoting news. The pairwise encoder and decoder endow images of the capability to carry along their imperceptible summarization. Our NIS reveals the underlying inconsistency, thereby, according to our experiments and investigations, contributes to the identification accuracy of fake news that engrafts untampered images.
Jizhe Zhou 0001, Chi-Man Pun, Yu Tong 0003
VCIP1