Jiwei Zhang 0007

dblp:61/7812-7 · DBLP profile ↗
← Back
25ranked-venue papers
6as first author
22since 2021 · last 2026
0000-0002-8910-2382ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 2 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 1 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Collaborative Transformers with Multi-Level Forensic Attention for Image Manipulation Localization
abstract
The proliferation of the tampered images on social media can pose serious societal risks, influencing public opinion and causing panic. Image Manipulation Localization technique has advanced to address this, but some methods focus on microscopic traces, overlooking macroscopic semantics that deceive viewers. To address this problem, we propose a novel Image Manipulation Localization framework called Collaborative Transformers (Co-Transformers), designed to fully explore and utilize the collaborative information between macroscopic semantics and microscopic traces. This framework is based on two Vision Transformer variants. The first variant captures the semantic logic of the image. The second variant delves into microscopic tampering traces. By dynamically fusing these two complementary features, the framework enables interaction between macroscopic semantic inconsistencies and microscopic abnormal traces, effectively coordinating their relationship in the latent space. Furthermore, we introduce a new Multi-Level Forensic Attention (MLF-Attention) mechanism to enhance the model's ability to extract various tampered traces, this mechanism can be integrated into our framework. Compared with existing methods, our proposed framework achieves state-of-the-art results in localization accuracy and shows good robustness against various attacks.
Jiwei Zhang 0007, Wenbo Feng, Feifei Kou, Shaozhang Niu
AAAI1
2026 MusicRec: Multi-modal Semantic-Enhanced Identifier with Collaborative Signals for Generative Recommendation
abstract
Generative recommendation as a new paradigm is influencing the current development of recommender systems. It aims to assign identifiers that capture richer semantic and collaborative information to items, and subsequently predict item identifiers via autoregressive generation using Large Language Models (LLMs). Existing approaches primarily tokenize item text into codebooks with preserved semantic IDs through RQ-VAE, or separately tokenize different modality features of items. However, existing tokenization methods face two major challenges: (1) Learning decoupled multi-modal features limits the quality of the semantic representation. (2) Ignoring collaborative signals from interaction history limits the comprehensiveness of identifiers. To address these limitations, we propose a multi-modal semantic-enhanced identifier with collaborative signals for generative recommendation, named MusicRec. In MusicRec, we propose a tokenization approach based on shared-specific modal fusion, enabling the generated identifiers to preserve semantic information more comprehensively from all modalities. In addition, we incorporate collaborative signals from user interactions to guide identifier generation, preserving collaborative patterns in the semantic representation space. Extensive experiments on three public datasets demonstrate that MusicRec achieves state-of-the-art performance compared to existing baseline methods.
Yuqiu Zhao, Lei Shi 0030, Yan Zhong 0001, Feifei Kou, Pengfei Zhang 0010, Jiwei Zhang 0007, Mingying Xu
AAAI6
2025 InpDiffusion: Image Inpainting Localization via Conditional Diffusion Models
abstract
As artificial intelligence advances rapidly, particularly with the advent of GANs and diffusion models, the accuracy of Image Inpainting Localization (IIL) has become increasingly challenging. Current IIL methods face two main challenges: a tendency towards overconfidence, leading to incorrect predictions; and difficulty in detecting subtle tampering boundaries in inpainted images. In response, we propose a new paradigm that treats IIL as a conditional mask generation task utilizing diffusion models. Our method, InpDiffusion, utilizes the denoising process enhanced by the integration of image semantic conditions to progressively refine predictions. During denoising, we employ edge conditions and introduce a novel edge supervision strategy to enhance the model's perception of edge details in inpainted objects. Balancing the diffusion model's stochastic sampling with edge supervision of tampered image regions mitigates the risk of incorrect predictions from overconfidence and prevents the loss of subtle boundaries that can result from overly stochastic processes. Furthermore, we propose an innovative Dual-stream Multi-scale Feature Extractor (DMFE) for extracting multi-scale features, enhancing feature representation by considering both semantic and edge conditions of the inpainted images. Extensive experiments across challenging datasets demonstrate that the InpDiffusion significantly outperforms existing state-of-the-art methods in IIL tasks, while also showcasing excellent generalization capabilities and robustness.
Shaozhang Niu, Qixian Hao, Jiwei Zhang 0007
AAAI4
2025 DcDsDiff: Dual-Conditional and Dual-Stream Diffusion Model for Generative Image Tampering Localization
abstract
Generative Image Tampering (GIT), due to its high diversity and realism, poses a significant challenge to traditional image tampering localization techniques. Consequently, this paper introduces a denoising diffusion probabilistic model-based DcDsDiff, which comprises a Dual-View Conditional Network (DVCN) and a Dual-Stream Denoising Network (DSDN). DVCN provides clues about the tampered areas. It extracts tampering features in the high-frequency view and integrates them with spatial domain features using attention mechanisms. DSDN jointly generates mask image and detail image, enhancing the generalization capability of the model against new tampering forms through iterative denoising. A multi-stream interaction mechanism enables the two generative tasks to promote each other, prompting the model to generate localization results that are rich in detail and complete. Experiments show that DcDsDiff outperforms mainstream methods in accurate localization, generalization, extensibility, and robustness. Code page: https://github.com/QixianHao/DcDsDiff-and-GIT10K.
Qixian Hao, Shaozhang Niu, Jiwei Zhang 0007
IJCAI3
2025 ForgDiffuser: General Image Forgery Localization with Diffusion Models
abstract
Current general image forgery localization (GIFL) methods confront two main challenges: decoder overconffdence causing misidentiffcation of the authentic regions or incomplete predicted masks, and limited accuracy in localizing forgery details. Recently, diffusion models have excelled as dominant approach for generative models, particularly effective in capturing complex scene details. However, their potential for GIFL remains underexplored. Therefore, we propose a GIFL framework named ForgDiffuser with diffusion models. The core of ForgDiffuser lies in leveraging diffusion models conditioned on the forgery image to efffciently generate the segmentation mask for tampered regions. Speciffcally, we introduce the attentionguided module (AGM) to aggregate and enhance image feature representations. Meanwhile, we design the boundary-driven module (BDM) with edge supervision to improve the localization accuracy of boundary details. Additionally, the probabilistic modeling and stochastic sampling mechanisms of diffusion models effectively alleviate the overconffdence issue commonly observed in traditional decoders. Experiments on six benchmark datasets demonstrate that ForgDiffuser outperforms existing mainstream GIFL methods in both localization accuracy and robustness, especially under challenging manipulation conditions.
Mengxi Wang, Shaozhang Niu, Jiwei Zhang 0007
IJCAI3
2025 OSTAR: Optimized Statistical Text-classifier with Adversarial Resistance
abstract
The advancements in generative models and the real-world attack of machine-generated text(MGT) create a demand for more robust detection methods. The existing MGT detection methods for adversarial environments primarily consist of manually designed statistical-based methods and fine-tuned classifier-based approaches. Statistical-based methods extract intrinsic features but suffer from rigid decision boundaries vulnerable to adaptive attacks, while fine-tuned classifiers achieve outstanding performance at the cost of overfitting to superficial textual feature. We argue that the key to detection in current adversarial environments lies in how to extract intrinsic invariant features and ensure that the classifier possesses dynamic adaptability. In that case, we propose OSTAR, a novel MGT detection framework designed for adversarial environments which composed of a statistical enhanced classifier and a Multi-Faceted Contrastive Learning(MFCL). In the classifier aspect, our Multi-Dimensional Statistical Profiling (MDSP) module extracts intrinsic difference between human and machine texts, complementing classifiers with useful stable features. In the model optimization aspect, the MFCL strategy enhances robustness by contrasting feature variations before and after text attacks, jointly optimizing statistical feature mapping and baseline pre-trained models. Experimental results on three public datasets under various adversarial scenarios demonstrate that our framework outperforms existing MGT detection methods, achieving state-of-the-art performance and robust against attacks.The code is available at https://github.com/BUPT-SN/OSTAR.
Yuhan Yao 0001, Feifei Kou, Lei Shi 0030, Zhongbao Zhang, Suguo Zhu, Jiwei Zhang 0007, Lirong Qiu, Hai-Sheng Li 0002
NeurIPS7
2025 UFG-Net: Uncertainty and frequency guided network for image forgery localization
Qixian Hao, Shaozhang Niu, Jiwei Zhang 0007
Neurocomputing4
2025 DualFocus GAN for Robust Watermarking in Transportation Cyber-Physical Systems
abstract
With the advancement of Transportation Cyber-Physical Systems (TCPS), information security has become increasingly critical. Invisible watermarking, which ensures reliable information traceability without compromising carrier quality, holds significant potential for TCPS. However, achieving high robustness in the real world while maintaining imperceptibility is a challenge. To address this, we propose DGWW (Dual-discriminator GAN-based WaveFusion Watermarking), a novel invisible watermarking method that balances robustness and imperceptibility. The GAN-based approach is well suited for TCPS, as it enables adaptive watermark embedding aligned with the dynamic and heterogeneous nature of transportation data, effectively handling diverse noise conditions and data types. DGWW integrates a WaveFusion Encoding Module, a Dual-Focus Discriminator, and a contrastive learning-based optimization strategy to enhance watermark embedding without degrading robustness. These components leverage multi-frequency information, assess local and global impacts on image quality, and guide model optimization. Experimental results show that DGWW outperforms state-of-the-art methods in visual quality and robustness under various noise conditions, offering a robust and scalable solution for image watermarking in TCPS environments. By maintaining data usability and strong resistance to noise attacks, DGWW advances digital watermarking in intelligent transportation systems.
Feifei Kou, Yuhan Yao 0001, Jideng Han, Hai-Sheng Li 0002, Jiwei Zhang 0007
IEEE Trans. Intell. Transp. Syst.7
2025 Intelligent Connected Vehicle Data Privacy and Security Transaction Sharing System Based on Blockchain
abstract
With the widespread application of Transportation Cyber Physical Systems (T-CPS), increasingly intelligent and interconnected vehicles are conducting extensive transportation activities. Compared with traditional transportation equipment, they integrate advanced information functions such as data collection, terminal communication, real-time computing, and remote coordination, which can generate and collect a large amount of real traffic data. The enormous value of these traffic data can be released through market-oriented transactions. Blockchain technology can support the transmission and collaborative control of information T-CPS, while protecting the privacy and data security of intelligent connected vehicles. This article proposes a blockchain based data trading system aimed at simplifying the transaction flow of traffic data for intelligent connected vehicle owners, while maintaining fairness, privacy, and sustainable market development. Our work introduces two key innovations: a two-stage availability verification process that reduces transaction costs while enhancing data reliability, and an efficient encryption confirmation mechanism that ensures privacy and security for data providers and buyers throughout the entire transaction lifecycle. Finally, we demonstrate the feasibility and overall performance of our system through comprehensive analysis including security and reliability assessment, market behavior analysis, and computational complexity modeling, as well as practical experiments based on the Ethereum blockchain network. The evaluation results indicate that this scheme can provide privacy and security data transaction services at lower transaction costs.
Jiwei Zhang 0007, Yufei Tu, Ziang Sun, Shaozhang Niu
IEEE Trans. Intell. Transp. Syst.1
2025 TEPR-Net: Image Inpainting Localization Network via Texture Enhancement and Progressive Refinement
abstract
To counter the security threats posed by the realism of image inpainting generated through diffusion models and GANs, in this paper, we propose a texture enhancement and progressive refinement network (TEPR-Net) for image inpainting localization (IIL). The IIL task is divided into two phases: coarse and fine locating. In the coarse locating phase, we utilize an anomaly texture encoder to capture tampering traces in textures, employ a texture–context feature interaction strategy to effectively integrate texture features with contextual features, and utilize a pixel-level contrastive learning strategy to enhance feature clustering and model generalization. In the fine locating phase, we first enhance the receptive field features in the frequency domain by transforming the features and separately enhancing the low- and high-frequency components. Then, we utilize the coarse localization result to augment the model's sensitivity to tampered regions. Additionally, we introduce a progressive edge distribution guidance and reconstruction strategy that progressively refines the edges of the tampered regions at each level, ultimately generating refined localization results. To support the research and evaluation of the IIL task, we create the Inpaint32K dataset, which is characterized by its large scale, diversity, comprehensiveness, high quality, and authenticity. Finally, extensive experiments demonstrate that TEPR-Net has significant advantages in terms of localization performance, generalizability, extensibility, and robustness.
Qixian Hao, Haoliang Cui, Jiwei Zhang 0007, Shaozhang Niu
IEEE Trans. Multim.4
2025 Potential Features Fusion Network for Multimodal Fake News Detection
abstract
With the popularization of social networks, fake news is also widely and rapidly spreading, which poses a great threat to the Internet. Therefore, how to detect fake news automatically and efficiently has become an urgent problem to be solved. However, the existing approaches mostly focus on the explicit features (images and text) and deep fusions, without considering potential features such as text emotion and image category. To find a solution to this issue, we propose a Potential Features Fusion Network (PFFN), which models the explicit and potential features at the same time. To exploit the potential image features, we introduce a mixture of experts structure to process the news image separately, which can best use the relationships between the news image category and fake news detection. Besides, we also extract emotion features as potential text features and fuse them with explicit text features. Finally, we establish an attention-based feature fusion network to fuse the potential features with the explicit features, which can obtain a multimodal fusion feature of a piece of news and thus further improve the performance. We make experiments on four public datasets (Weibo16, Weibo19, Twitter, and PolitiFact); the results compared with the baseline approaches demonstrate that our PFFN has a better performance. Our code is available at https://github.com/Wang-bupt/PFFN
Feifei Kou, Bingwei Wang, Hai-Sheng Li 0002, Chuangying Zhu, Lei Shi 0030, Jiwei Zhang 0007, Limei Qi
ACM Trans. Multim. Comput. Commun. Appl.6
2024 EMF-Net: An edge-guided multi-feature fusion network for text manipulation detection
Ruyong Ren, Qixian Hao, Feng Gu 0001, Shaozhang Niu, Jiwei Zhang 0007, Maosen Wang
Expert Syst. Appl.5
2024 EC-Net: General image tampering localization network based on edge distribution guidance and contrastive learning
Qixian Hao, Ruyong Ren, Shaozhang Niu, Jiwei Zhang 0007, Maosen Wang
Knowl. Based Syst.5
2024 UGEE-Net: Uncertainty-guided and edge-enhanced network for image splicing localization
Qixian Hao, Ruyong Ren, Shaozhang Niu, Maosen Wang, Jiwei Zhang 0007
Neural Networks6
2024 MFI-Net: Multi-Feature Fusion Identification Networks for Artificial Intelligence Manipulation
abstract
Tampered images can easily be used for illegal activities, such as spreading rumors, economic fraud, fabricating false news, and illegally obtaining experience benefits, etc. With the improvement and development of artificial intelligence (AI), image manipulation technology has also been further improved, more and more retouching software in daily life adopts AI technology. So far, there is no AI-based tampered dataset. To address this challenge, we propose a dataset-IPM15K. It utilizes the most advanced image processing technology and contains a total of 150,00 doctored vital images. This dataset also could serve as a catalyst for progressing many vision tasks, e.g., localization, segmentation, and alpha-matting, etc. Additionally, we propose an effective multi-feature fusion identification network (MFI-Net) to identify these challenging images. Our model consists of four modules: the detail extraction module (DEM), which utilizes different sizes of convolutions and perceptual fields to extract more valuable information of tampered locations; the multi-branch attention fusion module (MAFM), which fully exploits contextual information of different levels to capture subtle traces of tampering; the feature decoder component (FDC), which combines fused features to identify tampered regions; and the detail enhancement block (DEB), which continues to supplement the detailed information of the detected regions. Extensive experiments on three public datasets and the proposed dataset show that MFI-Net outperforms various state-of-the-art (SOTA) manipulation detection baselines.
Ruyong Ren, Qixian Hao, Shaozhang Niu, Keyang Xiong, Jiwei Zhang 0007, Maosen Wang
IEEE Trans. Circuits Syst. Video Technol.5
2023 Multi-scale attention context-aware network for detection and localization of image splicing
Ruyong Ren, Shaozhang Niu, Junfeng Jin, Jiwei Zhang 0007, Hua Ren
Appl. Intell.4
2023 An intelligent digital twin system for paper manufacturing in the paper industry
Jiwei Zhang 0007, Haoliang Cui, Andy L. Yang, Feng Gu 0001, Chengjie Shi, Shaozhang Niu
Expert Syst. Appl.1
2022 Adversarial Learning Enhancement for 3D Human Pose and Shape Estimation
abstract
Adversarial learning plays an important role in recovering 3D human pose and shape from monocular videos. However, the effectiveness of this process is not often considered. Hence we aim to improve the performance of adversarial learning in 3D human pose and shape estimation. The performance of adversarial learning is mainly influenced by two parts: generator and discriminator. For the generator, we utilize temporal information on a deeper level by adding an attention-based temporal encoder in generator to model the time series of features, which contributes to a more appropriate data representation for pose and shape regression. For the discriminator, we innovatively make use of human skeleton topology information when extracting features from the estimation results. To realize this, we base the discriminator’s design on the graph convolution network. In addition, to eliminate the jitter in the estimation results, we design a rotation disentangled smoothing module to process the estimated rotation parameters. We did adequate experiments on public in-the-wild datasets 3DPW and MPI-INF-3DHP. On both datasets, our method achieves higher accuracy and lower acceleration error compared with previous methods.
Yidian Sun, Jiwei Zhang 0007, Wendong Wang 0003
ICASSP2
2022 A face recognition algorithm based on feature fusion
abstract
Summary In the process of building a smart city, face recognition can be applied to the transformation of enterprises, communities, and parks. The combination of building security system and face recognition technology can improve the security experience of enterprises and citizens through the solution of hardware and software integration. Face recognition is still facing the challenges of illumination, occlusion, and attitude change in the actual application process. In addition, the end‐to‐end convolutional neural networks (CNN) seldom make use of the hierarchical feature of the network. So, we propose a hierarchy feature fusion method for face recognition, which uses supervisory information to learn shallow and deep facial features. The features are fused to enhance the recognition accuracy of face recognition against illumination and occlusion. The method is applied to the transformation of the visual geometry group network and Lightened CNN. The face recognition experiments are carried out using the hierarchy network. Our method has achieved good recognition results in the labeled faces in the wild (LFW) and AR face databases.
Jiwei Zhang 0007, Xiaodan Yan, Zelei Cheng, Xueqi Shen
Concurr. Comput. Pract. Exp.1
2022 AntiConcealer: Reliable Detection of Adversary Concealed Behaviors in EdgeAI-Assisted IoT
abstract
Internet of Things (IoT) is one of the rapidly developing technologies today that attract huge real-world applications. However, the reality is that IoT is easily vulnerable to numerous types of cyberattacks and anomalies. Detecting them is becoming increasingly challenging day by day due to limitations with IoT devices and threat intelligence. Particularly, one of the most challenging problems is to detect the existence of malicious adversaries that continuously adapt or conceal their behaviors in IoT to hide their actions and to make the IoT security protocol ineffective. In this article, we study this problem at the IoT device level that can be a great idea to avoid potential attacks. We presentAntiConcealer, an edge-aided IoT framework, and propose an edge artificial intelligence-enabled approach (EdgeAI) for detecting adversary concealed behaviors in the IoT. We first develop an adversary behavior model and use this to identify mid-attack temporal patterns by learning the multivariate Hawkes process (MHP), a kind of point process as a random and finite series of events (e.g., behaviors) controlled by a probabilistic model. Naturally, learning MHP processed on EdgeAI reveals the influence of the concealed behaviors of adversaries in the IoT. These concealed behaviors are then grouped using a nonnegative weighted influence matrix. To observe the performance of theAntiConcealerframework through evaluation, we employ honeypots integrated with edge servers and verify the usability and reliability of adversary behavioral identification.
Jiwei Zhang 0007, Md. Zakirul Alam Bhuiyan, Yang Xu 0013, Tian Wang 0001, Xuesong Xu, Thaier Hayajneh, Faiza Khan
IEEE Internet Things J.1
2022 Multi-attribute adaptive aggregation transformer for vehicle re-identification
Jiaming Pei, Mingpeng Zhu, Jiwei Zhang 0007, Jinhai Li 0002
Inf. Process. Manag.4
2022 Trustworthy Target Tracking With Collaborative Deep Reinforcement Learning in EdgeAI-Aided IoT
abstract
Mobile target tracking with artificial intelligence (AI) approaches such as deep reinforcement learning (DRL) in edge-assisted Internet of Things (Edge-IoT) platform can be promising. In this article, we proposeDRLTrack, a framework for target tracking with a collaborative DRL called C-DRL in Edge-IoT with the aim to obtain two major objectives: high quality of tracking (QoT) and resource-efficient network performance. InDRLTrack, a huge number of IoT devices are employed to collect data about a target of interest. One or two edge devices in the network coordinate with a group of IoT devices and collaboratively detect the target by using the C-DRL approach and form an area around the target by the group of IoT devices. To maintain such an area during the tracking time, we employ a deep Q-network to track the target from one group to another. An EdgeAI sitting on the top of the edge devices has the control of the C-DRL approach during tracking and can identify a sequence of tracks.DRLTrackis said to betrustworthyas it shows trustworthy performance in terms of QoT, dynamic environments, and even under certain cyberattacks. We validate the performance ofDRLTrackconsidering the objectives through simulations and it demonstrates superior performance compared with existing work.
Jiwei Zhang 0007, Md. Zakirul Alam Bhuiyan, Yang Xu 0013, Amit Kumar Singh 0001, D. Frank Hsu
IEEE Trans. Ind. Informatics1
2020 MRGAN: a generative adversarial networks model for global mosaic removal
abstract
In this study, the authors introduce a novel deep generative adversarial networks (GANs) model for global mosaic removal. The methods used in the proposed study consist of GANs model and a novel algorithm for maintaining and repairing (MR) images. The conventional mosaic removal algorithms all employ the correlation between the inserted pixel and its neighbouring pixels, which have a limited effect on the local mosaic removal but do not work well for the global mosaic removal. To respond to this difficulty, the authors introduce an MRGAN model with two novel parsing networks. Unlike previous GANs, the MR algorithm is used to calculate the pixel loss and content loss. The experimental comparison results show that the proposed MRGAN model has achieved leading results for the global mosaic removal task.
Zhiyi Cao, Shaozhang Niu, Jiwei Zhang 0007
IET Image Process.3
2019 Fast generative adversarial networks model for masked image restoration
abstract
The conventional masked image restoration algorithms all utilise the correlation between the masked region and its neighbouring pixels, which does not work well for the larger masked image. The latest research utilises Generative Adversarial Networks (GANs) model to generate a better result for the larger masked image but does not work well for the complex masked region. To get a better result for the complex masked region, the authors propose a novel fast GANs model for masked image restoration. The method used in authors’ research is based on GANs model and fast marching method (FMM). The authors trained an FMMGAN model which consists of a neighbouring network, a generator network, a discriminator network, and two parsing networks. A large number of experimental results on two open datasets show that the proposed model performs well for masked image restoration.
Zhiyi Cao, Shaozhang Niu, Jiwei Zhang 0007
IET Image Process.3
2019 Generative adversarial networks model for visible watermark removal
abstract
Previously visible watermark removal algorithms required the location of known watermarks. A corresponding removal algorithm is then proposed based on the location and the feature of the watermark. If the location of the watermark is random or the watermark has different angles, the watermark removal algorithm will encounter problems. The authors recommend a visible watermark removal algorithm based on generative adversarial networks (GANs) and self‐attention mechanisms. During the training, the authors introduce a GANs model to build mappings between watermarked images and real images. The authors observe that the feature of the watermarked region in different watermarked images is invariant in nature, and the other regions are changed. The self‐attention layer will automatically focus on this invariant feature. Experiments on two public datasets prove that the authors’ model has gained excellent performance. Compared with the other four most competitive watermark removal models, the authors improve the watermark removal rate indicator from 17 to 92%. For the other four evaluation indicators, the authors have improved performance by up to 20%.
Zhiyi Cao, Shaozhang Niu, Jiwei Zhang 0007
IET Image Process.3