Chunjie Ma

dblp:324/3976 · DBLP profile ↗
← Back
21ranked-venue papers
7as first author
21since 2021 · last 2026
0000-0002-6348-671XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 5 first-author · 13 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Security and privacy · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Amplifying Discrepancies: Exploiting Macro and Micro Inconsistencies for Image Manipulation Localization
abstract
The rapid development of image manipulation technologies poses significant challenges to multimedia forensics, especially in accurate localization of manipulated regions. Existing methods often fail to fully explore the intrinsic discrepancies between manipulated and authentic regions, resulting in sub-optimal performance. To address this limitation, we propose the Focus Region Discrepancy Network (FRD-Net), a novel and efficient framework that significantly enhances manipulation localization by amplifying discrepancies at both macro- and micro-levels. Specifically, our proposed Iterative Clustering Module (ICM) groups features into two discriminative clusters and refines representations via backward propagation from cluster centers, improving the distinction between tampered and authentic regions at the macro level. Thereafter, our Differential Progressive Module (DPM) is constructed to capture fine-grained structural inconsistencies within local neighborhoods and integrate them into a Central Difference Convolution, increasing sensitivity to subtle manipulation details at the micro level. Finally, these complementary modules are seamlessly integrated into a compact architecture that achieves a favorable balance between accuracy and efficiency. Extensive experiments on multiple benchmarks demonstrate that FRD-Net consistently surpasses state-of-the-art methods in terms of manipulation localization performance while maintaining a lower computational cost.
Shenghao Chen, Yibo Zhao 0001, Tianyi Wang 0006, Chunjie Ma, Weili Guan, Ming Li 0083, Zan Gao 0001
AAAI4
2026 Multi-domain coupled dynamics modeling and enhanced meta-transfer learning method for few-shot fault diagnosis of axial piston pumps
Chunjie Ma, Junlei Du, Haoran Han, Shengtao Liu
Adv. Eng. Informatics1
2026 Understanding complex queries: Multi-query comprehension network for temporal sentence grounding
Zan Gao 0001, Shengbo Xiong, Yibo Zhao 0001, Chunjie Ma, Tian Gan 0002, Riwei Wang
Expert Syst. Appl.4
2026 RMPSNet: Occluded person re-identification via regional masking and prompt-distribution synergy
Zan Gao 0001, Shuai Xie, Shengxun Wei, Yibo Zhao 0001, Chunjie Ma, Chen Li 0035
Pattern Recognit.5
2026 Prior-knowledge guidance and dual-domain representation refinement for deepfake detection
Zhiyong Cheng 0001, Tianyi Wang 0006, Chunjie Ma, Yibo Zhao 0001, Zan Gao 0001
Pattern Recognit.4
2026 A Novel Multi-View Perception and Shrinkage Aggregation Network for Inharmonious Region Localization
abstract
With the popularity of image editing techniques, synthetic images may have inharmonious regions due to color/illumination differences between the manipulated area and the background. The inharmonious region localization task aims to find these regions, which is crucial for blind image harmonization. Existing methods rely on single-view images and do not fully explore multi-scale fusion, which limits their performance. To address these issues, in this paper, we propose a novel multi-view perception and shrinkage aggregation network (MSANet) for the inharmonious region localization task that fully utilizes multi-view images and multi-scale fusion information and can mine subtle cues between candidate objects and the background. Specifically, we first design a multi-view ensemble encoder to fully perceive the inharmonious regions by multi-view interactive learning and then aggregate the feature representations of inharmonious regions. Moreover, we propose a multi-scale shrinkage fusion decoder, where multi-scale features with multi-view prior information are utilized to aggregate adjacent features, adaptively select high-quality information, reduce background interference and gradually locate inharmonious regions. Extensive experimental results on four public datasets (HDobe5K, HCOCO, HFlickr, and Hday2Night) demonstrate that the proposed MSANet can outperform all the SOTA methods in terms of average F1 and average IoU score, while maintaining a lower computational cost1.
Shenghao Chen, Chunjie Ma, Yibo Zhao 0001, Meng Liu 0006, Yanbing Xue, Zan Gao 0001
IEEE Trans. Circuits Syst. Video Technol.2
2026 Learning Generalizable Representations for Deepfake Detection With Realistic Sample Generation and Dual Augmentation
abstract
Deepfake detection aims to identify manipulated content generated by generative models such as GANs and diffusion models. Although many detection methods have been proposed in recent years, their performance often degrades significantly when the test data includes unknown images or novel forgery types. To address this challenge, we propose a Realistic Sample Generation and Dual Augmentation framework, abbreviated as RSG-DA, to enhance generalization in deepfake detection. The key idea is to explore and expand the forgery feature space in order to learn decision boundaries that can capture diverse forgery patterns. Specifically, we introduce a Dynamic Landmark Diffusion Generator (DLDG) that synthesizes hybrid forgery samples with high visual realism and structural diversity. Additionally, we design a Dual Data Augmentation (DDA) strategy composed of DW-Augmentation and Class-Augmentation, where DW-Augmentation strengthens the representation of authentic image features through multi-scale transformations, while Class-Augmentation enriches the forgery distribution by expanding it with varied manipulations. Finally, we present a Lightweight Generic Forgery Distillation (LGFD) module that integrates the above components into a unified encoder, enabling the learning of robust and transferable forgery representations. Extensive experiments show that our method consistently outperforms state-of-the-art approaches in both intra-dataset and cross-dataset evaluations.
Zan Gao 0001, Xinhai Zhu, Yibo Zhao 0001, Chunjie Ma, Chen Li 0035
IEEE Trans. Dependable Secur. Comput.5
2026 PVLM: Parsing-Aware Vision-Language Model With Dynamic Contrastive Learning for Zero-Shot Deepfake Attribution
abstract
The challenge of tracing the source attribution of forged faces has gained significant attention due to the rapid advancement of generative models. However, existing deepfake attribution (DFA) works primarily focus on the interaction among various domains in vision modality, and other modalities such as texts and face parsing are not fully explored. Besides, they tend to fail to assess the generalization performance of deepfake attributors to unseen advanced generators like diffusion in a fine-grained manner. In this paper, we propose a novelparsing-awarevisionlanguagemodel with dynamic contrastive learning (PVLM) method forzero-shotdeepfakeattribution (ZS-DFA), which facilitates effective and fine-grained traceability to unseen advanced generators. Specifically, we conduct a novel and fine-grained ZS-DFA benchmark to evaluate the attribution performance of deepfake attributors to unseen advanced generators like diffusion. Besides, we propose an innovative PVLM attributor based on the vision-language model to capture general and diverse attribution features. We are motivated by the observation that the preservation of source face attributes in facial images generated by GAN and diffusion models varies significantly. We propose to employ the inherent facial attributes preservation differences to capture face parsing-aware forgery representations. Therefore, we devise a novel parsing encoder to focus on global face attribute embeddings, enabling parsing-guided DFA representation learning via dynamic vision-parsing matching. Additionally, we present a novel deepfake attribution contrastive center loss to pull relevant generators closer and push irrelevant ones away, which can be introduced into DFA models to enhance traceability. Experimental results show that our model exceeds the state-of-the-art on the ZS-DFA benchmark via various protocol evaluations. Codes will be available at GitHub.
Chunjie Ma, Weili Guan, Tian Gan 0002, Zan Gao 0001
IEEE Trans. Dependable Secur. Comput.3
2026 LSGNet: A Local-Pattern Separation and Global-Aware Network for Temporal Action Detection
abstract
Temporal Action Detection aims to localize and classify action instances within untrimmed videos, yet it remains challenging due to background clutter, high intra-class similarity, and varied temporal scales in real-world scenarios. To address these issues, we propose the Local-Pattern Separation and Global-Aware Network (LSGNet) tailored for temporal action localization. Specifically, the core of LSGNet is the Local Pattern Separation Module (LPSM), which explicitly models both consistency and variation patterns of action segments within local temporal windows. Additionally, to capture comprehensive contextual information, we introduce the Global Context-Aware Representation Module (GCRM), which decouples temporal features across multiple granularities and enables robust modeling of long-range dependencies. Finally, we design the Multi-scale Feature Refinement Module (MFRM) to mitigate the degradation of fine-grained information by performing iterative reconstruction across temporal scales, thereby enriching semantic representations and preserving temporal details. Extensive experiments on THUMOS14, ActivityNet1.3, HACS, and EPIC-Kitchens-100 demonstrate the effectiveness of the proposed LSGNet method. Additional ablation studies on the QVHighlights dataset further confirm the generalization capability of LPSM module in video moment retrieval and highlight detection, achieving consistent improvements in retrieval accuracy and localization precision.
Zan Gao 0001, Weilin Yang, Yibo Zhao 0001, Chunjie Ma, Chen Li 0035, Riwei Wang
IEEE Trans. Image Process.4
2026 EFIN: A Novel Enhanced Feature Interaction Network for Temporal Sentence Grounding in Videos
abstract
Temporal sentence grounding in videos (TSGV) is a challenging task that aims to match text queries with semantically relevant segments in untrimmed videos. However, existing methods face limitations in modeling modality features, which constrains the expressive power of candidate moment features. To address this challenge, we propose a novel Enhanced Feature Interaction Network (EFIN) that effectively captures semantic information within each modality and aligns relationships between modalities. Additionally, EFIN enhances the fusion of information between candidate moments and modality features. Specifically, our model begins by extracting modality features to generate candidate moments as priors. Building upon these modality features, we introduce an enhanced feature encoder to extract semantic information within each modality, thereby improving intra-modality feature representation. Simultaneously, the encoder captures alignment relationships between modalities to optimize cross-modality feature representation, enhancing the overall modeling capacity of modality features. Moreover, we design an information fusion module to enrich the comprehension of modality information for candidate moments. Extensive experiments on four benchmark datasets demonstrate the superiority of our proposed EFIN model. Notably, EFIN achieves a maximum performance improvement of approximately 1.67% and 1.91% across different evaluation metrics on TACoS dataset.
Chongxu Hu, Xianbin Wen, Yibo Zhao 0001, Chunjie Ma, Weili Guan, Riwei Wang, Zan Gao 0001
IEEE Trans. Multim.4
2026 A Collaborative Hierarchical Aggregation Network for Weakly Supervised Temporal Action Localization
abstract
Temporal action localization is a fundamental task in video understanding that focuses on classifying and temporally localizing action instances in untrimmed videos. Compared to temporal action localization, the Weakly supervised Temporal Action Localization (WTAL) task presents greater challenges, as its training data lacks detailed information about action boundaries. Existing WTAL methods ignore the complementary relationship between modalities and the dependency between snippets, resulting in inaccurate localization results. To solve these issues, we propose a Collaborative Hierarchical Aggregation Network (CHA-Net). Specifically, we first use a modality complementary module to learn the synergies between modalities. Then, a collaborative enhance module is proposed to remove the information irrelevant to actions in RGB modality. Finally, a hierarchical aggregation module is proposed to capture the complete temporal information of action instances to better mine the temporal dependencies between snippets. Extensive experiments on THUMOS14, ActivityNet1.2, and ActivityNet1.3 datasets demonstrate the effectiveness of our method. Compared with F3-Net (TMM2024, Avg{0.1:0.5}) and SPCC-Net (TMM2024, Avg{0.1:0.7}) on the THUMOS14 dataset, the proposed method can achieve improvements of 3.2% and 2.4%, respectively.
Zan Gao 0001, Xiaoyi Xu, Yibo Zhao 0001, Chunjie Ma, Yanbing Xue, Riwei Wang
ACM Trans. Multim. Comput. Commun. Appl.4
2025 Lightweight Relational Proposal Network with Dual-Branch Distillation for Video Moment Retrieval
Yujia Zhu, Hao Yang 0067, Yibo Zhao 0001, Chunjie Ma, Weili Guan, Zan Gao 0001
ACM Multimedia4
2025 Survey on deep learning-based weakly supervised salient object detection
Lina Du, Chunjie Ma, Huimin Zheng, Xiushan Nie, Zan Gao 0001
Expert Syst. Appl.2
2025 Fine-Grained Modality Relation-Aware Network for Video Moment Retrieval
abstract
Video moment retrieval (VMR) involves localizing video segments semantically aligned with given queries within videos. Despite the development of numerous methods for VMR in recent years, there remains a need to better incorporate fine-grained modality relation-aware information both in intra-modality and cross-modality. To address these challenges, we propose a Fine-grained Modality Relation-Aware Network (FMRN) tailored for the video moment retrieval task. FMRN effectively explores fine-grained modality relation-aware information within text queries, videos, and proposals. Our approach begins with a semantic graph encoder to capture deep semantic relations in intra-modality. Besides, we introduce a novel fine-grained cross-modality interaction module comprising a cross-similarity weighting module, an intra-modality weighting module, and an adaptive fusion module. These components comprehensively exploit fine-grained relation information within intra-modality and cross-modality contexts. Specifically, the cross-similarity weighting module leverages similarities between text queries and video snippets, as well as between videos and query words. The intra-modality weighting module determines the importance of words and snippets, while the adaptive fusion module combines cross-similarity weighting and intra-modality weighting. Additionally, we design a proposal relation module to enhance retrieval by capturing fine-grained proposals-relation information in videos. Extensive experiments demonstrate that the proposed method can outperform all state-of-the-art methods on the TACoS dataset and obtain comparable results on the Charades-STA and ActivityNet-Captions datasets. Compared with MCMN (TCSVT2024) and DPHANet (TMM2024), FMRN can achieve average improvements of 3.61 % and 5.44 % on the TACoS dataset, respectively.
Yibo Zhao 0001, Zan Gao 0001, Chunjie Ma, Weili Guan, Riwei Wang, Shengyong Chen
IEEE Trans. Circuits Syst. Video Technol.3
2025 Multiple Information Prompt Learning for Cloth-Changing Person Re-Identification
abstract
Cloth-changing person re-identification is a subject closer to the real world, which focuses on solving the problem of person re-identification after pedestrians change clothes. The primary challenge in this field is to overcome the complex interplay between intra-class and inter-class variations and to identify features that remain unaffected by changes in appearance. Sufficient data collection for model training would significantly aid in addressing this problem. However, it is challenging to gather diverse datasets in practice. Current methods focus on implicitly learning identity information from the original image or introducing additional auxiliary models, which are largely limited by the quality of the image and the performance of the additional model. To address these issues, inspired by prompt learning, we propose a novel multiple information prompt learning (MIPL) scheme for cloth-changing person ReID, which learns identity robust features through the common prompt guidance of multiple messages. Specifically, the clothing information stripping (CIS) module is designed to decouple the clothing information from the original RGB image features to counteract the influence of clothing appearance. The bio-guided attention (BGA) module is proposed to increase the learning intensity of the model for key information. A dual-length hybrid patch (DHP) module is employed to make the features have diverse coverage to minimize the impact of feature bias. Extensive experiments demonstrate that the proposed method outperforms all state-of-the-art methods on the LTCC, CelebreID, Celeb-reID-light, and CSCC datasets, achieving rank-1 scores of 74.8%, 73.3%, 66.0%, and 88.1%, respectively. When compared to AIM (CVPR23), ACID (TIP23), and SCNet (MM23), MIPL achieves rank-1 improvements of 11.3%, 13.8%, and 7.9%, respectively, on the PRCC dataset.
Shengxun Wei, Zan Gao 0001, Chunjie Ma, Yibo Zhao 0001, Weili Guan, Shengyong Chen
IEEE Trans. Image Process.3
2024 A Coarse to Fine Detection Method for Prohibited Object in X-ray Images Based on Progressive Transformer Decoder
Chunjie Ma, Lina Du, Zan Gao 0001, Li Zhuo 0001, Meng Wang 0001
ACM Multimedia1
2024 MPLA-Net: Multiple Pseudo Label Aggregation Network for Weakly Supervised Video Salient Object Detection
abstract
Weakly Supervised Video Salient Object Detection (WSVSOD) only requires coarse-grained manual annotations, which can achieve a good trade-off between labeling efficiency and detection performance. In this paper, a Multiple Pseudo Label Aggregation Network (MPLA-Net) is proposed for WSVSOD. Firstly, the video frames that can obtain high-quality pseudo labels are selected to generate multiple pseudo labels, so as to avoid the prejudice of the single label. Moreover, the pseudo label with fine edge information is used to generate the Edge Information Map (EIM). Secondly, MPLA-Net is designed to adequately excavate and utilize the comprehensive saliency cues in multiple pseudo labels to improve the detection accuracy, in which ResNet-50 is adopted as the backbone network. Edge loss, pseudo label loss, self-supervised loss and fusion loss are exploited to jointly supervise and optimize the network training to obtain a robust detection model. Experimental results on five benchmark datasets demonstrate that, compared with existing weakly supervised methods, the proposed method can achieve state-of-the-art detection accuracy with less model parameters and higher detection speed. And the detected salient objects have fine boundaries.
Chunjie Ma, Lina Du, Li Zhuo 0001, Jiafeng Li 0001
IEEE Trans. Circuits Syst. Video Technol.1
2023 Occluded prohibited object detection in X-ray images with global Context-aware Multi-Scale feature Aggregation
Chunjie Ma, Li Zhuo 0001, Jiafeng Li 0001, Jing Zhang 0023
Neurocomputing1
2023 Cascade Transformer Decoder Based Occluded Pedestrian Detection With Dynamic Deformable Convolution and Gaussian Projection Channel Attention Mechanism
abstract
Occluded pedestrian detection is very challenging in computer vision, because the pedestrians are frequently occluded by various obstacles or persons, especially in crowded scenarios. In this article, an occluded pedestrian detection method is proposed under a basic DEtection TRansformer (DETR) framework. Firstly, Dynamic Deformable Convolution (DyDC) and Gaussian Projection Channel Attention (GPCA) mechanism are proposed and embedded into the low layer and high layer of ResNet50 respectively, to improve the representation capability of features. Secondly, Cascade Transformer Decoder (CTD) is proposed, which aims to generate high-score queries, avoiding the influence of low-score queries in the decoder stage, further improving the detection accuracy. The proposed method is verified on three challenging datasets, namely CrowdHuman, WiderPerson, and TJU-DHD-pedestrian. The experimental results show that, compared with the state-of-the-art methods, it can obtain a superior detection performance.
Chunjie Ma, Li Zhuo 0001, Jiafeng Li 0001, Jing Zhang 0023
IEEE Trans. Multim.1
2022 Prohibited Object Detection in X-ray Images with Dynamic Deformable Convolution and Adaptive IoU
abstract
Due to the variety and complexity of objects in X-ray images, how to detect the prohibited items automatically and accurately is a challenging problem. In this paper, an X-ray image prohibited object detection method based on Dynamic Deformable Convolution (DyDC) and adaptive Intersection over Union (IoU) is proposed based on Cascade R-CNN framework. The main contributions are as follows. First, DyDC is proposed to cope with the diversity of the prohibited objects in X-ray images and to improve the feature representation capability. Then, adaptive IoU mechanism is proposed, which can dynamically adjust the IoU threshold during the training process to generate high quality proposals. The proposed method is extensively evaluated on two publicly available benchmark datasets, namely SIXray and OPIXray, and the experimental results show that it can achieve the state-of-the-art detection accuracy, compared with other existing methods.
Chunjie Ma, Li Zhuo 0001, Jiafeng Li 0001, Jing Zhang 0023
ICIP1
2022 EAOD-Net: Effective anomaly object detection networks for X-ray images
abstract
Abstract Anomaly object detection is the core technology in the application for X‐ray images. However, the accuracy of current X‐ray anomaly object detection method still needs to be improved. In this paper, an effective anomaly object detection network is proposed to improve the detection accuracy of anomaly object for X‐ray images. Firstly, learnable Gabor convolution layer, deformable convolution, and spatial attention mechanism are introduced to enhance the representative ability of features in ResNeXt. Then, dense local regression is applied to predict the offset of multiple dense boxes in region proposal to locate the object accurately. At last, bigger discriminative RoI pooling is proposed to classify the candidate boxes to improve the accuracy of object classification. Experimental results on the SIXray and OPIXray datasets show that compared with the state‐of‐the‐art methods, the proposed EAOD‐Net can achieve the competitive detection performance.
Chunjie Ma, Li Zhuo 0001, Jiafeng Li 0001, Jing Zhang 0023
IET Image Process.1