Jiafeng Li 0005

dblp:144/1439-5 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
7since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021
YearPublicationVenuePosition
2025 M²RL-Net: Multi-View and Multi-Level Relation Learning Network for Weakly-Supervised Image Forgery Detection
abstract
As digital media manipulation becomes increasingly sophisticated, accurately detecting and localizing image forgeries with minimal supervision has become a critical challenge. Existing weakly supervised image forgery detection (W-IFD) methods often rely on convolutional neural networks (CNNs) and limited exploration of internal relationships, leading to poor detection and localization performance with only image-level labels. To address these limitations, we introduce a novel Multi-View and Multi-Level Relation Learning Network (M²RL-Net) for W-IFD. M²RL-Net effectively identifies forged images using only image-level annotations by exploring relationships between different views and hierarchical levels within images. Specifically, M²RL-Net achieves patch-level self-consistency learning (PSL) and feature-level contrastive learning (FCL) across different views, facilitating more generalized self-supervised learning of forgery features. In detail, PSL employs self-supervised learning to distinguish consistent and inconsistent regions within images, enhancing its ability to accurately locate tampered areas. FCL utilizes feature-level self-view and multi-view contrastive learning to differentiate between genuine and tampered image features, thereby improving the recognition of authentic and manipulated content across different views. Extensive experiments on various datasets demonstrate that M²RL-Net outperforms existing weakly-supervised methods in both detection and localization accuracy. This research sets a new benchmark for weakly-supervised image forgery detection and lays a robust foundation for future studies in this field.
Jiafeng Li 0005, Ying Wen 0003, Lianghua He
AAAI1
2025 Self-Optimization Training for Weakly Supervised Image Manipulation Localization
abstract
With the continuous evolution of image manipulation techniques, there is an urgent need for an effective method to detect and localize manipulated images. However, existing fully supervised methods require large amounts of costly pixel-level annotations, whereas weakly supervised methods often fall short in localization performance due to their inability to accurately localize tampered regions with precise boundaries. To tackle this issue, we propose a Self-Optimization Weakly Supervised Localization (SO-WSL) framework, which consists of two main components: a Pseudo-Label Generator (PLG) and a Self Iterative Optimization (SIO) module. The PLG employs Class Activation Maps (CAM) to guide the Segment Anything Model (SAM) in generating pseudo-labels with distinct edges, while the SIO module enhances the model’s focus on forgery-specific features by applying masks to suspected tampered regions and iteratively refining the pseudo-labels, thereby improving localization accuracy. Extensive experiments have shown that our SO-WSL framework significantly outperforms existing weakly supervised methods and can even compete with some fully supervised approaches.
Zhangchen Zhu, Jiafeng Li 0005, Ying Wen 0003
ICASSP2
2024 EAN: An Efficient Attention Module Guided by Normalization for Deep Neural Networks
abstract
Deep neural networks (DNNs) have achieved remarkable success in various fields, and two powerful techniques, feature normalization and attention mechanisms, have been widely used to enhance model performance. However, they are usually considered as two separate approaches or combined in a simplistic manner. In this paper, we investigate the intrinsic relationship between feature normalization and attention mechanisms and propose an Efficient Attention module guided by Normalization, dubbed EAN. Instead of using costly fully-connected layers for attention learning, EAN leverages the strengths of feature normalization and incorporates an Attention Generation (AG) unit to re-calibrate features. The proposed AG unit exploits the normalization component as a measure of the importance of distinct features and generates an attention mask using GroupNorm, L2 Norm, and Adaptation operations. By employing a grouping, AG unit and aggregation strategy, EAN is established, offering a unified module that harnesses the advantages of both normalization and attention, while maintaining minimal computational overhead. Furthermore, EAN serves as a plug-and-play module that can be seamlessly integrated with classic backbone architectures. Extensive quantitative evaluations on various visual tasks demonstrate that EAN achieves highly competitive performance compared to the current state-of-the-art attention methods while sustaining lower model complexity.
Jiafeng Li 0005, Zelin Li 0005, Ying Wen 0003
AAAI1
2024 IIRP-Net: Iterative Inference Residual Pyramid Network for Enhanced Image Registration
abstract
Deep learning-based image registration (DLIR) meth-ods have achieved remarkable success in deformable im-age registration. We observe that iterative inference can exploit the well-trained registration network to the fullest extent. In this work, we propose a novel Iterative Inference Residual Pyramid Network (IIRP-Net) to enhance registration performance without any additional training costs. In IIRP-Net, we construct a streamlined pyramid registration network consisting of a feature extractor and residual flow estimators (RP-Net) to achieve generalized capabilities in feature extraction and registration. Then, in the inference phase, IIRP-Net employs an iterative inference strategy to enhance RP-Net by iteratively reutilizing residual flow es-timators from coarse to fine. The number of iterations is adaptively determined by the proposed IterStop mecha-nism. We conduct extensive experiments on the FLARE and Mindboggle datasets and the results verify the effectiveness of the proposed method, outperforming state-of-the-art de-formable image registration methods. Our code is available at https://github.com/Torbjorn1997/IIRP-Net.
Tai Ma, Suwei Zhang, Jiafeng Li 0005, Ying Wen 0003
CVPR3
2024 InjectionNet: Realizing Information Injection for Medical Image Segmentation with Layer Relationships
Jiafeng Li 0005
ICPR (22)2
2023 SCConv: Spatial and Channel Reconstruction Convolution for Feature Redundancy
abstract
Convolutional Neural Networks (CNNs) have achieved remarkable performance in various computer vision tasks but this comes at the cost of tremendous computational resources, partly due to convolutional layers extracting redundant features. Recent works either compress well-trained large-scale models or explore well-designed lightweight models. In this paper, we make an attempt to exploit spatial and channel redundancy among features for CNN compression and propose an efficient convolution module, called SCConv (Spatial and Channel reconstruction Convolution), to decrease redundant computing and facilitate representative feature learning. The proposed SCConv consists of two units: spatial reconstruction unit (SRU) and channel reconstruction unit (CRU). SRU utilizes a separate-and-reconstruct method to suppress the spatial redundancy while CRU uses a split-transform-and-fuse strategy to diminish the channel redundancy. In addition, SCConv is a plug-and-play architectural unit that can be used to replace standard convolution in various convolutional neural networks directly. Experimental results show that SCConv-embedded models are able to achieve better performance by reducing redundant features with significantly lower complexity and computational costs.
Jiafeng Li 0005, Ying Wen 0003, Lianghua He
CVPR1
2022 KAConv: Kernel attention convolutions
Xinxin Shan, Tai Ma, YuTao Shen, Jiafeng Li 0005, Ying Wen 0003
Neurocomputing4