Qifeng Lin

dblp:119/0084 · DBLP profile ↗
← Back
25ranked-venue papers
12as first author
21since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 5 first-author · 8 since 2021Artificial intelligence and machine learning · 9 · 4 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2025 Multi-Attention Guided Knowledge Distillation For High-Performance Object Detection
abstract
Knowledge distillation is beneficial for improving the performance of object detection models. However, the existing methods utilizing attention maps for feature weighting embrace limited flexibility, and may result in the loss of crucial channel and location information. To fully utilize these critical information, this paper introduce a novel Multi-Attention Guided Distillation framework that aims to enrich the channel of detail attention representations by enhancing local feature expressions, while patching feature maps concurrently. Meanwhile, a global spatial attention map is utilized to supplement global feature information. To alleviate the discrepancy in the feature attention map between the teacher and student, we use the original student’s features to mimic the weighted teacher’s features. Extensive ablation experiments have proven the effectiveness of our method. Compared with other distillation methods, the SOTA results demostrate the advanced nature of our model, and some student models even outperform the teacher model.
Zhihao Kong, Qifeng Lin, Qishen Shen, Jiayi Qiu, Gang Fu 0003, Yuanlong Yu 0001
ICME2
2025 Byzantine-Resilient Decentralized Parallel Policy Gradient
abstract
Parallel reinforcement learning (RL) is an important approach to dealing with the challenge of data inefficiency in RL. The existing distributed framework requires a central server to collect messages from multiple agents for cooperation, and thus suffers from the communication bottleneck. To address this issue, we first develop a decentralized parallel RL algorithm, named as decentralized parallel policy gradient (DP-PG), within which the agents exchange learning parameters via a peer-to-peer network for cooperation. However, Byzantine attacks are ubiquitous in multi-agent systems, where malicious agents could send random or well-designed messages to their neighbors for the sake of hindering or destroying learning processes. Therefore, we further propose Byzantine-resilient decentralized parallel policy gradient (BRDP-PG) that replaces the vulnerable weighted mean aggregation in DP-PG with coordinate trimmed mean (CTM), a robust aggregation rule. Last but not the least, we conduct numerical experiments to confirm the effectiveness of the proposed DP-PG and BRDP-PG.
Qifeng Lin
IJCNN1
2025 EEG super-resolution with Laplacian Regularized Coupled Matrix Decomposition: A case study of Autism Spectrum Disorder EEG enhancement
Yunbo Tang, Qifeng Lin, Yuanlong Yu 0001, Dan Chen 0001
Artif. Intell. Medicine2
2025 Attention-Based Mean-Max Balance Assignment for Oriented Object Detection in Optical Remote Sensing Images
abstract
For objects with arbitrary angles in optical remote sensing (RS) images, the oriented bounding box regression task often faces the problem of ambiguous boundaries between positive and negative samples. The statistical analysis of existing label assignment strategies reveals that anchors with low Intersection over Union (IoU) between ground truth (GT) may also accurately surround the GT after decoding. Therefore, this article proposes an attention-based mean-max balance assignment (AMMBA) strategy, which consists of two parts: mean-max balance assignment (MMBA) strategy and balance feature pyramid with attention (BFPA). MMBA employs the mean-max assignment (MMA) and balance assignment (BA) to dynamically calculate a positive threshold and adaptively match better positive samples for each GT for training. Meanwhile, to meet the need of MMBA for more accurate feature maps, we construct a BFPA module that integrates spatial and scale attention mechanisms to promote global information propagation. Combined with S2ANet, our AMMBA method can effectively achieve state-of-the-art performance, with a precision of 80.91% on the DOTA dataset in a simple plug-and-play fashion. Extensive experiments on three challenging optical RS image datasets (DOTA-v1.0, HRSC, and DIOR-R) further demonstrate the balance between precision and speed in single-stage object detectors. Our AMMBA has enough potential to assist all existing RS models in a simple way to achieve better detection performance. The code is available athttps://github.com/promisekoloer/AMMBA.
Qifeng Lin, Daoye Zhu, Gang Fu 0003, Chuanxi Chen, Yuanlong Yu 0001
IEEE Trans. Geosci. Remote. Sens.1
2025 Multiple Region Proposal Experts Network for Wide-Scale Remote Sensing Object Detection
abstract
Faced with the wide-scale characteristics of objects in optical remote sensing images, the current object detection models are always unable to provide satisfactory detection capabilities for remote sensing tasks. To achieve better wide-scale coverage for various remote sensing regions of interest, this article introduces a multiprediction mechanism to build a novel region generation model, namely, a multiple region proposal experts network (MRPENet). Meanwhile, to achieve both region proposal coverage and receptive field coverage of wide-scale objects, we constructed a prior design of an anchor (PDA) module and an adaptive features compensation (AFC) module to achieve the coverage of wide-scale remote sensing objects. To better utilize the multiexpert characteristics of our model, we customized a new training sample allocation strategy, dynamic scale-assigned expert learning (DSAEL), to cultivate the ability of experts to deal with objects at various scales. To the best of our knowledge, this is the first time that a multiple region proposal network (RPN) mechanism has been used in the object detection of optical remote sensing images. Extensive experiments have shown the generality and effectiveness of our MRPENet. Without bells and whistles, MRPENet achieves a new state-of-the-art (SOTA) on standard benchmarks, i.e., DOTA-v1.0 [82.02% mean average precision (mAP)], HRSC2016 (98.16% mAP), and FAIR1M-v1.0 (48.80% mAP).
Qifeng Lin, Daoye Zhu, Gang Fu 0003, Yuanlong Yu 0001
IEEE Trans. Geosci. Remote. Sens.1
2025 FAME: Fusion of Alignment and Multiview Enhancement for Remote Sensing Image-Text Retrieval
abstract
Contemporary advancements in Earth observation technologies have generated substantial data resources for remote sensing image retrieval applications. However, existing models exhibit limitations in extracting local features from image-text pairs and establishing effective connections between local and global information. Furthermore, these models lack fine-grained alignment capabilities between different modalities. To address these challenges, we propose FAME, which integrates four key components: a Progressive Masking (ProgMask) module, an OmniView Fusion (OVF) module, a Global Feature Attention with Multi-view Feature Alignment (GFA-MFA) module, and Regional Clustering Alignment Loss (RCA-Loss). The ProgMask module employs a progressive training strategy with alternating masking of image and text modalities to enhance cross-modal learning robustness. The OVF module employs omni-view attention mechanisms with adaptive weighting to dynamically focus on different local semantic regions of remote sensing images and corresponding textual descriptions. The GFA-MFA module enhances global feature capture through selective filtering while ensuring precise multi-view feature alignment. These modules work synergistically to achieve fine-grained alignment between multi-scale local and global features across modalities. Additionally, RCA-Loss optimizes intra-class feature clustering while minimizing inter-class confusion through class center distance optimization and enhanced feature discrimination. Experimental validation on three benchmark datasets (RSITMD, RSICD, UCM-Captions) demonstrates that FAME achieves superior performance compared to existing state-of-the-art methods in remote sensing image-text retrieval tasks.
Yuanhao Su, Daoye Zhu, Zhande Dong, Qifeng Lin, Lingbo Liu, Shuming Bao
IEEE Trans. Geosci. Remote. Sens.4
2024 Towards High-Resolution Specular Highlight Detection
Gang Fu 0003, Qing Zhang 0006, Lei Zhu 0003, Qifeng Lin, Siyuan Fan, Chunxia Xiao
Int. J. Comput. Vis.4
2024 Controllable fundus image generation based on conditional generative adversarial networks with mask guidance
Xiaoxin Guo, Qifeng Lin, Xiaoying Hu, Songtian Che
Multim. Tools Appl.3
2024 Robust Reward-Free Actor-Critic for Cooperative Multiagent Reinforcement Learning
abstract
In this article, we consider centralized training and decentralized execution (CTDE) with diverse and private reward functions in cooperative multiagent reinforcement learning (MARL). The main challenge is that an unknown number of agents, whose identities are also unknown, can deliberately generate malicious messages and transmit them to the central controller. We term these malicious actions as Byzantine attacks. First, without Byzantine attacks, we propose a reward-free deep deterministic policy gradient (RF-DDPG) algorithm, in which gradients of agents' critics rather than rewards are sent to the central controller for preserving privacy. Second, to cope with Byzantine attacks, we develop a robust extension of RF-DDPG termed R2F-DDPG, which replaces the vulnerable average aggregation rule with robust ones. We propose a novel class of RL-specific Byzantine attacks that fail conventional robust aggregation rules, motivating the projection-boosted robust aggregation rules for R2F-DDPG. Numerical experiments show that RF-DDPG successfully trains agents to work cooperatively and that R2F-DDPG demonstrates robustness to Byzantine attacks.
Qifeng Lin, Qing Ling 0001
IEEE Trans. Neural Networks Learn. Syst.1
2023 Run and Chase: Towards Accurate Source-Free Domain Adaptive Object Detection
abstract
Recently, there has been increasing interest in the Source-Free Domain Adaptive Object Detection task, which involves training an object detector on the unlabeled target data using a pre-trained source model without accessing the source data. Most related methods are developed from the mean-teacher framework, which aims to train the student model closer to the teacher model via a pseudo labeling manner, where the teacher model is the exponential-moving-average of the student models at different time-steps. Following this line of works, we propose a Run-and-Chase Mutual-Learning method to strengthen the interactions between the student model and the teacher model in both feature and prediction levels. In our method, the student model is optimized to run away from the teacher model at the feature level, while chasing the teacher model at the prediction level. In this way, the student model is forced to be distinguishable at different time-steps, so that the teacher model can acquire more diverse task-related information and produce higher-accuracy pseudo labels. As the training goes, the student and teacher models are updated iteratively and promoted mutually, which can prevent the model collapse problem. Extensive experiments are conducted to validate the effectiveness of our method.
Luojun Lin, Zhifeng Yang, Qipeng Liu 0004, Yuanlong Yu 0001, Qifeng Lin
ICME5
2023 A Multiple Prediction Mechanisms Ensemble for Complex Remote Sensing Scenes
abstract
Facing complex remote sensing scenes, detection models with single detection mechanisms cannot always provide satisfactory detection capabilities. In order to obtain better detection performance in various remote sensing scenes, this paper constructs a novel ensemble model, namely: the multiple prediction mechanisms ensemble (MPME). In order to improve the feature representation ability and region recognition ability of the ensemble model, we build the ensemble of feature pyramids (EFP) and the ensemble of detection heads (EDH) respectively. In order to further improve the detection accuracy of the ensemble model, we propose a training strategy (k-Nearest Loss Learning), so that each sub-detector does not need to learn a trade-off among all training samples, and also reduces the possibility of model over-fitting. The experimental results show that our MPME is a more efficient and effective ensemble model. Compared with other ensemble models, our MPME has a faster detection speed and better detection accuracy. Compared with other state-of-the-art detectors, our detector also achieves superior detection performance.
Qifeng Lin, Luojun Lin, Yuanlong Yu 0001, Gang Fu 0003
ACM Multimedia1
2023 Joint grading of diabetic retinopathy and diabetic macular edema using an adaptive attention block and semisupervised learning
Xiaoxin Guo, Qifeng Lin, Xiaoying Hu, Songtian Che
Appl. Intell.3
2023 A Novel original feature fusion network for joint diabetic retinopathy and diabetic Macular edema grading
Xiaoxin Guo, Qifeng Lin, Haoren Wang, Xiaoying Hu, Songtian Che
Neural Comput. Appl.3
2022 Byzantine-Robust Federated Deep Deterministic Policy Gradient
abstract
Federated reinforcement learning (FRL) combines multi-agent reinforcement learning (MARL) and federated learning (FL) so that multiple agents can exchange messages with a central server for co-operatively learning their local policies. However, a number of malicious agents may deliberately modify the messages transmitted to the central server so as to hinder the learning process, which is often described by the Byzantine attacks model. To address this issue, we propose to employ robust aggregation to replace the simple average aggregation rule in FRL and enhance Byzantine robustness. To be specific, we focus on the episodic task where the environment and agents are reset in the beginning each episode. First, we extend deep deterministic policy gradient (DDPG) to FRL (termed as F-DDPG), which maintains a global critic and multiple local actors, and is thus computation- and communication-efficient. Then, we introduce geometric median and median to aggregate the gradients received from the agents and propose RF-DDPG, a class of Byzantine-robust FRL methods. Finally, we conduct numerical experiments to validate the robustness of RF-DDPG to Byzantine attacks.
Qifeng Lin
ICASSP1
2022 Counterfactual inference graph network for disease prediction
Baoliang Zhang, Xiaoxin Guo, Qifeng Lin, Haoren Wang, Songbai Xu
Knowl. Based Syst.3
2022 DDBN: Dual detection branch network for semantic diversity predictions
Qifeng Lin, Chengjiang Long, Jianhui Zhao 0001, Gang Fu 0003
Pattern Recognit.1
2022 MEDNet: Multiexpert Detection Network With Unsupervised Clustering of Training Samples
abstract
For various remote sensing objects, the current detection framework based on a single detection pipeline fails to provide satisfactory detection accuracy. In order to further improve the object recognition ability of the detection model, this article introduces the effective “multiexpert” mechanism into the field of remote sensing object detection and then constructs a multiexpert detection network (MEDNet). In this model, we first construct multiple feature pyramids (MFPs) to replace the traditional single feature pyramid to enrich the semantic representation ability of the model. Then, we equip multiple detection experts (MDEs) to leverage multiple kinds of features from MFP to perform different semantic predictions. As the first CNN-based multiexpert detection model for remote sensing images, we tailor a loss distance-based k-experts clustering (LD-kEC) strategy to assign training samples to different detection experts in an unsupervised fashion. By this strategy, we can directly use the existing remote sensing dataset without expert labels for end-to-end training of our multiexpert model. The experimental results prove that the proposed multiexpert-based detector can indeed significantly improve the object detection performance for remote sensing images.
Qifeng Lin, Jianhui Zhao 0001, Bo Du 0001, Gang Fu 0003
IEEE Trans. Geosci. Remote. Sens.1
2022 CRPN-SFNet: A High-Performance Object Detector on Large-Scale Remote Sensing Images
abstract
Limited by the GPU memory, the current mainstream detectors fail to directly apply to large-scale remote sensing images for object detection. Moreover, the scale range of objects in remote sensing images is much wider than that of general images, which also greatly hinders the existing methods to effectively detect geospatial objects of various scales. For achieving high-performance object detection on large-scale remote sensing images, this article proposes a much faster and more accurate detecting framework, called cropping region proposal network-based scale folding network (CRPN-SFNet). In our framework, the CRPN includes a weak semantic RPN for quickly locating interesting regions and a strategy of generating cropping regions to effectively filter out meaningless regions, which can greatly reduce the computation and storage burden. Meanwhile, the proposed SFNet leverages the scale folding-based training and testing methods to extend the valid detection range of existing detectors, which is beneficial for detecting remote sensing objects of various scales, including very small and very large geospatial objects. Extensive experiments on the public Dataset for Object deTection in Aerial images data set indicate that our CRPN can help our detector deal the larger image faster with the limited GPU memory; meanwhile, the SFNet is beneficial to achieve more accurate detection of geospatial objects with wide-scale range. For large-scale remote sensing images, the proposed detection framework outperforms the state-of-the-art object detection methods in terms of accuracy and speed.
Qifeng Lin, Jianhui Zhao 0001, Gang Fu 0003
IEEE Trans. Neural Networks Learn. Syst.1
2021 Self-Inference Of Others' Policies For Homogeneous Agents In Cooperative Multi-Agent Reinforcement Learning
abstract
Multi-agent reinforcement learning (MARL) has been widely applied in various cooperative tasks, where multiple agents are trained to collaboratively achieve global goals. During the training stage of MARL, inferring policies of other agents is able to improve the coordination efficiency. However, most of the existing policy inference methods require each agent to model all other agents separately, which results in quadratic growth of resource consumption as the number of agents increases. In addition, inferring the policy of an agent solely from its observations and actions may lead to failure of agent modeling. To address this issue, we propose to let each agent infer the others’ policies with its own model, given that the agents are homogeneous. This self-inference approach significantly reduces the computation and storage consumption, and guarantees the quality of agent modeling. Experimental results demonstrate effectiveness of the proposed approach.
Qifeng Lin
ICASSP1
2021 Decentralized TD(0) With Gradient Tracking
abstract
In this letter, we consider the policy evaluation problem with linear function approximation in the context of decentralized multi-agent reinforcement learning (MARL), where the agents with a fixed joint policy cooperate to estimate the global expected accumulative reward through a decentralized communication network. In the existing algorithms, every agent updates its local parameter by combining its neighboring local parameters and then running a local stochastic temporal-difference(0) (TD(0)) gradient step. However, due to the diversity of reward functions across the agents, the local stochastic TD(0) gradients can be very different, which hinders the agents from reaching the consensual and optimal parameter. Motivated by the gradient tracking strategy in decentralized optimization, we combine gradient tracking with decentralized TD(0) to accelerate the process of reaching consensus. We also propose two other acceleration strategies, one is gradient consensus while another jointly uses gradient tracking and gradient consensus. Numerical experiments demonstrate that the proposed algorithms attain faster convergence than the popular decentralized TD(0) method.
Qifeng Lin, Qing Ling 0001
IEEE Signal Process. Lett.1
2021 Deep Adversarial Data Augmentation for Extremely Low Data Regimes
abstract
Deep learning has revolutionized the performance of classification and object detection, but meanwhile demands sufficient labeled data for training. Given insufficient data, while many techniques have been developed to help combat overfitting, the challenge remains if one tries to train deep networks, especially in the ill-posedextremely low data regimes: only a small set of labeled data are available, and nothing – including unlabeled data – else. Such regimes arise from practical situations where not only data labeling but also data collection itself is expensive. We propose a deep adversarial data augmentation (DADA) technique to address the problem, in which we elaborately formulate data augmentation as a problem of training a class-conditional and supervised generative adversarial network (GAN). Specifically, a new discriminator loss is proposed to fit the goal of data augmentation, through which both real and augmented samples are enforced to contribute to and be consistent in finding the decision boundaries. Tailored training techniques are developed accordingly. To quantitatively validate its effectiveness, we first perform extensive simulations to show that DADA substantially outperforms both traditional data augmentation and a few GAN-based options. We then extend experiments to three real-world small labeled classification datasets where existing data augmentation and/or transfer learning strategies are either less effective or infeasible. We also demonstrate that DADA to can be extended to the detection task. We improve the pedestrian synthesis work by substitute for our discriminator and training scheme. Validation experiment shows that DADA can improve the detection mean average precision (mAP) compared with some traditional data augmentation techniques in object detection. Source code is available athttps://github.com/SchafferZhang/DADA.
Zhangyang Wang, Dong Liu 0002, Qifeng Lin, Qing Ling 0001
IEEE Trans. Circuits Syst. Video Technol.4
2020 Learning to Detect Specular Highlights from Real-world Images
abstract
Specular highlight detection is a challenging problem, and has many applications such as shiny object detection and light source estimation. Although various highlight detection methods have been proposed, they fail to disambiguate bright material surfaces from highlights, and cannot handle non-white-balanced images. Moreover, at present, there is still no benchmark dataset for highlight detection. In this paper, we present a large-scale real-world highlight dataset containing a rich variety of material categories, with diverse highlight shapes and appearances, in which each image is with an annotated ground-truth mask. Based on the dataset, we develop a deep learning-based specular highlight detection network (SHDNet) leveraging multi-scale context contrasted features to accurately detect specular highlights of varying scales. In addition, we design a binary cross-entropy (BCE) loss and an intersection-over-union edge (IoUE) loss for our network. Compared with existing highlight detection methods, our method can accurately detect highlights of different sizes, while effectively excluding the non-highlight regions, such as bright materials, non-specular as well as colored lighting, and even light sources.
Gang Fu 0003, Qing Zhang 0006, Qifeng Lin, Lei Zhu 0003, Chunxia Xiao
ACM Multimedia3
2020 A Novel Fusion Framework without Pooling for Noisy SAR Image Classification
abstract
Due to the particularity of SAR image, existing SAR image classification models often lack strong robustness against noise. Moreover, SAR images are naturally prone to speckle noise and sensitive to observed azimuth. To solve these problems, in this paper, we propose a novel fusion framework in which the convolutional layer with increased stride is used to replace the max pooling layer. Unlike max pooling layer roughly extracts the maximum pixel value in one region as its main feature, which is easy to introduce noise, convolution operation can update the weights and learn features more rationally by back-propagation. It also can achieve the same purpose of down sampling as pooling layers. In order to make full use of feature maps from different layers, our framework fuses the feature vectors extracted from different layers, which helps improve the performance of our classification model. For the problem of overfitting caused by the small MSTAR dataset of SAR images, we replace fully connected layers with convolution layers to relieve the overfitting of the convolution layers by reducing the number of parameters. In order to improve the robustness against observed azimuth angles of the dataset, we adopt the multi-channel calibration and superposition as model's input, which can be used in real flight platform. The extensive experiments conducted on the MSTAR dataset have clearly demonstrated that our framework achieves higher classification accuracy, stronger robustness against noise than other existing methods, as well as its excellent classification performance for the targets of the same category and different subcategories, which is more difficult to be classified.
Jianhui Zhao 0001, Qifeng Lin
SMC4
2019 Cropping Region Proposal Network Based Framework for Efficient Object Detection on Large Scale Remote Sensing Images
abstract
It is very difficult to directly detect objects on the entire large scale remote sensing image, due to the limited GPU memory. Moreover, there are no objects of interest in most areas of such a huge image, thus a lot of computational costs is wasted in dealing with these vain areas. Therefore, this paper proposes a Cropping Region Proposal Network (CRPN), which includes a weak semantic RPN for quickly locating interesting regions, and a dual-scale strategy for generating effective cropping regions. Cropping regions consist of small and large cropping scales for detecting various-scale objects including very small and very large objects, which is hard for existing methods. CRPN helps to detect effective regions of remote sensing image. Meanwhile, it is also modularized and can be easily connected with mainstream detectors to form an end-to-end detecting framework. Experiments on public DOTA dataset show that our CRPN is effective for filtering invalid regions to greatly reduce the computation burden, and helps to achieve more accurate object detection on large scale remote sensing images.
Qifeng Lin, Jianhui Zhao 0001, Qianqian Tong 0001, Guian Zhang, Gang Fu 0003
ICME1
2019 Specular Highlight Removal for Real-world Images
abstract
Abstract Removing specular highlight in an image is a fundamental research problem in computer vision and computer graphics. While various methods have been proposed, they typically do not work well for real‐world images due to the presence of rich textures, complex materials, hard shadows, occlusions and color illumination, etc. In this paper, we present a novel specular highlight removal method for real‐world images. Our approach is based on two observations of the real‐world images: (i) the specular highlight is often small in size and sparse in distribution; (ii) the remaining diffuse image can be represented by linear combination of a small number of basis colors with the sparse encoding coefficients. Based on the two observations, we design an optimization framework for simultaneously estimating the diffuse and specular highlight images from a single image. Specifically, we recover the diffuse components of those regions with specular highlight by encouraging the encoding coefficients sparseness using L0 norm. Moreover, the encoding coefficients and specular highlight are also subject to the non‐negativity according to the additive color mixing theory and the illumination definition, respectively. Extensive experiments have been performed on a variety of images to validate the effectiveness of the proposed method and its superiority over the previous methods.
Gang Fu 0003, Qing Zhang 0006, Chengfang Song, Qifeng Lin, Chunxia Xiao
Comput. Graph. Forum4