Fei Li 0022

dblp:87/3534-22 · DBLP profile ↗
← Back
17ranked-venue papers
3as first author
17since 2021 · last 2026
0000-0002-8440-359XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 2 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Underwater image enhancement via degradation information extraction and guidance
Fukuan Wang, Fei Li 0022, Chaojun Cen, Qingling Duan
Pattern Recognit.2
2025 Unified Multi-modal Salient Object Detection via Frequency Prompt and Adapter Tuning
abstract
Cost-effective salient object detection (SOD) in multi-modal scenarios is a crucial yet challenging task. This work explores a unified multi-modal SOD approach that demonstrate strong performance across various domain-specific datasets. However, existing methods often struggle with poor representation ability across data distributions and inefficient adaptiveness to multi-modal data. To address these challenges, we propose UMMSOD, a unified multi-modal salient object detection framework, which integrates a novel frequency-based prompt enhancement generator and a modified cross-modal adapter, achieving an effective balance between accuracy and efficiency. Concretely, this paper designs a novel frequency-based prompt enhancement generator with a spatial self-attention mechanism to extract salient features across different modalities, enhancing representation capability. Then, a modified cross-modal adapter utilizes the multi-scale features to facilitate modality knowledge integration while effectively bridging the gap between different modalities. Extensive experimental results on 15 major benchmarks in multi-modal SOD tasks demonstrate that UMMSOD achieves competitive performance while introducing only 1.89M cross-modality trainable parameters.
Chaojun Cen, Fei Li 0022
ICMR2
2025 TeMTG: Text-Enhanced Multi-Hop Temporal Graph Modeling for Audio-Visual Video Parsing
abstract
Audio-Visual Video Parsing (AVVP) task aims to parse the event categories and occurrence times from audio and visual modalities in a given video. Existing methods usually focus on implicitly modeling audio and visual features through weak labels, without mining semantic relationships for different modalities and explicit modeling of event temporal dependencies. This makes it difficult for the model to accurately parse event information for each segment under weak supervision, especially when high similarity between segmental modal features leads to ambiguous event boundaries. Hence, we propose a multimodal optimization framework, TeMTG, that combines text enhancement and multi-hop temporal graph modeling. Specifically, we leverage pre-trained multimodal models to generate modality-specific text embeddings, and fuse them with audio-visual features to enhance the semantic representation of these features. In addition, we introduce a multi-hop temporal graph neural network, which explicitly models the local temporal relationships between segments, capturing the temporal continuity of both short-term and long-range events. Experimental results demonstrate that our proposed method achieves state-of-the-art (SOTA) performance in multiple key indicators in the LLP dataset.
Yaru Chen 0003, Peiliang Zhang, Fei Li 0022, Faegheh Sardari, Ruohao Guo, Wenwu Wang 0001
ICMR3
2025 Towards salient object detection via parallel dual-decoder network
Chaojun Cen, Fei Li 0022, Yun Wang 0009
Eng. Appl. Artif. Intell.2
2025 Contrastive prototype learning with semantic patchmix for few-shot image classification
Mengping Dong, Fei Li 0022
Eng. Appl. Artif. Intell.2
2025 Plantformer: plant point cloud completion based on local-global feature aggregation and spatial context-aware transformer
Fei Li 0022, Yanyu Qi
Neural Comput. Appl.2
2025 PRSN: Prototype resynthesis network with cross-image semantic alignment for few-shot image classification
Mengping Dong, Fei Li 0022
Pattern Recognit.2
2024 UPFormer: U-sharped Perception lightweight Transformer for segmentation of field grape leaf diseases
Fei Li 0022, Haiying Zheng 0001, Weisong Mu
Expert Syst. Appl.2
2023 Multi-scale Attention Conditional GAN for Underwater Image Enhancement
Fei Li 0022
CGI (1)2
2023 OWS-Seg: Online Weakly Supervised Video Instance Segmentation via Contrastive Learning
Yuanxiang Ning, Fei Li 0022, Mengping Dong
ICANN (7)2
2023 Multi-Frequency Representation Enhancement with Privilege Information for Video Super-Resolution
abstract
CNN’s limited receptive field restricts its ability to capture long-range spatial-temporal dependencies, leading to unsatisfactory performance in video super-resolution (VSR). To tackle this challenge, this paper presents a novel multi-frequency representation enhancement module (MFE) that performs spatial-temporal information aggregation in the frequency domain. Specifically, MFE mainly includes a spatial-frequency representation enhancement branch which captures the long-range dependency in the spatial dimension, and an energy frequency representation enhancement branch to obtain the inter-channel feature relationship. Moreover, a novel model training method named privilege training is proposed to encode the privilege information from high-resolution videos to facilitate model training. With these two methods, we introduce a new VSR model named MFPI, which outperforms state-of-the-art methods by a large margin while maintaining good efficiency on various datasets, including REDS4, Vimeo, Vid4, and UDM10.
Fei Li 0022, Linfeng Zhang 0001, Zikun Liu 0001, Juan Lei
ICCV1
2023 CRFormer: Cross-Resolution Transformer for segmentation of grape leaf diseases with context mining
Chaojun Cen, Fei Li 0022, Weisong Mu
Expert Syst. Appl.3
2023 Dual-attention global domain adaptation for mariculture image enhancement
abstract
Abstract Mariculture image enhancement aims to recover degraded images and meet the requirements of various digital aquaculture systems. However, the existing underwater image enhancement (UIE) cannot suit diverse marine scenarios and leads to sub‐optimal results for the real mariculture images. To solve the aforementioned issues, a novel dual‐attention global domain‐adaptive mariculture image enhancement network (DAMIE) is proposed to improve the quality of degraded images. Specifically, the proposed method consists of two core parts: (1) an innovative depth transfer dual‐attention module to aggregate multiple features and bridge the difference between domains; (2) a modified encoder–decoder enhancement network with a global feature vector to reconstruct clean mariculture images. Meanwhile, a semi‐supervised adaptive training scheme is utilized to improve the model generalization in different mariculture domains. Extensive experiments demonstrate that the proposed DAMIE can achieve a good performance in terms of quantitative and qualitative metrics. In addition, an ablation study is conducted to analyse the contribution of the key components in the proposed model.
Fei Li 0022, Chaojun Cen
IET Image Process.1
2022 DRCNet: Dynamic Image Restoration Contrastive Network
Fei Li 0022, Lingfeng Shen, Yang Mi
ECCV (19)1
2022 Towards fusing fuzzy discriminative projection and representation learning for image classification
Yun Wang 0009, Fei Li 0022
Eng. Appl. Artif. Intell.3
2022 Fuzzy Discriminative Block Representation Learning for Image Feature Extraction
abstract
Representation learning is widely used to project high-dimensional data to low-dimensional subspace for feature extraction in image recognition tasks. However, many related methods barely explore the fuzziness and uncertainty between data classes. Besides, the classical unsupervised sparse constraint weakens the evaluation of feature importance and neglects the preservation of discriminant information during sparse representation. To solve these issues, a novel fuzzy discriminative block representation learning (FDBRL) algorithm is proposed for image feature extraction. FDBRL aims to enhance the discriminability of subspace by designing effective constraints for projection learning. Specifically, based on the label information and the fuzzy relation between data, we construct a fuzzy block weight matrix and embed it into the${l_{2,1}}$norm regularization term to realize supervised sparse constraint for the representation learning. Next, the low-rank constraint is used to capture the inherent global structure information of data. Finally, we introduce a classification loss term with transformation matrix for joint optimization, such that the projection learning is not limited to number of classes, and the discriminative ability is further improved. Comprehensive experimental results on six benchmarks verify that our method achieves promising performance with other state-of-the-arts in both robustness and effectiveness.
Yun Wang 0009, Fei Li 0022, Yang Mi
IEEE Trans. Image Process.3
2021 S-FPN: A shortcut feature pyramid network for sea cucumber detection in underwater images
Zheng Miao, Fei Li 0022
Expert Syst. Appl.3