Qiufu Li

dblp:133/7086 · DBLP profile ↗
← Back
26ranked-venue papers
10as first author
23since 2021 · last 2026
0000-0002-8120-6531ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 5 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 4 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2026 BCE3S: Binary Cross-Entropy Based Tripartite Synergistic Learning for Long-Tailed Recognition
abstract
For long-tailed recognition (LTR) tasks, high intra-class compactness and inter-class separability in both head and tail classes, as well as balanced separability among all the classifier vectors, are preferred. The existing LTR methods based on cross-entropy (CE) loss not only struggle to learn features with desirable properties but also couple imbalanced classifier vectors in the denominator of its Softmax, amplifying the imbalance effects in LTR. In this paper, for the LTR, we propose a binary cross-entropy (BCE)-based tripartite synergistic learning, termed BCE3S, which consists of three components: (1) BCE-based joint learning optimizes both the classifier and sample features, which achieves better compactness and separability among features than the CE-based joint learning, by decoupling the metrics between feature and the imbalanced classifier vectors in multiple Sigmoid; (2) BCE-based contrastive learning further improves the intra-class compactness of features; (3) BCE-based uniform learning balances the separability among classifier vectors and interactively enhances the feature properties by combining with the joint learning. The extensive experiments show that the LTR model trained by BCE3S not only achieves higher compactness and separability among sample features, but also balances the classifier's separability, achieving SOTA performance on various long-tailed datasets such as CIFAR10-LT, CIFAR100-LT, ImageNet-LT, and iNaturalist2018.
Weijia Fan, Qiufu Li, Jiajun Wen 0001, Xiaoyang Peng
AAAI2
2026 MNSeg: Mamba-based 3D neuron segmentation integrated with bidirectional attention mechanism and topological loss
Xinle Dai, Qiufu Li, LinLin Shen, Wenting Chen, Weijia Fan
Neurocomputing2
2026 AEPL: Adaptive empirical prototype learning with dynamic margins for deep face recognition
Weijia Fan, Zhixiang Cai, Chunsong Chen, Yanxi Liu 0004, Jiajun Wen 0001, Xi Jia, LinLin Shen, Jiancan Zhou, Qiufu Li
Pattern Anal. Appl.10
2026 Learning discriminative features within forward-Forward algorithm using convolutional prototype
Qiufu Li, LinLin Shen
Pattern Recognit.1
2026 Ranking-Based Self-Supervised Representation Learning for Skeleton-Based Action Recognition
abstract
Recently, researchers have achieved significant results in the skeleton-based action recognition. To better model the skeleton sequences, we drive the encoder to learn more discriminative representations in the self-supervised setting. We find that instead of clustering feature vectors to assign pseudo labels for samples as in DeepCluster, ranking them is a more reasonable, reliable, and efficient way to learn more effective feature representations. With this intuition, we propose a novel self-supervised learning framework,DeepRank. Specifically, we rank triplets of skeleton sequences with the ranking labels, obtained from the relative distances among them. Besides, to deeply mine complementary discriminative information that exists in different modalities of skeleton sequences, we further proposeMulti-ViewDeepRank(MV-DeepRank) to enable encoders to comprehensively learn complementary features from multiple modalities. Extensive experimental results on the NTU RGB+D, NTU RGB+D 120, PKU-MMD I, and PKU-MMD II datasets under various evaluation settings demonstrate the generality, transferability, and superiority of our proposed self-supervised learning frameworks. Notably, our frameworks surpass the previous methods that employ the same backbone networks as ours by at least 1.8% (ST-GCN) and 2.1% (STTFormer) under the finetuning setting. Additionally, DeepRank gains a significant advantage on computational complexities,$O(1)$, over the contrastive learning-based methods,$O(\rm{batch size})$, and the clustering-based methods,$O(\rm{number of clusters})$.
Bizhu Wu, Junliang Chen 0002, Jinheng Xie, Qiufu Li, Jianfeng Ren, Ruibin Bai, Rong Qu, LinLin Shen
IEEE Trans. Multim.4
2025 DAP-MAE: Domain-Adaptive Point Cloud Masked Autoencoder for Effective Cross-Domain Learning
Qiufu Li, LinLin Shen
ICCV2
2025 BCE vs. CE in Deep Feature Learning
abstract
When training classification models, it expects that the learned features are compact within classes, and can well separate different classes. As the dominant loss function for training classification models, minimizing cross-entropy (CE) loss maximizes the compactness and distinctiveness, i.e., reaching neural collapse (NC). The recent works show that binary CE (BCE) performs also well in multi-class tasks. In this paper, we compare BCE and CE in deep feature learning. For the first time, we prove that BCE can also maximize the intra-class compactness and inter-class distinctiveness when reaching its minimum, i.e., leading to NC. We point out that CE measures the relative values of decision scores in the model training, implicitly enhancing the feature properties by classifying samples one-by-one. In contrast, BCE measures the absolute values of decision scores and adjust the positive/negative decision scores across all samples to uniformly high/low levels. Meanwhile, the classifier biases in BCE present a substantial constraint on the decision scores to explicitly enhance the feature properties in the training. The experimental results are aligned with above analysis, and show that BCE could improve the classification and leads to better compactness and distinctiveness among sample features. The codes have be released.
Qiufu Li, Huibin Xiao, LinLin Shen
ICML1
2025 Distilled transformers with locally enhanced global representations for face forgery detection
Qiufu Li, Zitong Yu, LinLin Shen
Pattern Recognit.2
2025 A Wavelet-Guided Deep Unfolding Network for Single Image Reflection Removal
abstract
Removing unwanted reflections from images is a fundamental yet challenging problem in low-level computer vision. Recent deep learning-based Single Image Reflection Removal (SIRR) methods have made significant progress. However, separating reflections from transmission content remains difficult, particularly in complex scenes where the two exhibit high visual similarity. Upon careful analysis, we find that reflections predominantly reside in the high-frequency components of an image. These reflections tend to distort fine details in the high-frequency range, while the low-frequency information remains relatively less affected. This observation motivates us to explore a frequency-aware approach for SIRR by leveraging the Discrete Wavelet Transform (DWT). The wavelet decomposition enables us to distinguish and isolate reflective artifacts in the frequency domain while preserving the transmission information. Building on this insight, we propose a novel Wavelet-guided Deep Unfolding Network (WDUNet) that leverages the strengths of wavelet decomposition and deep unfolding techniques to improve interpretability and generalization in SIRR. Specifically, we formulate an optimization-based reflection removal model using DWT and convolutional dictionaries. The proposed model is optimized via a proximal gradient algorithm and then unfolded into a neural network architecture, where all parameters are learned end-to-end during training. By combining wavelet domain analysis with deep unfolding, WDUNet enhances both the interpretability and generalization of SIRR methods. Additionally, we design and integrate the Low-frequency Parameter Estimation Module (LPEM) and High-frequency Parameter Estimation Module (HPEM) modules into WDUNet, allowing the network to automatically learn and optimize the models' hyperparameters. Extensive experiments conducted on four benchmark datasets demonstrate that WDUNet consistently outperforms existing state-of-the-art methods in both objective evaluation metrics and subjective visual quality.
Qiufu Li, Xu Wu 0001, Nan Mu, LinLin Shen
IEEE Trans. Image Process.2
2024 PointFaceFormer: Local and Global Attention Based Transformer for 3D Point Cloud Face Recognition
abstract
Existing 3D point cloud-based facial recognition struggles to fully leverage both global and local information inherent in the 3D point cloud data. In this paper, we introduce the PointFaceFormer, the first Transformer model designed for 3D point cloud face recognition. It incorporates an attention mechanism based on dot product and cosine functions to construct a similarity Transformer architecture, which effectively extracts both local and global features from the point cloud data. Experimental results demonstrate that PointFaceFormer achieves a recognition accuracy of 89.08% and a verification accuracy of 76.93% on the large-scale facial point cloud dataset Lock3DFace, which is a new state-of-the-art in 3D face recognition. Furthermore, PointFaceFormer exhibits excellent generalization performance on cross-quality datasets. Additionally, we validate the effectiveness of the attention mechanism through ablation experiments, which justify the effectiveness of the proposed modules.
Qiufu Li, Gui Wang, LinLin Shen
FG2
2024 WiNet: Wavelet-Based Incremental Learning for Efficient Medical Image Registration
Xinxing Cheng, Xi Jia, Wenqi Lu 0001, Qiufu Li, LinLin Shen, Alexander Krull, Jinming Duan 0001
MICCAI (2)4
2024 3DFaceMAE: Pre-training of Masked Autoencoder Using Patch-Based Random Masking Reconstruction and Super-resolution for 3D Face Recognition
Qiufu Li, LinLin Shen, Junpeng Yang
PRCV (11)2
2024 Dual-branch interactive cross-frequency attention network for deep feature learning
Qiufu Li, LinLin Shen
Expert Syst. Appl.1
2024 DUGAN: Infrared and visible image fusion based on dual fusion paths and a U-type discriminator
Yongdong Huang, Qiufu Li, Yuduo Zhang, Qingjian Zhou
Neurocomputing3
2024 PointSurFace: Discriminative point cloud surface feature extraction for 3D face recognition
Junpeng Yang, Qiufu Li, LinLin Shen
Pattern Recognit.2
2023 Activation Template Matching Loss for Explainable Face Recognition
abstract
Can we construct an explainable face recognition network able to learn a facial part-based feature like eyes, nose, mouth and so forth, without any manual annotation or additionalsion datasets? In this paper, we propose a generic Explainable Channel Loss (ECLoss) to construct an explainable face recognition network. The explainable network trained with ECLoss can easily learn the facial part-based representation on the target convolutional layer, where an individual channel can detect a certain face part. Our experiments on dozens of datasets show that ECLoss achieves superior explainability metrics, and at the same time improves the performance of face verification without face alignment. In addition, our visualization results also illustrate the effectiveness of the proposed ECLoss.
Huawei Lin 0001, Qiufu Li, LinLin Shen
FG3
2023 UniFace: Unified Cross-Entropy Loss for Deep Face Recognition
abstract
As a widely used loss function in deep face recognition, the softmax loss cannot guarantee that the minimum positive sample-to-class similarity is larger than the maximum negative sample-to-class similarity. As a result, no unified threshold is available to separate positive sample-to-class pairs from negative sample-to-class pairs. To bridge this gap, we design a UCE (Unified Cross-Entropy) loss for face recognition model training, which is built on the vital constraint that all the positive sample-to-class similarities shall be larger than the negative ones. Our UCE loss can be integrated with margins for a further performance boost. The face recognition model trained with the proposed UCE loss, UniFace, was intensively evaluated using a number of popular public datasets like MFR, IJB-C, LFW, CFP-FP, AgeDB, and MegaFace. Experimental results show that our approach outperforms SOTA methods like SphereFace, CosFace, ArcFace, Partial FC, etc. Especially, till the submission of this work (Mar. 8, 2023), the proposed UniFace achieves the highest TAR@MR-All on the academic track of the MFR-ongoing challenge. $\color{Blue}{\mathbf{Code}}$ is publicly available.
Jiancan Zhou, Xi Jia, Qiufu Li, LinLin Shen, Jinming Duan 0001
ICCV3
2023 CDNet: Cross-frequency Dual-branch Network for Face Anti-Spoofing
abstract
Face anti-spoofing (FAS) defends the facial image recognition systems against the spoof attacks. While the imperceptible spoof cues in the facial images are usually represented in the images' high-frequency components, existing methods do not fully explore them. In this paper, we introduce wavelet into face anti-spoofing and propose a Cross-frequency Dual-branch network (CDNet), which mainly contains two frequency branches to explore spoof cues from the input facial images' high- and low-frequency components generated by wavelet transforms. In CDNet, we design Frequency Attention Module (FAM) to fuse different internal frequency features learned by two frequency branches, and propose a Complementary Learning Module (CLM) to aggregate the two final frequency features. In addition, we present a resolution-aware Binary Cross-Entropy Loss to balance the training samples with different resolutions. We conduct comprehensive experiments on four datasets, and the results shows that our CDNet performs better than the previous state-of-the-art methods on both intra- and inter-dataset testing.
Xiaobin Huang, Qiufu Li, LinLin Shen
IJCNN2
2023 UniTSFace: Unified Threshold Integrated Sample-to-Sample Loss for Face Recognition
abstract
Sample-to-class-based face recognition models can not fully explore the cross-sample relationship among large amounts of facial images, while sample-to-sample-based models require sophisticated pairing processes for training. Furthermore, neither method satisfies the requirements of real-world face verification applications, which expect a unified threshold separating positive from negative facial pairs. In this paper, we propose a unified threshold integrated sample-to-sample based loss (USS loss), which features an explicit unified threshold for distinguishing positive from negative pairs. Inspired by our USS loss, we also derive the sample-to-sample based softmax and BCE losses, and discuss their relationship. Extensive evaluation on multiple benchmark datasets, including MFR, IJB-C, LFW, CFP-FP, AgeDB, and MegaFace, demonstrates that the proposed USS loss is highly efficient and can work seamlessly with sample-to-class-based losses. The embedded loss (USS and sample-to-class Softmax loss) overcomes the pitfalls of previous approaches and the trained facial model UniTSFace exhibits exceptional performance, outperforming state-of-the-art methods, such as CosFace, ArcFace, VPL, AnchorFace, and UNPG. Our code is available at https://github.com/CVI-SZU/UniTSFace.
Qiufu Li, Xi Jia, Jiancan Zhou, LinLin Shen, Jinming Duan 0001
NeurIPS1
2022 Content and Gradient Model-driven Deep Network for Single Image Reflection Removal
abstract
Single image reflection removal (SIRR) is an extremely challenging, ill-posed problem with many application scenarios. In recent years, massive deep learning-based methods have been proposed to remove undesirable reflections from a single input image. However, these methods lack interpretability and do not fully utilize the intrinsic physical structure of reflection images. In this paper, we propose a content and gradient-guided deep network (CGDNet) for single image reflection removal, which is a full-interpretable and model-driven network. Firstly, using the multi-scale convolutional dictionary, we design a novel single image reflection removal model, which combines the image content prior and gradient prior information. Then, the model is optimized using an optimization algorithm based on the proximal gradient technique and unfolded into a neural network, i.e., CGDNet. All the parameters of CGDNet can be automatically learned by end-to-end training. Besides, we introduce a reflection detection module into CGDNet to obtain a probabilistic confidence map and ensure that the network pays attention to reflection regions. Extensive experiments on four benchmark datasets demonstrate that CGDNet is more efficient than state-of-the-art methods in terms of both subjective and objective evaluations. Code is available at https://github.com/zynwl/CGDNet.
LinLin Shen, Qiufu Li
ACM Multimedia3
2022 WaveSNet: Wavelet Integrated Deep Networks for Image Segmentation
Qiufu Li, LinLin Shen
PRCV (4)1
2022 Neuron segmentation using 3D wavelet integrated encoder-decoder network
abstract
MOTIVATION: 3D neuron segmentation is a key step for the neuron digital reconstruction, which is essential for exploring brain circuits and understanding brain functions. However, the fine line-shaped nerve fibers of neuron could spread in a large region, which brings great computational cost to the neuron segmentation. Meanwhile, the strong noises and disconnected nerve fibers bring great challenges to the task. RESULTS: In this article, we propose a 3D wavelet and deep learning-based 3D neuron segmentation method. The neuronal image is first partitioned into neuronal cubes to simplify the segmentation task. Then, we design 3D WaveUNet, the first 3D wavelet integrated encoder-decoder network, to segment the nerve fibers in the cubes; the wavelets could assist the deep networks in suppressing data noises and connecting the broken fibers. We also produce a Neuronal Cube Dataset (NeuCuDa) using the biggest available annotated neuronal image dataset, BigNeuron, to train 3D WaveUNet. Finally, the nerve fibers segmented in cubes are assembled to generate the complete neuron, which is digitally reconstructed using an available automatic tracing algorithm. The experimental results show that our neuron segmentation method could completely extract the target neuron in noisy neuronal images. The integrated 3D wavelets can efficiently improve the performance of 3D neuron segmentation and reconstruction. AVAILABILITYAND IMPLEMENTATION: The data and codes for this work are available at https://github.com/LiQiufu/3D-WaveUNet. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Qiufu Li, LinLin Shen
Bioinform.1
2021 WaveCNet: Wavelet Integrated CNNs to Suppress Aliasing Effect for Noise-Robust Image Classification
abstract
Though widely used in image classification, convolutional neural networks (CNNs) are prone to noise interruptions, i.e. the CNN output can be drastically changed by small image noise. To improve the noise robustness, we try to integrate CNNs with wavelet by replacing the common down-sampling (max-pooling, strided-convolution, and average pooling) with discrete wavelet transform (DWT). We firstly propose general DWT and inverse DWT (IDWT) layers applicable to various orthogonal and biorthogonal discrete wavelets like Haar, Daubechies, and Cohen, etc., and then design wavelet integrated CNNs (WaveCNets) by integrating DWT into the commonly used CNNs (VGG, ResNets, and DenseNet). During the down-sampling, WaveCNets apply DWT to decompose the feature maps into the low-frequency and high-frequency components. Containing the main information including the basic object structures, the low-frequency component is transmitted into the following layers to generate robust high-level features. The high-frequency components are dropped to remove most of the data noises. The experimental results show that WaveCNets achieve higher accuracy on ImageNet than various vanilla CNNs. We have also tested the performance of WaveCNets on the noisy version of ImageNet, ImageNet-C and six adversarial attacks, the results suggest that the proposed DWT/IDWT layers could provide better noise-robustness and adversarial robustness. When applying WaveCNets as backbones, the performance of object detectors (i.e., faster R-CNN and RetinaNet) on COCO detection dataset are consistently improved. We believe that suppression of aliasing effect, i.e. separation of low frequency and high frequency information, is the main advantages of our approach. The code of our DWT/IDWT layer and different WaveCNets are available at https://github.com/CVI-SZU/WaveCNet.
Qiufu Li, LinLin Shen, Sheng Guo 0005, Zhihui Lai 0001
IEEE Trans. Image Process.1
2020 Wavelet Integrated CNNs for Noise-Robust Image Classification
abstract
Convolutional Neural Networks (CNNs) are generally prone to noise interruptions, i.e., small image noise can cause drastic changes in the output. To suppress the noise effect to the final predication, we enhance CNNs by replacing max-pooling, strided-convolution, and average-pooling with Discrete Wavelet Transform (DWT). We present general DWT and Inverse DWT (IDWT) layers applicable to various wavelets like Haar, Daubechies, and Cohen, etc., and design wavelet integrated CNNs (WaveCNets) using these layers for image classification. In WaveCNets, feature maps are decomposed into the low-frequency and high-frequency components during the down-sampling. The low-frequency component stores main information including the basic object structures, which is transmitted into the subsequent layers to extract robust high-level features. The high-frequency components, containing most of the data noise, are dropped during inference to improve the noise-robustness of the WaveCNets. Our experimental results on ImageNet and ImageNet-C (the noisy version of ImageNet) show that WaveCNets, the wavelet integrated versions of VGG, ResNets, and DenseNet, achieve higher accuracy and better noise-robustness than their vanilla versions.
Qiufu Li, LinLin Shen, Sheng Guo 0005, Zhihui Lai 0001
CVPR1
2020 3D Neuron Reconstruction in Tangled Neuronal Image With Deep Networks
abstract
Digital reconstruction or tracing of 3D neuron is essential for understanding the brain functions. While existing automatic tracing algorithms work well for the clean neuronal image with a single neuron, they are not robust to trace the neuron surrounded by nerve fibers. We propose a 3D U-Net-based network, namely 3D U-Net Plus, to segment the neuron from the surrounding fibers before the application of tracing algorithms. All the images in BigNeuron, the biggest available neuronal image dataset, contain clean neurons with no interference of nerve fibers, which are not practical to train the segmentation network. Based upon the BigNeuron images, we synthesize a SYNethic TAngled NEuronal Image dataset (SYNTANEI) to train the proposed network, by fusing the neurons with extracted nerve fibers. Due to the adoption of dropout, àtrous convolution and Àtrous Spatial Pyramid Pooling (ASPP), experimental results on the synthetic and real tangled neuronal images show that the proposed 3D U-Net Plus network achieved very promising segmentation results. The neurons reconstructed by the tracing algorithm using the segmentation result match significantly better with the ground truth than that using the original images.
Qiufu Li, LinLin Shen
IEEE Trans. Medical Imaging1
2016 Generalization of SPIHT: Set Partition Coding System
abstract
This paper constructs a set partition coding system (SPACS) to combine the advantages of different types of set partition coding algorithms. General tree (GT) is an important conception introduced in this paper, which can represent tree set and square set simultaneously. With the help of GT, SPIHT is generalized to construct degree- k SPIHT based on the analysis of two kinds of set partition operations. Using the same coding mechanism, SPACS (k,p) is constructed, aided with virtual subbands that are generated by recursive division on the LL band. SPACS belongs to tree-set partition coding algorithms if k and p take smaller values. In particular, SPACS(2,1) is the classical SPIHT. SPACS tends toward a block-set partition coding algorithm as k,p increases. Location bit, amplitude bit, and unnecessary bit are presented, which can be used to analyze the coding efficiency of SPACS. We compress 256 images with 512×512 using SPACS. The numerical results show SPACS achieves some improvements in coding efficiency over SPIHT, especially at very low bitrate. On average, to code every image, SPACS(3,1) (at an average of 3.93 bpp) needs 7792 more location bits but saves 10 218 unnecessary bits, compared with SPIHT (3.94 bpp).
Qiufu Li, Derong Chen, Bingtai Liu, Jiulu Gong
IEEE Trans. Image Process.1