Fu Li 0002

dblp:37/4556-2 · DBLP profile ↗
← Back
36ranked-venue papers
4as first author
21since 2021 · last 2026
0000-0003-0319-0308ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 18 · 4 first-author · 6 since 2021Artificial intelligence and machine learning · 15 · 13 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 3 since 2021
YearPublicationVenuePosition
2026 NVFusion: Lightweight infrared and low-light night vision image fusion in dark environments
Fu Li 0002, Zhifu Zhao
Neurocomputing2
2026 A lightweight semantic decoding network with group to individual transfer learning for EEG-based visual recognition
Xiaotian Wang 0001, Doudou Zhang, Qimin Xu, Rongkai Zhang 0008, Yiming Jiang 0023, Fu Li 0002, Yang Li 0019, Guangming Shi
Neurocomputing7
2026 Temporal-adaptive sampling and artifact-guided reconstruction for compressed video super-resolution
Fu Zou, Mingming Ma, Qingyu Luo, Fu Li 0002
Neurocomputing5
2026 Two-phase collaborative model compression training for joint pruning and quantization
Chunxiao Fan 0002, Zhongqian Zhang, Fu Li 0002
Neural Networks4
2026 VGRF Signal-Based Gait Analysis for Parkinson's Disease Detection: A Multi-Scale Directed Graph Neural Network Approach
abstract
Parkinson's Disease (PD) is often characterized by abnormal gait patterns, which can be objectively and quantitatively diagnosed using Vertical Ground Reaction Force (VGRF) signals. Previous studies have demonstrated the effectiveness of deep learning in VGRF signal analysis. However, the inherent graph structure of VGRF signals has not been adequately considered, limiting the representation of dynamic gait characteristics. To address this, we propose a Multi-Scale Adaptive Directed Graph Neural Network (MS-ADGNN) approach to distinguish the gaits between Parkinson's patients and healthy controls. This method models the VGRF signal as a multi-scale directed graph, capturing the distribution relationships within the plantar sensors and the dynamic pressure conduction during walking. MS-ADGNN integrates an Adaptive Directed Graph Network (ADGN) unit and a Multi-Scale Temporal Convolutional Network (MSTCN) unit. ADGN extracts spatial features from three scales of the directed graph, effectively capturing local and global connectivity. MSTCN extracts multi-scale temporal features, capturing short to long-term dependencies. The proposed method outperforms existing methods on three widely used datasets. In cross-dataset experiments, the average improvements in terms of accuracy, F1-score, and geometric mean are 2.46$\%$, 1.25$\%$, and 1.11$\%$ respectively. Meanwhile, in 10-fold cross-validation experiments, the improvements are 0.78$\%$, 0.83$\%$, and 0.81$\%$ respectively.
Xiaotian Wang 0001, Xuanhang Xu, Zhifu Zhao, Fu Li 0002, Fei Qi 0001, Shuo Liang
IEEE J. Biomed. Health Informatics4
2026 Improved Spontaneous EEG Signal Decoding Efficiency by Function Predefined Convolutional Neural Network
abstract
A spontaneous electroencephalogram (EEG)-based brain-computer interface (BCI) is an ideal form of brain-computer interaction. The classical decoding methods can achieve classification by using meaningful manual features, but their performance is poor. The neural network (NN) methods have significantly improved the performance, but their interpretability and computational efficiency are much lower than those of the classical methods. This is because NN abandons the strong a priori knowledge of neuroscience and completely relies on training to extract EEG features. How to integrate the characteristics of neural signals into the design of the basic operator of the NNs while retaining its learning ability is the focus of this work. In this work, we proposed a function predefined convolutional NN (FPCNN) to search for the best frequency points and channel weights to decode spontaneous EEG signals. Among the FPCNN, a novel function predefined convolutional (FPC) layer adopts a learnable way to search for the key spatial-frequency parameters of spontaneous EEG, making its parameters have clear physical meanings. Furthermore, a trainable quadrature detector (TQD) based on FPC was constructed, and the quadrature characteristic was utilized to ensure the capture of complex phase change signals. The core contribution of our method lies in the proposal of a novel NN operator for decoding spontaneous EEG, and a quadrature scheme for handling the phase changes of signals. The experimental results show that the proposed FPCNN significantly improves the performance by 2.09% ( ${}^{\ast } $ ), 3.08% ( ${}^{\ast } $ ), and 3.41% ( ${}^{\ast \ast }$ ), respectively, compared with the state-of-the-art (SOTA) methods on three spontaneous EEG datasets. Moreover, the training and testing time cost of FPCNN in a non-GPU environment only takes 67.96 and 19.36 s per epoch. Its savings in computing resources and time are very beneficial for EEG processing in diverse environments. In addition, visualization experiments demonstrated the interpretability and stability of the proposed FPCNN. The experimental results show that our method is efficient, stable, and interpretable. This work has effectively improved the decoding efficiency of spontaneous EEG signals and demonstrated the power of combining traditional signal processing methods with NNs.
Boxun Fu, Fu Li 0002, Junkai Li, Youshuo Ji, Yang Li 0019, Yinghui Quan, Lijian Zhang, Guangming Shi
IEEE Trans. Neural Networks Learn. Syst.2
2025 Adaptive Progressive Attention Graph Neural Network for EEG Emotion Recognition
abstract
In recent years, numerous neuroscientific studies have demonstrated that specific brain regions are associated with human emotional responses, with these regions exhibiting variability across individuals and emotions. To effectively leverage these neural patterns, we propose an Adaptive Progressive Attention Graph Neural Network (APAGNN), which dynamically models the spatial relationships among brain regions during emotional processing. APAGNN employs three specialized expert modules that progressively analyze brain topology. The first expert captures global brain connectivity patterns, the second extracts localized regional features, and the third focuses on emotion-related channel interactions. This progressive refinement strategy enables hierarchical feature extraction from coarsegrained to fine-grained neural representations. Furthermore, a weight generator integrates the outputs of all three experts, adaptively balancing their contributions for final emotion recognition. Extensive experiments conducted on SEED, SEED-IV and MPED datasets demonstrate that our method significantly improves EEG emotion recognition performance, achieving superior results compared to baseline methods.
Tianzhi Feng, Chennan Wu, Fu Li 0002, Yang Li 0019, Boxun Fu, Zhifu Zhao, Xiaotian Wang 0001
BIBM4
2025 IMFR-Net: Interval Measurement and Full Recovery Network for video compressive sensing
Wanxin Zhang, Zhifu Zhao, Fu Li 0002, Jianan Li 0003
Neurocomputing4
2025 Enhancing Cross-Dataset EEG Emotion Recognition: A Novel Approach With Emotional EEG Style Transfer Network
abstract
Electroencephalogram (EEG)-based emotion recognition has achieved remarkable success in both subject-dependent and subject-independent scenarios. However, overcoming the challenges associated with reduced performance in EEG emotion recognition across devices, time, space, and subjects (i.e., cross-dataset) remains a significant obstacle for affective brain-computer interfaces (aBCIs). The key issue lies in the distributional mismatch between source and target domain EEG signals. To tackle the significant inter-domain differences in cross-dataset EEG emotion recognition, this paper introduces an innovative framework termed the Emotional EEG Style Transfer Network (E$^{2}$STN), which aims to effectively capture the emotional content information from the source domain and the style features from the target domain, facilitating the reconstruction of stylized emotion EEG representations. These stylized EEG representations significantly enhance the discriminative prediction performance in cross-dataset EEG emotion recognition. Specifically, E$^{2}$STN consists of three key modules: a Transfer Module for domain style transfer, a Transfer Evaluation Module for evaluating transfer quality, and a Discriminative Module for making discriminative predictions. Extensive experiments demonstrate that E$^{2}$STN achieves the state-of-the-art performance in cross-dataset emotion EEG recognition. To the best of our knowledge, this is the first work to explicitly address cross-dataset emotion EEG recognition. The experimental results provide a valuable benchmark for future research in this area.
Yijin Zhou, Fu Li 0002, Yang Li 0019, Youshuo Ji, Lijian Zhang, Yuanfang Chen, Huaning Wang
IEEE Trans. Affect. Comput.2
2024 Learning from the Web: Language Drives Weakly-Supervised Incremental Learning for Semantic Segmentation
Chang Liu 0047, Giulia Rizzoli, Pietro Zanuttigh, Fu Li 0002
ECCV (17)4
2024 A novel hybrid decoding neural network for EEG signal representation
Youshuo Ji, Fu Li 0002, Boxun Fu, Yijin Zhou, Yang Li 0019, Xiaoli Li 0002, Guangming Shi
Pattern Recognit.2
2023 Structure guided network for human pose estimation
Xuemei Xie, Bo'ao Li, Fu Li 0002
Appl. Intell.5
2023 Progressive graph convolution network for EEG emotion recognition
Yijin Zhou, Fu Li 0002, Yang Li 0019, Youshuo Ji, Guangming Shi, Wenming Zheng, Lijian Zhang, Yuanfang Chen, Rui Cheng 0010
Neurocomputing2
2023 GMSS: Graph-Based Multi-Task Self-Supervised Learning for EEG Emotion Recognition
abstract
Previous electroencephalogram (EEG) emotion recognition relies on single-task learning, which may lead to overfitting and learned emotion features lacking generalization. In this paper, a graph-based multi-task self-supervised learning model (GMSS) for EEG emotion recognition is proposed. GMSS has the ability to learn more general representations by integrating multiple self-supervised tasks, including spatial and frequency jigsaw puzzle tasks, and contrastive learning tasks. By learning from multiple tasks simultaneously, GMSS can find a representation that captures all of the tasks thereby decreasing the chance of overfitting on the original task, i.e., emotion recognition task. In particular, the spatial jigsaw puzzle task aims to capture the intrinsic spatial relationships of different brain regions. Considering the importance of frequency information in EEG emotional signals, the goal of the frequency jigsaw puzzle task is to explore the crucial frequency bands for EEG emotion recognition. To further regularize the learned features and encourage the network to learn inherent representations, contrastive learning task is adopted in this work by mapping the transformed data into a common feature space. The performance of the proposed GMSS is compared with several popular unsupervised and supervised methods. Experiments on SEED, SEED-IV, and MPED datasets show that the proposed model has remarkable advantages in learning more discriminative and general features for EEG emotional signals.
Yang Li 0019, Fu Li 0002, Boxun Fu, Youshuo Ji, Yijin Zhou, Guangming Shi, Wenming Zheng
IEEE Trans. Affect. Comput.3
2023 Toward Interactive Self-Supervised Denoising
abstract
Self-supervised denoising frameworks have recently been proposed to learn denoising models without noisy-clean image pairs, showing great potential in various applications. The denoising model is expected to produce visually pleasant images without noise patterns. However, it is non-trivial to achieve this goal using self-supervised methods because 1) the self-supervised model is difficult to restore the perceptual information due to the lack of clean supervision, and 2) perceptual quality is relatively subjective to users’ preferences. In this paper, we make the first attempt to build an interactive self-supervised denoising model to tackle the aforementioned problems. Specifically, we propose an interactive two-branch network to effectively restore perceptual information. The network consists of a denoising branch and an interactive branch, where the former focuses on efficient denoising, and the latter modulates the denoising branch. Based on the delicate architecture design, our network can produce various denoising outputs, allowing the user to easily select the most appealing outcome for satisfying the perceptual requirement. Moreover, to optimize the network with only noisy images, we propose a novel two-stage training strategy in a self-supervised way. Once the network is optimized, it can be interactively changed between noise reduction and texture restoration, providing more denoising choices for users. Existing self-supervised denoising methods can be integrated into our method to be user-friendly with interaction. Extensive experiments and comprehensive analyses are conducted to validate the effectiveness of the proposed method.
Mingde Yao, Dongliang He, Xin Li 0106, Fu Li 0002, Zhiwei Xiong
IEEE Trans. Circuits Syst. Video Technol.4
2022 ACO: lossless quality score compression based on adaptive coding order
abstract
BACKGROUND: With the rapid development of high-throughput sequencing technology, the cost of whole genome sequencing drops rapidly, which leads to an exponential growth of genome data. How to efficiently compress the DNA data generated by large-scale genome projects has become an important factor restricting the further development of the DNA sequencing industry. Although the compression of DNA bases has achieved significant improvement in recent years, the compression of quality score is still challenging. RESULTS: In this paper, by reinvestigating the inherent correlations between the quality score and the sequencing process, we propose a novel lossless quality score compressor based on adaptive coding order (ACO). The main objective of ACO is to traverse the quality score adaptively in the most correlative trajectory according to the sequencing process. By cooperating with the adaptive arithmetic coding and an improved in-context strategy, ACO achieves the state-of-the-art quality score compression performances with moderate complexity for the next-generation sequencing (NGS) data. CONCLUSIONS: The competence enables ACO to serve as a candidate tool for quality score compression, ACO has been employed by AVS(Audio Video coding Standard Workgroup of China) and is freely available at https://github.com/Yoniming/ACO.
Mingming Ma, Fu Li 0002, Guangming Shi
BMC Bioinform.3
2022 NL-CALIC Soft Decoding Using Strict Constrained Wide-Activated Recurrent Residual Network
Chang Liu 0047, Mingming Ma, Fu Li 0002, Zhiwen Chen 0002, Guangming Shi
IEEE Trans. Image Process.4
2021 A novel transferability attention neural network model for EEG emotion recognition
Yang Li 0019, Boxun Fu, Fu Li 0002, Guangming Shi, Wenming Zheng
Neurocomputing3
2021 Conditional generative adversarial network for EEG-based emotion fine-grained estimation and visualization
Boxun Fu, Fu Li 0002, Yang Li 0019, Guangming Shi
J. Vis. Commun. Image Represent.2
2021 A novel lossless compression framework for facial depth images in expression recognition
Chunxiao Fan 0002, Fu Li 0002, Xueliang Liu
Multim. Tools Appl.2
2021 Quality Index for View Synthesis by Measuring Instance Degradation and Global Appearance
abstract
Virtual view synthesis plays a vital role in the application of multi-view and free-viewpoint videos. Depth-image-based rendering (DIBR) is the most commonly used approach in view synthesis, and many DIBR algorithms have been proposed. However, how to evaluate the quality of DIBR-synthesized images and benchmark the DIBR algorithms are still very challenging, which may hinder the further development of the view synthesis technique. Hence, an effective quality metric for evaluating the distortions in view synthesis is urgently needed. With this motivation, this paper presents a quality index for view synthesis by simultaneously measuring local Instance DEgradation and global Appearance (IDEA). Due to the imperfection of rendering algorithms, local geometric distortions are easily introduced around instance contours, causing instance degradation, which is the dominant distortion in synthesized views. In this work, image instances are first detected and local instance degradation is measured based on discrete orthogonal moments. Meantime, we propose to measure the global appearance of synthesized images based on the superpixel representation. By integrating both local and global aspects of the distortions, a more accurate quality model is built for view synthesis. Extensive experiments and comparisons have demonstrated the superiority of the proposed method in evaluating the quality of DIBR-synthesized images and benchmarking the performance of view synthesis algorithms.
Leida Li, Yu Zhou 0009, Jinjian Wu, Fu Li 0002, Guangming Shi
IEEE Trans. Multim.4
2020 Channel-Grouping Based Patch Swap For Arbitrary Style Transfer
abstract
The basic principle of the patch-matching based style transfer is to substitute the patches of the content image feature maps by the closest patches from the style image feature maps. Since the finite features harvested from one single aesthetic style image are inadequate to represent the rich textures of the content natural image, existing techniques treat the full-channel style feature patches as simple signal tensors and create new style feature patches via signal-level fusion. In this paper, we propose a channel-grouping based patch swap technique to group the style feature maps into surface and texture channels, and the new features are created by the combination of these two groups, which can be regarded as a semantic-level fusion of the raw style features. Experimental results demonstrate that the proposed method outperforms the existing techniques in providing more style-consistent textures while keeping the content fidelity.
Fu Li 0002, Chunbo Zou, Guangming Shi
ICIP3
2020 Network pruning using sparse learning and genetic algorithm
Zhenyu Wang 0008, Fu Li 0002, Guangming Shi, Xuemei Xie, Fangyu Wang
Neurocomputing2
2019 Facial Attention based Convolutional Neural Network for 2D+3D Facial Expression Recognition
abstract
Discriminative facial parts are essential for facial expression recognition (FER) tasks because of small inter-class differences and large intra-class variations in expression images. Existing methods localize discriminative regions with the aid of extra facial landmarks, such as action units (AU). However, it consumes a lot of manpower in manually labeling. To address this problem, in this paper, we propose an advanced facial attention based convolutional neural network (FA-CNN) for 2D+3D FER. The main contribution of FA-CNN is the facial attention mechanism, which enables the network to localize the discriminative regions automatically from multi-modality expression images without dense landmark annotations. Experimental results conducted on BU-3DFE demonstrate that FA-CNN achieves state-of-the-art performance comparing with the existing 2D+3D FER techniques, and the discriminative facial parts estimated by the facial attention mechanism are highly interpretable and consistent with human perception.
Fu Li 0002, Chunbo Zou, Guangming Shi
VCIP4
2019 Depth acquisition with the combination of structured light and deep learning stereo matching
Fu Li 0002, Quanlu Li, Guangming Shi
Signal Process. Image Commun.1
2018 Depth sensing with coding-free pattern based on topological constraint
Guangming Shi, Ruodai Li, Fu Li 0002
J. Vis. Commun. Image Represent.3
2018 Dynamic Range Reduction of SAR Image via Global Optimum Entropy Maximization With Reflectivity-Distortion Constraint
abstract
The visualization of synthetic aperture radar (SAR) images plays a critical role in remote sensing applications. To effectively obtain the image suitable for human observation, this paper introduces a new SAR image visualization algorithm to map the high dynamic range SAR amplitude values to low dynamic range displays via reflectivity distortion preserved entropy maximization. Its designed objective is to present the maximal amount of information content in the displayed image, and being optimal in an information theoretical sense, as well as restricting the upper bound of the reflection distortion caused by tone mapping. The resulting optimization problem can be graph theoretically modeled as a K-edges maximum weight path problem in a directed acyclic graph, and it can be solved efficiently by dynamic programming in real time. Empirical evidences are provided to demonstrate the superior visual quality obtained by our new visualization technique.
Guanghui Zhao 0003, Guangming Shi, Fu Li 0002
IEEE Trans. Geosci. Remote. Sens.6
2017 Single-shot dense depth sensing with frequency-division multiplexing fringe projection
Fu Li 0002, Zhiwei Xiong, Guangming Shi, Ruodai Li
J. Vis. Commun. Image Represent.2
2017 A hierarchical multiplier-free architecture for HEVC transform
Chunxiao Fan 0002, Fu Li 0002, Guangming Shi, Fei Qi 0001, Xuemei Xie, Dandan Jiao
Multim. Tools Appl.2
2017 An AR based fast mode decision for H.265/HEVC intra coding
Fu Li 0002, Dandan Jiao, Guangming Shi, Chunxiao Fan 0002, Xuemei Xie
Multim. Tools Appl.1
2016 High quality impulse noise removal via non-uniform sampling and autoregressive modelling based super-resolution
abstract
The challenge of image impulse noise removal is to restore spatial details from damaged pixels using remaining ones in random locations. Most existing methods use all uncontaminated pixels within a local window to estimate the centred noisy one via a statistic way. These kinds of methods have two defects. First, all noisy pixels are treated as independent individuals and estimated by their neighbours one by one, with the correlation between their true values ignored. Second, the image structure as a natural feature is usually ignored. This study proposes a new denoising framework, in which all noisy pixels are jointly restored via non‐uniform sampling and supervised piecewise autoregressive modelling based super‐resolution. In this method, the noisy pixels are jointly estimated in groups through solving a well‐designed optimisation problem, in which image structure feature is considered as an important constraint. Another contribution is that piecewise autoregressive model is not simply adopted but carefully designed so that all noise‐free pixels can be used to supervise the model training and optimisation problem solving for higher accuracy. The experimental results demonstrate that the proposed method exhibits good denoising performance in a large noise density range (10–90%).
Xiaotian Wang 0001, Guangming Shi, Jinjian Wu, Fu Li 0002, Yantao Wang
IET Image Process.5
2013 Dense depth acquisition via one-shot stripe structured light
abstract
Depth acquisition for moving objects becomes increasingly critical for some applications such as human facial expression recognition. This paper presents a method for capturing the depth maps of moving objects that uses a one-shot black-and-white stripe pattern with the features of simplicity and easily generation. Considering the accuracy of a matching is crucial for a precise depth map but the matching of variant-width stripes is sparse and rough, the phase differences extracted by Gabor filter to achieve a pixel-wise matching with sub-pixel accuracy are used. The details of the derivation are presented to prove that this method based on the phase difference calculated by Gabor filter is valid. In addition, the periodic ambiguity of the encoded stripe is eliminated by the epipolar segment covering a given depth range at a camera-projector calibrating stage to decrease the calculation complexity. Experimental results show that our method can get a dense and accurate depth map of a moving object.
Fu Li 0002, Guangming Shi, Fei Qi 0001, Yuexin Shi
VCIP2
2013 Structure guided fusion for depth map inpainting
Fei Qi 0001, Junyu Han, Pengjin Wang, Guangming Shi, Fu Li 0002
Pattern Recognit. Lett.5
2013 Pattern Masking Estimation in Image With Structural Uncertainty
abstract
A model of visual masking, which reveals the visibility of stimuli in the human visual system (HVS), is useful in perceptual based image/video processing. The existing visual masking function mainly considers luminance contrast, which always overestimates the visibility threshold of the edge region and underestimates that of the texture region. Recent research on visual perception indicates that the HVS is sensitive to orderly regions that possess regular structures and insensitive to disorderly regions that possess uncertain structures. Therefore, structural uncertainty is another determining factor on visual masking. In this paper, we introduce a novel pattern masking function based on both luminance contrast and structural uncertainty. Through mimicking the internal generative mechanism of the HVS, a prediction model is firstly employed to separate out the unpredictable uncertainty from an input image. In addition, an improved local binary pattern is introduced to compute the structural uncertainty. Finally, combining luminance contrast with structural uncertainty, the pattern masking function is deduced. Experimental result demonstrates that the proposed pattern masking function outperforms the existing visual masking function. Furthermore, we extend the pattern masking function to just noticeable difference (JND) estimation and introduce a novel pixel domain JND model. Subjective viewing test confirms that the proposed JND model is more consistent with the HVS than the existing JND models.
Jinjian Wu, Weisi Lin, Guangming Shi, Xiaotian Wang 0001, Fu Li 0002
IEEE Trans. Image Process.5
2011 An efficient VLSI architecture for 4×4 intra prediction in the High Efficiency Video Coding (HEVC) standard
abstract
Intra prediction with fine directions is a critical feature in the new High Efficiency Video Coding (HEVC) standard because it provides significant performance gain. Different from the intra prediction in the H.264/AVC, this approach is more complicated in terms of computation and memory access, which makes the VLSI design very difficult. In this paper, we propose an efficient uniform architecture for all of the 4×4 intra directional modes. The architecture is implemented by a register array and a flexible reference sample selection technique. This novel architecture does not need to project the samples from the side reference to the main reference. Thus, it reduces the processing latency and the number of registers considerably. The proposed architecture has been implemented with TSMC 0.13μm CMOS technology. Simulation results show that the proposed architecture only needs 9020 logic gates for 17 directional modes and can run at 150 MHz operation frequency.
Fu Li 0002, Guangming Shi, Feng Wu 0001
ICIP1
2011 A pipelined architecture for 4×4 intra frame mode decision in the high efficiency video coding
abstract
Mode decision in High Efficient Video Coding (HEVC) is occupied more than half of the computational complexity in intra frame coding. Block size of 4×4 is the most frequently used block in HM. In this paper, we proposed a pipelined architecture for the 4×4 intra frame mode decision in HEVC to improve the computational capability. This novel architecture consists of six-stage pipelines, and each of the pipelines can be accomplished within 24 clock cycles. In the pipeline of prediction procedure, we proposed a folded project-skip architecture for prediction. It can save the processing latency and the registers considerably. We also proposed a simplified CAVLC with low complexity in the pipeline of bits estimation procedure. The architecture for mode decision has been evaluated with TSMC 0.13μm CMOS technology. Synthesized results show that the proposed architecture only needs 99K logic gates for modes decision and can run at 165 MHz operation frequency.
Fu Li 0002, Guangming Shi
MMSP1