EDBT 2026 Demo / reviewers in the wild / expert
Xiaofei Zhou 0003
dblp:58/1065-3
· DBLP profile ↗
83ranked-venue papers
18as first author
62since 2021 · last 2026
0000-0002-7977-9728ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 52 · 11 first-author · 37 since 2021Artificial intelligence and machine learning · 29 · 4 first-author · 23 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 first-author · 4 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Divide-and-Conquer Decoupled Network for Cross-Domain Few-Shot SegmentationabstractCross-domain few-shot segmentation (CD-FSS) aims to tackle the dual challenge of recognizing novel classes and adapting to unseen domains with limited annotations. However, encoder features often entangle domain-relevant and category-relevant information, limiting both generalization and rapid adaptation to new domains. To address this issue, we propose a Divide-and-Conquer Decoupled Network (DCDNet). In the training stage, to tackle feature entanglement that impedes cross-domain generalization and rapid adaptation, we propose the Adversarial-Contrastive Feature Decomposition (ACFD) module. It decouples backbone features into category-relevant private and domain-relevant shared representations via contrastive learning and adversarial learning. Then, to mitigate the potential degradation caused by the disentanglement, the Matrix-Guided Dynamic Fusion (MGDF) module adaptively integrates base, shared, and private features under spatial guidance, maintaining structural coherence. In addition, in the fine-tuning stage, to enhanced model generalization, the Cross-Adaptive Modulation (CAM) module is placed before the MGDF, where shared features guide private features via modulation ensuring effective integration of domain-relevant information. Extensive experiments on four challenging datasets show that DCDNet outperforms existing CD-FSS methods, setting a new state-of-the-art for cross-domain generalization and few-shot adaptation. Runmin Cong, Anpeng Wang, Bin Wan, Xiaofei Zhou 0003 |
AAAI | 5 |
| 2026 | SAM-DAQ: Segment Anything Model with Depth-guided Adaptive Queries for RGB-D Video Salient Object DetectionabstractRecently segment anything model (SAM) has attracted widespread concerns, and it is often treated as a vision foundation model for universal segmentation. Some researchers have attempted to directly apply the foundation model to the RGB-D video salient object detection (RGB-D VSOD) task, which often encounters three challenges, including the dependence on manual prompts, the high memory consumption of sequential adapters, and the computational burden of memory attention. To address the limitations, we propose a novel method, namely Segment Anything Model with Depth-guided Adaptive Queries (SAM-DAQ), which adapts SAM2 to pop-out salient objects from videos by seamlessly integrating depth and temporal cues within a unified framework. Firstly, we deploy a parallel adapter-based multi-modal image encoder (PAMIE), which incorporates several depth-guided parallel adapters (DPAs) in a skip-connection way. Remarkably, we fine-tune the frozen SAM encoder under prompt-free conditions, where the DPA utilizes depth cues to facilitate the fusion of multi-modal features. Secondly, we deploy a query-driven temporal memory (QTM) module, which unifies the memory bank and prompt embeddings into a learnable pipeline. Concretely, by leveraging both frame-level queries and video-level queries simultaneously, the QTM module can not only selectively extract temporal consistency features but also iteratively update the temporal representations of the queries. Extensive experiments are conducted on three RGB-D VSOD datasets, and the results show that the proposed SAM-DAQ consistently outperforms state-of-the-art methods in terms of all evaluation metrics. Xiaofei Zhou 0003, Runmin Cong, Guodao Zhang, Zhi Liu 0003, Jiyong Zhang 0001 |
AAAI | 2 |
| 2026 | Low-light light field image enhancement based on illumination-guided implicit gradient representation
Deyang Liu, Xiaofei Zhou 0003, Ping An 0001, Caifeng Shan, Hongbin Zha |
Neurocomputing | 3 |
| 2026 | Multiscale Guided Cross-Stage Refinement SAM2 for Seamless Steel Pipe Internal Surface Defect DetectionabstractIn recent years, vision foundation models (e.g., SAM2) have achieved remarkable success across various downstream tasks and have been increasingly adopted in Internet of Things (IoT)-enabled visual perception systems. However, the decoders of existing vision foundation models fail to effectively handle multi-scale contextual cues, leading to unsatisfactory performance in detecting defects of varying sizes in IoT-based inspection scenarios. Besides, the insufficient preservation of spatial information degrades detection accuracy, especially when handling small defect regions, which is critical for reliable perception in IoT devices. To address these issues, we propose a Multi-scale Guided Cross-stage Refinement SAM2 (MGCR-SAM2) for seamless steel pipe (SSP) internal surface defect detection. Specifically, a Multi-scale Information Compensation (MIC) module is designed for multi-scale feature aggregation in the decoder, which integrates contextual information from different receptive fields to enhance the representation of defects with diverse scales. Furthermore, a cross-stage refinement strategy is adopted to progressively enhance detection precision. In the first stage, we construct an encoder–decoder architecture based on SAM2 and the proposed MIC module to adequately capture multi-scale features from SSP images for coarse localization. In the second stage, a Spatial Information Compensation (SIC) module and CNN-based refinement network are employed to enhance local details and refine coarse predictions. Extensive experiments on SSP2000 and SD-Saliency-900 datasets demonstrate the effectiveness of our MGCR-SAM2. Code is available at https://github.com/Kunye-Shen/MGCR-SAM2. Kunye Shen, Xiaofei Zhou 0003, Zhi Liu 0003 |
IEEE Internet Things J. | 2 |
| 2026 | Pseudo-label enhanced consistency learning for semi-supervised medical image segmentation
Mei Xiao, Qingshan She, Songkai Sun, Xiaofei Zhou 0003, Yingchun Zhang |
Pattern Anal. Appl. | 4 |
| 2026 | A benchmark for robust salient object detection in adverse weather conditions
Bolun Zheng, Rongfeng Lu, Xiaokai Yang, Qianyu Zhang 0002, Yu Liu 0005, Xiaofei Zhou 0003 |
Pattern Recognit. | 8 |
| 2026 | SurfSyn: Domain foundation model for pixel-level surface defect detection driven by synthetic data
Kunye Shen, Xiaofei Zhou 0003, Zhi Liu 0003 |
Pattern Recognit. | 2 |
| 2026 | Dense multiscale inference network for lightweight salient object detection of strip steel surface defects
Yihan Qiu, Xiaofei Zhou 0003, Yong Wu 0007, Bin Wan, Juting Miu, Zhangping Chen, Deyang Liu |
Pattern Recognit. Lett. | 2 |
| 2026 | Global and local collaborative learning for no-reference omnidirectional image quality assessment
Deyang Liu, Lifei Wan, Xiaofei Zhou 0003, Caifeng Shan |
Signal Process. Image Commun. | 4 |
| 2026 | Text-Image Guided Retrieval Network for Triple-Modal Images Few-Shot Semantic SegmentationabstractThis letter presents a novel framework for triple-modal few-shot semantic segmentation involving visible, depth, and thermal modalities. While multi-modal approaches aim to provide complementary data, existing models often suffer from a heavy reliance on the visual appearance of support samples and lack effective semantic guidance during cross-modal fusion. To address these issues, we build a Text-Image Guided Retrieval Network (TIGRNet), where we leverage a Vision-Language Model (VLM) to inject semantic priors into the segmentation process. Specifically, we propose a Text Guided Fusion (TGF) module that dynamically regulates modality weights using global semantic cues, thereby suppressing noise and enhancing reliable features. Furthermore, we introduce a dual-branch retrieval strategy consisting of an Image-Guided Retrieval(IGR) module to capture visual appearance and a Text-Guided Retrieval(TGR) module to retrieve regions consistent with target semantics, reducing the burden of precise visual matching. Experiments on the VDT-2048-5$^{i}$dataset demonstrate that our TIGRNet outperforms existing state-of-the-art models in both 1-shot and 5-shot settings. The dataset and code are available athttps://github.com/FengshuoChen/TIGRNet. Fengshuo Chen, Liuxin Bao, Juting Miao, Xiaofei Zhou 0003 |
IEEE Signal Process. Lett. | 5 |
| 2026 | Degradation-Aware Blind Light-Field Image Quality Assessment With Linear AttentionabstractBlind Light-Field Image Quality Assessment (LFIQA) is challenging, as degradations are microlens-dependent and spatially non-uniform, while perceptual quality relies on both spatial fidelity and angular consistency. However, many existing methods either assume globally stationary distortions or adopt global pooling or self-attention, which can be biased by locally corrupted lenslets and become computationally prohibitive when modeling long-range spatial-angular dependencies. Therefore, in this letter, we present a degradation-aware framework that first predicts a microlens reliability map to make quality inference robust to spatially non-uniform, lenslet-varying corruption. It then extracts spatial and angular features, applies a shared Receptance Weighted Key Value (RWKV) module for linear-time long-range context, and fuses them to predict perceptual quality. Experiments show a higher correlation with subjective ratings with competitive efficiency. Youzhi Zhang 0004, Jianyu Qian, Deyang Liu, Xiaofei Zhou 0003, Hongbin Zha, Caifeng Shan |
IEEE Signal Process. Lett. | 4 |
| 2026 | Learning Implicit and Detail-Enhanced Network for Light Field Image Spatial-Angular Super-ResolutionabstractLight field (LF) imaging holds immense promise for applications such as post-capture refocusing and virtual reality. However, its inherent spatial-angular trade-off significantly limits both spatial and angular resolution, restricting its practicality in real-world scenarios. To address these limitations, spatial-angular super-resolution methods have been proposed to simultaneously enhance both dimensions. Yet, existing methods struggle to fully exploit the intertwined spatial-angular correlations and fail to effectively handle sparsely sampled LFs with low spatial resolution, often leading to cumulative errors during reconstruction. In this paper, we propose an Implicit and Detail-Enhanced Network (IDNet) to overcome these challenges. Our IDNet employs 3D convolution for the joint extraction of spatial and angular information, leveraging their interdependencies for more effective LF reconstruction. Additionally, we introduce an implicit detail restoration module that enhances features while encoding positional information to refine fine details. To overcome the limitations of sparse spatial and angular information on high-detail reconstruction and angular consistency in low-resolution LFs, we design a multi-representation enhancement block. This block enhances features by learning pixel differences across multiple directions in diverse representations, effectively capturing intricate details and complex correlations. Thanks to these designs, our IDNet reconstructs novel views with finer details, effectively learns occlusion relationships, and ensures geometric consistency. Experimental results on benchmark datasets demonstrate its superior quantitative and qualitative performance. The code is publicly available at https://github.com/ldyorchid/IDNet. Deyang Liu, Shizheng Li, Xiaofei Zhou 0003, Zeyu Xiao 0002, Caifeng Shan |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | Class Knowledge-Guided Lightweight Network for Salient Object Detection of Strip Steel Surface DefectabstractRecent lightweight salient object detection (SOD) models for strip steel surface defect typically adopt classification model as encoder to extract semantic features, followed by task-specific decoder for defect detection. Despite achieving promising results, these models often overlook the potential of leveraging class-level knowledge to enhance detection performance. In this paper, we propose a Class Knowledge-Guided Lightweight Network (CKLNet) for salient object detection of strip steel surface defect. Particularly, CKLNet employs a three-stage training strategy. In the first stage, a general SOD model is trained to perform initial detection of surface defect. Then, the second stage constructs a classification branch built upon the general SOD model to extract defect class knowledge. Embarking on this class knowledge, multiple class-specific decoders are trained in the third stage, where we can generate refined and class-aware defect predictions. Extensive experiments on two datasets demonstrate that CKLNet achieves superior detection performance and generalization capability when compared with the cutting-edge lightweight models, with real-time inference speed of 783 FPS on an RTX 2080Ti GPU. Moreover, unlike existing models that only enable defect localization, our CKLNet further provides accurate defect class identification. The source code is publicly available at https://github.com/Kunye-Shen/CKLNet. Kunye Shen, Xiaofei Zhou 0003, Zhi Liu 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | G2HFNet: GeoGran-Aware Hierarchical Feature Fusion Network for Salient Object Detection in Optical Remote Sensing ImagesabstractRemote sensing images captured from aerial perspectives often exhibit significant scale variations and complex backgrounds, posing challenges for salient object detection (SOD). Existing methods typically extract multi-level features at a single scale using uniform attention mechanisms, leading to suboptimal representations and incomplete detection results. To address these issues, we propose a GeoGran-Aware Hierarchical Feature Fusion Network (G2HFNet) that fully exploits geometric and granular cues in optical remote sensing images. Specifically, G2HFNet adopts Swin Transformer as the backbone to extract multi-level features and integrates three key modules: the multi-scale detail enhancement (MDE) module to handle object scale variations and enrich fine details, the dual-branch geo-gran complementary (DGC) module to jointly capture fine-grained details and positional information in mid-level features, and the deep semantic perception (DSP) module to refine high-level positional cues via self-attention. Additionally, a local-global guidance fusion (LGF) module is introduced to replace traditional convolutions for effective multi-level feature integration. Extensive experiments demonstrate that G2HFNet achieves high-quality saliency maps and significantly improves detection performance in challenging remote sensing scenarios. Bin Wan, Runmin Cong, Xiaofei Zhou 0003, Hao Fang 0010, Chengtao Lv, Sam Kwong |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | RSONet: Region-Guided Selective Optimization Network for RGB-T Salient Object DetectionabstractThis paper focuses on the inconsistency in salient regions between RGB and thermal images. To address this issue, we propose the Region-guided Selective Optimization Network for RGB-T Salient Object Detection, which consists of the region guidance stage and saliency generation stage. In the region guidance stage, three parallel branches with same encoder-decoder structure equipped with the context interaction (CI) module and spatial-aware fusion (SF) module are designed to generate the guidance maps which are leveraged to calculate similarity scores. Then, in the saliency generation stage, the selective optimization (SO) module fuses RGB and thermal features based on the previously obtained similarity values to mitigate the impact of inconsistent distribution of salient targets between the two modalities. After that, to generate high-quality detection result, the dense detail enhancement (DDE) module which adopts the multiple dense connections and visual state space blocks is applied to low-level features for optimizing the detail information. In addition, the mutual interaction semantic (MIS) module is placed in the high-level features to dig the location cues by the mutual fusion strategy. We conduct extensive experiments on the RGB-T dataset, and the results demonstrate that the proposed RSONet achieves competitive performance against 27 state-of-the-art SOD methods. Bin Wan, Runmin Cong, Xiaofei Zhou 0003, Hao Fang 0010, Chengtao Lv, Sam Kwong |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | Reparameterization-Driven Depthwise Separable Large-Kernel Network for Lightweight Salient Object Detection of Strip Steel Surface DefectsabstractWith the rapid development of neural networks, strip steel surface defect detection, as an important task in computer vision, has achieved remarkable progress. However, state-of-the-art methods still face a tradeoff between accuracy and efficiency. High-performing models are usually large and computationally expensive, whereas lightweight models often suffer from limited detection accuracy. To address this issue, we first propose a spatial channel enhancement (SCE) module, which consists of a reparameterizable depthwise large-kernel convolution and a reparameterizable pointwise (RepPw) convolution. The proposed SCE module enlarges the receptive field and strengthens long-range spatial and channel interactions while preserving computational efficiency. Based on the SCE module, we propose a novel lightweight saliency model for strip steel surface defects, namely, reparameterization-driven depthwise separable large-kernel network (RepDSLKNet). RepDSLKNet employs SCE modules to build an encoder and a decoder, and utilizes cascaded channel attention (CCA) modules for the feature fusion. The lightweight architecture can effectively extract and fuse the semantic information and detailed features of strip steel surface defects, thereby improving the accuracy and speed of detection with a small model size. With an input size of $224 \times 224$ , our RepDSLKNet has only 0.47 M parameters and 0.42 G FLOPs during inference. Compared to the current state-of-the-art methods, our approach achieves a 19-fold improvement in throughput and a twofold reduction in latency. Experiments on two public strip steel defect datasets demonstrate that RepDSLKNet delivers competitive performance against state-of-the-art methods. Xiaofei Zhou 0003, Zhenkun Mo, Gongyang Li, Liuxin Bao, Xiaobin Xu 0002, Jiyong Zhang 0001 |
IEEE Trans. Cybern. | 1 |
| 2026 | Probabilistic-Based Learning for Joint Light Field Image Compression and Enhancement Under Low-Light ConditionsabstractLight field (LF) imaging has attracted increasing research interest in challenging illumination conditions due to its ability to provide rich spatial and angular cues. However, such data present dual challenges: 1) the inherent multi-view structure introduces substantial data redundancy, creating high demands for efficient compression; 2) the insufficient illumination leads to severe quality degradation, which weakens inter-view consistency and visual perception. To address these coupled factors, we propose a Probabilistic-based learning for joint LF image compression and enhancement under low-light conditions (PrL-LFCE). The framework unifies structure-aware compression and feature enhancement mechanisms by introducing learnable probabilistic modeling into both feature coupling and latent distribution estimation to adaptively handle the uncertainty induced by illumination degradation and compression-related information loss. Specifically, we design a probability-based multi-directional feature coupling module that dynamically balances structural preservation and redundancy reduction across multiple directionally arranged sub-aperture images. Moreover, we introduce a swin-gated enhancement module that suppresses noise and highlights structurally salient regions in compression-aware feature representations through attention-guided gating. Extensive experiments show that PrL-LFCE consistently outperforms state-of-the-art methods, achieving at least 34.86% bitrate savings while maintaining excellent visual quality, demonstrating a strong joint compression and enhancement capability. Deyang Liu, Jimin Wang, Mounir Kaaniche, Xiaofei Zhou 0003, Gangyi Jiang, Caifeng Shan |
IEEE Trans. Image Process. | 5 |
| 2026 | Scale-Invariant Feature Matching Network for V-D-T Few-Shot Semantic SegmentationabstractMulti-modal few-shot semantic segmentation (FSS) aims to perform dense prediction from multiple modality images including visible image, depth image, and thermal image with a few annotated samples. However, some efforts treat the three modality information equally, where they don't incorporate the inherent differences among multiple modalities. Besides, the objects vary in size greatly, and the cutting-edge matching paradigms fail to establish an effective support-query connection. Therefore, we propose a novel scale-invariant feature matching network (i.e., SFM-Net), which consists of an encoder, a feature matching block, a feature elevation block, and a decoder, to conduct visible-depth-thermal (V-D-T) few-shot semantic segmentation. Firstly, in the encoder part, after the extraction of multi-level initial features, we fuse each level's RGB feature and thermal feature, yielding the support features and the query features. Secondly, in the feature matching block, a pixel-to-patch cross-attention (PTPCA) module is deployed to explore the correlation between each level's support feature and the query feature, where the pixel-to-patch pooling (PTP-pool) units are designed to build scale-invariant relationships, generating the coarse mask for the query image. Thirdly, in the feature elevation block, we employ the prior-related fusion (PF) module to integrate the depth image with a coarse mask via the cross-attention mechanism, yielding the enhanced coarse prediction result, which is further aggregated in a bottom-up way. Finally, in the decoder, we deploy a reverse attention (RA) unit to gradually explore the complementarity between object internal regions and spatial details, and further generate the final segmentation results via conventional convolution layers. Extensive experiments are conducted on the VDT-2048- $5^{i}$ dataset, and the results show that our model outperforms the state-of-the-art methods with a large margin. Xiaofei Zhou 0003, Deyang Liu, Jiyong Zhang 0001, Runmin Cong |
IEEE Trans. Image Process. | 1 |
| 2026 | Few-Shot Strip Steel Surface Defect Segmentation via Pre-Trained Variational Auto-Encoder-Based Latent Gaussian Process RegressionabstractRecently, few-shot strip steel surface defect segmentation has received more and more concerns. However, the existing few-shot segmentation methods usually adopt the frozen encoder, which is pre-trained on the classification task and can only provide class-related knowledge. Therefore, we propose a novel method, namely pre-trained variational auto-encoder based latent gaussian process regression (LGPR), to conduct few-shot strip steel surface defect segmentation. Firstly, different from previous methods, the frozen Variational Auto-Encoder (VAE) based encoder and decoder, which are pre-trained by using the pixel-level self-supervised task (i.e., image reconstruction), can provide rich image-related knowledge. This ensures the effective characterization of defect regions. Secondly, by deploying a gaussian process regression in the latent feature space generated by the VAE-based encoder, pixel-level correlation between support features and query features can be efficiently built. This operation is non-parametric and doesn't bring any training overhead. Besides, we deploy transformer-based projectors to dig long-range contextual cues of support and query features. Extensive experiments are performed on two public datasets, and the experimental results clearly show that our model consistently outperforms the state-of-the-art models with a large margin. Both the codes and results are publicly available at https://github.com/Hlao-hub/LGPR. Xiaofei Zhou 0003, Gongyang Li, Deyang Liu, Qingshan She, Xiaobin Xu 0002, Runmin Cong |
IEEE Trans. Image Process. | 1 |
| 2025 | Decoupled Motion Expression Video SegmentationabstractMotion expression video segmentation aims to segment objects based on input motion descriptions. Compared with traditional referring video object segmentation, it focuses on motion and multi-object expressions and is more challenging. Previous works achieved it by simply injecting text information into the video instance segmentation (VIS) model. However, this requires retraining the entire model and optimization is difficult. In this work, we propose DMVS, a simple framework constructed on the existing query-based VIS model, emphasizing decoupling the task into video instance segmentation and motion expression understanding. Firstly, we use a frozen video instance segmenter to extract object-specific contexts and convert them into frame-level and video-level queries. Secondly, we interact two levels of queries with static and motion cues, respectively, to further encode visually enhanced motion expressions. Furthermore, we propose a novel query initialization strategy that uses video queries guided by classification priors to initialize motion queries, greatly reducing the difficulty of optimization. Without bells and whistles, DMVS achieves state-of-the-art performance on the MeViS dataset at a lower training cost. Extensive experiments verify the effectiveness and efficiency of our framework. Hao Fang 0010, Runmin Cong, Xiankai Lu, Xiaofei Zhou 0003, Sam Kwong, Wei Zhang 0021 |
CVPR | 4 |
| 2025 | PLGMNet: Parallel Local-Global Mamba Network for Real-Time Steel Surface Defect Detection
Chenlei Li, Xiaofei Zhou 0003, Yong Wu 0007, Deyang Liu, Jiyong Zhang 0001, Zhi Liu 0003 |
PRCV (17) | 3 |
| 2025 | Multi-modal feature integration network for Visible-Depth-Thermal salient object detection
Fengyv Cui, Xiaofei Zhou 0003, Liuxin Bao, Bin Wan, Jiyong Zhang 0001 |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | Bidirectionally Guided Multi-Scale Feature Decoding Network for High-Resolution Salient Object DetectionabstractABSTRACT With the advancement of modern camera technology, the resolution and quality of images have been significantly improved. High‐resolution images can provide more detailed and clearer information, but they may also introduce more noise, which interferes with the detection and localization of salient objects. To address this issue, existing high‐resolution salient object detection methods either design complex network structures or adopt multi‐modal fusion. However, these approaches often consume significant computing and storage resources. This leads to redundancy of irrelevant features and loss of critical details. In this paper, we propose a network called bidirectionally guided multi‐scale feature decoding network for high‐resolution salient object detection. The model incorporates a bidirectional guidance method to explore the complementarity between encoding and decoding features, thereby achieving a comprehensive combination and enhancement of features. Additionally, in the decoder, multi‐scale encoding features are obtained and utilized sequentially to enhance feature learning and improve the accuracy of salient object detection. Specifically, our model consists of an encoder, a guided multi‐scale feature enhancement (GMFE) module, a guided feature fusion (GFF) module, and a multi‐scale feature decoder (MFD) module. First, multi‐scale encoding features are extracted through the encoder. These features are then fed into the GMFE module to enhance the multi‐scale encoding features under the guidance of saliency map derived from the decoding features of the previous layer. Subsequently, in the GFF module, the enhanced encoding features are fused with the decoding features from the previous layer. Finally, in the MFD module, the bidirectionally guided multi‐scale encoding features is integrated to generate an accurate saliency map. Experiments on two high‐resolution and two low‐resolution datasets demonstrate that our model outperforms on high‐resolution datasets while maintaining competitive performance on low‐resolution datasets, underscoring its effectiveness across varying image qualities. Jiangping Tang, Shuyao Guo, Xiaofei Zhou 0003, Liuxin Bao, Jiyong Zhang 0001 |
IET Image Process. | 3 |
| 2025 | Consistency perception network for 360° omnidirectional salient object detection
Hongfa Wen, Zunjie Zhu, Xiaofei Zhou 0003, Jiyong Zhang 0001, Chenggang Yan 0001 |
Neurocomputing | 3 |
| 2025 | Hierarchical spatiotemporal Feature Interaction Network for video saliency prediction
Yingjie Jin, Xiaofei Zhou 0003, Hao Fang 0010, Xiaobin Xu 0002 |
Image Vis. Comput. | 2 |
| 2025 | Deep hierarchical network for full-reference omnidirectional image quality assessment
Youzhi Zhang 0004, Lifei Wan, Xiaofei Zhou 0003, Deyang Liu |
Multim. Syst. | 4 |
| 2025 | Content adaptive JND profile by leveraging HVS inspired channel modeling and perception oriented energy allocation optimization
Haibing Yin, Xia Wang 0006, Guangtao Zhai, Xiaofei Zhou 0003, Chenggang Yan 0001 |
Signal Process. | 4 |
| 2025 | GLNet: Global-Local Fusion Network for Strip Steel Surface Defects DetectionabstractSurface defect detection in strip steel is a critical task in industrial quality control. However, existing methods struggle with capturing both local details and global context effectively. In this paper, we propose the Global-Local Fusion Network (GLNet) for strip steel surface defect detection, which combines the advantages of VMamba's global feature extraction and CNN's local feature modeling. GLNet employs an encoder-decoder structure, where the encoder consists of two parallel branches: one based on VMamba for capturing global features and the other using ResNet50 for extracting local features. In the decoder, a Global-Local Fusion (GLF) module integrates these features using the Cross Prototype Objective Enhancement (CPOE) and Selective Spatial and Channel Attention (SSCA) modules. The CPOE module facilitates the interaction and fusion between global and local features, while the SSCA module digs the multi-scale information from the global feature through dynamic attention to guide the feature aggregation. Extensive experiments on the ESDIs dataset, demonstrate that GLNet achieves state-of-the-art performance in defect detection, surpassing 13 existing methods in both quantitative and qualitative metrics. Liuxin Bao, Xiaofei Zhou 0003, Xiaobin Xu 0002 |
IEEE Signal Process. Lett. | 3 |
| 2025 | Learning Better Video Query With SAM for Video Instance SegmentationabstractRecently, Transformer-based offline video instance segmentation (VIS) solutions have made significant progress by decomposing the whole task into global segmentation map generation and instance discrimination. We argue that the quality of video queries that represent all instances in a video clip is crucial for offline VIS methods. Existing methods typically interact video queries with dense spatio-temporal features, resulting in significant computational complexity and redundant information. Thus, we propose a novel video instance segmentation framework, LBVQ, dedicated to learning better video queries. Specifically, we first obtain the frame queries for each frame independently without any complex inter-frame spatial-temporal association operations. Secondly, we propose an adaptive query initialization module (AQI), which adaptively integrates frame queries to initialize video queries instead of traditional random initialization strategies. This initialization method preserves rich instance clues and accelerates the optimization of the whole model. Finally, to enhance the quality of video queries, we propose a query propagation module (QPM) that captures relevant instance information in frame queries frame by frame, greatly improving the model’s understanding of long videos. By learning higher quality video queries, LBVQ achieves the state-of-the-art on VIS benchmarks with a ResNet-50 backbone: 52.2 AP, 44.8 AP on YouTube-VIS 2019 & 2021. Moreover, LBVQ achieves 39.7 AP on YouTube-VIS 2022 and 22.2 AP on OVIS, demonstrating superior potential for long videos. To further improve the quality of segmentation masks, a large-scale pretrained SAM is employed to refine the segmentation results. Code is available at https://github.com/fanghaook/LBVQ. Hao Fang 0010, Xiaofei Zhou 0003, Xinxin Zhang 0004 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | IFENet: Interaction, Fusion, and Enhancement Network for V-D-T Salient Object DetectionabstractVisible-depth-thermal (VDT) salient object detection (SOD) aims to highlight the most visually attractive object by utilizing the triple-modal cues. However, existing models don't give sufficient exploration of the multi-modal correlations and differentiation, which leads to unsatisfactory detection performance. In this paper, we propose an interaction, fusion, and enhancement network (IFENet) to conduct the VDT SOD task, which contains three key steps including the multi-modal interaction, the multi-modal fusion, and the spatial enhancement. Specifically, embarking on the Transformer backbone, our IFENet can acquire multi-scale multi-modal features. Firstly, the inter-modal and intra-modal graph-based interaction (IIGI) module is deployed to explore inter-modal channel correlation and intra-modal long-term spatial dependency. Secondly, the gated attention-based fusion (GAF) module is employed to purify and aggregate the triple-modal features, where multi-modal features are filtered along spatial, channel, and modality dimensions, respectively. Lastly, the frequency split-based enhancement (FSE) module separates the fused feature into high-frequency and low-frequency components to enhance spatial information (i.e., boundary details and object location) of the salient object. Extensive experiments are performed on VDT-2048 dataset, and the results show that our saliency model consistently outperforms 13 state-of-the-art models. Our code and results are available at https://github.com/Lx-Bao/IFENet. Liuxin Bao, Xiaofei Zhou 0003, Bolun Zheng, Runmin Cong, Haibing Yin, Jiyong Zhang 0001, Chenggang Yan 0001 |
IEEE Trans. Image Process. | 2 |
| 2025 | AGFNet: Adaptive Gated Fusion Network for RGB-T Semantic SegmentationabstractRGB-T semantic segmentation can effectively pop-out objects from challenging scenarios (e.g., low illumination and low contrast environments) by combining RGB and thermal infrared images. However, the existing cutting-edge RGB-T semantic segmentation methods often present insufficient exploration of multi-modal feature fusion, where they overlook the differences between the two modalities. In this paper, we propose an adaptive gated fusion network (AGFNet) to conduct RGB-T semantic segmentation, where the multi-modal features are combined via the gating mechanisms and the spatial details are enhanced via the introduction of edge information. Specifically, the AGFNet employs a cross-modal adaptive gated-attention fusion (CAGF) module to aggregate the RGB and thermal features, where we give a sufficient exploration of the complementarity between the two-modal features via the gated attention unit (GAU). Particularly, in GAU, the gates can be used to purify the features, and the channel and spatial attention mechanisms are further employed to enhance the two-modal features interactively. Then, we design an edge detection (ED) module to learn the object-related edge cues, which simultaneously incorporates local detail information from low-level features and global location information from high-level features. After that, we deploy the edge guidance (EG) module to emphasize the spatial details of the fused features. Next, we deploy the contextual elevation (CE) module to enrich the contextual information of features by iteratively introducing the sine and cosine functions. Finally, considering that the quality of thermal images is usually lower than that of RGB images, we progressively integrate the multi-level RGB encoder features with multi-level decoder features, thereby focusing more on appearance information. Following this way, we can acquire the final high-quality segmentation result. Extensive experiments are performed on three public datasets including MFNet, PST900 and FMB datasets, and the experimental results show that our method achieves competitive performance when compared with the 22 state-of-the-art methods. Xiaofei Zhou 0003, Liuxin Bao, Haibing Yin, Qiuping Jiang, Jiyong Zhang 0001 |
IEEE Trans. Intell. Transp. Syst. | 1 |
| 2024 | Quad Bayer Joint Demosaicing and Denoising Based on Dual Encoder Network with Joint Residual LearningabstractThe recent imaging technology Quad Bayer CFA brings better imaging PSNR and higher visual quality compared to traditional Bayer CFA, but also serious challenges for demosaicing and denoising during the ISP pipeline. In this paper, we propose a novel dual encoder network, namely DRNet, to achieve joint demosaicing and denoising for Quad Bayer CFA. The dual encoders are carefully designed in that one is mainly constructed by a joint residual block to jointly estimate the residuals for demosaicing and denoising separately. In contrast, the other one is started with a pixel modulation block which is specially designed to match the characteristics of Quad Bayer pattern for better feature extraction. We demonstrate the effectiveness of each proposed component through detailed ablation investigations. The comparison results on public benchmarks illustrate that our DRNet achieves an apparent performance gain~(0.38dB to the 2nd best) from the state-of-the-art method and balances performance and efficiency well. The experiments on real-world images show that the proposed method could enhance the reconstruction quality from the native ISP algorithm. Bolun Zheng, Haoran Li 0025, Tingyu Wang 0002, Xiaofei Zhou 0003, Chenggang Yan 0001 |
AAAI | 5 |
| 2024 | ADNet: Anti-noise dual-branch network for road defect detection
Bin Wan, Xiaofei Zhou 0003, Yaoqi Sun, Tingyu Wang 0002, Chengtao Lv, Shuai Wang 0003, Haibing Yin, Chenggang Yan 0001 |
Eng. Appl. Artif. Intell. | 2 |
| 2024 | GINet:Graph interactive network with semantic-guided spatial refinement for salient object detection in optical remote sensing images
Chenwei Zhu, Xiaofei Zhou 0003, Liuxin Bao, Hongkui Wang, Shuai Wang 0003, Zunjie Zhu, Chenggang Yan 0001, Jiyong Zhang 0001 |
J. Vis. Commun. Image Represent. | 2 |
| 2024 | TMNet: Triple-modal interaction encoder and multi-scale fusion decoder network for V-D-T salient object detection
Bin Wan, Chengtao Lv, Xiaofei Zhou 0003, Yaoqi Sun, Zunjie Zhu, Hongkui Wang, Chenggang Yan 0001 |
Pattern Recognit. | 3 |
| 2024 | Interactive Fusion and Correlation Network for Three-Modal Images Few-Shot Semantic SegmentationabstractThis letter presents a novel method for three-modal images few-shot semantic segmentation. Some previous efforts fuse multiple modalities before feature correlation, while this changes the original visual information that is useful to subsequent feature matching. Others are built based on early correlation learning, which can cause details loss and thereby defects multi-modal integration. To address these challenges, we build a novel interactive fusion and correlation network (IFCNet). Specifically, the proposed fusing and correlating (FC) module performs feature correlating and attention-based multi-modal fusing interactively, which establishes effective inter-modal complementarity and benefits intra-modal query-support correlation. Furthermore, we add a multi-modal correlation (MC) module, which leverages multi-layer cosine similarity maps to enrich multi-modal visual correspondence. Experiments on the VDT-2048-5$^{i}$dataset demonstrate the network's superior performance, which outperforms existing state-of-the-art methods in both 1-shot and 5-shot settings. The study also includes an ablation analysis to validate the contributions of the FC module and the MC module to the overall segmentation accuracy. Haolan He, Xianguo Dong, Xiaofei Zhou 0003, Bo Wang 0031, Jiyong Zhang 0001 |
IEEE Signal Process. Lett. | 3 |
| 2024 | Spatial Attention-Guided Light Field Salient Object Detection Network With Implicit Neural RepresentationabstractRecently, many Light Field Salient Object Detection (LF SOD) methods have been proposed. However, guaranteeing the integrality and recovering more high-frequency details of the generated salient object map still remain challenging. To this end, we propose a spatial attention-guided LF SOD network with implicit neural representation to further improve LF SOD performance. We adopt an encoder-decoder structure for model construction. In order to ensure the completeness of the generated salient object map, a multi-modal and multi-scale feature fusion module is designed in the encoder part to refine the salient regions within all-in-focus image and aggregate the focal stack and all-in-focus image in spatial attention-guided manner. In order to recover more high-frequency details of the obtained salient object map, an implicit detail restoration module is proposed in the decoder part. In virtue of implicit neural representation, we convert the detail restoration problem into a functional mapping problem. By further integrating the self-attention mechanism, the derived saliency map can be depicted at a more refined level. Comprehensive experimental results demonstrate the superiority of the proposed method. Ablation studies and visual comparisons further validate that the proposed method can guarantee the integrality and recover more high-frequency detail information of the obtained saliency map. The code is publicly available athttps://github.com/ldyorchid/LFSOD-Net. Xin Zheng 0006, Zhengqu Li, Deyang Liu, Xiaofei Zhou 0003, Caifeng Shan |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | MINet: Multiscale Interactive Network for Real-Time Salient Object Detection of Strip Steel Surface DefectsabstractThe automated surface defect detection is a fundamental task in industrial production, and the existing saliency-based works overcome the challenging scenes and give promising detection results. However, the cutting-edge efforts often suffer from large parameter size, heavy computational cost, and slow inference speed, which heavily limits the practical applications. To this end, we devise a multiscale interactive (MI) module, which employs depthwise convolution (DWConv) and pointwise convolution (PWConv) to independently extract and interactively fuse features of different scales, respectively. Particularly, the MI module can provide satisfactory characterization for defect regions with fewer parameters. Embarking on this module, we propose a lightweight multiscale interactive network (MINet) to conduct real-time salient object detection of strip steel surface defects. Comprehensive experimental results on SD-Saliency-900 dataset, which contains three kinds of strip steel surface defect detection images (i.e., inclusion, patches, and scratches), demonstrate that the proposed MINet presents comparable detection accuracy with the state-of-the-art methods while running at a GPU speed of 721 FPS and a CPU speed of 6.3 FPS for 368×368 images with only 0.28 M parameters. Kunye Shen, Xiaofei Zhou 0003, Zhi Liu 0003 |
IEEE Trans. Ind. Informatics | 2 |
| 2024 | Quality-Aware Selective Fusion Network for V-D-T Salient Object DetectionabstractDepth images and thermal images contain the spatial geometry information and surface temperature information, which can act as complementary information for the RGB modality. However, the quality of the depth and thermal images is often unreliable in some challenging scenarios, which will result in the performance degradation of the two-modal based salient object detection (SOD). Meanwhile, some researchers pay attention to the triple-modal SOD task, namely the visible-depth-thermal (VDT) SOD, where they attempt to explore the complementarity of the RGB image, the depth image, and the thermal image. However, existing triple-modal SOD methods fail to perceive the quality of depth maps and thermal images, which leads to performance degradation when dealing with scenes with low-quality depth and thermal images. Therefore, in this paper, we propose a quality-aware selective fusion network (QSF-Net) to conduct VDT salient object detection, which contains three subnets including the initial feature extraction subnet, the quality-aware region selection subnet, and the region-guided selective fusion subnet. Firstly, except for extracting features, the initial feature extraction subnet can generate a preliminary prediction map from each modality via a shrinkage pyramid architecture, which is equipped with the multi-scale fusion (MSF) module. Then, we design the weakly-supervised quality-aware region selection subnet to generate the quality-aware maps. Concretely, we first find the high-quality and low-quality regions by using the preliminary predictions, which further constitute the pseudo label that can be used to train this subnet. Finally, the region-guided selective fusion subnet purifies the initial features under the guidance of the quality-aware maps, and then fuses the triple-modal features and refines the edge details of prediction maps through the intra-modality and inter-modality attention (IIA) module and the edge refinement (ER) module, respectively. Extensive experiments are performed on VDT-2048 dataset, and the results show that our saliency model consistently outperforms 13 state-of-the-art methods with a large margin. Our code and results are available at https://github.com/Lx-Bao/QSFNet. Liuxin Bao, Xiaofei Zhou 0003, Xiankai Lu, Yaoqi Sun, Haibing Yin, Jiyong Zhang 0001, Chenggang Yan 0001 |
IEEE Trans. Image Process. | 2 |
| 2024 | MFFNet: Multi-Modal Feature Fusion Network for V-D-T Salient Object DetectionabstractThis article discusses the limitations of single- and two-modal salient object detection (SOD) methods and the emergence of multi-modal SOD techniques that integrate Visible, Depth, or Thermal information. However, current multi-modal methods often rely on simple fusion techniques such as addition, multiplication and concatenation, to combine the different modalities, which is ineffective for challenging scenes, such as low illumination and background messy. To address this issue, we propose a novel multi-modal feature fusion network (MFFNet) for V-D-T salient object detection, where the two key points are the triple-modal deep fusion encoder and the progressive feature enhancement decoder. The MFFNet's triple-modal deep fusion (TDF) module is designed to integrate the features of the three modalities and explore their complementarity by utilizing mutual optimization during the encoding phase. In addition, the progressive feature enhancement decoder consists of the weighted context-enhanced feature (WCF) module, region optimization (RO) module and boundary perception (BP) module to produce region-aware and contour-aware features. After that, a multi-scale fusion (MF) module is proposed to integrate these features and generate high-quality saliency maps. We conduct extensive experiments on the VDT-2048 dataset, and our results show that the proposed MFFNet outperforms 12 state-of-the-art multi-modal methods. Bin Wan, Xiaofei Zhou 0003, Yaoqi Sun, Tingyu Wang 0002, Chengtao Lv, Shuai Wang 0003, Haibing Yin, Chenggang Yan 0001 |
IEEE Trans. Multim. | 2 |
| 2024 | ADMNet: Attention-Guided Densely Multi-Scale Network for Lightweight Salient Object DetectionabstractRecently, benefitting from the rapid development of deep learning technology, the research of salient object detection has achieved great progress. However, the performance of existing cutting-edge saliency models relies on large network size and high computational overhead. This is unamiable to real-world applications, especially the practical platforms with low cost and limited computing resources. In this paper, we propose a novel lightweight saliency model, namely Attention-guided Densely Multi-scale Network (ADMNet), to tackle this issue. Firstly, we design the multi-scale perception (MP) module to acquire different contextual features by using different receptive fields. Embarking on MP module, we build the encoder of our model, where each convolutional block adopts a dense structure to connect MP modules. Following this way, our model can provide powerful encoder features for the characterization of salient objects. Secondly, we employ dual attention (DA) module to equip the decoder blocks. Particularly, in DA module, the binarized coarse saliency inference of the decoder block (i.e., a hard spatial attention map) is first employed to filter out interference cues from the decoder feature, and then by introducing large receptive fields, the enhanced decoder feature is used to generate a soft spatial attention map, which further purifies the fused features. Following this way, the deep features are steered to give more concerns to salient regions. Extensive experiments on five public challenging datasets including ECSSD, DUT-OMRON, DUTS-TE, HKU-IS, and PASCAL-S clearly show that our model achieves comparable performance with the state-of-the-art saliency models while running at a 219.4fps GPU speed and a 1.76fps CPU speed for a 368×368 image with only 0.84 M parameters. Xiaofei Zhou 0003, Kunye Shen, Zhi Liu 0003 |
IEEE Trans. Multim. | 1 |
| 2024 | Decoupling and Integration Network for Camouflaged Object DetectionabstractRecently, camouflaged object detection (COD), which suffers from numerous challenges such as low contrast between camouflaged objects and background and large variations of camouflaged object appearances, has received more and more concerns. However, the performance of existing camouflaged object detection methods is still unsatisfactory, especially when dealing with complex scenes. Therefore, in this paper, we propose a novel Decoupling and Integration Network (DINet) to detect camouflaged objects. Here, the depiction of camouflaged objects can be regarded as the iterative decoupling and integration of the body features and detail features, where the former focuses on the center of camouflaged objects and the latter contains pixels around edges. Concretely, firstly, we deploy two complementary decoder branches including a detail branch and a body branch to learn the decoupling features, namely body decoder features and detail decoder features. Particularly, each decoder block of the two branches incorporates features from three components,i.e., the previous interactive feature fusion (IFF) module, adjacent encoder layers, and corresponding encoder layer. Besides, to further elevate the body decoder features, the body blocks also introduce the global contextual information, which is the combination of all body encoder features via the global context (GC) unit, to provide coarse object location information. Secondly, to integrate the two decoupling decoder features, we deploy the interactive feature fusion (IFF) module based on the interactive combination and channel attention. Following this way, we can progressively provide a complete and accurate representation for camouflaged objects. Extensive experiments on three public challenging datasets, including CAMO, COD10K, and NC4K, show that our DINet presents competitive performance when compared with the state-of-the-art models. Xiaofei Zhou 0003, Zhicong Wu, Runmin Cong |
IEEE Trans. Multim. | 1 |
| 2023 | 360$^{\circ }$ Omnidirectional Salient Object Detection with Multi-scale Interaction and Densely-Connected Prediction
Haowei Dai, Liuxin Bao, Kunye Shen, Xiaofei Zhou 0003, Jiyong Zhang 0001 |
ICIG (1) | 4 |
| 2023 | Learning a Multilevel Cooperative View Reconstruction Network for Light Field Angular Super-ResolutionabstractRecently, many methods have been proposed to improve the angular resolution of sparsely-sampled Light Field (LF). However, the synthesized dense LF inevitably exhibits blurry edges and artifacts. This paper intents to model the global relations of LF views and quality degradation model by learning a multilevel cooperative view reconstruction network to further enhance LF angular Super-Resolution (SR) performance. The proposed LF angular SR network consists of three sub-networks including the Cooperative Angular Transformer Network (CATNet), the Deblurring Network (DBNet), and the Texture Repair Network (TRNet). The CATNet simultaneously captures global features of all LF views and local features within each view, which benefits in characterizing the inherent LF structure. The DBNet models a quality degradation model by estimating blur kernels to reduce the blurry edges and artifacts. The TRNet focuses on restoring fine-scale texture details. Experimental results over various LF datasets including large baseline LF images demonstrate the significant superiority of our method when compared with state-of-the-art ones. Deyang Liu, Xiaofei Zhou 0003, Ping An 0001, Yuming Fang 0001 |
ICME | 3 |
| 2023 | Frequency Perception Network for Camouflaged Object DetectionabstractCamouflaged object detection (COD) aims to accurately detect objects hidden in the surrounding environment. However,the existing COD methods mainly locate camouflaged objects in the RGB domain, their performance has not been fully exploited in many challenging scenarios. Considering that the features of the camouflaged object and the background are more discriminative in the frequency domain, we propose a novel learnable and separable frequency perception mechanism driven by the semantic hierarchy in the frequency domain. Our entire network adopts a two-stage model, including a frequency-guided coarse localization stage and a detail-preserving fine localization stage.With the multi-level features extracted by the backbone, we design a flexible frequency perception module based on octave convolution for coarse positioning. Then, we design the correction fusion module to step-by-step integrate the high-level features through the prior-guided correction and cross-layer feature channel association, and finally combine them with the shallow features to achieve the detailed correction of the camouflaged objects. Compared with the currently existing models, our proposed method achieves competitive performance in three popular benchmark datasets both qualitatively and quantitatively. The code will be released at https://github.com/rmcong/FPNet_ACMMM23. Runmin Cong, Mengyao Sun 0003, Sanyi Zhang, Xiaofei Zhou 0003, Wei Zhang 0021, Yao Zhao 0001 |
ACM Multimedia | 4 |
| 2023 | GFNet: gated fusion network for video saliency prediction
Songhe Wu, Xiaofei Zhou 0003, Yaoqi Sun, Zunjie Zhu, Jiyong Zhang 0001, Chenggang Yan 0001 |
Appl. Intell. | 2 |
| 2023 | SMINet: Semantics-aware multi-level feature interaction network for surface defect detection
Bin Wan, Xiaofei Zhou 0003, Yaoqi Sun, Zunjie Zhu, Haibing Yin, Ji Hu 0002, Jiyong Zhang 0001, Chenggang Yan 0001 |
Eng. Appl. Artif. Intell. | 2 |
| 2023 | Aggregating transformers and CNNs for salient object detection in optical remote sensing images
Liuxin Bao, Xiaofei Zhou 0003, Bolun Zheng, Haibing Yin, Zunjie Zhu, Jiyong Zhang 0001, Chenggang Yan 0001 |
Neurocomputing | 2 |
| 2023 | STI-Net: Spatiotemporal integration network for video saliency detection
Xiaofei Zhou 0003, Weipeng Cao, Hanxiao Gao, Zhong Ming 0001, Jiyong Zhang 0001 |
Inf. Sci. | 1 |
| 2023 | SRI-Net: Similarity retrieval-based inference network for light field salient object detection
Chengtao Lv, Xiaofei Zhou 0003, Deyang Liu, Bolun Zheng, Jiyong Zhang 0001, Chenggang Yan 0001 |
J. Vis. Commun. Image Represent. | 2 |
| 2023 | CANet: Context-aware Aggregation Network for Salient Object Detection of Surface Defects
Bin Wan, Xiaofei Zhou 0003, Mang Xiao, Yaoqi Sun, Bolun Zheng, Jiyong Zhang 0001, Chenggang Yan 0001 |
J. Vis. Commun. Image Represent. | 2 |
| 2023 | Depth-guided deep filtering network for efficient single image bokeh rendering
Bolun Zheng, Xiaofei Zhou 0003, Aiai Huang, Yaoqi Sun, Chuqiao Chen, Chenggang Yan 0001, Shanxin Yuan |
Neural Comput. Appl. | 3 |
| 2023 | Transformer-Based Multi-Scale Feature Integration Network for Video Saliency PredictionabstractMost cutting-edge video saliency prediction models rely on spatiotemporal features extracted by 3D convolutions due to its local contextual cues acquirement ability. However, the shortage of 3D convolutions is that it cannot effectively capture long-term spatiotemporal dependencies in videos. To address this limitation, we propose a novel Transformer-based Multi-scale Feature Integration Network (TMFI-Net) for video saliency prediction, where the proposed TMFI-Net consists of a semantic-guided encoder and a hierarchical decoder. Firstly, embarking on the Transformer-based multi-level spatiotemporal features, the semantic-guided encoder enhances the features by inserting the high-level feature into each level feature via a top-down pathway and a longitudinal connection, which endows the multi-level spatiotemporal features with rich contextual information. In this way, the features are steered to give more concerns to saliency regions. Secondly, the hierarchical decoder employs a multi-dimensional attention (MA) module to elevate features along channel, temporal, and spatial dimensions jointly. Successively, the hierarchical decoder deploys a progressive decoding block to conduct an initial saliency prediction, which provides a coarse localization of saliency regions. Lastly, considering the complementarity of different saliency predictions, we integrate all initial saliency prediction results into the final saliency map. Comprehensive experimental results on four video saliency datasets firmly demonstrate that our model achieves superior performance when compared with the state-of-the-art video saliency models. The code is available athttps://github.com/wusonghe/TMFI-Net. Xiaofei Zhou 0003, Songhe Wu, Bolun Zheng, Shuai Wang 0003, Haibing Yin, Jiyong Zhang 0001, Chenggang Yan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Edge-Guided Recurrent Positioning Network for Salient Object Detection in Optical Remote Sensing ImagesabstractOptical remote sensing images (RSIs) have been widely used in many applications, and one of the interesting issues about optical RSIs is the salient object detection (SOD). However, due to diverse object types, various object scales, numerous object orientations, and cluttered backgrounds in optical RSIs, the performance of the existing SOD models often degrade largely. Meanwhile, cutting-edge SOD models targeting optical RSIs typically focus on suppressing cluttered backgrounds, while they neglect the importance of edge information which is crucial for obtaining precise saliency maps. To address this dilemma, this article proposes an edge-guided recurrent positioning network (ERPNet) to pop-out salient objects in optical RSIs, where the key point lies in the edge-aware position attention unit (EPAU). First, the encoder is used to give salient objects a good representation, that is, multilevel deep features, which are then delivered into two parallel decoders, including: 1) an edge extraction part and 2) a feature fusion part. The edge extraction module and the encoder form a U-shape architecture, which not only provides accurate salient edge clues but also ensures the integrality of edge information by extra deploying the intraconnection. That is to say, edge features can be generated and reinforced by incorporating object features from the encoder. Meanwhile, each decoding step of the feature fusion module provides the position attention about salient objects, where position cues are sharpened by the effective edge information and are used to recurrently calibrate the misaligned decoding process. After that, we can obtain the final saliency map by fusing all position attention cues. Extensive experiments are conducted on two public optical RSIs datasets, and the results show that the proposed ERPNet can accurately and completely pop-out salient objects, which consistently outperforms the state-of-the-art SOD models. Xiaofei Zhou 0003, Kunye Shen, Li Weng, Runmin Cong, Bolun Zheng, Jiyong Zhang 0001, Chenggang Yan 0001 |
IEEE Trans. Cybern. | 1 |
| 2022 | DomainPlus: Cross Transform Domain Learning towards High Dynamic Range ImagingabstractHigh dynamic range (HDR) imaging by combining multiple low dynamic range (LDR) images of different exposures provides a promising way to produce high quality photographs. However, the misalignment between the input images leads to ghosting artifacts in the reconstructed HDR image. In this paper, we propose a cross-transform domain neural network for efficient HDR imaging. Our approach consists of two modules: a merging module and a restoration module. For the merging module, we propose a Multiscale Attention with Fronted Fusion (MAFF) mechanism to achieve coarse-to-fine spatial fusion. For the restoration module, we propose fronted Discrete Wavelet Transform (DWT) and Discrete Cosine Transform (DCT)-based learnable bandpass filters to formulate a cross-transform domain learning block, dubbed DomainPlus Block (DPB) for effective ghosting removal. Our ablation study and comprehensive experiments show that DomainPlus outperforms the existing state-of-the-art on several datasets. Bolun Zheng, Xiaokai Pan, Xiaofei Zhou 0003, Gregory Slabaugh, Chenggang Yan 0001, Shanxin Yuan |
ACM Multimedia | 4 |
| 2022 | Fully Squeezed Multiscale Inference Network for Fast and Accurate Saliency Detection in Optical Remote-Sensing ImagesabstractRecently, salient object detection in optical remote-sensing images (RSIs) has received more and more attention. To tackle the challenges of RSIs including large-scale variation of objects, cluttered background, irregular shape of objects, and big difference in illumination, the cutting-edge convolutional neural network (CNN)-based models are proposed and have achieved an encouraging performance. However, the performance of the top-level models usually depends on the large model size and high computational cost, which limits their practical applications. To remedy the issue, we introduce a fully squeezed multiscale (FSM) module to equip the entire network. Specifically, the FSM module squeezes the feature maps from high dimension to low dimension and introduces the multiscale strategy to endow the capability of feature characterization with different receptive fields and different contexts. Based on the FSM module, we build the FSM inference network (FSMI-Net) to pop-out salient objects from optical RSIs, which is with fewer parameters and fast inference speed. Particularly, the proposed FSMI-Net only contains 3.6M parameters, and its GPU running speed is about 28 fps for$384 \times 384$inputs, which is superior to the existing saliency models targeting optical RSIs. Extensive comparisons are performed on two public optical RSIs datasets, and our FSMI-Net achieves comparable detection accuracy when compared with the state-of-the-art models, where our model realizes a balance between the computational cost and detection performance. Kunye Shen, Xiaofei Zhou 0003, Bin Wan, Jiyong Zhang 0001 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2022 | FANet: Feature aggregation network for RGBD saliency detection
Xiaofei Zhou 0003, Hongfa Wen, Haibing Yin, Jiyong Zhang 0001, Chenggang Yan 0001 |
Signal Process. Image Commun. | 1 |
| 2022 | Edge-Aware Multiscale Feature Integration Network for Salient Object Detection in Optical Remote Sensing ImagesabstractThe optical remote sensing images (RSIs) show various spatial resolutions and cluttered background, where salient objects with different scales, types, and orientations are presented in diverse RSI scenes. Therefore, it is inappropriate to directly extend cutting-edge saliency detection methods for conventional RGB images to optical RSIs. Besides, the existing saliency models targeting RSIs often render imperfect saliency maps, where some of them are with coarse boundary details. To solve this problem, this article attempts to introduce the edge information to precisely detect salient objects in RSIs. Accordingly, we propose an edge-aware multiscale feature integration network (EMFI-Net) for salient object detection by conducting multiscale feature integration under the explicit and implicit assistance of salient edge cues. Specifically, our network contains two parts including the encoder and decoder. First, the encoder extracts multiscale deep features from three RSIs with different resolutions, where the high-level deep semantic features from three RSIs are integrated using a cascaded feature fusion module. Second, the encoder explicitly enriches the multiscale deep features by integrating the salient edge cues extracted by a salient edge extraction module. Meanwhile, we also implicitly deploy an edge-aware constraint to the supervision of the saliency map prediction by introducing a hybrid loss function. Finally, the decoder integrates the enriched multiscale deep features in a coarse-to-fine way, yielding a high-quality saliency map. The experiments conducted on two public optical RSI datasets clearly prove the effectiveness and superiority of the proposed EMFI-Net against the state-of-the-art saliency models. Xiaofei Zhou 0003, Kunye Shen, Zhi Liu 0003, Chen Gong 0002, Jiyong Zhang 0001, Chenggang Yan 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2021 | Personalized image observation behavior learning in fixation based personalized salient object segmentation
Gongyang Li, Weijie Wei 0001, Xiaofei Zhou 0003, Zhi Liu 0003 |
Neurocomputing | 4 |
| 2021 | Leveraging graph neural networks for point-of-interest recommendations
Jiyong Zhang 0001, Xin Liu 0027, Xiaofei Zhou 0003, Xiaowen Chu 0001 |
Neurocomputing | 3 |
| 2021 | ATCC: Accurate tracking by criss-cross location attention
Yong Wu 0007, Zhi Liu 0003, Xiaofei Zhou 0003, Linwei Ye, Yang Wang 0003 |
Image Vis. Comput. | 3 |
| 2021 | Dynamic Selective Network for RGB-D Salient Object DetectionabstractRGB-D saliency detection is receiving more and more attention in recent years. There are many efforts have been devoted to this area, where most of them try to integrate the multi-modal information, i.e. RGB images and depth maps, via various fusion strategies. However, some of them ignore the inherent difference between the two modalities, which leads to the performance degradation when handling some challenging scenes. Therefore, in this paper, we propose a novel RGB-D saliency model, namely Dynamic Selective Network (DSNet), to perform salient object detection (SOD) in RGB-D images by taking full advantage of the complementarity between the two modalities. Specifically, we first deploy a cross-modal global context module (CGCM) to acquire the high-level semantic information, which can be used to roughly locate salient objects. Then, we design a dynamic selective module (DSM) to dynamically mine the cross-modal complementary information between RGB images and depth maps, and to further optimize the multi-level and multi-scale information by executing the gated and pooling based selection, respectively. Moreover, we conduct the boundary refinement to obtain high-quality saliency maps with clear boundary details. Extensive experiments on eight public RGB-D datasets show that the proposed DSNet achieves a competitive and excellent performance against the current 17 state-of-the-art RGB-D SOD models. Hongfa Wen, Chenggang Yan 0001, Xiaofei Zhou 0003, Runmin Cong, Yaoqi Sun, Bolun Zheng, Jiyong Zhang 0001, Yongjun Bao, Guiguang Ding |
IEEE Trans. Image Process. | 3 |
| 2020 | Co-Saliency Detection Using Collaborative Feature Extraction And High-To-Low Feature IntegrationabstractCo-saliency detection, as a developing research branch of saliency detection, devotes to identify the common salient objects in a group of related images. The major challenge of co-saliency detection is how to effectively represent features considering both intra-image and inter-image information. In this paper, we propose a co-saliency detection model using collaborative feature extraction and high-to-low feature integration. We first feed the target image and its co-images into the Individual Feature Extraction Module (IFEM) to produce multi-level individual features. Then, to capture the collaborative inter-image information, the Collaborative Feature Extraction Module (CFEM) is applied to all highest-level individual features, generating the collaborative feature. Finally, we build a High-to-low Feature Integration Module (HFIM), which integrates the collaborative feature and multi-level individual features of the target image, to enrich the collaborative feature with individual intra-image information. Extensive experiments on two public datasets demonstrate that the proposed model achieves the state-of-the-art performance. Jingru Ren, Zhi Liu 0003, Gongyang Li, Xiaofei Zhou 0003, Cong Bai, Guangling Sun |
ICME | 4 |
| 2020 | Co-saliency detection via integration of multi-layer convolutional features and inter-image propagation
Jingru Ren, Zhi Liu 0003, Xiaofei Zhou 0003, Cong Bai, Guangling Sun |
Neurocomputing | 3 |
| 2020 | Cross-modal feature extraction and integration based RGBD saliency detection
Liang Pan, Xiaofei Zhou 0003, Jiyong Zhang 0001, Chenggang Yan 0001 |
Image Vis. Comput. | 2 |
| 2020 | Attention-guided RGBD saliency detection using appearance information
Xiaofei Zhou 0003, Gongyang Li, Chen Gong 0002, Zhi Liu 0003, Jiyong Zhang 0001 |
Image Vis. Comput. | 1 |
| 2020 | Depth-guided saliency detection via boundary information
Xiaofei Zhou 0003, Hongfa Wen, Haibing Yin, Chenggang Yan 0001 |
Image Vis. Comput. | 1 |
| 2020 | Saliency detection using adversarial learning networks
Yong Wu 0007, Zhi Liu 0003, Xiaofei Zhou 0003 |
J. Vis. Commun. Image Represent. | 3 |
| 2020 | FANet: Features Adaptation Network for 360$^{\circ }$ Omnidirectional Salient Object DetectionabstractSalient object detection (SOD) in 360° omnidirectional images has become an eye-catching problem because of the popularity of affordable 360° cameras. In this paper, we propose a Features Adaptation Network (FANet) to highlight salient objects in 360° omnidirectional images reliably. To utilize the feature extraction capability of convolutional neural networks and capture global object information, we input the equirectangular 360° images and corresponding cube-map 360° images to the feature extraction network (FENet) simultaneously to obtain multi-level equirectangular and cube-map features. Furthermore, we fuse these two kinds of features at each level of the FENetby a projection features adaptation (PFA) module, for selecting these two kinds of features adaptively. Finally, we combine the preliminary adaptation features at different levels by a multi-level features adaptation (MLFA) module, which weights these different-level features adaptively and produces the final saliency maps. Experiments show our FANet outperforms the state-of-the-art methods on the 360° omnidirectional SOD datasets. Mengke Huang, Zhi Liu 0003, Gongyang Li, Xiaofei Zhou 0003, Olivier Le Meur |
IEEE Signal Process. Lett. | 4 |
| 2019 | Learn Image Object Co-segmentation with Multi-scale Feature FusionabstractImage object co-segmentation aims to segment common objects in a group of images. This paper proposes a novel neural network, which extracts multi-scale convolutional features at multiple layers via a modified VGG network and fuses them both within and across images as the intra-image and the inter-image features. Then these two kinds of features are further fused at each scale as the multi-scale co-features of common objects, and finally the multi-scale co-features are summed up and upsampled to obtain the co-segmentation results. To simplify the network and reduce the rapidly rising resource cost along with the inputs, the reduced input size, less downsampling and dilation convolution are adopted in the proposed model. Experimental results on the public dataset demonstrate that the proposed model achieves a comparable performance to the state-of-the-art co-segmentation methods while the computation cost has been effectively reduced. Zhi Liu 0003, Jian Zhang 0002, Xiaofei Zhou 0003 |
VCIP | 4 |
| 2019 | Saliency detection via multi-level integration and multi-scale fusion neural networks
Mengke Huang, Zhi Liu 0003, Linwei Ye, Xiaofei Zhou 0003, Yang Wang 0003 |
Neurocomputing | 4 |
| 2019 | Deep fusion based video saliency detection
Hongfa Wen, Xiaofei Zhou 0003, Yaoqi Sun, Jiyong Zhang 0001, Chenggang Yan 0001 |
J. Vis. Commun. Image Represent. | 2 |
| 2019 | Effective online refinement for video object segmentation
Gongyang Li, Zhi Liu 0003, Xiaofei Zhou 0003 |
Multim. Tools Appl. | 3 |
| 2019 | Video co-segmentation based on directed graph
Zhi Liu 0003, Xiaofei Zhou 0003, Wei Liu 0044, Xuemei Zou |
Multim. Tools Appl. | 3 |
| 2018 | Video Saliency Detection Using Deep Convolutional Neural Networks
Xiaofei Zhou 0003, Zhi Liu 0003, Chen Gong 0002, Gongyang Li, Mengke Huang |
PRCV (2) | 1 |
| 2018 | Saliency integration driven by similar images
Jingru Ren, Zhi Liu 0003, Xiaofei Zhou 0003, Guangling Sun, Cong Bai |
J. Vis. Commun. Image Represent. | 3 |
| 2018 | Video saliency detection via bagging-based prediction and spatiotemporal propagation
Xiaofei Zhou 0003, Zhi Liu 0003, Kai Li 0016, Guangling Sun |
J. Vis. Commun. Image Represent. | 1 |
| 2018 | Spatiotemporal salient object detection by integrating with objectness
Tongbao Wu, Zhi Liu 0003, Xiaofei Zhou 0003, Kai Li 0016 |
Multim. Tools Appl. | 3 |
| 2018 | Improving Video Saliency Detection via Localized Estimation and Spatiotemporal RefinementabstractVideo saliency detection aims to pop out the most salient regions in every frame of a video. Up to now, many efforts have been made from various aspects for video saliency detection. Unfortunately, the existing video saliency models are very likely to fail in challenging videos with complicated motions and complex scenes. Therefore, in this paper, we propose a novel framework to improve the saliency detection results generated by existing video saliency models. The proposed framework consists of three key steps including localized estimation, spatiotemporal refinement, and saliency update. Specifically, the initial saliency map of each frame in a video is first generated by using an existing saliency model. Then, by considering the temporal consistency and strong correlation among adjacent frames, the localized estimation models, which are generated by training the random forest regressor within a local temporal window, are employed to generate the temporary saliency map. Finally, by taking the appearance and motion information of salient objects into consideration, the spatiotemporal refinement step is deployed to further improve the temporary saliency map and generate the final saliency map. Furthermore, such an improved saliency map is then utilized to update the initial saliency map and provide reliable cues for saliency detection in the next frame. The experimental results on four challenging datasets demonstrate that the proposed framework is able to consistently and significantly improve the saliency detection performance of various video saliency models, thereby achieving the state-of-the-art performance. Xiaofei Zhou 0003, Zhi Liu 0003, Chen Gong 0002, Wei Liu 0044 |
IEEE Trans. Multim. | 1 |
| 2017 | Adaptive saliency fusion based on quality assessment
Xiaofei Zhou 0003, Zhi Liu 0003, Guangling Sun, Xiangyang Wang 0003 |
Multim. Tools Appl. | 1 |
| 2016 | Facial descriptor for Kinect depth using inner-inter-normal components local binary patterns and tensor histograms
Guangling Sun, Yong Dong, Xiaofei Zhou 0003, Zhi Liu 0003 |
Mach. Vis. Appl. | 3 |
| 2016 | Saliency Detection Via Similar Image RetrievalabstractThis letter proposes a novel saliency detection framework by propagating saliency of similar images retrieved from large and diverse Internet image collections to boost saliency detection performance effectively. For the input image, a group of similar images is retrieved based on the saliency weighted color histograms and the Gist descriptor from Internet image collections. Then, a pixel-level correspondence process between images is performed to guide the saliency propagation from the retrieved images. Both initial saliency map and correspondence saliency map are exploited to select the training samples by using the graph cut-based segmentation. Finally, the training samples are input into a set of weak classifiers to learn the boosted classifier for generating the boosted saliency map, which is integrated with the initial saliency map to generate the final saliency map. Experimental results on two public image datasets demonstrate that the proposed model can achieve the better saliency detection performance than the state-of-the-art single-image saliency models and co-saliency models. Linwei Ye, Zhi Liu 0003, Xiaofei Zhou 0003, Liquan Shen, Jian Zhang 0002 |
IEEE Signal Process. Lett. | 3 |
| 2016 | Improving Saliency Detection Via Multiple Kernel Boosting and Adaptive FusionabstractThis letter proposes a novel framework to improve the saliency detection performance of an existing saliency model, which is used to generate the initial saliency map. First, a novel regional descriptor consisting of regional self-information, regional variance, and regional contrast on a number of features with local, global, and border context is proposed to describe the segmented regions at multiple scales. Then, regarding saliency computation as a regression problem, a multiple kernel boosting method based on support vector regression (MKB-SVR) is proposed to generate the complementary saliency map. Finally, an adaptive fusion method via learning a quality prediction model for saliency maps is proposed to effectively fuse the initial saliency map with the complementary saliency map and obtain the final saliency map with improvement on saliency detection performance. Experimental results on two public datasets with the state-of-the-art saliency models validate that the proposed method consistently improves the saliency detection performance of various saliency models. Xiaofei Zhou 0003, Zhi Liu 0003, Guangling Sun, Linwei Ye, Xiangyang Wang 0003 |
IEEE Signal Process. Lett. | 1 |