Gongyang Li

dblp:229/1266 · DBLP profile ↗
← Back
44ranked-venue papers
13as first author
34since 2021 · last 2026
0000-0001-7324-1196ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 28 · 7 first-author · 20 since 2021Artificial intelligence and machine learning · 14 · 4 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Interactive edge awareness network for salient object detection in optical remote sensing images
Xian Fang, Qiaohong Chen, Gongyang Li
Expert Syst. Appl.4
2026 IPDiff: Diffusion-driven ORSI Salient Object Detection with Information Reconstruction and Multi-Prior Guidance
Gongyang Li, Zhen Bai 0001, Runmin Cong, Dan Zeng 0001, Weisi Lin
Int. J. Comput. Vis.1
2026 Hierarchical prior-guided and channel-wise adaptation fusion network for RGB-D railway surface defect inspection
Jianlin Chen, Gongyang Li, Zhijiang Zhang, Dan Zeng 0001
J. Vis. Commun. Image Represent.2
2026 Diffusion-Driven RGB-D Salient Object Detection With Temporal Modulation
abstract
Existing RGB-D Salient Object Detection (SOD) methods are primarily built on the end-to-end prediction paradigm. Although these methods have achieved remarkable progress, they still struggle to generate accurate predictions in some complex scenes due to their lack of error correction capability. In this paper, we explore the use of conditional diffusion architectures for RGB-D SOD, producing saliency maps in a step-by-step generation paradigm. Accordingly, we proposeDiffRGBD, a novel diffusion-driven framework with temporal modulation. The core of DiffRGBD is using time steps to control the conditional information injected into the denoising network in a two-stage temporal modulation manner. Specifically, our DiffRGBD comprises a feature extractor, a conditional generator, two temporal modulators, and a denoising network. First, the SAM2 encoder with adapters is adopted to extract hierarchical cross-modal features. Then, the Mutual-Differential Attention Module is responsible for generating the conditional information via effective cross-modal fusion. Notably, the conditional information continuously achieves channel modulation and spatial modulation in the Temporal Channel Enhancement Module and the Temporal Spatial Refinement Module (i.e., two temporal modulators), resulting in comprehensive conditional information. Finally, conditional information is injected into the denoising network to guide the production of saliency maps. As the time step increases, our DiffRGBD can gradually correct errors and generate accurate saliency maps. Extensive experiments on seven public RGB-D SOD benchmarks demonstrate that our proposed DiffRGBD achieves superior performance over state-of-the-art methods. The code and results of our method are available at https://github.com/Shixiang02/DiffRGBD.
Shixiang Shi, Gongyang Li, Runmin Cong, Shunxin Xiao, Weisi Lin
IEEE Trans. Circuits Syst. Video Technol.2
2026 Reparameterization-Driven Depthwise Separable Large-Kernel Network for Lightweight Salient Object Detection of Strip Steel Surface Defects
abstract
With the rapid development of neural networks, strip steel surface defect detection, as an important task in computer vision, has achieved remarkable progress. However, state-of-the-art methods still face a tradeoff between accuracy and efficiency. High-performing models are usually large and computationally expensive, whereas lightweight models often suffer from limited detection accuracy. To address this issue, we first propose a spatial channel enhancement (SCE) module, which consists of a reparameterizable depthwise large-kernel convolution and a reparameterizable pointwise (RepPw) convolution. The proposed SCE module enlarges the receptive field and strengthens long-range spatial and channel interactions while preserving computational efficiency. Based on the SCE module, we propose a novel lightweight saliency model for strip steel surface defects, namely, reparameterization-driven depthwise separable large-kernel network (RepDSLKNet). RepDSLKNet employs SCE modules to build an encoder and a decoder, and utilizes cascaded channel attention (CCA) modules for the feature fusion. The lightweight architecture can effectively extract and fuse the semantic information and detailed features of strip steel surface defects, thereby improving the accuracy and speed of detection with a small model size. With an input size of $224 \times 224$ , our RepDSLKNet has only 0.47 M parameters and 0.42 G FLOPs during inference. Compared to the current state-of-the-art methods, our approach achieves a 19-fold improvement in throughput and a twofold reduction in latency. Experiments on two public strip steel defect datasets demonstrate that RepDSLKNet delivers competitive performance against state-of-the-art methods.
Xiaofei Zhou 0003, Zhenkun Mo, Gongyang Li, Liuxin Bao, Xiaobin Xu 0002, Jiyong Zhang 0001
IEEE Trans. Cybern.3
2026 MIST: A Benchmark and Baseline for Multi-Frame Infrared Small Target Detection in Complex Motion
abstract
Motion cues play a vital role in multi-frame infrared small target detection (MISTD). However, most targets in existing datasets exhibit regular and slow motion, which cannot reflect the complex and diverse motion patterns in real-world scenarios. This biased data distribution makes recent data-driven methods highly rely on simplified motion assumptions that tend to fail in irregular or fast motion, resulting in noisy feature representations cluttered with target-irrelevant factors. Hence, we stress that methods for MISTD should also work when targets are in complex motion. To enable this research, we propose a large-scale dataset called MIST for airborne infrared detection scenarios. The dataset is built on a synthetic data engine that models variations in pose, size, and intensity of moving targets while seamlessly blending them into real backgrounds for physical, geometric, and visual realism. Targets in MIST exhibit low signal-to-clutter ratios and complex motion, making it a promising yet challenging benchmark for developing algorithms focused on motion analysis. To tackle the challenges of MIST, we develop MISTNet, a robust baseline based on the Information Bottleneck theory. To handle irregular and fast motion, we propose a shifted neighborhood compensation block to efficiently model multi-scale correspondences for implicit motion compensation. To distill compact representations free from irrelevant cues, we design a progressive distillation decoder to hierarchically filter out redundancy while preserving target-relevant information. We benchmark 31 state-of-the-art methods and find that their performance on MIST drops significantly compared with that on the widely used NUDT-MIRSDT dataset. Our MISTNet outperforms all other methods by a large margin, with an over 6% gain in the IoU metric, demonstrating its superiority. The dataset, code, and model weights are available at https://github.com/GR-ray/MIST.
Meihong Zhang, Gongyang Li, Guanyi Li, Kai Zhao 0012, Xianchao Zhang 0002, Dan Zeng 0001
IEEE Trans. Image Process.3
2026 Few-Shot Strip Steel Surface Defect Segmentation via Pre-Trained Variational Auto-Encoder-Based Latent Gaussian Process Regression
abstract
Recently, few-shot strip steel surface defect segmentation has received more and more concerns. However, the existing few-shot segmentation methods usually adopt the frozen encoder, which is pre-trained on the classification task and can only provide class-related knowledge. Therefore, we propose a novel method, namely pre-trained variational auto-encoder based latent gaussian process regression (LGPR), to conduct few-shot strip steel surface defect segmentation. Firstly, different from previous methods, the frozen Variational Auto-Encoder (VAE) based encoder and decoder, which are pre-trained by using the pixel-level self-supervised task (i.e., image reconstruction), can provide rich image-related knowledge. This ensures the effective characterization of defect regions. Secondly, by deploying a gaussian process regression in the latent feature space generated by the VAE-based encoder, pixel-level correlation between support features and query features can be efficiently built. This operation is non-parametric and doesn't bring any training overhead. Besides, we deploy transformer-based projectors to dig long-range contextual cues of support and query features. Extensive experiments are performed on two public datasets, and the experimental results clearly show that our model consistently outperforms the state-of-the-art models with a large margin. Both the codes and results are publicly available at https://github.com/Hlao-hub/LGPR.
Xiaofei Zhou 0003, Gongyang Li, Deyang Liu, Qingshan She, Xiaobin Xu 0002, Runmin Cong
IEEE Trans. Image Process.3
2025 Few-shot fine-tuning with auxiliary tasks for video anomaly detection
Jing Lv, Zhi Liu 0003, Gongyang Li
Multim. Syst.3
2025 Feature Quality Assessment: A Database and A Lightweight Objective Method
abstract
In the era of Artificial Intelligence, visual data gathered by edge devices could be primarily utilized for machine vision tasks. The prominent coding frameworks accomplish this by extracting and compressing features extracted from input data. As such, the quality of these features is vital, as they reflect the performance of the coding framework. However, much less work has been dedicated to quality assessment on features, impeding the optimization of the coding system. In this work, we pioneer to explore the feature quality assessment by creating a novel database tailored for features, with the quality ground-truth for each feature. Then, we propose a lightweight feature quality assessment method, called Lightweight Feature Quality Assessment (LFQA). We analyze the feature characteristics from the perspective of spatial and channel thoroughly, and the framework of LFQA is designed based on the analysis results. Experimental results demonstrate that LFQA accurately evaluates the quality of features, reaching a notable Spearman Rank-Order Correlation Coefficient of 85.38%, and exhibits competitive performance in improving the performance of video coding for machine system. Furthermore, LFQA has fewer model parameters and faster inference speed, ensuring a wide range of promising applications.
Shipei Wang, Ping An 0001, Chao Yang 0021, Gongyang Li, Xinpeng Huang, Shiqi Wang 0001
IEEE Trans. Multim.4
2025 Dual-Guided Video Frame Interpolation With Spatial-Temporal Global Attention
abstract
Video frame interpolation technology improves visual experience with the development of deep learning. However, capturing large motions while synthesizing fine texture details remains a challenging task. Regarding large motion scenarios, some pioneering Transformer-based methods primarily rely on local attention, which does not fully leverage the global receptive field advantage. To address this issue, this paper proposes to further broaden the receptive field of the Transformer to capture more correlations in the video frame interpolation task. Specifically, we propose a global self-attention mechanism in the form of spatial-temporal separation. Regarding texture details, since roughly enlarging the receptive field results in the loss of details, we propose to use large motion information in both feature and pixel spaces as a dual-guided prior to enhance detail synthesis. The separable attention mechanism and the straightforward frame synthesis design significantly enhance the resource efficiency of our model. Extensive experiments show that our method achieves state-of-the-art performance, effectively capturing large motions and preserving texture details.
Baojun Zhou, Xinpeng Huang, Gongyang Li, Chao Yang 0021, Liquan Shen, Ping An 0001
IEEE Trans. Multim.3
2025 EMS: A Large-Scale Eye Movement Dataset, Benchmark, and New Model for Schizophrenia Recognition
abstract
Schizophrenia (SZ) is a common and disabling mental illness, and most patients encounter cognitive deficits. The eye-tracking technology has been increasingly used to characterize cognitive deficits for its reasonable time and economic costs. However, there is no large-scale and publicly available eye movement dataset and benchmark for SZ recognition. To address these issues, we release a large-scale Eye Movement dataset for SZ recognition (EMS), which consists of eye movement data from 104 schizophrenics and 104 healthy controls (HCs) based on the free-viewing paradigm with 100 stimuli. We also conduct the first comprehensive benchmark, which has been absent for a long time in this field, to compare the related 13 psychosis recognition methods using six metrics. Besides, we propose a novel mean-shift-based network (MSNet) for eye movement-based SZ recognition, which elaborately combines the mean shift algorithm with convolution to extract the cluster center as the subject feature. In MSNet, first, a stimulus feature branch (SFB) is adopted to enhance each stimulus feature with similar information from all stimulus features, and then, the cluster center branch (CCB) is utilized to generate the cluster center as subject feature and update it by the mean shift vector. The performance of our MSNet is superior to prior contenders, thus, it can act as a powerful baseline to advance subsequent study. To pave the road in this research field, the EMS dataset, the benchmark results, and the code of MSNet are publicly available at https://github.com/YingjieSong1/EMS.
Zhi Liu 0003, Gongyang Li, Qiang Wu 0001, Dan Zeng 0001, Lihua Xu, Tianhong Zhang, Jijun Wang 0003
IEEE Trans. Neural Networks Learn. Syst.3
2024 EFDCNet: Encoding fusion and decoding correction network for RGB-D indoor semantic segmentation
Jianlin Chen, Gongyang Li, Zhijiang Zhang, Dan Zeng 0001
Image Vis. Comput.2
2024 Audio-visual saliency prediction with multisensory perception and integration
Zhi Liu 0003, Gongyang Li
Image Vis. Comput.3
2024 Global semantic-guided network for saliency prediction
Zhi Liu 0003, Gongyang Li
Knowl. Based Syst.3
2024 Masked feature regeneration based asymmetric student-teacher network for anomaly detection
Haocheng Gu, Gongyang Li, Zhi Liu 0003
Multim. Tools Appl.2
2024 Light Field Salient Object Detection With Sparse Views via Complementary and Discriminative Interaction Network
abstract
4D light field data record the scene from multiple views, thus implicitly providing beneficial depth cue for salient object detection in challenging scenes. Existing light field salient object detection (LF SOD) methods usually use a large number of views to improve the detection accuracy. However, using so many views for LF SOD brings difficulties to its practical applications. Considering that adjacent views in a light field are actually with very similar contents, in this work, we propose defining a more efficient pattern of input views, i. e., key sparse views, and design a network to effectively explore the depth cue from sparse views for LF SOD. Specifically, we firstly introduce a low rank-based statistical analysis to the existing LF SOD datasets, which allows us to conclude a fixed yet universal pattern for our key sparse views, including the number and positions of views. These views maintain the sufficient depth cue, but greatly lower the number of views to be captured and processed, facilitating practical applications. Then, we propose an effective solution with a key Complementary and Discriminative Interaction Module (CDIM) for LF SOD from key sparse views, named CDINet. The CDINet follows a two-stream structure to extract the depth cue from the light field stream (i. e., sparse views) and the appearance cue from the RGB stream (i. e., center view), generating features and initial saliency maps for each stream. The CDIM is tailored for inter-stream interaction of both these features and saliency maps, using the depth cue to complement the missing salient regions in RGB stream and discriminate the background distraction, to enhance the final saliency map further. Extensive experiments on three LF multi-view datasets demonstrate that our CDINet not only outperforms the state-of-the-art 2D methods, but also achieves competitive performance as compared with the state-of-the-art 3D and 4D methods. The code and results of our method are available athttps://github.com/GilbertRC/LFSOD-CDINet.
Gongyang Li, Ping An 0001, Zhi Liu 0003, Xinpeng Huang, Qiang Wu 0001
IEEE Trans. Circuits Syst. Video Technol.2
2024 Context-Aware Interaction Network for RGB-T Semantic Segmentation
abstract
RGB-T semantic segmentation is a key technique for autonomous driving scenes understanding. For the existing RGB-T semantic segmentation methods, however, the effective exploration of the complementary relationship between different modalities is not implemented in the information interaction between multiple levels. To address such an issue, the Context-Aware Interaction Network (CAINet) is proposed for RGB-T semantic segmentation, which constructs interaction space to exploit auxiliary tasks and global context for explicitly guided learning. Specifically, we propose a Context-Aware Complementary Reasoning (CACR) module aimed at establishing the complementary relationship between multimodal features with the long-term context in both spatial and channel dimensions. Further, considering the importance of global contextual and detailed information, we propose the Global Context Modeling (GCM) module and Detail Aggregation (DA) module, and we introduce specific auxiliary supervision to explicitly guide the context interaction and refine the segmentation map. Extensive experiments on two benchmark datasets of MFNet and PST900 demonstrate that the proposed CAINet achieves state-of-the-art performance. The code is available athttps://github.com/YingLv1106/CAINet.
Zhi Liu 0003, Gongyang Li
IEEE Trans. Multim.3
2023 Exploring viewport features for semi-supervised saliency prediction in omnidirectional images
Mengke Huang, Gongyang Li, Zhi Liu 0003, Yong Wu 0007, Chen Gong 0002, Linchao Zhu, Yi Yang 0001
Image Vis. Comput.2
2023 Lightweight Distortion-Aware Network for Salient Object Detection in Omnidirectional Images
abstract
Compared with 2D image salient object detection (SOD), SOD in omnidirectional images (or 360° images) usually suffers from geometric distortion. Although existing omnidirectional image SOD (ODI-SOD) methods have improved the detection accuracy obviously, their application may be cumbersome in real scenes due to their high computational cost. To avoid distortion and reduce the computational cost simultaneously in ODI-SOD, we propose a novel lightweight distortion-aware network, named LDNet, in this letter. First, to extract features with less distortion from ODIs, we integrate the distortion-aware convolution and depth-wise separable convolution (DSConv) into distortion-aware DSConv (DDSConv) and replace the regular convolutions in the last two blocks of the ResNet-18 with DDSConvs to obtain our lightweight backbone network (LD-ResNet-18). To enhance spatial information in each channel of the extracted features at each level comprehensively, then, we propose a lightweight distortion-aware channel-wise enhancement (DCE) module (only 0.05M parameters) including DDSConvs with various dilation rates, channel shuffle operation and attention mechanism, and employ a high-to-low dense connection structure to modulate the enhanced multi-level features. Besides, we design a distortion-aware self-correlation (DSC) module (only 0.02M parameters) for mining the contextual dependency of the features via a coarse-fine strategy, and the correlated features are refined by DCE modules and integrated by another dense connection structure. The final saliency map is predicted from the densely integrated features. Compared with 12 state-of-the-art methods on two public datasets, our lightweight LDNet achieves competitive or even better performance with only 2.9M parameters and 3.4G FLOPs, which balances the efficiency and performance.
Mengke Huang, Gongyang Li, Zhi Liu 0003, Linchao Zhu
IEEE Trans. Circuits Syst. Video Technol.2
2023 RGB-T Semantic Segmentation With Location, Activation, and Sharpening
abstract
Semantic segmentation is important for scene understanding. To address the scenes of adverse illumination conditions of natural images, thermal infrared (TIR) images are introduced. Most existing RGB-T semantic segmentation methods follow three cross-modal fusion paradigms, i. e., encoder fusion, decoder fusion, and feature fusion. Some methods, unfortunately, ignore the properties of RGB and TIR features or the properties of features at different levels. In this paper, we propose a novel feature fusion-based network for RGB-T semantic segmentation, named LASNet, which follows three steps of location, activation, and sharpening. The highlight of LASNet is that we fully consider the characteristics of cross-modal features at different levels, and accordingly propose three specific modules for better segmentation. Concretely, we propose a Collaborative Location Module (CLM) for high-level semantic features, aiming to locate all potential objects. We propose a Complementary Activation Module for middle-level features, aiming to activate exact regions of different objects. We propose an Edge Sharpening Module (ESM) for low-level texture features, aiming to sharpen the edges of objects. Furthermore, in the training phase, we attach a location supervision and an edge supervision after CLM and ESM, respectively, and impose two semantic supervisions in the decoder part to facilitate network convergence. Experimental results on two public datasets demonstrate that the superiority of our LASNet over relevant state-of-the-art methods. The code and results of our method are available athttps://github.com/MathLee/LASNet.
Gongyang Li, Yike Wang 0003, Zhi Liu 0003, Xinpeng Zhang 0001, Dan Zeng 0001
IEEE Trans. Circuits Syst. Video Technol.1
2023 SGFNet: Semantic-Guided Fusion Network for RGB-Thermal Semantic Segmentation
abstract
Recently, semantic segmentation based on RGB and thermal infrared (TIR) images has become a research hotspot because of its stability in the weak light environment. However, most of the current methods ignore the differences between the two modalities of data and do not use semantic information in multi-modal fusion. In this paper, we propose a novel Semantic-Guided Fusion Network (SGFNet) for RGB-Thermal semantic segmentation, which makes full use of semantic information in the multi-modal fusion. Our SGFNet consists of an asymmetric encoder with TIR branch and RGB branch and a decoder. We concentrate on enhancing the multi-modal feature representation in the encoder with a pattern of fusion and enhancement. Specifically, considering that TIR images are stable under weak light conditions, we first propose a Semantic Guidance Head to extract semantic information in the TIR branch. In the RGB branch, we propose a Multi-modal Coordination and Distillation Unit to fuse multi-modal features first. Then, we propose a Cross-level and Semantic-guided Enhancement Unit to enhance the fused features with cross-level information and semantic information. We arrange these two units at all stages of the RGB branch to generate features with strong representation abilities at different levels. For the decoder, to obtain large receptive fields and fine edges, we improve the Lawin ASPP decoder by introducing edge information extracted from the low-level features, proposing the edge-aware Lawin ASPP decoder. With our encoder and decoder working together, our SGFNet can identify objects accurately and segment objects finely. Extensive experiments on the MFNet dataset demonstrate the superior performance of the proposed SGFNet compared with state-of-the-art methods. The code and results of our method are available athttps://github.com/kw717/SGFNet.
Yike Wang 0003, Gongyang Li, Zhi Liu 0003
IEEE Trans. Circuits Syst. Video Technol.2
2023 Adjacent Context Coordination Network for Salient Object Detection in Optical Remote Sensing Images
abstract
Salient object detection (SOD) in optical remote sensing images (RSIs), or RSI-SOD, is an emerging topic in understanding optical RSIs. However, due to the difference between optical RSIs and natural scene images (NSIs), directly applying NSI-SOD methods to optical RSIs fails to achieve satisfactory results. In this article, we propose a novel adjacent context coordination network (ACCoNet) to explore the coordination of adjacent features in an encoder-decoder architecture for RSI-SOD. Specifically, ACCoNet consists of three parts: 1) an encoder; 2) adjacent context coordination modules (ACCoMs); and 3) a decoder. As the key component of ACCoNet, ACCoM activates the salient regions of output features of the encoder and transmits them to the decoder. ACCoM contains a local branch and two adjacent branches to coordinate the multilevel features simultaneously. The local branch highlights the salient regions in an adaptive way, while the adjacent branches introduce global information of adjacent levels to enhance salient regions. In addition, to extend the capabilities of the classic decoder block (i.e., several cascaded convolutional layers), we extend it with two bifurcations and propose a bifurcation-aggregation block (BAB) to capture the contextual information in the decoder. Extensive experiments on two benchmark datasets demonstrate that the proposed ACCoNet outperforms 22 state-of-the-art methods under nine evaluation metrics, and runs up to 81 fps on a single NVIDIA Titan X GPU. The code and results of our method are available at https://github.com/MathLee/ACCoNet.
Gongyang Li, Zhi Liu 0003, Dan Zeng 0001, Weisi Lin, Haibin Ling
IEEE Trans. Cybern.1
2023 Lightweight Salient Object Detection in Optical Remote-Sensing Images via Semantic Matching and Edge Alignment
abstract
Recently, relying on convolutional neural networks (CNNs), many methods for salient object detection in optical remote-sensing images (ORSI-SOD) are proposed. However, most methods ignore the number of parameters and computational cost brought by CNNs, and only a few pay attention to portability and mobility. To facilitate practical applications, in this article, we propose a novel lightweight network for ORSI-SOD based on semantic matching and edge alignment, termed SeaNet. Specifically, SeaNet includes a lightweight MobileNet-V2 for feature extraction, a dynamic semantic matching module (DSMM) for high-level features, an edge self-alignment module (ESAM) for low-level features, and a portable decoder for inference. First, the high-level features are compressed into semantic kernels. Then, semantic kernels are used to activate salient object locations in two groups of high-level features through dynamic convolution operations in DSMM. Meanwhile, in ESAM, cross-scale edge information extracted from two groups of low-level features is self-aligned through$L_{2}$loss and used for detail enhancement. Finally, starting from the highest level features, the decoder infers salient objects based on the accurate locations and fine details contained in the outputs of the two modules. Extensive experiments on two public datasets demonstrate that our lightweight SeaNet not only outperforms most state-of-the-art lightweight methods, but also yields comparable accuracy with state-of-the-art conventional methods, while having only 2.76 M parameters and running with 1.7 G floating point operations (FLOPs) for$288 \times 288$inputs. Our code and results are available athttps://github.com/MathLee/SeaNet.
Gongyang Li, Zhi Liu 0003, Xinpeng Zhang 0001, Weisi Lin
IEEE Trans. Geosci. Remote. Sens.1
2023 Salient Object Detection in Optical Remote Sensing Images Driven by Transformer
abstract
Existing methods for Salient Object Detection in Optical Remote Sensing Images (ORSI-SOD) mainly adopt Convolutional Neural Networks (CNNs) as the backbone, such as VGG and ResNet. Since CNNs can only extract features within certain receptive fields, most ORSI-SOD methods generally follow the local-to-contextual paradigm. In this paper, we propose a novel Global Extraction Local Exploration Network (GeleNet) for ORSI-SOD following the global-to-local paradigm. Specifically, GeleNet first adopts a transformer backbone to generate four-level feature embeddings with global long-range dependencies. Then, GeleNet employs a Direction-aware Shuffle Weighted Spatial Attention Module (D-SWSAM) and its simplified version (SWSAM) to enhance local interactions, and a Knowledge Transfer Module (KTM) to further enhance cross-level contextual interactions. D-SWSAM comprehensively perceives the orientation information in the lowest-level features through directional convolutions to adapt to various orientations of salient objects in ORSIs, and effectively enhances the details of salient objects with an improved attention mechanism. SWSAM discards the direction-aware part of D-SWSAM to focus on localizing salient objects in the highest-level features. KTM models the contextual correlation knowledge of two middle-level features of different scales based on the self-attention mechanism, and transfers the knowledge to the raw features to generate more discriminative features. Finally, a saliency predictor is used to generate the saliency map based on the outputs of the above three modules. Extensive experiments on three public datasets demonstrate that the proposed GeleNet outperforms relevant state-of-the-art methods. The code and results of our method are available at https://github.com/MathLee/GeleNet.
Gongyang Li, Zhen Bai 0001, Zhi Liu 0003, Xinpeng Zhang 0001, Haibin Ling
IEEE Trans. Image Process.1
2023 Adaptive Group-Wise Consistency Network for Co-Saliency Detection
abstract
Co-saliency detection focuses on detecting common and salient objects among a group of images. With the application of deep learning in co-saliency detection, more accurate and more effective models are proposed in an end-to-end manner. However, two major drawbacks in these models hinder the further performance improvement of co-saliency detection: 1) the static manner-based inference, and 2) the constant quantity of input images. To address these limitations, we present a novel Adaptive Group-wise Consistency Network (AGCNet) with the ability of content-adaptive adjustment for a given image group with random quantity of images. In AGCNet, we first introduce intra-saliency priors generated from any off-the-shelf salient object detection model. Then, an Adaptive Group-wise Consistency (AGC) module is proposed to capture group consistency for each individual image, and is applied on three-scale features to capture the group consistency from different perspectives. This module is composed of two key components, where the content-adaptive group consistency block breaks the above limitations to adaptively capture the global group consistency with the assistance of intra-saliency priors and the ranking-based fusion block combines the consistency with individual attributes of each image feature to generate discriminative group consistency feature for each image. Following AGC modules, a specially designed Aggregated Decoder aggregates the three-scale group consistency features to adapt to co-salient objects with diverse scales for preliminary detection. Finally, we incorporate two normal decoders to progressively refine the preliminary detection and generate the final co-saliency maps. Extensive experiments on four benchmark datasets demonstrate that our AGCNet achieves competitive performance as compared with 19 state-of-the-art models, and the proposed modules experimentally show substantial practical merits.
Zhen Bai 0001, Zhi Liu 0003, Gongyang Li, Yang Wang 0003
IEEE Trans. Multim.3
2023 RINet: Relative Importance-Aware Network for Fixation Prediction
abstract
Fixation prediction aims to simulate human visual selection mechanism and estimate the visual saliency degree of regions in a scene. In semantically rich scenes, there are generally multiple salient regions. This condition requires a fixation prediction model to understand the relative importance relationship of multiple salient regions, that is, to identify which region is more important. In practice, existing fixation prediction models implicitly explore the relative importance relationship in the end-to-end training process while they do not work well. In this article, we propose a novel Relative Importance-aware Network (RINet) to explicitly explore the modeling of relative importance in fixation prediction. RINet perceives multi-scale local and global relative importance through the Hierarchical Relative Importance Enhancement (HRIE) module. Within a single scale subspace, on the one hand, HRIE module regards the similarity matrix as the local relative importance map to weight the input feature. On the other hand, HRIE module integrates a set of local relative importance maps into one map, defined as the global relative importance map, to grasp global relative importance. Moreover, we propose a Complexity-Relevant Focal (CRF) loss for network training. As such, we can progressively emphasize learning difficult samples for better handling the complicated scenarios, further improving the performance. The ablation studies confirm the contributions of key components of our RINet, and extensive experiments on five datasets demonstrate our RINet is superior to 28 relevant state-of-the-art models.
Zhi Liu 0003, Gongyang Li, Dan Zeng 0001, Tianhong Zhang, Lihua Xu, Jijun Wang 0003
IEEE Trans. Multim.3
2023 Spatio-Temporal Self-Attention Network for Video Saliency Prediction
abstract
3D convolutional neural networks have achieved promising results for video tasks in computer vision, including video saliency prediction that is explored in this paper. However, 3D convolution encodes visual representation merely on fixed local spacetime according to its kernel size, while human attention is always attracted by relational visual features at different time. To overcome this limitation, we propose a novel Spatio-Temporal Self-Attention 3D Network (STSANet) for video saliency prediction, in which multiple Spatio-Temporal Self-Attention (STSA) modules are employed at different levels of 3D convolutional backbone to directly capture long-range relations between spatio-temporal features of different time steps. Besides, we propose an Attentional Multi-Scale Fusion (AMSF) module to integrate multi-level features with the perception of context in semantic and spatio-temporal subspaces. Extensive experiments demonstrate the contributions of key components of our method, and the results on DHF1K, Hollywood-2, UCF, and DIEM benchmark datasets clearly prove the superiority of the proposed model compared with all state-of-the-art models.
Ziqiang Wang 0003, Zhi Liu 0003, Gongyang Li, Yang Wang 0003, Tianhong Zhang, Lihua Xu, Jijun Wang 0003
IEEE Trans. Multim.3
2022 Gaze Estimation via Modulation-Based Adaptive Network With Auxiliary Self-Learning
abstract
Given a face image, most of previous works in gaze estimation infer the gaze via a well-trained model with supervised training. However, the distribution of test data may be very different compared to that of training data since samples might be corrupted in real-world scenarios (e.g., taking a photo in strong light). This will lead to a gap between source domain (i.e., training data) and target domain (i.e., test data). In this paper, we first introduce self-supervised learning into our method for addressing challenging situations in gaze estimation. Moreover, existing appearance-based gaze estimation methods focus on directing towards the development of powerful regressors, which mainly utilize face and eye images simultaneously or face (eye) images only. However, the problem of inter cues between face and eye features has been largely overlooked. To this end, we propose a novel Modulation-based Adaptive Network (MANet) for gaze estimation, which uses high-level knowledge to filter the distractive information and bridges the intrinsic relationship between face and eye features. Further, we combine self-supervised learning and MANet to learn to adapt to challenging cases, such as abnormal lighting conditions and poor-quality images, by minimizing a self-supervised loss and a supervised loss jointly. The experimental results on several datasets demonstrate the effectiveness of our proposed approach with a real-time speed of 900fpson a PC with an NVIDIA Titan RTX GPU.
Yong Wu 0007, Gongyang Li, Zhi Liu 0003, Mengke Huang, Yang Wang 0003
IEEE Trans. Circuits Syst. Video Technol.2
2022 Lightweight Salient Object Detection in Optical Remote Sensing Images via Feature Correlation
abstract
Salient object detection in optical remote sensing images (ORSI-SOD) has been widely explored for understanding ORSIs. However, previous methods focus mainly on improving the detection accuracy while neglecting the cost in memory and computation, which may hinder their real-world applications. In this article, we propose a novel lightweight ORSI-SOD solution, named CorrNet, to address these issues. In CorrNet, we first lighten the backbone (VGG-16) and build a lightweight subnet for feature extraction. Then, following the coarse-to-fine strategy, we generate an initial coarse saliency map from high-level semantic features in a correlation module (CorrM). The coarse saliency map serves as the location guidance for low-level features. In CorrM, we mine the object location information between high-level semantic features through the cross-layer correlation operation. Finally, based on low-level detailed features, we refine the coarse saliency map in the refinement subnet equipped with dense lightweight refinement blocks (DLRBs) and produce the final fine saliency map. By reducing the parameters and computations of each component, CorrNet ends up having only 4.09M parameters and running with 21.09G FLOPs. Experimental results on two public datasets demonstrate that our lightweight CorrNet achieves competitive or even better performance compared with 26 state-of-the-art methods (including 16 large CNN-based methods and two lightweight methods), and meanwhile enjoys the clear memory and run-time efficiency. The code and results of our method are available athttps://github.com/MathLee/CorrNet.
Gongyang Li, Zhi Liu 0003, Zhen Bai 0001, Weisi Lin, Haibin Ling
IEEE Trans. Geosci. Remote. Sens.1
2022 Multi-Content Complementation Network for Salient Object Detection in Optical Remote Sensing Images
abstract
In the computer vision community, great progresses have been achieved in salient object detection from natural scene images (NSI-SOD); by contrast, salient object detection in optical remote sensing images (RSI-SOD) remains to be a challenging emerging topic. The unique characteristics of optical RSIs, such as scales, illuminations, and imaging orientations, bring significant differences between NSI-SOD and RSI-SOD. In this article, we propose a novel multi-content complementation network (MCCNet) to explore the complementarity of multiple content for RSI-SOD. Specifically, MCCNet is based on the general encoder–decoder architecture, and contains a novel key component named multi-content complementation module (MCCM), which bridges the encoder and the decoder. In MCCM, we consider multiple types of features that are critical to RSI-SOD, including foreground features, edge features, background features, and global image-level features, and exploit the content complementarity between them to highlight salient regions over various scales in RSI features through the attention mechanism. Besides, we comprehensively introduce pixel-level, map-level, and metric-aware losses in the training phase. Extensive experiments on two popular datasets demonstrate that the proposed MCCNet outperforms 23 state-of-the-art methods, including both NSI-SOD and RSI-SOD methods. The code and results of our method are available athttps://github.com/MathLee/MCCNet.
Gongyang Li, Zhi Liu 0003, Weisi Lin, Haibin Ling
IEEE Trans. Geosci. Remote. Sens.1
2021 Circular Complement Network for RGB-D Salient Object Detection
Zhen Bai 0001, Zhi Liu 0003, Gongyang Li, Linwei Ye, Yang Wang 0003
Neurocomputing3
2021 Personalized image observation behavior learning in fixation based personalized salient object segmentation
Gongyang Li, Weijie Wei 0001, Xiaofei Zhou 0003, Zhi Liu 0003
Neurocomputing2
2021 Hierarchical Alternate Interaction Network for RGB-D Salient Object Detection
abstract
Existing RGB-D Salient Object Detection (SOD) methods take advantage of depth cues to improve the detection accuracy, while pay insufficient attention to the quality of depth information. In practice, a depth map is often with uneven quality and sometimes suffers from distractors, due to various factors in the acquisition procedure. In this article, to mitigate distractors in depth maps and highlight salient objects in RGB images, we propose a Hierarchical Alternate Interactions Network (HAINet) for RGB-D SOD. Specifically, HAINet consists of three key stages: feature encoding, cross-modal alternate interaction, and saliency reasoning. The main innovation in HAINet is the Hierarchical Alternate Interaction Module (HAIM), which plays a key role in the second stage for cross-modal feature interaction. HAIM first uses RGB features to filter distractors in depth features, and then the purified depth features are exploited to enhance RGB features in turn. The alternate RGB-depth-RGB interaction proceeds in a hierarchical manner, which progressively integrates local and global contexts within a single feature scale. In addition, we adopt a hybrid loss function to facilitate the training of HAINet. Extensive experiments on seven datasets demonstrate that our HAINet not only achieves competitive performance as compared with 19 relevant state-of-the-art methods, but also reaches a real-time processing speed of 43 fps on a single NVIDIA Titan X GPU. The code and results of our method are available at https://github.com/MathLee/HAINet.
Gongyang Li, Zhi Liu 0003, Minyu Chen 0001, Zhen Bai 0001, Weisi Lin, Haibin Ling
IEEE Trans. Image Process.1
2021 Personal Fixations-Based Object Segmentation With Object Localization and Boundary Preservation
abstract
As a natural way for human-computer interaction, fixation provides a promising solution for interactive image segmentation. In this paper, we focus on Personal Fixations-based Object Segmentation (PFOS) to address issues in previous studies, such as the lack of appropriate dataset and the ambiguity in fixations-based interaction. In particular, we first construct a new PFOS dataset by carefully collecting pixel-level binary annotation data over an existing fixation prediction dataset, such dataset is expected to greatly facilitate the study along the line. Then, considering characteristics of personal fixations, we propose a novel network based on Object Localization and Boundary Preservation (OLBP) to segment the gazed objects. Specifically, the OLBP network utilizes an Object Localization Module (OLM) to analyze personal fixations and locates the gazed objects based on the interpretation. Then, a Boundary Preservation Module (BPM) is designed to introduce additional boundary information to guard the completeness of the gazed objects. Moreover, OLBP is organized in the mixed bottom-up and top-down manner with multiple types of deep supervision. Extensive experiments on the constructed PFOS dataset show the superiority of the proposed OLBP network over 17 state-of-the-art methods, and demonstrate the effectiveness of the proposed OLM and BPM components. The constructed PFOS dataset and the proposed OLBP network are available at https://github.com/MathLee/OLBPNet4PFOS.
Gongyang Li, Zhi Liu 0003, Weijie Wei 0001, Yong Wu 0007, Mengke Huang, Haibin Ling
IEEE Trans. Image Process.1
2020 Cross-Modal Weighting Network for RGB-D Salient Object Detection
Gongyang Li, Zhi Liu 0003, Linwei Ye, Yang Wang 0003, Haibin Ling
ECCV (17)1
2020 Co-Saliency Detection Using Collaborative Feature Extraction And High-To-Low Feature Integration
abstract
Co-saliency detection, as a developing research branch of saliency detection, devotes to identify the common salient objects in a group of related images. The major challenge of co-saliency detection is how to effectively represent features considering both intra-image and inter-image information. In this paper, we propose a co-saliency detection model using collaborative feature extraction and high-to-low feature integration. We first feed the target image and its co-images into the Individual Feature Extraction Module (IFEM) to produce multi-level individual features. Then, to capture the collaborative inter-image information, the Collaborative Feature Extraction Module (CFEM) is applied to all highest-level individual features, generating the collaborative feature. Finally, we build a High-to-low Feature Integration Module (HFIM), which integrates the collaborative feature and multi-level individual features of the target image, to enrich the collaborative feature with individual intra-image information. Extensive experiments on two public datasets demonstrate that the proposed model achieves the state-of-the-art performance.
Jingru Ren, Zhi Liu 0003, Gongyang Li, Xiaofei Zhou 0003, Cong Bai, Guangling Sun
ICME3
2020 Fixations based personal target objects segmentation
abstract
With the development of the eye-tracking technique, the fixation becomes an emergent interactive mode in many human-computer interaction study field. For a personal target objects segmentation task, although the fixation can be taken as a novel and more convenient interactive input, it induces a heavy ambiguity problem of the input's indication so that the segmentation quality is severely degraded. In this paper, to address this challenge, we develop an "extraction-to-fusion" strategy based iterative lightweight neural network, whose input is composed by an original image, a fixation map and a position map. Our neural network consists of two main parts: The first extraction part is a concise interlaced structure of standard convolution layers and progressively higher dilated convolution layers to better extract and integrate local and global features of target objects. The second fusion part is a convolutional long short-term memory component to refine the extracted features and store them. Depending on the iteration framework, current extracted features are refined by fusing them with stored features extracted in the previous iterations, which is a feature transmission mechanism in our neural network. Then, current improved segmentation result is generated to further adjust the fixation map and the position map in the next iteration. Thus, the ambiguity problem induced by the fixations can be alleviated. Experiments demonstrate better segmentation performance of our method and effectiveness of each part in our model.
Gongyang Li, Weijie Wei 0001, Zhi Liu 0003
MMAsia2
2020 Attention-guided RGBD saliency detection using appearance information
Xiaofei Zhou 0003, Gongyang Li, Chen Gong 0002, Zhi Liu 0003, Jiyong Zhang 0001
Image Vis. Comput.2
2020 Weakly supervised instance segmentation using multi-stage erasing refinement and saliency-guided proposals ordering
Zhi Liu 0003, Gongyang Li, Linwei Ye, Lei Zhou 0003, Yang Wang 0003
J. Vis. Commun. Image Represent.3
2020 FANet: Features Adaptation Network for 360$^{\circ }$ Omnidirectional Salient Object Detection
abstract
Salient object detection (SOD) in 360° omnidirectional images has become an eye-catching problem because of the popularity of affordable 360° cameras. In this paper, we propose a Features Adaptation Network (FANet) to highlight salient objects in 360° omnidirectional images reliably. To utilize the feature extraction capability of convolutional neural networks and capture global object information, we input the equirectangular 360° images and corresponding cube-map 360° images to the feature extraction network (FENet) simultaneously to obtain multi-level equirectangular and cube-map features. Furthermore, we fuse these two kinds of features at each level of the FENetby a projection features adaptation (PFA) module, for selecting these two kinds of features adaptively. Finally, we combine the preliminary adaptation features at different levels by a multi-level features adaptation (MLFA) module, which weights these different-level features adaptively and produces the final saliency maps. Experiments show our FANet outperforms the state-of-the-art methods on the 360° omnidirectional SOD datasets.
Mengke Huang, Zhi Liu 0003, Gongyang Li, Xiaofei Zhou 0003, Olivier Le Meur
IEEE Signal Process. Lett.3
2020 ICNet: Information Conversion Network for RGB-D Based Salient Object Detection
abstract
RGB-D based salient object detection (SOD) methods leverage the depth map as a valuable complementary information for better SOD performance. Previous methods mainly resort to exploit the correlation between RGB image and depth map in three fusion domains: input images, extracted features, and output results. However, these fusion strategies cannot fully capture the complex correlation between the RGB image and depth map. Besides, these methods do not fully explore the cross-modal complementarity and the cross-level continuity of information, and treat information from different sources without discrimination. In this paper, to address these problems, we propose a novel Information Conversion Network (ICNet) for RGB-D based SOD by employing the siamese structure with encoder-decoder architecture. To fuse high-level RGB and depth features in an interactive and adaptive way, we propose a novel Information Conversion Module (ICM), which contains concatenation operations and correlation layers. Furthermore, we design a Cross-modal Depth-weighted Combination (CDC) block to discriminate the cross-modal features from different sources and to enhance RGB features with depth features at each level. Extensive experiments on five commonly tested datasets demonstrate the superiority of our ICNet over 15 state-of-theart RGB-D based SOD methods, and validate the effectiveness of the proposed ICM and CDC block.
Gongyang Li, Zhi Liu 0003, Haibin Ling
IEEE Trans. Image Process.1
2019 Constrained fixation point based segmentation via deep neural network
Gongyang Li, Zhi Liu 0003, Weijie Wei 0001
Neurocomputing1
2019 Effective online refinement for video object segmentation
Gongyang Li, Zhi Liu 0003, Xiaofei Zhou 0003
Multim. Tools Appl.1
2018 Video Saliency Detection Using Deep Convolutional Neural Networks
Xiaofei Zhou 0003, Zhi Liu 0003, Chen Gong 0002, Gongyang Li, Mengke Huang
PRCV (2)4