Jing Sun 0012

dblp:181/2880-12 · DBLP profile ↗
← Back
26ranked-venue papers
5as first author
25since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 3 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 8 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 RS-GCL: Randomized SVD-based graph-enhanced contrastive learning for recommendation
Jing Sun 0012, Xingchen Peng, Ganghui Li, Fangmei Chen
Expert Syst. Appl.1
2026 MSINet: Mantis shrimp vision-inspired network for camouflaged object detection
Houjie Li, Jing Sun 0012, Fuming Sun
Expert Syst. Appl.4
2026 Infrared and visible image fusion based on multi-modal and multi-scale cross-compensation
Meitian Li, Jing Sun 0012, Fasheng Wang, Fuming Sun
Knowl. Based Syst.2
2026 LESOD: Lightweight and efficient network for RGB-D salient object detection
Mingyu Zhong, Jing Sun 0012, Fasheng Wang, Fuming Sun
Pattern Recognit.2
2026 DFD-Stereo: Dual-Domain Feature Decoupling for Stereo Matching
abstract
Accurate stereo matching under limited computational resources remains a central challenge in 3D perception tasks such as autonomous driving and robot navigation. Existing high-accuracy methods often rely on heavy architectures with significant memory and processing demands, while lightweight models typically compromise on feature expressiveness, leading to limited global understanding and detail loss. To bridge this gap, we propose DFD-Stereo, a lightweight and efficient stereo matching framework that delivers high-quality disparity estimation with reduced computational cost. The framework incorporates two key components: (1) a Decoupled Frequency-Spatial Learning (DFSL) module, which enables complementary spatial-frequency representation for enhanced global context modeling, and (2) a Stepwise Coupling Disparity Refinement (SCDR) module, which leverages multi-scale RGB-disparity fusion with Shuffle Attention to refine disparity predictions effectively. Experimental results across multiple benchmarks demonstrate that DFD-Stereo achieves superior accuracy with significantly improved efficiency, offering a promising solution for deployment in resource-constrained 3D vision systems.
Zhisheng Zhu, Fuming Sun, Jing Sun 0012
IEEE Trans. Circuits Syst. Video Technol.3
2025 Seeing the Unseen: A Semantic Alignment and Context-Aware Prompt Framework for Open-Vocabulary Camouflaged Object Segmentation
Tian Bai 0002, Jing Sun 0012, Fuming Sun
ICCV3
2025 Knowledge-guided and Collaborative Learning Network for Camouflaged Object Detection
Mengyin Wang, Jing Sun 0012
Eng. Appl. Artif. Intell.3
2025 Exploring a Lightweight and Efficient Network for Salient Object Detection in ORSI
abstract
In recent years, Optical Remote Sensing Image Salient Object Detection (ORSI-SOD) has made substantial progress. Nevertheless, it remains an open-ended research area with complex challenges. Most existing ORSI-SOD methods, aiming for high-performance detection, demand large-scale parameters and high computational costs. This significantly restricts their application on resource-constrained devices, which have limited computing power and memory capacity. To tackle this issue, we propose a lightweight and highly efficient ORSI-SOD network, termed RAMENet. With only 5.18M parameters and 8.72G FLOPs, RAMENet can achieve competitive detection accuracy compared to state-of-the-art methods. Specifically, we devise a Dynamic Region-aware Block (DRB) that can be nested within the encoder to realize plug-and-play functionality. This enables the network to learn ORSI domain-specific feature representations, thus more effectively locating salient object regions. Furthermore, we present a novel Multi-path Enhanced M-shaped Decoder (MED), which integrates both bottom-up and top-down paradigms. Comprising two feature extraction sub-branches and a master feature refinement branch, this architecture achieves multi-granularity feature aggregation via cross-level feature interaction. Consequently, it significantly improves the detailed representation capability while maintaining the integrity of the object structure. Extensive experimental results indicate that the RAMENet outperforms 5 state-of-the-art lightweight methods in terms ofSα,Fβmean,MAEon EORSSD and ORSSD datasets, with improvement reaching 0.68%, 0.92%, 0.13%, 0.60%, 1.13%, and 0.07%, respectively. The code and results are available at https://github.com/hjy0518/RAMENet/.
Jinyu Han, Fuming Sun, Yaoyao Hou, Jing Sun 0012
IEEE Trans. Geosci. Remote. Sens.4
2025 ORSIDiff: Diffusion Model for Salient Object Detection in Optical Remote Sensing Images
abstract
The unique imaging conditions of satellites introduce significant uncertainties in the structure and scale of ground objects, presenting a major challenge for Optical Remote Sensing Image Salient Object Detection (ORSI-SOD). Current ORSI-SOD methods often fail to effectively differentiate between salient objects and subtle background variations, leading to suboptimal prediction outcomes. Furthermore, ORSI-SOD is a dense pixel prediction task, and existing approaches frequently depend on pixel-level probabilities, which can result in overconfident and inaccurate predictions. To address these challenges, we reformulate the ORSI-SOD task as a mask-generation problem by introducing a novel paradigm and propose a diffusion model-based method for ORSI-SOD, termed ORSIDiff. Central to our approach is the design of a powerful denoising network that enhances the model’s refinement capabilities. This network leverages the strengths of both global and local modeling, improving the handling of salient object details and enabling a deeper understanding of the distinctions between salient objects and their surroundings. Additionally, we introduce a consistency assessment strategy that aggregates multiple potential predictions during the denoising process, effectively mitigating the issue of overconfident point estimation. Extensive experimental results on two widely used ORSI-SOD datasets demonstrate that ORSIDiff achieves significant performance improvements over 20 state-of-the-art methods.
Jinyu Han, Jing Sun 0012, Fasheng Wang, Fuming Sun
IEEE Trans. Geosci. Remote. Sens.2
2025 A UNet-Like Transformer Network for Camouflaged Object Detection
abstract
The role of Camouflaged Object Detection (COD) is to identify the objects that integrate seamlessly with the surrounding environment. Due to the high intrinsic similarity between the objects and their background, this task presents greater challenges than traditional object detection. Most existing COD methods often have a large number of parameters and high computational complexity in the pursuit of detection accuracy, which hinders the application of COD in practical scenarios. To address this issue, we propose a UNet-like Transformer Network for COD, termed UTNet, which achieves competitive detection accuracy with a smaller parameter set. Specifically, we propose a Camouflaged Region Awareness Module (CRAM) consisting of a Hierarchical Attention Mechanism (HAM) that groups features to reveal intrinsic consistency between sub-features. This CRAM can be embedded into the backbone network, giving it powerful modeling capabilities. And, we present a Contextual Knowledge Collector (CKC) that exploits a cross-aggregation approach for neighboring feature layers, promoting the flow of semantic information from high-level to low-level features, and ensuring the integrity of camouflaged objects at each level of features. Furthermore, we introduce a progressive decoder that utilizes a cascade of attention units to filter noise and explores knowledge aggregation to emphasize features from different levels, ensuring that camouflaged objects have complete spatial details at the local level. Extensive experimental results show that UTNet achieves competitive results compared to 20 state-of-the-art methods. Codes and results are released onhttps://github.com/hjy0518/UTNet.
Fuming Sun, Jinyu Han, Weiyi Wu, Jing Sun 0012, Mengyin Wang
IEEE Trans. Multim.4
2024 Behavior-Contextualized Item Preference Modeling for Multi-Behavior Recommendation
abstract
In recommender systems, multi-behavior methods have demonstrated their effectiveness in mitigating issues like data sparsity, a common challenge in traditional single-behavior recommendation approaches. These methods typically infer user preferences from various auxiliary behaviors and apply them to the target behavior for recommendations. However, this direct transfer can introduce noise to the target behavior in recommendation, due to variations in user attention across different behaviors. To address this issue, this paper introduces a novel approach, Behavior-Contextualized Item Preference Modeling (BCIPM), for multi-behavior recommendation. Our proposed Behavior-Contextualized Item Preference Network discerns and learns users' specific item preferences within each behavior. It then considers only those preferences relevant to the target behavior for final recommendations, significantly reducing noise from auxiliary behaviors. These auxiliary behaviors are utilized solely for training the network parameters, thereby refining the learning process without compromising the accuracy of the target behavior recommendations. To further enhance the effectiveness of BCIPM, we adopt a strategy of pre-training the initial embeddings. This step is crucial for enriching the item-aware preferences, particularly in scenarios where data related to the target behavior is sparse. Comprehensive experiments conducted on four real-world datasets demonstrate BCIPM's superior performance compared to several leading state-of-the-art models, validating the robustness and efficiency of our proposed approach.
Mingshi Yan, Fan Liu 0008, Jing Sun 0012, Fuming Sun, Zhiyong Cheng 0001, Yahong Han
SIGIR3
2024 Image deblurring method based on self-attention and residual wavelet transform
Bing Zhang 0023, Jing Sun 0012, Fuming Sun, Fasheng Wang
Expert Syst. Appl.2
2024 Cross-Modal Fusion and Progressive Decoding Network for RGB-D Salient Object Detection
Xihang Hu, Fuming Sun, Jing Sun 0012, Fasheng Wang
Int. J. Comput. Vis.3
2024 Enhancing Collaborative Information with Contrastive Learning for Session-based Recommendation
Guojia An, Jing Sun 0012, Fuming Sun
Inf. Process. Manag.2
2024 MAGNet: Multi-scale Awareness and Global fusion Network for RGB-D salient object detection
Mingyu Zhong, Jing Sun 0012, Fasheng Wang, Fuming Sun
Knowl. Based Syst.2
2024 MadFormer: multi-attention-driven image super-resolution method based on Transformer
Jing Sun 0012, Ting Li 0002, Fuming Sun
Multim. Syst.2
2024 Efficient Camouflaged Object Detection Network Based on Global Localization Perception and Local Guidance Refinement
abstract
Camouflaged Object Detection (COD) is a challenging visual task due to its complex contour, diverse scales, and high similarity to the background. Existing COD methods encounter two predicaments: One is that they are prone to falling into local perception, resulting in inaccurate object localization; Another issue is the difficulty in achieving precise object segmentation due to a lack of detailed information. In addition, most COD methods typically require larger parameter amounts and higher computational complexity in pursuit of better performance. To this end, we propose a global localization perception and local guidance refinement network (PRNet), that simultaneously addresses performance and computational costs. Through effective aggregation and use of semantic and details information, the PRNet can achieve accurate localization and refined segmentation of camouflaged objects. Specifically, with the help of a Cascaded Attention Perceptron (CAP) designed, we can effectively integrate and perceive multi-scale information to localize camouflaged objects. We also design a Guided Refinement Decoder (GRD) in a top-down manner to extract context information and aggregate details to further refine camouflaged prediction results. Extensive experimental results demonstrate that our PRNet outperforms 12 state-of-the-art models on 4 challenging datasets. Meanwhile, the PRNet has a smaller number of parameters (12.74M), lower computational complexity (10.24G), and real-time inference speed (105FPS). Source codes are available at https://github.com/hu-xh/PRNet.
Xihang Hu, Xiaoli Zhang 0001, Fasheng Wang, Jing Sun 0012, Fuming Sun
IEEE Trans. Circuits Syst. Video Technol.4
2024 Low-light image enhancement using transformer with color fusion and channel attention
Yinbang Sun, Jing Sun 0012, Fuming Sun, Fasheng Wang
J. Supercomput.2
2024 Cascading Residual Graph Convolutional Network for Multi-Behavior Recommendation
abstract
Multi-behavior recommendation exploits multiple types of user-item interactions, such as view and cart , to learn user preferences and has demonstrated to be an effective solution to alleviate the data sparsity problem faced by the traditional models that often utilize only one type of interaction for recommendation. In real scenarios, users often take a sequence of actions to interact with an item, in order to get more information about the item and thus accurately evaluate whether an item fits their personal preferences. Those interaction behaviors often obey a certain order, and more importantly, different behaviors reveal different information or aspects of user preferences towards the target item. Most existing multi-behavior recommendation methods take the strategy to first extract information from different behaviors separately and then fuse them for final prediction. However, they have not exploited the connections between different behaviors to learn user preferences. Besides, they often introduce complex model structures and more parameters to model multiple behaviors, largely increasing the space and time complexity. In this work, we propose a lightweight multi-behavior recommendation model named Cascading Residual Graph Convolutional Network ( CRGCN for short) for multi-behavior recommendation, which can explicitly exploit the connections between different behaviors into the embedding learning process without introducing any additional parameters (with comparison to the single-behavior based recommendation model). In particular, we design a cascading residual graph convolutional network (GCN) structure, which enables our model to learn user preferences by continuously refining the embeddings across different types of behaviors. The multi-task learning method is adopted to jointly optimize our model based on different behaviors. Extensive experimental results on three real-world benchmark datasets show that CRGCN can substantially outperform the state-of-the-art methods, achieving 24.76%, 27.28%, and 25.10% relative gains on average in terms of HR@K (K = {10,20,50,80}) over the best baseline across the three datasets. Further studies also analyze the effects of leveraging multi-behaviors in different numbers and orders on the final performance.
Mingshi Yan, Zhiyong Cheng 0001, Chen Gao 0001, Jing Sun 0012, Fan Liu 0008, Fuming Sun
ACM Trans. Inf. Syst.4
2024 Attention-guided Multi-modality Interaction Network for RGB-D Salient Object Detection
abstract
The past decade has witnessed great progress in RGB-D salient object detection (SOD). However, there are two bottlenecks that limit its further development. The first one is low-quality depth maps. Most existing methods directly use raw depth maps to perform detection, but low-quality depth images can bring negative impacts to the detection performance. Hence, it is not desirable to utilize depth maps indiscriminately. The other one is how to effectively predict salient maps with clear boundary and complete salient region. To address these problems, an Attention-Guided Multi-Modality Interaction Network (AMINet) is proposed. First, we propose a new quality enhancement strategy for unreliable depth images, named D epth E nhancement M odule ( DEM ). With respect to the second issue, we propose C ross- M odality A ttention M odule ( CMAM ) to rapidly locate salient region. The B oundary- A ware M odule ( BAM ) is designed to utilize high-level feature to guide the low-level feature generation in a top-down way to make up for the dilution of the boundary. To further improve the accuracy, we propose A trous R efined B lock ( ARB ) to adaptively compensate for the shortcoming of atrous convolution. By integrating these interactive modules, features from depth and RGB streams can be refined efficiently, which consequently boosts the detection performance. Experimental results demonstrate the proposed AMINet exceeds state-of-the-art (SOTA) methods on several public RGB-D datasets.
Fasheng Wang, Yiming Su, Jing Sun 0012, Fuming Sun
ACM Trans. Multim. Comput. Commun. Appl.4
2023 Attention-guided graph convolutional network for multi-behavior recommendation
Xingchen Peng, Jing Sun 0012, Mingshi Yan, Fuming Sun, Fasheng Wang
Knowl. Based Syst.2
2023 A cross-view geo-localization method guided by relation-aware global attention
Jing Sun 0012, Bing Zhang 0023, Fuming Sun
Multim. Syst.1
2022 Joint Adaptive Dual Graph and Feature Selection for Domain Adaptation
abstract
Domain adaptation aims to exploit domain-invariant features by aligning the cross-domain distributions in the manifold subspace for applying the classifier trained on the source domain to the target domain. However, two limitations may still deteriorate their performances: (1) the influences of noisy or irrelevant features in the original feature space are ignored, which may unexpectedly hurt the classification of target samples; (2) the graph constructed directly in the original data space cannot accurately capture the inherent local manifold structures of high-dimensional data due to the curse of dimensionality, which may seriously mislead the transferable features learning. In this paper, we propose a novel approach to address these problems, referred to as joint Adaptive Dual Graph and Feature Selection for domain adaptation (ADGFS). Specifically, feature selection can characterize the relative importance of different features through a scaling factor, which enables ADGFS to not only reduce the impacts of noisy or irrelevant features on knowledge transfer but also learn informative domain-invariant features. Meanwhile, ADGFS adaptively optimizes the dual graph by learning the similarity matrices of both instance-level and feature-level graphs in the projected low-dimensional manifold subspace rather than the original high-dimensional space, such that the intrinsic local manifold structures of data can be captured precisely. Moreover, ADGFS simultaneously aligns the marginal and conditional probability distributions in the nonnegative matrix factorization framework to narrow the distribution discrepancies between the two different domains, which can adequately transfer knowledge from the source domain to the target domain. Comprehensive experiments on four benchmark datasets can demonstrate that the effectiveness of the proposed approach in cross-domain image classification.
Jing Sun 0012, Zhihui Wang 0001, Wei Wang 0335, Fuming Sun, Zhengming Ding
IEEE Trans. Circuits Syst. Video Technol.1
2021 Domain adaptation with geometrical preservation and distribution alignment
Jing Sun 0012, Zhihui Wang 0001, Wei Wang 0335, Fuming Sun
Neurocomputing1
2021 Sparsely-labeled source assisted domain adaptation
Wei Wang 0335, Shenglun Chen, Yuankai Xiang, Jing Sun 0012, Zhihui Wang 0001, Fuming Sun, Zhengming Ding, Baopu Li
Pattern Recognit.4
2018 Sparse dual graph-regularized NMF for image co-clustering
Jing Sun 0012, Zhihui Wang 0001, Fuming Sun
Neurocomputing1