Fuming Sun

dblp:00/2551 · DBLP profile ↗
← Back
84ranked-venue papers
13as first author
64since 2021 · last 2027
0000-0003-3932-2712ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 40 · 5 first-author · 29 since 2021Graphics, computer vision, multimedia, augmented reality and games · 32 · 8 first-author · 24 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 7 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Computer networks · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2027 Multi-scale fusion diffusion network for salient object detection in optical remote sensing images
Jingyu Wu, Fuming Sun, Mingyu Lu
Expert Syst. Appl.2
2026 Breaking Alignment Barriers: TPS-Driven Semantic Correlation Learning for Alignment-Free RGB-T Salient Object Detection
abstract
Existing RGB-T salient object detection methods predominantly rely on manually aligned and annotated datasets, struggling to handle real-world scenarios with raw, unaligned RGB-T image pairs. In practical applications, due to significant cross-modal disparities such as spatial misalignment, scale variations, and viewpoint shifts, the performance of current methods drastically deteriorates on unaligned datasets. To address this issue, we propose an efficient RGB-T SOD method for real-world unaligned image pairs, termed Thin-Plate Spline-driven Semantic Correlation Learning Network (TPS-SCL). We employ a dual-stream MobileViT as the encoder, combined with efficient Mamba scanning mechanisms, to effectively model correlations between the two modalities while maintaining low parameter counts and computational overhead. To suppress interference from redundant background information during alignment, we design a Semantic Correlation Constraint Module (SCCM) to hierarchically constrain salient features. Furthermore, we introduce a Thin-Plate Spline Alignment Module (TPSAM) to mitigate spatial discrepancies between modalities. Additionally, a Cross-Modal Correlation Module (CMCM) is incorporated to fully explore and integrate inter-modal dependencies, enhancing detection performance. Extensive experiments on various datasets demonstrate that TPS-SCL attains state-of-the-art (SOTA) performance among existing lightweight SOD methods and outperforms mainstream RGB-T SOD approaches.
Lupiao Hu, Fasheng Wang, Fangmei Chen, Fuming Sun
AAAI4
2026 MSINet: Mantis shrimp vision-inspired network for camouflaged object detection
Houjie Li, Jing Sun 0012, Fuming Sun
Expert Syst. Appl.5
2026 Wild Animal Tracking with High-Quality Segment Anything Model and Domain Adaptation
Ganggang Huang, Fasheng Wang, Hanwei Li, Mingshu Zhang, Mengyin Wang, Fuming Sun
Int. J. Comput. Vis.7
2026 HGLFFNet: Hierarchical global-local feature fusion network for facial expression recognition
Houjie Li, Fuming Sun, Mengyin Wang
Neurocomputing5
2026 Modality Interaction Decoupling: Spatiotemporal Memory-Driven Unified Multimodal Tracking Framework
abstract
Current designs of multimodal tracking networks primarily focus on spatial feature interaction within the backbone and lack the exploitation of temporal information. Although some approaches incorporate temporal cues by introducing sequential information from adjacent frames or employ updated temporal features during the feature extraction stage, they struggle to capture dynamic object variations and motion information in complex scenarios. To address these limitations, this paper proposes a Spatiotemporal Memory-Driven Unified Multimodal Tracking Framework (SMMTrack). Unlike existing works relying on feature interaction paradigms within the backbone network, this study innovatively introduces a decoupled backbone feature extraction framework. It deploys parallel, independent Vision Transformer (ViT) networks dedicated to extracting information from RGB and X modalities (RGB-T, RGB-D, RGB-E), abandoning the conventional intra-backbone feature interaction. Furthermore, we introduce a memory mechanism during the feature extraction stage to enable long-term object modeling. In addition, a Long-term Memory storage and retrieval module is designed to dynamically update the Memory-list, thereby allowing the model to capture object appearance variations and motion trends comprehensively. SMMTrack is a unified framework across three tasks (RGB-T, RGB-D, and RGB-E tracking). Experimental results demonstrate that SMMTrack outperforms the state-of-the-art (SOTA) models, achieving outstanding performance in diverse multimodal tracking scenarios. Codes and results are released on https://github.com/qfxb/SMMTrack.
Fasheng Wang, Mengyin Wang, Fuming Sun
IEEE Internet Things J.5
2026 Multi-modal cooperative fusion network for dual-stream RGB-D salient object detection
Jingyu Wu, Fuming Sun, Mingyu Lu
Image Vis. Comput.2
2026 Infrared and visible image fusion based on multi-modal and multi-scale cross-compensation
Meitian Li, Jing Sun 0012, Fasheng Wang, Fuming Sun
Knowl. Based Syst.5
2026 Towards a prompt-driven framework with state space models for UAV object tracking
Ziqing Yan, Fasheng Wang, Fuming Sun
Knowl. Based Syst.5
2026 LESOD: Lightweight and efficient network for RGB-D salient object detection
Mingyu Zhong, Jing Sun 0012, Fasheng Wang, Fuming Sun
Pattern Recognit.4
2026 CERNet: Real-time stereo matching via collaborative enhancement and refinement
Zhisheng Zhu, Mengyin Wang, Fuming Sun
Pattern Recognit.3
2026 DFD-Stereo: Dual-Domain Feature Decoupling for Stereo Matching
abstract
Accurate stereo matching under limited computational resources remains a central challenge in 3D perception tasks such as autonomous driving and robot navigation. Existing high-accuracy methods often rely on heavy architectures with significant memory and processing demands, while lightweight models typically compromise on feature expressiveness, leading to limited global understanding and detail loss. To bridge this gap, we propose DFD-Stereo, a lightweight and efficient stereo matching framework that delivers high-quality disparity estimation with reduced computational cost. The framework incorporates two key components: (1) a Decoupled Frequency-Spatial Learning (DFSL) module, which enables complementary spatial-frequency representation for enhanced global context modeling, and (2) a Stepwise Coupling Disparity Refinement (SCDR) module, which leverages multi-scale RGB-disparity fusion with Shuffle Attention to refine disparity predictions effectively. Experimental results across multiple benchmarks demonstrate that DFD-Stereo achieves superior accuracy with significantly improved efficiency, offering a promising solution for deployment in resource-constrained 3D vision systems.
Zhisheng Zhu, Fuming Sun, Jing Sun 0012
IEEE Trans. Circuits Syst. Video Technol.2
2026 Depth-Assisted Camouflaged Object Segmentation via Frequency-Domain Fusion and High-Order Interaction
abstract
Recently, some studies have introduced depth cues to solve camouflaged object segmentation (COS) tasks and significantly improve segmentation performance. However, current methods still have two limitations: 1) they are confined to first-order or second-order interaction modeling; 2) they neglect the frequency-domain complementary characteristics between RGB and depth modalities. In this work, we propose a Frequency-domain Fusion and High-order Interaction framework, named F$^{2}$HI, to alleviate the above limitations. Specifically, F$^{2}$HI consists of two key components: Frequency Domain Interaction Fusion (FDIF) and the High-order Interaction Unit (HOIU). The FDIF decouples the frequency information of RGB and depth modalities into high-frequency and low-frequency components, utilizing the dominant frequency components of one modality to enhance the weaker frequency components of the other modality, thereby achieving complementary enhancement in the frequency domain. Moreover, it employs cross-attention to facilitate global multimodal interaction. The HOIU employs cascaded self-attention to parse cross-region features generated by multi-object camouflage scenarios, thereby achieving high-order feature interaction. Compared with 27 state-of-the-art COS methods, F$^{2}$HI achieves competitive performance on four mainstream COS benchmarks.
Fuming Sun, Tian Bai 0002
IEEE Trans. Multim.3
2026 Cross-Modal Fusion With Mixture-of-Experts for Efficient RGB-D Salient Object Detection
abstract
Current RGB-D salient object detection (SOD) models are plagued by issues, including excessive model parameters and high computational complexity. These drawbacks impede the model's efficient deployment and constrain the enhancement of model performance. This paper introduces an efficient and lightweight cross-modal feature cross-fusion network, termed CMFNet. In particular, we design an efficient model utilizing MobileViT as the dual-stream backbone network, thereby significantly reducing computational complexity while maintaining robust feature extraction capabilities. Firstly, we propose a Cross-Fusion Module (CFM) designed to integrate multi-scale semantic information from RGB and depth features dynamically. Furthermore, we design a Lightweight-Mixture-of-Experts Module (L-MoE) for multi-modal features, which enhances the representational capacity of fused features at various levels by employing a dynamic routing mechanism and balanced constraint strategy to allocate appropriate expert processing units to features at different scales. Additionally, we design a Multi-scale Feature Refinement Module (MSFR) that captures multi-scale context through a combination of channel-spatial attention mechanisms and depthwise separable convolution, and gradually eliminates feature distribution differences between modalities using a dual-path residual learning strategy. Abundant experimental findings verify that the proposed CMFNet outperforms the 25 existing State-of-the-art (SOTA) methods.
Jingyu Wu, Fuming Sun, Mingyu Lu
IEEE Trans. Multim.2
2025 Seeing the Unseen: A Semantic Alignment and Context-Aware Prompt Framework for Open-Vocabulary Camouflaged Object Segmentation
Tian Bai 0002, Jing Sun 0012, Fuming Sun
ICCV4
2025 Mirror Feature-Aware Generative Adversarial Network for RGB-T Salient Object Detection
abstract
Existing RGB-T salient object detection methods often employ asymmetric feature processing mechanisms, which lead to limited inter-modal interaction and difficulties in achieving cross-modal semantic alignment. Moreover, due to the inherent differences in imaging mechanisms, modality conflicts caused by feature prior mismatches severely restrict the efficient exploitation of complementary information. To address these challenges, we propose a Mirror Feature-aware Generative Adversarial Network (MFAGAN). We introduce adversarial learning into the multi-modal feature fusion process, transforming the implicit feature alignment assumptions in traditional fusion methods into explicit distribution consistency constraints through the dynamic game mechanism between the generator and discriminator. Specifically, MFAGAN designs a triple collaborative optimization component: 1) A symmetric two-stage encoder achieves a dynamic balance between pixel-level details and semantic-level representations through bidirectional alternating guidance; 2) A cross-modal residual decoder employs independent parameter paths to preserve modality-specific characteristics and suppress fusion bias; 3) A feature difference complementation module adaptively integrates differential information from regions with confidence conflicts. We conduct extensive experiments on three public datasets. The experimental results show that the MFAGAN achieves better performance than the competing methods. Codes and results are released on https://github.com/asd291614761/MFAGAN.
Fangkai Zhao, Fangmei Chen, Fasheng Wang, Fuming Sun
ICIP5
2025 ST-SAM: SAM-Driven Self-Training Framework for Semi-Supervised Camouflaged Object Detection
abstract
Semi-supervised Camouflaged Object Detection (SSCOD) aims to reduce reliance on costly pixel-level annotations by leveraging limited annotated data and abundant unlabeled data. However, existing SSCOD methods based on Teacher-Student frameworks suffer from severe prediction bias and error propagation under scarce supervision, while their multi-network architectures incur high computational overhead and limited scalability. To overcome these limitations, we propose ST-SAM, a highly annotation-efficient yet concise framework that breaks away from conventional SSCOD constraints. Specifically, ST-SAM employs Self-Training strategy that dynamically filters and expands high-confidence pseudo-labels to enhance a single-model architecture, thereby fundamentally circumventing inter-model prediction bias. Furthermore, by transforming pseudo-labels into hybrid prompts containing domain-specific knowledge, ST-SAM effectively harnesses the Segment Anything Model's potential for specialized tasks to mitigate error accumulation in self-training. Experiments on COD benchmark datasets demonstrate that ST-SAM achieves state-of-the-art performance with only 1% labeled data, outperforming existing SSCOD methods and even matching fully supervised methods. Remarkably, ST-SAM requires training only a single network, without relying on specific models or loss functions. This work establishes a new paradigm for annotation-efficient SSCOD. Codes will be available at https://github.com/hu-xh/ST-SAM.
Xihang Hu, Fuming Sun, Jiazhe Liu, Feilong Xu, Xiaoli Zhang 0001
ACM Multimedia2
2025 LIESA: Low-Light Image Enhancement with Semantic Awareness
Shijie Hao, Fuming Sun
MMM (2)3
2025 Beyond Whole Dialogue Modeling: Contextual Disentanglement for Conversational Recommendation
abstract
Conversational recommender systems aim to provide personalized recommendations by analyzing and utilizing contextual information related to dialogue. However, existing methods typically model the dialogue context as a whole, neglecting the inherent complexity and entanglement within the dialogue. Specifically, a dialogue comprises both focus information and background information, which mutually influence each other. Current methods tend to model these two types of information mixedly, leading to misinterpretation of users' actual needs, thereby lowering the accuracy of recommendations. To address this issue, this paper proposes a novel model to introduce contextual disentanglement for improving conversational recommender systems, named DisenCRS. The proposed model DisenCRS employs a dual disentanglement framework, including self-supervised contrastive disentanglement and counterfactual inference disentanglement, to effectively distinguish focus information and background information from the dialogue context under unsupervised conditions. Moreover, we design an adaptive prompt learning module to automatically select the most suitable prompt based on the specific dialogue context, fully leveraging the power of large language models. Experimental results on two widely used public datasets demonstrate that DisenCRS significantly outperforms existing conversational recommendation models, achieving superior performance on both item recommendation and response generation tasks.
Guojia An, Jie Zou 0001, Jiwei Wei, Chaoning Zhang, Fuming Sun, Yang Yang 0002
SIGIR5
2025 Facial attribute editing via a Balanced Simple Attention Generative Adversarial Network
Fanghui Ren, Wenpeng Liu, Fasheng Wang, Bo Wang 0167, Fuming Sun
Expert Syst. Appl.5
2025 SICFNet: Shared Information Interaction and Complementary Feature Fusion Network for RGB-T traffic scene parsing
Bing Zhang 0023, Zhenlong Li, Fuming Sun, Zhanghe Li, Xiaobo Dong, Xiaolan Zhao
Expert Syst. Appl.3
2025 ESNet: An Efficient Skeleton-guided Network for camouflaged object detection
Tian Bai 0002, Fuming Sun
Knowl. Based Syst.3
2025 Bio-inspired two-stage network for efficient RGB-D salient object detection
Tian Bai 0002, Fuming Sun
Neural Networks3
2025 Highly Efficient RGB-D Salient Object Detection With Adaptive Fusion and Attention Regulation
abstract
Existing RGB-D salient object detection (SOD) models have large numbers of parameters, high computational complexity, and slow inference speeds, limiting their deployment on edge devices. To address this issue, we propose a highly efficient network (HENet), focusing on developing lightweight RGB-D SOD models. Specifically, to fairly handle multimodal inputs and capture long-range dependencies of features, we employ a dual-stream structure and use MobileViT as the network encoder. We introduce the Adaptive Edge-Aware Fusion Module (AEFM) that adaptively adjusts the contribution of features during the fusion process based on the amount of feature information, and perceives the edges of the fused features at the pixel level. To compensate for the insufficient feature extraction capability of the lightweight backbone network, we propose the Dual-Branch Feature Enhancement Module (DFEM) to enhance the representation capability of the fused features. Finally, we design the Feature Attention Regulation Module (FARM) to adjust the model’s focus in real time. HENet has fewer parameters (11.9M) and lower computational complexity (10.7 GFLOPs), achieving an inference speed of 121 FPS for images with size$384\times 384$. Extensive experiments are conducted on seven challenging RGB-D SOD datasets. The experimental results demonstrate that HENet outperforms 16 state-of-the-art methods and shows great potential in downstream computer vision tasks. Codes and results are available onhttps://github.com/BojueGao/HENet.
Fasheng Wang, Mengyin Wang, Fuming Sun
IEEE Trans. Circuits Syst. Video Technol.4
2025 Rethinking How to Capture Long-Range Dependency in 3D Object Detection
abstract
LiDAR-based 3D object detection is essential for autonomous driving. Existing high-performance 3D object detectors usually design complex structures in the 3D backbone to capture long-range dependencies among features. However, introducing these complex structures into the 3D backbone significantly increases computational cost and inference latency, limiting the efficiency and feasibility of detectors in practical applications. In this work, we rethink the long-range dependency capturing problem from a new perspective, that is transferring this task from 3D backbone to 2D feature space. To accomplish this goal, we propose a Long-Range Dense Feature Capture Network (LDFCNet). LDFCNet retains the basic structure of the 3D backbone to extract preliminary 3D features but shifts the complex long-range dependency capturing task to be processed on a 2D dense feature map, thereby enhancing the detection performance while reducing the computational cost. Importantly, a robust 2D dense feature capture (2D-DFC) backbone is devised to effectively and efficiently capture the long-range dependencies. In addition, we introduce a re-parameterization technique to decouple the training and inference of the 2D backbone, further reducing inference latency. We conduct extensive experiments on the Waymo Open and nuScenes datasets and the experimental results show that LDFCNet demonstrates competitive performance. Notably, LDFCNet is$1.5\times $faster than the state-of-the-art hybrid detector HEDNet and$2.1\times $faster than the transformer-based detector DSVT. Codes and results are released onhttps://github.com/asd291614761/LDFCNet.
Fasheng Wang, Mengyin Wang, Fuming Sun
IEEE Trans. Circuits Syst. Video Technol.4
2025 Exploring a Lightweight and Efficient Network for Salient Object Detection in ORSI
abstract
In recent years, Optical Remote Sensing Image Salient Object Detection (ORSI-SOD) has made substantial progress. Nevertheless, it remains an open-ended research area with complex challenges. Most existing ORSI-SOD methods, aiming for high-performance detection, demand large-scale parameters and high computational costs. This significantly restricts their application on resource-constrained devices, which have limited computing power and memory capacity. To tackle this issue, we propose a lightweight and highly efficient ORSI-SOD network, termed RAMENet. With only 5.18M parameters and 8.72G FLOPs, RAMENet can achieve competitive detection accuracy compared to state-of-the-art methods. Specifically, we devise a Dynamic Region-aware Block (DRB) that can be nested within the encoder to realize plug-and-play functionality. This enables the network to learn ORSI domain-specific feature representations, thus more effectively locating salient object regions. Furthermore, we present a novel Multi-path Enhanced M-shaped Decoder (MED), which integrates both bottom-up and top-down paradigms. Comprising two feature extraction sub-branches and a master feature refinement branch, this architecture achieves multi-granularity feature aggregation via cross-level feature interaction. Consequently, it significantly improves the detailed representation capability while maintaining the integrity of the object structure. Extensive experimental results indicate that the RAMENet outperforms 5 state-of-the-art lightweight methods in terms ofSα,Fβmean,MAEon EORSSD and ORSSD datasets, with improvement reaching 0.68%, 0.92%, 0.13%, 0.60%, 1.13%, and 0.07%, respectively. The code and results are available at https://github.com/hjy0518/RAMENet/.
Jinyu Han, Fuming Sun, Yaoyao Hou, Jing Sun 0012
IEEE Trans. Geosci. Remote. Sens.2
2025 ORSIDiff: Diffusion Model for Salient Object Detection in Optical Remote Sensing Images
abstract
The unique imaging conditions of satellites introduce significant uncertainties in the structure and scale of ground objects, presenting a major challenge for Optical Remote Sensing Image Salient Object Detection (ORSI-SOD). Current ORSI-SOD methods often fail to effectively differentiate between salient objects and subtle background variations, leading to suboptimal prediction outcomes. Furthermore, ORSI-SOD is a dense pixel prediction task, and existing approaches frequently depend on pixel-level probabilities, which can result in overconfident and inaccurate predictions. To address these challenges, we reformulate the ORSI-SOD task as a mask-generation problem by introducing a novel paradigm and propose a diffusion model-based method for ORSI-SOD, termed ORSIDiff. Central to our approach is the design of a powerful denoising network that enhances the model’s refinement capabilities. This network leverages the strengths of both global and local modeling, improving the handling of salient object details and enabling a deeper understanding of the distinctions between salient objects and their surroundings. Additionally, we introduce a consistency assessment strategy that aggregates multiple potential predictions during the denoising process, effectively mitigating the issue of overconfident point estimation. Extensive experimental results on two widely used ORSI-SOD datasets demonstrate that ORSIDiff achieves significant performance improvements over 20 state-of-the-art methods.
Jinyu Han, Jing Sun 0012, Fasheng Wang, Fuming Sun
IEEE Trans. Geosci. Remote. Sens.4
2025 DINet: Depth-Guided and Iterative Refinement Network for Salient Object Detection in Optical Remote Sensing Images
abstract
Optical remote sensing images (ORSI) feature unique scenes and complex imaging conditions. Specifically, they exhibit substantial variations in object scale, quantity, structure, and distribution. Consequently, salient object detection in ORSI (ORSI-SOD) is pivotal in ORSI content perception and understanding. Additionally, the limitations of the single modality impede the advancement of ORSI-SOD. To tackle these issues, we propose a Depth-guided and Iterative Refinement Network (DINet) for ORSI-SOD. By incorporating depth information as auxiliary cues, we introduce a multi-modal strategy for ORSI-SOD, resulting in improved accuracy in the localization and segmentation of salient objects. To address the variability of salient objects, we design an Aggregation Perception Enhancement (APE) Module. This module integrates complementary cues from cross-modal features using multi-dimensional attention mechanisms. By fostering cross-modal interactions, the APE module effectively preserves both detail and spatial location information. Furthermore, we propose an Iterative Guidance Refinement Decoder to handle boundary uncertainty. The decoder uses initial predictions to guide the decoding phase and iteratively refine results. Simultaneously, it minimizes noise from depth cues, yielding predictions with more accurate boundaries. Experimental comparisons with 22 state-of-the-art methods show that DINet exhibits superior performance while maintaining lightweight (11.98M) and real-time (55FPS) capabilities.
Xihang Hu, Fuming Sun, Xiaoli Zhang 0001, Chuanmin Jia, Siwei Ma 0001
IEEE Trans. Geosci. Remote. Sens.2
2025 Lightweight Edge-Aware Mamba-Fusion Network for Weakly Supervised Salient Object Detection in Optical Remote Sensing Images
abstract
Despite the significant progress made in fully supervised salient object detection in optical remote sensing images (ORSI-SOD), these methods rely heavily on pixel-level annotations, which are time-consuming and labor-intensive. This situation has driven the development of weakly supervised ORSI-SOD methods. However, existing weakly supervised ORSI-SOD methods still face excessive model parameters and high computational complexity, hindering their flexibility and deployment in edge devices. To address these challenges, we propose the LightEMNet, a scribble-based, lightweight, and high-performance edge-aware network for ORSI-SOD. The network employs MobileNetV2 as its lightweight encoder backbone. To mitigate the suboptimal feature extraction performance caused by the lightweight architecture, we design a feature refinement layer (FRL) to refine the features extracted from the backbone, thereby generating guidance information while achieving better structural awareness and object localization. To realize better detail optimization, we introduce edge information extracted by a multiscale edge perception module (MEP) to regulate high-level features. Finally, considering the shortcomings of traditional convolution in global-awareness, we propose a Mamba-based cross-scale edge-semantic interaction (CESI) module to achieve efficient alignment of semantics and edges, which consequently enhances the representation consistency of the fused features and improves the model’s adaptability to complex scenes. We verify the effectiveness of the LightEMNet through extensive experiments. The results demonstrate that the proposed LightEMNet exhibits competitive detection performance with only 4.81 M parameters. Codes and results are available athttps://github.com/xingggao/LightEMNet
Gaojie Xing, Mengyin Wang, Fasheng Wang, Fuming Sun
IEEE Trans. Geosci. Remote. Sens.4
2025 A UNet-Like Transformer Network for Camouflaged Object Detection
abstract
The role of Camouflaged Object Detection (COD) is to identify the objects that integrate seamlessly with the surrounding environment. Due to the high intrinsic similarity between the objects and their background, this task presents greater challenges than traditional object detection. Most existing COD methods often have a large number of parameters and high computational complexity in the pursuit of detection accuracy, which hinders the application of COD in practical scenarios. To address this issue, we propose a UNet-like Transformer Network for COD, termed UTNet, which achieves competitive detection accuracy with a smaller parameter set. Specifically, we propose a Camouflaged Region Awareness Module (CRAM) consisting of a Hierarchical Attention Mechanism (HAM) that groups features to reveal intrinsic consistency between sub-features. This CRAM can be embedded into the backbone network, giving it powerful modeling capabilities. And, we present a Contextual Knowledge Collector (CKC) that exploits a cross-aggregation approach for neighboring feature layers, promoting the flow of semantic information from high-level to low-level features, and ensuring the integrity of camouflaged objects at each level of features. Furthermore, we introduce a progressive decoder that utilizes a cascade of attention units to filter noise and explores knowledge aggregation to emphasize features from different levels, ensuring that camouflaged objects have complete spatial details at the local level. Extensive experimental results show that UTNet achieves competitive results compared to 20 state-of-the-art methods. Codes and results are released onhttps://github.com/hjy0518/UTNet.
Fuming Sun, Jinyu Han, Weiyi Wu, Jing Sun 0012, Mengyin Wang
IEEE Trans. Multim.1
2025 Spatial-Frequency Collaborative Learning for Camouflaged Object Detection
abstract
Camouflaged object detection (COD) is a challenging task that struggles to accurately detect the objects concealed in the surrounding environment. This is largely attributed to the intrinsic similarity of the camouflaged objects with the surrounding environment. To address this challenge, we propose a Spatial-Frequency Collaborative Learning network for COD (SFCNet). Specifically, we propose a Domain Transformation Fusion (DTF) module to handle the similarity between the camouflaged objects and the background, because when processed in the frequency domain, the features of the camouflaged object and the background become easy to discriminate. Then, we design a Cross-domain Integration Unit (CIU) to integrate the high-level features progressively through a Spatial-Frequency Coordinated Fusion (SFCF) module and a Multi-scale Feature Enhancement (MFE) module. Finally, the low-level features are combined with the high-level features from different decoding stages to correct the camouflaged objects in detail. In addition, an Edge Amplification (EA) module is designed to enable the model to pay attention to the global contour of the camouflaged object. It can facilitate the generation of prediction maps with accurate object boundaries. Extensive experiments on four benchmark COD datasets show that SFCNet outperforms state-of-the-art (SOTA) COD models. Meanwhile, it also has the characteristics of low parameters (21.01 M) and low computational complexity (24.14 G). Codes and results are released onhttps://github.com/Zhaorui328/SFCNet.
Mengyin Wang, Fasheng Wang, Fuming Sun
IEEE Trans. Multim.4
2024 OmniStyleGAN for Style-Guided Image-to-Image Translation
Qianyi Zhao, Mengyin Wang, Fasheng Wang, Fuming Sun
PRCV (11)5
2024 Behavior-Contextualized Item Preference Modeling for Multi-Behavior Recommendation
abstract
In recommender systems, multi-behavior methods have demonstrated their effectiveness in mitigating issues like data sparsity, a common challenge in traditional single-behavior recommendation approaches. These methods typically infer user preferences from various auxiliary behaviors and apply them to the target behavior for recommendations. However, this direct transfer can introduce noise to the target behavior in recommendation, due to variations in user attention across different behaviors. To address this issue, this paper introduces a novel approach, Behavior-Contextualized Item Preference Modeling (BCIPM), for multi-behavior recommendation. Our proposed Behavior-Contextualized Item Preference Network discerns and learns users' specific item preferences within each behavior. It then considers only those preferences relevant to the target behavior for final recommendations, significantly reducing noise from auxiliary behaviors. These auxiliary behaviors are utilized solely for training the network parameters, thereby refining the learning process without compromising the accuracy of the target behavior recommendations. To further enhance the effectiveness of BCIPM, we adopt a strategy of pre-training the initial embeddings. This step is crucial for enriching the item-aware preferences, particularly in scenarios where data related to the target behavior is sparse. Comprehensive experiments conducted on four real-world datasets demonstrate BCIPM's superior performance compared to several leading state-of-the-art models, validating the robustness and efficiency of our proposed approach.
Mingshi Yan, Fan Liu 0008, Jing Sun 0012, Fuming Sun, Zhiyong Cheng 0001, Yahong Han
SIGIR4
2024 Unsupervised image-to-image translation with multiscale attention generative adversarial network
Fasheng Wang, Qianyi Zhao, Mengyin Wang, Fuming Sun
Appl. Intell.5
2024 OR2Net: Online Re-weighting Relation Network for kinship verification
abstract
Kinship verification aims to infer whether there is a kin relation between different individuals from facial images . However, popular kinship datasets are often small and suffer from data imbalance. Most existing methods build complex networks to extract features but ignore some implicit information, like family information. They use balanced datasets with fixed negative samples for training, which overlooks valuable information from multiple negative samples, leading to poor performance and robustness. To address these issues, we propose a novel end-to-end framework for kinship verification called Online Re-weighting Relation Network (OR 2 Net) based on an online re-weighting strategy of meta-learning and relation network. Our novel relation network aims to extract fine-grained features and reduce differences between generations by using multi-scale features and mining family information from kinship datasets. Additionally, we design a lightweight meta re-weighting network that uses a small, clean meta-set to guide the adaptive weighing of training examples. This is done using one-step stochastic gradient descent (SGD) based on an online re-weighting strategy from meta-learning. This helps find effective hard negative samples and reduces the imbalance problem. Extensive experiments on three public kinship verification datasets show that our proposed method is more effective compared to state-of-the-art methods. The code is publicly available on https://github.com/XinZhao-dlnu/OR2N .
Houjie Li, Mengyin Wang, Haiyu Song 0002, Fuming Sun
Expert Syst. Appl.5
2024 Image deblurring method based on self-attention and residual wavelet transform
Bing Zhang 0023, Jing Sun 0012, Fuming Sun, Fasheng Wang
Expert Syst. Appl.3
2024 Cross-Modal Fusion and Progressive Decoding Network for RGB-D Salient Object Detection
Xihang Hu, Fuming Sun, Jing Sun 0012, Fasheng Wang
Int. J. Comput. Vis.2
2024 Enhancing Collaborative Information with Contrastive Learning for Session-based Recommendation
Guojia An, Jing Sun 0012, Fuming Sun
Inf. Process. Manag.4
2024 MAGNet: Multi-scale Awareness and Global fusion Network for RGB-D salient object detection
Mingyu Zhong, Jing Sun 0012, Fasheng Wang, Fuming Sun
Knowl. Based Syst.5
2024 MadFormer: multi-attention-driven image super-resolution method based on Transformer
Jing Sun 0012, Ting Li 0002, Fuming Sun
Multim. Syst.5
2024 Unsupervised image style transformation of generative adversarial networks based on cyclic consistency
Jingyu Wu, Fuming Sun, Mingyu Lu
Multim. Syst.2
2024 Siamese Tracking Network with Multi-attention Mechanism
abstract
Object trackers based on Siamese networks view tracking as a similarity-matching process. However, the correlation operation operates as a local linear matching process, limiting the tracker’s ability to capture the intricate nonlinear relationship between the template and search region branches. Moreover, most trackers don’t update the template and often use the first frame of an image as the initial template, which will easily lead to poor tracking performance of the algorithm when facing instances of deformation, scale variation, and occlusion of the tracking target. To this end, we propose a Simases tracking network with a multi-attention mechanism, including a template branch and a search branch. To adapt to changes in target appearance, we integrate dynamic templates and multi-attention mechanisms in the template branch to obtain more effective feature representation by fusing the features of initial templates and dynamic templates. To enhance the robustness of the tracking model, we utilize a multi-attention mechanism in the search branch that shares weights with the template branch to obtain multi-scale feature representation by fusing search region features at different scales. In addition, we design a lightweight and simple feature fusion mechanism, in which the Transformer encoder structure is utilized to fuse the information of the template area and search area, and the dynamic template is updated online based on confidence. Experimental results on publicly tracking datasets show that the proposed method achieves competitive results compared to several state-of-the-art trackers.
Yuzhuo Xu, Ting Li 0002, Fasheng Wang, Fuming Sun
Neural Process. Lett.5
2024 Efficient Camouflaged Object Detection Network Based on Global Localization Perception and Local Guidance Refinement
abstract
Camouflaged Object Detection (COD) is a challenging visual task due to its complex contour, diverse scales, and high similarity to the background. Existing COD methods encounter two predicaments: One is that they are prone to falling into local perception, resulting in inaccurate object localization; Another issue is the difficulty in achieving precise object segmentation due to a lack of detailed information. In addition, most COD methods typically require larger parameter amounts and higher computational complexity in pursuit of better performance. To this end, we propose a global localization perception and local guidance refinement network (PRNet), that simultaneously addresses performance and computational costs. Through effective aggregation and use of semantic and details information, the PRNet can achieve accurate localization and refined segmentation of camouflaged objects. Specifically, with the help of a Cascaded Attention Perceptron (CAP) designed, we can effectively integrate and perceive multi-scale information to localize camouflaged objects. We also design a Guided Refinement Decoder (GRD) in a top-down manner to extract context information and aggregate details to further refine camouflaged prediction results. Extensive experimental results demonstrate that our PRNet outperforms 12 state-of-the-art models on 4 challenging datasets. Meanwhile, the PRNet has a smaller number of parameters (12.74M), lower computational complexity (10.24G), and real-time inference speed (105FPS). Source codes are available at https://github.com/hu-xh/PRNet.
Xihang Hu, Xiaoli Zhang 0001, Fasheng Wang, Jing Sun 0012, Fuming Sun
IEEE Trans. Circuits Syst. Video Technol.5
2024 Temporal Context and Environment-Aware Correlation Filter for UAV Object Tracking
abstract
In this article, we propose a temporal context and environment-aware correlation filter (CF) for unmanned aerial vehicle (UAV) object tracking. First, we exploit environmental residuals between two adjacent frames to enhance the discrimination ability and insensitivity of the tracker in complex tracking environments. Second, the historical filter model is utilized to build a temporal regularization term to prevent the filter degradation induced by the severe object appearance variations that affect the tracking performance while suppressing the boundary effect. Finally, we propose to use the predicted object position in the next frame to efficiently sample a predicted context patch for constructing the prediction context-aware regularization term, improving the ability of the tracker to cope with interference from unknown environment changes. Extensive experimental results on four challenging benchmarks show that our tracker achieves state-of-the-art (SOTA) performance compared with other advanced CF-based UAV trackers with real-time tracking speed (47 frames/s) on a single CPU.
Fasheng Wang, Fuming Sun
IEEE Trans. Geosci. Remote. Sens.5
2024 Low-light image enhancement using transformer with color fusion and channel attention
Yinbang Sun, Jing Sun 0012, Fuming Sun, Fasheng Wang
J. Supercomput.3
2024 CATNet: A Cascaded and Aggregated Transformer Network for RGB-D Salient Object Detection
abstract
Salient object detection (SOD) is an important preprocessing operation for various computer vision tasks. Most of existing RGB-D SOD models employ additive or connected strategies to directly aggregate and decode multi-scale features to predict salient maps. However, due to the large differences between the features of different scales, these aggregation strategies adopted may lead to information loss or redundancy, and few methods explicitly consider how to establish connections between features at different scales in the decoding process, which consequently deteriorates the detection performance of the models. To this end, we propose a cascaded and aggregated Transformer Network (CATNet) which consists of three key modules, i.e., attention feature enhancement module (AFEM), cross-modal fusion module (CMFM) and cascaded correction decoder (CCD). Specifically, the AFEM is designed on the basis of atrous spatial pyramid pooling to obtain multi-scale semantic information and global context information in high-level features through dilated convolution and multi-head self-attention mechanism, enhancing high-level features. The role of the CMFM is to enhance and thereafter fuse the RGB features and depth features, alleviating the problem of poor-quality depth maps. The CCD is composed of two subdecoders in a cascading fashion. It is designed to suppress noise in low-level features and mitigate the differences between features at different scales. Moreover, the CCD uses a feedback mechanism to correct and repair the output of the subdecoder by exploiting supervised features, so that the problem of information loss caused by the upsampling operation during the multi-scale features aggregation process can be mitigated. Extensive experimental results demonstrate that the proposed CATNet achieves superior performance over 14 state-of-the-art RGB-D methods on 7 challenging benchmarks.
Fuming Sun, Bowen Yin, Fasheng Wang
IEEE Trans. Multim.1
2024 Cascading Residual Graph Convolutional Network for Multi-Behavior Recommendation
abstract
Multi-behavior recommendation exploits multiple types of user-item interactions, such as view and cart , to learn user preferences and has demonstrated to be an effective solution to alleviate the data sparsity problem faced by the traditional models that often utilize only one type of interaction for recommendation. In real scenarios, users often take a sequence of actions to interact with an item, in order to get more information about the item and thus accurately evaluate whether an item fits their personal preferences. Those interaction behaviors often obey a certain order, and more importantly, different behaviors reveal different information or aspects of user preferences towards the target item. Most existing multi-behavior recommendation methods take the strategy to first extract information from different behaviors separately and then fuse them for final prediction. However, they have not exploited the connections between different behaviors to learn user preferences. Besides, they often introduce complex model structures and more parameters to model multiple behaviors, largely increasing the space and time complexity. In this work, we propose a lightweight multi-behavior recommendation model named Cascading Residual Graph Convolutional Network ( CRGCN for short) for multi-behavior recommendation, which can explicitly exploit the connections between different behaviors into the embedding learning process without introducing any additional parameters (with comparison to the single-behavior based recommendation model). In particular, we design a cascading residual graph convolutional network (GCN) structure, which enables our model to learn user preferences by continuously refining the embeddings across different types of behaviors. The multi-task learning method is adopted to jointly optimize our model based on different behaviors. Extensive experimental results on three real-world benchmark datasets show that CRGCN can substantially outperform the state-of-the-art methods, achieving 24.76%, 27.28%, and 25.10% relative gains on average in terms of HR@K (K = {10,20,50,80}) over the best baseline across the three datasets. Further studies also analyze the effects of leveraging multi-behaviors in different numbers and orders on the final performance.
Mingshi Yan, Zhiyong Cheng 0001, Chen Gao 0001, Jing Sun 0012, Fan Liu 0008, Fuming Sun
ACM Trans. Inf. Syst.6
2024 Attention-guided Multi-modality Interaction Network for RGB-D Salient Object Detection
abstract
The past decade has witnessed great progress in RGB-D salient object detection (SOD). However, there are two bottlenecks that limit its further development. The first one is low-quality depth maps. Most existing methods directly use raw depth maps to perform detection, but low-quality depth images can bring negative impacts to the detection performance. Hence, it is not desirable to utilize depth maps indiscriminately. The other one is how to effectively predict salient maps with clear boundary and complete salient region. To address these problems, an Attention-Guided Multi-Modality Interaction Network (AMINet) is proposed. First, we propose a new quality enhancement strategy for unreliable depth images, named D epth E nhancement M odule ( DEM ). With respect to the second issue, we propose C ross- M odality A ttention M odule ( CMAM ) to rapidly locate salient region. The B oundary- A ware M odule ( BAM ) is designed to utilize high-level feature to guide the low-level feature generation in a top-down way to make up for the dilution of the boundary. To further improve the accuracy, we propose A trous R efined B lock ( ARB ) to adaptively compensate for the shortcoming of atrous convolution. By integrating these interactive modules, features from depth and RGB streams can be refined efficiently, which consequently boosts the detection performance. Experimental results demonstrate the proposed AMINet exceeds state-of-the-art (SOTA) methods on several public RGB-D datasets.
Fasheng Wang, Yiming Su, Jing Sun 0012, Fuming Sun
ACM Trans. Multim. Comput. Commun. Appl.5
2023 DCMNet: Discriminant and cross-modality network for RGB-D salient object detection
Fasheng Wang, Fuming Sun
Expert Syst. Appl.3
2023 WATB: Wild Animal Tracking Benchmark
Fasheng Wang, Fuming Sun
Int. J. Comput. Vis.6
2023 Attention-guided graph convolutional network for multi-behavior recommendation
Xingchen Peng, Jing Sun 0012, Mingshi Yan, Fuming Sun, Fasheng Wang
Knowl. Based Syst.4
2023 Style transfer network for complex multi-stroke text
Fangmei Chen, Fasheng Wang, Fuming Sun
Multim. Syst.5
2023 A cross-view geo-localization method guided by relation-aware global attention
Jing Sun 0012, Bing Zhang 0023, Fuming Sun
Multim. Syst.5
2023 Deblurring transformer tracking with conditional cross-attention
Fuming Sun, Fasheng Wang
Multim. Syst.1
2023 SiamADT: Siamese Attention and Deformable Features Fusion Network for Visual Object Tracking
Fasheng Wang, Fuming Sun
Neural Process. Lett.5
2022 Learning Saliency-Aware Correlation Filters for Visual Tracking
abstract
Abstract Recently, visual object tracking has become a hot topic in computer vision community owing to its extensive applications and fundamental research significance. Many researchers have paid attention to the discriminative correlation filter (DCF) due to its excellent tracking performance. The background-aware CF (BACF) is developed to handle the inevitable boundary effects of DCFs that shows superior performance compared with other CF-based trackers. However, BACF has a poor occlusion handling ability. To overcome this defect, we propose a saliency-aware CF tracking (SACF) model. In SACF, a saliency map of the target is introduced on CFs to strengthen the ability to extract the target from a complex background. Meanwhile, to adjust to the rapid variation of the target appearance, we adaptively modify the learning rate of the filter template using adaptive learning rate. By incorporating the saliency map into the BACF objective function and adaptive learning rate adjustment, the robustness and accuracy of the tracker are boosted significantly. Extensive experiments validate that our method can efficaciously solve the occlusion problem and achieves state-of-the-art performance in terms of accuracy and speed on standard benchmarks, i.e. OTB-2013, OTB-2015 and LaSOT. Compared with other state-of-the-art CF tracking methods, the proposed method is competitive considering accuracy and success rate.
Yanbo Wang 0003, Fasheng Wang, Fuming Sun
Comput. J.4
2022 Aggregate interactive learning for RGB-D salient object detection
Jingyu Wu, Fuming Sun, Fasheng Wang
Expert Syst. Appl.2
2022 AMTSet: a benchmark for abrupt motion tracking
Fasheng Wang, Shuangshuang Yin, Fuming Sun, Junxing Zhang
Multim. Tools Appl.5
2022 Context and saliency aware correlation filter for visual tracking
Fasheng Wang, Shuangshuang Yin, Jimmy T. Mbelwa, Fuming Sun
Multim. Tools Appl.4
2022 Stacked Pyramid Attention Network for Object Detection
Shijie Hao, Fuming Sun
Neural Process. Lett.3
2022 Joint Adaptive Dual Graph and Feature Selection for Domain Adaptation
abstract
Domain adaptation aims to exploit domain-invariant features by aligning the cross-domain distributions in the manifold subspace for applying the classifier trained on the source domain to the target domain. However, two limitations may still deteriorate their performances: (1) the influences of noisy or irrelevant features in the original feature space are ignored, which may unexpectedly hurt the classification of target samples; (2) the graph constructed directly in the original data space cannot accurately capture the inherent local manifold structures of high-dimensional data due to the curse of dimensionality, which may seriously mislead the transferable features learning. In this paper, we propose a novel approach to address these problems, referred to as joint Adaptive Dual Graph and Feature Selection for domain adaptation (ADGFS). Specifically, feature selection can characterize the relative importance of different features through a scaling factor, which enables ADGFS to not only reduce the impacts of noisy or irrelevant features on knowledge transfer but also learn informative domain-invariant features. Meanwhile, ADGFS adaptively optimizes the dual graph by learning the similarity matrices of both instance-level and feature-level graphs in the projected low-dimensional manifold subspace rather than the original high-dimensional space, such that the intrinsic local manifold structures of data can be captured precisely. Moreover, ADGFS simultaneously aligns the marginal and conditional probability distributions in the nonnegative matrix factorization framework to narrow the distribution discrepancies between the two different domains, which can adequately transfer knowledge from the source domain to the target domain. Comprehensive experiments on four benchmark datasets can demonstrate that the effectiveness of the proposed approach in cross-domain image classification.
Jing Sun 0012, Zhihui Wang 0001, Wei Wang 0335, Fuming Sun, Zhengming Ding
IEEE Trans. Circuits Syst. Video Technol.5
2021 LEDet: A Single-Shot Real-Time Object Detector Based on Low-Light Image Enhancement
abstract
Abstract Recently, significant breakthroughs have been achieved in the field of object detection. However, existing methods mostly focus on the generic object detection task. Performance degradation can be unavoidable when applying the existing methods to some specific situations directly, e.g. a low-light environment. To address this issue, we propose a single-shot real-time object Detector based on Low-light image Enhancement, namely LEDet. LEDet adapts itself to the low-light detection task in three aspects. First, a low-light enhancement module is introduced as the image preprocessor, producing the augmented inputs from the low-light images. Second, two modules, i.e. low-light and enhanced features fusion module and the scale-aware channel attention dilated convolution module are designed. These two modules aim at learning robust and discriminative features from objects of various sizes hidden in the darkness. In experiments, we validate the effectiveness of each part of our LEDet model via several ablation studies. We also compare LEDet with various methods on the Exclusively Dark dataset, showing that our model achieves the state-of-the-art performance on the balance between speed and accuracy.
Shijie Hao, Fuming Sun
Comput. J.3
2021 Domain adaptation with geometrical preservation and distribution alignment
Jing Sun 0012, Zhihui Wang 0001, Wei Wang 0335, Fuming Sun
Neurocomputing5
2021 Sparsely-labeled source assisted domain adaptation
Wei Wang 0335, Shenglun Chen, Yuankai Xiang, Jing Sun 0012, Zhihui Wang 0001, Fuming Sun, Zhengming Ding, Baopu Li
Pattern Recognit.7
2020 Visual object tracking via iterative ant particle filtering
abstract
Visual object tracking remains a challenging task in computer vision although important progress has been made in the past decades. Particle filter (PF) is now a standard framework for solving non‐linear/non‐Gaussian problems, especially in visual object tracking. This study proposes an ant colony optimisation (ACO)‐based iterative PF for object tracking. In the proposed method, the basic idea of ACO is used to simulate the behaviour of a particle moving toward the posterior distribution. Such idea is incorporated into the particle filtering framework in order to overcome the well‐known particle impoverishment problem. An iterative unscented Kalman filter is used to design a proposal distribution for particle generation in order to generate better predicted sample states. For the likelihood model, the authors adopt the locality sensitive histogram to model the appearance of the target object, which can better handle the illumination variation during tracking. The experimental results demonstrate that the proposed tracker shows better performance than the other tracking methods.
Fasheng Wang, Yanbo Wang 0003, Fuming Sun, Xucheng Li, Junxing Zhang
IET Image Process.4
2018 Parameter Selection for Denoising Algorithms Using NR-IQA with CNN
Jianjun Li 0001, Lanlan Xu, Chin-Chen Chang 0001, Fuming Sun
MMM (1)5
2018 Sparse dual graph-regularized NMF for image co-clustering
Jing Sun 0012, Zhihui Wang 0001, Fuming Sun
Neurocomputing3
2018 Subspace learning by kernel dependence maximization for cross-modal retrieval
Meixiang Xu, Zhenfeng Zhu, Yao Zhao 0001, Fuming Sun
Neurocomputing4
2018 Multiscale Neighborhood Normalization-Based Multiple Dynamic PCA Monitoring Method for Batch Processes With Frequent Operations
abstract
This paper presents a novel multiscale neighborhood normalization-based multiple dynamic principal component analysis (MNN-MDPCA) method to detect the fault in complex batch processes with frequent operations. Since the difference between batches is larger under random frequent operations according to phase, the corresponding monitoring model should be changed accordingly. However, the data quantity is small under a single operation at each phase, the data with similar operations can be clustered together. Due to frequent operations, the data clustered follows non-Gaussian distribution. A normalization strategy called MNN is proposed to complete Gaussian distribution conversion so as to build multivariate statistical model. Subsequently, MDPCA is used to model the multioperation industry processes. Finally, to test the modeling and monitoring performance of the proposed method, a numerical example and the ladle furnace (LF) steelmaking process case are provided, where the comparison with Gaussian mixture model and MDPCA-based results is covered.
Yajun Wang 0003, Fuming Sun, Bo Li 0059
IEEE Trans Autom. Sci. Eng.2
2017 Adaptive NNs Fault-Tolerant Control for Nonstrict-Feedback Nonlinear Systems
Guowei Dong, Yongming Li 0002, Duo Meng, Fuming Sun, Rui Bai 0002
ISNN (2)4
2017 Image multi-label annotation based on supervised nonnegative matrix factorization with new matching measurement
Fuming Sun
Neurocomputing2
2016 Multi-level feature representations for video semantic concept detection
Fuming Sun, Chenxin Liu
Neurocomputing3
2016 Graph regularized and sparse nonnegative matrix factorization with hard constraints for data representation
Fuming Sun, Meixiang Xu, Xuekao Hu, Xiaojun Jiang
Neurocomputing1
2016 Social video annotation by combining features with a tri-adaptation approach
Fuming Sun, Meixiang Xu, Shijie Hao
Multim. Syst.1
2016 Active learning SVM with regularization path for image classification
Fuming Sun
Multim. Tools Appl.1
2015 Robust Multi-label Image Classification with Semi-Supervised Learning and Active Learning
Fuming Sun, Meixiang Xu, Xiaojun Jiang
MMM (2)1
2015 A novel traffic sign detection method via color segmentation and robust shape matching
Fuming Sun
Neurocomputing2
2014 Shape analysis based on feature-preserving Elastic Quadratic Patch Modeling
Fuming Sun, Shijie Hao
Neurocomputing1
2014 Multi-Label Image Categorization With Sparse Factor Representation
abstract
The goal of multilabel classification is to reveal the underlying label correlations to boost the accuracy of classification tasks. Most of the existing multilabel classifiers attempt to exhaustively explore dependency between correlated labels. It increases the risk of involving unnecessary label dependencies, which are detrimental to classification performance. Actually, not all the label correlations are indispensable to multilabel model. Negligible or fragile label correlations cannot be generalized well to the testing data, especially if there exists label correlation discrepancy between training and testing sets. To minimize such negative effect in the multilabel model, we propose to learn a sparse structure of label dependency. The underlying philosophy is that as long as the multilabel dependency cannot be well explained, the principle of parsimony should be applied to the modeling process of the label correlations. The obtained sparse label dependency structure discards the outlying correlations between labels, which makes the learned model more generalizable to future samples. Experiments on real world data sets show the competitive results compared with existing algorithms.
Fuming Sun, Jinhui Tang 0001, Guo-Jun Qi, Thomas S. Huang
IEEE Trans. Image Process.1
2013 Robust Detection and Localization of Human Action in Video
Fuming Sun, Yue Guan 0002
MMM (2)2
2013 Photo 4W: Mobile photo management on what, where, who and when
Fuming Sun, Xueming Wang
Neurocomputing1
2013 Towards tags ranking for social images
Fuming Sun, Yinghai Zhao, Xueming Wang, Dongxia Wang 0005
Neurocomputing1
2012 Optimizing social image search with multiple criteria: Relevance, diversity, and typicality
Fuming Sun, Meng Wang 0001, Dongxia Wang 0005, Xueming Wang
Neurocomputing1
2009 Texture Simulation and Implementation Based on Matlab and Simulink
abstract
In this paper, procedural texture mapping based on Perlin noise is firstly implemented and simulated in Matlab. And then the design is converted from float-point to fix-point in Simulink. Using the system modeling tool, System Generator from Xilinx company, the noise function can be directly mapped into FPGA hardware. From the experimental results, various graphical textures can be implemented in real time and with realistic effects in FPGA hardware.
Fuming Sun
ICIG1