Yun Xiao 0003

dblp:38/1284-3 · DBLP profile ↗
← Back
23ranked-venue papers
9as first author
18since 2021 · last 2026
0000-0002-5285-8565ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 4 first-author · 11 since 2021Artificial intelligence and machine learning · 7 · 5 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 6 since 2021Software engineering, systems software and programming languages · 1
YearPublicationVenuePosition
2026 Unaligned UAV RGBT Tracking: A Largescale Benchmark and a Novel Approach
abstract
With the rapid development of the low-altitude economy, multimodal visual tracking in UAV scenarios has attracted extensive attention. UAVs are typically equipped with independent visible (RGB) and thermal infrared (TIR) sensors, resulting in an inherent spatial misalignment between the two modalities. However, existing RGBT tracking methods generally rely on spatially aligned data inputs, making them unsuitable for unaligned RGBT tracking task in UAV scenarios. In this work, we introduce the new task called unaligned UAV RGBT tracking and construct the first large-scale unaligned RGB and TIR video dataset to promote the research and development of this field. The dataset contains 1,453 pairs of UAV-captured RGBT sequences with precise dual-modal bounding box annotations, and covers 42 object categories, 22 typical challenge attributes, and diverse spatial misalignment scales to better simulate real-world challenging scenarios. To address the limitations of existing methods that fail to handle the spatial misalignment issue in UAV scenarios, we propose the novel RGBT tracking approach. In particular, we design a mixture of shift estimation experts module to adaptively estimate the spatial shifts across two modalities at different scales, and a cross-modal alignment and fusion module to further compensate for nonlinear deformations and integrate multimodal information. Extensive experiments on the created dataset demonstrate that the proposed tracker significantly outperforms existing state-of-the-art tracking methods, validating its practicality and robustness in real-world unaligned UAV tracking scenarios.
Yun Xiao 0003, Jiandong Jin, Wankang Zhang, Chenglong Li 0002
AAAI1
2025 Cross-modulated Attention Transformer for RGBT Tracking
abstract
Existing Transformer-based RGBT trackers achieve remarkable performance benefits by leveraging self-attention to extract uni-modal features and cross-attention to enhance multi-modal feature interaction and search-template correlation. Nevertheless, the independent search-template correlation calculations are prone to be affected by low-quality data, which might result in contradictory and ambiguous correlation weights. It not only limits the intra-modal feature representation, but also harms the robustness of cross-attention for multi-modal feature interaction and search-template correlation computation. To address these issues, we propose a novel approach called Cross-modulated Attention Transformer (CAFormer), which innovatively integrates inter-modality interaction into the search-template correlation computation within typical attention mechanism, for RGBT tracking. In particular, we first independently generate correlation maps for each modality and feed them into the designed correlation modulated enhancement module, which can modify inaccurate correlation weights by seeking the consensus between modalities. Such kind of design unifies self-attention and cross-attention schemes, which not only alleviates inaccurate attention weight computation in self-attention but also eliminates redundant computation introduced by extra cross-attention scheme. In addition, we design a collaborative token elimination strategy to further improve tracking inference efficiency and accuracy. Experiments on five public RGBT tracking benchmarks show the outstanding performance of the proposed CAFormer against state-of-the-art methods.
Yun Xiao 0003, Jiacong Zhao, Andong Lu, Chenglong Li 0002, Yin Lin, Cong Liu 0006
AAAI1
2025 UAV Video Vehicle Detection: Benchmark and Baseline
abstract
With the increasing application of unmanned aerial vehicles (UAVs) in intelligent transportation systems, vehicle object detection in UAV videos has received increasing attention. Precise categorization and detection for vehicles in UAVs is important in many practical applications. However, existing object detection methods, tailored for natural images, often fall short of accurately identifying vehicle objects. Additionally, high-altitude UAV imaging mainly employs horizontal bounding box annotation, frequently leading to significant obstruction and overlapping. Hence, we propose a new task called UAV video vehicle detection (VVD) to achieve precise detection and categorization of vehicles in high-altitude UAV imaging environments. To facilitate the research and development of UAV VVD, we construct the first large-scale well-annotated benchmark UAV VVD dataset, which includes 70 UAV videos captured at a 500-m altitude, with 361489 vehicle instances annotated by the oriented bounding boxes and vehicle categories. Moreover, we introduce a novel category refinement network (CRNet) approach that extracts and refines vehicle object features from the bounding box of the detection results to classify vehicle categories. This approach effectively eliminates the interference of the background and other vehicle objects in candidate boxes. Notably, the vehicle object features are projected into subspace, enabling the category refinement module (CRM) to focus more on the distinctive characteristics of the vehicle object itself through normalization operations. We conduct extensive experiments on the proposed VVD dataset. Experimental results demonstrate the superiority and effectiveness of the proposed CRNet method. The relevant code and dataset are available athttps://github.com/mmic-lcl.
Yun Xiao 0003, Jinfa Wang, Zhicheng Zhao 0002, Bo Jiang 0002, Chenglong Li 0002, Jin Tang 0001
IEEE Trans. Geosci. Remote. Sens.1
2025 Multimodal Remote Sensing Image Registration via Modality Perception and Self-Supervised Position Estimation
abstract
Multi-modal remote sensing images registration ensures that images from different sensors or modalities are spatial and informational consistent for effective comparison and analysis. However, due to the non-linear modality gaps that exist between images, making it difficult to focus only on the spatial position differences of the images and ignore the modality gaps. In this paper, to address this issue, we propose a new framework for Multi-Modal remote sensing image Registration, named MMRNet. The proposed framework comprises the following main aspects. First, a novel self-supervised Positional Misalignment Estimator (PME) is designed for multi-modal image registration. PME is able to efficiently overcome the modality gaps and learn the positional differences between multi-modal images more reliably, optimizing the registration loss by minimizing the positional differences directly. Then, a new paradigm of modality translation, termed Modality Perception Module (MPM), is introduced to effectively learn modality gaps and perform modality translation in the case of positional misalignment. Finally, we further design the modality perception guidance loss to supervise the modality translation task, which can encourage the fidelity of the generated pseudo-modality images. Our registration network integrates both rigid registration model and non-rigid registration model. Experimental results demonstrate that the proposed registration framework can obtain obviously superior performance in both rigid and non-rigid image registration tasks on optical-SAR data, optical-map data and optical-infrared data. The code and relevant dataset will be made publicly available at https://github.com/Ahuer-Lei/MMRNet.
Yun Xiao 0003, Bo Jiang 0002, Yuan Chen 0012, Jin Tang 0001
IEEE Trans. Geosci. Remote. Sens.1
2025 Reflectance-Guided Progressive Feature Alignment Network for All-Day UAV Object Detection
abstract
Object detection using visible-infrared images has become increasingly crucial for all-day applications of unmanned aerial vehicle (UAV). However, existing multi-modal detection methods face significant challenges in low-light conditions, where degraded visible image quality exacerbates weak alignment issues and compromises feature fusion effectiveness. Although recent approaches have attempted to address these issues through cross-attention mechanisms or feature alignment strategies, they often suffer from unstable performance and limited generalization capability in challenging nighttime scenarios. To address these limitations, we propose a novel Reflectance-Guided Progressive Feature Alignment Network (RGFNet) for robust UAV object detection. Our proposed method leverages the illumination-invariant characteristic of reflectance features decomposed from visible images via Retinex theory to guide cross-modal alignment and fusion. Specifically, we design a Reflectance-Guided Collaborative Alignment Module (RCAM) that utilizes reflectance guidance to perform bidirectional feature alignment between visible and infrared modalities, effectively reducing position misalignment under varying lighting conditions. Furthermore, we introduce a Light-Aware Selective Fusion Module (LSFM) that maps multi-modal features into a shared hidden state space through selective state space mechanism, enabling efficient feature interaction while maintaining linear computational complexity. Extensive experiments on two challenging UAV detection benchmarks, DroneVehicle and DVTOD, demonstrate the superiority of our method. RGFNet achieves state-of-the-art performance with 81.4% mAP on DroneVehicle and 88.5% mAP on DVTOD. The code is available at https://github.com/uavdet/RGFNet.
Zhicheng Zhao 0002, Wei Zhang 0393, Yun Xiao 0003, Chenglong Li 0002, Jin Tang 0001
IEEE Trans. Geosci. Remote. Sens.3
2024 Breaking Modality Gap in RGBT Tracking: Coupled Knowledge Distillation
abstract
Modality gap between RGB and thermal infrared (TIR) images is a crucial issue but often overlooked in existing RGBT tracking methods. It can be observed that modality gap mainly lies in the image style difference. In this work, we propose a novel Coupled Knowledge Distillation framework called CKD, which pursues common styles of different modalities to break modality gap, for high performance RGBT tracking. In particular, we introduce two student networks and employ the style distillation loss to make their style features consistent as much as possible. Through alleviating the style difference of two student networks, we can break modality gap of different modalities well. However, the distillation of style features might harm to the content representations of two modalities in student networks. To handle this issue, we take original RGB and TIR networks as the teachers, and distill their content knowledge into two student networks respectively by the style-content orthogonal feature decoupling scheme. We couple the above two distillation processes in an online optimization framework to form new feature representations of RGB and thermal modalities without modality gap. In addition, we design a masked modeling strategy and a multi-modal candidate token elimination strategy into CKD to improve tracking robustness and efficiency respectively. Extensive experiments on five standard RGBT tracking datasets validate the effectiveness of the proposed method against state-of-the-art methods while achieving the fastest tracking speed of 96.4 FPS.
Andong Lu, Jiacong Zhao, Chenglong Li 0002, Yun Xiao 0003, Bin Luo 0001
ACM Multimedia4
2024 UAV-Ground Visual Tracking: A Unified Dataset and Collaborative Learning Approach
abstract
Visual tracking from the ground view and the UAV view has received increasing attention due to its wide range of practical applications. These two tasks have strong complementary benefits in the description of the target object, such as detailed appearance in the ground view and global motion information in the UAV view, and their combination has the potential to allow the tracking system to be more robust. However, no work has studied this problem in-depth, and it is challenging to accurately combine the ground view information and the UAV view information. To fill the gap and address the challenge, we propose a new computer vision task called UAV-Ground visual tracking. Considering the lack of relevant data and methods, we first propose a unified video dataset called UGVT, which includes 210 pairs of UAV and ground high-resolution video sequences with a total of more than 204K frames, which can be used as a comprehensive evaluation platform for relevant tracking methods. Secondly, based on the newly constructed dataset, we propose a co-learning method called MvCL to fuse the information of ground and UAV views. It first associates the same tracking target in the two views based on cross-attention operation and then fuses the complementary information of the two views. In particular, as a plug-and-play module based on Transformer structure, this method can be flexibly embedded into different tracking frameworks. Extensive experiments are conducted on the newly created dataset. The results demonstrate the effectiveness of the proposed method in improving the robustness of the tracking system compared with 10 state-of-the-art tracking methods and also indicate the prospect and significance of potential UAV-Ground visual tracking research. The dataset is available at:https://github.com/mmic-lcl/Datasets-and-benchmark-code/.
Dengdi Sun, Leilei Cheng, Chenglong Li 0002, Yun Xiao 0003, Bin Luo 0001
IEEE Trans. Circuits Syst. Video Technol.5
2024 ADRNet: Affine and Deformable Registration Networks for Multimodal Remote Sensing Images
abstract
Multi-modal remote sensing images registration ensures the consistency of the spatial positions for different images. It can provide the accurate geographic information and supports the fusion of multi-source data for geospatial analyses and applications. Rigid registration method shows high performance in dealing with large-scale deformation, but it is difficult to achieve high-precision image registration. In contrast, non-rigid registration method is suitable for processing local differences, but cannot effectively deal with large-scale deformation differences. Therefore, the combination of rigid and non-rigid registration methods becomes a necessary strategy to address such issues. In this paper, we propose a novel ADRNet method for multi-modal remote sensing images registration. The proposed ADRNet method contains three main modules: affine registration module, deformable registration module, and spatial transformer module that integrates the affine and deformable transformation parameters to obtain the final aligned images. Meanwhile, we design a new feature enhancement module and an attention module with dilated convolutions which have different dilation rates, which are used to alleviate the limitations imposed by receptive fields in the convolution operation. Moreover, we propose a specific symmetric loss function to optimize the whole network from the perspective of inverse consistency. To assess the efficiency and performance of the network, we extend the experimental data, ranging from cross-modal images in a conventional viewpoint to cross-modal images in a remote sensing viewpoint. The experimental results show that our method exhibits excellent performance for the images with different viewpoints and deformation scales. The relevant code will be released at: https://github.com/Ahuer-Lei/ADRNet.
Yun Xiao 0003, Yuan Chen 0012, Bo Jiang 0002, Jin Tang 0001
IEEE Trans. Geosci. Remote. Sens.1
2024 Dense Tiny Object Detection: A Scene Context Guided Approach and a Unified Benchmark
abstract
With the continuous advancement of remote sensing observation technology, wide-area observation and high-resolution imaging make remote sensing images contain a large number of dense tiny objects. The detection of dense tiny objects is a very challenging task since these objects are with very low resolution and might stick together. Existing work lacks further exploration of the contextual scene information and inherent characteristics of dense tiny objects, which are crucial for performance improvement of dense tiny object detection. In this work, we propose a novel Scene Contextualized Detection Network (SCDNet) by decoupling scene contextual information through a dedicated scene classification sub-network, thereby enabling an enhanced exploration of the relationship between tiny objects and their surrounding environments. In particular, we design a lightweight scene context guided fusion module in SCDNet to incorporate scene context information around dense tiny objects more effectively. Moreover, we further develop the scene context guided foreground enhancement module to suppress the background information while enhancing the foreground information based on the scene information. In addition, this research field still lacks a large-scale benchmark dataset with dense tiny objects, which is crucial for the training and comprehensive evaluation of detection methods. To this end, we construct a large-scale dataset for dense tiny object detection. It contains 11,600 images with 1,019,800 instances, the average absolute size of objects is smaller than 13 pixels, and each image contains 88 objects on average. Extensive experiments are conducted on the proposed dataset, and the results demonstrate the superiority and effectiveness of SCDNet compared to existing methods. The dataset and evaluation code are available at https://github.com/mmic-lcl.
Zhicheng Zhao 0002, Chenglong Li 0002, Yun Xiao 0003, Jin Tang 0001
IEEE Trans. Geosci. Remote. Sens.5
2024 RGBT Tracking via Challenge-Based Appearance Disentanglement and Interaction
abstract
RGB and thermal source data suffer from both shared and specific challenges, and how to explore and exploit them plays a critical role in representing the target appearance in RGBT tracking. In this paper, we propose a novel approach, which performs target appearance representation disentanglement and interaction via both modality-shared and modality-specific challenge attributes, for robust RGBT tracking. In particular, we disentangle the target appearance representations via five challenge-based branches with different structures according to their properties, including three parameter-shared branches to model modality-shared challenges and two parameter-independent branches to model modality-specific challenges. Considering the complementary advantages between modality-specific cues, we propose a guidance interaction module to transfer discriminative features from one modality to another one to enhance the discriminative ability of weak modality. Moreover, we design an aggregation interaction module to combine all challenge-based target representations, which could form more discriminative target representations and fit the challenge-agnostic tracking process. These challenge-based branches are able to model the target appearance under certain challenges so that the target representations can be learned by a few parameters even in the situation of insufficient training data. In addition, to relieve labor costs and avoid label ambiguity, we design a generation strategy to generate training data with different challenge attributes. Comprehensive experiments demonstrate the superiority of the proposed tracker against the state-of-the-art methods on four benchmark datasets.
Lei Liu 0049, Chenglong Li 0002, Yun Xiao 0003, Rui Ruan, Minghao Fan
IEEE Trans. Image Process.3
2023 Quality-Aware RGBT Tracking via Supervised Reliability Learning and Weighted Residual Guidance
abstract
RGB and thermal infrared (TIR) data have different visual properties, which make their fusion essential for effective object tracking in diverse environments and scenes. Existing RGBT tracking methods commonly use attention mechanisms to generate reliability weights for multi-modal feature fusion. However, without explicit supervision, these weights may be unreliably estimated, especially in complex scenarios. To address this problem, we propose a novel Quality-Aware RGBT Tracker (QAT) for robust RGBT tracking. QAT learns reliable weights for each modality in a supervised manner and performs weighted residual guidance to extract and leverage useful features from both modalities. We address the issue of the lack of labels for reliability learning by designing an efficient three-branch network that generates reliable pseudo labels, and a simple binary classification scheme that estimates high-accuracy reliability weights, mitigating the effect of noisy pseudo labels. To propagate useful features between modalities while reducing the influence of noisy modal features on the migrated information, we design a weighted residual guidance module based on the estimated weights and residual connections. We evaluate our proposed QAT on five benchmark datasets, including GTOT, RGBT210, RGBT234, LasHeR, and VTUAV, and demonstrate its excellent performance compared to state-of-the-art methods. Experimental results show that QAT outperforms existing RGBT tracking methods in various challenging scenarios, demonstrating its efficacy in improving the reliability and accuracy of RGBT tracking.
Lei Liu 0049, Chenglong Li 0002, Yun Xiao 0003, Jin Tang 0001
ACM Multimedia3
2023 Thermal UAV Image Super-Resolution Guided by Multiple Visible Cues
abstract
Unmanned aerial vehicle (UAV) thermal-imaging has received much attention, but the insufficient image resolution caused by thermal imaging systems is still a crucial problem that limits the understanding of thermal UAV images. However, high-resolution visible images are relatively easy to access, and it is thus valuable for exploring useful information from visible image to assist thermal UAV image super-resolution (SR). In this article, we propose a novel multiconditioned guidance network (MGNet) to effectively mine the information of visible images for thermal UAV image SR. High-resolution visible UAV images usually contain salient appearance, semantic, and edge information, which plays a critical role in boosting the performance of thermal UAV image SR. Therefore, we design an effective multicue guidance module (MGM) to leverage the appearance, edge, and semantic cues from visible images to guide thermal UAV image SR. In addition, we build the first benchmark dataset for the task of thermal UAV image SR guided by visible images. It is collected by a multimodal UAV platform and composes of 1025 pairs of manually aligned visible and thermal images. Extensive experiments on the built dataset show that our MGNet can effectively leverage useful information from visible images to improve the performance of thermal UAV image SR and perform well against several state-of-the-art methods. The dataset is available at:https://github.com/mmic-lcl/Datasets-and-benchmark-code.
Zhicheng Zhao 0002, Chenglong Li 0002, Yun Xiao 0003, Jin Tang 0001
IEEE Trans. Geosci. Remote. Sens.4
2022 Attribute-Based Progressive Fusion Network for RGBT Tracking
abstract
RGBT tracking usually suffers from various challenge factors, such as fast motion, scale variation, illumination variation, thermal crossover and occlusion, to name a few. Existing works often study fusion models to solve all challenges simultaneously, and it requires fusion models complex enough and training data large enough, which are usually difficult to be constructed in real-world scenarios. In this work, we disentangle the fusion process via the challenge attributes, and thus propose a novel Attribute-based Progressive Fusion Network (APFNet) to increase the fusion capacity with a small number of parameters while reducing the dependence on large-scale training data. In particular, we design five attribute-specific fusion branches to integrate RGB and thermal features under the challenges of thermal crossover, illumination variation, scale variation, occlusion and fast motion respectively. By disentangling the fusion process, we can use a small number of parameters for each branch to achieve robust fusion of different modalities and train each branch using the small training subset with the corresponding attribute annotation. Then, to adaptive fuse features of all branches, we design an aggregation fusion module based on SKNet. Finally, we also design an enhancement fusion transformer to strengthen the aggregated feature and modality-specific features. Experimental results on benchmark datasets demonstrate the effectiveness of our APFNet against other state-of-the-art methods.
Yun Xiao 0003, Chenglong Li 0002, Lei Liu 0049, Jin Tang 0001
AAAI1
2022 The First Challenge on Moving Object Detection and Tracking in Satellite Videos: Methods and Results
abstract
In this paper, we briefly summarize the first challenge on moving object detection and tracking in satellite videos (SatVideoDT). This challenge has three tracks related to satellite video analysis, including moving object detection (Track 1), single object tracking (Track 2), and multiple-object tracking (Track 3). 123, 89, and 70 participants successfully registered, while 37, 42, and 29 teams submitted their final results on the test datasets for Tracks 1-3, respectively. The top-performing methods and their results in each track are described with details. This challenge establishes a new benchmark for satellite video analysis.
Yulan Guo, Qingyong Hu, Feng Zhang 0046, Ye Zhang 0037, Hanyun Wang, Chenguang Dai, Weilong Guo, Xiyu Qi, Kelong Tu, Shudan Zhu, Lai Chen, Bin Lin 0013, Chaocan Xue, Jinlei Zheng, Limei Qin, Ying Li 0017, Manqi Zhao, Lu Ruan 0003, Mingpeng Cui, Guanchen Ding, Guangwei Jiang, Zhenzhong Chen 0001, Kaiyang Cao, Lingyu Kong, Shaodong Chen, Zhicheng Zhao 0001, Qin Shen, Lei Liu 0049, Chenglong Li 0002, Yun Xiao 0003
ICPR36
2022 Global-guided cross-reference network for co-salient object detection
Zhengyi Liu, Yun Xiao 0003
Mach. Vis. Appl.4
2022 AGRFNet: Two-stage cross-modal and multi-level attention gated recurrent fusion network for RGB-D saliency detection
Zhengyi Liu, Yacheng Tan, Yun Xiao 0003
Signal Process. Image Commun.5
2022 SwinNet: Swin Transformer Drives Edge-Aware RGB-D and RGB-T Salient Object Detection
abstract
Convolutional neural networks (CNNs) are good at extracting contexture features within certain receptive fields, while transformers can model the global long-range dependency features. By absorbing the advantage of transformer and the merit of CNN, Swin Transformer shows strong feature representation ability. Based on it, we propose a cross-modality fusion model,SwinNet, for RGB-D and RGB-T salient object detection. It is driven by Swin Transformer to extract the hierarchical features, boosted by attention mechanism to bridge the gap between two modalities, and guided by edge information to sharp the contour of salient object. To be specific, two-stream Swin Transformer encoder first extracts multi-modality features, and then spatial alignment and channel re-calibration module is presented to optimize intra-level cross-modality features. To clarify the fuzzy boundary, edge-guided decoder achieves inter-level cross-modality fusion under the guidance of edge features. The proposed model outperforms the state-of-the-art models on RGB-D and RGB-T datasets, showing that it provides more insight into the cross-modality complementarity task.
Zhengyi Liu, Yacheng Tan, Yun Xiao 0003
IEEE Trans. Circuits Syst. Video Technol.4
2021 TriTransNet: RGB-D Salient Object Detection with a Triplet Transformer Embedding Network
abstract
Salient object detection is the pixel-level dense prediction task which can highlight the prominent object in the scene. Recently U-Net framework is widely used, and continuous convolution and pooling operations generate multi-level features which are complementary with each other. In view of the more contribution of high-level features for the performance, we propose a triplet transformer embedding module to enhance them by learning long-range dependencies across layers. It is the first to use three transformer encoders with shared weights to enhance multi-level features. By further designing scale adjustment module to process the input, devising three-stream decoder to process the output and attaching depth features to color features for the multi-modal fusion, the proposed triplet transformer embedding network (TriTransNet) achieves the state-of-the-art performance in RGB-D salient object detection, and pushes the performance to a new level. Experimental results demonstrate the effectiveness of the proposed modules and the competition of TriTransNet.
Zhengyi Liu, Zhengzheng Tu, Yun Xiao 0003, Bin Tang 0003
ACM Multimedia4
2019 Saliency detection via multi-view graph based saliency optimization
abstract
Saliency detection is an important problem in computer vision and pattern recognition area. Many works have been proposed for addressing the saliency detection task. As a popular method, graph based saliency optimization has been widely studied. However, previous works have universally focussed on single graph optimization which fails to consider multi-view feature representation of image content . In this paper, we first provide a general framework for traditional graph based saliency optimization models. Then, we extend the general framework to the multi-view case and propose our general multi-view graph based saliency optimization model. Finally, we present a particular implementation of our general model and derive an effective updating algorithm to solve it. Experimental results using several benchmark datasets demonstrate the effectiveness of our proposed saliency model.
Yun Xiao 0003, Bo Jiang 0002, Aihua Zheng, Aiwu Zhou, Amir Hussain 0001, Jin Tang 0001
Neurocomputing1
2018 Multi-scale Cooperative Ranking for Saliency Detection
Bo Jiang 0002, Xingyue Jiang, Aihua Zheng, Yun Xiao 0003, Jin Tang 0001
PRCV (1)4
2018 A prior regularized multi-layer graph ranking model for image saliency computation
Yun Xiao 0003, Bo Jiang 0002, Zhengzheng Tu, Jixin Ma 0001, Jin Tang 0001
Neurocomputing1
2017 A new graph ranking model for image saliency detection problem
abstract
Saliency detection is an important problem in many computer vision applications. As a kind of popular method, graph based manifold ranking (GMR) has been successfully used in saliency detection problem. In traditional GMR saliency detection, it involves two main stages, i.e., ranking with background queries and ranking with foreground queries. However, in GMR method, these two stages are conducted separately, which ignores the correlation between background and foreground cues. In this paper, we propose a new graph ranking model, which aims to perform background and foreground ranking simultaneously by exploiting the correlation between background and foreground cues. We derive a closed-form solution for it. Experimental results on four benchmark datasets demonstrate that the proposed method performs better than some other state-of-art methods.
Yuanyuan Guan, Bo Jiang 0002, Yun Xiao 0003, Jin Tang 0001, Bin Luo 0001
SERA3
2017 A global and local consistent ranking model for image saliency computation
Yun Xiao 0003, Bo Jiang 0002, Zhengzheng Tu, Jin Tang 0001
J. Vis. Commun. Image Represent.1