Kun Sun 0002

dblp:30/3530-2 · DBLP profile ↗
← Back
39ranked-venue papers
6as first author
26since 2021 · last 2026
0000-0002-9503-3969ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 2 first-author · 12 since 2021Databases, data management, data science and information retrieval · 5 · 3 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2026 When Genes Speak: A Semantic-Guided Framework for Spatially Resolved Transcriptomics Data Clustering
abstract
Spatial transcriptomics enables gene expression profiling with spatial context, offering unprecedented insights into the tissue microenvironment. However, most computational models treat genes as isolated numerical features, ignoring the rich biological semantics encoded in their symbols. This prevents a truly deep understanding of critical biological characteristics. To overcome this limitation, we present SemST, a semantic-guided deep learning framework for spatial transcriptomics data clustering. SemST leverages Large Language Models (LLMs) to enable genes to "speak" through their symbolic meanings, transforming gene sets within each tissue spot into biologically informed embeddings. These embeddings are then fused with the spatial neighborhood relationships captured by Graph Neural Networks (GNNs), achieving a coherent integration of biological function and spatial structure. We further introduce the Fine-grained Semantic Modulation (FSM) module to optimally exploit these biological priors. The FSM module learns spot-specific affine transformations that empower the semantic embeddings to perform an element-wise calibration of the spatial features, thus dynamically injecting high-order biological knowledge into the spatial context. Extensive experiments on public spatial transcriptomics datasets show that SemST achieves state-of-the-art clustering performance. Crucially, the FSM module exhibits plug-and-play versatility, consistently improving the performance when integrated into other baseline methods.
Jiangkai Long, Yanran Zhu, Chang Tang, Kun Sun 0002, Yuanyuan Liu 0004
AAAI4
2026 SGAT: Learning Feature Matching with Singularity-enhanced Graph Attention Network
abstract
The task of image feature matching aims to establish correct correspondences between images from two different views. While approaches based on attention mechanisms have demonstrated remarkable advancements in image feature matching, they still encounter substantial limitations. Specifically, current graph attention network approaches face performance bottlenecks in complex scenarios, such as low-texture regions or occlusions. This limitation stems from the self-attention mechanism, which, when lacking effective guidance, can lead to divergent attention weights or incorrect focus on regions with low discriminability, resulting in matching failures in low-texture environments. Inspired by how humans focus on distinctive regions when performing cross-view matching, we enhance attention to singular points in images that are salient, unique and have high cross-view matching potential during information aggregation, thereby improving matching capability. To realize the aforementioned strategies, we develop a novel Singularity-enhanced Graph Attention Network (SGAT). SGAT leverages Co-potentiality and Multi-Scale Singularity as prior guidance, and designs a Singularity-aware Attention mechanism and a Co-potentiality Guided Attention mechanism , specifically enhancing the perception of singularity and matching potential during feature interaction. Experimental results on multiple datasets, including ScanNet1500, demonstrate that our method outperforms current state-of-the-art sparse matching methods. In particular, the improvement is most pronounced in complex scenarios such as low-texture environments, significantly enhancing the accuracy and robustness of image matching and its downstream tasks.
Kun Sun 0002, Chang Tang, Yuanyuan Liu 0004, Xin Li 0005
AAAI2
2026 MGC-Net: Learning feature matching with multi-geometry cooperation
Luxia Ai, Kun Sun 0002, Chen Zhang 0043, Nanjun Yuan, Qun Jiang, Wenbing Tao
Knowl. Based Syst.2
2026 Robust Context Modeling for Unsupervised Non-Rigid Point Cloud Correspondence
abstract
We address the “long-range ambiguity” problem for unsupervised non-rigid point cloud correspondence, where corresponding points own inconsistent features while different local regions are spatially or geometrically similar. Previous methods struggle with this problem, since local reference frames (LRF) or coordinate-based methods struggle to exclude locally similar or spatially near mismatches, and widely used independent geometric relations might be inconsistent under non-rigid deformation, introducing extra ambiguity. To this end, we propose a novel robust context modeling module (RCM) to alleviatelong-range ambiguityin two aspects: 1) RCM tackles the ambiguity problem by introducing inter-relation attention (IRA), which mines robust cues from the interplay between relative geometric relations. 2) RCM enhances features with accessiblelong-rangeinformation from IRAs, following a local-to-global manner. Our method shows significant improvements in multiple benchmarks, with accurate correspondence over rotation and large deformation perturbation. Specifically, our method achieves a new state-of-the-art performance with correspondence accuracy of 33.9% and mean error of 4.2 on the SURREAL benchmark.
Rui Li 0013, Jiaming Guo, Ya'nan He, Zhengbao Wang, Xian-Feng Han, Kun Sun 0002, Jiaqi Yang 0002
IEEE Trans. Circuits Syst. Video Technol.7
2025 Spatially Resolved Transcriptomics Data Clustering with Tailored Spatial-scale Modulation
abstract
Spatial transcriptomics, comprising spatial location and high-throughput gene expression information, provides revolutionary insights into disease discovery and cellular evolution. Spatial transcriptomic clustering, which pinpoints distinct spatial domains within tissues, reveals cellular interactions and enhances our understanding of the intricate architecture of tissues. Existing methods typically construct spatial graphs using a static radius based on spatial coordinates, which hinders the accurate identification of spatial domains and complicates the precise partitioning of boundary nodes within clusters. To address this issue, we introduce a novel spatially resolved transcriptomics data clustering network (TSstc). Specifically, we employ a tailored spatial-scale modulation approach, constructing different spatial graphs incrementally as the radius of the spatial domain expands, and a Spatiality-Aware Sampling (SAS) strategy is proposed to aggregate node representations by considering the spatial dependencies between spots. We then use GCN encoders to learn gene embedding with gene graph and multiple spatial embeddings with spatial graphs. During training, we incorporate cross-view correlation-based tailored spatial regularization constraints to preserve high-quality neighbor relationships across spatial embeddings at different scales. Finally, a zero-inflated negative binomial model is utilized to capture the global probability distribution of gene expression profiles. Extensive experimental results demonstrate that our approach surpasses existing state-of-the-art methods in clustering tasks and related downstream applications.
Yuang Xiao, Yanran Zhu, Chang Tang, Yuanyuan Liu 0004, Kun Sun 0002, Xinwang Liu 0002
IJCAI6
2025 Find True Collaborators: Banzhaf Index-based Cross View Alignment for Partially View-aligned Clustering
abstract
Partially view-aligned clustering (PVC) has emerged as a critical area in multi-view clustering, addressing the inherent instance misalignment across views during data collection. The primary challenge of PVC is accurately establishing correspondences between cross-view samples. The Banzhaf index in cooperative game theory serves as an effective tool for modeling complex relationships between multi-view samples by quantifying the marginal contributions of coalition members to collaborative benefits. To this end, we propose a Banzhaf Index-driven cross-view aligNment method, dubbed BIN, which systematically evaluates each view sample's contribution to joint decision-making within a game-theoretic framework. This approach overcomes the limitations of existing PVC methods reliant on prior alignment information and enhances the robustness of multi-view matching. Specifically, we model multi-view samples as players in a cooperative game and quantify their interactions using a payoff model. Simultaneously, we propose a dual-loss constraint: (1) Banzhaf gain loss, which dynamically captures the marginal contribution of key cross-view sample pairs to reinforce associations; (2) contrast loss, which applies exclusion constraints in the feature space to suppress interference from weakly correlated samples. Together, these losses form an effective optimization mechanism. This game-theoretic approach adaptively learns sample correspondences without pre-alignment and ensures robust matching in complex misalignment scenarios. Extensive experiments demonstrate that our method achieves competitive performance against eight state-of-the-art PVC algorithms.
Shanghui Deng, Chang Tang, Kun Sun 0002, Yuanyuan Liu 0004, Xinwang Liu 0002
ACM Multimedia4
2025 TPDepth: Leveraging Text Prompts with ControlNet to Boost Diffusion-based Depth Estimation
abstract
Recent diffusion-based methods have shown strong ability in the depth estimation task, but they largely overlook the rich textual priors embedded in pretrained diffusion models that can enhance both performance and robustness in diverse scenes. In this paper, we propose TPDepth, a diffusion-based, affine-invariant monocular depth estimator that incorporates textual semantics via a Text-Prompted ControlNet. While directly injecting text into the diffusion U-Net can cause the network to over-attend to local semantic cues and compromise global structural modeling, TPDepth processes textual features through a separate ControlNet branch, allowing semantic information to be incorporated without disrupting the spatial reasoning pipeline. Prompt-conditioned features are modulated by an Adaptive Control Scale Module(ACSM) and injected into decoder of the diffusion UNet with skip connections. The model is fine-tuned with a fixed timestep for deterministic prediction. TPDepth achieves state-of-the-art results on NYUv2, KITTI, and ScanNet, and demonstrates competitive performance on two additional zero-shot benchmarks using only 61K training images. Code and models can be found on our https://github.com/Lioely/TPDepth project page.
Yu Liu 0136, Kun Sun 0002, Chang Tang, Xin Li 0005
ACM Multimedia2
2025 SparseMVC: Probing Cross-view Sparsity Variations for Multi-view Clustering
abstract
Existing multi-view clustering methods employ various strategies to address data-level sparsity and view-level dynamic fusion. However, we identify a critical yet overlooked issue: varying sparsity across views. Cross-view sparsity variations lead to encoding discrepancies, heightening sample-level semantic heterogeneity and making view-level dynamic weighting inappropriate. To tackle these challenges, we propose Adaptive Sparse Autoencoders for Multi-View Clustering (SparseMVC), a framework with three key modules. Initially, the sparse autoencoder probes the sparsity of each view and adaptively adjusts encoding formats via an entropy-matching loss term, mitigating cross-view inconsistencies. Subsequently, the correlation-informed sample reweighting module employs attention mechanisms to assign weights by capturing correlations between early-fused global and view-specific features, reducing encoding discrepancies and balancing contributions. Furthermore, the cross-view distribution alignment module aligns feature distributions during the late fusion stage, accommodating datasets with an arbitrary number of views. Extensive experiments demonstrate that SparseMVC achieves state-of-the-art clustering performance. Our framework advances the field by extending sparsity handling from the data-level to view-level and mitigating the adverse effects of encoding discrepancies through sample-level dynamic weighting. The source code is publicly available at https://github.com/cleste-pome/SparseMVC.
Ruimeng Liu, Xin Zou 0001, Chang Tang, Xingchen Hu 0001, Kun Sun 0002, Xinwang Liu 0002
NeurIPS6
2025 Degradation-adaptive attack-robust self-supervised facial representation learning
Yuanyuan Liu 0004, Chang Tang, Kun Sun 0002, Yibing Zhan, Zhe Chen 0013
Neurocomputing4
2025 One-Step Multiview Clustering via Adaptive Graph Learning and Spectral Rotation
abstract
In graph based multiview clustering methods, the ultimate partition result is usually achieved by spectral embedding of the consistent graph using some traditional clustering methods, such as -means. However, optimal performance will be reduced by this multistep procedure since it cannot unify graph learning with partition generation closely. In this article, we propose a one-step multiview clustering method through adaptive graph learning and spectral rotation (AGLSR). For every view, AGLSR adaptively learns affinity graphs to capture similar relationships of samples. Then, a spectral embedding is designed to take advantage of the potential feature space shared by different views. In addition, AGLSR utilizes a spectral rotation strategy to obtain the discrete clustering labels from the learned spectral embeddings directly. An effective updating algorithm with proven convergence is derived to optimize the optimization problem. Sufficient experiments on benchmark datasets have clearly demonstrated the effectiveness of the proposed method in six metrics. The code of AGLSR is uploaded at https://github.com/tangchuan2000/AGLSR.
Chuan Tang, Minhui Wang, Kun Sun 0002
IEEE Trans. Neural Networks Learn. Syst.3
2024 3D Single-Object Tracking in Point Clouds with High Temporal Variation
Qiao Wu, Kun Sun 0002, Pei An, Mathieu Salzmann, Yanning Zhang 0001, Jiaqi Yang 0002
ECCV (7)2
2024 MAC: Maximal Cliques for 3D Registration
abstract
This paper presents a 3D registration method with maximal cliques (MAC) for 3D point cloud registration (PCR). The key insight is to loosen the previous maximum clique constraint and mine more local consensus information in a graph for accurate pose hypotheses generation: 1) A compatibility graph is constructed to render the affinity relationship between initial correspondences. 2) We search for maximal cliques in the graph, each representing a consensus set. 3) Transformation hypotheses are computed for the selected cliques by the SVD algorithm and the best hypothesis is used to perform registration. In addition, we present a variant of MAC if given overlap prior, called MAC-OP. Overlap prior further enhances MAC from many technical aspects, such as graph construction with re-weighted nodes, hypotheses generation from cliques with additional constraints, and hypothesis evaluation with overlap-aware weights. Extensive experiments demonstrate that both MAC and MAC-OP effectively increase registration recall, outperform various state-of-the-art methods, and boost the performance of deep-learned methods. For instance, MAC combined with GeoTransformer achieves a state-of-the-art registration recall of [Formula: see text] on 3DMatch / 3DLoMatch. We perform synthetic experiments on 3DMatch-LIR / 3DLoMatch-LIR, a dataset with extremely low inlier ratios for 3D registration in ultra-challenging cases.
Jiaqi Yang 0002, Xiyu Zhang 0001, Peng Wang 0015, Yulan Guo, Kun Sun 0002, Qiao Wu, Shikun Zhang, Yanning Zhang 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2024 TCTL-Net: Template-Free Color Transfer Learning for Self-Attention Driven Underwater Image Enhancement
abstract
Vision is an important source of information for underwater observations, but underwater images commonly suffer severe visual degradation due to the complexity of the underwater imaging environment and wavelength-dependent absorption effects. There is an urgent need for underwater image enhancement techniques to improve the visual quality of underwater images. Due to the scarcity of high-quality paired training samples, underwater image enhancement based on deep learning has never achieved success similar to other vision tasks. Instead of learning complicated distortion-to-clear mappings with deep networks, we design a template-free color transfer learning framework for predicting transfer parameters, which are more easily captured and described. In addition, we add attention-driven modules to learn differentiated transfer parameters for more flexible and robust enhancement. We verify the effectiveness of our method on multiple publicly available datasets and show its efficiency in enhancing high-resolution images. The source code and the trained models are available on the project homepage: https://trentqq.github.io/TCTL-Net.html.
Kunqian Li, Qi Qi 0008, Chi Yan, Kun Sun 0002, Q. M. Jonathan Wu
IEEE Trans. Circuits Syst. Video Technol.5
2024 Learning Scribbles for Dense Depth: Weakly Supervised Single Underwater Image Depth Estimation Boosted by Multitask Learning
abstract
Estimating depth from a single underwater image is one of the main tasks of underwater visual perception. However, data-driven underwater depth estimation methods have long been challenging to make breakthroughs due to the difficulty of obtaining a large number of true-value references. This is partly due to the high cost of acquisition equipment, which is difficult to be applied to diverse ocean scenes by a wide range of users, and therefore sample diversity is difficult to guarantee; on the other hand, manual annotation of dense depth relationships is almost impossible to achieve. In this paper, we establish a new underwater relative depth estimation benchmark, namely SUIM-SDA, by extending the SUIM dataset with more than 6,000 manually annotated depth trendlines, 25 million pixels with paired depth-ranking labels and 14 million depth-ranked pixel pairs. Using the sparse depth relation annotation provided by SUIM-SDA and the semantic information provided by SUIM, we design a new multi-stage multi-task learning framework to predict a dense relative depth map for a single underwater image. Comprehensive comparison and ablation study on the publicly available dataset and our new benchmark demonstrate the effectiveness of the proposed weakly-supervised strategy for dense relative depth estimation. The new benchmark, source code, and trained models are available on the project home page: https://wangxy97.github.io/WsUIDNet.
Kunqian Li, Xiya Wang, Qi Qi 0008, Guojia Hou, Zhiguo Zhang 0005, Kun Sun 0002
IEEE Trans. Geosci. Remote. Sens.7
2024 Efficient and Effective One-Step Multiview Clustering
abstract
Multiview clustering algorithms have attracted intensive attention and achieved superior performance in various fields recently. Despite the great success of multiview clustering methods in realistic applications, we observe that most of them are difficult to apply to large-scale datasets due to their cubic complexity. Moreover, they usually use a two-stage scheme to obtain the discrete clustering labels, which inevitably causes a suboptimal solution. In light of this, an efficient and effective one-step multiview clustering (E2OMVC) method is proposed to directly obtain clustering indicators with a small-time burden. Specifically, according to the anchor graphs, the smaller similarity graph of each view is constructed, from which the low-dimensional latent features are generated to form the latent partition representation. By introducing a label discretization mechanism, the binary indicator matrix can be directly obtained from the unified partition representation which is formed by fusing all latent partition representations from different views. In addition, by coupling the fusion of all latent information and the clustering task into a joint framework, the two processes can help each other and obtain a better clustering result. Extensive experimental results demonstrate that the proposed method can achieve comparable or better performance than the state-of-the-art methods. The demo code of this work is publicly available at https://github.com/WangJun2023/EEOMVC.
Jun Wang 0118, Chang Tang, Zhiguo Wan, Wei Zhang 0049, Kun Sun 0002, Albert Y. Zomaya
IEEE Trans. Neural Networks Learn. Syst.5
2023 MixCycle: Mixup Assisted Semi-Supervised 3D Single Object Tracking with Cycle Consistency
abstract
3D single object tracking (SOT) is an indispensable part of automated driving. Existing approaches rely heavily on large, densely labeled datasets. However, annotating point clouds is both costly and time-consuming. Inspired by the great success of cycle tracking in unsupervised 2D SOT, we introduce the first semi-supervised approach to 3D SOT. Specifically, we introduce two cycle-consistency strategies for supervision: 1) Self tracking cycles, which leverage labels to help the model converge better in the early stages of training; 2) forward-backward cycles, which strengthen the tracker’s robustness to motion variations and the template noise caused by the template update strategy. Furthermore, we propose a data augmentation strategy named SOTMixup to improve the tracker’s robustness to point cloud diversity. SOTMixup generates training samples by sampling points in two point clouds with a mixing rate and assigns a reasonable loss weight for training according to the mixing rate. The resulting MixCycle approach generalizes to appearance matching-based trackers. On the KITTI benchmark, based on the P2B tracker [16], MixCycle trained with 10% labels outperforms P2B trained with 100% labels, and achieves a 28.4% precision improvement when using 1% labels. Our code will be released at https://github.com/Mumuqiao/MixCycle.
Qiao Wu, Jiaqi Yang 0002, Kun Sun 0002, Chu'ai Zhang, Yanning Zhang 0001, Mathieu Salzmann
ICCV3
2023 Hierarchical Attention Learning for Multimodal Classification
abstract
Multimodal learning aims to integrate complementary information from different modalities for more reliable decisions. However, existing multimodal classification methods simply integrate the learned local features, which ignore the underlying structure of each modality and the higher-order correlation across modalities. In this paper, we propose a novel Hierarchical Attention Learning Network (HALNet) for multimodal classification. Specifically, HALNet has three merits: 1) A hierarchical feature fusion module is proposed to learn multilevel features, aggregating multi-level features for a global feature representation with the attention mechanism and progressive fusion tactics. 2) A cross-modal higher-order fusion module is introduced to capture the prospective cross-modal correlations at label space. 3) A dual prediction pattern is designed to generate credible decisions. Extensive experiments on three real-world multimodal datasets demonstrate that HALNet achieves competitive performance compared to the state-of-the-art.
Xin Zou 0001, Chang Tang, Wei Zhang 0049, Kun Sun 0002, Liangxiao Jiang
ICME4
2023 Mutual structure learning for multiple kernel clustering
Zhenglai Li, Chang Tang, Zhiguo Wan, Kun Sun 0002, Wei Zhang 0049, Xinzhong Zhu
Inf. Sci.5
2023 Inclusivity induced adaptive graph learning for multi-view clustering
Xin Zou 0001, Chang Tang, Kun Sun 0002, Wei Zhang 0049, Deqiong Ding
Knowl. Based Syst.4
2023 Towards Accurate Image Matching by Exploring Redundancy Between Multiple Descriptors
abstract
Finding correspondences between a pair of images is the key ingredient for many applications such as localization and panorama. However, due to a variety of challenges between multi-view images in practice, the results of using a single kind of descriptor may vary significantly across different scenes. In this paper, we treat the assignment task as a clustering problem and propose an image matching method that fuses multiple descriptors to tackle the above difficulties. First, we extract multiple descriptors at the keypoints on two images. Then, we compute a pairwise similarity matrix for each kind of descriptor. Afterwards, we compute a weighted combination of these similarity matrices, and use it to build correspondences via a modified multi-kernel clustering module. The proposed method is tested on three public image datasets: two ground image sets and an Unmanned Aerial Vehicle (UAV) image set. Experiments show that the proposed method can adapt to different number of descriptors. It significantly improves the matching accuracy in a variety of scenarios and downstream tasks.
Jinhong Yu, Kun Sun 0002, Kunqian Li, Chuan Tang, Ruyi Feng
IEEE Geosci. Remote. Sens. Lett.2
2023 Multi-view subspace clustering via adaptive graph learning and late fusion alignment
Chuan Tang, Kun Sun 0002, Chang Tang, Xinwang Liu 0002, Junjie Huang 0001, Wei Zhang 0049
Neural Networks2
2023 SC$^{2}$2-PCR++: Rethinking the Generation and Selection for Efficient and Robust Point Cloud Registration
abstract
Outlier removal is a critical part of feature-based point cloud registration. In this paper, we revisit the model generation and selection of the classic RANSAC approach for fast and robust point cloud registration. For the model generation, we propose a second-order spatial compatibility (SC$^{2}$) measure to compute the similarity between correspondences. It takes into account global compatibility instead of local consistency, allowing for more distinctive clustering between inliers and outliers at an early stage. The proposed measure promises to find a certain number of outlier-free consensus sets using fewer samplings, making the model generation more efficient. For the model selection, we propose a new Feature and Spatial consistency constrained Truncated Chamfer Distance (FS-TCD) metric for evaluating the generated models. It considers the alignment quality, the feature matching properness, and the spatial consistency constraint simultaneously, enabling the correct model to be selected even when the inlier rate of the putative correspondence set is extremely low. Extensive experiments are carried out to investigate the performance of our method. In addition, we also experimentally prove that the proposed SC$^{2}$measure and the FS-TCD metric are general and can be easily plugged into deep learning based frameworks. The code will be available athttps://github.com/ZhiChen902/SC2-PCR-plusplus.
Zhi Chen 0011, Kun Sun 0002, Fan Yang 0088, Lin Guo 0019, Wenbing Tao
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 Object Detection in Hyperspectral Image via Unified Spectral-Spatial Feature Aggregation
abstract
Deep learning-based hyperspectral image (HSI) classification and object detection techniques have gained significant attention due to their vital role in image content analysis, interpretation, and broader HSI applications. However, current hyperspectral object detection approaches predominantly emphasize spectral or spatial information, overlooking the valuable complementary relationship between these two aspects. In this study, we present a novel Spectral-Spatial Aggregation (S2ADet) object detector that effectively harnesses the rich spectral and spatial complementary information inherent in the hyperspectral image. S2ADet comprises a hyperspectral information decoupling (HID) module, a two-stream feature extraction network, and a one-stage detection head. The HID module processes hyperspectral data by aggregating spectral and spatial information via band selection and principal components analysis, consequently reducing redundancy. Based on the acquired spectral and spatial aggregation information, we propose a feature aggregation two-stream network for interacting spectral-spatial features. Furthermore, to address the limitations of existing databases, we annotate an extensive dataset, designated as HOD3K, containing 3,242 hyperspectral images captured across diverse real-world scenes and encompassing three object classes. These images possess a resolution of 512×256 pixels and cover 16 bands ranging from 470 nm to 620 nm. Comprehensive experiments on two datasets demonstrate that S2ADet surpasses existing state-of-the-art methods, achieving robust and reliable results. The demo code and dataset of this work are publicly available at https://github.com/hexiao-cs/S2ADet.
Xiao He 0010, Chang Tang, Xinwang Liu 0002, Wei Zhang 0049, Kun Sun 0002, Jiangfeng Xu
IEEE Trans. Geosci. Remote. Sens.5
2022 SC2-PCR: A Second Order Spatial Compatibility for Efficient and Robust Point Cloud Registration
abstract
In this paper, we present a second order spatial compat-ibility (SC2) measure based method for efficient and robust point cloud registration (PCR), called SC2-PCR 1. Firstly, we propose a second order spatial compatibility (SC2) mea-sure to compute the similarity between correspondences. It considers the global compatibility instead of local consis-tency, allowing for more distinctive clustering between in-liers and outliers at early stage. Based on this measure, our registration pipeline employs a global spectral technique to find some reliable seeds from the initial correspondences. Then we design a two-stage strategy to expand each seed to a consensus set based on the SC2measure matrix. Finally, we feed each consensus set to a weighted SVD algorithm to generate a candidate rigid transformation and select the best model as the final result. Our method can guarantee to find a certain number of outlier-free consensus sets using fewer samplings, making the model estimation more ef-ficient and robust. In addition, the proposed SC2measure is general and can be easily plugged into deep learning based frameworks. Extensive experiments are carried out to in-vestigate the performance of our method.
Zhi Chen 0011, Kun Sun 0002, Fan Yang 0088, Wenbing Tao
CVPR2
2022 Robust consensus-aware network for 3D point registration
Fan Yang 0088, Zhi Chen 0011, Kun Sun 0002, Liman Liu, Wenbing Tao
Neurocomputing3
2022 SGUIE-Net: Semantic Attention Guided Underwater Image Enhancement With Multi-Scale Perception
abstract
Due to the wavelength-dependent light attenuation, refraction and scattering, underwater images usually suffer from color distortion and blurred details. However, due to the limited number of paired underwater images with undistorted images as reference, training deep enhancement models for diverse degradation types is quite difficult. To boost the performance of data-driven approaches, it is essential to establish more effective learning mechanisms that mine richer supervised information from limited training sample resources. In this paper, we propose a novel underwater image enhancement network, called SGUIE-Net, in which we introduce semantic information as high-level guidance via region-wise enhancement feature learning. Accordingly, we propose semantic region-wise enhancement module to better learn local enhancement features for semantic regions with multi-scale perception. After using them as complementary features and feeding them to the main branch, which extracts the global enhancement features on the original image scale, the fused features bring semantically consistent and visually superior enhancements. Extensive experiments on the publicly available datasets and our proposed dataset demonstrate the impressive performance of SGUIE-Net. The code and proposed dataset are available at https://trentqq.github.io/SGUIE-Net.html.
Qi Qi 0008, Kunqian Li, Haiyong Zheng, Xiang Gao 0009, Guojia Hou, Kun Sun 0002
IEEE Trans. Image Process.6
2020 R²MRF: Defocus Blur Detection via Recurrently Refining Multi-Scale Residual Features
abstract
Defocus blur detection aims to separate the in-focus and out-of-focus regions in an image. Although attracting more and more attention due to its remarkable potential applications, there are still several challenges for accurate defocus blur detection, such as the interference of background clutter, sensitivity to scales and missing boundary details of defocus blur regions. In order to address these issues, we propose a deep neural network which Recurrently Refines Multi-scale Residual Features (R2MRF) for defocus blur detection. We firstly extract multi-scale deep features by utilizing a fully convolutional network. For each layer, we design a novel recurrent residual refinement branch embedded with multiple residual refinement modules (RRMs) to more accurately detect blur regions from the input image. Considering that the features from bottom layers are able to capture rich low-level features for details preservation while the features from top layers are capable of characterizing the semantic information for locating blur regions, we aggregate the deep features from different layers to learn the residual between the intermediate prediction and the ground truth for each recurrent step in each residual refinement branch. Since the defocus degree is sensitive to image scales, we finally fuse the side output of each branch to obtain the final blur detection map. We evaluate the proposed network on two commonly used defocus blur detection benchmark datasets by comparing it with other 11 state-of-the-art methods. Extensive experimental results with ablation studies demonstrate that R2MRF consistently and significantly outperforms the competitors in terms of both efficiency and accuracy.
Chang Tang, Xinwang Liu 0002, Xinzhong Zhu, En Zhu, Kun Sun 0002, Pichao Wang, Lizhe Wang 0001, Albert Y. Zomaya
AAAI5
2020 Guide to Match: Multi-Layer Feature Matching With a Hybrid Gaussian Mixture Model
abstract
As a fundamental yet challenging task in computer vision, finding correspondences between two sets of feature points has received extensive attention. Among all the proposed methods, the Gaussian Mixture Model (GMM) based algorithms show their great power in formulating such problems. However, they are vulnerable to large portion of outliers in the extracted feature points. In this paper, a new Hybrid Gaussian Mixture Model (HGMM) combined with a multi-layer matching framework is proposed. Different from existing GMM based methods, HGMM uses a set of seed correspondences to guide the matching procedure. To automatically find seed correspondences, the feature points are divided into multiple layers according to their matching potential. With the help of Locality Sensitive Hashing, this can be done economically and efficiently. Correspondences found in lower layers which contain few outliers will be used as hard constraint when matching features in higher layers where a large portion of outliers exist. Extensive experiments show that the proposed method is efficient and more robust to outliers when images have large viewpoint difference or small scene overlap.
Kun Sun 0002, Wenbing Tao
IEEE Trans. Multim.1
2019 A center-driven image set partition algorithm for efficient structure from motion
Kun Sun 0002, Wenbing Tao
Inf. Sci.1
2018 Rolling Guidance Based Scaled-Aware Spatial Sparse Unmixing for Hyperspectral Remote Sensing Imagery
abstract
Spatial regularization based sparse unmixing has been attracted much attention and has achieved improved fractional abundance results. However, the traditional approach to spatial consideration can only suppress discrete wrong unmixing points and smooth an abundance map with low-contrast changes, and it has no concept of scale difference. As the different levels of structures and edges in remote sensing have different meanings and importance, to better extract the different levels of spatial details, rolling guidance based scale-aware spatial sparse unmixing (RGSU), is proposed in this paper to extract and recover the different levels important structures and details in the hyperspectral remote sensing image unmixing procedure. Differing from the existing spatial regularization based sparse unmixing approaches, the proposed method considers the different levels of edges by combining a Gaussian filter-like method to realize small-scale structure removal with a joint bilateral filtering process to account for the spatial domain and range domain correlations. The experimental results obtained with both simulated and real hyperspectral images show that the proposed method achieves a better performance and produces more accurate abundance maps, as well as higher quantitative results, when compared to the current state-of-the-art sparse unmixing algorithms
Ruyi Feng, Tian Tian 0007, Xianju Li, Kun Sun 0002
IGARSS4
2018 A Constrained Radial Agglomerative Clustering Algorithm for Efficient Structure From Motion
abstract
Building three-dimensional models effectively and accurately is an important issue. In this letter, a new image set partition method for efficient structure from motion (SfM) from a set of unevenly distributed images is proposed. Given the largest connected component in the image matching graph, we first reconstruct a base model from a set of images with large overlap and sufficient feature correspondences. Then, a novel constrained radial agglomerative clustering algorithm is proposed to divide the remaining images, so that each image cluster could be independently added to the base model in parallel. Finally, all the partial models are merged into a complete scene. Experiment results show that the proposed method works better than the popular normalized-cuts-based SfM method.
Kun Sun 0002, Wenbing Tao
IEEE Signal Process. Lett.1
2017 Image Matching via Feature Fusion and Coherent Constraint
abstract
The Gaussian mixture model (GMM)-based methods have achieved great success in point set registration. However, they cannot be directly applied to image matching, because the features extracted from two images usually contain a large portion of outliers. In this letter, we propose a new method to extend the powerful GMM to the field of image feature points matching. The algorithm consists of two main steps. In the first step, points extracted from the images are mapped into a new subspace, in which feature similarity information is fused to get the new representation of the points. The second step performs an improved progressive process with the GMM to find correspondences satisfying the coherent constraint. In this way, finding correspondences among large outliers is feasible and the iteration converges faster. Experimental results on benchmark data sets show that the proposed method can find more correct matches with high accuracy.
Kun Sun 0002, Liman Liu, Wenbing Tao
IEEE Geosci. Remote. Sens. Lett.1
2016 Progressive match expansion via coherent subspace constraint
Kun Sun 0002, Liman Liu, Wenbing Tao
Inf. Sci.1
2016 Pedestrian detection aided by fusion of binocular information
Zhiguo Zhang 0005, Wenbing Tao, Kun Sun 0002
Pattern Recognit.3
2015 Feature Guided Biased Gaussian Mixture Model for image matching
Kun Sun 0002, Wenbing Tao, Yuan Yan Tang
Inf. Sci.1
2015 SaCoseg: Object Cosegmentation by Shape Conformability
abstract
In this paper, an object cosegmentation method based on shape conformability is proposed. Different from the previous object cosegmentation methods which are based on the region feature similarity of the common objects in image set, our proposed SaCoseg cosegmentation algorithm focuses on the shape consistency of the foreground objects in image set. In the proposed method, given an image set where the implied foreground objects may be varied in appearance but share similar shape structures, the implied common shape pattern in the image set can be automatically mined and regarded as the shape prior of those unsatisfactorily segmented images. The SaCoseg algorithm mainly consists of four steps: 1) the initial Grabcut segmentation; 2) the shape mapping by coherent point drift registration; 3) the common shape pattern discovery by affinity propagation clustering; and 4) the refinement by Grabcut with common shape constraint. To testify our proposed algorithm and establish a benchmark for future work, we built the CoShape data set to evaluate the shape-based cosegmentation. The experiments on CoShape data set and the comparison with some related cosegmentation algorithms demonstrate the good performance of the proposed SaCoseg algorithm.
Wenbing Tao, Kunqian Li, Kun Sun 0002
IEEE Trans. Image Process.3
2015 Robust Point Sets Matching by Fusing Feature and Spatial Information Using Nonuniform Gaussian Mixture Models
abstract
Most of the traditional methods that handle the point sets matching between two images are based on local feature descriptors and the succedent mismatch eliminating strategies, which usually suffers from the sparsity of the initial match set because some correct ambiguous associations are easily filtered out by the ratio test of SIFT matching due to their second ranking in feature similarity. In this paper, we propose a nonuniform Gaussian mixture model (NGMM) for point sets matching between a pair of images which combines feature with position information of the local feature points extracted from the image pair to achieve point sets matching in a GMM framework. The proposed point set matching using an NGMM is able to change the correspondence assignments throughout the matching process and has the potential to match up even ambiguous matches correctly. The proposed NGMM framework can be either used to directly find matches between two point sets obtained from two images or applied to remove outliers in a match set. When finding matches, NGMM tries to learn a nonrigid transformation between the two point sets and provide a probability for every found match to measure the reliability of the match. Then, a probability threshold can be used to get the final robust match set. When removing outliers, NGMM requires that the vector field formed by the correct matches to be coherent and the matches contradicting the coherent vector field will be regarded as mismatches to be removed. A number of comparison and evaluation experiments reveal the good performance of the proposed NGMM framework in both finding matches and discarding mismatches.
Wenbing Tao, Kun Sun 0002
IEEE Trans. Image Process.2
2014 Asymmetrical Gauss Mixture Models for Point Sets Matching
abstract
The probabilistic methods based on Symmetrical Gauss Mixture Model(SGMM)[4, 13, 8] have achieved great success in point sets registration, but are seldom used to find the correspondences between two images due to the complexity of the non-rigid transformation and too many outliers. In this paper we propose an Asymmetrical GMM(AGMM) for point sets matching between a pair of images. Different from the previous SGMM, the AGMM gives each Gauss component a different weight which is related to the feature similarity between the data point and model point, which leads to two effective algorithms: the Single Gauss Model for Mismatch Rejection(SGMR) algorithm and the AGMM algorithm for point sets matching. The SGMR algorithm iteratively filters mismatches by estimating a non-rigid transformation between two images based on the spatial coherence of point sets. The AGMM algorithm combines the feature information with position information of the SIFT feature points extracted from the images to achieve point sets matching so that much more correct correspondences with high precision can be found. A number of comparison and evaluation experiments reveal the excellent performance of the proposed SGMR algorithm and AGMM algorithm.
Wenbing Tao, Kun Sun 0002
CVPR2
2014 Spatial adjacent bag of features with multiple superpixels for object segmentation and classification
Wenbing Tao, Yicong Zhou, Liman Liu, Kunqian Li, Kun Sun 0002, Zhiguo Zhang 0005
Inf. Sci.5