Kun Ding 0001

dblp:81/7632-1 · DBLP profile ↗
← Back
32ranked-venue papers
11as first author
19since 2021 · last 2026
0000-0002-2256-8815ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 21 · 6 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 6 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 2 since 2021
YearPublicationVenuePosition
2026 LookFlow: Training-Free and Efficient High-Resolution Image Synthesis via Dynamic Lookahead Guidance Flow
abstract
Rectification flow Transformers (RFTs) have shown promising performance in diffusion-based image synthesis but are typically confined to lower-resolution scenarios, limiting their ability to generate high-resolution images. Existing resolution extrapolation approaches often suffer from excessive computational overhead, resulting in prolonged inference times. We propose LookFlow, a training-free high-resolution synthesis framework that accelerates inference while preserving visual quality. Building on pretrained text-to-image RFTs, LookFlow employs a dynamic lookahead guidance flow mechanism to refine high-resolution velocity predictions by leveraging multi-timestep lookahead information extracted from a low-resolution flow. Additionally, reusing temporally similar features across consecutive timesteps drastically reduces computation and significantly decreases inference time overhead. Extensive experiments on COCO demonstrate that LookFlow robustly scales resolutions from 4× to 25×, achieving up to a maximum speedup of 2.01× while maintaining competitive visual fidelity.
Jianlong Chang, Ying Wang 0008, Kun Ding 0001, Shiming Xiang
AAAI6
2026 Beyond Counting: Evaluating Abstract and Emotional Reasoning in Vision-Language Models
abstract
Despite the rapid progress of Vision Language Models (VLMs), existing benchmarks still concentrate on coarse-grained object recognition or simple relational reasoning, leaving the fine-grained and higher-order reasoning abilities of these systems largely unexamined. To bridge this critical evaluation gap, we introduce EmojiGrid, a novel diagnostic benchmark specifically designed to probe these fine-grained and higher-order skills. Leveraging the universal and semantically rich nature of emojis, we synthesize a grid‑based visual dataset paired with 29,000+ QA pairs. Each pair is explicitly anchored in a three-level cognitive taxonomy comprising (i) Perception and Information Extraction, (ii) Relational and Structural Reasoning, and (iii) Abstraction and Advanced Cognition. These dimensions further decompose into nine categories covering a broad range of cognitive skills, including counting, spatial relations, compositional logic, semantic sentiment, and related higher-order reasoning tasks. Our extensive evaluation of 25 state-of-the-art open-source and proprietary VLMs reveals a significant performance gap between foundational perceptual tasks and higher-level cognitive abilities, particularly in abstraction and advanced emotional reasoning. Notably, all models struggle with compositional logic, spatial consistency, and especially emotional and semantic understanding. EmojiGrid provides a quantifiable, fine-grained benchmark to diagnose VLM limitations and guides future progress toward models that can truly perceive, reason about, and interpret complex, symbol-rich visual scenes.
Jianlong Chang, Ying Wang 0008, Kun Ding 0001, Shiming Xiang
AAAI6
2026 Efficient redundancy reduction for open-vocabulary semantic segmentation
Qi Yang 0015, Kun Ding 0001, Qiyuan Cao, Shiming Xiang
Neurocomputing3
2026 AtmosOceanNet: SST forecast method driven by atmospheric-oceanic multimodal data
Qinxuan Wang, Kun Ding 0001, Ying Wang 0008, Yineng Li, Shiming Xiang, Xiaoqing Chu
Pattern Recognit.3
2026 USVTrack: A Benchmark for Multi-Object Tracking in Complex Water Surface Scenes
abstract
Multi-object tracking (MOT) in water surface scenes is crucial for the autonomous navigation of Unmanned Surface Vehicles (USVs). However, existing MOT datasets rarely focus on these scenes. Moreover, the few available water surface MOT datasets contain limited data shot onboard and concentrate narrowly on specific marine scenes, creating a significant gap from real-world USV navigation applications. To promote research on USV autonomous navigation, we introduce USVTrack, a fully onboard-shot MOT benchmark that covers diverse and complex water surface scenes, characterized by a high proportion of small objects and varied backgrounds. Then, we propose an innovative end-to-end method specifically designed for MOT in complex water surface scenes, termed as USVMOT. It improves tracking performance through four key contributions: 1) integrating mask information via knowledge distillation to boost feature discriminability; 2) deploying task-specific auxiliary pathways to alleviate the competition between detection and re-identification (ReID) in end-to-end MOT methods; 3) employing an adaptive high-quality mask generation strategy based on the Segment Anything Model (SAM) that obviates extensive manual annotation; and 4) introducing an object-aware association method that dynamically tailors the tracking strategy according to object size and motion speed. Extensive experiments on the USVTrack benchmark demonstrate that USVMOT outperforms existing methods. Our analysis reveals that MOT in complex water surface scenes remains challenging, highlighting the need for further advancements.
Yuwei Cheng, Kun Ding 0001, Chunhong Pan, Shiming Xiang
IEEE Trans. Circuits Syst. Video Technol.3
2025 UNIP: Rethinking Pre-trained Attention Patterns for Infrared Semantic Segmentation
abstract
Pre-training techniques significantly enhance the performance of semantic segmentation tasks with limited training data. However, the efficacy under a large domain gap between pre-training (e.g. RGB) and fine-tuning (e.g. infrared) remains underexplored. In this study, we first benchmark the infrared semantic segmentation performance of various pre-training methods and reveal several phenomena distinct from the RGB domain. Next, our layerwise analysis of pre-trained attention maps uncovers that: (1) There are three typical attention patterns (local, hybrid, and global); (2) Pre-training tasks notably influence pattern distribution across layers; (3) The hybrid pattern is crucial for semantic segmentation as it attends to both nearby and foreground elements; (4) The texture bias impedes model generalization in infrared tasks. Building on these insights, we propose UNIP, a UNified Infrared Pre-training framework, to enhance the pre-trained model performance. This framework uses the hybrid-attention distillation NMI-HAD as the pre-training target, a large-scale mixed dataset InfMix for pre-training, and a last-layer feature pyramid network LL-FPN for fine-tuning. Experimental results show that UNIP outperforms various pre-training methods by up to 13.5% in average mIoU on three infrared segmentation tasks, evaluated using fine-tuning and linear probing metrics. UNIP-S achieves performance on par with MAE-L while requiring only 1/10 of the computational cost. Furthermore, with fewer parameters, UNIP significantly surpasses state-of-the-art (SOTA) infrared or RGB segmentation methods and demonstrates the broad potential for application in other modalities, such as RGB and depth. Our code is available at https://github.com/casiatao/UNIP.
Jinyong Wen, Kun Ding 0001, Shiming Xiang, Chunhong Pan
ICLR4
2025 Re-ranking Reasoning Context with Tree Search Makes Large Vision-Language Models Stronger
abstract
Recent advancements in Large Vision Language Models (LVLMs) have significantly improved performance in Visual Question Answering (VQA) tasks through multimodal Retrieval-Augmented Generation (RAG). However, existing methods still face challenges, such as the scarcity of knowledge with reasoning examples and erratic responses from retrieved knowledge. To address these issues, in this study, we propose a multimodal RAG framework, termed RCTS, which enhances LVLMs by constructing a Reasoning Context-enriched knowledge base and a Tree Search re-ranking method. Specifically, we introduce a self-consistent evaluation mechanism to enrich the knowledge base with intrinsic reasoning patterns. We further propose a Monte Carlo Tree Search with Heuristic Rewards (MCTS-HR) to prioritize the most relevant examples. This ensures that LVLMs can leverage high-quality contextual reasoning for better and more consistent responses. Extensive experiments demonstrate that our framework achieves state-of-the-art performance on multiple VQA datasets, significantly outperforming In-Context Learning (ICL) and Vanilla-RAG methods. It highlights the effectiveness of our knowledge base and re-ranking method in improving LVLMs.
Qi Yang 0015, Chenghao Zhang 0003, Lubin Fan, Kun Ding 0001, Jieping Ye, Shiming Xiang
ICML4
2025 Agent Reviewers: Domain-specific Multimodal Agents with Shared Memory for Paper Review
abstract
Feedback from peer review is essential to improve the quality of scientific articles. However, at present, many manuscripts do not receive sufficient external feedback for refinement before or during submission. Therefore, a system capable of providing detailed and professional feedback is crucial for enhancing research efficiency. In this paper, we have compiled the largest dataset of paper reviews to date by collecting historical open-access papers and their corresponding review comments and standardizing them using LLM. We then developed a multi-agent system that mimics real human review processes, based on LLMs. This system, named Agent Reviewers, includes the innovative introduction of multimodal reviewers to provide feedback on the visual elements of papers. Additionally, a shared memory pool that stores historical papers' metadata is preserved, which supplies reviewer agents with background knowledge from different fields. Our system is evaluated using ICLR 2024 papers and achieves superior performance compared to existing AI-based review systems. Comprehensive ablation studies further demonstrate the effectiveness of each module and agent in this system.
Shixiong Xu, Jinqiu Li, Kun Ding 0001, Gaofeng Meng
ICML4
2025 EvoVLMA: Evolutionary Vision-Language Model Adaptation
Kun Ding 0001, Ying Wang 0008, Shiming Xiang
ACM Multimedia1
2025 Knowledge-based Visual Question Answer with Multimodal Processing, Retrieval and Filtering
abstract
The task of Knowlegde-Based Visual Question Answering (KB-VQA) requires the model to understand visual features and retrieve external knowledge. Retrieval-Augmented Generation (RAG) have been employed to address this problem through knowledge base querying. However, existing work demonstrate two limitations: insufficient interactivity during knowledge retrieval and ineffective organization of retrieved information for Visual-Language Model (VLM). To address these challenges, we propose a three-stage visual language model with Process, Retrieve and Filter (VLM-PRF) framework. For interactive retrieval, VLM-PRF uses reinforcement learning (RL) to guide the model to strategically process information via tool-driven operations. For knowledge filtering, our method trains the VLM to transform the raw retrieved information into into task-specific knowledge. With a dual reward as supervisory signals, VLM-PRF successfully enable model to optimize retrieval strategies and answer generation capabilities simultaneously. Experiments on two datasets demonstrate the effectiveness of our framework.
Yuyang Hong, Qi Yang 0015, Lubin Fan, Ying Wang 0008, Kun Ding 0001, Shiming Xiang, Jieping Ye
NeurIPS7
2025 Diffusion Model as a Noise-Aware Latent Reward Model for Step-Level Preference Optimization
abstract
Preference optimization for diffusion models aims to align them with human preferences for images. Previous methods typically use Vision-Language Models (VLMs) as pixel-level reward models to approximate human preferences. However, when used for step-level preference optimization, these models face challenges in handling noisy images of different timesteps and require complex transformations into pixel space. In this work, we show that pre-trained diffusion models are naturally suited for step-level reward modeling in the noisy latent space, as they are explicitly designed to process latent images at various noise levels. Accordingly, we propose the **Latent Reward Model (LRM)**, which repurposes components of the diffusion model to predict preferences of latent images at arbitrary timesteps. Building on LRM, we introduce **Latent Preference Optimization (LPO)**, a step-level preference optimization method conducted directly in the noisy latent space. Experimental results indicate that LPO significantly improves the model's alignment with general, aesthetic, and text-image alignment preferences, while achieving a 2.5-28x training speedup over existing preference optimization methods.
Cheng Da, Kun Ding 0001, Huan Yang 0005, Yan Li 0043, Tingting Gao, Di Zhang 0026, Shiming Xiang, Chunhong Pan
NeurIPS3
2025 Instructed fine-tuning based on semantic consistency constraint for deep multi-view stereo
Hongping Yan, Kun Ding 0001, Tingting Cai, Yueyue Zhou
Appl. Intell.3
2025 CLIP-MoA: Visual-Language Models With Mixture of Adapters for Multitask Remote Sensing Image Classification
abstract
In recent years, the rapid growth of remote sensing (RS) datasets has greatly enhanced the difficulty of RS image classification tasks. Given the remarkable success of large visual-language models (VLMs) in various domains, their potential for RS image classification tasks has attracted significant attention. In this study, we focus on Contrastive Language-Image Pre-training (CLIP), a widely recognized contrastive learning model known for its strong zero-shot transferability. However, the domain gap between the pre-training data (e.g., natural images) of CLIP and RS data can significantly hinder the application of CLIP. First, the scarcity of labeled RS data poses a significant challenge for fine-tuning CLIP. Second, the complexity and diversity of RS images substantially increase the fine-tuning time and computational resources needed to individually adapt CLIP to new tasks. Third, independently fine-tuning CLIP for each RS dataset ignores the inter-task relationships that frequently exist among remote sensing images from various sources, leading to reduced classification accuracy. To address these challenges, we propose a learnable Mixture of Adapters (MoA) module in the visual branch of CLIP, resulting in the CLIP-MoA framework. Inspired by the Mixture of Experts (MoE), the MoA module models inter-dataset relationships by integrating multiple adapters through a gating network, while preserving original information via a residual connection between the MoA input and the fine-tuned features. This module enables CLIP-MoA to effectively capture both inter-task commonalities and task-specific characteristics by fine-tuning a minimal set of MoA parameters across multiple datasets simultaneously, thereby reducing computational overhead and fine-tuning time while improving classification accuracy. Comprehensive experiments conducted on a multi-task benchmark involving 11 remote sensing image classification datasets demonstrate that CLIP-MoA achieves significant improvements in both classification accuracy and generalization capability compared to zero-shot CLIP, linear probe CLIP, CLIP-Adapter, and CoOp.
Zhongzheng Fu, Hongping Yan, Kun Ding 0001
IEEE Trans. Geosci. Remote. Sens.3
2024 Weak Distribution Detectors Lead to Stronger Generalizability of Vision-Language Prompt Tuning
abstract
We propose a generalized method for boosting the generalization ability of pre-trained vision-language models (VLMs) while fine-tuning on downstream few-shot tasks. The idea is realized by exploiting out-of-distribution (OOD) detection to predict whether a sample belongs to a base distribution or a novel distribution and then using the score generated by a dedicated competition based scoring function to fuse the zero-shot and few-shot classifier. The fused classifier is dynamic, which will bias towards the zero-shot classifier if a sample is more likely from the distribution pre-trained on, leading to improved base-to-novel generalization ability. Our method is performed only in test stage, which is applicable to boost existing methods without time-consuming re-training. Extensive experiments show that even weak distribution detectors can still improve VLMs' generalization ability. Specifically, with the help of OOD detectors, the harmonic mean of CoOp and ProGrad increase by 2.6 and 1.5 percentage points over 11 recognition datasets in the base-to-novel setting.
Kun Ding 0001, Haojian Zhang, Ying Wang 0008, Shiming Xiang, Chunhong Pan
AAAI1
2024 Zero-shot Generalizable Incremental Learning for Vision-Language Object Detection
abstract
This paper presents Incremental Vision-Language Object Detection (IVLOD), a novel learning task designed to incrementally adapt pre-trained Vision-Language Object Detection Models (VLODMs) to various specialized domains, while simultaneously preserving their zero-shot generalization capabilities for the generalized domain. To address this new challenge, we present the Zero-interference Reparameterizable Adaptation (ZiRa), a novel method that introduces Zero-interference Loss and reparameterization techniques to tackle IVLOD without incurring a significant increase in memory usage. Comprehensive experiments on COCO and ODinW-13 datasets demonstrate that ZiRa effectively safeguards the zero-shot generalization ability of VLODMs while continuously adapting to new tasks. Specifically, after training on ODinW-13 datasets, ZiRa exhibits superior performance compared to CL-DETR and iDETR, boosting zero-shot generalizability by substantial $\textbf{13.91}$ and $\textbf{8.74}$ AP, respectively. Our code is available at https://github.com/JarintotionDin/ZiRaGroundingDINO.
Jieren Deng, Haojian Zhang, Kun Ding 0001, Xingxuan Zhang, Yunkuan Wang
NeurIPS3
2024 Deep convolutional neural network based on self-distillation for tool wear recognition
Yi Pan 0005, Ling Hao, Jianliang He, Kun Ding 0001, Yulin Wang 0004
Eng. Appl. Artif. Intell.4
2024 Compositional Kronecker Context Optimization for vision-language models
Kun Ding 0001, Ying Wang 0008, Haojian Zhang, Shiming Xiang
Neurocomputing1
2024 Multi-task prompt tuning with soft context sharing for vision-language models
Kun Ding 0001, Ying Wang 0008, Pengzhang Liu, Haojian Zhang, Shiming Xiang, Chunhong Pan
Neurocomputing1
2022 Train in Dense and Test in Sparse: A Method for Sparse Object Detection in Aerial Images
abstract
Applications of aerial imaging, especially based on unmanned aerial vehicles (UAVs) platform, rapidly explode in recent years. Meanwhile, vision-based sensing, e.g., detection and recognition, for UAVs becomes increasingly important. Objects in aerial images are usually of tiny size, hence occupying a limited area. Terminology speaking, the images are very sparse in spatial. However, existing work in aerial object detection commonly ignores this point. Conversely, we explore the availability of such a property in improving the detection performance of aerial images. Specifically, we propose a general method, train in dense and test in sparse (TDTS), to exploit sparsity in aerial object detection: 1) in the training stage, the possible positions of object are learned by training a fully convolutional network (called prophet head) and 2) in the testing stage, prophet head identifies the possible object locations to reduce redundant computation in classification and box prediction head by sparse convolution. By extensive experiments on the VisDrone2019-Det data set, we find that the sparsity can not only help to speed up inference but also to improve accuracy. Thus, we argue that the sparsity deserves more attention.
Kun Ding 0001, Guojin He, Huxiang Gu, Zisha Zhong, Shiming Xiang, Chunhong Pan
IEEE Geosci. Remote. Sens. Lett.1
2020 PackDet: Packed Long-Head Object Detector
Kun Ding 0001, Guojin He, Huxiang Gu, Zisha Zhong, Shiming Xiang, Chunhong Pan
ECCV (13)1
2019 Nonlinear Asymmetric Multi-Valued Hashing
abstract
Most existing hashing methods resort to binary codes for large scale similarity search, owing to the high efficiency of computation and storage. However, binary codes lack enough capability in similarity preservation, resulting in less desirable performance. To address this issue, we propose Nonlinear Asymmetric Multi-Valued Hashing (NAMVH) supported by two distinct non-binary embeddings. Specifically, a real-valued embedding is used for representing the newly-coming query by an ideally nonlinear transformation. Besides, a multi-integer-embedding is employed for compressing the whole database, which is modeled by Binary Sparse Representation (BSR) with fixed sparsity. With these two non-binary embeddings, NAMVH preserves more precise similarities between data points and enables access to the incremental extension with database samples evolving dynamically. To perform meaningful asymmetric similarity computation for efficient semantic search, these embeddings are jointly learnt by preserving the pairwise label-based similarity. Technically, this results in a mixed integer programming problem, which is efficiently solved by a well-designed alternative optimization method. Extensive experiments on seven large scale datasets demonstrate that our approach not only outperforms the existing binary hashing methods in search accuracy, but also retains their query and storage efficiency.
Cheng Da, Gaofeng Meng, Shiming Xiang, Kun Ding 0001, Shibiao Xu, Qing Yang 0002, Chunhong Pan
IEEE Trans. Pattern Anal. Mach. Intell.4
2019 Dense semantic embedding network for image captioning
Xinyu Xiao, Lingfeng Wang 0002, Kun Ding 0001, Shiming Xiang, Chunhong Pan
Pattern Recognit.3
2019 Deep Hierarchical Encoder-Decoder Network for Image Captioning
abstract
Encoder-decoder models have been widely used in image captioning, and most of them are designed via single long short term memory (LSTM). The capacity of single-layer network, whose encoder and decoder are integrated together, is limited for such a complex task of image captioning. Moreover, how to effectively increase the “vertical depth” of encoder-decoder remains to be solved. To deal with these problems, a novel deep hierarchical encoder-decoder network is proposed for image captioning, where a deep hierarchical structure is explored to separate the functions of encoder and decoder. This model is capable of efficiently exerting the representation capacity of deep networks to fuse high level semantics of vision and language in generating captions. Specifically, visual representations in top levels of abstraction are simultaneously considered, and each of these levels is associated to one LSTM. The bottom-most LSTM is applied as the encoder of textual inputs. The application of the middle layer in encoder-decoder is to enhance the decoding ability of top-most LSTM. Furthermore, depending on the introduction of semantic enhancement module of image feature and distribution combine module of text feature, variants of architectures of our model are constructed to explore the impacts and mutual interactions among the visual representation, textual representations, and the output of the middle LSTM layer. Particularly, the framework is training under a reinforcement learning method to address the exposure bias problem between the training and the testing by the policy gradient optimization. Qualitative analyses indicate the process that our model “translates” image to sentence and further visualization presents the evolution of the hidden states from different hierarchical LSTMs over time. Extensive experiments demonstrate that our model outperforms current state-of-the-art models on three benchmark datasets: Flickr8K, Flickr30K, and MSCOCO. On both image captioning and retrieval tasks, our method achieves the best results. On MSCOCO captioning Leaderboard, our method also achieves superior performance.
Xinyu Xiao, Lingfeng Wang 0002, Kun Ding 0001, Shiming Xiang, Chunhong Pan
IEEE Trans. Multim.3
2018 In Defense of Locality-Sensitive Hashing
abstract
Hashing-based semantic similarity search is becoming increasingly important for building large-scale content-based retrieval system. The state-of-the-art supervised hashing techniques use flexible two-step strategy to learn hash functions. The first step learns binary codes for training data by solving binary optimization problems with millions of variables, thus usually requiring intensive computations. Despite simplicity and efficiency, locality-sensitive hashing (LSH) has never been recognized as a good way to generate such codes due to its poor performance in traditional approximate neighbor search. We claim in this paper that the true merit of LSH lies in transforming the semantic labels to obtain the binary codes, resulting in an effective and efficient two-step hashing framework. Specifically, we developed the locality-sensitive two-step hashing (LS-TSH) that generates the binary codes through LSH rather than any complex optimization technique. Theoretically, with proper assumption, LS-TSH is actually a useful LSH scheme, so that it preserves the label-based semantic similarity and possesses sublinear query complexity for hash lookup. Experimentally, LS-TSH could obtain comparable retrieval accuracy with state of the arts with two to three orders of magnitudes faster training speed.
Kun Ding 0001, Chunlei Huo, Bin Fan 0001, Shiming Xiang, Chunhong Pan
IEEE Trans. Neural Networks Learn. Syst.1
2017 AMVH: Asymmetric Multi-Valued hashing
abstract
Most existing hashing methods resort to binary codes for similarity search, owing to the high efficiency of computation and storage. However, binary codes lack enough capability in similarity preservation, resulting in less desirable performance. To address this issue, we propose an asymmetric multi-valued hashing method supported by two different non-binary embeddings. (1) A real-valued embedding is used for representing the newly-coming query. (2) A multi-integer-embedding is employed for compressing the whole database, which is modeled by binary sparse representation with fixed sparsity. With these two non-binary embeddings, the similarities between data points can be preserved precisely. To perform meaningful asymmetric similarity computation for efficient semantic search, these embeddings are jointly learnt by preserving the label-based similarity. Technically, this results in a mixed integer programming problem, which is efficiently solved by alternative optimization. Extensive experiments on three multilabel datasets demonstrate that our approach not only outperforms the existing binary hashing methods in search accuracy, but also retains their query and storage efficiency.
Cheng Da, Shibiao Xu, Kun Ding 0001, Gaofeng Meng, Shiming Xiang, Chunhong Pan
CVPR3
2017 Cross-Modal Hashing via Rank-Order Preserving
abstract
Due to the query effectiveness and efficiency, cross-modal similarity search based on hashing has acquired extensive attention in the multimedia community. Most existing methods do not explicitly employ the ranking information when learning hash functions, which is quite important for building practical retrieval systems. To solve this issue, this paper proposes a rank-order preserving hashing (RoPH) method with a novel regression-based rank-order preserving loss that has provable large margin property and is easy to optimize. Moreover, we jointly learn the binary codes and hash functions instead of using any relaxation trick. To solve the induced optimization problem, the alternating descent technique is adopted and each subproblem can be solved conveniently. Specifically, we show that the involved binary quadratic programming subproblem with respect to an introduced auxiliary binary variable satisfies submodularity, enabling us to use the off-the-shelf graph-cut algorithms to solve it exactly and efficiently. Extensive experiments on three benchmarks demonstrate that RoPH significantly improves the ranking quality over the state of the arts.
Kun Ding 0001, Bin Fan 0001, Chunlei Huo, Shiming Xiang, Chunhong Pan
IEEE Trans. Multim.1
2016 Efficient Multiple Feature Fusion With Hashing for Hyperspectral Imagery Classification: A Comparative Study
abstract
Due to the complementary properties of different features, multiple feature fusion has a large potential for hyperspectral imagery classification. At the meantime, hashing is promising in representing a high-dimensional float-type feature with extremely low bit binary codes while maintaining the performance. In this paper, we study the possibility of using hashing to fuse multiple features for hyperspectral imagery classification. For this purpose, we propose a multiple feature fusion framework to evaluate the performance of using different hashing methods. For comparison and completeness, we also have an extensive comparison to five subspace-based dimension reduction methods and six fusion-based methods which are popular solutions to deal with multiple features in hyperspectral image classification. Experimental results on four benchmark hyperspectral data sets demonstrate that using hashing to fuse multiple features can achieve comparable or better performance with the traditional subspace-based dimension reduction methods and fusion-based methods. Moreover, the binary features obtained by using hashing need much less storage and are faster to compute distances with the help of machine instructions.
Zisha Zhong, Bin Fan 0001, Kun Ding 0001, Haichang Li, Shiming Xiang, Chunhong Pan
IEEE Trans. Geosci. Remote. Sens.3
2015 kNN Hashing with Factorized Neighborhood Representation
abstract
Hashing is very effective for many tasks in reducing the processing time and in compressing massive databases. Although lots of approaches have been developed to learn data-dependent hash functions in recent years, how to learn hash functions to yield good performance with acceptable computational and memory cost is still a challenging problem. Based on the observation that retrieval precision is highly related to the kNN classification accuracy, this paper proposes a novel kNN-based supervised hashing method, which learns hash functions by directly maximizing the kNN accuracy of the Hamming-embedded training data. To make it scalable well to large problem, we propose a factorized neighborhood representation to parsimoniously model the neighborhood relationships inherent in training data. Considering that real-world data are often linearly inseparable, we further kernelize this basic model to improve its performance. As a result, the proposed method is able to learn accurate hashing functions with tolerable computation and storage cost. Experiments on four benchmarks demonstrate that our method outperforms the state-of-the-arts.
Kun Ding 0001, Chunlei Huo, Bin Fan 0001, Chunhong Pan
ICCV1
2015 Sparse Hierarchical Clustering for VHR Image Change Detection
abstract
The traditional clustering approaches are limited for the unsupervised change detection of very high resolution images due to the multimodal distribution of change features. To overcome this difficulty, a sparse hierarchical clustering approach is proposed. Discriminative change features are generated by stacking bitemporal multiscale center-symmetric local binary pattern features. In order to explore the multimodal and hierarchical distribution of the change features, a tree-structured dictionary is learned from the pseudotraining set and the unlabeled data. The sparse reconstruction error, a more robust distance compared to the Euclidean distance, is used to determine the label of each change feature. Comparative experiments demonstrate the effectiveness of the proposed method.
Kun Ding 0001, Chunlei Huo, Zisha Zhong, Chunhong Pan
IEEE Geosci. Remote. Sens. Lett.1
2015 Multicluster Spatial-Spectral Unsupervised Feature Selection for Hyperspectral Image Classification
abstract
A new unsupervised spatial-spectral feature selection method for hyperspectral images has been proposed in this letter. The key idea is to select the features that better preserve the multicluster structure of the multiple spatial-spectral features. Specifically, the multicluster structure information is obtained through spectral clustering utilizing a weighted combination of the multiple features. Then, such information is preserved in a group-sparsity-based robust linear regression model. The features that contribute more in preserving the multicluster structure information are selected. Comparative experiments on two popular real hyperspectral images validate the effectiveness of the proposed method, showing higher classification accuracy.
Haichang Li, Shiming Xiang, Zisha Zhong, Kun Ding 0001, Chunhong Pan
IEEE Geosci. Remote. Sens. Lett.4
2015 Discriminant Tensor Spectral-Spatial Feature Extraction for Hyperspectral Image Classification
abstract
We propose to integrate spectral-spatial feature extraction and tensor discriminant analysis for hyperspectral image classification. First, we apply remarkable spectral-spatial feature extraction approaches in the hyperspectral cube to extract a feature tensor for each pixel. Then, based on class label information, local tensor discriminant analysis is used to remove redundant information for subsequent classification procedure. The approach not only extracts sufficient spectral-spatial features from original hyperspectral images but also gets better feature representation owing to tensor framework. Comparative results on two benchmarks demonstrate the effectiveness of our method.
Zisha Zhong, Bin Fan 0001, Jiangyong Duan, Lingfeng Wang 0002, Kun Ding 0001, Shiming Xiang, Chunhong Pan
IEEE Geosci. Remote. Sens. Lett.5
2013 VHR image change detection based on discriminative dictionary learning
abstract
The difficulty of Very High Resolution (VHR) image change detection is mainly due to the low separability between the changed and unchanged class. The traditional approaches usually address the problem by solving the feature extraction and classification separately, which cannot ensure that the classification algorithm makes the best use of the features. Considering this, we propose a novel approach that combines the feature extraction and the classification task by utilizing the sparse representation algorithm with discriminative dictionary. Experiments on real data sets show that our method achieves effective results.
Kun Ding 0001, Chunlei Huo, Chunhong Pan
ICASSP1