Lu Jin 0001

dblp:28/2680-1 · DBLP profile ↗
← Back
21ranked-venue papers
10as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 17 · 7 first-author · 9 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
YearPublicationVenuePosition
2026 FineG-RAG: Fine-Grained Retrieval-Augmented Generation for Multimodal Large Language Models
abstract
Fine-grained visual recognition refers to the ability to distinguish subtle differences between visually similar objects— a fundamental yet challenging capability for Multimodal Large Language Models (MLLMs). In this paper, we observe that even strong open-source MLLMs, such as Qwen2-VL and InternVL2, still struggle with accurately identifying fine-grained categories. These models often fail to attend to subtle but critical details for precise discrimination. To unlock this potential, we propose FineG-RAG, a retrieval-augmented generation pipeline designed to enhance the fine-grained recognition capabilities of MLLMs. FineG-RAG integrates external fine-grained knowledge into the recognition process via a generalized retriever. To support this, we construct fine-grained visual-language knowledge database containing representative images with wide visual diversity and expert-crafted attribute descriptions from multiple perspectives. Relevant fine-grained knowledge is retrieved from this database and fed into a visual-language augmented prompt, which provides rich multimodal context to guide MLLMs in generating accurate labels. To better evaluate the fine-grained recognition capabilities of MLLMs, we design a multiple-choice evaluation strategy based on publicly four fine-grained datasets. Extensive experiments demonstrate that FineG-RAG consistently outperforms baseline methods, achieving superior recognition accuracy across a range of off-the-shelf, open-source MLLMs.
Lu Jin 0001, Xinguang Xiang, Yanpeng Sun, Zechao Li, Jinhui Tang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2026 Multi-Modal Knowledge Distillation Hashing Based on CLIP for Weakly Supervised Image Retrieval
abstract
Existing weakly supervised hashing often suffers from the imprecision of user-provided tags and over-reliance on textual knowledge from pre-trained word embeddings, neglecting crucial visual knowledge associated with image labels. As a result, this leads to unsatisfactory performance in closed-vocabulary tasks and limited generalization in open-vocabulary scenarios. To address this issue, we propose Multi-modal Knowledge Distillation Hashing (MKDH), a novel method leveraging visual and language pre-training (VLP) model such as CLIP to learn robust hash codes. Our method designs a dual-layer attention adapter to generate joint representations by capturing fine-grained visual and textual knowledge from the CLIP teacher network. Additionally, we introduce a knowledge extraction contrastive loss to enhance the robustness of joint representations and a knowledge distillation contrastive loss to transfer the extracted multi-modal knowledge to the hash codes. To further mitigate the negative impact of false negative pairs in these contrastive losses, we introduce false negative weighting strategy that reduces the weights assigned to such pairs. Extensive experiments on three widely used datasets demonstrate that our method achieves robust retrieval performance with significant improvements in both closed- and open-vocabulary settings. The source code is available athttps://github.com/IMAG-LZY/MKDH.
Zhengyun Lu, Lu Jin 0001, Zechao Li, Jinhui Tang 0001
IEEE Trans. Multim.2
2025 ReDet: Effective Real-time Object Detection via Efficient Multi-scale Extraction Aggregation
abstract
Real-time object detection demands detectors that excel in both speed and accuracy. However, existing methods rely on complex Feature Pyramid Networks and computationally intensive post-processing to boost performance, often struggling to balance efficiency and accuracy. In this paper, we propose ReDet, an efficient real-time end-to-end object detection framework that improves detection capability while maintaining low computational cost. Specifically, we propose a Multi-scale Extraction Aggregation module to enhance feature fusion ability with minimal overhead, boosting feature representation across scales. Additionally, a Regression Enhancement Module is incorporated to mitigate the assignment inconsistency between localization and classification, further enhancing detection accuracy with negligible computational cost. Moreover, we introduce a Multi-label Auxiliary Strategy to eliminate the reliance on post-processing by enabling one-to-one label assignment. The experimental results demonstrate that ReDet-T achieves 39.2% AP and 500 FPS on the COCO val2017 dataset, while ReDet-S achieves 45.3% AP and 294 FPS, outperforming many other detectors in both speed and accuracy.
Xin Jiang 0010, Lu Jin 0001, Zechao Li
ICME3
2025 Causal Inference Hashing for Long-Tailed Image Retrieval
abstract
In hashing-based long-tailed image retrieval, the dominance of data-rich head classes often hinders the learning of effective hash codes for data-poor tail classes due to inherent long-tailed bias. Interestingly, this bias also contains valuable prior knowledge by revealing inter-class dependencies, which can be beneficial for hash learning. However, previous methods have not thoroughly analyzed this tangled negative and positive effects of long-tailed bias from a causal inference perspective. In this paper, we propose a novel hash framework that employs causal inference to disentangle detrimental bias effects from beneficial ones. To capture good bias in long-tailed datasets, we construct hash mediators that conserve valuable prior knowledge from class centers. Furthermore, we propose a de-biased hash loss To enhance the beneficial bias effects while mitigating adverse ones, leading to more discriminative hash codes. Specifically, this loss function leverages the beneficial bias captured by hash mediators to support accurate class label prediction, while mitigating harmful bias by blocking its causal path to the hash codes and refining predictions through backdoor adjustment. Extensive experimental results on four widely used datasets demonstrate that the proposed method improves retrieval performance against the state-of-the-art methods by large margins. The source code is available at https://github.com/IMAG-LuJin/CIH.
Lu Jin 0001, Zhengyun Lu, Zechao Li, Yonghua Pan, Longquan Dai, Jinhui Tang 0001, Ramesh Jain 0001
IEEE Trans. Image Process.1
2025 Relational Consistency Induced Self-Supervised Hashing for Image Retrieval
abstract
This article proposes a new hashing framework named relational consistency induced self-supervised hashing (RCSH) for large-scale image retrieval. To capture the potential semantic structure of data, RCSH explores the relational consistency between data samples in different spaces, which learns reliable data relationships in the latent feature space and then preserves the learned relationships in the Hamming space. The data relationships are uncovered by learning a set of prototypes that group similar data samples in the latent feature space. By uncovering the semantic structure of the data, meaningful data-to-prototype and data-to-data relationships are jointly constructed. The data-to-prototype relationships are captured by constraining the prototype assignments generated from different augmented views of an image to be the same. Meanwhile, these data-to-prototype relationships are preserved to learn informative compact hash codes by matching them with these reliable prototypes. To accomplish this, a novel dual prototype contrastive loss is proposed to maximize the agreement of prototype assignments in the latent feature space and Hamming space. The data-to-data relationships are captured by enforcing the distribution of pairwise similarities in the latent feature space and Hamming space to be consistent, which makes the learned hash codes preserve meaningful similarity relationships. Extensive experimental results on four widely used image retrieval datasets demonstrate that the proposed method significantly outperforms the state-of-the-art methods. Besides, the proposed method achieves promising performance in out-of-domain retrieval tasks, which shows its good generalization ability. The source code and models are available at https://github.com/IMAG-LuJin/RCSH.
Lu Jin 0001, Zechao Li, Yonghua Pan, Jinhui Tang 0001
IEEE Trans. Neural Networks Learn. Syst.1
2024 Context Disentangling and Prototype Inheriting for Robust Visual Grounding
abstract
Visual grounding (VG) aims to locate a specific target in an image based on a given language query. The discriminative information from context is important for distinguishing the target from other objects, particularly for the targets that have the same category as others. However, most previous methods underestimate such information. Moreover, they are usually designed for the standard scene (without any novel object), which limits their generalization to the open-vocabulary scene. In this paper, we propose a novel framework with context disentangling and prototype inheriting for robust visual grounding to handle both scenes. Specifically, the context disentangling disentangles the referent and context features, which achieves better discrimination between them. The prototype inheriting inherits the prototypes discovered from the disentangled visual features by a prototype bank to fully utilize the seen data, especially for the open-vocabulary scene. The fused features, obtained by leveraging Hadamard product on disentangled linguistic and visual features of prototypes to avoid sharp adjusting the importance between the two types of features, are then attached with a special token and feed to a vision Transformer encoder for bounding box regression. Extensive experiments are conducted on both standard and open-vocabulary scenes. The performance comparisons indicate that our method outperforms the state-of-the-art methods in both scenarios.
Wei Tang 0011, Liang Li 0003, Xuejing Liu, Lu Jin 0001, Jinhui Tang 0001, Zechao Li
IEEE Trans. Pattern Anal. Mach. Intell.4
2024 Self-Paced Relational Contrastive Hashing for Large-Scale Image Retrieval
abstract
Supervised deep hashing aims to learn hash functions using label information. Existing methods learn hash functions by employing either pairwise/triplet loss to explore the point-to-point relation or center loss to explore the point-to-class relation. However, these methods overlook the collaboration between the above two kinds of relations and the hardness of pairs. In this work, we propose a novel Self-Paced Relational Contrastive Hashing (SPRCH) method with a single learning objective to capture valuable discriminative information from hard pairs using both the point-to-point and point-to-class relations. To exploit the above two kinds of relations, the Relational Contrastive Hash (RCH) loss is proposed, which ensures that each data anchor is closer to all similar data points and corresponding class centers in the Hamming space compared to dissimilar ones. Moreover, the proposed RCH loss reduces the drastic imbalance between point-to-point pairs and point-to-class pairs by rebalancing their weights. To prioritize hard pairs, a self-paced learning schedule is proposed, assigning higher weights to these pairs in the RCH loss. The self-paced learning schedule assigns dynamic weights to pairs according to their similarities and the training process. In this way, deep hash model can initially learn universal patterns from the entire set of pairs and then gradually acquire more valuable discriminative information from hard pairs. Experimental results on four widely-used image retrieval datasets demonstrate that our proposed SPRCH method significantly outperforms the state-of-the-art supervised deep hash methods.
Zhengyun Lu, Lu Jin 0001, Zechao Li, Jinhui Tang 0001
IEEE Trans. Multim.2
2024 Unbiased Visual Question Answering by Leveraging Instrumental Variable
abstract
Existing unbiased VQA models reduce the spurious correlation between questions and answers to force the models to focus on visual information. However, the visual information captured by these unbiased models is irrelevant to the correct answer, resulting in leveraging spurious correlation to predict incorrect answers. This makes these unbiased methods fail to obtain critical visual information, thus performing poorly on questions dominated by the visual information. To capture the valuable visual information, this paper proposes a novel unbiased VQA model based on causal inference, leveraging Instrumental Variable (IVar) to increase the causal effect between visual features and answers. First, to obtain suitable instrumental variables, the noise generator is proposed according to the constraints of IVar. The generated noise can be regarded as IVar, which is used to pollute the original visual features. Then, this paper proposes IVar loss which utilizes the generated IVar to increase the causal effect between visual features and answers. When the visual feature is polluted by IVar, IVar loss guides the model to predict incorrect answers to enhance the correlation between IVar and the answer. Since the correlation between IVar and the answer is proportional to the causal effect between the visual feature and the answer, IVar loss enhances the importance of the visual information, thereby rectifying the model to capture critical visual information. The extensive experimental results on widely-used benchmarks demonstrate the advantages of the proposed method. The proposed method gains the best accuracy on answer typeOtherof VQA-CP v2. These results demonstrate the superiority of the proposed method in capturing critical visual information since most questions on the answer typeOtherare dominated by visual information.
Yonghua Pan, Jing Liu 0001, Lu Jin 0001, Zechao Li
IEEE Trans. Multim.3
2024 Alleviating Over-Fitting in Hashing-Based Fine-Grained Image Retrieval: From Causal Feature Learning to Binary-Injected Hash Learning
abstract
Hashing-based fine-grained image retrieval pursues learning diverse local features to generate inter-class discriminative hash codes. However, existing fine-grained hash methods with attention mechanisms usually tend to just focus on a few obvious areas, which misguides the network to over-fit some salient features. Such a problem raises two main limitations. 1) It overlooks some subtle local features, degrading the generalization capability of learned embedding. 2) It causes the over-activation of some hash bits correlated to salient features, which breaks the binary code balance and further weakens the discrimination abilities of hash codes. To address these limitations of the over-fitting problem, we propose a novel hash framework fromCausalFeature learning toBinary-injectedHash learning (CFBH), which captures various local information and suppresses over-activated hash bits simultaneously. For causal feature learning, we adopt causal inference theory to alleviate the bias towards the salient regions in fine-grained images. In detail, we obtain local features from the feature map and combine this local information with original image information followed by this theory. Theoretically, these fused embeddings help the network to re-weight the retrieval effort of each local feature and exploit more subtle variations without observational bias. For binary-injected hash learning, we propose a Binary Noise Injection (BNI) module inspired by Dropout. The BNI module not only mitigates over-activation to particular bits, but also makes hash codes uncorrelated and balanced in the Hamming space. Extensive experimental results on six popular fine-grained image datasets demonstrate the superiority of CFBH over several State-of-the-Art methods.
Xinguang Xiang, Xinhao Ding, Lu Jin 0001, Zechao Li, Jinhui Tang 0001, Ramesh Jain 0001
IEEE Trans. Multim.3
2023 Deep Semantic Multimodal Hashing Network for Scalable Image-Text and Video-Text Retrievals
abstract
Hashing has been widely applied to multimodal retrieval on large-scale multimedia data due to its efficiency in computation and storage. In this article, we propose a novel deep semantic multimodal hashing network (DSMHN) for scalable image-text and video-text retrieval. The proposed deep hashing framework leverages 2-D convolutional neural networks (CNN) as the backbone network to capture the spatial information for image-text retrieval, while the 3-D CNN as the backbone network to capture the spatial and temporal information for video-text retrieval. In the DSMHN, two sets of modality-specific hash functions are jointly learned by explicitly preserving both intermodality similarities and intramodality semantic labels. Specifically, with the assumption that the learned hash codes should be optimal for the classification task, two stream networks are jointly trained to learn the hash functions by embedding the semantic labels on the resultant hash codes. Moreover, a unified deep multimodal hashing framework is proposed to learn compact and high-quality hash codes by exploiting the feature representation learning, intermodality similarity-preserving learning, semantic label-preserving learning, and hash function learning with different types of loss functions simultaneously. The proposed DSMHN method is a generic and scalable deep hashing framework for both image-text and video-text retrievals, which can be flexibly integrated with different types of loss functions. We conduct extensive experiments for both single-modal- and cross-modal-retrieval tasks on four widely used multimodal-retrieval data sets. Experimental results on both image-text- and video-text-retrieval tasks demonstrate that the DSMHN significantly outperforms the state-of-the-art methods.
Lu Jin 0001, Zechao Li, Jinhui Tang 0001
IEEE Trans. Neural Networks Learn. Syst.1
2022 Sub-Region Localized Hashing for Fine-Grained Image Retrieval
abstract
Fine-grained image hashing is challenging due to the difficulties of capturing discriminative local information to generate hash codes. On the one hand, existing methods usually extract local features with the dense attention mechanism by focusing on dense local regions, which cannot contain diverse local information for fine-grained hashing. On the other hand, hash codes of the same class suffer from large intra-class variation of fine-grained images. To address the above problems, this work proposes a novel sub-Region Localized Hashing (sRLH) to learn intra-class compact and inter-class separable hash codes that also contain diverse subtle local information for efficient fine-grained image retrieval. Specifically, to localize diverse local regions, a sub-region localization module is developed to learn discriminative local features by locating the peaks of non-overlap sub-regions in the feature map. Different from localizing dense local regions, these peaks can guide the sub-region localization module to capture multifarious local discriminative information by paying close attention to dispersive local regions. To mitigate intra-class variations, hash codes of the same class are enforced to approach one common binary center. Meanwhile, the gram-schmidt orthogonalization is performed on the binary centers to make the hash codes inter-class separable. Extensive experimental results on four widely used fine-grained image retrieval datasets demonstrate the superiority of sRLH to several state-of-the-art methods. The source code of sRLH will be released at https://github.com/ZhangYajie-NJUST/sRLH.git.
Xinguang Xiang, Lu Jin 0001, Zechao Li, Jinhui Tang 0001
IEEE Trans. Image Process.3
2021 Self-Adaptive Hashing for Fine-Grained Image Retrieval
abstract
The main challenge of fine-grained image hashing is how to learn highly discriminative hash codes to distinguish the within and between class variations. On the one hand, most of the existing methods treat sample pairs as equivalent in hash learning, ignoring the more discriminative information contained in hard sample pairs. On the other hand, in the testing phase, these methods ignore the influence of outliers on retrieval performance. In order to solve the above issues, this paper proposes a novel Self-Adaptive Hashing method, which learns discriminative hash codes by mining hard sample pairs, and improves retrieval performance by correcting outliers in the testing phase. In particular, to improve the discriminability of hash codes, a pair-weighted based loss function is proposed to enhance the learning of hash functions of hard sample pairs. Furthermore, in the testing phase, a self-adaptive module is proposed to discover and correct outliers by generating self-adaptive boundaries, thereby improving the retrieval performance. Experimental results on two widely-used fine-grained datasets demonstrate the effectiveness of the proposed method.
Yuxuan Dai, Wei Tang 0011, Lu Jin 0001, Xinguang Xiang
MMAsia4
2020 Weakly-Supervised Image Hashing through Masked Visual-Semantic Graph-based Reasoning
abstract
With the popularization of social websites, many methods have been proposed to explore the noisy tags for weakly-supervised image hashing.The main challenge lies in learning appropriate and sufficient information from those noisy tags. To address this issue, this work proposes a novel Masked visual-semantic Graph-based Reasoning Network, termed as MGRN, to learn joint visual-semantic representations for image hashing. Specifically, for each image, MGRN constructs a relation graph to capture the interactions among its associated tags and performs reasoning with Graph Attention Networks (GAT). MGRN randomly masks out one tag and then make GAT to predict this masked tag. This forces the GAT model to capture the dependence between the image and its associated tags, which can well address the problem of noisy tags. Thus it can capture key tags and visual structures from images to learn well-aligned visual-semantic representations. Finally, the auto-encoders is leveraged to learn hash codes that can preserve the local structure of the joint space. Meanwhile, the joint visual-semantic representations are reconstructed from those hash codes by using a decoder. Experimental results on two widely-used benchmark datasets demonstrate the superiority of the proposed method for image retrieval compared with several state-of-the-art methods.
Lu Jin 0001, Zechao Li, Yonghua Pan, Jinhui Tang 0001
ACM Multimedia1
2019 Attention-Aware Feature Pyramid Ordinal Hashing for Image Retrieval
abstract
Due to the effectiveness of representation learning, deep hashing methods have attracted increasing attention in image retrieval. However, most existing deep hashing methods merely encode the raw information of the last layer for hash learning, which result in the following deficiencies: (1) the useful information from the preceding-layer is not fully exploited; (2) the local salient information of the image is neglected. To this end, we propose a novel deep hashing method, called Attention-Aware Feature Pyramid Ordinal Hashing (AFPH), which explores both the visual structure information and semantic information from different convolutional layers. Specifically, two feature pyramids based on spatial and channel attention are well constructed to capture the local salient structure from multiple scales. Moreover, a multi-scale feature fusion strategy is proposed to aggregate the feature maps from multi-level pyramidal layers to generate the discriminative feature for ranking-based hashing. The experimental results conducted on two widely-used image retrieval datasets demonstrate the superiority of our method.
Xie Sun, Lu Jin 0001, Zechao Li
MMAsia2
2019 Multimedia retrieval by deep hashing with multilevel similarity learning
Qiuli Liu, Lu Jin 0001, Zechao Li, Jinhui Tang 0001
J. Vis. Commun. Image Represent.2
2019 Deep Ordinal Hashing With Spatial Attention
abstract
Hashing has attracted increasing research attention in recent years due to its high efficiency of computation and storage in image retrieval. Recent works have demonstrated the superiority of simultaneous feature representations and hash functions learning with deep neural networks. However, most existing deep hashing methods directly learn the hash functions by encoding the global semantic information, while ignoring the local spatial information of images. The loss of local spatial structure makes the performance bottleneck of hash functions, therefore limiting its application for accurate similarity retrieval. In this paper, we propose a novel deep ordinal hashing (DOH) method, which learns ordinal representations to generate ranking-based hash codes by leveraging the ranking structure of feature space from both local and global views. In particular, to effectively build the ranking structure, we propose to learn the rank correlation space by exploiting the local spatial information from fully convolutional network and the global semantic information from the convolutional neural network simultaneously. More specifically, an effective spatial attention model is designed to capture the local spatial information by selectively learning well-specified locations closely related to target objects. In such hashing framework, the local spatial and global semantic nature of images is captured in an end-to-end ranking-to-hashing manner. Experimental results conducted on three widely used datasets demonstrate that the proposed DOH method significantly outperforms the state-of-the-art hashing methods.
Lu Jin 0001, Xiangbo Shu, Kai Li 0005, Zechao Li, Guo-Jun Qi, Jinhui Tang 0001
IEEE Trans. Image Process.1
2019 Deep Semantic-Preserving Ordinal Hashing for Cross-Modal Similarity Search
abstract
Cross-modal hashing has attracted increasing research attention due to its efficiency for large-scale multimedia retrieval. With simultaneous feature representation and hash function learning, deep cross-modal hashing (DCMH) methods have shown superior performance. However, most existing methods on DCMH adopt binary quantization functions (e.g., [Formula: see text]) to generate hash codes, which limit the retrieval performance since binary quantization functions are sensitive to the variations of numeric values. Toward this end, we propose a novel end-to-end ranking-based hashing framework, in this paper, termed as deep semantic-preserving ordinal hashing (DSPOH), to learn hash functions with deep neural networks by exploring the ranking structure of feature dimensions. In DSPOH, the ordinal representation, which encodes the relative rank ordering of feature dimensions, is explored to generate hash codes. Such ordinal embedding benefits from the numeric stability of rank correlation measures. To make the hash codes discriminative, the ordinal representation is expected to well predict the class labels so that the ranking-based hash function learning is optimally compatible with the label predicting. Meanwhile, the intermodality similarity is preserved to guarantee that the hash codes of different modalities are consistent. Importantly, DSPOH can be effectively integrated with different types of network architectures, which demonstrates the flexibility and scalability of our proposed hashing framework. Extensive experiments on three widely used multimodal data sets show that DSPOH outperforms state of the art for cross-modal retrieval tasks.
Lu Jin 0001, Kai Li 0005, Zechao Li, Fu Xiao 0001, Guo-Jun Qi, Jinhui Tang 0001
IEEE Trans. Neural Networks Learn. Syst.1
2018 Semantic Neighbor Graph Hashing for Multimodal Retrieval
abstract
Hashing methods have been widely used for approximate nearest neighbor search in recent years due to its computational and storage effectiveness. Most existing multimodal hashing methods try to preserve the similarity relationship based on either metric distances or semantic labels in a procrustean way, while ignoring the intra-class and inter-class variations inherent in the metric space. In this paper, we propose a novel multimodal hashing method, termed as semantic neighbor graph hashing (SNGH), which aims to preserve the fine-grained similarity metric based on the semantic graph that is constructed by jointly pursuing the semantic supervision and the local neighborhood structure. Specifically, the semantic graph is constructed to capture the local similarity structure for the image modality and the text modality, respectively. Furthermore, we define a function based on the local similarity in particular to adaptively calculate multi-level similarities by encoding the intra-class and inter-class variations. After obtaining the unified hash codes, the logistic regression with kernel trick is employed to learn view-specific hash functions independently for each modality. Extensive experiments are conducted on four widely used multimodal data sets. The experimental results demonstrate the superiority of the proposed SNGH method compared with the state-of-the-art multimodal hashing methods.
Lu Jin 0001, Kai Li 0005, Hao Hu 0010, Guo-Jun Qi, Jinhui Tang 0001
IEEE Trans. Image Process.1
2015 Partially Common-Semantic Pursuit for RGB-D Object Recognition
abstract
For the RGB-D object recognition task, the robust and rich representations can boost the performance. Most works employ feature learning approaches to learn specific representation for the RGB and depth modalities independently, while some directly learn common property. Different from them, this paper proposes a novel supervised feature learning method for RGB-D object recognition, named Partially Common-Semantic Learning (PCSL), which jointly captures the complementary and consistency semantic information from RGB and depth modalities. The complementary information is revealed by the individual modality, while the consistency is exploited by both modalities simultaneously. In PCSL, Reconstruction Independent Component Analysis (RICA) is extended to integrate the supervised information and learn both of the complementary and partially shared common semantic information. The proposed approach is evaluated on two public RGB-D datasets and achieves better performance than several state-of-the-art methods.
Lu Jin 0001, Zechao Li, Xiangbo Shu, Shenghua Gao, Jinhui Tang 0001
ACM Multimedia1
2015 RGB-D Object Recognition via Incorporating Latent Data Structure and Prior Knowledge
abstract
For the task of RGB-D object recognition, it is important to identify suitable representations of images, which can boost the performance of object recognition. In this work, we propose a novel representation learning method for RGB-D images by jointly incorporating the underlying data structure and the prior knowledge of the data. Specifically, the convolutional neural networks (CNN) are employed to learn image representation by exploiting the underlying data structure. To handle the problem of the limited RGB and depth images for object recognition, the multi-level hierarchies of features trained on ImageNet from the CNN are transferred to learn rich generic feature representation for RGB and depth images while the labeled images are leveraged. On the other hand, we propose a novel deep auto-encoders (DAE) to exploit the prior knowledge, which can overcome the expensive computational cost of optimization in feature encoding. The expected representations of images are obtained by integrating the two types of image representations. To verify the effectiveness of the proposed method, we thoroughly conduct extensive experiments on two publicly available RGB-D datasets. The encouraging experimental results compared with the state-of-the-art approaches demonstrate the advantages of the proposed method.
Jinhui Tang 0001, Lu Jin 0001, Zechao Li, Shenghua Gao
IEEE Trans. Multim.2
2014 Hand-Crafted Features or Machine Learnt Features? Together They Improve RGB-D Object Recognition
abstract
RGB-D object recognition is an important research topic in computer version, and seeking a robust image representation is the most important sub problem for RGB-D object recognition. On the one hand, the recently emerging deep learning methods, which learns image representations automatically by capturing the data structure, have demonstrated the impressive performance for object recognition. On the other hand, the previously commonly used hand-crafted features also encodes the prior knowledge about the data. By realizing that the hand-crafted features and machine learnt features actually characterize the different aspects of image data, rather than only using one type of feature, we propose to jointly use the machine learnt features and hand-crafted features for RGB-D object recognition. Specifically, we use the Convolution Neural Networks (CNNs) to extract the machine learnt representation, and use Locality-constrained Linear Coding (LLC) based spatial pyramid matching for hand-crafted features. We evaluated our proposed approach on three publicly available RGB-D datasets. Experimental results show that our method achieves the best performance under all the cases, which demonstrates the effectiveness of our method.
Lu Jin 0001, Shenghua Gao, Zechao Li, Jinhui Tang 0001
ISM1