Jiachen Li 0002

dblp:137/8316-2 · DBLP profile ↗
← Back
25ranked-venue papers
4as first author
24since 2021 · last 2026
0000-0002-0602-9360ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 15 · 3 first-author · 15 since 2021Databases, data management, data science and information retrieval · 9 · 9 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 6 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2026 A Sketch+Text Composed Image Retrieval Dataset for Thangka
abstract
Composed Image Retrieval (CIR) enables image retrieval by combining multiple query modalities, but existing benchmarks predominantly focus on general-domain imagery and rely on reference images with short textual modifications. As a result, they provide limited support for retrieval scenarios that require fine-grained semantic reasoning, structured visual understanding, and domain-specific knowledge. In this work, we introduce CIRThan, a sketch+text composed image retrieval dataset for Thangka imagery, a culturally grounded and knowledge-specific visual domain characterized by complex structures, dense symbolic elements, and domain-dependent semantic conventions. CIRThan contains 2,287 high-quality Thangka images, each paired with a human-drawn sketch and hierarchical textual descriptions at three semantic levels, enabling composed queries that jointly express structural intent and multi-level semantic specification. We provide standardized data splits, comprehensive dataset analysis, and benchmark evaluations of representative supervised and zero-shot CIR methods. Experimental results reveal that existing CIR approaches, largely developed for general-domain imagery, struggle to effectively align sketch-based abstractions and hierarchical textual semantics with fine-grained Thangka images, particularly without in-domain supervision. We believe CIRThan offers a valuable benchmark for advancing sketch+text CIR, hierarchical semantic modeling, and multimodal retrieval in cultural heritage and other knowledge-specific visual domains. The dataset is publicly available at https://github.com/jinyuxu-whut/CIRThan.
Jinyu Xu 0001, Jiangling Zhang, Qing Xie 0002, Daomin Ji, Zhifeng Bao, Jiachen Li 0002, Yanchun Ma, Yongjian Liu
SIGIR7
2026 SDR-CIR: Semantic Debias Retrieval Framework for Training-Free Zero-Shot Composed Image Retrieval
abstract
Composed Image Retrieval (CIR) aims to retrieve a target image from a query composed of a reference image and modification text. Recent training-free zero-shot methods often employ Multimodal Large Language Models (MLLMs) with Chain-of-Thought (CoT) to compose a target image description for retrieval. However, due to the fuzzy matching nature of ZS-CIR, the generated description is prone to semantic bias relative to the target image. We propose SDR-CIR, a training-free Semantic Debias Ranking method based on CoT reasoning. First, Selective CoT guides the MLLM to extract visual content relevant to the modification text during image understanding, thereby reducing visual noise at the source. We then introduce a Semantic Debias Ranking with two steps, Anchor and Debias, to mitigate semantic bias. In the Anchor step, we fuse reference image features with target description features to reinforce useful semantics and supplement omitted cues. In the Debias step, we explicitly model the visual semantic contribution of the reference image to the description and incorporate it into the similarity score as a penalty term. By supplementing omitted cues while suppressing redundancy, SDR-CIR mitigates semantic bias and improves retrieval performance. Experiments on three standard CIR benchmarks show that SDR-CIR achieves state-of-the-art results among one-stage methods while maintaining high efficiency. The code is publicly available at https://github.com/suny105/SDR-CIR.
Jinyu Xu 0001, Qing Xie 0002, Jiachen Li 0002, Yanchun Ma, Yongjian Liu
WWW4
2026 Guided by Principles of Composition: A Domain-Specific Priors Based Detector for Recognizing Ritual Implements in Thangka
abstract
ABSTRACT Detecting ritual implements in Thangka paintings—such as swords and scriptures—remains challenging due to their intricate visual composition and symbolic complexity. Existing object detection models, typically trained on natural scenes, tend to perform poorly in this domain. To address this limitation, we summarize the principles of composition in Thangka and identify key spatial and co‐occurrence priors specific to ritual implements. Based on these insights, we propose GPCDet: a guided by principles of composition detector that integrates domain‐specific priors into the detection process. Specifically, we introduce a spatial coordinate attention module to emphasize critical spatial regions where implements frequently appear. In addition, we design a graph convolution network‐auxiliary detection module to model inter‐category co‐occurrence, thereby enhancing feature representation and improving classification performance. Experiments on the newly curated ritual implements in Thangka (RITK) dataset show that GPCDet achieves substantial improvements over existing methods, establishing a new state‐of‐the‐art baseline for this challenging task.
Jiachen Li 0002, Hongyun Wang, Xiaolong Peng, Jinyu Xu 0001, Qing Xie 0002, Yanchun Ma, Wenbo Jiang 0001, Mengzi Tang
IET Image Process.1
2026 TrojanEdit: Multimodal backdoor attack against image editing model
Ji Guo, Runjia Zhang, Wenbo Jiang 0001, Yiting Zhu, Jiachen Li 0002, Jiaming He, Hongwei Li 0001
Neurocomputing6
2026 LGD: Leveraging generative descriptions for zero-shot referring image segmentation
abstract
Zero-shot referring image segmentation aims to locate and segment the target region based on a referring expression, with the primary challenge of aligning and matching semantics across visual and textual modalities without training. Previous works address this challenge by utilizing Vision-Language Models and mask proposal networks for region-text matching. However, this paradigm may lead to incorrect target localization due to the inherent ambiguity and diversity of free-form referring expressions. To alleviate this issue, we present LGD (Leveraging Generative Descriptions), a framework that utilizes the advanced language generation capabilities of Multi-Modal Large Language Models to enhance region-text matching performance in Vision-Language Models. Specifically, we first design two kinds of prompts, the attribute prompt and the surrounding prompt, to guide the Multi-Modal Large Language Models in generating descriptions related to the crucial attributes of the referent object and the details of surrounding objects, referred to as attribute description and surrounding description, respectively. Secondly, three visual-text matching scores are introduced to evaluate the similarity between instance-level visual features and textual features, which determines the mask most associated with the referring expression. The proposed method achieves new state-of-the-art performance on three public datasets RefCOCO, RefCOCO+ and RefCOCOg, with maximum improvements of 9.97 % in oIoU and 11.29 % in mIoU compared to previous methods
Jiachen Li 0002, Qing Xie 0002, Renshu Gu, Jinyu Xu 0001, Yongjian Liu, Xiaohan Yu 0001
Pattern Recognit.1
2026 Enhancing Fine-Grained Sketch-Based Image Retrieval Through Contextual Information
abstract
Fine-grained sketch-based image retrieval aims to retrieve matched natural photos at instance level based on hand-drawn sketches. Understanding sketches and aligning them with photos in the joint embedding space are the primary challenges. Previous works have attempted to improve the understanding and alignment of fine-grained semantics by leveraging a selected subset of local features based on the local matching strategy. However, these approaches rely heavily on local features to capture regional visual details, which introduces unavoidable uncertainty and noise in the matching process. In addition, existing works lack the ability to effectively separate unmatched pairs with similar overall semantics in the common space, which may disrupt fine-grained semantic alignment. In this paper, we presentCIGSPA: aContextual Information Guided Sketch-Photo Alignmentframework that leverages latent contextual information to enhance sketch understanding and sketch-photo matching without additional knowledge and prior supervision. Specifically, we first design anAdaptive Local Feature Filtermodule to refine feature similarity based on the context of local features and select more appropriate features, which ensure the semantic fidelity and spatial compactness of salient features. Secondly, we design aContextual Weighted Contrastiveloss to guide global sketch-photo matching by leveraging discriminative contextual information from photos, and hence aligns the sketch and photo at the instance level. Our comprehensive experiments on six public datasets demonstrate the effectiveness of our proposed CIGSPA. The code is available athttps://github.com/jinyuxu-whut/CIGSPA4FG-SBIR.
Jinyu Xu 0001, Qing Xie 0002, Jiachen Li 0002, Zhifeng Bao, Yanchun Ma, Yongjian Liu
IEEE Trans. Multim.3
2026 Relation-Aware Proxy Hashing for Cross-Modal Retrieval
abstract
Proxy hashing methods have attracted increasing attention in cross-modal retrieval, because they are able to learn the mapping of different modalities into a common low-dimensional hash space by leveraging global proxies. However, existing approaches typically suffer from two limitations: (1) They solely utilize prior labels to capture the global semantic information, and hence lack the ability to explore necessary fine-grained semantic information to bridge the modality gap effectively. (2) They seldom consider the guidance information of intrinsic semantic similarity on proxy-centered space, and thus fail to leverage the similarity relations among instances sufficiently. To mitigate these limitations, we propose RAPH, a novel Relation-Aware Proxy Hashing framework that learns the semantic relations between different modalities and different semantic levels to enhance the discriminative capability of hash codes. Specifically, we first propose a Local Semantic Interaction (LSI) module based on masked language modeling to achieve the interaction of multi-modal fine-grained semantic features. Second, a relation-aware hashing learning scheme is designed to simultaneously explore the intrinsic semantic relationships and the global semantic information based on proxies. This is achieved by minimizing the reconstruction error between the multi-modal affinity matrices derived from learned features and the cross-modal similarity matrix of the hash codes. The proposed framework is able to learn more discriminative hash codes and achieves superior performance to many baselines on three public datasets.
Jinyu Xu 0001, Qing Xie 0002, Jiachen Li 0002, Yanchun Ma, Yongjian Liu, Zhifeng Bao
ACM Trans. Multim. Comput. Commun. Appl.3
2026 Enhancing image vectorization fidelity and editability through adaptive layered techniques
Qing Xie 0002, Guixiang Nie, Anshu Hu, Yanchun Ma, Jinyu Xu 0001, Jiachen Li 0002
Vis. Comput.6
2025 PKI-SSM: Prior Knowledge Integrated Self-supervised Model for Point Cloud Completing
Lingli Tang, Jiachen Li 0002, Yanchun Ma, Qing Xie 0002, Yongjian Liu
ICIC (9)3
2025 Stealthy Backdoor Attack against Object Detection
abstract
Recent research has revealed that object detectors are highly susceptible to backdoor attacks, which can introduce detection errors during inference, such as detecting non-existent objects or failing to detect existing objects. Though several backdoor attacks targeting object detection have been proposed to achieve high attack success rates, these methods often involve visible triggers, which can be detected by human inspection or backdoor defenses. To enhance the attack stealthiness, we introduce a stealthy backdoor attack for object detection. Specifically, it employs a uniform shift on each pixel within images as the trigger. The particle swarm optimization is utilized to effectively find the optimal uniform shift to accomplish different attack targets in object detection, including object disappearance, object generation, and object misclassification. To achieve these targets and preserve stealthy, we design corresponding objective functions to maintain a balance between attack stealthiness and attack effectiveness. We have conducted comprehensive experiments to demonstrate the effectiveness of our proposed attack across the three attack targets in object detection, as well as its robustness against existing defense methods.
Xiaoyang Ning, Qing Xie 0002, Jinyu Xu 0001, Wenbo Jiang 0001, Xiaoyuan Liu 0002, Jiachen Li 0002, Yanchun Ma
IJCNN6
2025 When Hallucinated Concepts Cross Modals: Unveiling Backdoor Vulnerability in Multi-modal In-context Learning
abstract
Due to the remarkable performance of multi-modal large language models (MLLMs) in multi-modal capabilities, multi-modal in-context learning (M-ICL) has garnered widespread attention for fast adapting MLLMs to downstream tasks. However, the vulnerability of M-ICL to attacks remains largely unexplored. In this work, we take the first step to explore the backdoor vulnerability of M-ICL, which allows the adversary only to manipulate the multi-modal demonstration examples to mislead the victim model. We propose a multi-modal backdoor strategy on M-ICL via cross-modal concept mis-matching under black-box attack setting. Extensive experimental results demonstrate that our attacks exhibit high attack effectiveness while preserving the normal functionality of the victim model. Moreover, we further conduct experiments to prove our attacks are robust against backdoor defenses and still remain effective in various real-world conditions.
Guanyu Hou, Jiaming He, Yitong Qiao, Jiachen Li 0002, Qiyang Song, Ji Guo, Wenbo Jiang 0001
MMAsia4
2025 TOVect: Topology-Optimized Vectorization for Intangible Cultural Heritage Thangka Element Line Art
abstract
Thangka art, part of the UNESCO Intangible Cultural Heritage of Humanity, is visually characterized by complex junctions and intricate corners, demand high-fidelity vectorization to preserve its structural integrity and smooth curvilinear aesthetics. Conventional line art vectorization algorithms applied to Thangka element line art face challenges: (1) hard to fit complex junctions that leads to spurious spikes and discontinuous strokes; and (2) unnatural distortions in long curves due to insufficient smoothness constraints. To address these challenges, we propose a skeleton-guided vectorization framework to optimize the topology of vectorized Thangka element line art, and a multilayer perceptual loss as a smoothness regulation to improve curve continuity. Experimental results on manually annotated Thangka element line art dataset demonstrate that our method surpasses state-of-the-art approaches in preserving topological integrity and achieving visual smoothness, offering a robust foundation for digitizing cultural heritage artworks with complex topologies and similar aesthetic requirements.
Anshu Hu, Yifei Sun 0018, Jiachen Li 0002, Yanchun Ma, Qing Xie 0002, Yongjian Liu
MMAsia3
2025 You Are Out of My Focus: A Defocus-Blur Backdoor Attack against Deep Learning Models
abstract
With the widespread adoption of deep learning in image recognition, backdoor attacks have emerged as a significant security threat, drawing increasing attention from the research community. Traditional backdoor attacks are often limited to the digital domain, while few existing physical-world attacks suffer from a lack of stealthiness. In this paper, inspired by the natural defocus blur commonly caused by camera optics in real-world environments, we propose a physically-aware backdoor attack method called DBBA based on the defocus blur phenomenon. By leveraging Gaussian blur to simulate this natural phenomenon, the proposed method enhances both the stealthiness and plausibility of the trigger. To further optimize the attack effectiveness while maintaining stealthiness, we introduce a Particle Swarm Optimization (PSO) algorithm to automatically search for the optimal Gaussian blur parameters that best simulate the defocus phenomenon. We conduct extensive experiments on multiple mainstream image classification datasets and across various model architectures. Experimental results demonstrate that the proposed defocus-blur based trigger achieves a high attack effectiveness with minimal degradation in the classification accuracy of the model. In addition, evaluations against representative defense techniques reveal that the proposed method exhibits strong stealthiness and robustness.
Hongwei Li 0001, Wenbo Jiang 0001, Jiaming He, Rui Zhang 0090, Ji Guo, Jiachen Li 0002
MMAsia8
2025 Robust Dual Embedding Contrastive Learning for Text-to-Image Person Re-identification with Noisy Correspondence
abstract
Text-to-Image person re-identification (TIReID) aims to retrieve pedestrian images from a gallery based on textual descriptions, thus bridging vision and language modalities for practical retrieval scenarios. Despite recent advances leveraging various cross-modal alignment strategies, existing methods typically assume all image-text pairs in training datasets are correctly matched, overlooking the pervasive Noisy Correspondence (NC) problem—erroneous image-text associations that degrade model robustness. Prior approaches either lack noise identification mechanisms or rely on direct filtering of detected noisy samples, which only partially mitigates the adverse effects of noise and cannot fully prevent overfitting to incorrect correspondences during training. Addressing this challenge, we propose Robust Dual Embedding Contrastive Learning (RDECL), which consists of two main components: 1) A Dual-View Cumulative Trust Division (DCTD) progressively constructs a high-confidence clean sample repository via adaptive sample selection, ensuring reliable image-text correspondence learning under uncertain noise detection.2) A Robust Generalized Contrastive Loss (RGCL) further enhances robustness by leveraging all negative samples and maximizing the loss distribution discrepancy between clean and noisy samples, thereby suppressing overfitting to noisy labels. We conduct extensive experiments on three public benchmark datasets, namely CUHK-PEDES, ICFG-PEDES, and RSTPReID, to evaluate the performance and robustness of our RDECL.
Jingjie Zhang, Lingli Tang, Jiachen Li 0002, Jinyu Xu 0001, Yanchun Ma, Qing Xie 0002
MMAsia3
2025 GroupRF: Panoptic Scene Graph Generation with group relation tokens
abstract
Panoptic Scene Graph Generation (PSG) aims to predict a variety of relations between pairs of objects within an image, and indicate the objects by panoptic segmentation masks instead of bounding boxes . Existing PSG methods attempt to straightforwardly fuse the object tokens for relation prediction, thus failing to fully utilize the interaction between the pairwise objects. To address this problem, we propose a novel framework named Group R elation F ormer (GroupRF) to capture the fine-grained inter-dependency among all instances. Our method introduce a set of learnable tokens termed group rln tokens, which exploit fine-grained contextual interaction between object tokens with multiple attentive relations. In the process of relation prediction, we adopt multiple triplets to take advantage of the fine-grained interaction included in group rln tokens. We conduct comprehensive experiments on OpenPSG dataset, which show that our method outperforms the previous state-of-the-art method. Furthermore, we also show the effectiveness of our framework by ablation studies. Our code is available at https://github.com/WHY-student/GroupRF .
Hongyun Wang, Jiachen Li 0002, Xiang Xiang 0001, Qing Xie 0002, Yanchun Ma, Yongjian Liu
J. Vis. Commun. Image Represent.2
2025 Backdoor attacks against Hybrid Classical-Quantum Neural Networks
Ji Guo, Wenbo Jiang 0001, Rui Zhang 0090, Wenshu Fan, Jiachen Li 0002, Guoming Lu, Hongwei Li 0001
Neural Networks5
2024 DPA-RCNN: Dual Position Aware 3D Object Detector for Point Cloud
abstract
In this paper, we explore the impact of the spatial properties of point clouds on 3D object detection in autonomous driving scenarios. To reduce the memory and computational costs, existing point-based models typically use random sampling or the farthest point sampling strategy to retain the foreground points or points closer to the center of the object. However, they treated points with different spatial properties equally, which led to the loss of potential relative geometric information in the point cloud. To this end, we design a two-stage detector that fully exploits the spatial properties of point clouds, termed Dual Position Aware 3D Object Detector(DPA-RCNN). Specifically, we first sample more points close to the object center by exploiting the relative position information between points and the object center. These sampling points can generate more precisely positioned proposals, which can reduce the difficulty of subsequent bounding box regression stages. In addition, we use the distance from the point to the edge of the bounding box to learn edge features in proposals, which further improves the regression accuracy of the bounding box. Experiments on KITTI and ONCE datasets validate the superiority and universality of our approach. Our method outperforms all point-based detectors, and the proposed Edge Point-aware Segmentation Module(EPAS) can effectively improve the detection accuracy of two-stage detectors.
Yidong Jiang, Qing Xie 0002, Jiachen Li 0002, Jinyu Xu 0001, Yongjian Liu, Yanchun Ma
IJCNN3
2024 HcaNet: Haze-concentration-aware Network for Real-scene Dehazing with Codebook Priors
abstract
In the task of image dehazing, it has been proven that high-quality codebook priors can be used to compensate for the distribution differences between real-world hazy images and synthetic hazy images, thereby helping the model improve its performance. However, because the concentration and distribution of haze in the image are irregular, the manners those simply replacing or blending the prior information in the codebook with the original image features are inconsistent with this irregularity, which leads to a non-ideal dehazing performance. To this end, we propose a haze concentration aware network (HcaNet), its haze-concentration-aware module (HcaM) can reduce the information loss in the vector quantization stage and achieve an adaptive domain transfer for regions with different degrees of degradation. To further capture the detailed texture information, we develop a frequency selective fusion module (FSFM) to facilitate the transmission of shallow information retained in haze areas to deeper layers, thereby enhancing the fusion with high-quality feature priors. Extensive evaluations demonstrate that the proposed model can be merely trained on synthetic hazy-clean pairs and effectively generalize to real-world data. Several experimental results confirm that the proposed dehazing model outperforms state-of-the-art methods significantly on real-world images.
Jiachen Li 0002, Yanchun Ma, Qing Xie 0002, Yongjian Liu
ACM Multimedia2
2024 Dlpp-Net: Degradation Location Prior Prediction Network for Image Restoration
Yongjian Liu, Shunwei Zhang, Jinyu Xu 0001, Jiachen Li 0002, Yanchun Ma, Qing Xie 0002
MMAsia4
2024 Mutually-Guided Hierarchical Multi-Modal Feature Learning for Referring Image Segmentation
abstract
Referring image segmentation aims to locate and segment the target region based on a given textual expression query. The primary challenge is to understand semantics from visual and textual modalities and achieve alignment and matching. Prior works have attempted to address this challenge by leveraging separately pretrained unimodal models to extract global visual and textual features and perform straightforward fusion to establish cross-modal semantic associations. However, these methods often concentrate solely on the global semantics, disregarding the hierarchical semantics of expression and image and struggling with complex and open real scenarios, thus failing to capture critical cross-modal information. To address these limitations, this article introduces an innovative mutually-guided hierarchical multi-modal feature learning scheme. By leveraging the guidance of global visual features, the model mines hierarchical text features from different stages of the text encoder. Simultaneously, the guidance of global textual features is leveraged to aggregate multi-scale visual features. This mutually guided hierarchical feature learning effectively addresses the semantically inaccurate cause by free-form text and naturally occurring scale variations. Furthermore, a Segment Detail Refinement (SDR) module is designed to enhance the model’s spatial detail awareness through attention mapping of low-level visual features and cross-modal features. To evaluate the effectiveness of the proposed approach, extensive experiments are conducted on three widely used referring image object segmentation datasets. The results demonstrate the superiority of the presented method in accurately locating and segmenting objects in images.
Jiachen Li 0002, Qing Xie 0002, Xiaojun Chang, Jinyu Xu 0001, Yongjian Liu
ACM Trans. Multim. Comput. Commun. Appl.1
2023 A Multi-scale and Dense Object Detector for Tibetan Thangka Images
abstract
Thangka cultural elements detection aims to locate and identify instances in Thangka. However, as a unique form of pictorial art, Thangka exhibits distinct spatial structures that deviate significantly from general images in scale and density. Therefore, it is challenging for most state-of-the-art detectors designed for natural scenes to handle Thangka cultural elements detection effectively. To overcome this issue, we propose a multi-scale and dense object detector referred as MDDet. It embeds a multi-scale receptive field fusion module (MRF) that enlarges the receptive field while capturing the spatial and channel relationships at different scales, which significantly enriches the multi-scale features extracted from the backbone. In addition, we introduce a threshold-slicing aided hyper inference (T-SAHI) scheme, which adaptively slices images in dense scenarios to aid with dense object detection in the test time. We thoroughly evaluate our method, and MDDet outperforms the prior art by a clear margin on the Thangka dataset, achieving an absolute improvement of 1.9% in average precision (AP). For the challenging medium and small objects in Thangka, MDDet obtains wide margins of 12% and 3.7% in accuracy improvement, respectively. It also shows strong generalization ability when evaluated on general scenarios, e.g., Pascal VOC 2007 and MS COCO, validating the role of MDDet in object detection.
Gaohuan Dong, Qing Xie 0002, Jiachen Li 0002, Yanchun Ma, Yuhan Liu 0001, Yongjian Liu
MMAsia3
2023 Mixture of Experts Residual Learning for Hamming Hashing
Jinyu Xu 0001, Qing Xie 0002, Jiachen Li 0002, Yanchun Ma, Yuhan Liu 0001
Neural Process. Lett.3
2021 Gat-Assisted Deep Hashing For Multi-Label Image Retrieval
abstract
Multi-Label hash methods have achieved excellent performance in multi-label image retrieval, but how to leverage the semantic information of label to improve retrieval quality is still a challenge in this field. This paper proposes GAT-Assisted Deep Hashing (DHGAT). Our model uses Convolutional Neural Network (CNN) to extract image-level features, along with graph attention network (GAT) to extract label-level features. Assisted by GAT, DHGAT is able to pay more attention on the co-occurrence of label. In order to solve the problem of feature fusion, we propose Multi-modal Max-Pooling Bilinear (MMB) mechanism, which MMB fuses image-level feature and label-level feature to generate abundant semantic features, so that the model can output discriminative hash code. Extensive experiments demonstrate that the proposed method can generate hash codes which achieve better retrieval performance on two benchmark datasets, NUS-WIDE and MS-COCO.
Jiachen Li 0002, Yanchun Ma, Qing Xie 0002, Yongjian Liu
ICIP1
2021 Visible-infrared Person Re-identification with Human Body Parts Assistance
abstract
Person re-identification (re-id) has received ever-increasing research focus, because of its important role in video surveillance applications. This paper addresses the re-id problem between visible images of color cameras and infrared images of infrared cameras, which is significant in case that the appearance information is insufficient in poor illumination conditions. In this field, there are two key challenges, i.e., the difficulty to locate the discriminative information to re-identify the same person between visible and infrared images, and the difficulty to learn a robust metric for such large-scale cross-modality retrieval. In this paper, we propose a novel human body parts assistance network (BANet) to tackle the two challenges above. BANet mainly focuses on extracting discriminative information and learning robust features by leveraging the human body part cues. Extensive experiments demonstrate that the proposed approach outperforms the baseline and the state-of-the-art methods.
Huangpeng Dai, Qing Xie 0002, Jiachen Li 0002, Yanchun Ma, Lin Li 0001, Yongjian Liu
ICMR3
2019 Deep Multi-label Hashing for Image Retrieval
abstract
Due to its low storage cost and fast query speed, hashing has been widely applied to approximate nearest neighbor search for large-scale image retrieval, while deep hashing further improves the retrieval quality by learning a good image representation. However, existing deep hash methods simplify multi-label images into single-label processing, so the rich semantic information from multi-label is ignored. Meanwhile, the imbalance of similarity information leads to the wrong sample weight in the loss function, which makes unsatisfactory training performance and lower recall rate. In this paper, we propose Deep Multi-Label Hashing (DMLH) model that generates binary hash codes which retain the semantic relationship of multi-label of the image. The contributions of this new model mainly include the following two aspects: (1) A novel sample weight calculation model adaptively adjusts the weight of the sample pair by calculating the semantic similarity of the multi-label image pairs. (2) The sample weight cross-entropy loss function, which is designed according to the similarity of the image, adjusts the balance of similar image pairs and dissimilar image pairs. Extensive experiments demonstrate that the proposed method can generate hash codes which achieve better retrieval performance on two benchmark datasets, NUS-WIDE and MS-COCO.
Xian Zhong, Jiachen Li 0002, Wenxin Huang
ICTAI2