VLDB 2026 Research / reviewers in the wild / expert
Dayan Wu
dblp:200/3052
· DBLP profile ↗
66ranked-venue papers
7as first author
56since 2021 · last 2026
0000-0002-8604-7226ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 53 · 7 first-author · 43 since 2021Artificial intelligence and machine learning · 21 · 2 first-author · 18 since 2021Databases, data management, data science and information retrieval · 7 · 1 first-author · 4 since 2021Computer networks · 3 · 1 first-author · 3 since 2021Security and privacy · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Discretization Is Not Always Better: Rethinking Deep Quantization for Asymmetric Image RetrievalabstractAsymmetric image retrieval (AIR), which typically employs a compact model for the query side and a large model for the database server, has garnered significant attention in resource-constrained environments. While deep hashing methods have shown great potential in large-scale image retrieval, current attempts for the asymmetric image retrieval overlook the differences in quantization capabilities between query and gallery networks. In AIR, the conventional quantization scheme forces the outputs of small query models to approximate the discrete outputs of large models, imposing overly rigid and stringent constraints that severely limit the optimization of small query models. Furthermore, existing deep hashing methods for AIR necessitate labeled datasets from large models, which also limits their practical applicability. To this end, we reconsider the necessity of strict discretization in AIR and propose a novel asymmetric hashing method, named Deep Correlation Alignment Hashing (DCAH). Rather than explicitly quantizing continuous query features to match discrete gallery representations, we distill the correlation across both models and introduce a Correlation Alignment based Quantization (CAQ) scheme, thereby implicitly accomplishing quantization. To preserve the similarity consistency between the query and gallery models, we further employ a correlation alignment-based knowledge distillation strategy which is intrinsically compatible with the CAQ. Notably, the proposed quantization scheme can function as a plug-and-play module that seamlessly integrates with existing AIR methods. Comprehensive evaluations on three real-world benchmark datasets demonstrate the effectiveness of the proposed quantization scheme CAQ, and also show that DCAH achieves state-of-the-art performance in asymmetric image retrieval scenarios. Dayan Wu, Hengjie Zhu, Chenming Wu, Pengwen Dai |
AAAI | 2 |
| 2026 | Beyond Vector Search: MLLM-Driven Generative Composed Image Retrieval
Ruicheng Xiong, Gengqi Yang, Dayan Wu |
ICIC (13) | 3 |
| 2026 | Correspondence Matters: Explicit Visual-Knowledge Alignment for Retrieval-Augmented Image Captioning
Gengqi Yang, Dayan Wu |
ICIC (10) | 2 |
| 2026 | Seedcap: Semantic Expansion and Entity-Driven Zero-Shot Image CaptioningabstractZero-shot image captioning (ZIC) generates natural language descriptions for images using only textual data during training. Recent progress leverages text-to-image models to construct pseudo image-text pairs, alleviating the modality mismatch between text-only training and image-based inference. However, the distribution gap between synthetic and real images can compromise performance when models trained on generated images are deployed in real-world scenarios. To address this challenge, we propose SEEDCap, a zero-shot captioning framework featuring SEED (Semantic Expansion and Entity-Driven), a parameter-free module that performs dual-level semantic expansion. SEED enhances the correspondence between image features and retrieved textual cues at both entity and sentence levels, producing enriched semantic embeddings that effectively reduce the distributional gap between synthetic and real inputs. These enhanced features are then integrated with the global visual representation through a modality fusion module, which is subsequently decoded by a large language model to generate accurate and contextually rich captions. Extensive experiments demonstrate that SEEDCap achieves state-of-the-art zero-shot results on both in-domain and cross-domain benchmarks, underscoring its robustness and practical utility. Binbin Li 0003, Shupei Xiao, Dayan Wu, Gengqi Yang, Siyu Jia, Zisen Qi |
ICMR | 3 |
| 2026 | Planning forward: Deep incremental hashing by gradually defrosting bits
Qinghang Su, Dayan Wu, Chenming Wu, Bo Li 0063, Weiping Wang 0005 |
Neural Networks | 2 |
| 2026 | H2R-BM: Can leveraging human videos enhance performance and generalizability in robotic bimanual manipulation?
Xiaoshuai Hao, Huaihai Lyu, Dayan Wu, Jing Zhang 0037, Long Chen 0015 |
Pattern Recognit. | 5 |
| 2026 | TrustSearch: Toward Secure and Efficient Reverse Image Search via SGXabstractOutsourcing image management to a cloud should not only protect the confidentiality of image data, but also maintain the capability of reverse image search, which requires identifying the existing stored images that are similar to an input image. Previous studies build on cryptographic approaches to realize reverse image search on encrypted images, yet failing to achieve either security or performance. This paper explores trusted image search, which uses Intel SGX to realize reverse image search in an enclave, in order to provide security guarantees via SGX while performing search on plain data (inside the enclave) for performance. However, due to the resource limits of SGX, directly realizing the search process in the enclave incurs high performance overhead. We present TRUSTSEARCH, which implements various design approaches to mitigate the resource overhead of SGX. We evaluate TRUSTSEARCH using real-world image datasets, and show that it outperforms state-of-the-art approaches for search performance while preserving space efficiency for the enclave. Fang Zou, Jingwei Li 0001, Dayan Wu, Xiong Li 0002, Hongwei Li 0001, Ting Chen 0002, Xiaosong Zhang 0001 |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2025 | Corer: Concept Residue Erasing in Text-to-Image Diffusion ModelsabstractThe remarkable development of text-to-image generation models has raised notable security concerns, such as the infringement of portrait rights and the generation of inappropriate content. Concept erasure has been proposed to remove the model’s knowledge about protected or inappropriate concepts. Although many methods have tried to balance the efficacy (erasing target concepts) and specificity (retaining irrelevant concepts), they can still generate abundant erasure concepts under the steering of semantically related inputs. In this work, we propose Corer to address this "concept residue" issue. Specifically, we first introduce the mechanism of neighbor-concept mining to dig out the associated concepts and expand the erasing range. Furthermore, to mitigate the negative impact on the generation of irrelevant concepts caused by the expansion of erasure scope, Corer preserves the specificity through the beyond-concept regularization. We also employ the closed-form solution to optimize weights of U-Net, as well as the prediction noise alignment with the LoRA module. Extensive experiments on multiple benchmarks demonstrate that Corer outperforms previous concept-erasing methods in terms of superior erasing efficacy, specificity, and generality. Yufan Liu 0002, Jinyang An, Huashan Chen, Wanqian Zhang, Dayan Wu, Jingzi Gu, Zheng Lin 0001, Weiping Wang 0005 |
ICME | 6 |
| 2025 | Categorical Attention: Fine-grained Language-guided Noise Filtering Network for Occluded Person Re-IdentificationabstractPerson Re-Identification (ReID) aims to match individuals across different camera views, but occlusions in real-world scenarios, such as vehicles or crowds, hinder feature extraction and matching. Current occluded ReID methodologies typically leverage visual augmentation techniques in an attempt to mitigate the disruptive effects of occlusion-induced noise. However, relying solely on visual data fail to effectively filter out occlusion noise. In this paper, we introduce the Fine-grained Language-guided Noise Filtering Network (FLaN-Net) for occluded ReID. FLaN-Net innovatively employs categorical attention mechanism to generate adaptive tokens that capture the following three distinct types of visual information: comprehensive descriptions of individuals, detailed visible attributes, and characteristics of occluding objects. Subsequently, a cross-attention mechanism aligns these prompts with the image, guiding the model to focus on relevant regions. To generate robust and discriminative features for occluded pedestrians, we further introduce a dynamic weighting fusion module that integrates visual, textual, and cross-attention features based on their reliability. Experimental results demonstrate that FLaN-Net outperforms existing methods on occluded ReID benchmarks, offering a robust solution for challenging real-world conditions. Dayan Wu, Chenxu Yang, Qinghang Su, Zheng Lin 0001 |
IJCAI | 2 |
| 2025 | Endogenous Recovery via Within-modality Prototypes for Incomplete Multimodal HashingabstractMultimodal hashing projects multimodal data into compact binary codes, enabling rapid and storage-efficient retrieval of large-scale multimedia content. In practical scenarios, the issue of missing modality frequently arises when dealing with multimodal data. Existing incomplete multimodal hashing techniques directly recover missing modalities by neural networks, resulting in a disjointed representation space between the recovered and true data. In this paper, we present a novel recovery paradigm, namely Prototype-based Modality Completion Hashing (PMCH). Instead of directly synthesizing it from available modalities, PMCH adaptively aggregates associated within-modality prototypes to recover missing modality data. Specifically, PMCH introduces an within-modality prototype learning module to optimize representative prototypes for each modality. These prototypes act as recovery anchors and reside within the same representation space as their corresponding modality data. Subsequently, PMCH adaptively aggregates the associated within-modality prototypes with coefficients derived from the modality-specific Weight-Net. By utilizing prototypes from the same modality, the semantic disparity between the reconstructed and authentic data can be substantially diminished. Extensive experiments on three widely used benchmark datasets demonstrate that PMCH can effectively recover the missing modality, and attain state-of-the-art performance in both complete and incomplete multimodal retrieval scenarios. Code is available at https://github.com/Sasa77777779/PMCH.git. Sa Zhu, Dayan Wu, Chenming Wu, Pengwen Dai, Bo Li 0063 |
IJCAI | 2 |
| 2025 | Beyond Similarity: Aligning Retrieval with Generation for Enhanced Image CaptioningabstractImage captioning is a crucial task that bridges computer vision and natural language processing, with applications ranging from assistive technologies to content creation. Recently, retrieval-augmented generation (RAG) methods have been introduced to enhance captioning models by incorporating external knowledge and enabling training-free domain transfer. These approaches utilize pre-trained retrievers to fetch relevant information for input images, providing additional context to the caption generation process. However, we observe that high-similarity retrieved results do not always enhance the generation process, revealing a misalignment between retrieval and generation in existing methods. To address this issue, we propose Retrieval Alignment for Caption Enhancement (RACE), a novel RAG-based image captioning method. Instead of merely enhancing the similarity between the input image and the retrieved results, our RACE method focuses on retrieving results that more effectively benefit the final generation outcomes. Specifically, the retriever is optimized to align with the image captioning model. We introduce a generation-oriented similarity correction loss to align the similarity-based distribution with the CIDEr-based evaluation distribution of the generated captions. In this proposed loss function, greater emphasis is placed on training samples with higher divergence. This ensures that the retriever retrieves more relevant results preferred by the language decoder, while preserving pre-trained knowledge. Extensive experiments on both in-domain (COCO) and out-of-domain (NoCaps, Flickr30k, VizWiz) benchmarks demonstrate that RACE outperforms other RAG-based image captioning methods. By effectively bridging the gap between retrieval and generation, RACE enhances overall captioning quality. Gengqi Yang, Dayan Wu, Jinyang An, Shupei Xiao, Zheng Lin 0001 |
IJCNN | 2 |
| 2025 | Mitigating the Evolving Semantic Entanglement in Continual Learning of Vision-Language Models
Yiliang Zhu 0002, Dayan Wu, Qinghang Su, Zexian Yang, Zheng Lin 0001, Weiping Wang 0005 |
ACM Multimedia | 2 |
| 2025 | NeuS-PIR: Learning Relightable Neural Surface Using Pre-Integrated RenderingabstractIn this paper, we propose NeuS-PIR, a novel approach for learning relightable neural surfaces using pre-integrated rendering from multi-view image observations. Unlike traditional methods based on NeRFs or discrete mesh representations, our approach employs an implicit neural surface representation to reconstruct high-quality geometry. This representation enables the factorization of the radiance field into two components: a spatially varying material field and an all-frequency lighting model. By jointly optimizing this factorization with a differentiable pre-integrated rendering framework, and material encoding regularization, our method effectively addresses the ambiguity in geometry reconstruction, leading to improved disentanglement and refinement of scene properties. Furthermore, we introduce a technique to distill indirect illumination fields, capturing complex lighting effects such as inter-reflections. As a result, NeuS-PIR enables advanced applications like relighting, which can be seamlessly integrated into modern graphics engines. Extensive qualitative and quantitative experiments on both synthetic and real datasets demonstrate that NeuS-PIR outperforms existing methods across various tasks. Source code is available at https://github.com/Sheldonmao/NeuSPIR. Shi Mao, Chenming Wu, Zhelun Shen, Dayan Wu, Liangjun Zhang |
Comput. Vis. Media | 5 |
| 2025 | Enhancing facial privacy protection in customized diffusion models via masked attention erasure
Yisu Liu, Lin Wang 0108, Wanqian Zhang, Jinyang An, Huashan Chen, Dayan Wu, Zheng Lin 0001, Weiping Wang 0005 |
Knowl. Based Syst. | 6 |
| 2025 | Boundary-aware Prototype Augmentation and Dual-level Knowledge Distillation for Non-Exemplar Class-Incremental Hashing
Qinghang Su, Dayan Wu, Bo Li 0063 |
Knowl. Based Syst. | 2 |
| 2025 | TextSafety: Visual Text Vanishing via Hierarchical Context-Aware Interaction ReconstructionabstractPrivacy information existing in the scene text will be leaked with the spread of images in cyberspace. Vanishing the scene text from the image is a simple yet effective method to prevent privacy disclosure to the machine and the human. Previous visual text vanishing methods have achieved promising results but the performance still fell short of expectations for complicated-shape scene texts with various scales. In this paper, we propose a novel hierarchical context-aware interaction reconstruction method to make the visual text vanish in the natural scene image. To avoid the interference of the non-text regions, we narrow down the reconstruction regions by the guidance of the hierarchical refined text region masks, helping provide accurate position information. Meanwhile, we propose to learn the long-range context-aware interaction in a lightweight way, which can ensure the smoothing of the artifacts that are easily generated by the convolutional layers. To be more specific, we first simultaneously generate the coarse text region mask and the initially vanishing scene text image. Then, we obtain more accurate refined masks to better capture the locations of complicated-shape texts via a hierarchical mask generation network. Next, based on the refined masks, we exploit a channel-wise context-aware interaction mechanism to model the long-range relationships between the reconstruction region and the backgrounds for better removing the artifacts. Finally, we fuse the reconstructed text regions with the non-masked regions to obtain the ultimate protected image. Experiments on two frequently-used benchmarks SCUT-EnsText and SCUT-Syn demonstrate that our proposed method outperforms previous related methods by a large margin. Pengwen Dai, Dayan Wu, Peijia Zheng, Xiaochun Cao |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2024 | Pairwise-Label-Based Deep Incremental Hashing with Simultaneous Code ExpansionabstractDeep incremental hashing has become a subject of considerable interest due to its capability to learn hash codes in an incremental manner, eliminating the need to generate codes for classes that have already been learned. However, accommodating more classes requires longer hash codes, and regenerating database codes becomes inevitable when code expansion is required. In this paper, we present a unified deep hash framework that can simultaneously learn new classes and increase hash code capacity. Specifically, we design a triple-channel asymmetric framework to optimize a new CNN model with a target code length and a code projection matrix. This enables us to directly generate hash codes for new images, and efficiently generate expanded hash codes for original database images from the old ones with the learned projection matrix. Meanwhile, we propose a pairwise-label-based incremental similarity-preserving loss to optimize the new CNN model, which can incrementally preserve new similarities while maintaining the old ones. Additionally, we design a double-end quantization loss to reduce the quantization error from new and original query images. As a result, our method efficiently embeds both new and original similarities into the expanded hash codes, while keeping the original database codes unchanged. We conduct extensive experiments on three widely-used image retrieval benchmarks, demonstrating that our method can significantly reduce the time required to expand existing database codes, while maintaining state-of-the-art retrieval performance. Dayan Wu, Qinghang Su, Bo Li 0063, Weiping Wang 0005 |
AAAI | 1 |
| 2024 | A Pedestrian is Worth One Prompt: Towards Language Guidance Person Re- IdentificationabstractExtensive advancements have been made in person ReID through the mining of semantic information. Nevertheless, existing methods that utilize semantic-parts from a single image modality do not explicitly achieve this goal. Whiteness the impressive capabilities in multimodal understanding of Vision Language Foundation Model CLIP, a recent two-stage CLIP-based method employs automated prompt engineering to obtain specific textual labels for classifying pedestrians. However, we note that the predefined soft prompts may be inadequate in expressing the entire visual context and struggle to generalize to unseen classes. This paper presents an end-to-end Prompt-driven Semantic Guidance (PromptSG) framework that harnesses the rich semantics inherent in CLIP. Specifically, we guide the model to attend to regions that are semantically faithful to the prompt. To provide personalized language descriptions for specific individuals, we propose learning pseudo tokens that represent specific visual contexts. This design not only facilitates learning fine-grained attribute information but also can inherently leverage language prompts during inference. Without requiring additional labeling efforts, our PromptSG achieves state-of-the-art by over 10% on MSMTI7 and nearly 5% on the Market-I50I benchmark. The codes will be available at h t tps: / / gi th ub. com/ YzXian16/PromptSG Zexian Yang, Dayan Wu, Chenming Wu, Zheng Lin 0001, Jingzi Gu, Weiping Wang 0005 |
CVPR | 2 |
| 2024 | Prediction Exposes Your Face: Black-Box Model Inversion via Prediction Alignment
Yufan Liu 0002, Wanqian Zhang, Dayan Wu, Zheng Lin 0001, Jingzi Gu, Weiping Wang 0005 |
ECCV (36) | 3 |
| 2024 | Exploring Targeted Universal Adversarial Attack for Deep HashingabstractAlthough image-dependent adversarial attacks have been studied, the more challenging image-agnostic adversarial attack for deep hashing remains an unexplored territory. In this paper, we take the first attempt on the more efficient and malicious targeted universal adversarial attack (TUAA) for deep hashing. When previous image-dependent attacks are directly applied to TUAA task, they usually face two main issues. Firstly, existing anchor code generation methods generate anchor code with inferior representative semantic-preserving ability. Secondly, previous methods simply minimize the distance between hash codes and anchor code in Hamming space, which tends to optimize targeted universal adversarial perturbation (TUAP) in a coarse-grained manner. To tackle the above problems, we propose a Semantic-enhanced and Stabilized Targeted Universal Adversarial Attack (SS-TUAA) method. Specifically, we first propose a new Candidate Anchor code Evaluation (CAE) method to generate anchor code with superior semantic-preserving ability. Then, to enhance the ‘dominant role’ and the stability of TUAP, we propose a Feature Consistency Loss (FCL) to align the fine-grained feature representation between TUAP and adversarial examples. Extensive experiments demonstrate the effectiveness of each component within our SS-TUAA method, and our method can achieve the state-of-the-art TUAA performance for deep hashing. Wanqian Zhang, Dayan Wu, Lin Wang 0108, Bo Li 0063, Weiping Wang 0005 |
ICASSP | 3 |
| 2024 | SD4Privacy: Exploiting Stable Diffusion for Protecting Facial PrivacyabstractRecently, adversarial examples are introduced to protect personal images from being identified by unauthorized face recognition systems. Existing approaches follow the transfer-based adversarial attack paradigm, where local surrogate models are utilized to generate protected images. However, these surrogate models can neither be necessary nor efficient for generating adversarial examples. In this paper, we propose SD4Privacy, i.e., Stable Diffusion for Privacy, which exploits the latent space of Stable Diffusion Model to synthesize adversarial examples. First, we learn an optimal textual embedding of target image to preserve its representative semantics, directly guiding the sampling process of synthesized image. Then, we utilize the encoder of UNet in Stable Diffusion as the substitution of surrogate classification models, which enables the efficient adversarial guidance by semantic h-space of UNet for adversarial example generation. Experiments show the state-of-the-art protection performance, as well as high-quality protected images with visual naturalness and imperceptible perturbations. Jinyang An, Wanqian Zhang, Dayan Wu, Zheng Lin 0001, Jingzi Gu, Weiping Wang 0005 |
ICME | 3 |
| 2024 | Exploiting Vision-Language Model for Visible-Infrared Person Re-identification via Textual Modality AlignmentabstractVisible-Infrared Person Re-identification (VI-ReID) aims at matching the images of specific person captured by different modality cameras. Previous methods introduce a synthesized auxiliary modality to relieve the modality discrepancy. However, they directly fuse the raw pixels of visible and infrared images, ignoring the high-level semantic patterns. Additionally, the huge modality gap can’t be bridged up closely, which leads to the oscillations in the feature space. Thus, in this paper, we propose a novel Textual Modality Alignment Learning method, named TMAL, which tackles these two issues in a unified two-stage framework. Specifically, we first exploit the semantic alignment in CLIP model through learnable text tokens, which are then encoded to form semantic representations of each identity. In the second stage, we propose the Modality Alignment Module, empowering the image encoder with modality-shared and modality-specific features. We also introduce the Identity Enhancement module (IEM) to extract more informative modality-specific features. Experiments on two benchmarks demonstrate the efficacy of our method. Bingyu Duan, Wanqian Zhang, Dayan Wu, Zheng Lin 0001, Jingzi Gu, Weiping Wang 0005 |
ICME | 3 |
| 2024 | Privacy-Preserving Replay and Adaptive Relation Distillation for Camera Incremental Person Re-IdentificationabstractTraditional person re-identification (ReID) methods trained on static data are ill-suited to real-world dynamic surveillance systems. Recently, a more desirable setting "Camera Incremental Person ReID (CIPR)", has been proposed to continually adapt to new cameras and accumulate knowledge. However, prior work on relation distillation heavily constrains intra-class relations for all identities, while under-exploring the credibility of different identities in knowledge transfer. Besides, their rehearsal-free setting sidesteps privacy concerns but compromises performance. In this paper, we present a novel framework, P2-ARD, designed specifically for CIPR. Firstly, we propose an innovative Adaptive Relation Distillation loss that automatically selects more crucial identities for distillation. Additionally, we introduce the privacy-preserving replay scheme to effectively retain semantic information while ensuring the privacy of the identity. Finally, we incorporate a cycle-consistent correlation method to address the class overlap issue in CIPR. Extensive experiments demonstrate our method outperforming the state-of-the-art. Zexian Yang, Dayan Wu, Wanqian Zhang, Jingzi Gu, Zheng Lin 0001, Weiping Wang 0005 |
ICME | 2 |
| 2024 | Disrupting Diffusion: Token-Level Attention Erasure Attack against Diffusion-based CustomizationabstractWith the development of diffusion-based customization methods like DreamBooth, individuals now have access to train the models that can generate their personalized images. Despite the convenience, malicious users have misused these techniques to create fake images, thereby triggering a privacy security crisis. In light of this, proactive adversarial attacks are proposed to protect users against customization. The adversarial examples are trained to distort the customization model's outputs and thus block the misuse. In this paper, we propose DisDiff (Disrupting Diffusion), a novel adversarial attack method to disrupt the diffusion model outputs. We first delve into the intrinsic image-text relationships, well-known as cross-attention, and empirically find that the subject-identifier token plays an important role in guiding image generation. Thus, we propose the Cross-Attention Erasure module to explicitly "erase" the indicated attention maps and disrupt the text guidance. Besides, we analyze the influence of the sampling process of the diffusion model on Projected Gradient Descent (PGD) attack and introduce a novel Merit Sampling Scheduler to adaptively modulate the perturbation updating amplitude in a step-aware manner. Our DisDiff outperforms the state-of-the-art methods by 12.75% of FDFR scores and 7.25% of ISM scores across two facial benchmarks and two commonly used prompts on average. Yisu Liu, Jinyang An, Wanqian Zhang, Dayan Wu, Jingzi Gu, Zheng Lin 0001, Weiping Wang 0005 |
ACM Multimedia | 4 |
| 2024 | Central similarity consistency hashing for asymmetric image retrievalabstractAsymmetric image retrieval methods have drawn much attention due to their effectiveness in resource-constrained scenarios. They try to learn two models in an asymmetric paradigm, i.e., a small model for the query side and a large model for the gallery. However, we empirically find that the mutual training scheme (learning with each other) will inevitably degrade the performance of the large gallery model, due to the negative effects exerted by the small query one. In this paper, we propose Central Similarity Consistency Hashing (CSCH), which simultaneously learns a small query model and a large gallery model in a mutually promoted manner, ensuring both high retrieval accuracy and efficiency on the query side. To achieve this, we first introduce heuristically generated hash centers as the common learning target for both two models. Instead of randomly assigning each hash center to its corresponding category, we introduce the Hungarian algorithm to optimally match each of them by aligning the Hamming similarity of hash centers to the semantic similarity of their classes. Furthermore, we introduce the instance-level consistency loss, which enables the explicit knowledge transfer from the gallery model to the query one, without the sacrifice of gallery performance. Guided by the unified learning of hash centers and the distilled knowledge from gallery model, the query model can be gradually aligned to the Hamming space of the gallery model in a decoupled manner. Extensive experiments demonstrate the superiority of our CSCH method compared with current state-of-the-art deep hashing methods. The open-source code is available at https://github.com/dubanx/CSCH . Zhaofeng Xuan, Dayan Wu, Wanqian Zhang, Qinghang Su, Bo Li 0063, Weiping Wang 0005 |
Comput. Vis. Media | 2 |
| 2024 | From Data to Optimization: Data-Free Deep Incremental Hashing With Data Disambiguation and Adaptive ProxiesabstractDeep incremental hashing methods require a large number of original training samples to preserve old knowledge. However, the old training samples are not always available. This “data-free” setting poses great challenges for learning discriminative codes for new classes (plasticity) and maintaining the code invariance of old ones (stability). On the one hand, the presence of ambiguous data in new-emerging classes, which is highly similar to that in old classes, further aggravates catastrophic forgetting. On the other hand, although well-separated hash codes of new classes can be learned by forcing them towards fixed hash centers, it may significantly change the learned parameters of the old model, leading to severe forgetting on old classes. To alleviate the stability-plasticity dilemma in data-free situations, this paper presents a novel deep incremental hashing method called Data-Free Deep Incremental Hashing (DFIH) from the data to the optimization aspect. We start from the data aspect and propose a data disambiguation module to reveal and discard ambiguous data, especially pixels to alleviate the forgetting issues. Subsequently, we introduce a set of trainable hash proxies during the optimization process. These proxies are optimized adaptively as well as the hash codes, not only guiding the model to learn discriminative hash codes for new classes but also avoiding the dramatic modification of the model’s parameters, thus improving plasticity and maintaining stability. Extensive experiments on six widely-used image retrieval benchmarks and sixteen incremental learning situations show the superiority of DFIH. Ablation analysis further confirms the effectiveness of the components in DFIH. The code of this work is released athttps://github.com/SuQinghang/DFIH. Qinghang Su, Dayan Wu, Chenming Wu, Bo Li 0063, Weiping Wang 0005 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2023 | Dual Alignment Unsupervised Domain Adaptation for Video-Text RetrievalabstractVideo-text retrieval is an emerging stream in both computer vision and natural language processing communities, which aims to find relevant videos given text queries. In this paper, we study the notoriously challenging task, i.e., Unsupervised Domain Adaptation Video-text Retrieval (UDAVR), wherein training and testing data come from different distributions. Previous works merely alleviate the domain shift, which however overlook the pairwise misalignment issue in target domain, i.e., there exist no semantic relationships between target videos and texts. To tackle this, we propose a novel method named Dual Alignment Domain Adaptation (DADA). Specifically, we first introduce the cross-modal semantic embedding to generate discriminative source features in a joint embedding space. Besides, we utilize the video and text domain adaptations to smoothly balance the minimization of the domain shifts. To tackle the pairwise misalignment in target domain, we propose the Dual Alignment Consistency (DAC) to fully exploit the semantic information of both modalities in target domain. The proposed DAC adaptively aligns the video-text pairs which are more likely to be relevant in target domain, enabling that positive pairs are increasing progressively and the noisy ones will potentially be aligned in the later stages. To that end, our method can generate more truly aligned target pairs and ensure the discriminability of target features. Compared with the state-of-the-art methods, DADA achieves 20.18% and 18.61% relative improvements on R@1 under the setting of TGIF→MSR-VTT and TGIF→MSVD respectively, demonstrating the superiority of our method. Xiaoshuai Hao, Wanqian Zhang, Dayan Wu, Bo Li 0063 |
CVPR | 3 |
| 2023 | Cross-Camera Prototype Learning for Intra-camera Supervised Person Re-identification
Bingyu Duan, Wanqian Zhang, Dayan Wu, Lin Wang 0108, Bo Li 0063, Weiping Wang 0005 |
ICANN (7) | 3 |
| 2023 | TeAw: Text-Aware Few-Shot Remote Sensing Image Scene ClassificationabstractThe recent advance has shown that few-shot learning may be a promising way to alleviate the data reliance of remote sensing image scene classification. However, most existing works focus on extracting distinguishable features only from visual modality, while the problem of learning knowledge from multiple modalities has barely been visited. In this work, we propose a text-aware framework for few-shot remote sensing image scene classification (TeAw). Specifically, TeAw converts the class names to more detailed text descriptions and extracts text features using a pre-trained text encoder. Mean-while, TeAw obtains image features via an image encoder. Then we compute the correlation between the text and the image features, which helps the model grasp the core concept of the input image. Finally, TeAw calculates the similarity of local features between supports and queries to get the predictions. Extensive experiments show the outperformance of our TeAw compared with other SOTA methods. Kaihui Cheng, Chule Yang, Zunlin Fan, Dayan Wu, Naiyang Guan |
ICASSP | 4 |
| 2023 | AREA: Adaptive Reweighting via Effective Area for Long-Tailed ClassificationabstractLarge-scale data from the real-world usually follow a long-tailed distribution (i.e., a few majority classes occupy plentiful training data, while most minority classes have few samples), making the hyperplanes heavily skewed to the minority classes. Traditionally, reweighting is adopted to make the hyperplanes fairly split the feature space, where the weights are designed according to the number of samples. However, we find that the number of samples in a class can not accurately measure the size of its spanned space, especially for the majority class, where the size of its spanned space is usually larger than the samples’ number because of the high diversity. Therefore, weights designed based on the samples’ number will still compress the space of minority classes. In this paper, we reconsider reweighting from a totally new perspective of analyzing the spanned space of each class. We argue that, besides statistical numbers, relations between samples are also significant for sufficiently depicting the spanned space. Consequently, we estimate the size of the spanned space for each category, namely effective area, by detailedly analyzing its samples’ distribution. By treating samples of a class as identically distributed random variables and analyzing their correlations, a simple and non-parametric formula is derived to estimate the effective area. Then, the weight simply calculated inversely proportional to the effective area of each class is adopted to achieve fairer training. Note that our weights are more flexible as they can be adaptively adjusted along with the optimizing features during training. Experiments on four long-tailed datasets show that the proposed weights outperform the state-of-the-art reweighting methods. Moreover, our method can also achieve better results on statistically balanced CIFAR-10/100. Code is available at https://github.com/xiaohua-chen/AREA. Xiaohua Chen 0002, Yucan Zhou, Dayan Wu, Chule Yang, Bo Li 0063, Qinghua Hu, Weiping Wang 0005 |
ICCV | 3 |
| 2023 | Digging into Depth Priors for Outdoor Neural Radiance FieldsabstractNeural Radiance Fields (NeRFs) have demonstrated impressive performance in vision and graphics tasks, such as novel view synthesis and immersive reality. However, the shape-radiance ambiguity of radiance fields remains a challenge, especially in the sparse viewpoints setting. Recent work resorts to integrating depth priors into outdoor NeRF training to alleviate the issue. However, the criteria for selecting depth priors and the relative merits of different priors have not been thoroughly investigated. Moreover, the relative merits of selecting different approaches to use the depth priors is also an unexplored problem. In this paper, we provide a comprehensive study and evaluation of employing depth priors to outdoor neural radiance fields, covering common depth sensing technologies and most application ways. Specifically, we conduct extensive experiments with two representative NeRF methods equipped with four commonly-used depth priors and different depth usages on two widely used outdoor datasets. Our experimental results reveal several interesting findings that can potentially benefit practitioners and researchers in training their NeRF models with depth priors. Project page: https://cwchenwang.github.io/outdoor-nerf-depth Chen Wang 0049, Jiadai Sun, Lina Liu 0010, Chenming Wu, Zhelun Shen, Dayan Wu, Yuchao Dai, Liangjun Zhang |
ACM Multimedia | 6 |
| 2023 | Handling Label Uncertainty for Camera Incremental Person Re-IdentificationabstractIncremental learning for person re-identification (ReID) aims to develop models that can be trained with a continuous data stream, which is a more practical setting for real-world applications. However, the existing incremental ReID methods make two strong assumptions that the cameras are fixed and the new-emerging data is class-disjoint from previous classes. This is unrealistic as previously observed pedestrians may re-appear and be captured again by new cameras. In this paper, we investigate person ReID in an unexplored scenario named Camera Incremental Person ReID (CIPR), which advances existing lifelong person ReID by taking into account the class overlap issue. Specifically, new data collected from new cameras may probably contain an unknown proportion of identities seen before. This subsequently leads to the lack of cross-camera annotations for new data due to privacy concerns. To address these challenges, we propose a novel framework ExtendOVA. First, to handle the class overlap issue, we introduce an instance-wise seen-class identification module to discover previously seen identities at the instance level. Then, we propose a criterion for selecting confident ID-wise candidates and also devise an early learning regularization term to correct noise issues in pseudo labels. Furthermore, to compensate for the lack of previous data, we resort prototypical memory bank to create surrogate features, along with a cross-camera distillation loss to further retain the inter-camera relationship. The comprehensive experimental results on multiple benchmarks show that ExtendOVA significantly outperforms the state-of-the-arts with remarkable advantages. Zexian Yang, Dayan Wu, Wanqian Zhang, Bo Li 0063, Weiping Wang 0005 |
ACM Multimedia | 2 |
| 2023 | Targeted Transferable Attack against Deep Hashing RetrievalabstractWith the extensive utilization of deep hashing, there exists a surging interest in studying adversarial attacks against it. Previous methods have demonstrated the superior white-box attack performance against deep hashing. However, the more challenging and realistic targeted black-box attack has not yet been explored sufficiently, which will result in an over-estimation on model robustness. In this paper, we focus on targeted black-box attack based on transferability, and propose a novel Targeted Transferable Attack method against deep hashing with Generative Adversarial Network (TTA-GAN). Specifically, we first propose a new Iterative Anchor code Optimization (IAO) method to generate anchor code with superior representative semantics of target label, which can improve both targeted white-box and black-box performances. Then, we propose a generation-based method to directly generate targeted transferable adversarial example by training a conditional generator and a discriminator. Moreover, to further promote the targeted transferability, we conduct multiple input transformations on the generated adversarial example to alleviate the overfitting phenomenon on source model. Finally, we extend our method to a novel model ensemble attack method TTA-GANens to preserve the representative semantics on multiple models, specialized for deep hashing. Extensive experiments demonstrate the superior targeted black-box attack performance than the state-of-the-art methods. Wanqian Zhang, Dayan Wu, Lin Wang 0108, Bo Li 0063, Weiping Wang 0005 |
MMAsia | 3 |
| 2023 | Decoupled Contrastive Learning for Long-Tailed Distribution
Xiaohua Chen 0002, Yucan Zhou, Lin Wang 0108, Dayan Wu, Wanqian Zhang, Bo Li 0063, Weiping Wang 0005 |
PRCV (9) | 4 |
| 2023 | Deep Uncoupled Discrete Hashing via Similarity Matrix DecompositionabstractHashing has been drawing increasing attention in the task of large-scale image retrieval owing to its storage and computation efficiency, especially the recent asymmetric deep hashing methods. These approaches treat the query and database in an asymmetric way and can take full advantage of the whole training data. Though it has achieved state-of-the-art performance, asymmetric deep hashing methods still suffer from the large quantization error and efficiency problem on large-scale datasets due to the tight coupling between the query and database. In this article, we propose a novel asymmetric hashing method, called D eep U ncoupled D iscrete H ashing (DUDH), for large-scale approximate nearest neighbor search. Instead of directly preserving the similarity between the query and database, DUDH first exploits a small similarity-transfer image set to transfer the underlying semantic structures from the database to the query and implicitly keep the desired similarity. As a result, the large similarity matrix is decomposed into two relatively small ones and the query is decoupled from the database. Then both database codes and similarity-transfer codes are directly learned during optimization. The quantization error of DUDH only exists in the process of preserving similarity between the query and similarity-transfer set. By uncoupling the query from the database, the training cost of optimizing the CNN model for the query is no longer related to the size of the database. Besides, to further accelerate the training process, we propose to optimize the similarity-transfer codes with a constant-approximation solution. In doing so, the training cost of optimizing similarity-transfer codes can be almost ignored. Extensive experiments on four widely used image retrieval benchmarks demonstrate that DUDH can achieve state-of-the-art retrieval performance with remarkable training cost reduction (30× - 50× relative). Dayan Wu, Qi Dai 0001, Bo Li 0063, Weiping Wang 0005 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2022 | Imagine by Reasoning: A Reasoning-Based Implicit Semantic Data Augmentation for Long-Tailed ClassificationabstractReal-world data often follows a long-tailed distribution, which makes the performance of existing classification algorithms degrade heavily. A key issue is that the samples in tail categories fail to depict their intra-class diversity. Humans can imagine a sample in new poses, scenes and view angles with their prior knowledge even if it is the first time to see this category. Inspired by this, we propose a novel reasoning-based implicit semantic data augmentation method to borrow transformation directions from other classes. Since the covariance matrix of each category represents the feature transformation directions, we can sample new directions from similar categories to generate definitely different instances. Specifically, the long-tailed distributed data is first adopted to train a backbone and a classifier. Then, a covariance matrix for each category is estimated, and a knowledge graph is constructed to store the relations of any two categories. Finally, tail samples are adaptively enhanced via propagating information from all the similar categories in the knowledge graph. Experimental results on CIFAR-LT-100, ImageNet-LT, and iNaturalist 2018 have demonstrated the effectiveness of our proposed method compared with the state-of-the-art methods. Xiaohua Chen 0002, Yucan Zhou, Dayan Wu, Wanqian Zhang, Yu Zhou 0015, Bo Li 0063, Weiping Wang 0005 |
AAAI | 3 |
| 2022 | Deep Piecewise Hashing for Efficient Hamming Space RetrievalabstractHamming space retrieval can achieve constant-time image search, which is more efficient than linear scan. In Hamming space retrieval, the data points inside the Hamming ball imply retrievable while the data points outside are irretrievable. Therefore, it is crucial to explicitly characterize the Hamming ball. However, for the existing Hamming space retrieval methods, many similar points are found close to the outside of the Hamming ball while many dissimilar points are found close to the query point, leading to the decline of both retrieval accuracy and recall. In this paper, we present a novel method named Deep Piecewise Hashing (DPH), for Efficient Hamming Space Retrieval. A piecewise loss is elaborately designed to guide the learning of hash codes. Meanwhile, a piecewise probability distribution is introduced in the proposed loss function. The piecewise probability distribution pays more attention to the learning of those "marginal" similar points. It considers both discrimination and robustness for the dissimilar points inside the Hamming ball. Comprehensive experiments on two datasets, MS-COCO and NUS-WIDE, demonstrate that DPH can yield state-of-the-art Hamming space retrieval performance. Jingzi Gu, Dayan Wu, Peng Fu 0008, Bo Li 0063, Weiping Wang 0005 |
ICASSP | 2 |
| 2022 | Prototype-Based Inter-Camera Learning for Person Re-IdentificationabstractPerson re-identification (ReID) aims at retrieving images of the same person across non-overlapping camera views. The prior works focus on either fully supervised or unsupervised ReID settings, and achieve remarkable performances. In real scenarios, however, the major annotation cost comes from matching identity classes across camera views, thus leading to the Intra-Camera Supervised (ICS) ReID problem. In this work, we propose a Prototype-based Inter-camera ReID (PIRID) method, which tackles the ICS setting through the lens of prototype learning. Specifically, we first introduce the intra-camera learning with non-parametric classifiers to separately generate discriminative features within each camera view. Moreover, the inter-camera prototype learning provides prototypes as the representatives of each class in the common space, making the learned features to be camera-agnostic. Experiments conducted on three benchmarks, i.e., Market-1501, DukeMTMC-ReID, and MSMT17, show the superiority of our method. Lin Wang 0108, Wanqian Zhang, Dayan Wu, Pingting Hong, Bo Li 0063 |
ICASSP | 3 |
| 2022 | Clustering and Separating Similarities for Deep Unsupervised HashingabstractThe lack of supervised information is the pivotal problem in unsupervised hashing. Most methods leverage deep features extracted from pre-trained models to generate semantic similarities as supervised information. These fixed features are, however, neither designed originally for retrieval nor updated adaptively during training. In this paper, we propose a novel deep Unsupervised Cluster and Separate Hashing (UCSH) to address these issues. Specifically, we introduce a fully end-to-end deep hashing network with a binary latent Variational AutoEncoder (VAE), which enables hash codes capable of reconstructing deep features as well as preserving semantic relations. Moreover, a ‘Cluster and Separate’ scheme is proposed to jointly cluster deep features and separate semantic similarities. Both the implicit feature clustering and the explicit similarity separating loss encourage the separation of similar and dissimilar pairs, enabling the iteratively updated similarities to better excavate semantic relations. Experiments conducted on three benchmarks show the superiority of UCSH. Wanqian Zhang, Dayan Wu, Chule Yang, Bo Li 0063, Weiping Wang 0005 |
ICASSP | 2 |
| 2022 | Listen and Look: Multi-Modal Aggregation and Co-Attention Network for Video-Audio RetrievalabstractVideo is a natural source of multi-modal data with intrinsic correlations between different modalities, such as objects, motions and captions. Though intuitive, such inherent supervision has not been well explored in previous video-audio retrieval works. Besides, existing methods exploit the video stream and the audio stream seperately, whereas ignoring the mutual interactions between them. In this paper, we propose a two-stream model named Multi-modal Aggregation and Co-attention network (MAC), which processes video and audio inputs with co-attentional interactions. Specifically, our method takes raw videos as inputs and extracts aggregated features from multiple modalities to benefit the video representation learning. Then, we introduce the self-attention mechanism to make videos adaptively assign higher weights to the representative modalities. Moreover, we introduce a coattention transformer module to better capture the relations among videos and audios. By exchanging key-value pairs in the multi-headed attention, this module enables video-attended audio features to be incorporated into video representations and vice versa. Experiments show that our method significantly outperform other state-of-the-arts. Xiaoshuai Hao, Wanqian Zhang, Dayan Wu, Bo Li 0063 |
ICME | 3 |
| 2022 | Camera-specific Informative Data Augmentation Module for Unbalanced Person Re-identificationabstractPerson re-identification~(Re-ID) aims at retrieving the same person across the non-overlapped camera networks. Recent works have achieved impressive performance due to the rapid development of deep learning techniques. However, most existing methods have ignored the practical unbalanced property in real-world Re-ID scenarios. In fact, the number of pedestrian images in different cameras vary a lot. Some cameras cover thousands of images while others only have a few. As a result, the camera-unbalanced problem will reduce intra-camera diversity, then the model cannot learn camera-invariant features to distinguish pedestrians from "poor" cameras. In this paper, we design a novel camera-specific informative data augmentation module~(CIDAM) to alleviate the proposed camera-unbalanced problem. Specifically, we first calculate the camera-specific distribution online, then refine the "poor" camera-specific covariance matrix with similar cameras defined in the prototype-based similarity matrix. Consequently, informative augmented samples are generated by combining original samples with sampled random vectors in feature space. To ensure these augmented samples can better benefit the model training, we further propose a dynamic-threshold-based contrastive loss. Since augmented samples may not be as real as original ones, we calculate a threshold for each original one dynamically and only push hard negative augmented samples away. Moreover, our CIDAM can be compatible with a variety of existing Re-ID methods. Extensive experiments prove the effectiveness of our method. Pingting Hong, Dayan Wu, Bo Li 0063, Weiping Wang 0005 |
ACM Multimedia | 2 |
| 2022 | TPSNet: Reverse Thinking of Thin Plate Splines for Arbitrary Shape Scene Text RepresentationabstractThe research focus of scene text detection and recognition has shifted to arbitrary shape text in recent years, where the text shape representation is a fundamental problem. An ideal representation should be compact, complete, efficient, and reusable for subsequent recognition in our opinion. However, previous representations have flaws in one or more aspects. Thin-Plate-Spline (TPS) transformation has achieved great success in scene text recognition. Inspired by this, we reversely think of its usage and sophisticatedly take TPS as an exquisite representation for arbitrary shape text representation. The TPS representation is compact, complete, and efficient. With the predicted TPS parameters, the detected text region can be directly rectified to a near-horizontal one to assist the subsequent recognition. To further exploit the potential of the TPS representation, the Border Alignment Loss is proposed. Based on these designs, we implement the text detector TPSNet, which can be extended to a text spotter conveniently. Extensive evaluation and ablation of several public benchmarks demonstrate the effectiveness and superiority of the proposed method for text representation and spotting. Particularly, TPSNet achieves the detection F-Measure improvement of 4.4% (78.4% vs. 74.0%) on Art dataset and the end-to-end spotting F-Measure improvement of 5.0% (78.5% vs. 73.5%) on Total-Text, which are large margins with no bells and whistles. The source code will be available. Wei Wang 0315, Yu Zhou 0015, Jiahao Lyu 0002, Dayan Wu, Weiping Wang 0005 |
ACM Multimedia | 4 |
| 2022 | Attack is the Best Defense: Towards Preemptive-Protection Person Re-IdentificationabstractPerson Re-IDentification (ReID) aims at retrieving images of the same person across multiple camera views. Despite its popularity in surveillance and public safety, the leakage of identity information is still at risk. For example, once obtaining the illegal access to ReID systems, malicious user can accurately retrieve the target person, leading to the exposure of private information. Recently, some pioneering works protect private images with adversarial examples by adding imperceptible perturbations to target images. However, in this paper, we argue that directly applying adversary-based methods to protect the ReID system is sub-optimal due to the 'overlap identity' issue. Specifically, merely pushing the adversarial image away from its original label would probably make it move into the vicinity of other identities. This leads to the potential risk of being retrieved when querying with all the other identities exhaustively. We thus propose a novel preemptive-Protection person Re-IDentification (PRIDE) method. By explicitly constraining the adversarial image to an isolated location, the target person is far away from neither the original identity nor any other identities, which protects him from being retrieved by illegal queries. Moreover, we further propose two crucial attack scenarios (Random Attack and Order Attack) and a novel Success Protection Rate (SPR) metric to quantify the protection ability. Experiments show consistent outperformance of our method over other baselines across different ReID models, datasets and attack scenarios. Lin Wang 0108, Wanqian Zhang, Dayan Wu, Bo Li 0063 |
ACM Multimedia | 3 |
| 2022 | Efficient Hash Code Expansion by Recycling Old BitsabstractDeep hashing methods have been intensively studied and successfully applied in large-scale multimedia retrieval. In real-world scenarios, code length can not be set once for all if retrieval accuracy is not satisfying. However, when code length increases, conventional deep hashing methods have to retrain their models and regenerate the whole database codes, which is impractical for large-scale retrieval system. In this paper, we propose an interesting deep hashing method from a brand new perspective, called Code Expansion oriented Deep Hashing (CEDH). Different from conventional deep hashing methods, our CEDH focuses on the fast expansion of existing hash codes. Instead of regenerating all bits from raw images, the new bits in CEDH can be incrementally learned by recycling the old ones. Specifically, we elaborately design an end-to-end asymmetric framework to simultaneously optimize a CNN model for query images and a code projection matrix for database images. With the learned code projection matrix, hash codes can achieve fast expansion through simple matrix multiplication. Subsequently, a novel code expansion hashing loss is proposed to preserve the similarities between query codes and expanded database codes. Due to the loose coupling in our framework, our CEDH is compatible with a variety of deep hashing methods. Moreover, we propose to adopt smooth similarity matrix to solve "similarity contradiction" problem existing in multi-label image datasets, thus further improving our performance on multi-label datasets. Extensive experiments on three widely used image retrieval benchmarks demonstrate that CEDH can significantly reduce the cost for expanding database codes (about 100,000x faster with GPU and 1,000,000x faster with CPU) when code length increases while keeping the state-of-the-art retrieval accuracy. Our code is available at https://github.com/IIE-MMR/2022MM-CEDH. Dayan Wu, Qinghang Su, Bo Li 0063, Weiping Wang 0005 |
ACM Multimedia | 1 |
| 2022 | Cross-view vehicle re-identification based on graph matching
Chao Zhang 0074, Chule Yang, Dayan Wu, Hongbin Dong |
Appl. Intell. | 3 |
| 2022 | Multi-View correlation distillation for incremental object detection
Dongbao Yang, Yu Zhou 0015, Aoting Zhang, Xurui Sun, Dayan Wu, Weiping Wang 0005, Qixiang Ye |
Pattern Recognit. | 5 |
| 2022 | Uncertainty-Aware and Multigranularity Consistent Constrained Model for Semi-Supervised HashingabstractRecently, deep semi-supervised hashing methods have attracted increasing attention, which can significantly improve retrieval performance by leveraging abundant unlabeled data. These methods usually generate surrogate supervision signals to learn with unlabeled data, such as neighborhood information and augmentation invariant requirements. However, an essential issue of these methods is that the supervised signals are not always reliable, which may damage the performance. In this paper, we propose a novel Uncertainty-Aware and Multi-Granularity Consistent Constrained Semi-Supervised Hashing (UMCSH) method to alleviate the negative effects of noisy supervised signals and enlarge the inter-class distance. Specifically, our UMCSH mainly consists of an Uncertainty-Aware Instance-Level Consistency (UAILC) model and a Cluster-Based Class-Level Consistency (CBCLC) model. UAILC introduces an uncertainty estimation method to select reliable supervised signals to extract discriminative features for each unlabeled data. CBCLC establishes connections between labeled data and unlabeled data by encouraging each unlabeled sample to be close to the hash center (calculated with the labeled data) according to its pseudo-label. Extensive experimental results demonstrate the superior performance of our proposed approach compared with several state-of-the-art semi-supervised hashing methods. Shuai Cheng 0002, Yucan Zhou, Wanqian Zhang, Dayan Wu, Chule Yang, Bo Li 0063, Weiping Wang 0005 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Exploring Relations in Untrimmed Videos for Self-Supervised LearningabstractExisting video self-supervised learning methods mainly rely on trimmed videos for model training. They apply their methods and verify the effectiveness on trimmed video datasets including UCF101 and Kinetics-400, among others. However, trimmed datasets are manually annotated from untrimmed videos. In this sense, these methods are not truly unsupervised. In this article, we propose a novel self-supervised method, referred to as Exploring Relations in Untrimmed Videos (ERUV), which can be straightforwardly applied to untrimmed videos (real unlabeled) to learn spatio-temporal features. ERUV first generates single-shot videos by shot change detection. After that, some designed sampling strategies are used to model relations for video clips. The strategies are saved as our self-supervision signals. Finally, the network learns representations by predicting the category of relations between the video clips. ERUV is able to compare the differences and similarities of video clips, which is also an essential procedure for video-related tasks. We validate our learned models with action recognition, video retrieval, and action similarity labeling tasks with four kinds of 3D convolutional neural networks. Experimental results show that ERUV is able to learn richer representations with untrimmed videos, and it outperforms state-of-the-art self-supervised methods with significant margins. Dezhao Luo, Yu Zhou 0015, Bo Fang 0003, Yucan Zhou, Dayan Wu, Weiping Wang 0005 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2022 | RD-IOD: Two-Level Residual-Distillation-Based Triple-Network for Incremental Object DetectionabstractAs a basic component in multimedia applications, object detectors are generally trained on a fixed set of classes that are pre-defined. However, new object classes often emerge after the models are trained in practice. Modern object detectors based on Convolutional Neural Networks (CNN) suffer from catastrophic forgetting when fine-tuning on new classes without the original training data. Therefore, it is critical to improve the incremental learning capability on object detection. In this article, we propose a novel Residual-Distillation-based Incremental learning method on Object Detection (RD-IOD). Our approach rests on the creation of a triple-network based on Faster R-CNN. To enable continuous learning from new classes, we use the original model as well as a residual model to guide the learning of the incremental model on new classes while maintaining the previous learned knowledge. To better maintain the discrimination between the features of old and new classes, the residual model is jointly trained with the incremental model on new classes in the incremental learning procedure. In addition, a two-level distillation scheme is designed to guide the training process, which consists of (1) a general distillation for imitating the original model in feature space along with a residual distillation on the features in both image level and instance level, and (2) a joint classification distillation on the output layers. To well preserve the learned knowledge, we design a 2-threshold training strategy to guide the learning of a Region Proposal Network and a detection head. Extensive experiments conducted on VOC2007 and COCO demonstrate that the proposed method can effectively learn to incrementally detect objects of new classes, and the problem of catastrophic forgetting is mitigated. Our code is available at https://github.com/yangdb/RD-IOD. Dongbao Yang, Yu Zhou 0015, Wei Shi 0001, Dayan Wu, Weiping Wang 0005 |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2021 | FC2RN: A Fully Convolutional Corner Refinement Network for Accurate Multi-Oriented Scene Text DetectionabstractAccurate detection of multi-oriented text that accounts for a large proportion in real practice is of great significance. The performance has improved rapidly on common benchmarks in recent years. However, dense long text case and the quality of detection are easy to be overlooked. Direct regression may produce low-quality and incomplete detections due to the constrain of the receptive field; proposal-based methods could alleviate this but might introduce redundant context due to RoI operation, degrading the performance. To address the dilemma, a novel proposed corner-aware convolution in which the sampling positions tightly cover the text area is utilized to encode an initial corner prediction into the feature maps, which can be further used to produce a refined corner prediction. We embed the proposed module into an anchor-free baseline model, leading to a simple and effective fully convolutional corner refinement network (FC2RN). Experimental results on four public datasets including MSRATD500, ICDAR2015, RCTW-17, and COCO-Text demonstrate that FC2RN can outperform state-of-the-art methods. Xugong Qin, Yu Zhou 0015, Youhui Guo, Dayan Wu, Weiping Wang 0005 |
ICASSP | 4 |
| 2021 | Disturbance Consistent Self-Ensembling for Semi-Supervised HashingabstractRecently, deep semi-supervised hashing methods have attracted increasing attention, where the visual similarity of unlabeled data is usually adopted to guide the hash codes learning. However, samples with similar appearance may come from different categories, making their hash codes similar will lead to sub-optimal retrieval results. In this paper, we propose a novel Disturbance Consistent Self-Ensembling (DCSE) method to alleviate the drawback of visual similarity constraint. Specially, DCSE forms consensus hash codes for the same sample under different augmentations. These ensemble hash codes can capture the discriminative characteristics of a sample. Therefore, as more augmented data is involved, more ensemble hash codes in one category can become similar gradually. Then, we design a disturbance consistent loss to learn the discriminative hash codes by minimizing the distance between outputs of the hash layer and ensemble hash codes. Extensive experiments show that our proposed approach significantly outperforms state-of-the-art semi-supervised hashing methods. Shuai Cheng 0002, Yucan Zhou, Dayan Wu, Haisu Zhang, Bo Li 0063, Weiping Wang 0005 |
ICME | 3 |
| 2021 | What Matters: Attentive and Relational Feature Aggregation Network for Video-Text RetrievalabstractCross-modal video-text retrieval has been an emerging task due to the rapid growth of user-generated videos on the Internet. Most existing approaches focus on extracting visual feature for the video, while audio and caption on the screen containing rich information are ignored. Recently, the aggregations of multi-modal features in videos boost the benchmark of video-text retrieval. However, since these multi-modal features are high-dimensional and heterogeneous, their intrinsically structural relations have not been attached with enough importance and are often overlooked in previous methods. To address this issue, we propose a novel Attentive and Relational Feature Aggregation Network (ARFAN). Specifically, we introduce the self-attention mechanism to make videos adaptively assign higher weights to the representative modalities. Then, the graph convolutional layers are inserted to capture the relations among the multi-modal features to combine them. Our method achieves 15% and 12.9% relative improvements on R@1 when compared with the state-of-the-art method on MSR-VTT and MSVD datasets, respectively. Xiaoshuai Hao, Yucan Zhou, Dayan Wu, Wanqian Zhang, Bo Li 0063, Weiping Wang 0005, Dan Meng 0002 |
ICME | 3 |
| 2021 | Rescuing Deep Hashing from Dead Bits ProblemabstractDeep hashing methods have shown great retrieval accuracy and efficiency in large-scale image retrieval. How to optimize discrete hash bits is always the focus in deep hashing methods. A common strategy in these methods is to adopt an activation function, e.g. sigmoid() or tanh(), and minimize a quantization loss to approximate discrete values. However, this paradigm may make more and more hash bits stuck into the wrong saturated area of the activation functions and never escaped. We call this problem "Dead Bits Problem (DBP)". Besides, the existing quantization loss will aggravate DBP as well. In this paper, we propose a simple but effective gradient amplifier which acts before activation functions to alleviate DBP. Moreover, we devise an error-aware quantization loss to further alleviate DBP. It avoids the negative effect of quantization loss based on the similarity between two images. The proposed gradient amplifier and error-aware quantization loss are compatible with a variety of deep hashing methods. Experimental results on three datasets demonstrate the efficiency of the proposed gradient amplifier and the error-aware quantization loss. Shu Zhao 0006, Dayan Wu, Yucan Zhou, Bo Li 0063, Weiping Wang 0005 |
IJCAI | 2 |
| 2021 | Multi-Feature Graph Attention Network for Cross-Modal Video-Text RetrievalabstractCross-modal retrieval between videos and texts has attracted growing attention due to the rapid growth of user-generated videos on the web. To solve this problem, most approaches try to learn a joint embedding space to measure the cross-modal similarities, while paying little attention to the representation of each modality. Video is more complicated than the commonly used visual feature, since the audio and caption on the screen also contain rich information. Recently, the aggregations of multiple features in videos boost the benchmark of the video-text retrieval system. However, they usually handle each feature independently, which ignores the interchange of high-level semantic relations among these multiple features. Moreover, despite the inter-modal ranking constraint where semantically-similar texts and videos should stay closer, the modality-specific requirement, i.e. two similar videos/texts should have similar representations, is also significant. In this paper, we propose a novel Multi-Feature Graph ATtention Network (MFGATN) for cross-modal video-text retrieval. Specifically, we introduce a multi-feature graph attention module, which enriches the representation of each feature in videos with the interchange of high-level semantic information among them. Moreover, we elaborately design a novel Dual Constraint Ranking Loss (DCRL), which simultaneously considers the inter-modal ranking constraint and the intra-modal structure constraint to preserve both the cross-modal semantic similarity and the modality-specific consistency in the embedding space. Experiments on two datasets, i.e. MSR-VTT and MSVD, demonstrate that our method achieves significant performance gain compared with the state-of-the-arts. Xiaoshuai Hao, Yucan Zhou, Dayan Wu, Wanqian Zhang, Bo Li 0063, Weiping Wang 0005 |
ICMR | 3 |
| 2021 | Mask is All You Need: Rethinking Mask R-CNN for Dense and Arbitrary-Shaped Scene Text DetectionabstractDue to the large success in object detection and instance segmentation, Mask R-CNN attracts great attention and is widely adopted as a strong baseline for arbitrary-shaped scene text detection and spotting. However, two issues remain to be settled. The first is dense text case, which is easy to be neglected but quite practical. There may exist multiple instances in one proposal, which makes it difficult for the mask head to distinguish different instances and degrades the performance. In this work, we argue that the performance degradation results from the learning confusion issue in the mask head. We propose to use an MLP decoder instead of the "deconv-conv" decoder in the mask head, which alleviates the issue and promotes robustness significantly. And we propose instance-aware mask learning in which the mask head learns to predict the shape of the whole instance rather than classify each pixel to text or non-text. With instance-aware mask learning, the mask branch can learn separated and compact masks. The second is that due to large variations in scale and aspect ratio, RPN needs complicated anchor settings, making it hard to maintain and transfer across different datasets. To settle this issue, we propose an adaptive label assignment in which all instances especially those with extreme aspect ratios are guaranteed to be associated with enough anchors. Equipped with these components, the proposed method named MAYOR achieves state-of-the-art performance on five benchmarks including DAST1500, MSRA-TD500, ICDAR2015, CTW1500, and Total-Text. Xugong Qin, Yu Zhou 0015, Youhui Guo, Dayan Wu, Zhihong Tian 0001, Weiping Wang 0005 |
ACM Multimedia | 4 |
| 2021 | Binary Neural Network Hashing for Image RetrievalabstractHashing has become increasingly important for large-scale image retrieval, of which the low storage cost and fast searching are two key properties. However, existing methods adopt large neural networks, which are hard to be deployed in resource-limited devices due to the unacceptable memory and runtime overhead. We address that this huge overhead of neural networks somewhatviolates the appealing properties of hashing. In this paper, we propose a novel deep hashing method, called Binary Neural Network Hashing (BNNH) for fast image retrieval. Specifically, we construct an efficient binarized network architecture to provide lighter model and faster inference, which directly generates binary outputs as the desired hash codes without introducing the quantization loss. Besides, in order to circumvent the huge performance degradation caused by the extremely quantized activations, we introduce a simple yet effective activation-aware loss to explicitly guide the updating of activations in intermediate layers. Extensive experiments conducted on three benchmarks show that the proposed method outperforms the state-of-the-art binarization methods by large margins and validate the efficiency of BNNH. Wanqian Zhang, Dayan Wu, Yu Zhou 0015, Bo Li 0063, Weiping Wang 0005, Dan Meng 0002 |
SIGIR | 2 |
| 2020 | Marginalized Graph Attention Hashing for Zero-Shot Image Retrieval
Meixue Huang, Dayan Wu, Wanqian Zhang, Bo Li 0063, Weiping Wang 0005 |
BMVC | 2 |
| 2020 | Deep Discrete Attention Guided Hashing for Face Image RetrievalabstractRecently, face image hashing has been proposed in large-scale face image retrieval due to its storage and computational efficiency. However, owing to the large intra-identity variation (same identity with different poses, illuminations, and facial expressions) and the small inter-identity separability (different identities look similar) of face images, existing face image hashing methods have limited power to generate discriminative hash codes. In this work, we propose a deep hashing method specially designed for face image retrieval named deep Discrete Attention Guided Hashing (DAGH). In DAGH, the discriminative power of hash codes is enhanced by a well-designed discrete identity loss, where not only the separability of the learned hash codes for different identities is encouraged, but also the intra-identity variation of the hash codes for the same identities is compacted. Besides, to obtain the fine-grained face features, DAGH employs a multi-attention cascade network structure to highlight discriminative face features. Moreover, we introduce a discrete hash layer into the network, along with the proposed modified backpropagation algorithm, our model can be optimized under discrete constraint. Experiments on two widely used face image retrieval datasets demonstrate the inspiring performance of DAGH over the state-of-the-art face image hashing methods. Dayan Wu, Wen Gu, Haisu Zhang, Bo Li 0063, Weiping Wang 0005 |
ICMR | 2 |
| 2020 | Deep Semantic-Alignment Hashing for Unsupervised Cross-Modal RetrievalabstractDeep hashing methods have achieved tremendous success in cross-modal retrieval, due to its low storage consumption and fast retrieval speed. In real cross-modal retrieval applications, it's hard to obtain label information. Recently, increasing attention has been paid to unsupervised cross-modal hashing. However, existing methods fail to exploit the intrinsic connections between images and their corresponding descriptions or tags (text modality). In this paper, we propose a novel Deep Semantic-Alignment Hashing (DSAH) for unsupervised cross-modal retrieval, which sufficiently utilizes the co-occurred image-text pairs. DSAH explores the similarity information of different modalities and we elaborately design a semantic-alignment loss function, which elegantly aligns the similarities between features with those between hash codes. Moreover, to further bridge the modality gap, we innovatively propose to reconstruct features of one modality with hash codes of the other one. Extensive experiments on three cross-modal retrieval datasets demonstrate that DSAH achieves the state-of-the-art performance. Dejie Yang, Dayan Wu, Wanqian Zhang, Haisu Zhang, Bo Li 0063, Weiping Wang 0005 |
ICMR | 2 |
| 2020 | Deep Unsupervised Hybrid-similarity Hadamard HashingabstractHashing has become increasingly important for large-scale image retrieval. Recently, deep supervised hashing has shown promising performance, yet little work has been done under the more realistic unsupervised setting. The most challenging problem in unsupervised hashing methods is the lack of supervised information. Besides, existing methods fail to distinguish image pairs with different similarity degrees, which leads to a suboptimal construction of similarity matrix. In this paper, we propose a simple yet effective unsupervised hashing method, dubbed Deep Unsupervised Hybrid-similarity Hadamard Hashing (DU3H), which tackles these issues in an end-to-end deep hashing framework. DU3H employs orthogonal Hadamard codes to provide auxiliary supervised information in unsupervised setting, which can maximally satisfy the independence and balance properties of hash codes. Moreover, DU3H utilizes both highly and normally confident image pairs to jointly construct a hybrid-similarity matrix, which can magnify the impacts of different pairs to better preserve the semantic relations between images. Extensive experiments conducted on three widely used benchmarks validate the superiority of DU3H. Wanqian Zhang, Dayan Wu, Yu Zhou 0015, Bo Li 0063, Weiping Wang 0005, Dan Meng 0002 |
ACM Multimedia | 2 |
| 2020 | Asymmetric Deep Hashing for Efficient Hash Code CompressionabstractBenefiting from recent advances in deep learning, deep hashing methods have achieved promising performance in large-scale image retrieval. To improve storage and computational efficiency, existing hash codes need to be compressed accordingly. However, previous deep hashing methods have to retrain their models and then regenerate the whole database codes using the new models when code length changes, which is time consuming especially for large image databases. In this paper, we propose a novel deep hashing method, called Code Compression oriented Deep Hashing (CCDH), for efficiently compressing hash codes. CCDH learns deep hash functions for query images, while learning a one-hidden-layer Variational Autoencoder (VAE) from existing hash codes. With such asymmetric design, CCDH can efficiently compress database codes only using the learned encoder of VAE. Furthermore, CCDH is flexible enough to be used with a variety of deep hashing methods. Extensive experiments on three widely used image retrieval benchmarks demonstrate that CCDH can significantly reduce the cost for compressing database codes when code length changes while keeping the state-of-the-art retrieval accuracy. Shu Zhao 0006, Dayan Wu, Wanqian Zhang, Yu Zhou 0015, Bo Li 0063, Weiping Wang 0005 |
ACM Multimedia | 2 |
| 2019 | Fast and Multilevel Semantic-Preserving Discrete Hashing
Wanqian Zhang, Dayan Wu, Jing Liu 0034, Bo Li 0063, Xiaoyan Gu 0001, Weiping Wang 0005, Dan Meng 0002 |
BMVC | 2 |
| 2019 | Deep Incremental Hashing Network for Efficient Image RetrievalabstractHashing has shown great potential in large-scale image retrieval due to its storage and computation efficiency, especially the recent deep supervised hashing methods. To achieve promising performance, deep supervised hashing methods require a large amount of training data from different classes. However, when images of new categories emerge, existing deep hashing methods have to retrain the CNN model and generate hash codes for all the database images again, which is impractical for large-scale retrieval system. In this paper, we propose a novel deep hashing framework, called Deep Incremental Hashing Network (DIHN), for learning hash codes in an incremental manner. DIHN learns the hash codes for the new coming images directly, while keeping the old ones unchanged. Simultaneously, a deep hash function for query set is learned by preserving the similarities between training points. Extensive experiments on two widely used image retrieval benchmarks demonstrate that the proposed DIHN framework can significantly decrease the training time while keeping the state-of-the-art retrieval accuracy. Dayan Wu, Qi Dai 0001, Jing Liu 0034, Bo Li 0063, Weiping Wang 0005 |
CVPR | 1 |
| 2018 | Deep Uniqueness-Aware Hashing for Fine-Grained Multi-Label Image RetrievalabstractDeep supervised hashing methods for multi-label image retrieval have achieved great success nowadays. However, these methods only take the similarity between the database images and the query images into account, but they ignore the uniqueness of the database images when deciding on their rankings. Here we present a novel Deep Uniqueness-Aware Hashing (DUAH) method for learning hash functions that preserve not only multilevel semantic similarity between multi-label images, but also the unique semantic structure of each image. In our approach, both the pairwise label information and the classification information are fully exploited to maximize the discriminability of the output binary codes within one stream framework. Extensive evaluations conducted on three widely used multi-label image benchmarks demonstrate that DUAH can support fine-grained multi-label image retrieval better. Dayan Wu, Zheng Lin 0001, Bo Li 0063, Jing Liu 0034, Weiping Wang 0005 |
ICASSP | 1 |
| 2018 | Deep Index-Compatible Hashing for Fast Image RetrievalabstractDeep hashing methods have achieved promising results for large-scale image retrieval recently. To accelerate the subsequent Hamming ranking process, the multi-index approach has been proposed to reduce the computations for the Hamming distance. However, the binary codes output by the previous deep hashing methods may not be optimally compatible with the multi-index approach. In this paper, we present a novel Deep Index-Compatible Hashing (DICH) method for fast image retrieval, which can learn similarity-preserving binary codes that are more compatible with the multi-index approach. With the learned binary codes, both the size of the intermediate result set produced by the multi-index approach and the number of the candidate images can be reduced, which can accelerate the Hamming ranking process. By taking advantage of the unique feature of DICH, we further propose a block-based ranking strategy to quickly rank the candidate images without calculating the Hamming distance. Extensive evaluations demonstrate that the proposed method can significantly reduce the retrieval time with almost no loss of retrieval accuracy. Dayan Wu, Jing Liu 0034, Bo Li 0063, Weiping Wang 0005 |
ICME | 1 |
| 2017 | Deep Supervised Hashing for Multi-Label and Large-Scale Image RetrievalabstractOne of the most challenging tasks in large-scale multi-label image retrieval is to map images into binary codes while preserving multilevel semantic similarity. Recently, several deep supervised hashing methods have been proposed to learn hash functions that preserve multilevel semantic similarity with deep convolutional neural networks. However, these triplet label based methods try to preserve the ranking order of images according to their similarity degrees to the queries while not putting direct constraints on the distance between the codes of very similar images. Besides, the current evaluation criteria are not able to measure the performance of existing hashing methods on preserving fine-grained multilevel semantic similarity. To tackle these issues, we propose a novel Deep Multilevel Semantic Similarity Preserving Hashing (DMSSPH) method to learn compact similarity-preserving binary codes for the huge body of multi-label image data with deep convolutional neural networks. In our approach, we make the best of the supervised information in the form of pairwise labels to maximize the discriminability of output binary codes. Extensive evaluations conducted on several benchmark datasets demonstrate that the proposed method significantly outperforms the state-of-the-art supervised and unsupervised hashing methods at the accuracies of top returned images, especially for shorter binary codes. Meanwhile, the proposed method shows better performance on preserving fine-grained multilevel semantic similarity according to the results under the Jaccard coefficient based evaluation criteria we propose. Dayan Wu, Zheng Lin 0001, Bo Li 0063, Mingzhen Ye, Weiping Wang 0005 |
ICMR | 1 |