EDBT 2026 Demo / reviewers in the wild / expert
Lei Zhang 0119
dblp:97/8704-119
· DBLP profile ↗
46ranked-venue papers
6as first author
35since 2021 · last 2026
0000-0002-2839-8693ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 37 · 6 first-author · 27 since 2021Artificial intelligence and machine learning · 18 · 1 first-author · 16 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LayerEdit: Disentangled Multi-Object Editing via Conflict-Aware Multi-Layer LearningabstractText-driven multi-object image editing which aims to precisely modify multiple objects within an image based on text descriptions, has recently attracted considerable interest. Existing works primarily follow the localize-editing paradigm, focusing on independent object localization and editing while neglecting critical inter-object interactions. However, this work points out that the neglected attention entanglements in inter-object conflict regions, inherently hinder disentangled multi-object editing, leading to either inter-object editing leakage or intra-object editing constraints. We thereby propose a novel multi-layer disentangled editing framework LayerEdit, a training-free method which, for the first time, through precise object-layered decomposition and coherent fusion, enables conflict-free object-layered editing. Specifically, LayerEdit introduces a novel “decompose-editing-fusion” framework, consisting of: (1) Conflict-aware Layer Decomposition module, which utilizes an attention-aware IoU scheme and time-dependent region removing, to enhance conflict awareness and suppression for layer decomposition. (2) Object-layered Editing module, to establish coordinated intra-layer text guidance and cross-layer geometric mapping, achieving disentangled semantic and structural modifications. (3) Transparency-guided Layer Fusion module, to facilitate structure-coherent inter-object layer fusion through precise transparency guidance learning. Extensive experiments verify the superiority of LayerEdit over existing methods, showing unprecedented intra-object controllability and inter-object coherence in complex multi-object scenarios. Fengyi Fu, Mengqi Huang, Lei Zhang 0119, Zhendong Mao 0001 |
AAAI | 3 |
| 2026 | Toward Accurate Image Generation via Dynamic Generative Image TransformerabstractExisting generative image transformers follow a two-stage generation paradigm, where the first stage learns a codebook to encode images into discrete codes via vector quantization, and the second stage completes the image generation based on the learned codebook. However, existing methods ignore the naturally varying information densities across different image regions and indiscriminately encode fixed-size regions into fixed-length codes, resulting in insufficient encoding in important regions and redundant encoding in unimportant ones, which degrades both the image generation quality and speed. To address this challenge, we propose a novel information-density-based variable-length image coding and generation framework. In the first stage, our Dynamic Quantization VAE++ (DQVAE++) performs information-adaptive encoding by assigning variable-length codes to image regions according to their information densities, yielding more accurate and robust code representations. In the second stage, the Dynamic Generative Image Transformer (DGiT) enables information-adaptive image generation in both autoregressive and non-autoregressive manners. Specifically, for autoregressive (AR) generation, DGiT-AR generates images autoregressively from coarse-grained regions (smooth areas with fewer codes) to fine-grained regions (detailed areas with more codes). This is accomplished through a novel stacked-transformer architecture that alternately models the position and content of image codes, and a novel heterogeneous embedding scheme to distinguish codes of different granularities. Similarly, for non-autoregressive (NAR) generation, DGiT-NAR introduces a novel information-prioritized mask scheduling mechanism, prioritizing the generation of key structural regions with higher information density. This enables more coherent modeling of global structures initially, followed by a more effective synthesis of local details subsequently. Comprehensive experiments on unconditional and conditional image generation validate the superiority of our proposed variable-length coding in both effectiveness and efficiency. Zhendong Mao 0001, Mengqi Huang, Yijing Lin, Quan Wang 0002, Lei Zhang 0119, Yongdong Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2026 | Subjective-Objective Emotion-Correlated Generation Network for Subjective Video CaptioningabstractThe emotional video captioning (EVC) task, which aims to generate factual descriptions based on the perceived subtle visual emotion cues, has received more and more attention and research. However, EVC is essentially an objective video captioning task, and ignores the subjective emotional reactions of video viewers, which cannot reflect personalized affective understandings of different viewers on the same video. To fill the research gap, we investigate the subjective video captioning (SVC) task in this paper, which aims to generate emotional captions by incorporating viewers' personalized emotional reactions upon the EVC task. SVC is extremely challenging, which lies in two aspects: 1) the correlative emotion perception between subjective and objective emotions and 2) the collaborative generation between emotional and factual information. To this end, we propose the Subjective-Objective Emotion-Correlated Generation Network (SO-ECGN) in this paper. Specifically, our SO-ECGN leverages the proposed dynamic mask attention and emotion domain shifting module to achieve the objective emotion incremental learning, and then, a subjective-objective emotions correlation module is proposed to adaptively combine two perspective emotions to provide accurate emotion guidance (i.e., emotional polarity and intensity) for each generation step. Furthermore, an emotion-correlated decoder is proposed to generate subjective captions by adaptively referring to factual information and emotional information. Extensive experiments on three challenging datasets demonstrate the superiority of our approach and each proposed module, i.e., reaching 79.2%, 45.1% on BLEU-1, CIDEr metrics on EmVidCap-L dataset. Weidong Chen 0013, Cheng Ye 0004, Peipei Song, Lei Zhang 0119, Yongdong Zhang 0001, Zhendong Mao 0001 |
IEEE Trans. Image Process. | 4 |
| 2025 | CustomContrast: A Multilevel Contrastive Perspective for Subject-Driven Text-to-Image CustomizationabstractSubject-driven text-to-image (T2I) customization has drawn significant interest in academia and industry. This task enables pre-trained models to generate novel images based on unique subjects. Existing studies adopt a self-reconstructive perspective, focusing on capturing all details of a single image, which will misconstrue the specific image's irrelevant attributes (e.g., view, pose, and background) as the subject intrinsic attributes. This misconstruction leads to both overfitting or underfitting of irrelevant and intrinsic attributes of the subject, i.e., these attributes are over-represented or under-represented simultaneously, causing a trade-off between similarity and controllability. In this study, we argue an ideal subject representation can be achieved by a cross-differential perspective, i.e., decoupling subject intrinsic attributes from irrelevant attributes via contrastive learning, which allows the model to focus more on intrinsic attributes through intra-consistency (features of the same subject are spatially closer) and inter-distinctiveness (features of different subjects have distinguished differences). Specifically, we propose CustomContrast, a novel framework, which includes a Multilevel Contrastive Learning (MCL) paradigm and a Multimodal Feature Injection (MFI) Encoder. The MCL paradigm is used to extract intrinsic features of subjects from high-level semantics to low-level appearance through crossmodal semantic contrastive learning and multiscale appearance contrastive learning. To facilitate contrastive learning, we introduce the MFI encoder to capture cross-modal representations. Extensive experiments show the effectiveness of CustomContrast in subject similarity and text controllability. Mengqi Huang, Zhuowei Chen, Lei Zhang 0119, Zhendong Mao 0001 |
AAAI | 5 |
| 2025 | FeedEdit: Text-Based Image Editing with Dynamic Feedback RegulationabstractText-based image editing which aims at generating rigid or non-rigid changes to images conditioned on the given text, has recently attracted considerable interest. Previous works mainly follow the multi-step denoising diffusion paradigm, which adopts a fixed text guidance intensity (i.e., editing intensity) to inject textual features, while ignoring the step-specific editing requirements. This work argues that the editing intensity at each denoising step should be adaptively adjusted conditioned on the historical editing degree, to provide accurate text guidance for the whole denoising process. We thereby propose a novel feedback editing framework (FeedEdit), a training-free method which, explicitly exploits the feedback regulation on editing intensity to ensure precise and harmonious editing at all steps. Specifically, we design (1) Dynamic Editing Degree Perceiving module, which is based on specific frequency-domain filtering, to enhance and exploit the correlation between feature differences and editing degree for perceiving. (2) Proportional-Integral feedback controller, to automatically map the perceived editing errors into appropriate feedback control signals. (3) Phrase-level Regulating Strategy, to achieve fine-grained function-specific regulation of textual features. Extensive experiments demonstrate the superiority of FeedEdit over existing methods in both editability and quality, especially for multi-function editing scenarios. Fengyi Fu, Lei Zhang 0119, Mengqi Huang, Zhendong Mao 0001 |
CVPR | 2 |
| 2025 | D^2iT: Dynamic Diffusion Transformer for Accurate Image GenerationabstractDiffusion models are widely recognized for their ability to generate high-fidelity images. Despite the excellent performance and scalability of the Diffusion Transformer (DiT) architecture, it applies fixed compression across different image regions during the diffusion process, disregarding the naturally varying information densities present in these regions. However, large compression leads to limited local realism, while small compression increases computational complexity and compromises global consistency, ultimately impacting the quality of generated images. To address these limitations, we propose dynamically compressing different image regions by recognizing the importance of different regions, and introduce a novel two-stage framework designed to enhance the effectiveness and efficiency of image generation: (1) Dynamic VAE (DVAE) at first stage employs a hierarchical encoder to encode different image regions at different downsampling rates, tailored to their specific information densities, thereby providing more accurate and natural latent codes for the diffusion process. (2) Dynamic Diffusion Transformer (D2iT) at second stage generates images by predicting multi-grained noise, consisting of coarse-grained (less latent code in smooth regions) and fine-grained (more latent codes in detailed regions), through an novel combination of the Dynamic Grain Transformer and the Dynamic Content Transformer. The strategy of combining rough prediction of noise with detailed regions correction achieves a unification of global consistency and local realism. Comprehensive experiments on various generation tasks validate the effectiveness of our approach. Code will be released at https://github.com/jiawn-creator/Dynamic-DiT. Weinan Jia, Mengqi Huang, Lei Zhang 0119, Zhendong Mao 0001 |
CVPR | 4 |
| 2025 | Hierarchy-Aware Pseudo Word Learning with Text Adaptation for Zero-Shot Composed Image Retrieval
Zhe Li 0028, Lei Zhang 0119, Zheren Fu, Kun Zhang 0040, Zhendong Mao 0001 |
ICCV | 2 |
| 2025 | On-the-fly Preference Alignment via Principle-Guided DecodingabstractWith the rapidly expanding landscape of large language models, aligning model generations with human values and preferences is becoming increasingly important. Popular alignment methods, such as Reinforcement Learning from Human Feedback, have shown significant success in guiding models with greater control. However, these methods require considerable computational resources, which is inefficient, and substantial collection of training data to accommodate the diverse and pluralistic nature of human preferences, which is impractical. These limitations significantly constrain the scope and efficacy of both task-specific and general preference alignment methods. In this work, we introduce On-the-fly Preference Alignment via Principle-Guided Decoding (OPAD) to directly align
model outputs with human preferences during inference, eliminating the need for fine-tuning. Our approach involves first curating a surrogate solution to an otherwise infeasible optimization problem and then designing a principle-guided reward function based on this surrogate. The final decoding policy is derived by maximizing this customized reward, which exploits the discrepancy between the
constrained policy and its unconstrained counterpart. OPAD directly modifies the model’s predictions during inference, ensuring principle adherence without incurring the computational overhead of retraining or fine-tuning. Experiments show that OPAD achieves competitive or superior performance in both general and personalized alignment tasks, demonstrating its efficiency and effectiveness compared to state-of-the-art baselines. Mingye Zhu, Yi Liu 0148, Lei Zhang 0119, Junbo Guo, Zhendong Mao 0001 |
ICLR | 3 |
| 2025 | Multi-round Mutual Emotion-Cause Pair Extraction for Emotion-Attributed Video CaptioningabstractEmotional Video Captioning (EVC) is an emerging task that aims to describe factual content with the intrinsic emotions expressed in videos. Existing EVC methods perceive global emotional cues through visual features at first, and then combine them with the video features to guide the emotional caption generation, which ignores the critical characteristic of the EVC task that emotional cues have intrinsic motivational causes reflected in the video content. Such video causes have a facilitative effect on both emotion perception and emotion-attributed caption generation. To this end, a multi-round mutual emotion-cause pair extraction network (MM-ECPE) is proposed in this paper for the joint extraction of emotional cues and visual causes through iterative mutual refinement. Specifically, in the 1st-round mutual learning, we propose a spatio-temporal disentangled visual adaptive refinement (ST-DVAR) and a multi-level video-guided emotion affine transformation (MV-EAT) to achieve preliminary refinement on video features and emotion lexicon to eliminate the noise caused by emotion-irrelevant visual information and video-irrelevant emotional information. Then, in the 2nd-round mutual learning, we exploit the cross-attention of the preliminary refined features and the original features to obtain the ultimate emotional cues and visual causes, and couple them in pair-wise extraction through contrastive loss. Overall, our approach optimizes complex semantic understanding and emotion perception of videos, leading to a promising performance in emotional captioning. Extensive experiments on three challenging datasets demonstrate the superiority of our approach and each proposed module, e.g., improving the latest records by +97.5% and +76.2% w.r.t. CIDEr and CFS, respectively, on the EVC-MSVD dataset. Cheng Ye 0004, Weidong Chen 0013, Peipei Song, Xinyan Liu 0008, Lei Zhang 0119, Zhendong Mao 0001 |
ACM Multimedia | 5 |
| 2025 | Rethinking Pseudo Word Learning in Zero-Shot Composed Image Retrieval: From an Object-Aware PerspectiveabstractComposed Image Retrieval (CIR) takes a composed query of a reference image and a text describing the user's intention, with the aim to retrieve the target image under both conditions. Conventional CIR approaches heavily rely on massive annotated triplets, which often comes at a considerable cost. Zero-Shot CIR (ZS-CIR) offers a new solution that can perform diverse CIR tasks without training on the triplet datasets. The key to the ZS-CIR task is to make specified changes to specific objects in the reference image based on the text. Previous works utilize a projection module to map the reference image into single or multiple pseudo words. However, they are either only applicable to single-object scenarios, or naively convert entire image features into multiple pseudo words and fail to focus on the desired target objects specified by the text description. In this work, we rethink how to learn pseudo words based on the objects attended by the text and propose a Multi-Object Aware ZS-CIR framework (MOA). Specifically, a multi-object recognizer first recognizes valid objects in the reference image guided by a set of learnable object queries. Then, we devise an object filtering strategy, which utilizes contextual prompts comprised of noun categories to guide the model in precisely screening out the objects that need to be modified. Finally, the pseudo word learning branch adaptively converts the screened objects into multiple pseudo words for accurate ZS-CIR. Although simple, our MOA consistently outperforms previous state-of-the-art methods across diverse benchmarks and even achieves competitive results with many supervised methods. Zhe Li 0028, Lei Zhang 0119, Kun Zhang 0040, Weidong Chen 0013, Yongdong Zhang 0001, Zhendong Mao 0001 |
SIGIR | 2 |
| 2025 | Fully Semantic Gap Recovery for End-to-End Image CaptioningabstractImage captioning (IC) involves the comprehension of images from the visual domain to generate descriptions that are grounded in visual elements within the linguistic domain. Current image captioning methods typically rely on pre-trained unimodal visual backbones or vision-language models to identify visual entities. Subsequently, these methods employ unimodal self-attention fusion to uncover high-level semantic associations. However, we uncover this paradigm suffers from the inherent intra-modal semantic gap from the input features. Unimodal pre-trained visual features lack sufficient linguistic semantic information due to the modality misalignment. Furthermore, contrastive pre-trained vision-language models, such as CLIP, confine to the global cross-modal alignment, leading to local visual features belonging to the same object exhibiting distinct semantics. Given the semantically insufficient visual features, unimodal self-attention fusion struggles to accurately capture semantic associations among visual patches, thereby exacerbating the semantic gap. This gap results in inaccurate visual entities and associations in the generated captions. Therefore, we propose a novel Fully Semantic Gap Recovery (FSGR) method to broaden the robust cross-modal bridge of CLIP into a fine-grained level and consolidate vision-language semantic associations for more precise visual comprehension. Technically, we first propose a local contrastive learning method to aggregate the semantically similar visual patches. Next, we design a semantic quantification module to abstract the language-bridged visual map from the enhanced local visual features. Finally, fine-grained cross-modal interaction consolidates the image patches with their corresponding linguistic semantics, allowing the generation of plausible captions based on the aggregated features. Extensive experiments on comprehensive metrics demonstrate that our model has achieved new state-of-the-art performance on the MSCOCO dataset, while also exhibiting competitive cross-domain capability on the Nocaps dataset. Source code released at https://github.com/gjc0824/FSGR. Jingchun Gao, Lei Zhang 0119, Zhendong Mao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Matryoshka Learning With Metric Transfer for Image-Text MatchingabstractImage-text matching is a significant technology for vision-language tasks, as it bridges the semantic gap between visual and text modalities. Although existing methods have achieved remarkable progress, high-dimensional embeddings or ensemble methods are often used to achieve sufficiently good recall or accuracy, which significantly increase the computational and storage costs in practical applications. Knowledge distillation can help achieve resource-efficient deployment, however, existing techniques are not directly applicable to cross-modal matching scenarios. The main difficulties arise from two aspects: (1) the distillation from teacher model to student model is usually conducted in two separate stages, and this inconsistency in learning objectives may lead to sub-optimal compression results. (2) distilling knowledge from each modality independently cannot ensure the preservation of cross-modal alignment established in the original embeddings, which can lead to the compressed ones failing to achieve accurate alignment. To address these issues, we propose a novel Matryoshka Learning with Metric Transfer framework (MAMET) for image-text matching. After capturing multi-granularity information through multiple high-dimensional embeddings, we propose an efficient Matryoshka training process with shared backbone to compress the different granularity information into a low-dimensional embedding, facilitating the integration of cross-modal matching and knowledge distillation in one single stage. Meanwhile, a novel metric transfer criterion is innovated to diversely align the metric relations across embedding spaces of different dimensions and modalities, ensuring a good cross-modal alignment after distillation. In this way, our MAMET transfers strong representation and generalization capability from the high-dimensional ensemble models to a basic network, which not only can get great performance boost, but also introduce no extra overhead during online inference. Extensive experiments on benchmark datasets demonstrate the superior effectiveness and efficiency of our MAMET, consistently achieving an average of 2%-20% performance improvement over state-of-the-art methods across various backbones and domains. Pengzhe Wang, Lei Zhang 0119, Zhendong Mao 0001, Nenan Lyu, Yongdong Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Boosting Faithful Multi-Modal LLMs via Complementary Visual GroundingabstractMultimodal Large Language Models (MLLMs) exhibit impressive performance across vision-language tasks, but still face the hallucination challenges, where generated texts are factually inconsistent with visual input. Existing mitigation methods focus on surface symptoms of hallucination and heavily rely on post-hoc corrections, extensive data curation, or costly inference schemes. In this work, we identify two key factors of MLLM hallucination: Insufficient Visual Context, where ambiguous visual contexts lead to language speculation, and Progressive Textual Drift, where model attention strays from visual inputs in longer responses. To address these problems, we propose a novel Complementary Visual Grounding (CVG) framework. CVG exploits the intrinsic architecture of MLLMs, without requiring any external tools, models, or additional data. CVG first disentangles visual context into two complementary branches based on query relevance, then maintains steadfast visual grounding during the auto-regressive generation. Finally, it contrasts the output distributions of two branches to produce a faithful response. Extensive experiments on various hallucination and general benchmarks demonstrate that CVG achieves state-of-the-art performances across MLLM architectures and scales. Zheren Fu, Zhendong Mao 0001, Lei Zhang 0119, Yongdong Zhang 0001 |
IEEE Trans. Image Process. | 3 |
| 2025 | Improving Video Summarization by Exploring the Coherence Between Corresponding CaptionsabstractVideo summarization aims to generate a compact summary of the original video by selecting and combining the most representative parts. Most existing approaches only focus on recognizing key video segments to generate the summary, which lacks holistic considerations. The transitions between selected video segments are usually abrupt and inconsistent, making the summary confusing. Indeed, the coherence of video summaries is crucial to improve the quality and user viewing experience. However, the coherence between video segments is hard to measure and optimize from a pure vision perspective. To this end, we propose a Language-guided Segment Coherence-Aware Network (LS-CAN), which integrates entire coherence considerations into the key segment recognition. The main idea of LS-CAN is to explore the coherence of corresponding text modality to facilitate the entire coherence of the video summary, which leverages the natural property in the language that contextual coherence is easy to measure. In terms of text coherence measures, specifically, we propose the multi-graph correlated neural network module (MGCNN), which constructs a graph for each sentence based on three key components, i.e., subject, attribute, and action words. For each sentence pair, the node features are then discriminatively learned by incorporating neighbors of its own graph and information of its dual graph, reducing the error of synonyms or reference relationships in measuring the correlation between sentences, as well as the error caused by considering each component separately. In doing so, MGCNN utilizes subject agreement, attribute coherence, and action succession to measure text coherence. Besides, with the help of large language models, we augment the original text coherence annotations, improving the ability of MGCNN to judge coherence. Extensive experiments on three challenging datasets demonstrate the superiority of our approach and each proposed module, especially improving the latest records by +3.8%, +14.2% and +12% w.r.t. F1 scores, $\tau $ and $\rho $ metrics on the BLiSS dataset. Cheng Ye 0004, Weidong Chen 0013, Bo Hu 0036, Lei Zhang 0119, Yongdong Zhang 0001, Zhendong Mao 0001 |
IEEE Trans. Image Process. | 4 |
| 2025 | Exploiting Pre-Trained Language Models for Black-Box Attack against Knowledge Graph EmbeddingsabstractDespite the emerging research on adversarial attacks against knowledge graph embedding (KGE) models, most of them focus on white-box attack settings. However, white-box attacks are difficult to apply in practice compared to black-box attacks since they require access to model parameters that are unlikely to be provided. In this article, we propose a novel black-box attack method that only requires access to knowledge graph data, making it more realistic in real-world attack scenarios. Specifically, we utilize pre-trained language models (PLMs) to encode text features of the knowledge graphs, an aspect neglected by previous research. We then employ these encoded text features to identify the most influential triples for constructing corrupted triples for the attack. To improve the transferability of the attack, we further propose to fine-tune the PLM model by enriching triple embeddings with structure information. Extensive experiments conducted on two knowledge graph datasets illustrate the effectiveness of our proposed method. Guangqian Yang, Lei Zhang 0119, Yi Liu 0148, Hongtao Xie 0001, Zhendong Mao 0001 |
ACM Trans. Knowl. Discov. Data | 2 |
| 2024 | Gradual Residuals Alignment: A Dual-Stream Framework for GAN Inversion and Image Attribute EditingabstractGAN-based image attribute editing firstly leverages GAN Inversion to project real images into the latent space of GAN and then manipulates corresponding latent codes. Recent inversion methods mainly utilize additional high-bit features to improve image details preservation, as low-bit codes cannot faithfully reconstruct source images, leading to the loss of details. However, during editing, existing works fail to accurately complement the lost details and suffer from poor editability. The main reason is they inject all the lost details indiscriminately at one time, which inherently induces the position and quantity of details to overfit source images, resulting in inconsistent content and artifacts in edited images. This work argues that details should be gradually injected into both the reconstruction and editing process in a multi-stage coarse-to-fine manner for better detail preservation and high editability. Therefore, a novel dual-stream framework is proposed to accurately complement details at each stage. The Reconstruction Stream is employed to embed coarse-to-fine lost details into residual features and then adaptively add them to the GAN generator. In the Editing Stream, residual features are accurately aligned by our Selective Attention mechanism and then injected into the editing process in a multi-stage manner. Extensive experiments have shown the superiority of our framework in both reconstruction accuracy and editing quality compared with existing methods. Hao Li 0189, Mengqi Huang, Lei Zhang 0119, Bo Hu 0036, Yi Liu 0148, Zhendong Mao 0001 |
AAAI | 3 |
| 2024 | Identification of Necessary Semantic Undertakers in the Causal View for Image-Text MatchingabstractImage-text matching bridges vision and language, which is a fundamental task in multimodal intelligence. Its key challenge lies in how to capture visual-semantic relevance. Fine-grained semantic interactions come from fragment alignments between image regions and text words. However, not all fragments contribute to image-text relevance, and many existing methods are devoted to mining the vital ones to measure the relevance accurately. How well image and text relate depends on the degree of semantic sharing between them. Treating the degree as an effect and fragments as its possible causes, we define those indispensable causes for the generation of the degree as necessary undertakers, i.e., if any of them did not occur, the relevance would be no longer valid. In this paper, we revisit image-text matching in the causal view and uncover inherent causal properties of relevance generation. Then we propose a novel theoretical prototype for estimating the probability-of-necessity of fragments, PN_f, for the degree of semantic sharing by means of causal inference, and further design a Necessary Undertaker Identification Framework (NUIF) for image-text matching, which explicitly formalizes the fragment's contribution to image-text relevance by modeling PN_f in two ways. Extensive experiments show our method achieves state-of-the-art on benchmarks Flickr30K and MSCOCO. Huatian Zhang 0001, Lei Zhang 0119, Kun Zhang 0040, Zhendong Mao 0001 |
AAAI | 2 |
| 2024 | Visual-Linguistic Dependency Encoding for Image-Text RetrievalabstractImage-text retrieval is a fundamental task to bridge the semantic gap between natural language and vision. Recent works primarily focus on aligning textual meanings with visual appearance. However, they often overlook the semantic discrepancy caused by syntactic structure in natural language expressions and relationships among visual entities. This oversight would lead to sub-optimal alignment and degraded retrieval performance, since the underlying semantic dependencies and object interactions remain inadequately encoded in both textual and visual embeddings. In this paper, we propose a novel Visual-Linguistic Dependency Encoding (VL-DE) framework, which explicitly models the dependency information among textual words and interaction patterns between image regions, improving the discriminative power of cross-modal representations for more accurate image-text retrieval. Specifically, VL-DE enhances textual representations by considering syntactic relationships and dependency types, and visual representations by attending to its spatially neighboring regions. Cross-attention mechanism is then introduced to aggregate aligned region-word pairs into image-text similarities. Analysis on Winoground, a dataset specially designed to measure vision-linguistic compositional structure reasoning, shows that VL-DE outperforms existing methods, demonstrating its effectiveness at this task. Comprehensive experiments on two benchmarks, Flickr30K and MS-COCO, further validates the competitiveness of our approach. Wenxin Guo, Lei Zhang 0119, Kun Zhang 0040, Yi Liu 0148, Zhendong Mao 0001 |
LREC/COLING | 2 |
| 2024 | Linguistic-Aware Patch Slimming Framework for Fine-Grained Cross-Modal AlignmentabstractCross-modal alignment aims to build a bridge connecting vision and language. It is an important multi-modal task that efficiently learns the semantic similarities be-tween images and texts. Traditional finegrained alignment methods heavily rely on pre-trained object detectors to extract region features for subsequent region-word alignment, thereby incurring substantial computational costs for region detection and error propagation issues for two-stage training. In this paper, we focus on the mainstream vision transformer, incorporating patch features for patch-word alignment, while addressing the resultant issue of visual patch redundancy and patch ambiguity for semantic alignment. We propose a novel Linguistic-Aware Patch Slim-ming (LAPS) framework for fine-grained alignment, which explicitly identifies redundant visual patches with language supervision and rectifies their semantic and spatial information to facilitate more effective and consistent patchword alignment. Extensive experiments on various evaluation benchmarks and model backbones show LAPS out-performs the state-of-the-art fine-grained alignment methods by 5%-15% rSum. Our code is available at https://github.com/CrossmodalGroup/LAPS. Zheren Fu, Lei Zhang 0119, Hou Xia, Zhendong Mao 0001 |
CVPR | 2 |
| 2024 | Neighborhood-Adaptive Context Enhancement Learning For Scene Graph GenerationabstractConventional scene graph generation methods primarily attempt to implicitly memorize data knowledge within model parameters, yet struggle to excel across all categories, particularly under the influence of long-tailed distributions. Complementing this parameter-based approach, we suggest gathering triplets from similar neighbor instances as extra knowledge. Based on this, we propose a novel Neighborhood-adaptive Context Enhancement Learning (NACEL) network to dynamically select helpful knowledge and integrate it with contextual features for enhanced adaptability. Our method exposes the model to more instances beyond the input, boosting the efficiency and performance of relation prediction. By applying our method on various baselines, extensive experiments on VG dataset have shown that category-level metric mRecall has been significantly improved while instance-level metric Recall has not excessively degraded, which demonstrates our plug-and-play method effectively alleviates biased problem and has the best comprehensive performance. Licheng Zhang 0002, Lei Zhang 0119, Zhendong Mao 0001 |
ICME | 3 |
| 2024 | Dual-path Collaborative Generation Network for Emotional Video CaptioningabstractEmotional Video Captioning (EVC) is an emerging task that aims to describe factual content with the intrinsic emotions expressed in videos. The essential of the EVC task is to effectively perceive subtle and ambiguous visual emotional cues during the caption generation, which is neglected by the traditional video captioning. Existing emotional video captioning methods perceive global visual emotional cues at first, and then combine them with the video features to guide the emotional caption generation, which neglects two characteristics of the EVC task. Firstly, their methods neglect the dynamic subtle changes in the intrinsic emotions of the video, which makes it difficult to meet the needs of common scenes with diverse and changeable emotions. Secondly, as their methods incorporate emotional cues into each step, the guidance role of emotion is overemphasized, which makes factual content more or less ignored during generation. To this end, we propose a dual-path collaborative generation network, which dynamically perceives visual emotional cues evolutions while generating emotional captions by collaborative learning. The two paths promote each other and significantly improve the generation performance. Specifically, in the dynamic emotion perception path, we propose a dynamic emotion evolution module, which first aggregates visual features and historical caption features to summarize the global visual emotional cues, and then dynamically selects emotional cues required to be re-composed at each stage as well as re-composed them to achieve emotion evolution by dynamically enhancing or suppressing different granularity subspace's semantics. Besides, in the adaptive caption generation path, to balance the description of factual content and emotional cues, we propose an emotion adaptive decoder, which firstly estimates emotion intensity via the alignment of emotional features and historical caption features at each generation step, and then, emotional guidance adaptively incorporate into the caption generation based on the emotional intensity. Thus, our methods can generate emotion-related words at the necessary time step, and our caption generation balances the guidance of factual content and emotional cues well. Extensive experiments on three challenging datasets demonstrate the superiority of our approach and each proposed module. Cheng Ye 0004, Weidong Chen 0013, Lei Zhang 0119, Zhendong Mao 0001 |
ACM Multimedia | 4 |
| 2024 | Homology Consistency Constrained Efficient Tuning for Vision-Language ModelsabstractEfficient transfer learning has shown remarkable performance in tuning large-scale vision-language models (VLMs) toward downstream tasks with limited data resources. The key challenge of efficient transfer lies in adjusting image-text alignment to be task-specific while preserving pre-trained general knowledge. However, existing methods adjust image-text alignment merely on a set of observed samples, e.g., data set and external knowledge base, which cannot guarantee to keep the correspondence of general concepts between image and text latent manifolds without being disrupted and thereby a weak generalization of the adjusted alignment. In this work, we propose a Homology Consistency (HC) constraint for efficient transfer on VLMs, which explicitly constrains the correspondence of image and text latent manifolds through structural equivalence based on persistent homology in downstream tuning. Specifically, we build simplicial complex on the top of data to mimic the topology of latent manifolds, then track the persistence of the homology classes of topological features across multiple scales, and guide the directions of persistence tracks in image and text manifolds to coincide each other, with a deviating perturbation additionally. For practical application, we tailor the implementation of our proposed HC constraint for two main paradigms of adapter tuning. Extensive experiments on few-shot learning over 11 datasets and domain generalization demonstrate the effectiveness and robustness of our method. Huatian Zhang 0001, Lei Zhang 0119, Yongdong Zhang 0001, Zhendong Mao 0001 |
NeurIPS | 2 |
| 2024 | Fast, Accurate, and Lightweight Memory-Enhanced Embedding Learning Framework for Image-Text RetrievalabstractImage-text retrieval is a fundamental task in bridging the semantics between vision and language. The key challenge lies in accurately and efficiently learning the semantic alignment between two heterogeneous modalities. Existing image-text retrieval approaches can be roughly classified into two paradigms. The first independent-embedding paradigm is to learn the global embeddings of two modalities, which can achieve efficient retrieval while failing to effectively capture the cross-modal fine-grained interaction information between images and texts. The second interactive-embedding paradigm is to learn fine-grained alignment between regions and words, which can achieve accurate retrieval while sacrificing retrieval efficiency. In this paper, we propose a novel Independent Memory-Enhanced emBedding learning framework (IMEB), which introduces a lightweight middleware,i.e., memory network, into the independent-embedding approaches to simultaneously exploit the complementary of both paradigms. Specifically, first, in the training stage, we propose a novel cross-modal association graph to learn cross-modal fine-grained interaction information. Then, we delicately design a memory-assisted embedding learning network to store these prototypical features after interaction as agents, and effectively update the memory network via two learning strategies. Finally, in the inference stage, we directly interact with these agent-level prototypical features from the memory bank, thus efficiently obtaining cross-modal memory-enhanced embeddings. In this way, our model not only effectively learns cross-modal interaction information, but also maintains the retrieval efficiency. Extensive experimental results on two benchmarks,i.e., Flickr30K and MS-COCO, demonstrate that our IMEB performs favorably against state-of-the-art methods. Zhe Li 0028, Lei Zhang 0119, Kun Zhang 0040, Yongdong Zhang 0001, Zhendong Mao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Cascade Semantic Prompt Alignment Network for Image CaptioningabstractImage captioning (IC) takes an image as input and generates open-form descriptions in the domain of natural language. IC requires the detection of objects, modeling of relations between them, an assessment of the semantics of the scene and representing the extracted knowledge in a language space. Previous detector-based models suffer from limited semantic perception capability due to predefined object detection classes and semantic inconsistency between visual region features and numeric labels of the detector. Inspired by the fact that text prompts in pre-trained multi-modal models contain specific linguistic knowledge rather than discrete labels, and excel at an open-form semantic understanding of visual inputs and their representation in the domain of natural language. We aim to distill and leverage the transferable language knowledge from the pre-trained RegionCLIP model to remedy the detector for generating rich image captioning. In this paper, we propose a novel Cascade Semantic Prompt Alignment Network (CSA-Net) to produce an aligned fine-grained regional semantic-visual space where rich and consistent textual semantic details are automatically incorporated to region features. Specifically, we first align the object semantic prompt and region features to produce semantic grounded object features. Then, we employ these object features and relation semantic prompt to predict the relations between objects. Finally, these enhanced object and relation features are fed into the language decoder, generating rich descriptions. Extensive experiments conducted on the MSCOCO dataset show that our method achieves a new state-of-the-art performance with 145.2% (single model) and 147.0% (ensemble of 4 models) CIDEr scores on the ‘Karpathy’ split, 141.6% (c5) and 144.1% (c40) CIDEr scores on the official online test server. Significantly, CSA-Net outperforms in generating captions with higher quality and diversity, achieving a RefCLIP-S score of 83.2. Moreover, we expand the testbeds to other challenging captioning benchmarks, i.e., nocaps datasets, CSA-Net demonstrates superior zero-shot capability. Source codes released at https://github.com/CrossmodalGroup/CSA-Net. Lei Zhang 0119, Kun Zhang 0040, Bo Hu 0036, Hongtao Xie 0001, Zhendong Mao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Improving Image-Text Matching With Bidirectional Consistency of Cross-Modal AlignmentabstractImage-text matching is a fundamental task in bridging the semantics between vision and language. The key challenge lies in establishing accurate alignment between two heterogeneous modalities. Existing cross-modal fine-grained matching methods normally include two alignment directions, “word to region” and “region to word”, and the overall image-text similarity is calculated from the alignments. However, the alignment of these two directions is typically independent, that is, the alignment of “word to region” and “region to word” is irrelevant, so the alignment consistency cannot be guaranteed in two directions, which inevitably introduces inconsistent alignments, leading to potential inaccurate image-text matching results. In this paper, we propose a novel Bidirectional cOnsistency netwOrks for cross-Modal alignment (BOOM), which achieves more accurate cross-modal semantic alignments by imposing explicit consistency constraints in both directions. Specifically, according to three aspects reflected by alignment consistency,i.e., significance, wholeness, and alignment orderliness, we design a novel systematic multi-granularity consistency constraints: point-wise consistency, which enforces consistency of the most significant single word item in bidirectional alignments; set-wise consistency, which maintains more comprehensive and accurate bidirectional entire alignment values consistent and order-wise consistency, which ensures order consistency of bidirectional alignment results. Bidirectional cross-modal alignment between words and regions is corrected from three different perspectives: maximum, distribution, and order. Extensive experiments on two benchmarks,i.e., Flickr30K and MS-COCO, demonstrate that our BOOM achieves state-of-the-art performance. Zhe Li 0028, Lei Zhang 0119, Kun Zhang 0040, Yongdong Zhang 0001, Zhendong Mao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Semantic-Enhanced Proxy-Guided Hashing for Long-Tailed Image RetrievalabstractHashing has been studied extensively for large-scale image retrieval due to its efficient computation and storage. Deep hashing methods typically train models with category-balanced data and suffer from a serious performance deterioration when dealing with long-tailed training samples. Recently, several long-tailed hashing methods focus on this newly emerging field for practical purpose. However, existing methods still face challenges that fixed category centers with limited semantic information cannot effectively improve the discriminative ability of tail-category hash codes. To tackle the issue, we propose a novel method called Semantic-enhanced Proxy-guided Hashing in this paper. We leverage two sets of learnable category proxies in the feature space and the Hamming space respectively, which can describe category semantics by getting updated continuously along with the whole model via back-propagation. Based on this, we introduce the Mahalanobis distance metric to characterize relationships accurately and enhance the semantic representation of both proxies and samples concurrently, improving the hash learning process. Moreover, we capture the multilateral correlations between proxies and samples in the feature space and extend a hypergraph neural network to transfer semantic knowledge from proxies to samples in the Hamming space. Extensive experiments show that our method achieves the state-of-the-art performance and surpasses existing methods by 1.47%–7.56% MAP on long-tailed benchmarks, demonstrating the superiority of learnable category proxies and the effectiveness of our proposed learning algorithm for long-tailed hashing. Hongtao Xie 0001, Lei Zhang 0119, Pandeng Li, Dongming Zhang 0004, Yongdong Zhang 0001 |
IEEE Trans. Multim. | 3 |
| 2023 | E-CORE: Emotion Correlation Enhanced Empathetic Dialogue GenerationabstractAchieving empathy is a crucial step toward humanized dialogue systems.Current approaches for empathetic dialogue generation mainly perceive an emotional label to generate an empathetic response conditioned on it, which simply treat emotions independently, but ignore the intrinsic emotion correlation in dialogues, resulting in inaccurate emotion perception and unsuitable response generation.In this paper, we propose a novel emotion correlation enhanced empathetic dialogue generation framework, which comprehensively realizes emotion correlation learning, utilization, and supervising.Specifically, a multi-resolution emotion graph is devised to capture context-based emotion interactions from different resolutions, further modeling emotion correlation.Then we propose an emotion correlation enhanced decoder, with a novel correlation-aware aggregation and soft/hard strategy, respectively improving the emotion perception and response generation.Experimental results on the benchmark dataset demonstrate the superiority of our model in both empathetic perception and expression. Fengyi Fu, Lei Zhang 0119, Quan Wang 0002, Zhendong Mao 0001 |
EMNLP | 2 |
| 2023 | Dual Dynamic Proxy Hashing Network for Long-tailed Image RetrievalabstractDeep hashing has been extensively explored for image retrieval due to fast computation and efficient storage. Since conventional deep hashing methods are not suitable for the common scenario in real life that data exhibits a long-tailed distribution, several long-tailed hashing methods have been proposed recently. However, existing long-tail hashing methods seek to utilize fixed class centroids and cannot fully develop the discriminative ability of hash codes for tail-class samples. Specifically, fixed class centroids cannot characterize authentic semantics of tail classes or provide effective semantic information for hash codes learning under the long-tailed setting. To this end, we propose a novel Dual Dynamic Proxy Hashing Network (DDPHN) with two sets of learnable dynamic proxies, i.e. hash proxies and feature proxies, to improve the discrimination of hash codes for tail-class samples. Compared with fixed class centroids, learnable proxies can be optimized constantly via the proxy learning loss and depict accurate class semantics despite the scarcity of tail-class samples. Apart from low-dimensional binary hash proxies, we introduce high-dimensional continuous feature proxies that can describe semantic relationships more precisely, contributing to hash codes learning as well. To further leverage semantic information carried by proxies, we build a hypergraph by exploring neighborhood relationships in the feature space and then introduce a hypergraph neural network to transfer knowledge from proxies to samples in the Hamming space. Extensive experiments show the superiority of our learnable dynamic proxies and demonstrate that our method outperforms numerous deep hashing models and recent state-of-the-art long-tailed hashing methods. Hongtao Xie 0001, Lei Zhang 0119, Pandeng Li, Dongming Zhang 0004, Yongdong Zhang 0001 |
ACM Multimedia | 3 |
| 2023 | Reducing Intrinsic and Extrinsic Data Biases for Moment Localization with Natural LanguageabstractMoment Localization with Natural Language (MLNL) aims to locate the target moment from an untrimmed video by a linguistic query. Recent works reveal the severe data bias problem in MLNL and point out that the multi-modal content may not be understood by fitting the timestamp distribution. In this paper, we study the data biases on the intrinsic and extrinsic aspects: the former is mainly caused by the ambiguity of the moment boundary and the information imbalance between input and output; The latter results from the long-tail distribution of moments in MLNL datasets. To alleviate this, we propose a hybrid multi-modal debiasing network with temporal consistency constraint for MLNL. Specifically, we first design the multi-temporal Transformer to mitigate the ambiguity of boundary by integrating frame-wise features into segment-wise and dynamically matching with moment boundaries. Then, we introduce the temporal consistency constraint that highlights the action information in complex moment content to overcome the intrinsic bias from information imbalance.Furthermore, we design the hybrid linguistic activating module with external knowledge to relieve the extrinsic bias, which introduces a prior guidance to focus the discriminative information from the tail samples. Extensive experiments on three public datasets demonstrate that our model outperforms the existing methods. Jiong Yin, Liang Li 0003, Chenggang Yan 0001, Lei Zhang 0119, Zunjie Zhu |
ACM Multimedia | 5 |
| 2023 | Frequency-based Zero-Shot Learning with Phase AugmentationabstractZero-Shot Learning (ZSL) aims to recognize images from seen and unseen classes by aligning visual and semantic knowledge (e.g., attribute descriptions). However, the fine-grained attributes in the RGB domain can be easily affected by background noise (e.g., the grey bird tail blending with the ground), making it difficult to effectively distinguish them. Analyzing the features in the frequency domain assists in better distinguishing the attributes since their patterns remain consistent across different images, unlike noise which may be more variable. Nevertheless, existing ZSL methods typically learn visual features directly from the RGB domain, which can impede the recognition of certain attributes. To overcome this limitation, we propose a novel ZSL method named Frequency-based Phase Augmentation (FPA) network, which learns an effective representation of the attributes in the frequency domain. Specifically, we introduce a Hybrid Phase Augmentation (HPA) module to transform visual features into the frequency domain and augment the phase component for better retention of semantic information of the attributes. The use of phase-augmented features enables FPA to capture more semantic knowledge that can be challenging to distinguish in the RGB domain, suppress noise, and highlight significant attributes. Our extensive experiments show that FPA achieves state-of-the-art performance across four standard datasets. Wanting Yin, Hongtao Xie 0001, Lei Zhang 0119, Jiannan Ge, Pandeng Li, Chuanbin Liu 0001, Yongdong Zhang 0001 |
ACM Multimedia | 3 |
| 2023 | Unlocking the Power of Cross-Dimensional Semantic Dependency for Image-Text MatchingabstractImage-text matching, as a fundamental cross-modal task, bridges vision and language. The key challenge lies in accurately learning the semantic similarity of these two heterogeneous modalities. To determine the semantic similarity between visual and textual features, existing paradigm typically first maps them into a d-dimensional shared representation space, then independently aggregates all dimensional correspondences of cross-modal features to reflect it, e.g., the inner product. However, in this paper, we are motivated by an insightful finding that dimensions are not mutually independent, but there are intrinsic dependencies among dimensions to jointly represent latent semantics. Ignoring this intrinsic information probably leads to suboptimal aggregation for semantic similarity, impairing cross-modal matching learning. To solve this issue, we propose a novel cross-dimensional semantic dependency-aware model (called X-Dim), which explicitly and adaptively mines the semantic dependencies between dimensions in the shared space, enabling dimensions with joint dependencies to be enhanced and utilized. X-Dim (1) designs a generalized framework to learn dimensions' semantic dependency degrees, and (2) devises the adaptive sparse probabilistic learning to autonomously make the model capture precise dependencies. Theoretical analysis and extensive experiments demonstrate the superiority of X-Dim over state-of-the-art methods, achieving 5.9%-7.3% rSum improvements on Flickr30K and MS-COCO benchmarks. Kun Zhang 0040, Lei Zhang 0119, Bo Hu 0036, Mengxiao Zhu 0001, Zhendong Mao 0001 |
ACM Multimedia | 2 |
| 2023 | MomentDiff: Generative Video Moment Retrieval from Random to RealabstractVideo moment retrieval pursues an efficient and generalized solution to identify the specific temporal segments within an untrimmed video that correspond to a given language description.
To achieve this goal, we provide a generative diffusion-based framework called MomentDiff, which simulates a typical human retrieval process from random browsing to gradual localization.
Specifically, we first diffuse the real span to random noise, and learn to denoise the random noise to the original span with the guidance of similarity between text and video.
This allows the model to learn a mapping from arbitrary random locations to real moments, enabling the ability to locate segments from random initialization.
Once trained, MomentDiff could sample random temporal segments as initial guesses and iteratively refine them to generate an accurate temporal boundary.
Different from discriminative works (e.g., based on learnable proposals or queries), MomentDiff with random initialized spans could resist the temporal location biases from datasets.
To evaluate the influence of the temporal location biases, we propose two ``anti-bias'' datasets with location distribution shifts, named Charades-STA-Len and Charades-STA-Mom.
The experimental results demonstrate that our efficient framework consistently outperforms state-of-the-art methods on three public benchmarks, and exhibits better generalization and robustness on the proposed anti-bias datasets.
The code, model, and anti-bias evaluation datasets will be released publicly. Pandeng Li, Chen-Wei Xie, Hongtao Xie 0001, Lei Zhang 0119, Deli Zhao, Yongdong Zhang 0001 |
NeurIPS | 5 |
| 2022 | Neighborhood-Adaptive Structure Augmented Metric LearningabstractMost metric learning techniques typically focus on sample embedding learning, while implicitly assume a homogeneous local neighborhood around each sample, based on the metrics used in training ( e.g., hypersphere for Euclidean distance or unit hyperspherical crown for cosine distance). As real-world data often lies on a low-dimensional manifold curved in a high-dimensional space, it is unlikely that everywhere of the manifold shares the same local structures in the input space. Besides, considering the non-linearity of neural networks, the local structure in the output embedding space may not be homogeneous as assumed. Therefore, representing each sample simply with its embedding while ignoring its individual neighborhood structure would have limitations in Embedding-Based Retrieval (EBR). By exploiting the heterogeneity of local structures in the embedding space, we propose a Neighborhood-Adaptive Structure Augmented metric learning framework (NASA), where the neighborhood structure is realized as a structure embedding, and learned along with the sample embedding in a self-supervised manner. In this way, without any modifications, most indexing techniques can be used to support large-scale EBR with NASA embeddings. Experiments on six standard benchmarks with two kinds of embeddings, i.e., binary embeddings and real-valued embeddings, show that our method significantly improves and outperforms the state-of-the-art methods. Pandeng Li, Yan Li 0068, Hongtao Xie 0001, Lei Zhang 0119 |
AAAI | 4 |
| 2022 | Dual-Stream Knowledge-Preserving Hashing for Unsupervised Video Retrieval
Pandeng Li, Hongtao Xie 0001, Jiannan Ge, Lei Zhang 0119, Shaobo Min, Yongdong Zhang 0001 |
ECCV (14) | 4 |
| 2022 | Learning Common Dependency Structure for Unsupervised Cross-Domain NerabstractUnsupervised cross-domain NER task aims to solve the issues when data in a new domain are fully-unlabeled. It leverages labeled data from source domain to predict entities in unlabeled target domain. Since training models on large domain corpus is time-consuming, in this paper, we consider an alternative way by introducing syntactic dependency structure. Such information is more accessible and can be shared between sentences from different domains. We propose a novel framework with dependency-aware GNN (DGNN) to learn these common structures from source domain and adapt them to target domain, alleviating the data scarcity issue and bridging the domain gap. Experimental results show that our method outperforms state-of-the-art methods. Luchen Liu, Xixun Lin, Peng Zhang 0001, Lei Zhang 0119, Bin Wang 0004 |
ICASSP | 4 |
| 2015 | What is the next step of binary features?abstractVarious binary features have been recently proposed in literature, aiming at improving the computational efficiency and storage efficiency of image retrieval applications. However, the most common way of using binary features is voting strategy based on brute-force matching, since binary features are discrete data points distributed in Hamming space, so that models based on clustering such as BoW are unsuitable for them. Although indexing mechanism substantially decreases the time cost, the brute-force matching strategy becomes a bottleneck that restricts the performance of binary features. To address this issue, we propose a simple but effective method, namely COIP (Coding by Order-independent Projection), which projects binary features into a binary code of limited bits. As a result, each image is represented by one single binary code that can be indexed for computational and storage efficiency. We prove that the similarity between the COIP codes of two images with probability proportional to the ratio of their matched features. A comprehensive evaluation with several state-of-the-art binary features is performed on benchmark dataset. Experimental results reveal that for binary feature based image retrieval, our approach improves the storage/time efficiency by one/two orders of magnitude, while the retrieval performance remains almost unchanged. Zhendong Mao 0001, Lei Zhang 0119, Bin Wang 0004, Li Guo 0001 |
ICME | 2 |
| 2015 | Full-Space Local Topology Extraction for Cross-Modal RetrievalabstractWith the ever increasing availability of various kinds of multimedia data, cross-modal retrieval, which enables information retrieval from various types of data given various types of query, has become a research hotspot. Hashing-based techniques have been developed to solve this problem, however, most previous works cannot capture the shared underlying structure of real-world multimodal data, which degrades their retrieval performances. In this paper, we propose a novel hashing method based on the extraction of the common manifold structure shared among different feature spaces. To faithfully represent the common structure, two kinds of local topology information are exploited in our method. Local angles are incorporated within the extraction of local topology of each feature space, which is then used to learn a common intermediate subspace. After heterogeneous features being embedded into this subspace, local similarities are exploited to extract the local topology between different feature spaces, and learn compact Hamming embeddings to facilitate cross-modal retrieval. The proposed method is referred to as full-space local topology extraction for hashing. Extensive comparisons with other state-of-the-art methods on three benchmark multimedia data sets demonstrate the superiority of our proposed method in terms of retrieval recall and search accuracy. Lei Zhang 0119, Yongdong Zhang 0001, Richang Hong, Qi Tian 0001 |
IEEE Trans. Image Process. | 1 |
| 2014 | Scalable Similarity Search With Topology Preserving HashingabstractHashing-based similarity search techniques is becoming increasingly popular in large data sets. To capture meaningful neighbors, the topology of a data set, which represents the neighborhood relationships between its subregions and the relative proximities between the neighbors of each subregion, e.g., the relative neighborhood ranking of each subregion, should be exploited. However, most existing hashing methods are developed to preserve neighborhood relationships while ignoring the relative neighborhood proximities. Moreover, most hashing methods lack in providing a good result ranking, since there are often lots of results sharing the same Hamming distance to a query. In this paper, we propose a novel hashing method to solve these two issues jointly. The proposed method is referred to as topology preserving hashing (TPH). TPH is distinct from prior works by also preserving the neighborhood ranking. Based on this framework, we present three different TPH methods, including linear unsupervised TPH, semisupervised TPH, and kernelized TPH. Particularly, our unsupervised TPH is capable of mining semantic relationship between unlabeled data without supervised information. Extensive experiments on four large data sets demonstrate the superior performances of the proposed methods over several state-of-the-art unsupervised and semisupervised hashing techniques. Lei Zhang 0119, Yongdong Zhang 0001, Xiaoguang Gu, Jinhui Tang 0001, Qi Tian 0001 |
IEEE Trans. Image Process. | 1 |
| 2014 | A Prior-Free Weighting Scheme for Binary Code RankingabstractFast similarity search has been a research focus in recent years. Binary hashing, which embeds high-dimensional data points into Hamming space, is a promising way to accelerate similarity search, since its search process can be performed in real-time by using Hamming distance as similarity metric. However, as Hamming distance is discrete and bounded by code length, its resolution is limited. In practice, there are often many results sharing the same Hamming distance to a query, which poses a critical issue for problems where ranking is important. This paper proposes a weighted Hamming distance ranking algorithm (WhRank) to give a better ranking of results with equal Hamming distances to a query. By assigning different bit-level weights to different bits, WhRank is able to distinguish between the relative importance of different bits, and to rank the results at a finer-grained hash code level rather than the original integer Hamming distance level. We show that an effective weight is not only data-adaptive but also query-sensitive, and give a simple yet effective prior-free weight learning algorithm. Evaluations on three large-scale image datasets containing up to one million points demonstrate the efficacy of the proposed algorithm. Yongdong Zhang 0001, Lei Zhang 0119, Qi Tian 0001 |
IEEE Trans. Multim. | 2 |
| 2013 | Binary Code Ranking with Weighted Hamming DistanceabstractBinary hashing has been widely used for efficient similarity search due to its query and storage efficiency. In most existing binary hashing methods, the high-dimensional data are embedded into Hamming space and the distance or similarity of two points are approximated by the Hamming distance between their binary codes. The Hamming distance calculation is efficient, however, in practice, there are often lots of results sharing the same Hamming distance to a query, which makes this distance measure ambiguous and poses a critical issue for similarity search where ranking is important. In this paper, we propose a weighted Hamming distance ranking algorithm (WhRank) to rank the binary codes of hashing methods. By assigning different bit-level weights to different hash bits, the returned binary codes are ranked at a finer-grained binary code level. We give an algorithm to learn the data-adaptive and query-sensitive weight for each hash bit. Evaluations on two large-scale image data sets demonstrate the efficacy of our weighted Hamming distance for binary code ranking. Lei Zhang 0119, Yongdong Zhang 0001, Jinhui Tang 0001, Ke Lu 0002, Qi Tian 0001 |
CVPR | 1 |
| 2013 | A video copy detection algorithm combining local feature's robustness and global feature's speedabstractThis paper presents a novel algorithm for fast and robust video copy detection. The idea is to use local features to estimate the copy transformation parameters first and then use the estimated parameters to guide the global-feature-based matching at a later stage. It is based on the fact that the copy transformations generally remain unchanged in a continuous video clip even in the whole video. Local-feature-based matching can find the candidates which are difficult to be detected only using global features. Furthermore, the matched local feature points can provide enough information to estimate the copy transformations. After the copy transformations are estimated, the subsequent detection can be accelerated by doing global-feature-based matching. The experimental results show that the proposed algorithm can get the same good robustness as the local-feature-based method but the faster detection speed. Xiaoguang Gu, Dongming Zhang 0004, Yongdong Zhang 0001, Jintao Li 0001, Lei Zhang 0119 |
ICASSP | 5 |
| 2013 | Topology preserving hashing for similarity searchabstractBinary hashing has been widely used for efficient similarity search. Learning efficient codes has become a research focus and it is still a challenge. In many cases, the real-world data often lies on a low-dimensional manifold, which should be taken into account to capture meaningful neighbors with hashing. The importance of a manifold is its topology, which represents the neighborhood relationships between its subregions and the relative proximities between the neighbors of each subregion, e.g. the relative ranking of neighbors of each subregion. Most existing hashing methods try to preserve the neighborhood relationships by mapping similar points to close codes, while ignoring the neighborhood rankings. Moreover, most hashing methods lack in providing a good ranking for query results since they use Hamming distance as the similarity metric, and in practice, there are often a lot of results sharing the same distance to a query. In this paper, we propose a novel hashing method to solve these two issues jointly. The proposed method is referred to as Topology Preserving Hashing (TPH). TPH is distinct from prior works by preserving the neighborhood rankings of data points in Hamming space. The learning stage of TPH is formulated as a generalized eigendecomposition problem with closed form solutions. Experimental comparisons with other state-of-the-art methods on three noted image benchmarks demonstrate the efficacy of the proposed method. Lei Zhang 0119, Yongdong Zhang 0001, Jinhui Tang 0001, Xiaoguang Gu, Jintao Li 0001, Qi Tian 0001 |
ACM Multimedia | 1 |
| 2013 | Distribution-Aware Locality Sensitive Hashing
Lei Zhang 0119, Yongdong Zhang 0001, Dongming Zhang 0004, Qi Tian 0001 |
MMM (2) | 1 |
| 2013 | An improved method of locality sensitive hashing for indexing large-scale and high-dimensional features
Xiaoguang Gu, Yongdong Zhang 0001, Lei Zhang 0119, Dongming Zhang 0004, Jintao Li 0001 |
Signal Process. | 3 |
| 2012 | Query Range Sensitive Probability Guided Multi-probe Locality Sensitive HashingabstractLocality Sensitive Hashing (LSH) is proposed to construct indexes for high-dimensional approximate similarity search. Multi-Probe LSH (MPLSH) is a variation of LSH which can reduce the number of hash tables. Based on the idea of MPLSH, this paper proposes a novel probability model and a query-adaptive algorithm to generate the optimal multi-probe sequence for range queries. Our probability model takes the query range into account to generate the probe sequence which is optimal for range queries. Furthermore, our algorithm does not use a fixed number of probe steps but a query-adaptive threshold to control the search quality. We do the experiments on an open dataset to evaluate our method. The experimental results show that our method can probe fewer points than MPLSH for getting the same recall. As a result, our method can get an average acceleration of 10% compared to MPLSH. Xiaoguang Gu, Lei Zhang 0119, Dongming Zhang 0004, Yongdong Zhang 0001, Jintao Li 0001, Ning Bao |
SNPD | 2 |
| 2011 | A pivot-based filtering algorithm for enhancing query performance of LSHabstractIn recent years, Locality Sensitive Hashing (LSH) (and its variant Euclidean LSH) has become a popular index structure for large-scale and high-dimensional similarity search problem. In this paper, we analyze a phenomenon we called "Non-Uniform" that degrades the query performance of LSH and propose a pivot-based algorithm to improve the query performance. We also provide a method to get optimal pivot for even larger improvement. Experiments show that our algorithm significantly improves the query performance of LSH. Lei Zhang 0119, Xiaoguang Gu, Yongdong Zhang 0001, Dongming Zhang 0004, Jintao Li 0001 |
VCIP | 1 |