EDBT 2026 Demo / reviewers in the wild / expert
Donglin Zhang 0001
dblp:16/493-1
· DBLP profile ↗
37ranked-venue papers
23as first author
36since 2021 · last 2026
0000-0002-0716-0918ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 18 · 10 first-author · 17 since 2021Artificial intelligence and machine learning · 14 · 9 first-author · 13 since 2021Databases, data management, data science and information retrieval · 6 · 4 first-author · 6 since 2021Computer networks · 3 · 3 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SketchMind: Understanding Abstract Sketches with MLLMs for Fine-Grained Sketch-Based Image Retrieval
Changxing Li, Donglin Zhang 0001, Zhikai Hu, Xiaojun Wu 0001, Josef Kittler |
WWW | 2 |
| 2026 | HACG: Leveraging Hierarchical Alignment and Caption Generation for Text-Video Retrieval
Donglin Zhang 0001, Xiaojun Wu 0001, Josef Kittler |
Int. J. Comput. Vis. | 1 |
| 2026 | Task-specialized multi-modal fusion network for food nutrition estimation
Shuying Hong, Donglin Zhang 0001 |
Multim. Syst. | 2 |
| 2026 | CTNet: learning local details and global context for food recognition
Boyuan Ma, Donglin Zhang 0001 |
Multim. Syst. | 2 |
| 2026 | RSPH: Robust self-paced hashing for cross-modal retrieval
Donglin Zhang 0001, Zhikai Hu, Xiaojun Wu 0001, Josef Kittler |
Pattern Recognit. | 1 |
| 2026 | Supervised discrete cross-modal hashing with exploiting semantic correlations
Donglin Zhang 0001, Zhikai Hu, Xiaojun Wu 0001, Josef Kittler |
Pattern Recognit. | 1 |
| 2026 | HTVR: Hierarchical text-to-video retrieval based on relative similarity
Donglin Zhang 0001, Zhikai Hu |
Pattern Recognit. | 1 |
| 2026 | VIP-CAN: Vision prompt-based cross-attention network for zero-shot sketch-based image retrieval
Donglin Zhang 0001 |
Pattern Recognit. Lett. | 2 |
| 2026 | Cross-Modality Neighborhood Attention Network for Fine-Grained Sketch-Based Image RetrievalabstractFine-Grained Sketch-Based Image Retrieval (FG-SBIR) is a challenging task that focuses on retrieving relevant images based on sketch queries at the instance level. Most existing works tend to make cross-modality interactions between features at the global level and treat each feature equally when calculating similarity. However, this equal treatment may reduce the contribution of fine-grained information. To address these challenges, we propose a novel Cross-Modality Neighborhood Attention Network, named CMNA-Net. Specifically, we develop a novel CMNA mechanism in the proposed method to adjust the importance of modality-agnostic features, which not only performs neighborhood attention operation on the current modality but also performs cross-neighborhood attention on both modalities. Additionally, the edge map is innovatively introduced to reduce the modality gap. Furthermore, this work proposes a Neighborhood Weighted module that integrates the local feature weight scheme and top-k feature selection strategy, yielding more discriminative features tailored to fine-grained retrieval. Extensive experiments on three benchmark datasets demonstrate that our proposed CMNA-Net can achieve better performance compared to some existing state-of-the-art methods in FG-SBIR scenarios. The code of our method will be available at https://github.com/li1changxing/CMNA-Net. Donglin Zhang 0001, Changxing Li, Xiaojun Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2026 | Probabilistic Embeddings With Evidence Learning and Refinement for Text-Video RetrievalabstractThis paper studies the problem of text-video retrieval, where the goal is to learn accurate cross-modal alignment between videos and text. This problem is challenging because of the matching ambiguity caused by the inherent gap between the heterogeneous video and text modalities. In particular, the differences in the information granularity and abstraction levels between the two modalities hinder a reliable sample-level alignment. Moreover, redundant visual content, sparse textual descriptions, and temporal variability in videos introduce additional uncertainty, resulting in ambiguous matching and suboptimal performance. In this paper, we propose a novel method named Probabilistic Embeddings with Evidence Learning and Refinement (PE2LR), which models video-text pairs as probability distributions and captures uncertainty through the evidence theory. Specifically, we perform distribution-level representation learning to resolve the semantic ambiguity of video-text pairs. To improve the alignment further, we introduce a distribution-based embedding refinement module to ameliorate the semantic consistency across modalities. The proposed PE2LR is able to pull positive sample pairs closer in the embedding space, while pushing the negative pairs apart. Comprehensive experiments on several benchmark datasets (including MSRVTT, DiDeMo, and ActivityNet Captions) demonstrate that our PE2LR achieves state-of-the-art search performance. Donglin Zhang 0001, Zheng Rao, Xing Xu 0001, Xiaojun Wu 0001, Josef Kittler |
IEEE Trans. Image Process. | 1 |
| 2026 | Prototype-Based Asymmetric Hierarchical Matching for Text-Video RetrievalabstractIn recent years, CLIP-based text-video retrieval has made significant progress. Most existing text-video retrieval methods usually directly align global video-level and sentence-level features without establishing intrinsic connections between different modalities. Furthermore, due to the asymmetry between visual and textual information, the symmetrical modeling also has inherent limitations that may affect the search performance. To mitigate these problems, we propose an asymmetric hierarchical text-video retrieval framework based on clustering prototypes. Specifically, beyond the commonly used global features, we generate video instance prototypes using clustering algorithms and derive video object prototypes from video patch tokens. The text instance prototypes can be obtained by the developed Text Prototype Generation module (TPG). Besides, the Joint-Interaction Module (JIM) and the coarse-grained instance alignment scheme are proposed to coordinate different modalities. Furthermore, the cross-correlation matrix among instances is proposed, which can further reduce the heterogeneity between video and text modalities. Extensive experiments on several datasets (including MSRVTT, and MSVD) demonstrate that our method outperforms some existing state-of-the-art methods. The source code is available at:https://github.com/rzheng77/PAHM-text-video-retrieval. Donglin Zhang 0001, Zheng Rao, Xiaojun Wu 0001 |
IEEE Trans. Multim. | 1 |
| 2026 | Beyond Mere Tuning: Harnessing the Full Potential of Prompts for Text-Video RetrievalabstractText-video retrieval has advanced significantly by thanks to the merits of large-scale contrastive languageimage pre-training. Recently, prompt tuning has emerged as a parameter-efficient strategy to solve the substantial training costs associated with these models. However, existing studies are generally confined to simple tuning strategies, without fully exploring the potential of prompts. Unlike its application to singlemodal tasks, prompt tuning in cross-modal learning involves not only the adaptation to specific tasks but also the challenge of dealing with the intricate interaction of semantic information. These aspects remain insufficiently explored, thereby limiting retrieval performance. To this end, we introduce a comprehensive framework for harnessing the full potential of prompts in textvideo retrieval. Specifically, we explore prompts in three respects: i) Task adaptation: we design appropriate prompts to guide the fine-tuning of pre-trained models to perform text-video retrieval tasks. ii) Temporal modelling: to enable temporal modelling, while protecting the original knowledge from being corrupted, we propose a prompt rolling module that enables effective temporal processing without adding extra training parameters. iii) The sampling scope prediction: We introduce a query expansion strategy that exploits the rich semantic information conveyed by multimodal prompt features to define the sampling scope more precisely. Extensive experiments on four mainstream datasets, i.e., MSR-VTT, ActivityNet, DiDeMo, and LSMDC, demonstrate that our method achieves state-of-the-art performance. Donglin Zhang 0001, Xing Xu 0001, Xiaojun Wu 0001, Josef Kittler |
IEEE Trans. Multim. | 1 |
| 2026 | Resilient Semantic Pseudo-Text Embedding for Zero-Shot Video Moment RetrievalabstractWith the explosive growth of video data, video moment retrieval (VMR) has attracted increasing attention due to its ability to localize semantically relevant moments in untrimmed videos. However, existing VMR approaches usually rely on annotated video-text correspondences or temporal annotations, both of which require significant human effort and are costly to scale. Even worse, the inherent subjectivity in manual labeling often introduces inconsistencies into the training data, further complicating the issue. In this article, we investigate the problem of Zero-Shot Video Moment Retrieval (ZS-VMR) and develop a novel method, Resilient Semantic Pseudo-Text Modeling (RSPT). The core of RSPT is to construct semantically rich pseudo-text embeddings through visually guided perturbations. Specifically, RSPT first generates initial pseudo-texts by injecting random noise into visual features and then learns adaptive noise weights by modeling the correlations between these pseudo-texts and visual features. This enables the generation of diverse and semantically aligned representations from multiple perspectives. To ensure alignment with visual semantics and suppress irrelevant noise, RSPT introduces a quality-aware contrastive loss that regularizes the semantic boundaries of pseudo-texts. Extensive experiments on Charades-STA and ActivityNet-Captions show that RSPT outperforms existing competitive baselines, validating its efficacy. Code is available at https://github.com/dmcsy/RSPT . Donglin Zhang 0001, Weixiang Shi, Xiaojun Wu 0001, Josef Kittler |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2025 | Ingredients-Guided and Nutrients-Prompted Network for Food Nutrition EstimationabstractFood plays a vital role in human health, and accurate nutrition estimation is crucial for guiding healthy dietary choices. Traditional biochemical-based assessment methods are often inefficient, costly, and impractical for daily use. With the continuous progress in computer vision, some vision-based nutrition estimation approaches have emerged, typically relying on RGB images alone or in combination with depth images to infer nutritional information. These methods have achieved promising performance and garnered considerable attention. However, these methods often ignore visually imperceptible ingredients such as oil, sugar, and salt, which may significantly influence the estimation of nutritional content. Besides, existing methods lack explicit mechanisms for modeling nutrient-specific information and guiding attention toward nutrition-relevant semantics. To solve the above two issues, we propose a novel ingredients-guided and nutrients-prompted nutrition estimation method. Our method adopts multi-scale feature fusion and integrates RGB and depth modalities to enhance visual representation learning. To account for invisible ingredients, we introduce an ingredients-guided strategy, which enhances the sensitivity to non-visible nutritional factors. Moreover, a nutrient-prompt mechanism is introduced to explicitly guide the focus of the model toward nutrient-relevant attributes during estimation. We validate our method on Nutrition5k, where it consistently outperforms existing state-of-the-art methods, demonstrating its efficacy. Donglin Zhang 0001, Boyuan Ma, Xiaojun Wu 0001, Josef Kittler |
ACM Multimedia | 1 |
| 2025 | Multi-level Encoding with Hierarchical Alignment for Sketch-Based 3D Shape RetrievalabstractSketch-based 3D shape retrieval (SBSR) aims to retrieve 3D shapes using hand-drawn sketches as query inputs. Although existing SBSR methods have achieved promising results, several challenges still require further investigation. First, most existing approaches usually leverage simple aggregation schemes, often failing to capture the intrinsic relationships between views, which limits the effectiveness of 3D shape feature extraction. Second, conventional SBSR primarily focuses on instance-level alignment while ignoring multi-level alignment, which may neglect complex hierarchical relationships. To address these limitations, we propose a novel Multi-level Encoding with Hierarchical Alignment (MEHA) method for SBSR. Specifically, we adopt spatial encoding and view encoding for multiple views of 3D shapes. The proposed aggregation scheme then integrates these multi-level embedded local features to enhance the representation of 3D shape features. Considering the complexity of 3D shapes, MEHA adopts a two-stage training process: the first stage focuses on learning 3D shape features, while the second stage emphasizes modality alignment. Furthermore, we introduce a hierarchical alignment strategy that bridges the modality gap through instance-level, prototype-level, and centre-level alignment. Extensive experiments on two public benchmark datasets demonstrate the superiority of our method, showing that MEHA outperforms the state-of-the-art baselines. Donglin Zhang 0001, Changxing Li, Xiaojun Wu 0001 |
SIGIR | 1 |
| 2025 | Food nutrition estimation with RGB-D fusion module and bidirectional feature pyramid network
Boyuan Ma, Donglin Zhang 0001, Xiaojun Wu 0001 |
Multim. Syst. | 2 |
| 2025 | Modality Fused Class-Proxy With Knowledge Distillation for Zero-Shot Sketch-Based Image RetrievalabstractIn recent years, zero-shot sketch-based image retrieval (ZS-SBIR) task has attracted considerable attention. Although some ZS-SBIR approaches have been proposed, it remains challenging to handle the inherent linkages between the sketch and image domains. Moreover, how to transfer semantic knowledge from seen categories to unseen categories is still an open problem, significantly affecting retrieval performance. In this article, we propose a novel approach Modality Fused Class-Proxy with Knowledge Distillation, named MFCPKD, which develops two novel schemes to remedy the above issues. Specifically, MFCPKD leverages a Modality Fusion Model to learn modality-fused feature embeddings and class proxies. The knowledge distillation is employed for student to learn the feature from seen categories and infer the unknown category through class proxies. Furthermore, three losses constrain the student network to narrow the modality gap between sketch and image domains. Finally, we conduct extensive experiments on three benchmark datasets (Sketchy Ext, TU-Berlin Ext, and QuickDraw Ext) and demonstrate that our MFCPKD method can achieve excellent performance compared to some existing methods in ZS-SBIR scenarios. The code for our project is available athttps://github.com/li1changxing/MFCPKD. Changxing Li, Donglin Zhang 0001, Zhikai Hu, Xiaojun Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Discrete Elective Hashing with Incomplete Labels for Efficient Cross-Modal RetrievalabstractRecently, supervised cross-modal hashing methods have gained considerable attention due to their ability to mine credible semantic relationships between multi-modal data. These methods typically rely on labels to explore semantic relationships provided that labels are always reliable, which, however, may not be true from the practical perspective. In fact, labels may be incomplete, i.e., true-label incomplete and fine-grained incomplete, which makes the performance of the existing methods deteriorated. To this end, we propose a method called Discrete Elective Hashing with Incomplete Labels (DEH-IL), which is designed to alleviate the impact of incomplete labels. Specifically, we introduce a relaxed label scheme that allows the algorithm to automatically mine potential missing information from incomplete labels, which is beneficial for exploring intra-class relationships. Moreover, we propose a novel elective loss that aggregates all estimations from incomplete labels to mine inter-class relationships. Since elective loss does not rely on any single estimation, it can effectively mitigate estimation errors arising from incomplete labels. By combining these two components, DEH-IL can effectively explore both intra-class and inter-class relationships through incomplete labels. Experimental results on benchmark datasets demonstrate the effectiveness of the proposed method. Donglin Zhang 0001, Changxing Li, Mengke Li 0001, Zhikai Hu |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2024 | Key Points Centered Sparse Hashing for Cross-Modal RetrievalabstractSupervised cross-modal hashing methods usually construct a massive undirected weighted graph based on labels for training data, with the aim of learning more structured hash codes by preserving relationships within this graph. However, as the volume of data increases, such an approach demands substantial computational and storage resources and tends to aggregate all data points with paths, even semantically unrelated ones, which undermines the retrieval performance. In this paper, we propose to prune less crucial paths from this graph to obtain a clearer representation of relationships among data points. This not only reduces computational resources but separates semantically unrelated data points. Specifically, we define key points within the graph and retain relationships only between all data points and these key points, resulting in a simplified and more transparent graph that is used to supervise hash code learning. Experimental results on three datasets demonstrate that removing unimportant paths from the relationship graph can lead to the learning of more structured hash codes, thereby improving retrieval performance. Zhikai Hu, Yiu-Ming Cheung, Mengke Li 0001, Weichao Lan, Donglin Zhang 0001 |
ICASSP | 5 |
| 2024 | Novel Clustering Aggregation and Multi-grained Alignment for Image-Text Matching
Xiaojun Wu 0001, Tianyang Xu 0001, Donglin Zhang 0001 |
ICPR (21) | 4 |
| 2024 | Joint Semantic Preserving Sparse Hashing for Cross-Modal RetrievalabstractSupervised cross-modal hashing has received wide attention in recent years. However, existing methods primarily rely on sample-wise semantic relationships to evaluate the semantic similarity between samples, overlooking the impact of label distribution on enhancing retrieval performance. Moreover, the limited representation capability of traditional dense hash codes hinders the preservation of semantic relationship. To overcome these challenges, we propose a new method, Joint Semantic Preserving Sparse Hashing (JSPSH). Specifically, we introduce a new concept of cluster-wise semantic relationship, which leverages label distribution to indicate which samples are more suitable for clustering. Then, we jointly utilize sample-wise and cluster-wise semantic relationships to supervise the learning of hash codes. In this way, JSPSH preserves both kinds of semantic relationships to ensure that more samples with similar semantics are clustered together, thereby achieving better retrieval results. Furthermore, we utilize high-dimensional sparse hash codes that offer stronger representation capability to preserve such more complex semantics. Finally, an interaction term is introduced in hash functions learning stage to further narrow the gap between modalities. Experimental results on three large-scale datasets demonstrate the effectiveness of JSPSH in achieving superior retrieval performance. Codes are available at https://github.com/hutt94/JSPSH. Zhikai Hu, Yiu-Ming Cheung, Mengke Li 0001, Weichao Lan, Donglin Zhang 0001, Qiang Liu 0018 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | VLNet: A Multi-task Network for Joint Vehicle and Lane Detection
Aiqi Feng, Tianyang Xu 0001, Donglin Zhang 0001, Xiaojun Wu 0001 |
ICIG (2) | 4 |
| 2023 | Online supervised collective matrix factorization hashing for cross-modal retrieval
Zhenqiu Shu, Jun Yu 0011, Donglin Zhang 0001, Zhengtao Yu 0001, Xiaojun Wu 0001 |
Appl. Intell. | 4 |
| 2023 | Robust supervised matrix factorization hashing with application to cross-modal retrieval
Zhenqiu Shu, Kailing Yong, Donglin Zhang 0001, Jun Yu 0011, Zhengtao Yu 0001, Xiaojun Wu 0001 |
Neural Comput. Appl. | 3 |
| 2023 | ONION: Online Semantic Autoencoder Hashing for Cross-Modal RetrievalabstractCross-modal hashing (CMH) has recently received increasing attention with the merit of speed and storage in performing large-scale cross-media similarity search. However, most existing cross-media approaches utilize the batch-based mode to update hash functions, without the ability to efficiently handle the online streaming multimedia data. Online hashing can effectively address the preceding issue by using the online learning scheme to incrementally update the hash functions. Nevertheless, the existing online CMH approaches still suffer from several challenges, such as (1) how to efficiently and effectively utilize the supervision information, (2) how to learn more powerful hash functions, and (3) how to solve the binary constraints. To mitigate these limitations, we present a novel online hashing approach named ONION ( ON line semant I c aut O encoder hashi N g). Specifically, it leverages the semantic autoencoder scheme to establish the correlations between binary codes and labels, delivering the power to obtain more discriminative hash codes. Besides, the proposed ONION directly utilizes the label inner product to build the connection between existing data and newly coming data. Therefore, the optimization is less sensitive to the newly arriving data. Equipping a discrete optimization scheme designed to solve the binary constraints, the quantization errors can be dramatically reduced. Furthermore, the hash functions are learned by the proposed autoencoder strategy, making the hash functions more powerful. Extensive experiments on three large-scale databases demonstrate that the performance of our ONION is superior to several recent competitive online and offline cross-media algorithms. Donglin Zhang 0001, Xiaojun Wu 0001 |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2023 | WATCH: Two-Stage Discrete Cross-Media HashingabstractDue to the explosive growth of multimedia data in recent years, cross-media hashing (CMH) approaches have recently received increasing attention. To learn the hash codes, most existing supervised CMH algorithms employ the strict binary label information, which has small margins between the incorrect labels (0) and the true labels (1), increasing the risk of classification error. Besides, most existing CMH approaches are one-stage algorithms, in which the hash functions and binary codes can be learned simultaneously, complicating the optimization. To avoid NP-hard optimization, many approaches utilize a relaxation strategy. However, this optimisation trick may cause large quantization errors. To address this, we present a novel tWo-stAge discreTe Cross-media Hashing method based on smooth matrix factorization and label relaxation, named WATCH. The proposed WATCH controls the margins adaptively by the novel label relaxation strategy. This innovation reduces the quantization error significantly. Besides, WATCH is a two-stage model. In stage 1, we employ a discrete smooth matrix factorization model. Then, the hash codes can be generated discretely, reducing the large quantization loss greatly. In stage 2, we adopt an effective hash function learning strategy, which produces more effective hash functions. Comprehensive experiments on several datasets demonstrate that WATCH outperforms some state-of-the-art methods. Donglin Zhang 0001, Xiaojun Wu 0001, Tianyang Xu 0001, Josef Kittler |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2023 | DAH: Discrete Asymmetric Hashing for Efficient Cross-Media RetrievalabstractGiven the merits in high computational efficiency and low storage cost, hashing techniques have been widely studied in cross-media retrieval. Existing methods usually adopt the equal length encoding scheme to represent the multimedia data. However, the strictly equal length scheme maybe not optimal because the dimension of different modalities is often various. Besides, there exists other challenges in designing a cross-media retrieval system, e.g., how to address the discrete constraints, how to avoid using the n*n similarity matrix, and how to effectively exploit the discriminative label information. To conquer the above challenges, we propose a novel method, i.e., discrete asymmetric hashing (DAH). Specifically, DAH exploits a flexible model, which can seamlessly deal with equal or unequal length encoding scenarios. Moreover, DAH constructs a supervised semantic embedding framework by jointly minimizing the distance-distance difference and label reconstructing error, significantly reducing the computational complexity. An asymmetric strategy is employed to establish the connection between hash codes and the latent subspace. Furthermore, the hash codes can be learned discretely by the designed optimization algorithm. In the training stage2, a semantic intersection scheme is proposed to learn more powerful hash functions. Experiments show that our DAH is effective in equal and unequal scenarios. Donglin Zhang 0001, Xiaojun Wu 0001, Tianyang Xu 0001, He-Feng Yin |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2022 | Adaptive multi-modal fusion hashing via Hadamard matrix
Jun Yu 0011, Donglin Zhang 0001, Zhenqiu Shu |
Appl. Intell. | 2 |
| 2022 | Specific class center guided deep hashing for cross-modal retrieval
Zhenqiu Shu, Yibing Bai, Donglin Zhang 0001, Jun Yu 0011, Zhengtao Yu 0001, Xiaojun Wu 0001 |
Inf. Sci. | 3 |
| 2022 | Robust and discrete matrix factorization hashing for cross-modal retrieval
Donglin Zhang 0001, Xiaojun Wu 0001 |
Pattern Recognit. | 1 |
| 2022 | Scalable Discrete Matrix Factorization and Semantic Autoencoder for Cross-Media RetrievalabstractHashing methods have sparked great attention on multimedia tasks due to their effectiveness and efficiency. However, most existing methods generate binary codes by relaxing the binary constraints, which may cause large quantization error. In addition, most supervised cross-modal approaches preserve the similarity relationship by constructing an n×n large-size similarity matrix, which requires huge computation, making these methods unscalable. To address the above challenges, this article presents a novel algorithm, called scalable discrete matrix factorization and semantic autoencoder method (SDMSA). SDMSA is a two-stage method. In the first stage, the matrix factorization scheme is utilized to learn the latent semantic information, the label matrix is incorporated into the loss function instead of the similarity matrix. Thereafter, the binary codes can be generated by the latent representations. During optimization, we can avoid manipulating a large n×n similarity matrix, and the hash codes can be generated directly. In the second stage, a novel hash function learning scheme based on the autoencoder is proposed. The encoder-decoder paradigm aims to learn projections, the feature vectors are projected to code vectors by encoder, and the code vectors are projected back to the original feature vectors by the decoder. The encoder-decoder scheme ensures the embedding can well preserve both the semantic and feature information. Specifically, two algorithms SDMSA-lin and SDMSA-ker are developed under the SDMSA framework. Owing to the merit of SDMSA, we can get more semantically meaningful binary hash codes. Extensive experiments on several databases show that SDMSA-lin and SDMSA-ker achieve promising performance. Donglin Zhang 0001, Xiaojun Wu 0001 |
IEEE Trans. Cybern. | 1 |
| 2022 | Two-Stage Supervised Discrete Hashing for Cross-Modal RetrievalabstractRecently, hashing-based multimodal learning systems have received increasing attention due to their query efficiency and parsimonious storage costs. However, impeded by the quantization loss caused by numerical optimization, the existing cross-media hashing approaches are unable to capture all the discriminative information present in the original multimodal data. Besides, most cross-modal methods belong to the one-step paradigm, which learn the binary codes and hash function simultaneously, increasing the complexity of optimization. To address these issues, we propose a novel two-stage approach, named the two-stage supervised discrete hashing (TSDH) method. In particular, in the first phase, TSDH generates a latent representation for each modality. These representations are then mapped to a common Hamming space to generate the binary codes. In addition, TSDH directly endows the hash codes with the semantic labels, enhancing the discriminatory power of the learned binary codes. A discrete hash optimization approach is developed to learn the binary codes without relaxation, avoiding the large quantization loss. The proposed hash function learning scheme reuses the semantic information contained by the embeddings, endowing the hash functions with enhanced discriminability. Extensive experiments on several databases demonstrate the effectiveness of the developed TSDH, outperforming several recent competitive cross-media algorithms. Donglin Zhang 0001, Xiaojun Wu 0001, Tianyang Xu 0001, Josef Kittler |
IEEE Trans. Syst. Man Cybern. Syst. | 1 |
| 2021 | Discrete Bidirectional Matrix Factorization Hashing for Zero-Shot Cross-Media Retrieval
Donglin Zhang 0001, Xiaojun Wu 0001, Jun Yu 0011 |
PRCV (2) | 1 |
| 2021 | Learning latent hash codes with discriminative structure preserving for cross-modal retrieval
Donglin Zhang 0001, Xiaojun Wu 0001, Jun Yu 0011 |
Pattern Anal. Appl. | 1 |
| 2021 | MOON: Multi-hash codes joint learning for cross-media retrieval
Donglin Zhang 0001, Xiaojun Wu 0001, He-Feng Yin, Josef Kittler |
Pattern Recognit. Lett. | 1 |
| 2021 | Label Consistent Flexible Matrix Factorization Hashing for Efficient Cross-modal RetrievalabstractHashing methods have sparked a great revolution on large-scale cross-media search due to its effectiveness and efficiency. Most existing approaches learn unified hash representation in a common Hamming space to represent all multimodal data. However, the unified hash codes may not characterize the cross-modal data discriminatively, because the data may vary greatly due to its different dimensionalities, physical properties, and statistical information. In addition, most existing supervised cross-modal algorithms preserve the similarity relationship by constructing an n × n pairwise similarity matrix, which requires a large amount of calculation and loses the category information. To mitigate these issues, a novel cross-media hashing approach is proposed in this article, dubbed label flexible matrix factorization hashing (LFMH). Specifically, LFMH jointly learns the modality-specific latent subspace with similar semantic by the flexible matrix factorization. In addition, LFMH guides the hash learning by utilizing the semantic labels directly instead of the large n × n pairwise similarity matrix. LFMH transforms the heterogeneous data into modality-specific latent semantic representation. Therefore, we can obtain the hash codes by quantifying the representations, and the learned hash codes are consistent with the supervised labels of multimodal data. Then, we can obtain the similar binary codes of the corresponding modality, and the binary codes can characterize such samples flexibly. Accordingly, the derived hash codes have more discriminative power for single-modal and cross-modal retrieval tasks. Extensive experiments on eight different databases demonstrate that our model outperforms some competitive approaches. Donglin Zhang 0001, Xiaojun Wu 0001, Jun Yu 0011 |
ACM Trans. Multim. Comput. Commun. Appl. | 1 |
| 2020 | Fast Discrete Cross-Modal Hashing Based on Label Relaxation and Matrix FactorizationabstractIn recent years, cross-media retrieval has drawn considerable attention due to the exponential growth of multimedia data. Many hashing approaches have been proposed for the cross-media search task. However, there are still open problems that warrant investigation. For example, most existing supervised hashing approaches employ a binary label matrix, which achieves small margins between wrong labels (0) and true labels (1). This may affect the retrieval performance by generating many false negatives and false positives. In addition, some methods adopt a relaxation scheme to solve the binary constraints, which may cause large quantization errors. There are also some discrete hashing methods that have been presented, but most of them are time-consuming. To conquer these problems, we present a label relaxation and discrete matrix factorization method (LRMF) for cross-modal retrieval. It offers a number of innovations. First of all, the proposed approach employs a novel label relaxation scheme to control the margins adaptively, which has the benefit of reducing the quantization error. Second, by virtue of the proposed discrete matrix factorization method designed to learn the binary codes, large quantization errors caused by relaxation can be avoided. The experimental results obtained on two widely-used databases demonstrate that LRMF outperforms state-of-the-art cross- media methods. Donglin Zhang 0001, Xiaojun Wu 0001, Zhen Liu 0015, Jun Yu 0011, Josef Kittler |
ICPR | 1 |