EDBT 2026 Demo / reviewers in the wild / expert
Zhikai Hu
dblp:210/4997
· DBLP profile ↗
23ranked-venue papers
7as first author
20since 2021 · last 2026
0000-0001-7278-9977ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 2 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 6 first-author · 10 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DRFGD: Disentangled Representation-Focused Generative Defense for Attack-Tolerant Cross-Modal HashingabstractWith the widespread deployment of cross-modal retrieval in real-world scenarios, ensuring robustness against adversarial attacks is increasingly critical. Remarkably, deep cross-modal hashing is highly vulnerable to adversarial attacks due to its discrete nature and low-dimensional hash codes, while existing defense methods often fail to suppress perturbations embedded in vulnerable features and lack the capacity to model modality-specific structural differences, resulting in suboptimal adversarial robustness. To address these challenges, we propose a novel Disentangled Representation-Focused Generative Defense (DRFGD) framework for attack-tolerant cross-modal hashing. Without altering the structure of retrieval model, DRFGD defends against adversarial attacks by disentangling input representations into adversarial-robust and adversarial-vulnerable components, by an efficient dual-branch semantic-aware encoder. Guided by such disentangled robust features, an attack-tolerant generative module is seamlessly designed to synthesize semantically aligned and perturbation-resilient examples for robust adversarial training, thereby significantly promoting collaborative defense robustness to attackers. Consequently, the semantically consistent hash codes can be well obtained to enhance adversarial robustness in complex cross-modal attacking scenarios. Extensive experiments on public benchmarks demonstrate that DRFGD substantially improves retrieval robustness under various attacking scenarios, and shows its improved defense performance in comparison with the SOTA works. Zhongqing Yu, Xin Liu 0011, Yiu-Ming Cheung, Zhikai Hu, Wentao Fan 0001, Pan Zhou 0001 |
AAAI | 4 |
| 2026 | SketchMind: Understanding Abstract Sketches with MLLMs for Fine-Grained Sketch-Based Image Retrieval
Changxing Li, Donglin Zhang 0001, Zhikai Hu, Xiaojun Wu 0001, Josef Kittler |
WWW | 3 |
| 2026 | MLCA: Multi-level Correlative Attacks against Deep Cross-Modal Hashing
Xiaohang Fang, Xin Liu 0011, Zhikai Hu, Yiu-Ming Cheung, Shu-Juan Peng, Xing Xu 0001 |
Pattern Recognit. | 3 |
| 2026 | RSPH: Robust self-paced hashing for cross-modal retrieval
Donglin Zhang 0001, Zhikai Hu, Xiaojun Wu 0001, Josef Kittler |
Pattern Recognit. | 2 |
| 2026 | Supervised discrete cross-modal hashing with exploiting semantic correlations
Donglin Zhang 0001, Zhikai Hu, Xiaojun Wu 0001, Josef Kittler |
Pattern Recognit. | 2 |
| 2026 | HTVR: Hierarchical text-to-video retrieval based on relative similarity
Donglin Zhang 0001, Zhikai Hu |
Pattern Recognit. | 3 |
| 2026 | Parameterization of Volume Sampling for Active Learning of Radiance FieldabstractRadiance field techniques have demonstrated remarkable performance in reconstructing photorealistic real- world scenes. However, a widely recognized limitation of these methods is their reliance on densely captured images for training. Active learning offers a promising solution by selecting the most informative images to capture. An effective measure of data informativeness is the mutual information between the unknown image and the model parameters, known as the expected information gain. Naive estimation of this quantity requires inferring both the model and predictive distributions, which is computationally intractable for high-dimensional parameter spaces. In this work, we propose a computationally tractable method to estimate information gain. Instead of sampling in the high-dimensional model parameter space, we leverage the ray sampling process inherent in volume rendering to approximate the expected information gain. Specifically, we parameterize volume sampling with perturbed ray directions and learn the predictive distribution to infer optimal perturbation patterns. Furthermore, we derive an empirical risk decomposition that demonstrates how our method effectively explores the volume sampling space to enhance diversity, leading to an efficient and informative view selection. Experiments show that our approach achieves state-of-the-art performance in image quality for active learning of radiance fields, outperforming previous methods across various datasets, including both forward-facing and object-centric scenes. Yiu-Ming Cheung, Zhikai Hu, Yonggang Zhang 0003 |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2025 | Component-Level Segmentation for Oracle Bone Inscription DeciphermentabstractOracle Bone Inscriptions (OBIs), as the earliest systematically organized pictographic script in China, hold significant importance in the study of the origins of Chinese civilization. Of the approximately 4,500 excavated OBI characters, only about one-third have been deciphered, leaving the remaining characters shrouded in mystery. Over the past decade, an increasing number of researchers have attempted to leverage artificial intelligence to assist in deciphering OBIs, but these efforts have not yet fully met the demands of this challenging objective. In this paper, we identify a key task—Component-Level OBI Segmentation—based on a successful deciphering case from 2018. This task aims to help experts quickly identify specific components within OBIs, thereby accelerating the deciphering process. Accordingly, we propose a new model to accomplish this task. Our model leverages a small amount of annotated data and a large amount of weakly annotated data and incorporates expert-provided prior knowledge, i.e., stroke rules, to automatically segment OBI components. Additionally, we train a series of auxiliary classifiers to evaluate the segmentation results during the test stage. We also invite experts to conduct a professional assessment of the results, which we cross-validated against our proposed evaluation metrics. Experimental results demonstrate that our method can accurately and clearly present the segmented components to experts. Zhikai Hu, Yiu-Ming Cheung, Yonggang Zhang 0003, Peiying Zhang 0003, Puiling Tang |
AAAI | 1 |
| 2025 | Modality Fused Class-Proxy With Knowledge Distillation for Zero-Shot Sketch-Based Image RetrievalabstractIn recent years, zero-shot sketch-based image retrieval (ZS-SBIR) task has attracted considerable attention. Although some ZS-SBIR approaches have been proposed, it remains challenging to handle the inherent linkages between the sketch and image domains. Moreover, how to transfer semantic knowledge from seen categories to unseen categories is still an open problem, significantly affecting retrieval performance. In this article, we propose a novel approach Modality Fused Class-Proxy with Knowledge Distillation, named MFCPKD, which develops two novel schemes to remedy the above issues. Specifically, MFCPKD leverages a Modality Fusion Model to learn modality-fused feature embeddings and class proxies. The knowledge distillation is employed for student to learn the feature from seen categories and infer the unknown category through class proxies. Furthermore, three losses constrain the student network to narrow the modality gap between sketch and image domains. Finally, we conduct extensive experiments on three benchmark datasets (Sketchy Ext, TU-Berlin Ext, and QuickDraw Ext) and demonstrate that our MFCPKD method can achieve excellent performance compared to some existing methods in ZS-SBIR scenarios. The code for our project is available athttps://github.com/li1changxing/MFCPKD. Changxing Li, Donglin Zhang 0001, Zhikai Hu, Xiaojun Wu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Convolution Filter Compression via Sparse Linear Combinations of Quantized BasisabstractConvolutional neural networks (CNNs) have achieved significant performance on various real-life tasks. However, the large number of parameters in convolutional layers requires huge storage and computation resources, making it challenging to deploy CNNs on memory-constrained embedded devices. In this article, we propose a novel compression method that generates the convolution filters in each layer using a set of learnable low-dimensional quantized filter bases. The proposed method reconstructs the convolution filters by stacking the linear combinations of these filter bases. By using quantized values in weights, the compact filters can be represented using fewer bits so that the network can be highly compressed. Furthermore, we explore the sparsity of coefficients through $L_{1}$ -ball projection when conducting linear combination to further reduce the storage consumption and prevent overfitting. We also provide a detailed analysis of the compression performance of the proposed method. Evaluations of image classification and object detection tasks using various network structures demonstrate that the proposed method achieves a higher compression ratio with comparable accuracy compared with the existing state-of-the-art filter decomposition and network quantization methods. Weichao Lan, Yiu-Ming Cheung, Liang Lan, Juyong Jiang, Zhikai Hu |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2025 | Discrete Elective Hashing with Incomplete Labels for Efficient Cross-Modal RetrievalabstractRecently, supervised cross-modal hashing methods have gained considerable attention due to their ability to mine credible semantic relationships between multi-modal data. These methods typically rely on labels to explore semantic relationships provided that labels are always reliable, which, however, may not be true from the practical perspective. In fact, labels may be incomplete, i.e., true-label incomplete and fine-grained incomplete, which makes the performance of the existing methods deteriorated. To this end, we propose a method called Discrete Elective Hashing with Incomplete Labels (DEH-IL), which is designed to alleviate the impact of incomplete labels. Specifically, we introduce a relaxed label scheme that allows the algorithm to automatically mine potential missing information from incomplete labels, which is beneficial for exploring intra-class relationships. Moreover, we propose a novel elective loss that aggregates all estimations from incomplete labels to mine inter-class relationships. Since elective loss does not rely on any single estimation, it can effectively mitigate estimation errors arising from incomplete labels. By combining these two components, DEH-IL can effectively explore both intra-class and inter-class relationships through incomplete labels. Experimental results on benchmark datasets demonstrate the effectiveness of the proposed method. Donglin Zhang 0001, Changxing Li, Mengke Li 0001, Zhikai Hu |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2024 | Feature Fusion from Head to Tail for Long-Tailed Visual RecognitionabstractThe imbalanced distribution of long-tailed data presents a considerable challenge for deep learning models, as it causes them to prioritize the accurate classification of head classes but largely disregard tail classes. The biased decision boundary caused by inadequate semantic information in tail classes is one of the key factors contributing to their low recognition accuracy. To rectify this issue, we propose to augment tail classes by grafting the diverse semantic information from head classes, referred to as head-to-tail fusion (H2T). We replace a portion of feature maps from tail classes with those belonging to head classes. These fused features substantially enhance the diversity of tail classes. Both theoretical analysis and practical experimentation demonstrate that H2T can contribute to a more optimized solution for the decision boundary. We seamlessly integrate H2T in the classifier adjustment stage, making it a plug-and-play module. Its simplicity and ease of implementation allow for smooth integration with existing long-tailed recognition methods, facilitating a further performance boost. Extensive experiments on various long-tailed benchmarks demonstrate the effectiveness of the proposed H2T. The source code is available at https://github.com/Keke921/H2T. Mengke Li 0001, Zhikai Hu, Yang Lu 0009, Weichao Lan, Yiu-Ming Cheung, Hui Huang 0004 |
AAAI | 2 |
| 2024 | Key Points Centered Sparse Hashing for Cross-Modal RetrievalabstractSupervised cross-modal hashing methods usually construct a massive undirected weighted graph based on labels for training data, with the aim of learning more structured hash codes by preserving relationships within this graph. However, as the volume of data increases, such an approach demands substantial computational and storage resources and tends to aggregate all data points with paths, even semantically unrelated ones, which undermines the retrieval performance. In this paper, we propose to prune less crucial paths from this graph to obtain a clearer representation of relationships among data points. This not only reduces computational resources but separates semantically unrelated data points. Specifically, we define key points within the graph and retain relationships only between all data points and these key points, resulting in a simplified and more transparent graph that is used to supervise hash code learning. Experimental results on three datasets demonstrate that removing unimportant paths from the relationship graph can lead to the learning of more structured hash codes, thereby improving retrieval performance. Zhikai Hu, Yiu-Ming Cheung, Mengke Li 0001, Weichao Lan, Donglin Zhang 0001 |
ICASSP | 1 |
| 2024 | Component-Level Oracle Bone Inscription RetrievalabstractOracle Bone Inscriptions (OBIs) represent the early pictographic writing system of the matured Chinese civilization, documenting the history of the Shang dynasty. Deciphering them holds significant importance for unraveling the origins of civilization. Recently, an increasing number of algorithms have been proposed to assist in deciphering OBIs. However, most of these efforts have focused on the character level, thus offering limited assistance. Considering the presence of many similar components within OBI characters, associating different OBI characters with the same component will facilitate OBI decipherment. In this paper, we therefore propose a component-level OBI retrieval task, i.e., using an OBI component to retrieve all OBI characters containing this component. We accordingly collect a dataset, termed OBI component 20, containing 10,257 OBIs, which is annotated by OBI experts. Then, we propose a dual-stream attention-based model and two types of triplets based on components and characters as anchors to model the relationships between components and characters. Specifically, these two types of triplets ensure that characters containing different components are further apart, while those containing the same component are closer to the corresponding component. Experimental results demonstrate the effectiveness of our proposed model. Zhikai Hu, Yiu-Ming Cheung, Yonggang Zhang 0003, Peiying Zhang 0003, Puiling Tang |
ICMR | 1 |
| 2024 | Cross-Modal Hashing Method With Properties of Hamming Space: A New PerspectiveabstractCross-modal hashing (CMH) has attracted considerable attention in recent years. Almost all existing CMH methods primarily focus on reducing the modality gap and semantic gap, i.e., aligning multi-modal features and their semantics in Hamming space, without taking into account the space gap, i.e., difference between the real number space and the Hamming space. In fact, the space gap can affect the performance of CMH methods. In this paper, we analyze and demonstrate how the space gap affects the existing CMH methods, which therefore raises two problems: solution space compression and loss function oscillation. These two problems eventually cause the retrieval performance deteriorating. Based on these findings, we propose a novel algorithm, namely Semantic Channel Hashing (SCH). First, we classify sample pairs into fully semantic-similar, partially semantic-similar, and semantic-negative ones based on their similarity and impose different constraints on them, respectively, to ensure that the entire Hamming space is utilized. Then, we introduce a semantic channel to alleviate the issue of loss function oscillation. Experimental results on three public datasets demonstrate that SCH outperforms the state-of-the-art methods. Furthermore, experimental validations are provided to substantiate the conjectures regarding solution space compression and loss function oscillation, offering visual evidence of their impact on the CMH methods. Zhikai Hu, Yiu-Ming Cheung, Mengke Li 0001, Weichao Lan |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | Compact Neural Network via Stacking Hybrid UnitsabstractAs an effective tool for network compression, pruning techniques have been widely used to reduce the large number of parameters in deep neural networks (NNs). Nevertheless, unstructured pruning has the limitation of dealing with the sparse and irregular weights. By contrast, structured pruning can help eliminate this drawback but it requires complex criteria to determine which components to be pruned. Therefore, this paper presents a new method termed BUnit-Net, which directly constructs compact NNs by stacking designed basic units, without requiring additional judgement criteria anymore. Given the basic units of various architectures, they are combined and stacked systematically to build up compact NNs which involve fewer weight parameters due to the independence among the units. In this way, BUnit-Net can achieve the same compression effect as unstructured pruning while the weight tensors can still remain regular and dense. We formulate BUnit-Net in diverse popular backbones in comparison with the state-of-the-art pruning methods on different benchmark datasets. Moreover, two new metrics are proposed to evaluate the trade-off of compression performance. Experiment results show that BUnit-Net can achieve comparable classification accuracy while saving around 80% FLOPs and 73% parameters. That is, stacking basic units provides a new promising way for network compression. Weichao Lan, Yiu-Ming Cheung, Juyong Jiang, Zhikai Hu, Mengke Li 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | Joint Semantic Preserving Sparse Hashing for Cross-Modal RetrievalabstractSupervised cross-modal hashing has received wide attention in recent years. However, existing methods primarily rely on sample-wise semantic relationships to evaluate the semantic similarity between samples, overlooking the impact of label distribution on enhancing retrieval performance. Moreover, the limited representation capability of traditional dense hash codes hinders the preservation of semantic relationship. To overcome these challenges, we propose a new method, Joint Semantic Preserving Sparse Hashing (JSPSH). Specifically, we introduce a new concept of cluster-wise semantic relationship, which leverages label distribution to indicate which samples are more suitable for clustering. Then, we jointly utilize sample-wise and cluster-wise semantic relationships to supervise the learning of hash codes. In this way, JSPSH preserves both kinds of semantic relationships to ensure that more samples with similar semantics are clustered together, thereby achieving better retrieval results. Furthermore, we utilize high-dimensional sparse hash codes that offer stronger representation capability to preserve such more complex semantics. Finally, an interaction term is introduced in hash functions learning stage to further narrow the gap between modalities. Experimental results on three large-scale datasets demonstrate the effectiveness of JSPSH in achieving superior retrieval performance. Codes are available at https://github.com/hutt94/JSPSH. Zhikai Hu, Yiu-Ming Cheung, Mengke Li 0001, Weichao Lan, Donglin Zhang 0001, Qiang Liu 0018 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2023 | Few-Shot Lip-Password Based Speaker VerificationabstractLip-password has provided a promising solution for speaker verification (Liu and Cheung 2014). Despite the potential of this technology, there are few related studies, largely attributed to the lack of corresponding public datasets. Furthermore, previous works in this field generally demand a substantial amount of training samples and negative samples, impeding their applications from a practical perspective. Therefore, this paper collects a lip-password dataset and proposes a novel few-shot lip-password based speaker verification model, which can be effectively deployed in real-world scenarios because only a small number of data are required for training. Specifically, with an analysis of lip-password features, a down-sampling strategy is presented to generate more training samples. To compensate for the information loss caused by this strategy, a few-shot model, consisting of global and local models, is designed to simultaneously verify the global and local information of the lip-password. Speaker identity is verified only if both stages are passed. The efficacy of the proposed method is demonstrated using the newly collected dataset. Zhikai Hu, Yiu-Ming Cheung, Mengke Li 0001, Weichao Lan |
ICIP | 1 |
| 2023 | Key Point Sensitive Loss for Long-Tailed Visual RecognitionabstractFor long-tailed distributed data, existing classification models often learn overwhelmingly on the head classes while ignoring the tail classes, resulting in poor generalization capability. To address this problem, we thereby propose a new approach in this paper, in which a key point sensitive (KPS) loss is presented to regularize the key points strongly to improve the generalization performance of the classification model. Meanwhile, in order to improve the performance on tail classes, the proposed KPS loss also assigns relatively large margins on tail classes. Furthermore, we propose a gradient adjustment (GA) optimization strategy to re-balance the gradients of positive and negative samples for each class. By virtue of the gradient analysis of the loss function, it is found that the tail classes always receive negative signals during training, which misleads the tail prediction to be biased towards the head. The proposed GA strategy can circumvent excessive negative signals on tail classes and further improve the overall classification accuracy. Extensive experiments conducted on long-tailed benchmarks show that the proposed method is capable of significantly improving the classification accuracy of the model in tail classes while maintaining competent performance in head classes. Mengke Li 0001, Yiu-Ming Cheung, Zhikai Hu |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2021 | MTFH: A Matrix Tri-Factorization Hashing Framework for Efficient Cross-Modal RetrievalabstractHashing has recently sparked a great revolution in cross-modal retrieval because of its low storage cost and high query speed. Recent cross-modal hashing methods often learn unified or equal-length hash codes to represent the multi-modal data and make them intuitively comparable. However, such unified or equal-length hash representations could inherently sacrifice their representation scalability because the data from different modalities may not have one-to-one correspondence and could be encoded more efficiently by different hash codes of unequal lengths. To mitigate these problems, this paper exploits a related and relatively unexplored problem: encode the heterogeneous data with varying hash lengths and generalize the cross-modal retrieval in various challenging scenarios. To this end, a generalized and flexible cross-modal hashing framework, termed Matrix Tri-Factorization Hashing (MTFH), is proposed to work seamlessly in various settings including paired or unpaired multi-modal data, and equal or varying hash length encoding scenarios. More specifically, MTFH exploits an efficient objective function to flexibly learn the modality-specific hash codes with different length settings, while synchronously learning two semantic correlation matrices to semantically correlate the different hash representations for heterogeneous data comparable. As a result, the derived hash codes are more semantically meaningful for various challenging cross-modal retrieval tasks. Extensive experiments evaluated on public benchmark datasets highlight the superiority of MTFH under various retrieval scenarios and show its competitive performance with the state-of-the-arts. Xin Liu 0011, Zhikai Hu, Haibin Ling, Yiu-Ming Cheung |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2019 | Fast Semantic Preserving Hashing for Large-Scale Cross-Modal RetrievalabstractMost Cross-modal hashing methods do not sufficiently exploit the discrimination power of semantic information when learning hash codes, while often involving time-consuming training procedures for large-scale dataset. To tackle these issues, we first formulate the learning of similarity-preserving hash codes in terms of orthogonally rotating the semantic data to hamming space, and then propose a novel Fast Semantic Preserving Hashing (FSePH) approach to large-scale cross-modal retrieval. Specifically, FSePH introduces an orthonormal basis to regress the targeted hash codes of training examples to their corresponding reasonably relaxed class labels, featuring significantly reducing the quantization error. Meanwhile, an effective optimization algorithm is derived for modality-specific projection function learning and an efficient closed-form solution for hash code learning, which are computationally tractable. Extensive experiments have shown that the proposed FSePH approach runs sufficiently fast, and also significantly improves the retrieval performances over the state-of-the-arts. Xingzhi Wang, Xin Liu 0011, Shu-Juan Peng, Yiu-Ming Cheung, Zhikai Hu, Nannan Wang 0001 |
ICDM | 5 |
| 2019 | Semi-Supervised Semantic-Preserving Hashing for Efficient Cross-Modal RetrievalabstractCross-modal hashing has recently gained significant popularity to facilitate retrieval across different modalities. With limited label available, this paper presents a novel Semi-Supervised Semantic-Preserving Hashing (S3PH) for flexible cross-modal retrieval. In contrast to most semi-supervised cross-modal hashing works that need to predict the label of unlabeled data, our proposed approach groups the labeled and unlabeled data together, and integrates the relaxed latent subspace learning and semantic-preserving regularization across different modalities. Accordingly, an efficient relaxed objective function is proposed to learn the latent subspaces for both labeled and unlabeled data. Further, an orthogonal rotation matrix is efficiently learned to transform the latent subspace to hash space by minimizing the quantization error. Without sacrificing the retrieval performance, the proposed S3PH method can benefit various kinds of retrieval tasks, i.e., unsupervised, semi-supervised and supervised. Experimental results compared with several competitive algorithms show the effectiveness of the proposed method and its superiority over state-of-the-arts. Xingzhi Wang, Xin Liu 0011, Zhikai Hu, Nannan Wang 0001, Wentao Fan 0001, Jixiang Du |
ICME | 3 |
| 2019 | Triplet Fusion Network Hashing for Unpaired Cross-Modal RetrievalabstractWith the dramatic increase of multi-media data on the Internet, cross-modal retrieval has become an important and valuable task in searching systems. The key challenge of this task is how to build the correlation between multi-modal data. Most existing approaches only focus on dealing with paired data. They use pairwise relationship of multi-modal data for exploring the correlation between them. However, in practice, unpaired data are more common on the Internet but few methods pay attention to them. To utilize both paired and unpaired data, we propose a one-stream framework triplet fusion network hashing (TFNH), which mainly consists of two parts. The first part is a triplet network which is used to handle both kinds of data, with the help of zero padding operation. The second part consists of two data classifiers, which are used to bridge the gap between paired and unpaired data. In addition, we embed manifold learning into the framework for preserving both inter and intra modal similarity, exploring the relationship between unpaired and paired data and bridging the gap between them in learning process. Extensive experiments show that the proposed approach outperforms several state-of-the-art methods on two datasets in paired scenario. We further evaluate its ability of handling unpaired scenario and robustness in regard to pairwise constraint. The results show that even we discard 50% data under the setting in [19], the performance of TFNH is still better than that of other unpaired approaches and that only 70% pairwise relationships are preserved, TFNH can still outperform almost all paired approaches. Zhikai Hu, Xin Liu 0011, Xingzhi Wang, Yiu-Ming Cheung, Nannan Wang 0001, Yewang Chen |
ICMR | 1 |