EDBT 2026 Demo / reviewers in the wild / expert
Chaoqun Zheng
dblp:187/5769
· DBLP profile ↗
24ranked-venue papers
7as first author
22since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021Databases, data management, data science and information retrieval · 7 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 5 since 2021Computer networks · 3 · 3 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Depression Detection from Social Media: A Mutual Guidance Multi-modal Network with Complementary Graph LearningabstractDepression has become a critical global public health challenge, creating an urgent need for automated and scalable screening solutions. Social media platforms, which capture rich and spontaneous multi-modal behavioral data, offer a promising avenue for detecting early signs of mental distress. However, existing depression detection methods predominantly rely on static multi-modal fusion strategies and frequently fail to effectively tackle cross-modal semantic gaps. To address these limitations, we propose a Mutual Guidance Multi-modal Network with Complementary Graph Learning (MGMN) for depression detection by observing individuals' behavioral performance on social media. Specifically, a cross-modal mutual guidance mechanism is designed to dynamically construct a complementary graph by using mutual similarities within and across visual and acoustic modalities common in social media. More specifically, based on this complementary graph, a modality-specific adaptive residual learning module is applied to each modality to stabilize deep feature learning and preserve modality-specific and complementary information via graph-conditioned adaptive residual fusion. Furthermore, the refined uni-modal features are subsequently fed into a joint-modal fusion and prediction module to output the final disease prediction probability. Extensive experiments on the MUD3, LMVD, and D-vlog datasets demonstrate our proposed method's superiority over state-of-the-art methods, confirming that the proposed framework provides a robust and effective solution for mental health monitoring. Codes are available at https://github.com/Petofi-romance/MGMN Guocheng Hu, Chaoqun Zheng, Ruifan Zuo, Fengling Li 0001, Dan Shi 0003, Xiaofeng Qu, Wenpeng Lu |
SIGIR | 2 |
| 2026 | Scaffolding thought: Imposing logical structure on LLMs with knowledge graphs for counterfactual generation
Jiasheng Si, Yingjie Zhu, Yeqing Teng, Rui Wang 0043, Tianyi Wang 0006, Weiyu Zhang 0001, Chaoqun Zheng, Wenpeng Lu |
Knowl. Based Syst. | 7 |
| 2025 | Learning Together Securely: Prototype-Based Federated Multi-Modal Hashing for Safe and Efficient Multi-Modal RetrievalabstractWith the proliferation of multi-modal data, safe and efficient multi-modal hashing retrieval has become a pressing research challenge, particularly due to concerns over data privacy during centralized processing. To address this, we propose Prototype-based Federated Multi-modal Hashing (PFMH), an innovative framework that seamlessly integrates federated learning with multi-modal hashing techniques. PFMH achieves fine-grained fusion of heterogeneous multi-modal data, enhancing retrieval accuracy while ensuring data privacy through prototype-based communication, thereby reducing communication costs and mitigating risks of data leakage. Furthermore, using a prototype completion strategy, PFMH tackles class imbalance and statistical heterogeneity in multi-modal data, improving model generalization and performance across diverse data distributions. Extensive experiments demonstrate the efficiency and effectiveness of PFMH within the federated learning framework, enabling distributed training for secure and precise multi-modal retrieval in real-world scenarios. Ruifan Zuo, Chaoqun Zheng, Lei Zhu 0002, Wenpeng Lu, Yuanyuan Xiang, Xiaofeng Qu |
AAAI | 2 |
| 2025 | MADAWSD: Multi-Agent Debate Framework for Adversarial Word Sense DisambiguationabstractKaiyuan Zhang, Qian Liu, Luyang Zhang, Chaoqun Zheng, Shuaimin Li, Bing Xu, Muyun Yang, Xinxiao Qiao, Wenpeng Lu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Qian Liu 0012, Chaoqun Zheng, Shuaimin Li, Muyun Yang, Xinxiao Qiao, Wenpeng Lu |
EMNLP | 4 |
| 2025 | Plan Dynamically, Express Rhetorically: A Debate-Driven Rhetorical Framework for Argumentative WritingabstractArgumentative essay generation (AEG) is a complex task that requires advanced semantic understanding, logical reasoning, and organized integration of perspectives.Despite showing a promising performance, current efforts often overlook the dynamical and hierarchical nature of structural argumentative planning, and struggle with flexible rhetorical expression, leading to limited argument divergence and rhetorical optimization.Inspired by human debate behavior and Bitzer's rhetorical situation theory, we propose a debate-driven rhetorical framework for argumentative writing.The uniqueness lies in three aspects: (1) it dynamically assesses the divergence of viewpoints and progressively reveals the hierarchical outline of arguments based on a depththen-breadth paradigm, improving the perspective divergence within argumentation; (2) simulates human debate through iterative defenderattacker interactions, improving the logical coherence of arguments; (3) incorporates Bitzer's rhetorical situation theory to flexibly select appropriate rhetorical techniques, enabling the rhetorical expression.Experiments on four benchmarks validate that our approach significantly improves logical depth, argumentative diversity, and rhetorical persuasiveness over existing state-of-the-art models 1 .* Corresponding authors. 1 Code and data are available at https://github.com/ zxg-x/DARE Social media affects attention and distracts people.Social media affects attention and distracts people.Social media affects attention and distracts people.Social media is a waste of time.Social media is a waste of time.Social media is a waste of time. Xueguan Zhao, Wenpeng Lu, Chaoqun Zheng, Weiyu Zhang 0001, Jiasheng Si |
EMNLP | 3 |
| 2025 | Contrastive Learning-Based Standard-Free Calibration Transfer for Near-Infrared SpectroscopyabstractThis paper proposes a standard-free calibration transfer method for near-infrared (NIR) spectroscopy based on contrastive learning, aiming to address the challenge of calibration transfer caused by distribution shifts in spectral data across different instruments. We adapt and modify the original Contrastive Unpaired Translation (CUT) framework, initially designed for image domains, to accommodate one-dimensional spectral data, leveraging its strengths in one-way transfer and unsupervised learning. Experiments are conducted on a cross-device tobacco NIR dataset collected from two instruments. Quantitative results demonstrate that the proposed method achieves superior performance across multiple evaluation metrics, including Pearson correlation coefficient (PCC), cosine similarity (CS), mean absolute percentage error (MAPE) and root mean square error (RMSE). The proposed method outperforms existing standard-free methods such as Multiplicative Signal Correction (MSC) and Finite Impulse Response (FIR) filtering. Moreover, its performance, when compared to these methods, is closer to that of the existing standard calibration transfer method, Direct Standardization (DS). The method requires only standard-free data, significantly reducing the cost and complexity of calibration transfer in industrial applications where standard samples are scarce. These results highlight the potential of contrastive learning as a scalable and effective solution for real-world calibration transfer tasks. Chaoting Li, Chaoqun Zheng, Xinyao Lu, Guohao Zong, Weihua Feng, Sanying Feng |
IECON | 3 |
| 2025 | Research on Bearing Fault Diagnosis Based on IWOA-CNNLSTMabstractAs a key component of industrial equipment, bearings are prone to failure under complex operating conditions. Bearing fault diagnosis can detect potential dangers at an early stage, thus ensuring the stability and efficiency of production. Existing research mainly focuses on the structural improvement of the fault classification model and often neglects the optimization of the algorithm parameters. To address this deficiency, this paper proposed a hybrid convolutional neural network-long short-term memory (CNN-LSTM) model for bearing fault diagnosis. The model parameters were optimized using the improved whale optimization algorithm (IWOA). Experimental results show that the proposed method has superior performance in bearing fault diagnosis and has broad application prospects. Chaoqun Zheng, Weihua Feng, Guohao Zong, Chenhao Cui |
INDIN | 1 |
| 2025 | Toward Accurate Federated Graph Learning Via Layer-Wised Clustering for Social Internet of Thingsabstractfederated graph learning (FGL) has emerged as a promising paradigm for privacy-preserving collaborative learning in Social Internet of Things (SIoT), where nodes form complex interconnected networks. Existing FGL approaches face significant challenges including model degradation in handling nonindependent and identically distributed (non-IID) data and maintaining model performance across heterogeneous nodes. This article proposes framework via layer-wised clustering (FedLWC), a novel layer-wised clustering framework inspired by evolutionary processes is proposed to enhance the effectiveness of FGL. FedLWC designs three key aspects: 1) a fisher information matrix-based layer selection mechanism that identifies and evaluates critical model layers, which can reduce parameter redundancy; 2) a layer intersection clustering algorithm that preserves common key layers while accommodating local features; and 3) an adaptive layer merge strategy that effectively combines global shared layers with clustered key layers. To make sure that the proposed approach is rigorous, we conduct theoretical convergence analysis for the proposed framework under non-IID conditions. Extensive experiments on multiple benchmark graph datasets demonstrate FedLWC’s performance, achieving an average accuracy improvement of 7.01% compared to state-of-the-art federated learning methods. Yuru Liu, Yuange Liu, Weishan Zhang, Qiao Qiao, Daobin Luo, Chaoqun Zheng, Shaohua Cao, Lingzhao Meng, Tao Chen 0023 |
IEEE Internet Things J. | 7 |
| 2025 | Symbiosis Rather Than Aggregation: Toward Generalized Federated Learning via Model SymbiosisabstractFederated learning (FL) faces significant challenges in scenarios with nonindependent and identically distributed (non-IID) data distributions across participating clients. Traditional aggregation-based approaches often struggle with the inherent misalignment between local and global optimization objectives, which leads to gradient divergence and suboptimal generalization performance. This article proposes a novel FL framework that replaces conventional aggregation with a biologically inspired model symbiosis approach called FedSym, which employs a dual-level symbiotic mechanism. Ectosymbiosis performs coarse-grained hierarchical parameter recombinations through random layer-wise model combination, while endosymbiosis enables fine-grained intralayer parameter fusion through weighted averaging, collectively steering model updates toward flatter loss landscapes. Our theoretical analysis demonstrates that FedSym’s convergence rate is$O({}{1}/{T})$under non-IID conditions, which matches the convergence properties of FedAvg. Extensive evaluations across multiple datasets and model architectures show that FedSym achieves substantial improvements over state-of-the-art FL methods, particularly in challenging scenarios with high data heterogeneity, and demonstrates robust performance across varying numbers of participating clients and federation scales. Yuange Liu, Yuru Liu, Weishan Zhang, Chaoqun Zheng, Daobin Luo, Qiao Qiao, Lingzhao Meng, Su Yang 0001 |
IEEE Internet Things J. | 4 |
| 2025 | Multi-View Gait Recognition With Joint Local Multi-Scale and Global Contextual Spatio-Temporal FeaturesabstractExisting gait recognition methods are capable of extracting rich spatial gait information but often overlook fine-grained temporal features within local regions and temporal contextual information across different sub-regions. Considering gait recognition as a fine-grained recognition task and each individual exhibits uniqueness in their movements across different temporal sequences, we propose a local multi-scale and global contextual spatio-temporal (LMGCS) network for gait recognition. It divides the whole gait sequence into sub-sequences with multiple spatio resolutions and extracts multi-scale temporal features. We extract the temporal context information of different sub-sequences with the transformer, and all sub-sequences are fused to form global features. Furthermore, the loss function that combines the triplet loss function and cross-entropy loss function is utilized to prompt the proposed model to fulfill the gait recognition. The proposed method achieved state-of-the-art results on two popular public datasets. It achieved rank-1 accuracy of 98.0%, 95.4%, and 85.0% on the three walk states of the CASIA-B dataset and 90.9% on the OU-MVLP dataset. Wenzhe Zhai, Haomiao Li, Chaoqun Zheng, Xianglei Xing |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Compact-Yet-Separate: Proto-Centric Multi-Modal Hashing With Pronounced Category Differences for Multi-Modal RetrievalabstractMulti-modal hashing achieves low storage costs and high retrieval speeds by using compact hash codes to represent complex and heterogeneous multi-modal data, effectively addressing the inefficiency and resource intensiveness challenges faced by the traditional multi-modal retrieval methods. However, balancing intraclass compactness and interclass separability remains a struggle in existing works due to coarse-grained feature limitations, simplified fusion strategies that overlook semantic complementarity, and neglect of the structural information within the multi-modal data. To address these limitations comprehensively, we propose a Proto-centric Multi-modal Hashing with Pronounced Category Differences (PMH-PCD) model. Specifically, PMH-PCD first learns modality-specific prototypes by deeply exploring within-modality class information, ensuring effective fusion of each modality's unique characteristics. Furthermore, it learns multi-modal integrated class prototypes that seamlessly incorporate semantic information across modalities to effectively capture and represent the intricate relationships and complementary semantic content embedded within the multi-modal data. Additionally, to generate more discriminative and representative binary hash codes, PMH-PCD integrates multifaceted semantic information, encompassing both low-level pairwise relations and high-level structural patterns, holistically capturing intricate data details and leveraging underlying structures. The experimental results demonstrate that, compared with existing advanced methods, PMH-PCD achieves superior and consistent performances in multi-modal retrieval tasks. To promote further research and reproducibility, we have publicly released the source code of PMH-PCD at https://github.com/vindahi/PMH-PCD. Ruifan Zuo, Chaoqun Zheng, Lei Zhu 0002, Wenpeng Lu, Jiasheng Si, Weiyu Zhang 0001 |
IEEE Trans. Multim. | 2 |
| 2025 | EPFL: Toward Elastic Personalized Federated Learning With Seamless Client Joining and Quitting
Yuange Liu, Daobin Luo, Weishan Zhang, Chaoqun Zheng, Yuru Liu, Qiao Qiao, Tao Chen 0023, Su Yang 0001, Fei-Yue Wang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 5 |
| 2024 | Semantic Reconstruction Guided Missing Cross-modal HashingabstractCross-modal hashing is favored in the field of large-scale cross-modal retrieval for its fast query and low storage cost. Existing cross-modal hashing methods are typically based on an idealized assumption that data from different modalities are fully and completely paired. However, in practical scenarios, the occurrence of missing modal data is common due to technical difficulties, hardware failures, and challenges in data collection, contradicting the aforementioned assumption. In this paper, we propose an innovative unsupervised cross-modal hashing frame-work, named Semantic Reconstruction Guided Missing Hashing (SRGMH). Specifically, we utilize Dual-Variational Autoencoders (D-VAEs) to map features of different modalities into a shared low-dimensional latent representation, better bridging the gaps between modalities. Notably, we generate pseudo-representations of other modalities corresponding to the missing modality, effectively solving the problem of missing modal data. Moreover, we construct a refined adjacency similarity matrix to align the latent representations of different modalities, enhancing the quality of completed data and ensuring semantic consistency across modalities. Finally, we embed the adjacency similarity matrices in a shared latent representation space to jointly learn hash functions and hash codes, which enriches the semantic content and maintains structural similarity between hash functions and codes. Experiments on three real-world datasets demonstrate the superior performance of the proposed method on both retrieval accuracy and efficiency. Yafang Li, Chaoqun Zheng, Ruifan Zuo, Wenpeng Lu |
IJCNN | 2 |
| 2024 | LCEMH: Label Correlation Enhanced Multi-modal Hashing for efficient multi-modal retrieval
Chaoqun Zheng, Lei Zhu 0002, Zheng Zhang 0006, Wenjun Duan, Wenpeng Lu |
Inf. Sci. | 1 |
| 2024 | Multi-Modal Hashing for Efficient Multimedia Retrieval: A SurveyabstractWith the explosive growth of multimedia contents, multimedia retrieval is facing unprecedented challenges on both storage cost and retrieval speed. Hashing technique can project the high-dimensional data into compact binary hash codes. With it, the most time-consuming semantic similarity computation during the multimedia retrieval process can be significantly accelerated with fast Hamming distance computation, and meanwhile the storage cost can be reduced greatly by the binary embedding. In the light of this, multi-modal hashing has recently received considerable attention to support large-scale multimedia retrieval. Different from uni-modal hashing, the multi-modal hashing focuses on modeling the multi-modal semantics and further preserving them into binary hash codes with hash learning. In this paper, we first systematically review the existing learning to hash methods for efficient multimedia retrieval, categorizing them according to the multimedia retrieval tasks, the specific multi-modal semantic modeling techniques, and hash learning strategies. Thereafter, we present the performance comparison results. We ultimately discuss the challenges and potential research directions that may require further investigation in multi-modal hash learning. Lei Zhu 0002, Chaoqun Zheng, Weili Guan, Jingjing Li 0001, Yang Yang 0002, Heng Tao Shen |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2023 | Intention-Aware User Modeling for Personalized News Recommendation
Rongyao Wang, Shoujin Wang, Wenpeng Lu, Xueping Peng, Weiyu Zhang 0001, Chaoqun Zheng, Xinxiao Qiao |
DASFAA (2) | 6 |
| 2023 | News Recommendation via Jointly Modeling Event Matching and Style Matching
Shoujin Wang, Wenpeng Lu, Xueping Peng, Weiyu Zhang 0001, Chaoqun Zheng, Yonggang Huang 0001 |
ECML/PKDD (4) | 6 |
| 2023 | One for more: Structured Multi-Modal Hashing for multiple multimedia retrieval tasks
Chaoqun Zheng, Fengling Li 0001, Lei Zhu 0002, Zheng Zhang 0006, Wenpeng Lu |
Expert Syst. Appl. | 1 |
| 2022 | Charger and receiver deployment with delay constraint in mobile wireless rechargeable sensor networks
Haiqing Yao, Chaoqun Zheng, Xiuwen Fu, Ioan Ungurean |
Ad Hoc Networks | 2 |
| 2022 | Efficient Semi-Supervised Multimodal Hashing With Importance Differentiation RegressionabstractMulti-modal hashing learns compact binary hash codes by collaborating heterogeneous multi-modal features at both the model training and online retrieval stages to support large-scale multimedia retrieval. Previous multi-modal hashing methods mainly focus on supervised and unsupervised hashing. The performance of supervised hashing largely relies on the number of labeled data, which is practically expensive to obtain. Unsupervised hashing methods cannot effectively capture the semantic correlations of multi-modal data without any labels for supervision. In this paper, we propose an Efficient Semi-supervised Multi-modal Hashing with Importance Differentiation Regression (ESMH-IDR) model, which can alleviate the existing problems by learning from both labeled and unlabeled data. Specifically, in this paper, we develop an efficient semi-supervised multi-modal hash code learning module. It learns the hash codes for labeled data in an efficient asymmetric way, and simultaneously performs nonlinear regression using the same projection matrix as the labeled samples to preserve the intrinsic data structure of unlabeled data. Besides, different from existing methods, we propose an importance differentiation regression strategy to learn hash functions by specially considering the different importance of hash codes learned from the labeled and unlabeled samples. Finally, we develop an efficient discrete optimization method guaranteed with convergence to iteratively solve the hash optimization problem. Experiments on several public multimedia retrieval datasets demonstrate the superiority of our proposed method on both retrieval effectiveness and efficiency. Our source codes and testing datasets can be obtained at https://github.com/ChaoqunZheng/ESMH. Chaoqun Zheng, Lei Zhu 0002, Zheng Zhang 0006, Jingjing Li 0001, Xiaomei Yu |
IEEE Trans. Image Process. | 1 |
| 2022 | Efficient Multi-modal Hashing with Online Query Adaption for Multimedia RetrievalabstractMulti-modal hashing supports efficient multimedia retrieval well. However, existing methods still suffer from two problems: (1) Fixed multi-modal fusion. They collaborate the multi-modal features with fixed weights for hash learning, which cannot adaptively capture the variations of online streaming multimedia contents. (2) Binary optimization challenge. To generate binary hash codes, existing methods adopt either two-step relaxed optimization that causes significant quantization errors or direct discrete optimization that consumes considerable computation and storage cost. To address these problems, we first propose a Supervised Multi-modal Hashing with Online Query-adaption method. A self-weighted fusion strategy is designed to adaptively preserve the multi-modal features into hash codes by exploiting their complementarity. Besides, the hash codes are efficiently learned with the supervision of pair-wise semantic labels to enhance their discriminative capability while avoiding the challenging symmetric similarity matrix factorization. Further, we propose an efficient Unsupervised Multi-modal Hashing with Online Query-adaption method with an adaptive multi-modal quantization strategy. The hash codes are directly learned without the reliance on the specific objective formulations. Finally, in both methods, we design a parameter-free online hashing module to adaptively capture query variations at the online retrieval stage. Experiments validate the superiority of our proposed methods. Lei Zhu 0002, Chaoqun Zheng, Xu Lu 0004, Zhiyong Cheng 0001, Liqiang Nie, Huaxiang Zhang 0001 |
ACM Trans. Inf. Syst. | 2 |
| 2021 | Adaptive Partial Multi-View Hashing for Efficient Social Image RetrievalabstractSocial networks allow users to actively upload images and descriptive tags, which has led to an explosive growth in the number of social images. Multi-view hashing is an efficient technique for supporting large-scale social image retrieval because of its desirable capabilities of encoding multi-view features into compact binary hash codes with extremely low storage costs and fast retrieval speeds. However, existing methods require multi-view features to be fully paired at both the offline model training and online query stages. This requirement cannot be easily satisfied for social image retrieval, where social images that lack descriptive tags are common in social networks. In this paper, we propose anUnsupervised Adaptive Partial Multi-view Hashing(UAPMH) method to handle the partial-view hashing problem for efficient social image retrieval. Specifically, the shared and view-specific latent representations of fully paired and partial-view images, respectively, are learned separately by an adaptive partial multi-view matrix factorization module within the identical semantic space. In particular, instead of adopting simple fixed view combination weights, we develop a parameter-free weight learning scheme to adaptively learn the weights to capture the view variations and the discriminative capabilities of different views. With such a design, our model can sufficiently exploit the available partial-view samples with separate hash code learning and effectively preserve the latent relations of images and tags in hash codes with semantic space sharing. Moreover, to avoid relaxing errors and improve the learning efficiency, binary hash codes are directly learned in a fast mode with simple and efficient operations. Finally, we extend UAPMH to the supervised learning paradigm asSupervised Adaptive Partial Multi-view Hashing(SAPMH) with the supervision of pair-wise semantic labels to further enhance the discriminative capability of hash codes. The experiments demonstrate the state-of-the-art performance of the proposed approaches on public social image retrieval datasets. Our source codes and testing datasets can be obtained athttps://github.com/ChaoqunZheng/APMH. Chaoqun Zheng, Lei Zhu 0002, Zhiyong Cheng 0001, Jingjing Li 0001, Anan Liu |
IEEE Trans. Multim. | 1 |
| 2020 | Efficient Parameter-Free Adaptive Multi-Modal HashingabstractUnsupervised multi-modal hashing has recently attracted broad attention in research area of large-scale multimedia retrieval for its low storage cost, high retrieval speed, and independence on semantic labels. However, the model learning process of existing methods still suffer from the problem of low efficiency: 1) Many existing methods measure the contributions of different modalities using fixed modality weights. In order to avoid over-fitting, they need an inefficient hyper-parameter adjustment process. 2) Most existing methods adopt inefficient optimization strategies to learn hash codes. In this letter, we propose an unsupervised Efficient Parameter-free Adaptive Multi-modal Hashing (EPAMH) model to adaptively capture the modality variations and preserve the discriminative semantics of multi-modal features into the binary hash codes. Moreover, we directly learn the binary codes with simple and efficient operations, which prevents the relaxing quantization errors and improves the model learning efficiency. Experiments prove the superior performance of EPAMH on three public multimedia retrieval datasets. Our source codes and testing datasets can be obtained at https://github.com/ChaoqunZheng/EPAMH. Chaoqun Zheng, Lei Zhu 0002, Shusen Zhang, Huaxiang Zhang 0001 |
IEEE Signal Process. Lett. | 1 |
| 2020 | Fast Discrete Collaborative Multi-Modal Hashing for Large-Scale Multimedia RetrievalabstractMany achievements have been made on learning to hash for uni-modal and cross-modal retrieval. However, it is still an unsolved problem that how to directly and efficiently learn discriminative discrete hash codes for the multimedia retrieval, where both query and database samples are represented with heterogeneous multi-modal features. With this motivation, we propose a Fast Discrete Collaborative Multi-modal Hashing (FDCMH) method in this paper. We first propose an efficient collaborative multi-modal mapping that first transforms heterogeneous multi-modal features into the unified factors to exploit the complementarity of multi-modal features and preserve the semantic correlations in multiple modalities with linear computation and space complexity. Such shared factors also bridge the heterogeneous modality gap and remove the inter-modality redundancy. Further, we develop an asymmetric hashing learning module to simultaneously correlate the learned hash codes with low-level data distribution and high-level semantics. In particular, this design could avoid the challenging symmetric semantic matrix factorization and O(n2) memory cost (n is the number of training samples). It can support both computation and memory efficient discrete hash optimization. Experiments on several public multimedia retrieval datasets demonstrate the superiority of the proposed approach compared with state-of-the-art hashing techniques, in terms of both model learning efficiency and retrieval accuracy. Chaoqun Zheng, Lei Zhu 0002, Xu Lu 0004, Jingjing Li 0001, Zhiyong Cheng 0001, Hanwang Zhang |
IEEE Trans. Knowl. Data Eng. | 1 |