EDBT 2026 Demo / reviewers in the wild / expert
Yuan Cao 0005
dblp:52/4472-5
· DBLP profile ↗
33ranked-venue papers
14as first author
31since 2021 · last 2026
0000-0002-1445-8210ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 19 · 7 first-author · 18 since 2021Artificial intelligence and machine learning · 13 · 2 first-author · 12 since 2021Systems, architecture and hardware · 4 · 4 first-author · 4 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Security and privacy · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multiplex Heterogeneous Graph Neural Networks with Euclidean-Riemannian Mutual Space SynergyabstractMultiplex heterogeneous networks are common in real-world scenarios, where entities interact through diverse types of relations across multiple semantic layers. Recent advances in multiplex heterogeneous graph neural networks have achieved remarkable results by incorporating node and relation types into message passing and designing relation-aware architectures. However, most existing methods either decouple relations and risk losing complex semantics or require handcrafted relation patterns, which limit scalability. Moreover, prevailing models are typically restricted to Euclidean space, making it difficult to capture non-Euclidean topologies and to distinguish complex interactions among heterogeneous nodes and relations. Standard GNN message passing, grounded in the homophily assumption, also proves inadequate for the intricate, coupled structures in multiplex heterogeneous graphs. To address these challenges, we propose MRiemGNN, a novel multiplex heterogeneous graph neural network that synergizes Euclidean and Riemannian spaces through a geometry-aware, relation-specific message passing scheme and cross-space mutual learning. Experiments on multiple real-world datasets show that MRiemGNN achieves superior performance, efficiency, and scalability on both node classification and link prediction tasks. Xiang Li 0111, Yuan Cao 0005, Zhongying Zhao 0001, Guoqing Chao, Yanwei Yu |
AAAI | 2 |
| 2026 | Automatic Channel Pruning by Searching with Structure Embedding for Hash NetworkabstractDeep hash networks are widely used in tasks such as large-scale image retrieval due to high search efficiency and low storage costs through binary hash codes. With the growing demand for deploying deep hash networks on resource-constrained devices, it is crucial to perform network compression on them, in which automatic pruning constitutes a priority option owing to efficacy maintenance. However, existing pruning methods are mostly designed for image classification, while hashing networks must generate compact binary codes, making each channel more sensitive to retrieval objectives. As a result, their performance often degrades when applied to image retrieval tasks. In this paper, we propose a novel Automatic Channel Pruning framework by Searching with Structure Embedding (ACP-SSE). To the best of our knowledge, this is the first study to explore pruning techniques for deep hash networks and the first automatic pruning method by searching based on network topology structure. Specifically, we first design a structure encoding model by Graph Convolutional Networks (GCNs) whose graph is constructed by hash network and nodes' features are initialized by pruning strategies. The model is trained by contrastive learning loss efficiently without accuracy supervision by fine-tuning pruned models. In addition, we introduce a dynamic pruning search space in consideration of the resource constraints. By converting the automatic channel pruning task into searching the pruned structure with effect similar to the unpruned structure, it enables the method to adapt to various network architectures. Finally, the optimal networks are selected from the candidate set according to their performance in specific downstream tasks. Extensive experiments demonstrate that ACP-SSE indeed works in the automatic channel pruning area, outperforming state-of-the-art baselines in hashing-based image retrieval, while maintaining competitive accuracy in image classification. Zifan Liu, Yuan Cao 0005, Yanwei Yu, Heng Qi |
AAAI | 2 |
| 2026 | Self-Supervised Cross-City Trajectory Representation Learning Based on Meta-LearningabstractTrajectory representation learning transforms complex spatio-temporal features of trajectories into dense, low-dimensional embeddings, enabling applications in intelligent transportation systems. With advances in this field and the availability of large-scale traffic data, intelligent urban systems have been widely deployed in major cities. However, existing methods heavily rely on large volumes of trajectory data, limiting their transferability to cities with sparse data, especially small or less-developed ones. Moreover, most current approaches learn representations within a single city, overlooking the shared travel patterns across regions and cities with similar geographic contexts. To address these issues, we propose MetaTRL, a self-supervised cross-city trajectory representation learning method based on meta-learning. Specifically, we introduce a Shared and Private Parameterized Cross-city Meta-learning Framework to support knowledge sharing and transfer across cities. We further design a Meta-knowledge Enhanced Road Segment Encoder and a Trajectory Encoder that integrates private and shared knowledge to learn and fuse spatio-temporal trajectory features. Extensive experiments on two real-world datasets and multiple downstream tasks demonstrate the significant superiority of MetaTRL over state-of-the-art baselines and achieves a remarkable average improvement of 134.66% in Macro-F1 on destination prediction task. Yanwei Yu, Hong Xia, Shaoxuan Gu, Xingyu Zhao 0006, Yuan Cao 0005 |
AAAI | 6 |
| 2026 | Proxy Zero-Shot Hashing with Multimodal Fusion via Stable DiffusionabstractWith the rapid growth of visual content in open-world environments, zero-shot hashing image retrieval (ZSHIR) has emerged to tackle the challenge of recognizing novel classes using attribute-level and semantic information. However, existing methods often rely on shallow fusion of multi-source cues (e.g., attributes, labels, and visual features) through external supervision or feature concatenation, failing to capture the underlying semantic structure in a generative way. Particularly, current bridging strategies between modalities suffer from information fragmentation and weak alignment, hindering the model's ability to fully understand complex attribute-visual relations. Moreover, subtle semantic gaps or “semantic drift” between seen and unseen classes further degrade inter-class separability and the scalability of hashing models. To address these issues, we propose a novel framework called Proxy Zero-Shot Hashing with Multimodal Fusion via Stable Diffusion (PZSH), which integrates generative modeling and contrastive learning. PZSH leverages a pre-trained Stable Diffusion (SD) model to synthesize multimodal content, and uses dual BLIP encoders to enhance semantic alignment across modalities. We further design a proxy hashing loss to enforce discriminative binary representations. Extensive experiments on benchmark datasets show that PZSH achieves state-of-the-art performance with stronger generalization to unseen classes. Weikang Gao, Yuan Cao 0005 |
AAAI | 4 |
| 2026 | TrajAgg: Dual-Scale Feature Aggregation with Hybrid Training for Trajectory Similarity Computation in Free SpaceabstractWith the widespread use of location-tracking technologies, large volumes of trajectory data are continuously generated. Trajectory similarity computation is a core task in trajectory mining with broad applications. However, existing methods still face two key challenges: (1) the difficulty of balancing efficiency and representation quality, and (2) the reliance on a single training paradigm, which limits the ability to capture both pairwise similarity and batch-level coherence. To address these challenges, we propose a trajectory similarity computation framework named TrajAgg. Specifically, our framework incorporates a novel Aggregation Transformer that efficiently aggregates GPS and grid features through two stages of direct interaction and enhances the expressiveness of the resulting trajectory embeddings. In addition, by integrating two distinct training paradigms, our model captures both fine-grained pairwise relationships and global structural consistency. We further analyze its effectiveness from the perspective of mutual information. Extensive experiments on three publicly available datasets show that TrajAgg consistently outperforms state-of-the-art baselines. Our method achieves average improvements of 15.11%, 16.49%, 10.41%, and 40.15% in HR@1 under four distance measures across three datasets, respectively. Xingyu Zhao 0006, Yuan Cao 0005, Bin Wang 0045, Guiyuan Jiang, Yanwei Yu |
AAAI | 3 |
| 2026 | ScaleGNN: Towards Scalable Graph Neural Networks via Adaptive High-order Neighboring Feature FusionabstractGraph Neural Networks (GNNs) have demonstrated impressive performance across diverse graph-based tasks by leveraging message passing to capture complex node relationships. However, on large-scale real-world graphs, GNNs face two major challenges: (1) GNNs struggle to ensure scalability and efficiency as repeated aggregation of large neighborhoods incurs significant computational overhead; (2) GNNs suffer from over-smoothing, where excessive propagation makes node representations indistinguishable, hindering model expressiveness. To tackle these, we propose ScaleGNN, which adaptively fuses multi-hop node features for scalable and effective graph learning. We first compute per-hop pure-neighbor matrices to isolate exclusive structural signals, then apply lightweight fusion to balance low- and high-order information, preserving both local detail and global correlations. To curb redundancy and over-smoothing, we introduce Local Contribution Score (LCS)–based masking to prune low-relevance high-order neighbors, and impose learnable sparsity to selectively integrate valuable multi-hop features. Extensive experiments on real-world datasets show that ScaleGNN consistently outperforms state-of-the-art GNNs in both predictive accuracy and computational efficiency. The source code is available at https://github.com/lx970414/ScaleGNN. Xiang Li 0111, Jianpeng Qi, Haobing Liu 0001, Yuan Cao 0005, Guoqing Chao, Zhongying Zhao 0001, Junyu Dong, Xinwang Liu 0002, Yanwei Yu |
WWW | 4 |
| 2026 | Object-guided multi-granularity unsupervised hashing for image retrieval
Zifan Liu, Yuan Cao 0005, Peng Luan, Yanwei Yu |
Neural Networks | 2 |
| 2026 | Trajectory Similarity Hash Learning With Spatio-Temporal GRUabstractTrajectory similarity computation plays a critical role in a wide range of trajectory-related applications, including transportation optimization and behavior study. Most studies aim at learning discriminative real-valued trajectory representations. However, these methods struggle to scale to large datasets due to their linear time complexity. To address this problem, only one hypergraph hash learning approach (HHL-Traj) has been proposed to realize efficient trajectory similarity computation by calculating Hamming distances among the trajectory hash codes. Nevertheless, it fails to effectively integrate the spatial and temporal information of trajectory data with semantic relevance. In this paper, we present a novel Trajectory Similarity Hash Learning method with Spatio-Temporal GRU (TrajH-ST), which fuses the spatial and temporal features through reset and update gates within the network architecture. Additionally, we design an alternating sampling strategy to generate two sub-trajectories for contrastive learning, which enhances both generalization and robustness compared to nonuniform sampling techniques. To optimize the proposed end-to-end model, we develop an objective function that incorporates InfoNCE loss, alignment loss, and quantization loss. Extensive experiments on two widely-used trajectory datasets demonstrate that the proposed model consistently outperforms state-of-the-art baselines, achieving accuracy improvements of up to 5.58% and 4.33% with real-valued and binary features, respectively. Our code is available athttps://github.com/caoyuan618/Traj-ST Yuan Cao 0005, Zifan Liu, Lei Li 0071, Bin Wang 0045, Yanwei Yu |
IEEE Trans. Big Data | 1 |
| 2026 | Long-Tailed Approaching Cross-Modal Hashing With Multi-Expert Collaborative LearningabstractCross-modal hashing enables efficient retrieval across different modalities by mapping heterogeneous data into compact binary codes within a shared Hamming space. However, most existing methods assume that data from each class are evenly distributed, which contradicts the long-tailed nature of real-world data. Consequently, these approaches often exhibit suboptimal performance when handling imbalanced datasets. The only existing cross-modal hashing method that considers long-tailed data attempts to mine both the individuality and commonality across modalities, yet it relies on a negative log-likelihood pairwise loss that tends to bias the model toward head categories. To address this issue, we propose a novel Long-tailed Approaching Cross-modal Hashing (LACH) framework based on multi-expert collaborative learning. Specifically, LACH constructs a multi-expert architecture with a Graph Convolutional Network (GCN) backbone to facilitate knowledge transfer. Unlike conventional multi-expert models that either share identical data copies or employ entirely distinct data subsets, we introduce a partial data replication strategy that ensures each expert receives a balanced yet overlapping training set. Furthermore, we design a proxy-based pointwise loss to treat head and tail categories equitably, along with an inter-modal approaching loss to enhance the alignment of hash codes across modalities within each class. Extensive experiments demonstrate that LACH achieves accuracy improvements of up to 4.2% and 6.3% over state-of-the-art baselines on balanced and long-tailed datasets, respectively. Our code is available at https://github.com/caoyuan618/LACH. Yuan Cao 0005, Zifan Liu, Weikang Gao, Jie Gui, Yanwei Yu |
IEEE Trans. Image Process. | 1 |
| 2025 | Deep Graph Online Hashing for Multi-Label Image RetrievalabstractOnline hashing has attracted much research attention for large-scale image retrieval in a streaming way. The main challenge lies in keeping balance between high retrieval accuracy and low training time. Existing online hashing methods almost rely on shallow models rather than deep networks due to high training costs, because it is unacceptable to update hash functions on an order of hours. In addition, the multi-label supervision information is not fully utilized to guide the hash learning process and the affinity matrix is always fixed once constructed. In this paper, we propose a novel Deep Graph Online Hashing (DGOH) method, which for the first time introduces inductive graph neural networks (GNNs) to realize deep online hashing with acceptable training costs on an order of seconds. Furthermore, we mine the multi-label information of the images by constructing a label network and learn label-wise weights dynamically to help to update the affinity matrix. In addition, we provide a strategy to obtain examples from the old data to solve the catastrophic forgetting problem. An integrated objective function is designed to train the entire architecture. Extensive experiments on two common benchmarks demonstrate that the proposed method achieves up to 13.3% accuracy gains over state-of-the-art baselines and shows competitive performance on training time. Yuan Cao 0005, Xiangru Chen 0001, Zifan Liu, Wenzhe Jia, Fanlei Meng, Jie Gui |
AAAI | 1 |
| 2025 | A Privacy-Preserving Cross-Modal Retrieval Scheme Based on CLIP and Deep HashingabstractWith the massive growth of multimedia data, local devices gradually cannot meet the data processing needs, thus utilizing cloud server resources becomes better choice. To prevent privacy leakage, user data can only be stored in ciphertext. Existing cross-media retrieval schemes are only for plaintext data and cannot protect privacy. Privacy-preserving cross-modal retrieval of cloud data becomes a research priority. In this paper, we propose a CLIP-based privacy-preserving cross-modal retrieval scheme. First, CLIP encoder is utilized to transform images and texts into public space features to enhance retrieval accuracy. Subsequently, the original feature relevance is reduced by feature transformation network and mutual information loss to enhance privacy preservation. Finally, a transformer encoder-based hash learning method is used to embed the transformed features into compact hash codes, which is combined with comparison learning to enhance the discriminative properties. Experimental results on cross-modal datasets verify the efficiency and accuracy of the proposed scheme. Yuan Cao 0005, Xinzheng Shang |
ICASSP | 1 |
| 2025 | Self-Supervised Trajectory Representation Learning with Multi-Scale Spatio-Temporal Feature ExplorationabstractTrajectory representation learning transforms the complex spatio-temporal features of trajectories into a dense, low-dimensional embedding, which supports various downstream analytics tasks such as trajectory classification, travel time estimation, and similar trajectory search. Existing trajectory representation learning methods treat trajectories merely as general point sequences and use sequence models to learn the correlations between points. However, the complex spatio-temporal features of trajectories are multi-scale, meaning they are not only reflected in the correlations between trajectory points but also in the correlations between trajectory segments. Moreover, most existing methods do not sufficiently capture the multi-faceted temporal features within trajectories. To fill these gaps, we propose a novel self-supervised Trajectory$R$epresentation$L$earning model with multi-scale spatio-temporal features exploration called TrajRL. Specifically, we utilize trajectory augmentation to generate two views to achieve self-supervised pre-training exploiting multiple self-supervisory signals. In each view, we can learn the multi-scale spatio-temporal correlations both within and between road segments and road segment sequences in trajectories through the proposed multi-scale trajectory encoder. Additionally, we perform multi-faceted temporal information encoding, especially leveraging time intervals to learn multi-scale context-aware time patterns within the trajectories. Extensive experiments demonstrate the superiority of our TrajRL as compared to state-of-the-art baselines on two real-world datasets across various downstream tasks. The source code of our model is available at https://github.com/Xfc30/TrajRL. Hong Xia, Yuan Cao 0005, Lei Cao 0004, Yanwei Yu, Junyu Dong |
ICDE | 3 |
| 2025 | Transformer Based Unsupervised Cross-Modal Hashing for Normal and Remote Sensing RetrievalabstractWith the rapid expansion of online information, cross-modal retrieval has emerged as a crucial and dynamic research focus. Deep hashing has gained significant traction in this field due to its efficiency in storage and retrieval speed, making it particularly valuable for remote sensing multi-modal retrieval. However, existing deep cross-modal hashing techniques often rely on parallel network structures for processing different modalities, overlooking a unified representation that captures cross-modal visual information. To address this limitation, we introduce a novel unsupervised cross-modal hashing framework that incorporates two modality-specific encoders and a fusion module. This fusion module facilitates modality interaction, enabling the extraction of meaningful semantic relationships across different data types. To ensure comprehensive similarity preservation, we design an integrated objective function that incorporates inter-modal and intra-modal constraints, joint consistency, and binary alignment losses. Furthermore, instead of conventional convolutional networks, we adopt the Swin Transformer as the backbone to enhance the discriminative power of image features. Our approach achieves an average 2.3% improvement in mAP on remote sensing cross-modal retrieval tasks compared to existing methods. The implementation is available at https://github.com/sellaner/TUCH. Weikang Gao, Zifan Liu, Yuan Cao 0005, Zuojin Huang, Yaru Gao |
IEEE Signal Process. Lett. | 3 |
| 2025 | A Privacy-Preserving Large-Scale Image Retrieval Framework With Vision GNN HashingabstractWith the growing popularity of cloud services, companies and individuals outsource images to cloud servers to reduce storage and computing burdens. The images are encrypted before outsourcing for privacy protection. It has become urgent to solve the privacy-preserving image retrieval problem on the cloud. There are three main challenges in this area. First, how can we achieve high retrieval accuracy on the encryption domain? Second, how can we improve efficiency in large-scale encrypted image retrieval? Third, how can we ensure the reliability of the retrieval results? The existing schemes only consider some of these characteristics and the retrieval accuracy is insufficient. In this paper, we propose a privacy-preserving large-scale image retrieval framework with vision graph convolutional neural network hashing (ViGH). To the best of our knowledge, this is the first framework that is able to address all the above challenges with more advanced accuracy performance. To be specific, cycle-consistent adversarial networks and vision graph convolutional networks (ViG) are utilized to increase retrieval accuracy. By embedding encrypted images into hash codes, we can obtain high retrieval efficiency by Hamming distances. Cloud servers store the hash codes on the blockchain (Ethereum). The retrieval algorithm on the smart contracts and the consensus mechanism of blockchain ensure reliability of the retrieval results. The experimental results on three common datasets verify the effectiveness and efficiency of the proposed privacy-preserving image retrieval framework. The reliability of the retrieval results is ensured by the consensus mechanism of blockchain with no need for verification. Yuan Cao 0005, Fanlei Meng, Xinzheng Shang, Jie Gui, Yuan Yan Tang |
IEEE Trans. Big Data | 1 |
| 2025 | Improving Fast Adversarial Training via Self-Knowledge GuidanceabstractAdversarial training has achieved remarkable advancements in defending against adversarial attacks. Among them, fast adversarial training (FAT) is gaining attention for its ability to achieve competitive robustness with fewer computing resources. Existing FAT methods typically employ a uniform strategy that optimizes all training data equally without considering the influence of different examples, which leads to an imbalanced optimization. However, this imbalance remains unexplored in the field of FAT. In this paper, we conduct a comprehensive study of the imbalance issue in FAT and observe an obvious class disparity regarding their performances. This disparity could be embodied from a perspective of alignment between clean and robust accuracy. Based on the analysis, we mainly attribute the observed misalignment and disparity to the imbalanced optimization in FAT, which motivates us to optimize different training data adaptively to enhance robustness. Specifically, we take disparity and misalignment into consideration. First, we introduce self-knowledge guided regularization, which assigns differentiated regularization weights to each class based on its training state, alleviating class disparity. Additionally, we propose self-knowledge guided label relaxation, which adjusts label relaxation according to the training accuracy, alleviating the misalignment and improving robustness. By combining these methods, we formulate the Self-Knowledge Guided FAT (SKG-FAT), leveraging naturally generated knowledge during training to enhance the adversarial robustness without compromising training efficiency. Extensive experiments on four standard datasets demonstrate that the SKG-FAT improves the robustness and preserves competitive clean accuracy, outperforming the state-of-the-art methods. Code and checkpoints are available at SFG-FAT Code Implementation. Chengze Jiang, Minjing Dong, Jie Gui, Xinli Shi, Yuan Cao 0005, Yuan Yan Tang, James T. Kwok |
IEEE Trans. Inf. Forensics Secur. | 6 |
| 2025 | Divide and Conquer: Frequency-Aware Contrastive Adversarial Training for Robust Point Cloud ClassificationabstractContrastive adversarial training has shown great potential in enhancing model robustness and has been adopted in point cloud classification. There are varying spatial distributions and densities across different regions in point cloud data, which makes adversarial perturbations always exhibit non-uniform patterns of attack intensity and distribution in different regions. However, existing approaches always rely on uniform feature contrast without considering the granularity in the context of point cloud data, limiting their capacities to counter adversarial perturbations effectively. To address this issue, we propose a novel frequency-aware contrastive adversarial training framework, which considers feature contrast via a “divide-and-conquer” method. Specifically, we systematically “divide” point clouds into distinct frequency components and “conquer” feature contrast within each frequency band, which fosters fine-grained feature consistency learning and leads to more informative as well as robust representations. Besides, existing methods typically apply group-level contrastive learning, which emphasizes category-wise similarity but often overlooks the nuanced structural variations among instances. To remedy this, we incorporate instance-level contrastive learning to capture per-instance geometric variations. Moreover, a frequency-specific hard-masked sample generation module is designed to construct challenging sample pairs by masking keypoint features in each frequency band, thereby promoting the model to learn more robust feature representations. Extensive experiments on multiple benchmark datasets demonstrate that our proposed method significantly outperforms existing state-of-the-art approaches in adversarial robustness for point cloud classification. The code is available on DiCon-FAT. Yu-Xin Zhang 0004, Jie Gui, Minjing Dong, Xiaofeng Cong, Yuan Cao 0005, Xin Gong 0001, Yuan Yan Tang, James T. Kwok |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2024 | Taxonomy Driven Fast Adversarial TrainingabstractAdversarial training (AT) is an effective defense method against gradient-based attacks to enhance the robustness of neural networks. Among them, single-step AT has emerged as a hotspot topic due to its simplicity and efficiency, requiring only one gradient propagation in generating adversarial examples. Nonetheless, the problem of catastrophic overfitting (CO) that causes training collapse remains poorly understood, and there exists a gap between the robust accuracy achieved through single- and multi-step AT. In this paper, we present a surprising finding that the taxonomy of adversarial examples reveals the truth of CO. Based on this conclusion, we propose taxonomy driven fast adversarial training (TDAT) which jointly optimizes learning objective, loss function, and initialization method, thereby can be regarded as a new paradigm of single-step AT. Compared with other fast AT methods, TDAT can boost the robustness of neural networks, alleviate the influence of misclassified examples, and prevent CO during the training process while requiring almost no additional computational and memory resources. Our method achieves robust accuracy improvement of 1.59%, 1.62%, 0.71%, and 1.26% on CIFAR-10, CIFAR-100, Tiny ImageNet, and ImageNet-100 datasets, when against projected gradient descent PGD10 attack with perturbation budget 8/255. Furthermore, our proposed method also achieves state-of-the-art robust accuracy against other attacks. Code is available at https://github.com/bookman233/TDAT. Kun Tong, Chengze Jiang, Jie Gui, Yuan Cao 0005 |
AAAI | 4 |
| 2024 | Hypergraph Hash Learning for Efficient Trajectory Similarity ComputationabstractTrajectory similarity computation is a fundamental problem in various applications (e.g., transportation optimization, behavioral study). Recent researches learn trajectory representations instead of point matching to realize more accurate and efficient trajectory similarity computation. However, these methods can still not be scaled to large datasets due to high computational cost. In this paper, we propose a novel hash learning method to encode the trajectories into binary hash codes and compute trajectory similarities by Hamming distances which is much more efficient. To the best of our knowledge, this is the first work to conduct hash learning for trajectory similarity computation. Furthermore, unlike the Word2Vec model based on random walk strategy, we utilize hypergraph neural networks for the first time to learn the representations for the grids by constructing the hyperedges according to the real-life trajectories, resulting in more representative grid embeddings. In addition, we design a residual network into the multi-layer GRU to learn more discriminative trajectory representations. The proposed Hypergraph Hash Learning for Trajectory similarity commutation is an end-to-end framework and named HHL-Traj. Experimental results on two real-world trajectory datasets (i.e., Porto and Beijing) demonstrate that the proposed framework achieves up to 6.23% and 15.42% accuracy gains compared with state-of-the-art baselines in unhashed and hashed cases, respectively. The efficiency of trajectory similarity computation based on hash codes is also verified. Our code is available at https://github.com/caoyuan57/HHL-Traj. Yuan Cao 0005, Lei Li 0071, Xiangru Chen 0001, Zuojin Huang, Yanwei Yu |
CIKM | 1 |
| 2024 | Self-supervised Adversarial Hashing for Large-scale Image RetrievalabstractHashing has been widely used in large-scale image retrieval due to its high storage and search efficiency. Unsupervised learning saves a significant amount of labor costs compared to supervised learning. Existing unsupervised methods mostly convert unsupervised problems into supervised problems by reconstructing semantic information. In this paper, we propose a novel Self-supervised Adversarial Hashing (SAH) method which utilizes unsupervised semantic reconstruction methods along with self-supervised generative adversarial and contrastive learning methods to improve the model’s accuracy and robustness. In addition, we propose a novel method for semantic relationship reconstruction, taking into account the similarity between different categories. Based on this, we have designed a multi-joint loss to further achieve reasonable intra-class aggregation and inter-class differentiation. The experimental results show that the proposed method outperforms state-of-the-art hashing methods for large-scale image retrieval. Yuan Cao 0005, Xiangru Chen 0001 |
ICPADS | 1 |
| 2024 | Targeted Universal Adversarial Attack on Deep Hash NetworksabstractDeep hash networks have garnered significant attention due to their efficiency and ability to learn discriminative embeddings for approximate nearest neighbor search. However, it is observed that deep hash networks are vulnerable to adversarial interference, which is an important security problem. Despite the growing interest in targeted attack on deep hash networks, it suffers from a scarcity of research on generating universal adversarial perturbations which are unrelated to the specific images. In this paper, we introduce a novel Targeted Universal adversarial Attack (TUA) on deep hash networks. Our framework consists of two key components: a ReferenceNet and a universal generative adversarial network. Specifically, ReferenceNet is designed to generate category-level representative reference codes for the target labels by introducing a cosine similarity based reference loss. Additionally, we feed the fixed random noise and target labels into the generator to learn universal adversarial perturbations. Particularly, the reference codes are used to optimize the generator by minimizing the Hamming distances between the hash codes of the adversarial examples and the reference codes. Extensive experiments on three common datasets validate the superior targeted attack performance, transferability, and universality of our method compared with state-of-the-art targeted attack methods on deep hash networks. Fanlei Meng, Xiangru Chen 0001, Yuan Cao 0005 |
ICMR | 3 |
| 2023 | Fast Online Hashing with Multi-Label ProjectionabstractHashing has been widely researched to solve the large-scale approximate nearest neighbor search problem owing to its time and storage superiority. In recent years, a number of online hashing methods have emerged, which can update the hash functions to adapt to the new stream data and realize dynamic retrieval. However, existing online hashing methods are required to update the whole database with the latest hash functions when a query arrives, which leads to low retrieval efficiency with the continuous increase of the stream data. On the other hand, these methods ignore the supervision relationship among the examples, especially in the multi-label case. In this paper, we propose a novel Fast Online Hashing (FOH) method which only updates the binary codes of a small part of the database. To be specific, we first build a query pool in which the nearest neighbors of each central point are recorded. When a new query arrives, only the binary codes of the corresponding potential neighbors are updated. In addition, we create a similarity matrix which takes the multi-label supervision information into account and bring in the multi-label projection loss to further preserve the similarity among the multi-label data. The experimental results on two common benchmarks show that the proposed FOH can achieve dramatic superiority on query time up to 6.28 seconds less than state-of-the-art baselines with competitive retrieval accuracy. Wenzhe Jia, Yuan Cao 0005, Jie Gui |
AAAI | 2 |
| 2023 | Graph Structure Learning on User Mobility Data for Social Relationship InferenceabstractWith the prevalence of smart mobile devices and location-based services, uncovering social relationships from human mobility data is of great value in real-world spatio-temporal applications ranging from friend recommendation, advertisement targeting to transportation scheduling. While a handful of sophisticated graph embedding techniques are developed for social relationship inference, they are significantly limited to the sparse and noisy nature of user mobility data, as they all ignore the essential problem of the existence of a large amount of noisy data unrelated to social activities in such mobility data. In this work, we present Social Relationship Inference Network (SRINet), a novel Graph Neural Network (GNN) framework, to improve inference performance by learning to remove noisy data. Specifically, we first construct a multiplex user meeting graph to model the spatial-temporal interactions among users in different semantic contexts. Our proposed SRINet tactfully combines the representation learning ability of Graph Convolutional Networks (GCNs) with the power of removing noisy edges of graph structure learning, which can learn effective user embeddings on the multiplex user meeting graph in a semi-supervised manner. Extensive experiments on three real-world datasets demonstrate the superiority of SRINet against state-of-the-art techniques in inferring social relationships from user mobility data. The source code of our method is available at https://github.com/qinguangming1999/SRINet. Guangming Qin, Lexue Song, Yanwei Yu, Chao Huang 0001, Wenzhe Jia, Yuan Cao 0005, Junyu Dong |
AAAI | 6 |
| 2023 | Generative Adversarial Network Based Asymmetric Deep Cross-Modal Unsupervised Hashing
Yuan Cao 0005, Yaru Gao, Jiacheng Lin, Sheng Chen 0015 |
ICA3PP (1) | 1 |
| 2023 | Deep Hash Learning of Feature-Invariant Representation for Single-Label and Multi-label Retrieval
Yuan Cao 0005, Xinzheng Shang, Chengzhi Qian, Sheng Chen 0015 |
ICA3PP (1) | 1 |
| 2023 | Generative Enhancement-based Similarity Prediction Hashing for Image RetrievalabstractHashing is frequently used in approximate nearest neighbor search due to its storage and search efficiency. On account of the bottleneck of traditional learning to hash methods, deep-based learning to hash has gained quite a popularity among researchers recently. While such methods show a promising performance gain by utilizing deep neural networks into its end-to-end training process to generate compact binary codes, the intrinsic connections between components make it unfeasible to optimize the architecture significantly. Subject to noise interference and absence of similarity labels of training data, normal unsupervised deep models even carry a noticeable deviation at representation learning stage. By integrating Generative Adversarial Network, this paper presents a novel architecture for generating compact hash codes from an extended set derived from original images while using the benefit of similarity prediction. Extensive experiments conducted on three benchmarks, NUS-WIDE, CIFAR-10, and MS-COCO illustrate that our method outperforms state-of-the-art image retrieval models and generates high-quality binary hash codes seamlessly. Yuan Cao 0005, Fanlei Meng |
ICPADS | 1 |
| 2022 | Hash Learning With Variable Quantization for Large-Scale RetrievalabstractApproximate Nearest Neighbor(ANN) search is the core problem in many large-scale machine learning and computer vision applications such as multimodal retrieval. Hashing is becoming increasingly popular, since it can provide efficient similarity search and compact data representations suitable for handling such large-scale ANN search problems. Most hashing algorithms concentrate on learning more effective projection functions. However, the accuracy loss in the quantization step has been ignored and barely studied. In this paper, we analyse the importance of various projected dimensions, distribute them into several groups and quantize them with two types of values which can both better preserve the neighborhood structure among data. One is Variable Integer-based Quantization (VIQ) that quantizes each projected dimension with integer values. The other is Variable Codebook-based Quantization (VCQ) that quantizes each projected dimension with corresponding codebook values. We conduct experiments on five common public data sets containing up to one million vectors. The results show that the proposed VCQ and VIQ algorithms can both achieve much higher accuracy than state-of-the-art quantization methods. Furthermore, although VCQ performs better than VIQ, ANN search with VIQ provides much higher search efficiency. Yuan Cao 0005, Sheng Chen 0015, Jie Gui, Heng Qi, Zhiyang Li 0001, Chao Liu 0008 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2022 | Scalable Distributed Hashing for Approximate Nearest Neighbor SearchabstractHashing has been widely applied to the large-scale approximate nearest neighbor search problem owing to its high efficiency and low storage requirement. Most investigations concentrate on learning hashing methods in a centralized setting. However, in existing big data systems, data is often stored across different nodes. In some situations, data is even collected in a distributed manner. A straightforward way to solve this problem is to aggregate all the data into the fusion center to obtain the search result (aggregating method). However, this strategy is not feasible because of the prohibitive communication cost. Although a few distributed hashing methods have been proposed to reduce this cost, they only focus on designing a distributed algorithm for a specific global optimization objective without considering scalability. Moreover, existing distributed hashing methods aim at finding a distributed solution to hashing, meanwhile avoiding accuracy loss, rather than improving accuracy. To address these challenges, we propose a Scalable Distributed Hashing (SDisH) model in which most existing hashing methods can be extended to process distributed data with no changes. Furthermore, to improve accuracy, we utilize the search radius as a global variable across different nodes to achieve a global optimum search result for every iteration. In addition, a voting algorithm is presented based on the results produced by multiple iterations to further reduce search errors. Theoretical analyses of communication, computation, and accuracy demonstrate the superiority of the proposed model. Numerical simulations on three large-scale and two relatively small benchmark datasets also show that the SDisH model achieves up to 44.75% and 10.23% accuracy gains compared to the aggregating method and state-of-the-art distributed hashing methods, respectively. Yuan Cao 0005, Heng Qi, Jie Gui, Keqiu Li, Jieping Ye, Chao Liu 0008 |
IEEE Trans. Image Process. | 1 |
| 2021 | A Comprehensive Survey on Image Dehazing Based on Deep LearningabstractThe presence of haze significantly reduces the quality of images. Researchers have designed a variety of algorithms for image dehazing (ID) to restore the quality of hazy images. However, there are few studies that summarize the deep learning (DL) based dehazing technologies. In this paper, we conduct a comprehensive survey on the recent proposed dehazing methods. Firstly, we conclude the commonly used datasets, loss functions and evaluation metrics. Secondly, we group the existing researches of ID into two major categories: supervised ID and unsupervised ID. The core ideas of various influential dehazing models are introduced. Finally, the open issues for future research on ID are pointed out. Jie Gui, Xiaofeng Cong, Yuan Cao 0005, Wenqi Ren, Jun Zhang 0011, Jing Zhang 0037, Dacheng Tao |
IJCAI | 3 |
| 2021 | Deep Cross-Modal Supervised Hashing Based on Joint Semantic Matrix
Yuan Cao 0005, Chao Liu 0008 |
NSS | 2 |
| 2021 | Fast kNN Search in Weighted Hamming Space With Multiple TablesabstractHashing methods have been widely used in Approximate Nearest Neighbor (ANN) search for big data due to low storage requirements and high search efficiency. These methods usually map the ANN search for big data into the k -Nearest Neighbor ( k NN) search problem in Hamming space. However, Hamming distance calculation ignores the bit-level distinction, leading to confusing ranking. In order to further increase search accuracy, various bit-level weights have been proposed to rank hash codes in weighted Hamming space. Nevertheless, existing ranking methods in weighted Hamming space are almost based on exhaustive linear scan, which is time consuming and not suitable for large datasets. Although Multi-Index hashing that is a sub-linear search method has been proposed, it relies on Hamming distance rather than weighted Hamming distance. To address this issue, we propose an exact k NN search approach with Multiple Tables in Weighted Hamming space named WHMT, in which the distribution of bit-level weights is incorporated into the multi-index building. By WHMT, we can get the optimal candidate set for exact k NN search in weighted Hamming space without exhaustive linear scan. Experimental results show that WHMT can achieve dramatic speedup up to 69.8 times over linear scan baseline without losing accuracy in weighted Hamming space. Jie Gui, Yuan Cao 0005, Heng Qi, Keqiu Li, Jieping Ye, Chao Liu 0008, Xiaowei Xu 0005 |
IEEE Trans. Image Process. | 2 |
| 2021 | Learning to Hash With Dimension Analysis Based Quantizer for Image RetrievalabstractThe last few years have witnessed the rise of the big data era in which approximate nearest neighbor search is a fundamental problem in many applications, such as large-scale image retrieval. Recently, many research results have demonstrated that hashing can achieve promising performance due to its appealing storage and search efficiency. Since complex optimization problems for loss functions are difficult to solve, most hashing methods decompose the hash code learning problem into two steps: projection and quantization. In the quantization step, binary codes are widely used because ranking them by the Hamming distance is very efficient. However, the massive information loss produced by the quantization step should be reduced in applications where high search accuracy is required, such as in image retrieval. Since many two-step hashing methods produce uneven projected dimensions in the projection step, in this paper, we propose a novel dimension analysis-based quantization (DAQ) on two-step hashing methods for image retrieval. We first perform an importance analysis of the projected dimensions and select a subset of them that are more informative than others, and then we divide the selected projected dimensions into several regions with our quantizer. Every region is quantized with its corresponding codebook. Finally, the similarity between two hash codes is estimated by the Manhattan distance between their corresponding codebooks, which is also efficient. We conduct experiments on three public benchmarks containing up to one million descriptors and show that the proposed DAQ method consistently leads to significant accuracy improvements over state-of-the-art quantization methods. Yuan Cao 0005, Heng Qi, Jie Gui, Keqiu Li, Yuan Yan Tang, James T. Kwok |
IEEE Trans. Multim. | 1 |
| 2020 | Scale balance for prototype-based binary quantization
Zhiyang Li 0001, Wenyu Qu, Yuan Cao 0005, Heng Qi, Milos Stojmenovic, Jia Hu 0001 |
Pattern Recognit. | 3 |
| 2019 | General Distributed Hash Learning on Image Descriptors for $k$-Nearest Neighbor SearchabstractHashing methods have attracted much attention due to their superior time and storage properties for image retrieval. To learn similarity-preserving hash function, most existing methods are designed for the centralized setting. However, the current data storage systems are distributed to increase scalability. Obviously, it is infeasible to aggregate all the data into a fusion center because of the prohibitively expensive communication and computation overhead. Motivated by this, some methods are proposed to achieve hashing for distributed data. However, these methods mostly focus on extending one specific hashing to a distributed model without considering the generality. In this letter, we propose a novel general distributed hash learning model, which can be viewed as an effective distributed model of most hashing methods. The proposed model can achieve up to 15.2% accuracy gains over state-of-the-art distributed hashing methods, while the communication cost is independent on the data size. Yuan Cao 0005, Heng Qi, Jie Gui, Shuai Li 0002, Keqiu Li |
IEEE Signal Process. Lett. | 1 |