Tianshi Wang 0001

dblp:147/8926-1 · DBLP profile ↗
← Back
15ranked-venue papers
4as first author
12since 2021 · last 2026
0000-0002-8013-5188ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Unified stable and generalizable online hashing for cross-modal retrieval
Tianshi Wang 0001, Fengling Li 0001, Guohua Dong, Lei Zhu 0002
Pattern Recognit.2
2026 Prompt-Driven Bit Extension Hashing for Continual Cross-Modal Retrieval
abstract
Continual cross-modal hashing is critical for efficient retrieval across heterogeneous modalities in dynamic environments. Yet, existing approaches primarily focus on mitigating catastrophic forgetting, while overlooking two key challenges: 1) the hash collision arising from the excessive utilization of the Hamming space across tasks, and 2) the absence of consistency modeling for cross-modal dynamic associations. To address these challenges, we introduce Prompt-driven Bit Extension Hashing (PBEH), a novel framework that dynamically extends hash codes to prevent hash collisions and capture evolving modality-aligned semantics in continuously expanding multi-modal data. Specifically, PBEH first adaptively initializes a set of modality-shared prompts for each task, which are jointly optimized with the hashing functions to enhance model plasticity and retain task-specific knowledge, enabling continual cross-modal semantic alignment. In parallel, a dynamic Hamming space extension mechanism allocates dedicated capacity per task, alleviating bottlenecks and inter-task collisions. During retrieval, queries are encoded via the extended hash functions and matched to stored codes using a truncated strategy for compatibility. To ensure efficiency and semantic stability, only the prompts and hashing functions are updated while the pre-trained backbone remains frozen. Extensive experiments demonstrate that PBEH achieves superior and stable performance in continual cross-modal retrieval with substantially reduced computational overhead. The source codes and datasets are available at https://github.com/Liuwwhh/PBEH.
Tianshi Wang 0001, Fengling Li 0001, Jingjing Li 0001, Lei Zhu 0002
IEEE Trans. Circuits Syst. Video Technol.2
2026 PIC-CMH: Efficient Prompt-Infused Continual Cross-Modal Hashing
abstract
Cross-modal hashing models face significant challenges in handling continuous data growth, particularly in balancing the plasticity to learn new knowledge and the stability to retain prior cross-modal knowledge. Existing studies partially address this by maintaining previous mappings or extending hash codes, but struggle to reconcile plasticity and stability while requiring heavy parameter optimization. To tackle this, we propose an efficient Prompt-Infused Continual Cross-Modal Hashing (PIC-CMH) approach designed for hash learning with the continuous growth of multi-modal data and emerging knowledge. Specifically, PIC-CMH introduces a finite set of learnable multi-modal prompts, including global and task-specific expert prompts, which work in synergy with multi-modal representations. All prompts are optimized with the hash functions via backpropagation after Gaussian initialization. Global prompts stay learnable throughout, linking tasks, while expert prompts are updated only within their tasks, facilitating knowledge acquisition and mitigating catastrophic forgetting in continual learning. By freezing the pre-trained models used for multi-modal representations, continual learning is confined to the lightweight multi-modal prompts and hash functions, significantly reducing computational overhead. Extensive experiments demonstrate that PIC-CMH effectively addresses the stability-plasticity trade-off in cross-modal hash learning, delivering high retrieval accuracy with low computational cost and a simple yet efficient architecture. The source codes and datasets are available athttps://github.com/styx29-0/PIC-CMH.
Fengling Li 0001, Tianshi Wang 0001, Lei Zhu 0002, Xiaojun Chang
IEEE Trans. Multim.3
2026 Noise-Robust Generative Hashing for Cross-Modal Retrieval
abstract
Deep hashing has proven remarkable effectiveness for large-scale cross-modal retrieval, yet its performance is highly vulnerable to supervisory noise, such as mismatched cross-modal correspondences and incorrect category labels. Such noise is prevalent in real-world scenarios, where correspondence mismatches and label inaccuracies often coexist, posing significant challenges for learning accurate multimodal representations. Existing methods typically address only a single type of noise in isolation and neglect the potential value of noisy data, resulting in limited performance gains. To address these challenges, we propose Noise-Robust Generative Hashing (NRGH), a unified framework designed to accommodate various forms of noise inherent in cross-modal retrieval. Specifically, NRGH introduces a hash-driven noise estimation module that computes the confidence score for each multimodal sample by combining frozen auxiliary hash functions with a Gaussian mixture model. Guided by these confidence scores, NRGH performs data correction through two stages: generative text refinement and multi-label probability calibration. The former leverages a pre-trained vision-language model to generate descriptive captions that refine noisy textual information, while the latter corrects noisy labels using confidence-aware soft labels. Furthermore, a dynamic margin contrastive loss adaptively modulates the data contribution of each sample based on its confidence, enabling sample-level adaptive learning. Extensive experiments on benchmark datasets demonstrate that NRGH significantly exceeds state-of-the-art baselines in various noisy scenarios, delivering superior robustness and accuracy. Our source codes and datasets are available at https://github.com/xiaolaohuuu/NRGH .
Tianshi Wang 0001, Fengling Li 0001, Jingjing Li 0001, Lei Zhu 0002
ACM Trans. Multim. Comput. Commun. Appl.2
2025 Attack as Defense: Proactive Adversarial Multi-Modal Learning to Evade Retrieval
abstract
With growing concerns about information security, protecting the privacy of user-sensitive data has become crucial. The rapid development of multi-modal retrieval technologies poses new threats, making sensitive data more vulnerable to leakage and malicious mining. To address this, we introduce a Proactive Adversarial Multi-modal Learning (PAML) approach that transforms sensitive data into adversarial counterparts, evading malicious multi-modal retrieval and ensuring privacy. Our method starts by sending queries to a knowledge-agnostic retrieval system and analyzing the results to understand the retrieval feedback mechanism. Using a U-Net-based diffusion model, we create a semantic perturbation network that subtly alters the implicit semantics of sensitive data. This, combined with multi-modal retrieved results and random noise, shifts the data's semantics towards outliers, preventing retrieval as neighbors to relevant queries. Additionally, a discriminator and pre-trained model enhance the visual realism and outlier generalization of protected data. Extensive experiments show that PAML outperforms potential baselines in data privacy protection. Ablation analysis validates each component's effectiveness, and our approach's variants are applicable to diverse retrieval systems.
Fengling Li 0001, Tianshi Wang 0001, Lei Zhu 0002, Jingjing Li 0001, Heng Tao Shen
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 Generative Augmentation Hashing for Few-Shot Cross-Modal Retrieval
abstract
Deep cross-modal hashing has demonstrated strong performance in large-scale retrieval but remains challenging in few-shot scenarios due to limited data and weak cross-modal alignment. We propose Generative Augmentation Hashing (GAH), a new framework that synergizes Visual-Language Models (VLMs) and generation-driven hashing to address these limitations. GAH first introduces a cycle generative augmentation mechanism: VLMs generate descriptive textual captions for images, which, combined with label semantics, guide diffusion models to synthesize semantically aligned images via inconsistency filtering. These images then regenerate coherent textual descriptions through VLMs, forming a self-reinforcing cycle that iteratively expands cross-modal data. To resolve the diversity-alignment trade-off in augmentation, we design cross-modal perturbation enhancement, injecting synchronized perturbations with controlled noise to preserve inter-modal semantic relationships while enhancing robustness. Finally, GAH employs dual-level adversarial hash learning, where adversarial alignment of modality-specific and shared latent spaces optimizes both cross-modal consistency and discriminative hash code generation, effectively bridging heterogeneous gaps. Extensive experiments on benchmark datasets show that GAH outperforms state-of-the-art methods in few-shot cross-modal retrieval, achieving significant improvements in retrieval accuracy. Our source codes and datasets are available at https://github.com/xiaolaohuuu/GAH.
Fengling Li 0001, Tianshi Wang 0001, Lei Zhu 0002, Xiaojun Chang
IEEE Trans. Circuits Syst. Video Technol.3
2025 Fast Partial-Modal Online Cross-Modal Hashing
abstract
Cross-Modal Hashing (CMH) has become a powerful technique for large-scale cross-modal retrieval, offering benefits like fast computation and efficient storage. However, most CMH models struggle to adapt to streaming multimodal data in real-time once deployed. Although recent online CMH studies have made progress in this area, they often overlook two key challenges: 1) learning effectively from streaming partial-modal multimodal data, and 2) avoiding the high costs associated with frequent hash function re-training and large-scale updates to database hash codes. To address these issues, we propose Fast Partial-modal Online Cross-Modal Hashing (FPO-CMH), the first approach to tackle online cross-modal hash learning with partial-modal data. This marks a significant shift from previous methods that rely on fully-available multimodal data. Specifically, our approach introduces a multimodal dual-tier anchor bank, initialized using offline training data, which allows offline-trained CMH models to adapt seamlessly to partial-modal data while progressively updating the anchor bank. By leveraging gradient accumulation and asynchronous optimization, FPO-CMH facilitates efficient online cross-modal hash learning. Additionally, an initial-anchor rehearsal strategy is employed to prevent model catastrophic forgetting during online optimization, ensuring the code invariance of database hash codes and eliminating the need for frequent hash function re-training. Extensive experiments validate the superiority of FPO-CMH, especially in handling streaming partial-modal multimodal data, a more realistic scenario. The source codes and datasets are available at https://github.com/DandelionWow/FPO-CMH.
Fengling Li 0001, Tianshi Wang 0001, Lei Zhu 0002, Xiaojun Chang
IEEE Trans. Image Process.3
2024 Cross-Modal Retrieval: A Systematic Review of Methods and Future Directions
abstract
With the exponential surge in diverse multimodal data, traditional unimodal retrieval methods struggle to meet the needs of users seeking access to data across various modalities. To address this, cross-modal retrieval has emerged, enabling interaction across modalities, facilitating semantic matching, and leveraging complementarity and consistency between heterogeneous data. Although prior literature has reviewed the field of cross-modal retrieval, it suffers from numerous deficiencies in terms of timeliness, taxonomy, and comprehensiveness. This article conducts a comprehensive review of cross-modal retrieval’s evolution, spanning from shallow statistical analysis techniques to vision-language pretraining (VLP) models. Commencing with a comprehensive taxonomy grounded in machine learning paradigms, mechanisms, and models, this article delves deeply into the principles and architectures underpinning existing cross-modal retrieval methods. Furthermore, it offers an overview of widely used benchmarks, metrics, and performances. Lastly, this article probes the prospects and challenges that confront contemporary cross-modal retrieval, while engaging in a discourse on potential directions for further progress in the field. To facilitate the ongoing research on cross-modal retrieval, we develop a user-friendly toolbox and an open-source repository athttps://cross-modal-retrieval.github.io.
Tianshi Wang 0001, Fengling Li 0001, Lei Zhu 0002, Jingjing Li 0001, Zheng Zhang 0006, Heng Tao Shen
Proc. IEEE1
2024 Invisible Black-Box Backdoor Attack against Deep Cross-Modal Hashing Retrieval
abstract
Deep cross-modal hashing has promoted the field of multi-modal retrieval due to its excellent efficiency and storage, but its vulnerability to backdoor attacks is rarely studied. Notably, current deep cross-modal hashing methods inevitably require large-scale training data, resulting in poisoned samples with imperceptible triggers that can easily be camouflaged into the training data to bury backdoors in the victim model. Nevertheless, existing backdoor attacks focus on the uni-modal vision domain, while the multi-modal gap and hash quantization weaken their attack performance. In addressing the aforementioned challenges, we undertake an invisible black-box backdoor attack against deep cross-modal hashing retrieval in this article. To the best of our knowledge, this is the first attempt in this research field. Specifically, we develop a flexible trigger generator to generate the attacker’s specified triggers, which learns the sample semantics of the non-poisoned modality to bridge the cross-modal attack gap. Then, we devise an input-aware injection network, which embeds the generated triggers into benign samples in the form of sample-specific stealth and realizes cross-modal semantic interaction between triggers and poisoned samples. Owing to the knowledge-agnostic of victim models, we enable any cross-modal hashing knockoff to facilitate the black-box backdoor attack and alleviate the attack weakening of hash quantization. Moreover, we propose a confusing perturbation and mask strategy to induce the high-performance victim models to focus on imperceptible triggers in poisoned samples. Extensive experiments on benchmark datasets demonstrate that our method has a state-of-the-art attack performance against deep cross-modal hashing retrieval. Besides, we investigate the influences of transferable attacks, few-shot poisoning, multi-modal poisoning, perceptibility, and potential defenses on backdoor attacks. Our codes and datasets are available at https://github.com/tswang0116/IB3A.
Tianshi Wang 0001, Fengling Li 0001, Lei Zhu 0002, Jingjing Li 0001, Zheng Zhang 0006, Heng Tao Shen
ACM Trans. Inf. Syst.1
2023 Targeted Adversarial Attack Against Deep Cross-Modal Hashing Retrieval
abstract
Deep cross-modal hashing has achieved excellent retrieval performance with the powerful representation capability of deep neural networks. Regrettably, current methods are inevitably vulnerable to adversarial attacks, especially well-designed subtle perturbations that can easily fool deep cross-modal hashing models into returning irrelevant or the attacker’s specified results. Although adversarial attacks have attracted increasing attention, there are few studies on specialized attacks against deep cross-modal hashing. To solve these issues, we propose a targeted adversarial attack method against deep cross-modal hashing retrieval in this paper. To the best of our knowledge, this is the first work in this research field. Concretely, we first build a progressive fusion module to extract fine-grained target semantics through a progressive attention mechanism. Meanwhile, we design a semantic adaptation network to generate the target prototype code and reconstruct the category label, thus realizing the semantic interaction between the target semantics and the implicit semantics of the attacked model. To bridge modality gaps and preserve local example details, a semantic translator seamlessly translates the target semantics and then embeds them into benign examples in collaboration with a U-Net framework. Moreover, we construct a discriminator for adversarial training, which enhances the visual realism and category discrimination of adversarial examples, thus improving their targeted attack performance. Extensive experiments on widely tested cross-modal retrieval datasets demonstrate the superiority of our proposed method. Also, transferable attacks show that our generated adversarial examples have well generalization capability on targeted attacks. The source codes and datasets are available athttps://github.com/tswang0116/TA-DCH.
Tianshi Wang 0001, Lei Zhu 0002, Zheng Zhang 0006, Huaxiang Zhang 0001, Junwei Han 0001
IEEE Trans. Circuits Syst. Video Technol.1
2023 Efficient Query-based Black-box Attack against Cross-modal Hashing Retrieval
abstract
Deep cross-modal hashing retrieval models inherit the vulnerability of deep neural networks. They are vulnerable to adversarial attacks, especially for the form of subtle perturbations to the inputs. Although many adversarial attack methods have been proposed to handle the robustness of hashing retrieval models, they still suffer from two problems: (1) Most of them are based on the white-box settings, which is usually unrealistic in practical application. (2) Iterative optimization for the generation of adversarial examples in them results in heavy computation. To address these problems, we propose an Efficient Query-based Black-Box Attack (EQB 2 A) against deep cross-modal hashing retrieval, which can efficiently generate adversarial examples for the black-box attack. Specifically, by sending a few query requests to the attacked retrieval system, the cross-modal retrieval model stealing is performed based on the neighbor relationship between the retrieved results and the query, thus obtaining the knockoffs to substitute the attacked system. A multi-modal knockoffs-driven adversarial generation is proposed to achieve efficient adversarial example generation. While the entire network training converges, EQB 2 A can efficiently generate adversarial examples by forward-propagation with only given benign images. Experiments show that EQB 2 A achieves superior attacking performance under the black-box setting.
Lei Zhu 0002, Tianshi Wang 0001, Jingjing Li 0001, Zheng Zhang 0006, Jialie Shen 0001, Xinhua Wang 0003
ACM Trans. Inf. Syst.2
2021 CGNet: A Cascaded Generative Network for dense point cloud reconstruction from a single image
Li Liu 0031, Huaxiang Zhang 0001, Tianshi Wang 0001
Knowl. Based Syst.4
2020 A background-induced generative network with multi-level discriminator for text-to-image generation
abstract
Most existing text-to-image generation methods focus on synthesizing images using only text descriptions, but this cannot meet the requirement of generating desired objects with given backgrounds. In this paper, we propose a Background-induced Generative Network (BGNet) that combines attention mechanisms, background synthesis, and multi-level discriminator to generate realistic images with given backgrounds according to text descriptions. BGNet takes a multi-stage generation as the basic framework to generate fine-grained images and introduces a hybrid attention mechanism to capture the local semantic correlation between texts and images. To adjust the impact of the given backgrounds on the synthesized images, synthesis blocks are added at each stage of image generation, which appropriately combines the foreground objects generated by the text descriptions with the given background images. Besides, a multi-level discriminator and its corresponding loss function are proposed to optimize the synthesized images. The experimental results on the CUB bird dataset demonstrate the superiority of our method and its ability to generate realistic images with given backgrounds.
Li Liu 0031, Huaxiang Zhang 0001, Tianshi Wang 0001
MMAsia4
2020 A multi-label text classification method via dynamic semantic representation model and deep neural network
Tianshi Wang 0001, Li Liu 0031, Naiwen Liu, Huaxiang Zhang 0001
Appl. Intell.1
2020 Fuzzy weighted sparse reconstruction error-steered semi-supervised learning for face recognition
Li Liu 0031, Xiuxiu Chen, Tianshi Wang 0001
Vis. Comput.4