Yu-Wei Zhan

dblp:274/6929 · DBLP profile ↗
← Back
24ranked-venue papers
5as first author
23since 2021 · last 2026
0000-0002-5822-5646ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 15 · 3 first-author · 14 since 2021Artificial intelligence and machine learning · 9 · 1 first-author · 9 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Reflective Cross-Granularity Grounding with Preference Optimization for Long Video Understanding
abstract
Video Large Language Models (video LLMs) have demonstrated remarkable capabilities in video understanding tasks, such as video question answering and temporal localization. However, understanding long videos still remains a significant challenge. Existing video LLMs adopt uni-granularity tokens for long videos, failing to simultaneously understand both high-level semantics and low-level visual details in videos. To tackle this problem, we propose ReCrossVLLM, a reflective cross-granularity grounding framework for video LLM with preference optimization to collaboratively achieve long video understanding, which not only retains the capabilities of high-level video semantics understanding, but also strengthens the fine-grained understanding abilities. Specifically, we propose the coarse-to-fine grounding and fine-to-coarse reflection strategies for long video understanding. In the coarse-to-fine grounding strategy, the video LLM with a coarse-grained module first locates the key video segments from the long video by tackling massive frames of the long video with fewer per-frame tokens. And then video LLM adapted with the fine-grained module further analyzes the key video segments with more per-frame tokens so that it can understand fine-grained information. In case the video LLM locates the wrong key video segments, during the inference stage, our designed fine-to-coarse reflection strategy instructs the fine-grained module to reflect the effectiveness of the locating result and decide whether to return to the coarse-to-fine grounding strategy with reflection feedback. Additionally, during the training stage, the coarse-to-fine grounding strategy is optimized with our proposed cross-granularity preference optimization strategy to further improve grounding efficiency. Extensive experiments for long video question answering and temporal video grounding tasks demonstrate that our proposed ReCrossVLLM framework can significantly improve the Video Large Language Model for long video understanding.
Xin Wang 0019, Hong Chen 0011, Yu-Wei Zhan, Zihan Song 0003, Bin Huang 0004, Kecheng Zheng, Wenwu Zhu 0001
ICMR4
2026 Federated class-incremental learning with prompting
Xin Luo 0006, Fang-Yi Liang, Yu-Wei Zhan, Zhen-Duo Chen 0001, Xin-Shun Xu
Expert Syst. Appl.4
2026 FusionHash: Empowering cross-modal hashing with fine-grained knowledge fusion
Bo-Lin Zhang, Yu-Wei Zhan, Zhen-Duo Chen 0001, Xin Luo 0006, Xin-Shun Xu
Knowl. Based Syst.5
2026 Dynamic Clustering-driven weakly-supervised online hashing with enhanced similarity
Chong-Yu Zhang, Yu-Wei Zhan, Zhen-Duo Chen 0001, Xin Luo 0006, Xin-Shun Xu
Pattern Recognit.3
2026 Gleaning Wisdom from the Past: Towards Label Incremental Learning for Online Hashing with a Plug-and-Play Framework
abstract
Benefiting from low storage requirements, high search efficiency, and promising performance, online hashing has attracted considerable attention for large-scale streaming data retrieval. However, most existing online hashing methods implicitly assume that the label space is fixed and fail to effectively handle the label incremental problem. With this motivation, in this article, we introduce a novel supervised online hashing framework named Incremental Plug-and-play Online Hashing (IPOH), which endows prevailing methods with the ability to solve the above issues. Specifically, IPOH could extract the patterns of old classes from the learned hash codes at the last round and then generate the class-wise representations for unseen new labels. With these representations, we could learn the hash codes for streaming data with new labels. Experimental results on two widely used datasets demonstrate the effectiveness of our method in handling class incremental data. We hope this article could convince more researchers to look into this interesting problem, and the source code for the implementation of our IPOH framework has been released at https://github.com/ZCyueternal/IPOH.git .
Chong-Yu Zhang, Xin Luo 0006, Yu-Wei Zhan, Zhen-Duo Chen 0001, Xin-Shun Xu
ACM Trans. Multim. Comput. Commun. Appl.3
2025 Tag-Aware Weakly-Supervised Online Hashing with Enhanced Joint Representation
abstract
Weakly-supervised online hashing has garnered significant attention recently, yet several challenges remain unresolved, such as how to effectively denoise tags, and how to efficiently learn hash functions in dynamic online scenarios. To tackle these challenges, we propose a novel method named Tag-Aware Weakly-supervised Online Hashing with enhanced joint representation (TA-WOH). Our method creates an enhanced joint representation with CLIP-based features in order to reduce the tag noise. Additionally, we introduce a tag association and noise model for improved similarity matrix and hash code learning. A novel mapping mechanism is developed to align joint representations with the optimal tag space, enhancing both the accuracy and robustness of the model. The computational complexity of TA-WOH is dependent solely on the size of the incoming data, ensuring scalability and efficiency for large-scale datasets. Extensive experiments on two datasets demonstrate that our method surpasses several state-of-the-art methods in both accuracy and efficiency.
Yu-Wei Zhan, Zhen-Duo Chen 0001, Yongxin Wang 0001, Xin Luo 0006, Xin-Shun Xu
ICASSP2
2025 Improving Compositional Generalization in Cross-Embodiment Learning via Mixture of Disentangled Prototypes
Ren Wang 0011, Xin Wang 0019, Tongtong Feng, Xinyue Gong, Guangyao Li 0001, Yu-Wei Zhan, Qing Li 0046, Wenwu Zhu 0001
ACM Multimedia6
2025 Enhancing HOI Detection with Contextual Cues from Large Vision-Language Models
Yu-Wei Zhan, Fan Liu 0008, Xin Luo 0006, Xin-Shun Xu, Liqiang Nie, Mohan Kankanhalli
ACM Multimedia1
2025 OH-CMH: Towards cross-modal hashing for streaming data with hierarchical labels and label increment scenario
Chong-Yu Zhang, Yu-Wei Zhan, Zhen-Duo Chen 0001, Yongxin Wang 0001, Xin Luo 0006, Xin-Shun Xu
Knowl. Based Syst.4
2025 Hypergraph-based CLIP hashing for unsupervised cross-modal retrieval
Jia-Rui Zhao, Yu-Wei Zhan, Zhen-Duo Chen 0001, Xin Luo 0006, Xin-Shun Xu
Knowl. Based Syst.4
2025 Class-Aware Prompting for Federated Few-Shot Class-Incremental Learning
abstract
Few-Shot Class-Incremental Learning (FSCIL) aims to continuously learn new classes from limited samples while preventing catastrophic forgetting. With the increasing distribution of learning data across different clients and privacy concerns, FSCIL faces a more realistic scenario where few learning samples are distributed across different clients, thereby necessitating a Federated Few-Shot Class-Incremental Learning (FedFSCIL) scenario. However, this integration faces challenges from non-IID problem, which affects model generalization and training efficiency. The communication overhead in federated settings also presents a significant challenge. To address these issues, we propose Class-Aware Prompting for Federated Few-Shot Class-Incremental Learning (FedCAP). Our framework leverages pre-trained models enhanced by a class-wise prompt pool, where shared class-wise keys enable clients to utilize global class information during training. This unifies the understanding of base class features across clients and enhances model consistency. We further incorporate a class-level information fusion module to improve class representation and model generalization. Our approach requires very few parameter transmission during model aggregation, ensuring communication efficiency. To our knowledge, this is the first study to explore the scenario of FedFSCIL. Consequently, we designed comprehensive experimental setups and made the code publicly available.
Fang-Yi Liang, Yu-Wei Zhan, Chong-Yu Zhang, Zhen-Duo Chen 0001, Xin Luo 0006, Xin-Shun Xu
IEEE Trans. Circuits Syst. Video Technol.2
2025 Distributed Learning for Privacy-Preserving Semi-Supervised Video Anomaly Detection
abstract
Semi-supervised video anomaly detection (SS-VAD) is essential for intelligent monitoring. However, collecting large-scale surveillance videos from various organizations raises significant privacy concerns regarding sensitive information. Federated learning offers a promising solution by enabling distributed learning among multiple participants while safeguarding privacy. Despite its potential, research on applying federated learning to SS-VAD remains unexplored due to the inherent challenges of this task. In this paper, we solve this task via proposing DLPP, a novel distributed learning framework for privacy-preserving SS-VAD. It addresses the issue of statistical heterogeneity among data from different participants in real-world federated SS-VAD applications, particularly focusing on non-independent and identically distributed (non-IID) data and imbalanced data volumes. In specific, it addresses these challenges in two key innovations: 1) For the non-IID data challenge, it dynamically updates the client model based on the overall gradient at the client of the previous training round and the degree of divergence between the server model and the client model. In this way, it can better adapt the server model to each client and promote convergence. 2) For the imbalanced data volumes challenge, it adaptively allocates client aggregation weights by comprehensively considering the data volumes, model quality, and learning efficiency of clients. This means a more robust server model can be obtained, and model bias reduced. We conduct extensive experiments to evaluate the performance of DLPP on benchmark datasets by partitioning data to simulate various degrees of non-IID environments. The results show that DLPP significantly outperforms both Baseline and SOTA methods, achieving up to a 3.89% improvement, and its communication efficiency is 3x better than FedAvg.
Xiao-Dong Xie, Yu-Wei Zhan, Zhen-Xiang Ma, Hong-Mei Liu, Zhen-Duo Chen 0001, Xin Luo 0006, Xin-Shun Xu
IEEE Trans. Circuits Syst. Video Technol.2
2024 FedCAFE: Federated Cross-Modal Hashing with Adaptive Feature Enhancement
abstract
Deep Cross-Modal Hashing (CMH) has become one of the most popular solutions for cross-modal retrieval. Existing methods need to first collect data and then be trained with these accumulated data. However, in real world, data may be generated and possessed by different owners. Considering the concerns about privacy, data may not be shared or transmitted, leading to the failure of sufficient training of CMH. To solve the problem, we propose a new framework called Federated Cross-modal Hashing with Adaptive Feature Enhancement (FedCAFE). FedCAFE is a federated method which could use distributed data to train existing CMH methods under the privacy protection. To overcome the data heterogeneity challenge of distributed data and improve the generalization ability of global model, FedCAFE is endowed with a novel adaptive feature enhancement module and a new weighted aggregation strategy. Besides, it could fully utilize the rich global information carried in the global model to constrain the model during the local training process. We have conducted extensive experiments on four widely-used datasets in CMH domain with both IID and non-IID settings. The reported results demonstrate that the proposed FedCAFE achieves better performance than several state-of-the-art baselines.
Yu-Wei Zhan, Chong-Yu Zhang, Xin Luo 0006, Zhen-Duo Chen 0001, Yongxin Wang 0001, Xun Yang 0001, Xin-Shun Xu
ACM Multimedia2
2024 POLISH: Adaptive Online Cross-Modal Hashing for Class Incremental Data
abstract
In recent years, hashing-based online cross-modal retrieval has garnered growing attention. This trend is motivated by the fact that web data is increasingly delivered in a streaming manner as opposed to batch processing. Simultaneously, the sheer scale of web data sometimes makes it impractical to fully load for the training of hashing models. Despite the evolution of online cross-modal hashing techniques, several challenges remain: 1) Most existing methods learn hash codes by considering the relevance among newly arriving data or between new data and the existing data, often disregarding valuable global semantic information. 2) A common but limiting assumption in many methods is that the label space remains constant, implying that all class labels should be provided within the first data chunk. This assumption does not hold in real-world scenarios, and the presence of new labels in incoming data chunks can severely degrade or even break these methods.
Yu-Wei Zhan, Xin Luo 0006, Zhen-Duo Chen 0001, Yongxin Wang 0001, Yinwei Wei, Xin-Shun Xu
WWW1
2024 Multiple Information Embedded Hashing for Large-Scale Cross-Modal Retrieval
abstract
Recently, many efforts have been devoted to improving the retrieval performance of supervised cross-modal hashing; however, current methods are gradually reaching a performance bottleneck, especially when dealing with real-world multimedia data. This is mainly due to their application of coarse-grained semantics, unrobust hash functions, and inflexible workflows. Therefore, discovering refined semantics hidden in data, designing robust hash functions, and creating a non-interfering but facilitative learning workflow are much more significant. With this motivation, in this paper, we propose a novel supervised cross-modal hashing method, i.e., Multiple Information Embedded Hashing, MIEH for short. It consists of a three-step working flow that flexibly handles multiple information mining, hash code learning, and hash function learning. First, it explores the multimedia data from multiple perspectives such as modal-level consistency, class-level discriminability, and instance-level similarity to mine comprehensive semantic information, which not only contributes to the generation of discriminative hash codes, but also accelerates convergence. Subsequently, MIEH is committed to embed the refined semantics into targeted hash codes with an efficient discrete optimization algorithm. Finally, it improves the learning ability of linear hash function by noisy example erasing and deviation correcting. Considering this, MIEH is able to garner more robust hash function. Extensive experiments conducted on three popular benchmark datasets highlight the superiority of our MIEH on large-scale cross-modal retrieval tasks and demonstrate its competitive performance against state-of-the-art approaches. The source code is available1.
Yongxin Wang 0001, Yu-Wei Zhan, Zhen-Duo Chen 0001, Xin Luo 0006, Xin-Shun Xu
IEEE Trans. Circuits Syst. Video Technol.2
2023 Prototype-Based Layered Federated Cross-Modal Hashing
abstract
Recently, deep cross-modal hashing has gained increasing attention. However, in many practical cases, data are distributed and cannot be collected due to privacy concerns, which greatly reduces the cross-modal hashing performance on each client. And due to the problems of statistical heterogeneity, model heterogeneity, and forcing each client to accept the same parameters, applying federated learning to cross-modal hash learning becomes very tricky. In this paper, we propose a novel method called prototype-based layered federated cross-modal hashing. Specifically, the prototype is introduced to learn the similarity between instances and classes on server, reducing the impact of statistical heterogeneity (non-IID) on different clients. And we monitor the distance between local and global prototypes to further improve the performance. To realize personalized federated learning, a hypernetwork is deployed on server to dynamically update different layers’ weights of local model. Experimental results on benchmark datasets show that our method outperforms state-of-the-art methods.
Yu-Wei Zhan, Xin Luo 0006, Zhen-Duo Chen 0001, Yongxin Wang 0001, Xin-Shun Xu
ICASSP2
2023 Self-Distillation Dual-Memory Online Hashing with Hash Centers for Streaming Data Retrieval
abstract
With the continuous generation of massive amounts of multimedia data nowadays, hashing has demonstrated significant potentials for large-scale search. To handle the emerging needs for streaming data retrieval, online hashing is drawing more and more attention. For online scenario, data distribution may change and concept drifts may occur as new data is continuously added to the database. Inevitably, hashing models may lose or disrupt the previously obtained knowledge when learning from new information, which is called the problem of catastrophic forgetting. In this paper, we propose a new online hashing method called Self-distillation Dual-memory Online Hashing with Hash Centers, which is abbreviated to SDOH-HC, to overcome this challenge. Specifically, SDOH-HC contains replay and distillation modules. For replay, a dual-memory mechanism is proposed which involves hash centers and exemplars. For knowledge distillation, we let hash centers distill information from themselves, i.e., the version of last round. Additionally, a new objective function is further built on above modules and is solved discretely to learn hash codes. Extensive experiments on three benchmark datasets demonstrate the effectiveness of our method.
Chong-Yu Zhang, Xin Luo 0006, Yu-Wei Zhan, Peng-Fei Zhang 0001, Zhen-Duo Chen 0001, Yongxin Wang 0001, Xun Yang 0001, Xin-Shun Xu
ACM Multimedia3
2023 Expansion window local alignment weighted network for fine-grained sketch-based image retrieval
Zi-Chao Zhang 0002, Zhen-Yu Xie, Zhen-Duo Chen 0001, Yu-Wei Zhan, Xin Luo 0006, Xin-Shun Xu
Pattern Recognit.4
2022 Online Enhanced Semantic Hashing: Towards Effective and Efficient Retrieval for Streaming Multi-Modal Data
abstract
With the vigorous development of multimedia equipments and applications, efficient retrieval of large-scale multi-modal data has become a trendy research topic. Thereinto, hashing has become a prevalent choice due to its retrieval efficiency and low storage cost. Although multi-modal hashing has drawn lots of attention in recent years, there still remain some problems. The first point is that existing methods are mainly designed in batch mode and not able to efficiently handle streaming multi-modal data. The second point is that all existing online multi-modal hashing methods fail to effectively handle unseen new classes which come continuously with streaming data chunks. In this paper, we propose a new model, termed Online enhAnced SemantIc haShing (OASIS). We design novel semantic-enhanced representation for data, which could help handle the new coming classes, and thereby construct the enhanced semantic objective function. An efficient and effective discrete online optimization algorithm is further proposed for OASIS. Extensive experiments show that our method can exceed the state-of-the-art models. For good reproducibility and benefiting the community, our code and data are already publicly available.
Xiao-Ming Wu 0002, Xin Luo 0006, Yu-Wei Zhan, Chenlu Ding, Zhen-Duo Chen 0001, Xin-Shun Xu
AAAI3
2022 Weakly-Supervised Online Hashing with Refined Pseudo Tags
abstract
With the rapid development of social media, various types of tags uploaded by social users are attached to the images. Compared to clean labels marked by experts, although user-provided tags are imperfect, e.g., wrong tags, reduplicative tags, or missing tags, they are more diverse, fine-grained, and informative. Currently, there exist several weakly-supervised hashing methods attempting to learn hash codes using tags as supervision. Although they could benefiting from the rich information contained in tags, most of them may defy the nature of social media data. In real scenarios, social media data appears in streaming fashion, but most weakly-supervised hashing methods are just batch-based which cannot effectively handle streaming data. To this end, only one weakly-supervised online hashing method has been proposed, but it is still far from enough to alleviate the negative effects of tags.
Chenlu Ding, Xin Luo 0006, Xiao-Ming Wu 0002, Yu-Wei Zhan, Rui Li 0090, Xin-Shun Xu
CIKM4
2022 Discrete online cross-modal hashing
Yu-Wei Zhan, Yongxin Wang 0001, Xiao-Ming Wu 0002, Xin Luo 0006, Xin-Shun Xu
Pattern Recognit.1
2021 Weakly-Supervised Online Hashing
abstract
With the rapid development of social websites, recent years have witnessed an explosive growth of social images with user-provided tags. Most existing hashing methods for social image retrieval are batch-based which may violate the nature of social images, i.e., social images are usually generated periodically or collected in a stream fashion. Although there exist many online hashing methods, they either adopt unsupervised learning which ignore the relevant tags, or are designed in the supervised manner which needs high-quality labels. In this paper, to overcome the above limitations, we propose a new method named Weakly-supervised Online Hashing (WOH). In order to learn high-quality hash codes, WOH exploits the weak supervision, i.e., tags, by considering the semantics of tags and removing the noise. Besides, we develop a discrete online optimization algorithm, which is efficient and scalable. Extensive experiments conducted on two real-world datasets demonstrate the superiority of WOH.
Yu-Wei Zhan, Xin Luo 0006, Yongxin Wang 0001, Zhen-Duo Chen 0001, Xin-Shun Xu
ICME1
2021 TEACH: Attention-Aware Deep Cross-Modal Hashing
abstract
Hashing methods for cross-modal retrieval have recently been widely investigated due to the explosive growth of multimedia data. Generally, real-world data is imperfect and has more or less redundancy, making cross-modal retrieval task challenging. However, most existing cross-modal hashing methods fail to deal with the redundancy, leading to unsatisfactory performance on such data. In this paper, to address this issue, we propose a novel cross-modal hashing method, namely aTtEntion-Aware deep Cross-modal Hashing (TEACH). It could perform feature learning and hash-code learning simultaneously. Besides, with designed attention modules for different modalities, one for each, TEACH can effectively highlight the useful information of data while suppressing the redundant information. Extensive experiments on benchmark datasets demonstrate that our method outperforms some state-of-the-art hashing methods in cross-modal retrieval tasks.
Honglei Yao, Yu-Wei Zhan, Zhen-Duo Chen 0001, Xin Luo 0006, Xin-Shun Xu
ICMR2
2020 Supervised Hierarchical Deep Hashing for Cross-Modal Retrieval
abstract
Cross-modal hashing has attracted much attention in the large-scale multimedia search area. In many real applications, labels of samples have hierarchical structure which also contains much useful information for learning. However, most existing methods are originally designed for non-hierarchical labeled data and thus fail to exploit the rich information of the label hierarchy. In this paper, we propose an effective cross-modal hashing method, named Supervised Hierarchical Deep Cross-modal Hashing, SHDCH for short, to learn hash codes by explicitly delving into the hierarchical labels. Specifically, both the similarity at each layer of the label hierarchy and the relatedness across different layers are implanted into the hash-code learning. Besides, an iterative optimization algorithm is proposed to directly learn the discrete hash codes instead of relaxing the binary constraints. We conducted extensive experiments on two real-world datasets and the experimental results show the superior performance of SHDCH over several state-of-the-art methods.
Yu-Wei Zhan, Xin Luo 0006, Yongxin Wang 0001, Xin-Shun Xu
ACM Multimedia1