Zhulin Tao

dblp:253/7942 · DBLP profile ↗
← Back
23ranked-venue papers
5as first author
20since 2021 · last 2027
0000-0001-9011-8464ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 3 first-author · 8 since 2021Databases, data management, data science and information retrieval · 6 · 2 first-author · 4 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2027 MMVUF: Improving video-based human value understanding in multimodal large language models
Libiao Jin, Zhulin Tao
Expert Syst. Appl.5
2026 Flow-guided cascaded transformer for consistent video colorization
Yan Zhai, Zishan Li, Zhulin Tao, Longquan Dai, Xianglin Huang
Pattern Recognit.3
2026 UDMMColor: A Unified Diffusion Model for Multi-Modal Colorization
abstract
Diffusion model-based networks have been widely applied in the field of image generation and have gradually demonstrated a strong potential in image colorization tasks. However, despite the emergence of various colorization diffusion models, two major challenges remain: (1) the lack of effective control over the colorization process and (2) the prevalent issue of color bleeding. Integrating suitable conditional control can effectively alleviate these challenges. To this end, we propose a unified multi-modal diffusion model that harnesses diverse modality information to achieve flexible and high-quality colorization. Specifically, we introduce a Stroke-Adapter that extracts and integrates stroke prompt, enhancing user control over color distribution. Additionally, we design an Edge-Guided Attention mechanism to effectively inject edge information into the colorization process, significantly reducing color bleeding artifacts. Extensive comparative experiments demonstrate that our method outperforms state-of-the-art image colorization approaches in both qualitative and quantitative evaluations, achieving superior colorization results with enhanced controllability.
Yan Zhai, Zerui Han, Zhulin Tao, Xianglin Huang, Jinshan Pan, Jinhui Tang 0001
IEEE Trans. Circuits Syst. Video Technol.3
2026 Lightweight Multi-Dilated Transformer for Image Deblurring
abstract
Window-based Transformers have achieved promising results in image deblurring. However, their limited ability to capture nonlocal information hinders further improvement in deblurring performance. In this article, we develop an effective multi-dilated Transformer, named MDFormer, to address this issue. Specifically, we first develop a multi-dilated feature aggregation (MDFA) module, which aims to extract and aggregate nonlocal information with reduced computational costs. As commonly used feed-forward networks are pixelwise operations, we propose a dilated feed-forward network (DiFFN) module to enhance the information interaction between pixels further. Moreover, to fully utilize the features of different scales, we introduce a multiscale feature fusion (MSFF) module to provide improved guidance for image reconstruction. Extensive experiments demonstrate that the proposed method generates comparable results against state-of-the-art approaches with reduced computational costs.
Zhulin Tao, Jinshan Pan
IEEE Trans. Neural Networks Learn. Syst.2
2025 Cross-Stain Contrastive Learning for Paired Immunohistochemistry and Histopathology Slide Representation Learning
abstract
Universal, transferable whole-slide image (WSI) representations are central to computational pathology. Incorporating multiple markers (e.g., immunohistochemistry, IHC) alongside H&E enriches H&E-based features with diverse, biologically meaningful information. However, progress is limited by the scarcity of well-aligned multi-stain datasets. Inter-stain Misalignment shifts corresponding tissue across slides, hindering consistent patch-level features and degrading slide-level embeddings. To address this, we curated a slide-level aligned, five-stain dataset (H&E, HER2, KI67, ER, PGR) to enable paired H&E-IHC learning and robust cross-stain representation. Leveraging this dataset, we propose Cross-Stain Contrastive Learning (CSCL), a two-stage pretraining framework: a lightweight adapter trained with patch-wise contrastive alignment to improve the compatibility of H&E features with corresponding IHC-derived contextual cues; and slide-level representation learning with Multiple Instance Learning (MIL), which uses a cross-stain attention fusion module to integrate stain-specific patch features and a crossstain global alignment module to enforce consistency among slide-level embeddings across different stains. Experiments on cancer subtype classification, IHC biomarker status classification, and survival prediction, show consistent gains by yielding high-quality, transferable H&E slide-level representations. The code and data are available at: https://github.com/lily-zyz/CSCL.
Yizhi Zhang, Lei Fan 0007, Zhulin Tao, Donglin Di, Yang Song 0001, Sidong Liu, Cong Cong 0001
BIBM3
2025 Fine-tuning Multimodal Large Language Models for Product Bundling
abstract
Recent advances in product bundling have leveraged multimodal information through sophisticated encoders, but remain constrained by limited semantic understanding and a narrow scope of knowledge. Therefore, some attempts employ In-context Learning (ICL) to explore the potential of large language models (LLMs) for their extensive knowledge and complex reasoning abilities. However, these efforts are inadequate in understanding mulitmodal data and exploiting LLMs' knowledge for product bundling. To bridge the gap, we introduce Bundle-MLLM, a novel framework that fine-tunes LLMs through a hybrid item tokenization approach within a well-designed optimization strategy. Specifically, we integrate textual, media, and relational data into a unified tokenization, introducing a soft separation token to distinguish between textual and non-textual tokens. Additionally, a streamlined yet powerful multimodal fusion module is employed to embed all non-textual features into a single, informative token, significantly boosting efficiency. To tailor product bundling tasks for LLMs, we reformulate the task as a multiple-choice question with candidate items as options. We further propose a progressive optimization strategy that fine-tunes LLMs for disentangled objectives: learning bundle patterns and enhancing multimodal semantic understanding specific to product bundling. Extensive experiments demonstrate that our approach outperforms a range of state-of-the-art (SOTA) methods. Codes are available at https://github.com/Xiaohao-Liu/Bundle-MLLM
Xiaohao Liu, Zhulin Tao, Yunshan Ma 0002, Yinwei Wei, Tat-Seng Chua
KDD (1)3
2025 Can Multimodal Large Language Models Understand Human Values in Videos?
abstract
Human values are core principles to determine what is right, desirable, and important for individuals and societies. The deep integration of large language models (LLMs) into human life has facilitated their remarkable performance in value understanding. Multimodal content is rich in value-laden information. The development of multimodal large language models (MLLMs) provides new perspectives for multimodal value understanding. However, MLLMs’ capacity for understanding basic human values in the video domain remains underexplored. To bridge this gap, our work focuses on evaluating their ability to understand specific human values embedded in videos. We assess 11 advanced MLLMs based on the VVALUES video dataset, using a four-module strategy designed to answer two questions: Do MLLMs understand human values, and Do their design and training elements impact the performance? Based on the evaluation results, we derive valuable findings across eight aspects that provide deeper insights into the value understanding of MLLMs and offer useful guides for future research in this field.
Junbin Xiao, Zhulin Tao, Libiao Jin
MMAsia5
2025 EverybodyDance: Bipartite Graph-Based Identity Correspondence for Multi-Character Animation
abstract
Consistent pose‐driven character animation has achieved remarkable progress in single‐character scenarios. However, extending these advances to multi‐character settings is non‐trivial, especially when position swap is involved. Beyond mere scaling, the core challenge lies in enforcing correct Identity Correspondence (IC) between characters in reference and generated frames. To address this, we introduce EverybodyDance, a systematic solution targeting IC correctness in multi-character animation. EverybodyDance is built around the **Identity Matching Graph (IMG)**, which models characters in the generated and reference frames as two node sets in a weighted complete bipartite graph. Edge weights, computed via our proposed Mask–Query Attention (MQA), quantify the affinity between each pair of characters. Our key insight is to formalize IC correctness as a graph structural metric and to optimize it during training. We also propose a series of targeted strategies tailored for multi-character animation, including identity-embedded guidance, a multi-scale matching strategy, and pre-classified sampling, which work synergistically. Finally, to evaluate IC performance, we curate the **Identity Correspondence Evaluation** benchmark, dedicated to multi‐character IC correctness. Extensive experiments demonstrate that EverybodyDance substantially outperforms state‐of‐the‐art baselines in both IC and visual fidelity.
Haotian Ling, Zequn Chen, Qiuying Chen, Donglin Di, Yongjia Ma, Hao Li 0030, Zhulin Tao, Xun Yang 0001
NeurIPS8
2025 EgoBlind: Towards Egocentric Visual Assistance for the Blind
abstract
We present EgoBlind, the first egocentric VideoQA dataset collected from blind individuals to evaluate the assistive capabilities of contemporary multimodal large language models (MLLMs). EgoBlind comprises 1,392 first-person videos from the daily lives of blind and visually impaired individuals. It also features 5,311 questions directly posed or verified by the blind to reflect their in-situation needs for visual assistance. Each question has an average of 3 manually annotated reference answers to reduce subjectiveness.Using EgoBlind, we comprehensively evaluate 16 advanced MLLMs and find that all models struggle. The best performers achieve an accuracy near 60\%, which is far behind human performance of 87.4\%. To guide future advancements, we identify and summarize major limitations of existing MLLMs in egocentric visual assistance for the blind and explore heuristic solutions for improvement. With these efforts, we hope that EgoBlind will serve as a foundation for developing effective AI assistants to enhance the independence of the blind and visually impaired. Data and code are available at \url{https://github.com/doc-doc/EgoBlind}.
Junbin Xiao, Nanxin Huang, Zhulin Tao, Xun Yang 0001, Richang Hong, Meng Wang 0001, Angela Yao
NeurIPS4
2025 VideoQA in the Era of LLMs: An Empirical Study
Junbin Xiao, Nanxin Huang, Hangyu Qin, Yicong Li 0004, Fengbin Zhu, Zhulin Tao, Jianxing Yu, Tat-Seng Chua, Angela Yao
Int. J. Comput. Vis.7
2025 Multimodal understanding of human values in videos: A benchmark dataset and PLM-based method
abstract
Multimodal content has become the mainstream communication medium in video sharing platforms such as TikTok and Twitter, containing rich values information. Understanding human values is of great significance to multimodal content analysis and can be applied to downstream tasks such as recommendation systems and value alignment. However, current studies on human values mainly focus on text and lack a multimodal perspective. In this work, we present a new multimodal human values video dataset called VVALUES, which contains 5,104 annotated videos along with their titles. The dataset is labeled with coarse-grained polarity tags of positive and neutral, and fine-grained tags including 13 classes of value vocabulary. Based on VVALUES, we further develop a pre-trained language model (PLM)-based multimodal method adopting a dual-transformer variant for value recognition, MMVR. Extensive experiments demonstrate that our method significantly improves the performance of understanding values in videos. To the best of our knowledge, we are the first to try to incorporate human values in video understanding , and VVALUES is the first multimodal video dataset for human values.
Zhulin Tao, Nanxin Huang, Libiao Jin, Xiaofang Luo
Neurocomputing2
2025 Deep Frequency-Separable Temporal Network for Efficient Video Denoising
Zhulin Tao, Jinjuan Wang, Lifang Yang, Jinshan Pan, Jinhui Tang 0001
IEEE Trans. Multim.1
2025 Multimodal Consistency Suppression Factor for Fake News Detection
abstract
Recent multimodal fake news detection methods often use the consistency between textual and visual contents to determine the truth or fake of news information. Higher levels of textual-visual consistency typically lead to a greater likelihood of classifying a news item as real. However, a critical observation reveals that creators of most fake news intentionally select images that align with the textual content, thereby enhancing the credibility of the news. Consequently, high consistency between textual and visual contents alone cannot guarantee the authenticity of the information. To address this problem, we introduce a novel approach termed Multimodal Consistency-based Suppression Factor to modulate the significance of textual-visual consistency in information assessment. When the textual-visual matching is high, this suppression factor reduces the influence of consistency during the judgment process. Moreover, we use contrastive language-image pre-training (CLIP) model to extract features and measure the consistency level between modalities to guide multimodal fusion. In addition, we also use a method of compressing and fusing modal information based on variational autoencoder (VAE) to reconstruct CLIP features, learning the shared representation of different modal information of CLIP. Finally, extensive experiments were conducted on three publicly datasets, Weibo, Twitter, and Weibo21, and the results confirmed that our method outperformed the state-of-the-art methods in the field and had 0.8%, 2.6%, and 4.1% effect improvement on the accuracy rate.
Zhulin Tao, Xingyu Gao 0001, Xi Wang 0014, Xianglin Huang
ACM Trans. Multim. Comput. Commun. Appl.1
2024 Leveraging Multimodal Features and Item-level User Feedback for Bundle Construction
abstract
Automatic bundle construction is a crucial prerequisite step in various bundle-aware online services. Previous approaches are mostly designed to model the bundling strategy of existing bundles. However, it is hard to acquire large-scale well-curated bundle dataset, especially for those platforms that have not offered bundle services before. Even for platforms with mature bundle services, there are still many items that are included in few or even zero bundles, which give rise to sparsity and cold-start challenges in the bundle construction models. To tackle these issues, we target at leveraging multimodal features, item-level user feedback signals, and the bundle composition information, to achieve a comprehensive formulation of bundle construction. Nevertheless, such formulation poses two new technical challenges: 1) how to learn effective representations by unifying multiple features optimally, and 2) how to address the problems of modality missing, noise, and sparsity problems induced by the incomplete query bundles. In this work, to address these technical challenges, we propose a Contrastive Learning-enhanced Hierarchical Encoder method (CLHE). Specifically, we use self-attention modules to combine the multimodal and multi-item features, and then leverage both item- and bundle-level contrastive learning to enhance the representation learning, thus to counter the modality missing, noise, and sparsity problems. Extensive experiments on four datasets in two application domains demonstrate that our method outperforms a list of SOTA methods. The code and dataset are available at https://github.com/Xiaohao-Liu/CLHE.
Yunshan Ma 0002, Xiaohao Liu, Yinwei Wei, Zhulin Tao, Xiang Wang 0010, Tat-Seng Chua
WSDM4
2024 Attention-enhanced joint learning network for micro-video venue classification
Bing Wang 0013, Xianglin Huang, Gang Cao 0001, Lifang Yang, Zhulin Tao
Multim. Tools Appl.5
2024 BiSTNet: Semantic Image Prior Guided Bidirectional Temporal Feature Fusion for Deep Exemplar-Based Video Colorization
abstract
How to effectively explore the colors of exemplars and propagate them to colorize each frame is vital for exemplar-based video colorization. In this article, we present a BiSTNet to explore colors of exemplars and utilize them to help video colorization by a bidirectional temporal feature fusion with the guidance of semantic image prior. We first establish the semantic correspondence between each frame and the exemplars in deep feature space to explore color information from exemplars. Then, we develop a simple yet effective bidirectional temporal feature fusion module to propagate the colors of exemplars into each frame and avoid inaccurate alignment. We note that there usually exist color-bleeding artifacts around the boundaries of important objects in videos. To overcome this problem, we develop a mixed expert block to extract semantic information for modeling the object boundaries of frames so that the semantic image prior can better guide the colorization process. In addition, we develop a multi-scale refinement block to progressively colorize frames in a coarse-to-fine manner. Extensive experimental results demonstrate that the proposed BiSTNet performs favorably against state-of-the-art methods on the benchmark datasets and real-world scenes. Moreover, the BiSTNet obtains one champion in NTIRE 2023 video colorization challenge (Kang et al. 2023).
Yixin Yang 0005, Jinshan Pan, Zhongzheng Peng, Xiaoyu Du 0002, Zhulin Tao, Jinhui Tang 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2023 Flow-Guided Transformer for Video Colorization
abstract
Video colorization aims to add color to black-and-white films. However, propagating color information to the whole video clip accurately is a challenging task. In this paper, we propose Flow-Guided Transformer for Video Colorization (FGTVC), consisting of a Global Motion Aggregation (GMA) module, Residual modules, Flow-Guided Attention blocks (FGAB) based on encoder and decoder, to exploit the information from the neighbor patch with high similarity for each video patch colorization. Specifically, we employ Transformer to capture the long-distance dependencies between frames and learn non-local self-similarity in the frame. To overcome the shortcomings of previous optical flow-based methods, FGAB enjoys the guidance of optical flow to sample elements from spatio-temporal adjacent frames when calculating self-attention. Experiments show that the proposed FGTVC has an outstanding performance than the state-of-the-art methods. In addition, comprehensive findings demonstrate the superiority of our framework in real-world video colorization tasks.
Yan Zhai, Zhulin Tao, Longquan Dai, He Wang 0054, Xianglin Huang, Lifang Yang
ICIP2
2023 UA-FedRec: Untargeted Attack on Federated News Recommendation
abstract
News recommendation is essential for personalized news distribution. Federated news recommendation, which enables collaborative model learning from multiple clients without sharing their raw data, is a promising approach for preserving users' privacy. However, the security of federated news recommendation is still unclear. In this paper, we study this problem by proposing an untargeted attack on federated news recommendation called UA-FedRec. By exploiting the prior knowledge of news recommendation and federated learning, UA-FedRec can effectively degrade the model performance with a small percentage of malicious clients. First, the effectiveness of news recommendation highly depends on user modeling and news modeling. We design a news similarity perturbation method to make representations of similar news farther and those of dissimilar news closer to interrupt news modeling, and propose a user model perturbation method to make malicious user updates in opposite directions of benign updates to interrupt user modeling. Second, updates from different clients are typically aggregated with a weighted average based on their sample sizes. We propose a quantity perturbation method to enlarge sample sizes of malicious clients in a reasonable range to amplify the impact of malicious updates. Extensive experiments on two real-world datasets show that UA-FedRec can effectively degrade the accuracy of existing federated news recommendation methods, even when defense is applied. Our study reveals a critical security issue in existing federated news recommendation systems and calls for research efforts to address the issue. Our code is available at https://github.com/yjw1029/UA-FedRec.
Jingwei Yi, Fangzhao Wu, Bin B. Zhu, Jing Yao 0003, Zhulin Tao, Guangzhong Sun, Xing Xie 0001
KDD5
2023 Self-Supervised Learning for Multimedia Recommendation
abstract
Learning representations for multimedia content is critical for multimedia recommendation. Current representation learning methods roughly fall into two groups: (1) using the historical interactions to create ID embeddings of users and items, and (2) treating multi-modal data as the side information of items to enrich their ID embeddings. Each user-item interaction offers the supervisory signal to optimize the representation learning by the traditional supervised learning paradigm. Due to the overlook of the multi-modal patterns ($e.g.$, co-occurrence of visual, acoustic, textual features in micro-videos a user saw before, and her behavioral features) hidden in the data, these methods are insufficient to create powerful representations and obtain satisfactory recommendation accuracy. To capture multi-modal patterns in the data itself, we go beyond the supervised learning paradigm, and incorporate the idea of self-supervised learning (SSL) into multimedia recommendation. Specifically, SSL consists of two components: (1) data augmentation upon multi-modal contents, where we design three operators — feature dropout (FD), feature masking (FM), feature fine and coarse spaces (FAC) — to generate multiple views of individual items; and (2) contrastive learning, which differentiates the views of an item from the others’ to distill additional supervisory signals. Clearly, SSL enables us to explore and exhibit the underlying relations among modalities, thereby resulting in powerful representations. We denote the generic framework by Self-supervised Learning-guided Multimedia Recommendation (SLMRec). Extensive experiments are performed on three real-world datasets, showing that SLMRec achieves significant improvements over several state-of-the-art baselines like LightGCN [1], MMGCN [2]. Further analysis shows how SSL affects recommendation performance.
Zhulin Tao, Xiaohao Liu, Yewei Xia, Xiang Wang 0010, Lifang Yang, Xianglin Huang, Tat-Seng Chua
IEEE Trans. Multim.1
2022 EliMRec: Eliminating Single-modal Bias in Multimedia Recommendation
abstract
The main idea of multimedia recommendation is to introduce the profile content of multimedia documents as an auxiliary, so as to endow recommenders with generalization ability and gain better performance. However, recent studies using non-uniform datasets roughly fuse single-modal features into multi-modal features and adopt the strategy of directly maximizing the likelihood of user preference scores, leading to the single-modal bias. Owing to the defect in architecture, there is still room for improvement for recent multimedia recommendation.
Xiaohao Liu, Zhulin Tao, Jiahong Shao, Lifang Yang, Xianglin Huang
ACM Multimedia2
2020 HoAFM: A High-order Attentive Factorization Machine for CTR Prediction
Zhulin Tao, Xiang Wang 0010, Xiangnan He 0001, Xianglin Huang, Tat-Seng Chua
Inf. Process. Manag.1
2020 MGAT: Multimodal Graph Attention Network for Recommendation
Zhulin Tao, Yinwei Wei, Xiang Wang 0010, Xiangnan He 0001, Xianglin Huang, Tat-Seng Chua
Inf. Process. Manag.1
2019 Better Word Representations with Word Weight
abstract
As a fundamental task of natural language processing, text classification has been widely used in various applications such as sentiment analysis and spam detection. In recent years, the continuous-valued word embedding learned by neural network attaches extensive attentions. Although word embedding achieves impressive results in capturing similarities and regularities between words, it fails to highlight important words for identifying text category. Such deficiency could be attenuated by word weight, which conveys word contribution in text categorization. Toward this end, we propose an effective text classification scheme by incorporating word weight into word embedding in this paper. Specifically, in order to enrich word representation, the bidirectional gated recurrent units (Bi-GRU) is first employed to grasp context information of words. Then the word weights yielded by term frequency (TF) are used to modulate the word representation of Bi-GRU for constructing text representation. Extensive experimental results on several large text datasets verify that the accuracy of our proposed text classification scheme outperforms the state-of-the-art ones.
Gege Song, Xianglin Huang, Gang Cao 0001, Zhulin Tao, Wei Liu 0084, Lifang Yang
MMSP4