VLDB 2026 Research / reviewers in the wild / expert
Fan Zhang 0111
dblp:21/3626-111
· DBLP profile ↗
14ranked-venue papers
10as first author
14since 2021 · last 2026
0000-0001-5250-7258ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 9 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | S³-MSD: Large Vision-Language Model for Explainable and Generalizable Multi-modal Sarcasm DetectionabstractMultimodal sarcasm detection (MSD) aims to identify sarcasm polarity from diverse modalities (i.e., image–text pairs), a task that has received increasing attention. While significant progress has been made, existing approaches still face two major issues: lack of explainability and weak generalizability. In this paper, we introduce a new large vision–language model (LVLM) dubbed S³-MSD for explainable and generalizable MSD through three key components. For explainability, we develop (1) a self-training paradigm that automatically bootstraps answers with explanations, and (2) a self-calibrating mechanism that rectifies flawed explanations. For generalizability, we design (3) a self-focusing module that amplifies visual semantic entities through preference optimization, thereby mitigating textual over-reliance. Experimental results on both in-distribution and out-of-distribution (OOD) benchmarks demonstrate that S³-MSD consistently outperforms state-of-the-art methods in detection performance. Furthermore, the proposed S³-MSD provides persuasive explanations, as verified by both quantitative metrics and human evaluations. Zhihong Zhu 0001, Fan Zhang 0111, Yunyan Zhang, Jinghan Sun, Guimin Hu, Hao Wu 0094, Yuyan Chen, Xian Wu 0001 |
AAAI | 2 |
| 2026 | CMID: Towards Medical Visual Question Answering via Contrastive Mutual Information DecodingabstractMedical Visual Question Answering (Med-VQA) aims to generate accurate answers for clinical questions grounded in medical images, which has attracted increasing research attention due to its potential to streamline diagnostics and reduce clinical burden. Recent advances in Large Vision-Language Models (LVLMs) have shown great promise for Med-VQA, but still suffer from two inference-time issues: (1) attention shift, where the LVLM over-relies on textual priors; and (2) attention dispersion, where it fails to focus on critical diagnostic regions. To tackle these issues, we propose Contrastive Mutual Information Decoding (CMID), a training-free inference-time intervention grounded in information theory for Med-VQA. Concretely, CMID first identifies the Principal Focus Area (PFA) from decoder attention maps, then constructs focus-preserving and focus-excluding views to derive dual contrastive signals that simultaneously amplify salient visual cues and suppress background noise. Crucially, these corrective signals are adaptively scaled by a reliability-gated self-correction mechanism, based on the distributional shift induced by the PFA. Extensive experiments on three Med-VQA benchmarks demonstrate the effectiveness of CMID. Further analyses showcase its robust generalizability across diverse medical architectures and tasks. Zhihong Zhu 0001, Yunyan Zhang, Fan Zhang 0111, Xian Wu 0001 |
AAAI | 3 |
| 2026 | 3D landmark detection on human point clouds: A benchmark and a dual cascade point transformer framework
Fan Zhang 0111, Shuyi Mao, Xiaojiang Peng |
Expert Syst. Appl. | 1 |
| 2025 | DREAM: Decoupled Discriminative Learning with Bigraph-aware Alignment for Semi-supervised 2D-3D Cross-modal RetrievalabstractWith the burst of big data, 2D-3D cross-modal retrieval has received increasing attention, which aims to retrieve relevant data from one modality given the query from the other modality. In this paper, we study an underexplored yet practical problem of semi-supervised 2D-3D cross-modal retrieval, which could suffer from serious label scarcity in real-world applications. Moreover, the huge heterogeneous gap could deteriorate the process of learning from unlabeled data. In this work, we propose a novel approach named Decoupled Discriminative Learning with Bigraph-aware Alignment (DREAM) for semi-supervised 2D-3D cross-modal retrieval. The core of our DREAM is to decouple the label prediction and reliability measurement processes to reduce overconfident samples in discriminative learning. In particular, we enhance a label prediction module with label propagation from labeled samples and additionally introduce a reliability measurement module to learn the scores of predicted labels. To reduce class-related bias, we compare reliability scores with class-specific adaptive thresholds to identify samples for additional learning. In addition, negative labels are estimated for unselected samples, which guides soft semantic learning to make the best use of all the information. To further minimize the heterogeneous gap, we build a bigraph graph that connects cross-modal similar examples and then conduct learning to cluster with most edges kept for alignment. Extensive experiments on several benchmark datasets validate the superiority of the proposed DREAM. Fan Zhang 0111, Changhu Wang, Zebang Cheng, Xiaojiang Peng, Dongjie Wang 0001, Yijia Xiao, Chong Chen 0002, Xian-Sheng Hua 0001, Xiao Luo 0001 |
AAAI | 1 |
| 2025 | A Survey on Foundation Language Models for Single-cell BiologyabstractFan Zhang, Hao Chen, Zhihong Zhu, Ziheng Zhang, Zhenxi Lin, Ziyue Qiao, Yefeng Zheng, Xian Wu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Fan Zhang 0111, Hao Chen 0011, Zhihong Zhu 0001, Zhenxi Lin, Ziyue Qiao, Yefeng Zheng 0001, Xian Wu 0001 |
ACL (1) | 1 |
| 2025 | CellVerse: Do Large Language Models Really Understand Cell Biology?abstractRecent studies have demonstrated the feasibility of modeling single-cell data as natural languages and the potential of leveraging powerful large language models (LLMs) for understanding cell biology. However, a comprehensive evaluation of LLMs' performance on language-driven single-cell analysis tasks still remains unexplored. Motivated by this challenge, we introduce CellVerse, a unified language-centric question-answering benchmark that integrates four types of single-cell multi-omics data and encompasses three hierarchical levels of single-cell analysis tasks: cell type annotation (cell-level), drug response prediction (drug-level), and perturbation analysis (gene-level). Going beyond this, we systematically evaluate the performance across 14 open-source and closed-source LLMs ranging 160M $\rightarrow$ 671B on CellVerse. Remarkably, the experimental results reveal: (1) Existing specialist models (C2S-Pythia) fail to make reasonable decisions across all sub-tasks within CellVerse, while generalist models such as Qwen, Llama, GPT, and DeepSeek family models exhibit preliminary understanding capabilities within the realm of cell biology. (2) The performance of current LLMs falls short of expectations and has substantial room for improvement. Notably, in the widely studied drug response prediction task, none of the evaluated LLMs demonstrate significant performance improvement over random guessing. CellVerse offers the first large-scale empirical demonstration that significant challenges still remain in applying LLMs to cell biology. By introducing CellVerse, we lay the foundation for advancing cell biology through natural languages and hope this paradigm could facilitate next-generation single-cell analysis. Project Page: https://cellverse-cuhk.github.io Fan Zhang 0111, Tianyu Liu 0001, Zhihong Zhu 0001, Hao Wu 0094, Haixin Wang 0003, Yefeng Zheng 0001, Kun Wang 0056, Xian Wu 0001, Pheng-Ann Heng |
NeurIPS | 1 |
| 2025 | LEAF: Unveiling two sides of the same coin in semi-supervised facial expression recognition
Fan Zhang 0111, Zhi-Qi Cheng, Jian Zhao 0006, Xiaojiang Peng, Xuelong Li 0001 |
Comput. Vis. Image Underst. | 1 |
| 2025 | Facial Action Units as a Joint Dataset Training Bridge for Facial Expression RecognitionabstractLabel biases in facial expression recognition (FER) datasets, caused by annotators' subjectivity, pose challenges in improving the performance of target datasets when auxiliary labeled data are used. Moreover, training with multiple datasets can lead to visible degradations in the target dataset. To address these issues, we propose a novel framework called the AU-aware Vision Transformer (AU-ViT), which leverages unified action unit (AU) information and discards expression annotations of auxiliary data. AU-ViT integrates an elaborately designed AU branch in the middle part of a master ViT to enhance representation learning during training. Through qualitative and quantitative analyses, we demonstrate that AU-ViT effectively captures expression regions and is robust to real-world occlusions. Additionally, we observe that AU-ViT also yields performance improvements on the target dataset, even without auxiliary data, by utilizing pseudo AU labels. Our AU-ViT achieves performances superior to, or comparable to, that of the state-of-the-art methods on FERPlus, RAFDB, AffectNet, LSD and the other three occlusion test datasets. Shuyi Mao, Xinpeng Li 0004, Fan Zhang 0111, Xiaojiang Peng, Yang Yang 0002 |
IEEE Trans. Multim. | 3 |
| 2024 | Fine-grained Prototypical Voting with Heterogeneous Mixup for Semi-supervised 2D-3D Cross-modal RetrievalabstractThis paper studies the problem of semi-supervised 2D-3D retrieval, which aims to align both labeled and unla-beled 2D and 3D data into the same embedding space. The problem is challenging due to the complicated heteroge-neous relationships between 2D and 3D data. Moreover, label scarcity in real-world applications hinders from gen-erating discriminative representations. In this paper, we propose a semi-supervised approach named Fine-grained Prototypcical ⊻oting with Heterogeneous Mixup (FIVE), which maps both 2D and 3D data into a common embed-ding space for cross-modal retrieval. Specifically, we gen-erate fine-grained prototypes to model intra-class variation for both 2D and 3D data. Then, considering each unlabeled sample as a query, we retrieve relevant prototypes to vote for reliable and robust pseudo-labels, which serve as guid-ance for discriminative learning under label scarcity. Fur-thermore, to bridge the semantic gap between two modali-ties, we mix cross-modal pairs with similar semantics in the embedding space and then perform similarity learning for cross-modal discrepancy reduction in a soft manner. The whole FIVE is optimized with the consideration of sharp-ness to mitigate the impact of potential label noise. Exten-sive experiments on benchmark datasets validate the supe-riority of FIVE compared with a range of baselines in differ-ent settings. On average, FIVE outperforms the second-best approach by 4.74% on 3D MNIST, 12.94% on ModelNet10, and 22.10% on ModelNet40. Fan Zhang 0111, Xian-Sheng Hua 0001, Chong Chen 0002, Xiao Luo 0001 |
CVPR | 1 |
| 2024 | DEMO: A Statistical Perspective for Efficient Image-Text MatchingabstractFan Zhang, Xian-Sheng Hua, Chong Chen, Xiao Luo. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Fan Zhang 0111, Xian-Sheng Hua 0001, Chong Chen 0002, Xiao Luo 0001 |
NAACL-HLT | 1 |
| 2024 | Semi-supervised Knowledge Transfer Across Multi-omic Single-cell DataabstractKnowledge transfer between multi-omic single-cell data aims to effectively transfer cell types from scRNA-seq data to unannotated scATAC-seq data. Several approaches aim to reduce the heterogeneity of multi-omic data while maintaining the discriminability of cell types with extensive annotated data. However, in reality, the cost of collecting both a large amount of labeled scRNA-seq data and scATAC-seq data is expensive. Therefore, this paper explores a practical yet underexplored problem of knowledge transfer across multi-omic single-cell data under cell type scarcity. To address this problem, we propose a semi-supervised knowledge transfer framework named Dual label scArcity elimiNation with Cross-omic multi-samplE Mixup (DANCE). To overcome the label scarcity in scRNA-seq data, we generate pseudo-labels based on optimal transport and merge them into the labeled scRNA-seq data. Moreover, we adopt a divide-and-conquer strategy which divides the scATAC-seq data into source-like and target-specific data. For source-like samples, we employ consistency regularization with random perturbations while for target-specific samples, we select a few candidate labels and progressively eliminate incorrect cell types from the label set for additional supervision. Next, we generate virtual scRNA-seq samples with multi-sample Mixup based on the class-wise similarity to reduce cell heterogeneity. Extensive experiments on many benchmark datasets suggest the superiority of our DANCE over a series of state-of-the-art methods. Fan Zhang 0111, Tianyu Liu 0005, Xiaojiang Peng, Chong Chen 0002, Xian-Sheng Hua 0001, Xiao Luo 0001, Hongyu Zhao 0003 |
NeurIPS | 1 |
| 2024 | HOPE: A Hierarchical Perspective for Semi-Supervised 2D-3D Cross-Modal RetrievalabstractWith the emergence of AI generated content, cross-modal retrieval of 2D and 3D data has obtained increasing research attention. In practical applications, massive amounts of 2D and 3D data need expensive annotation, which would make labels scarce. Even worse, complicated heterogeneous relationships between 2D and 3D data make the problem more challenging. In this research, we study the problem of semi-supervised 2D and 3D cross-modal retrieval and provide a novel method namedHierarchical Alignment with AmbiguousPseudo-labeling (HOPE) for this problem. The core of HOPE is to align two modalities in the common space from a hierarchical perspective. Specifically, HOPE not only enforces each sample to approach its respective modality-invariant anchors from an individual view, but also measures both prototypes and distribution for both modalities for discrepancy reduction from a group view. To handle label scarcity with limited error accumulation, HOPE employs two branches of perturbed networks to generate ambiguous candidates, which guides the cross-branch supervision using a margin-based ranking objective. In addition, we retrieve reliable unlabeled samples for each anchor with curriculum learning and class balance, which are added into labeled datasets to clear ambiguity. Extensive experiments on various benchmark datasets validate the superiority of the proposed HOPE. Fan Zhang 0111, Hang Zhou 0008, Xian-Sheng Hua 0001, Chong Chen 0002, Xiao Luo 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | FATE: Learning Effective Binary Descriptors With Group FairnessabstractHashing has received significant interest in large-scale data retrieval due to its outstanding computational efficiency. Of late, numerous deep hashing approaches have emerged, which have obtained impressive performance. However, these approaches can contain ethical risks during image retrieval. To address this, we are the first to study the problem of group fairness within learning to hash and introduce a novel method termed Fairness-aware Hashing with Mixture of Experts (FATE). Specifically, FATE leverages the mixture-of-experts framework as the hashing network, where each expert contributes knowledge from an individual viewpoint, followed by aggregation using the gating mechanism. This strongly enhances the model capability, facilitating the generation of both discriminative and unbiased binary descriptors. We also incorporate fairness-aware contrastive learning, combining sensitive labels with feature similarities to ensure unbiased hash code learning. Furthermore, an adversarial learning objective condition on both deep features and hash codes is employed to further eliminate group biases. Extensive experiments on several benchmark datasets validate the superiority of the proposed FATE compared with various state-of-the-art approaches. Fan Zhang 0111, Chong Chen 0002, Xian-Sheng Hua 0001, Xiao Luo 0001 |
IEEE Trans. Image Process. | 1 |
| 2023 | Semi-Supervised Multimodal Emotion Recognition with Expression MAEabstractThe Multimodal Emotion Recognition (MER 2023) challenge aims to recognize emotion with audio, language, and visual signals, facilitating innovative technologies of affective computing. This paper presents our submission approach on the Semi-Supervised Learning Sub-Challenge (MER-SEMI). First, with large-scale unlabeled emotional videos, we train both image-based and video-based Masked Autoencoders to extract visual features, which termed as expression MAE (expMAE) for simplicity. The expMAE features are found to be largely complementary with other official baseline features. Second, since there is only a few labeled data, we use a classifier to generate pseudo labels for unlabeled videos which have high confidence for a certain category. In addition, we also explore several advanced large models for cross-feature extraction like CLIP, and apply factorized bilinear pooling (FBP) for multimodal feature fusion. Our methods finally achieved 88.55% in F1 score on MER-SEMI, ranking second place among all participating teams. Zebang Cheng, Zhaoru Chen, Xiang Li 0130, Shuyi Mao, Fan Zhang 0111, Daijun Ding, Bowen Zhang 0005, Xiaojiang Peng |
ACM Multimedia | 6 |