VLDB 2026 Research / reviewers in the wild / expert
Youcai Zhang
dblp:64/10215
· DBLP profile ↗
12ranked-venue papers
4as first author
7since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 4 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Image recognition and object detection · 24% Representation and self-supervised learning · 24% Vision and language · 24% |
Topics — the 8 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Image recognition and object detection
image annotation |
1.6 | 2 | 2025 | Open-Set Image Tagging with Multi-Grained Text Supervision · ACM Multimedia 2025 Tag2Text: Guiding Vision-Language Model via Image Tagging · ICLR 2024 |
Computer vision › Vision and language
vision-language pretraining |
1.3 | 2 | 2024 | Tag2Text: Guiding Vision-Language Model via Image Tagging · ICLR 2024 IDEA: Increasing Text Diversity via Online Multi-Label Recognition for Vision-Language Pre-training · ACM Multimedia 2022 |
Machine learning › Representation and self-supervised learning
contrastive learning |
0.6 | 1 | 2022 | On the Efficacy of Small Self-Supervised Contrastive Models without Distillation Signals · AAAI 2022 |
Machine learning › Efficient and distributed learning
model compression |
0.6 | 1 | 2022 | On the Efficacy of Small Self-Supervised Contrastive Models without Distillation Signals · AAAI 2022 |
Machine learning › Learning paradigms
multi-label classification |
0.6 | 1 | 2022 | IDEA: Increasing Text Diversity via Online Multi-Label Recognition for Vision-Language Pre-training · ACM Multimedia 2022 |
Machine learning › Efficient and distributed learning › model compression
knowledge distillation |
0.4 | 1 | 2020 | Prime-Aware Adaptive Distillation · ECCV (19) 2020 |
Computer vision › Vision and language
vision-language model |
0.2 | 1 | 2024 | Tag2Text: Guiding Vision-Language Model via Image Tagging · ICLR 2024 |
Machine learning › Representation and self-supervised learning
multimodal representation learning |
0.2 | 1 | 2022 | IDEA: Increasing Text Diversity via Online Multi-Label Recognition for Vision-Language Pre-training · ACM Multimedia 2022 |
Methods — techniques the papers use, named apart from their topics
vision-language alignment · 0.9multi-grained supervision · 0.9large language model · 0.9vision-language pretraining · 0.8image tagging · 0.8over-clustering mitigation · 0.6online tag extraction · 0.6multi-label learning · 0.6contrastive learning · 0.6prime-aware adaptive distillation · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Knowledge distillation from single-label to multi-label with class activation maps
Boyu Guan, Hengwei Liu, Youcai Zhang, Yuzhuo Qin, Xiaodong Gu 0001 |
Expert Syst. Appl. | 3 |
| 2025 | Open-Set Image Tagging with Multi-Grained Text SupervisionabstractThis paper introduces the Recognize Anything Plus Model (RAM++), an open-set image tagging model effectively leveraging multi-grained text supervision. Previous approaches (e.g., CLIP) primarily utilize global text supervision paired with images, leading to sub-optimal performance in recognizing multiple individual semantic tags. In contrast, RAM++ seamlessly integrates individual tag supervision with global text supervision, all within a unified alignment framework. This integration not only ensures efficient recognition of predefined tag categories, but also enhances generalization capabilities for diverse open-set categories. Furthermore, RAM++ employs large language models (LLMs) to convert semantically constrained tag supervision into more expansive tag description supervision, thereby enriching the scope of open-set visual description concepts. Comprehensive evaluations on various image recognition benchmarks demonstrate RAM++ exceeds existing state-of-the-art (SOTA) open-set image tagging models on most aspects. Specifically, for predefined commonly used tag categories, RAM++ showcases 10.2 mAP and 15.4 mAP enhancements over CLIP on OpenImages and ImageNet. For open-set categories beyond predefined, RAM++ records improvements of 5.0 mAP and 6.4 mAP over CLIP and RAM respectively on OpenImages. For diverse human-object interaction phrases, RAM++ achieves 7.8 mAP and 4.7 mAP improvements on the HICO benchmark. Yi-Jie Huang, Youcai Zhang, Rui Feng 0001, Yuejie Zhang, Yanchun Xie, Lei Zhang 0001 |
ACM Multimedia | 3 |
| 2024 | Tag2Text: Guiding Vision-Language Model via Image TaggingabstractThis paper presents Tag2Text, a vision language pre-training (VLP) framework, which introduces image tagging into vision-language models to guide the learning of visual-linguistic features. In contrast to prior works which utilize object tags either manually labeled or automatically detected with a limited detector, our approach utilizes tags parsed from its paired text to learn an image tagger and meanwhile provides guidance to vision-language models. Given that, Tag2Text can utilize large-scale annotation-free image tags in accordance with image-text pairs, and provides more diverse tag categories beyond objects. Strikingly, Tag2Text showcases the ability of a foundational image tagging model, with superior zero-shot performance even comparable to full supervision manner. Moreover, by leveraging tagging guidance, Tag2Text effectively enhances the performance of vision-language models on both generation-based and alignment-based tasks. Across a wide range of downstream benchmarks, Tag2Text achieves state-of-the-art results with similar model sizes and data scales, demonstrating the efficacy of the proposed tagging guidance. Youcai Zhang, Jinyu Ma, Rui Feng 0001, Yuejie Zhang, Yandong Guo, Lei Zhang 0001 |
ICLR | 2 |
| 2022 | On the Efficacy of Small Self-Supervised Contrastive Models without Distillation SignalsabstractIt is a consensus that small models perform quite poorly under the paradigm of self-supervised contrastive learning. Existing methods usually adopt a large off-the-shelf model to transfer knowledge to the small one via distillation. Despite their effectiveness, distillation-based methods may not be suitable for some resource-restricted scenarios due to the huge computational expenses of deploying a large model. In this paper, we study the issue of training self-supervised small models without distillation signals. We first evaluate the representation spaces of the small models and make two non-negligible observations: (i) the small models can complete the pretext task without overfitting despite their limited capacity and (ii) they universally suffer the problem of over clustering. Then we verify multiple assumptions that are considered to alleviate the over-clustering phenomenon. Finally, we combine the validated techniques and improve the baseline performances of five small architectures with considerable margins, which indicates that training small self-supervised contrastive models is feasible even without distillation signals. The code is available at https://github.com/WOWNICE/ssl-small. Haizhou Shi, Youcai Zhang, Siliang Tang, Wenjie Zhu 0003, Yandong Guo, Yueting Zhuang |
AAAI | 2 |
| 2022 | IDEA: Increasing Text Diversity via Online Multi-Label Recognition for Vision-Language Pre-trainingabstractVision-Language Pre-training (VLP) with large-scale image-text pairs has demonstrated superior performance in various fields. However, the image-text pairs co-occurrent on the Internet typically lack explicit alignment information, which is suboptimal for VLP. Existing methods proposed to adopt an off-the-shelf object detector to utilize additional image tag information. However, the object detector is time-consuming and can only identify the pre-defined object categories, limiting the model capacity. Inspired by the observation that the texts incorporate incomplete fine-grained image information, we introduce IDEA, which stands for increasing text diversity via online multi-label recognition for VLP. IDEA shows that multi-label learning with image tags extracted from the texts can be jointly optimized during VLP. Moreover, IDEA can identify valuable image tags online to provide more explicit textual supervision. Comprehensive experiments demonstrate that IDEA can significantly boost the performance on multiple downstream datasets with a small extra computational cost. Youcai Zhang, Ying Cheng 0005, Rui Feng 0001, Yuejie Zhang, Yandong Guo |
ACM Multimedia | 2 |
| 2021 | IIT-GAT: Instance-level image transformation via unsupervised generative attention networks with disentangled representations
Ming-Wen Shao, Youcai Zhang, Wangmeng Zuo, Deyu Meng |
Knowl. Based Syst. | 2 |
| 2021 | DMDIT: Diverse multi-domain image-to-image translation
Ming-Wen Shao, Youcai Zhang, Huan Liu 0012, Chao Wang 0102, Xun Shao |
Knowl. Based Syst. | 2 |
| 2020 | Prime-Aware Adaptive Distillation
Youcai Zhang, Zhonghao Lan, Yuchen Dai, Fangao Zeng |
ECCV (19) | 1 |
| 2019 | Adversarial Learning for Cross-Modal Retrieval with Wasserstein Distance
Qingrong Cheng, Youcai Zhang, Xiaodong Gu 0001 |
ICONIP (1) | 2 |
| 2018 | Two-Stream Convolutional Neural Network for Multimodal Matching
Youcai Zhang, Yiwei Gu, Xiaodong Gu 0001 |
ICANN (1) | 1 |
| 2018 | Cross-modal Metric Learning with Graph EmbeddingabstractMetric learning with neural networks has exhibited promising improvements in representation learning. Yet cross-modal retrieval poses a unique challenge to metric learning: how to compute the distance across different modalities such as image and text. Existing neural network based methods tend to establish two branches for images and texts respectively to bridge the modal gap. Also, most of them cannot fully exploit the structure embedded in the multimodal data. This paper introduces embedding layer to provide cross-modal shared representation with non-linearity and reformulates the cross-modal retrieval problem as a graph embedding problem by constructing a multimodal graph. To learn the graph embedding, training pairs and triplets are uniformly generated from random walk sequences on the graph. Then graph pair and triplet constraints are imposed on the embedding layer for structure preservation. Meanwhile, a classifier is trained with labeled data to ensure the learned embedding is coupled with semantic information. For optimization, graph pair and triplet constraints are integrated into a unified multi-task learning with the supervised classifier. Experimental results on the Wiki and NUS-WIDE datasets demonstrate the effectiveness and superiority of the learned embedding for cross-modal retrieval. Youcai Zhang |
IJCNN | 1 |
| 2017 | Graph Embedding Learning for Cross-Modal Information Retrieval
Youcai Zhang, Xiaodong Gu 0001 |
ICONIP (3) | 1 |