Youcai Zhang

dblp:64/10215 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
7since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 4 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Image recognition and object detection · 24% Representation and self-supervised learning · 24% Vision and language · 24%

Topics — the 8 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Image recognition and object detection
image annotation
1.622025
Open-Set Image Tagging with Multi-Grained Text Supervision · ACM Multimedia 2025
Tag2Text: Guiding Vision-Language Model via Image Tagging · ICLR 2024
Computer vision › Vision and language
vision-language pretraining
1.322024
Tag2Text: Guiding Vision-Language Model via Image Tagging · ICLR 2024
IDEA: Increasing Text Diversity via Online Multi-Label Recognition for Vision-Language Pre-training · ACM Multimedia 2022
Machine learning › Representation and self-supervised learning
contrastive learning
0.612022
On the Efficacy of Small Self-Supervised Contrastive Models without Distillation Signals · AAAI 2022
Machine learning › Efficient and distributed learning
model compression
0.612022
On the Efficacy of Small Self-Supervised Contrastive Models without Distillation Signals · AAAI 2022
Machine learning › Learning paradigms
multi-label classification
0.612022
IDEA: Increasing Text Diversity via Online Multi-Label Recognition for Vision-Language Pre-training · ACM Multimedia 2022
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
0.412020
Prime-Aware Adaptive Distillation · ECCV (19) 2020
Computer vision › Vision and language
vision-language model
0.212024
Tag2Text: Guiding Vision-Language Model via Image Tagging · ICLR 2024
Machine learning › Representation and self-supervised learning
multimodal representation learning
0.212022
IDEA: Increasing Text Diversity via Online Multi-Label Recognition for Vision-Language Pre-training · ACM Multimedia 2022

Methods — techniques the papers use, named apart from their topics

vision-language alignment · 0.9multi-grained supervision · 0.9large language model · 0.9vision-language pretraining · 0.8image tagging · 0.8over-clustering mitigation · 0.6online tag extraction · 0.6multi-label learning · 0.6contrastive learning · 0.6prime-aware adaptive distillation · 0.4
YearPublicationVenuePosition
2026 Knowledge distillation from single-label to multi-label with class activation maps
Boyu Guan, Hengwei Liu, Youcai Zhang, Yuzhuo Qin, Xiaodong Gu 0001
Expert Syst. Appl.3
2025 Open-Set Image Tagging with Multi-Grained Text Supervision
abstract
This paper introduces the Recognize Anything Plus Model (RAM++), an open-set image tagging model effectively leveraging multi-grained text supervision. Previous approaches (e.g., CLIP) primarily utilize global text supervision paired with images, leading to sub-optimal performance in recognizing multiple individual semantic tags. In contrast, RAM++ seamlessly integrates individual tag supervision with global text supervision, all within a unified alignment framework. This integration not only ensures efficient recognition of predefined tag categories, but also enhances generalization capabilities for diverse open-set categories. Furthermore, RAM++ employs large language models (LLMs) to convert semantically constrained tag supervision into more expansive tag description supervision, thereby enriching the scope of open-set visual description concepts. Comprehensive evaluations on various image recognition benchmarks demonstrate RAM++ exceeds existing state-of-the-art (SOTA) open-set image tagging models on most aspects. Specifically, for predefined commonly used tag categories, RAM++ showcases 10.2 mAP and 15.4 mAP enhancements over CLIP on OpenImages and ImageNet. For open-set categories beyond predefined, RAM++ records improvements of 5.0 mAP and 6.4 mAP over CLIP and RAM respectively on OpenImages. For diverse human-object interaction phrases, RAM++ achieves 7.8 mAP and 4.7 mAP improvements on the HICO benchmark.
Yi-Jie Huang, Youcai Zhang, Rui Feng 0001, Yuejie Zhang, Yanchun Xie, Lei Zhang 0001
ACM Multimedia3
2024 Tag2Text: Guiding Vision-Language Model via Image Tagging
abstract
This paper presents Tag2Text, a vision language pre-training (VLP) framework, which introduces image tagging into vision-language models to guide the learning of visual-linguistic features. In contrast to prior works which utilize object tags either manually labeled or automatically detected with a limited detector, our approach utilizes tags parsed from its paired text to learn an image tagger and meanwhile provides guidance to vision-language models. Given that, Tag2Text can utilize large-scale annotation-free image tags in accordance with image-text pairs, and provides more diverse tag categories beyond objects. Strikingly, Tag2Text showcases the ability of a foundational image tagging model, with superior zero-shot performance even comparable to full supervision manner. Moreover, by leveraging tagging guidance, Tag2Text effectively enhances the performance of vision-language models on both generation-based and alignment-based tasks. Across a wide range of downstream benchmarks, Tag2Text achieves state-of-the-art results with similar model sizes and data scales, demonstrating the efficacy of the proposed tagging guidance.
Youcai Zhang, Jinyu Ma, Rui Feng 0001, Yuejie Zhang, Yandong Guo, Lei Zhang 0001
ICLR2
2022 On the Efficacy of Small Self-Supervised Contrastive Models without Distillation Signals
abstract
It is a consensus that small models perform quite poorly under the paradigm of self-supervised contrastive learning. Existing methods usually adopt a large off-the-shelf model to transfer knowledge to the small one via distillation. Despite their effectiveness, distillation-based methods may not be suitable for some resource-restricted scenarios due to the huge computational expenses of deploying a large model. In this paper, we study the issue of training self-supervised small models without distillation signals. We first evaluate the representation spaces of the small models and make two non-negligible observations: (i) the small models can complete the pretext task without overfitting despite their limited capacity and (ii) they universally suffer the problem of over clustering. Then we verify multiple assumptions that are considered to alleviate the over-clustering phenomenon. Finally, we combine the validated techniques and improve the baseline performances of five small architectures with considerable margins, which indicates that training small self-supervised contrastive models is feasible even without distillation signals. The code is available at https://github.com/WOWNICE/ssl-small.
Haizhou Shi, Youcai Zhang, Siliang Tang, Wenjie Zhu 0003, Yandong Guo, Yueting Zhuang
AAAI2
2022 IDEA: Increasing Text Diversity via Online Multi-Label Recognition for Vision-Language Pre-training
abstract
Vision-Language Pre-training (VLP) with large-scale image-text pairs has demonstrated superior performance in various fields. However, the image-text pairs co-occurrent on the Internet typically lack explicit alignment information, which is suboptimal for VLP. Existing methods proposed to adopt an off-the-shelf object detector to utilize additional image tag information. However, the object detector is time-consuming and can only identify the pre-defined object categories, limiting the model capacity. Inspired by the observation that the texts incorporate incomplete fine-grained image information, we introduce IDEA, which stands for increasing text diversity via online multi-label recognition for VLP. IDEA shows that multi-label learning with image tags extracted from the texts can be jointly optimized during VLP. Moreover, IDEA can identify valuable image tags online to provide more explicit textual supervision. Comprehensive experiments demonstrate that IDEA can significantly boost the performance on multiple downstream datasets with a small extra computational cost.
Youcai Zhang, Ying Cheng 0005, Rui Feng 0001, Yuejie Zhang, Yandong Guo
ACM Multimedia2
2021 IIT-GAT: Instance-level image transformation via unsupervised generative attention networks with disentangled representations
Ming-Wen Shao, Youcai Zhang, Wangmeng Zuo, Deyu Meng
Knowl. Based Syst.2
2021 DMDIT: Diverse multi-domain image-to-image translation
Ming-Wen Shao, Youcai Zhang, Huan Liu 0012, Chao Wang 0102, Xun Shao
Knowl. Based Syst.2
2020 Prime-Aware Adaptive Distillation
Youcai Zhang, Zhonghao Lan, Yuchen Dai, Fangao Zeng
ECCV (19)1
2019 Adversarial Learning for Cross-Modal Retrieval with Wasserstein Distance
Qingrong Cheng, Youcai Zhang, Xiaodong Gu 0001
ICONIP (1)2
2018 Two-Stream Convolutional Neural Network for Multimodal Matching
Youcai Zhang, Yiwei Gu, Xiaodong Gu 0001
ICANN (1)1
2018 Cross-modal Metric Learning with Graph Embedding
abstract
Metric learning with neural networks has exhibited promising improvements in representation learning. Yet cross-modal retrieval poses a unique challenge to metric learning: how to compute the distance across different modalities such as image and text. Existing neural network based methods tend to establish two branches for images and texts respectively to bridge the modal gap. Also, most of them cannot fully exploit the structure embedded in the multimodal data. This paper introduces embedding layer to provide cross-modal shared representation with non-linearity and reformulates the cross-modal retrieval problem as a graph embedding problem by constructing a multimodal graph. To learn the graph embedding, training pairs and triplets are uniformly generated from random walk sequences on the graph. Then graph pair and triplet constraints are imposed on the embedding layer for structure preservation. Meanwhile, a classifier is trained with labeled data to ensure the learned embedding is coupled with semantic information. For optimization, graph pair and triplet constraints are integrated into a unified multi-task learning with the supervised classifier. Experimental results on the Wiki and NUS-WIDE datasets demonstrate the effectiveness and superiority of the learned embedding for cross-modal retrieval.
Youcai Zhang
IJCNN1
2017 Graph Embedding Learning for Cross-Modal Information Retrieval
Youcai Zhang, Xiaodong Gu 0001
ICONIP (3)1