Jiarui Ouyang

dblp:371/5303 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2026
0009-0001-4716-4235ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Vision and language · 87% Representation and self-supervised learning · 13%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 4 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language
image captioning
0.912025
Zero-Shot Image Captioning with Multi-type Entity Representations · AAAI 2025
Computer vision › Vision and language › image captioning › low-shot image captioning
zero-shot image captioning
0.912025
Zero-Shot Image Captioning with Multi-type Entity Representations · AAAI 2025
Information retrieval › search engines › semantic search
entity retrieval
0.912025
Zero-Shot Image Captioning with Multi-type Entity Representations · AAAI 2025
Machine learning › Representation and self-supervised learning
contrastive learning
0.312025
Zero-Shot Image Captioning with Multi-type Entity Representations · AAAI 2025

Methods — techniques the papers use, named apart from their topics

contrastive learning · 1.7GPT-2 · 1.7CLIP · 1.7
YearPublicationVenuePosition
2026 GenAR: Next-scale autoregressive generation for spatial gene expression prediction
Jiarui Ouyang, Yihui Wang 0002, Yihang Gao, Yingxue Xu, Shu Yang 0004, Hao Chen 0011
Medical Image Anal.1
2025 Zero-Shot Image Captioning with Multi-type Entity Representations
abstract
As data and computational resources continue to expand, incorporating a variety of knowledge during the pre-training phase enhances large models, providing them with strong zero-shot capabilities. Due to the alignment of modal features by visual language models, zero-shot image captioning no longer necessitates pre-training on paired image-text labeled data, enabling accurate text description generation for images not encountered before. While recent research focuses on methods utilizing entity retrieval as anchors to bridge the gap between different modalities, these approaches often fall short of thoroughly analyzing the impact of entity retrieval recall on the zero-shot generation capabilities. To address this issue, we propose MERCap, a zero-shot image captioning method employing Multi-type Entity representation Retrieval. More specifically, we first approximate image representation using the CLIP representation of text and Gaussian noise to address the modality gap. Then, we train a GPT-2 decoder to reconstruct text using entities as hard prompts and CLIP representations as soft prompts. Additionally, we construct a domain-specific entity set, assigning multiple representations to each entity and refining their representation vectors through contrastive learning. During inference, we retrieve entities and input them into the decoder to generate corresponding captions. Extensive experiments validate that our approach is efficient, achieving a new state-of-the-art level in cross-domain captioning and demonstrating strong competitiveness in in-domain captioning compared to existing methods.
Delong Zeng, Ying Shen 0001, Man Lin, Zihao Yi, Jiarui Ouyang
AAAI5
2024 MSIF: Multi-source Information Fusion for Financial Question Answering
Man Lin, Delong Zeng, Jiarui Ouyang, Ying Shen 0001
ICANN (9)3
2024 Explore the Textual Perception Ability on the Images for Multimodal Large Language Models
Jiayi Kuang, Jiarui Ouyang, Ying Shen 0001
NLPCC (5)2