Jianbo Ouyang

dblp:304/4448 · DBLP profile ↗
← Back
3ranked-venue papers
2as first author
3since 2021 · last 2025
0000-0001-5464-7668ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
1 paper
Visual content generation and editing · 100%
Databases, data mining, and information retrieval
2 papers
Information retrieval · 100%
Artificial intelligence
1 paper
Generative modeling · 100%

Topics — the 8 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval › image retrieval › web image search
image re-ranking
1.022021
Collaborative Image Relevance Learning for Visual Re-Ranking · IEEE Trans. Multim. 2021
Contextual Similarity Aggregation with Self-attention for Visual Re-ranking · NeurIPS 2021
Visual content generation and editing › image generation › text-to-image generation
identity customization
0.912025
Inv-Adapter: ID Customization Generation via Image Inversion and Lightweight Parameter Adapter · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Visual content generation and editing › image generation
text-to-image generation
0.912025
Inv-Adapter: ID Customization Generation via Image Inversion and Lightweight Parameter Adapter · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Information retrieval
image retrieval
0.512021
Collaborative Image Relevance Learning for Visual Re-Ranking · IEEE Trans. Multim. 2021
Information retrieval › reranking
search result re-ranking
0.512021
Contextual Similarity Aggregation with Self-attention for Visual Re-ranking · NeurIPS 2021
Machine learning › Generative modeling
diffusion model
0.312025
Inv-Adapter: ID Customization Generation via Image Inversion and Lightweight Parameter Adapter · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Information retrieval › image retrieval
content-based image retrieval
0.112021
Contextual Similarity Aggregation with Self-attention for Visual Re-ranking · NeurIPS 2021
Information retrieval › query reformulation
query expansion
0.112021
Collaborative Image Relevance Learning for Visual Re-Ranking · IEEE Trans. Multim. 2021

Methods — techniques the papers use, named apart from their topics

diffusion model · 1.7attention adapter · 1.7DDIM inversion · 1.7weighted MSE loss · 0.5transformer encoder · 0.5self-attention · 0.5data augmentation · 0.5correlation matrix · 0.5convolutional neural network · 0.5
YearPublicationVenuePosition
2025 Inv-Adapter: ID Customization Generation via Image Inversion and Lightweight Parameter Adapter
abstract
The remarkable advancement in text-to-image generation models significantly boosts the research in ID customization generation. However, existing personalization methods cannot simultaneously satisfy high-fidelity and low-costs requirements. Their main bottleneck lies in the additional prompt image encoder (i.e., CLIP vision encoder), which produces weak alignment signals with the text-to-image model that may lose face information and is not well 'absorbed' by the text-to-image model. Towards this end, we propose Inv-Adapter, which first introduces a more reasonable and efficient token representation of ID image features and introduces a lightweight parameter adaptor to inject ID features. Specifically, our Inv-Adapter extracts diffusion-domain representations of ID images utilizing a pre-trained text-to-image model via DDIM image inversion, without an additional image encoder. Benefiting from the high alignment of the extracted ID prompt features and the intermediate features of the text-to-image model, we then introduce a lightweight attention adapter to embed them efficiently into the base text-to-image model. We conduct extensive experiments on different text-to-image models to assess ID fidelity, generation loyalty, speed, training costs, model scale and generalization ability in scenarios of general object, all of which show that the proposed Inv-Adapter is highly competitive in ID customization generation and model scale.
Peng Xing, Ning Wang 0020, Jianbo Ouyang, Zechao Li
IEEE Trans. Pattern Anal. Mach. Intell.3
2021 Contextual Similarity Aggregation with Self-attention for Visual Re-ranking
abstract
In content-based image retrieval, the first-round retrieval result by simple visual feature comparison may be unsatisfactory, which can be refined by visual re-ranking techniques. In image retrieval, it is observed that the contextual similarity among the top-ranked images is an important clue to distinguish the semantic relevance. Inspired by this observation, in this paper, we propose a visual re-ranking method by contextual similarity aggregation with self-attention. In our approach, for each image in the top-K ranking list, we represent it into an affinity feature vector by comparing it with a set of anchor images. Then, the affinity features of the top-K images are refined by aggregating the contextual information with a transformer encoder. Finally, the affinity features are used to recalculate the similarity scores between the query and the top-K images for re-ranking of the latter. To further improve the robustness of our re-ranking model and enhance the performance of our method, a new data augmentation scheme is designed. Since our re-ranking model is not directly involved with the visual feature used in the initial retrieval, it is ready to be applied to retrieval result lists obtained from various retrieval algorithms. We conduct comprehensive experiments on four benchmark datasets to demonstrate the generality and effectiveness of our proposed visual re-ranking method.
Jianbo Ouyang, Min Wang 0019, Wengang Zhou 0001, Houqiang Li
NeurIPS1
2021 Collaborative Image Relevance Learning for Visual Re-Ranking
abstract
In content-based image retrieval, the initial retrieval result may be unsatisfactory, which can be refined with visual re-ranking techniques, such as query expansion, geometric verification,etc. In this work, we approach visual re-ranking from a novel perspective. Observing that the contextual similarity of images from a retrieval result list exhibits strong visual relevance, we propose to collaboratively learn the semantic relevance among images for visual re-ranking. In our approach, we represent the image set of a fixed-length retrieval list into a correlation matrix, and learn the relevance of all image pairs simultaneously with a lightweight CNN model. To optimize the CNN model, a weighted MSE loss is defined, which takes into account the sparsity of labels. To find the optimal length of retrieval result list for different queries, we present a query sensitive selection method. We conduct comprehensive experiments on five benchmark datasets, and demonstrate the generality, and effectiveness of the proposed visual re-ranking method.
Jianbo Ouyang, Wengang Zhou 0001, Min Wang 0019, Qi Tian 0001, Houqiang Li
IEEE Trans. Multim.1