Jihao Dong

dblp:364/2072 · DBLP profile ↗
← Back
4ranked-venue papers
3as first author
4since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Vision and language · 67% Representation and self-supervised learning · 33%

Topics — the 2 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Representation and self-supervised learning
contrastive learning
0.912025
Discovering Clone Negatives via Adaptive Contrastive Learning for Image-Text Matching · ICLR 2025
Computer vision › Vision and language › cross-modal matching
image-text matching
0.912025
Discovering Clone Negatives via Adaptive Contrastive Learning for Image-Text Matching · ICLR 2025

Methods — techniques the papers use, named apart from their topics

margin modulation · 0.9adaptive contrastive learning · 0.9
YearPublicationVenuePosition
2025 Semantic Context Re-Mining for Multimodal Guided Human-Object Interaction Detection
abstract
Human-Object Interaction (HOI) detection aims to identify interactions between humans and objects in visual scenes. While vision-language models like CLIP have shown promising results in zero-shot HOI detection, challenges persist, including the limitation in spatial feature modeling and the semantic diversity of verbs in textual representations. To address these issues, we introduce a novel approach that utilizes multimodal prompts to guide interaction prediction. First, we re-mine the semantic context within visual features to generate a more comprehensive interaction representation. Second, we utilize pre-trained models to generate both visual and textual prompts, effectively transferring the prior knowledge to the HOI detection task. Our method achieves competitive performance on standard HOI benchmarks, Especially under the zero-shot setting, demonstrating its potential to advance the field of HOI detection.
Jihao Dong
ICIP1
2025 Discovering Clone Negatives via Adaptive Contrastive Learning for Image-Text Matching
abstract
In this paper, we identify a common yet challenging issue in image-text matching, i.e., clone negatives: negative image-text pairs that semantically resemble positive pairs, leading to ambiguous and sub-optimal matching outcomes. To tackle this issue, we propose Adaptive Contrastive Learning (AdaCL), which introduces two margin parameters along with a modulating anchor to dynamically strengthen the compactness between positives and mitigate the influence of clone negatives. The modulating anchor is selected based on the distribution of negative samples without the need for explicit training, allowing for progressive tuning and advanced in-batch supervision. Extensive experiments across several tasks demonstrate the effectiveness of AdaCL in image-text matching. Furthermore, we extend AdaCL to weakly-supervised image-text matching by replacing human-annotated descriptions with automatically generated captions, thereby increasing the number of potential clone negatives. AdaCL maintains robustness in this setting, alleviating the reliance on crowd-sourced annotations and laying a foundation for scalable vision-language contrastive learning.
Renjie Pan 0001, Jihao Dong, Hua Yang 0001
ICLR2
2025 RM-BGNN: A weakly informative Bayesian graph neural network based on residual mechanism
Jihao Dong, Zhaowei Liu 0001, Peng Song 0002, Jinglei Liu, Anzuo Jiang
Neurocomputing1
2024 Exploring Interactive Semantic Alignment for Efficient HOI Detection with Vision-language Model
abstract
Human-Object Interaction (HOI) detection aims to localize human-object pairs and comprehend their interactions. Recently, two-stage transformer-based methods have demonstrated competitive performance. However, these methods frequently focus on object appearance features and ignore global contextual information. Besides, vision-language model CLIP which effectively aligns visual and text embeddings has shown great potential in zero-shot HOI detection. Based on the former facts, We introduce a novel HOI detector named ISA-HOI, which extensively leverages knowledge from CLIP, aligning interactive semantics between visual and textual features. We first extract global context of image and local features of object to Improve interaction Features in images (IF). On the other hand, we propose a Verb Semantic Improvement (VSI) module to enhance textual features of verb labels via cross-modal fusion. Ultimately, our method achieves competitive results on the HICO-DET and V-COCO benchmarks with much fewer training epochs, and outperforms the state-of-the-art under zero-shot settings.
Jihao Dong, Hua Yang 0001, Renjie Pan 0001
ICME1