Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Jinhao Li 0004

dblp:309/6695-4 · DBLP profile ↗
← Back
2ranked-venue papers
2as first author
2since 2021 · last 2026
0009-0006-9301-5579ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
1 paper
Knowledge graphs · 87% Information retrieval · 13%
Artificial intelligence
1 paper
Vision and language · 67% Transfer learning and domain adaptation · 33%

Topics — the 5 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Knowledge graphs
knowledge graph construction
1.012026
MuseKG: An Interactive Knowledge Graph Over Museum Collections · SIGIR 2026
Knowledge graphs › knowledge graph querying
knowledge graph question answering
1.012026
MuseKG: An Interactive Knowledge Graph Over Museum Collections · SIGIR 2026
Computer vision › Vision and language › cross-modal matching
cross-modal similarity
0.812024
Visual-Text Cross Alignment: Refining the Similarity Score in Vision-Language Models · ICML 2024
Computer vision › Vision and language
vision-language model
0.812024
Visual-Text Cross Alignment: Refining the Similarity Score in Vision-Language Models · ICML 2024
Machine learning › Transfer learning and domain adaptation › zero-shot learning
zero-shot classification
0.812024
Visual-Text Cross Alignment: Refining the Similarity Score in Vision-Language Models · ICML 2024

Methods — techniques the papers use, named apart from their topics

natural language query grounding · 1.0evidence retrieval · 1.0similarity matrix · 0.8localized visual prompting · 0.8large language model · 0.8
YearPublicationVenuePosition
2026 MuseKG: An Interactive Knowledge Graph Over Museum Collections
abstract
Digitisation in the cultural heritage sector has produced large but fragmented repositories of museum collection data, spanning structured catalogue records, images, and unstructured descriptions. Existing museum information systems often make it difficult to integrate these sources into a unified, queryable representation that supports relation-aware exploration. We present MuseKG, an interactive knowledge graph system that organises heterogeneous museum data into a typed graph that links objects, people, organisations, images, image-derived labels, and extracted semantic entities within a coherent schema. MuseKG supports natural-language queries by grounding user questions to graph entities and retrieving a compact neighbourhood of evidence for answer generation. Through an interactive demonstration on real museum collections, we show that MuseKG supports common exploration tasks such as attribute lookup, relation exploration, and relation-aware retrieval, with answers that remain inspectable via explicit graph structures.
Jinhao Li 0004, Jianzhong Qi 0001, Soyeon Caren Han, Eun-Jung Holden
SIGIR1
2024 Visual-Text Cross Alignment: Refining the Similarity Score in Vision-Language Models
abstract
It has recently been discovered that using a pre-trained vision-language model (VLM), e.g., CLIP, to align a whole query image with several finer text descriptions generated by a large language model can significantly enhance zero-shot performance. However, in this paper, we empirically find that the finer descriptions tend to align more effectively with local areas of the query image rather than the whole image, and then we theoretically validate this finding. Thus, we present a method called weighted visual-text cross alignment (WCA). This method begins with a localized visual prompting technique, designed to identify local visual areas within the query image. The local visual areas are then cross-aligned with the finer descriptions by creating a similarity matrix using the pre-trained VLM. To determine how well a query image aligns with each category, we develop a score function based on the weighted similarities in this matrix. Extensive experiments demonstrate that our method significantly improves zero-shot performance across various datasets, achieving results that are even comparable to few-shot learning methods.
Jinhao Li 0004, Sarah M. Erfani, Lei Feng 0006, James Bailey 0001, Feng Liu 0003
ICML1