Linh Ly

dblp:425/4785 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2025
0009-0006-0390-412XORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%
Artificial intelligence
1 paper
Vision and language · 100%

Topics — the 3 heaviest of 3, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Information retrieval
cross-modal retrieval
0.912025
Enhancing Endoscopic Image Retrieval via Self-Supervised Learning and Large VLM-Based Re-ranking · ACM Multimedia 2025
Information retrieval › image retrieval
medical image retrieval
0.912025
Enhancing Endoscopic Image Retrieval via Self-Supervised Learning and Large VLM-Based Re-ranking · ACM Multimedia 2025
Computer vision › Vision and language
vision-language model
0.312025
Enhancing Endoscopic Image Retrieval via Self-Supervised Learning and Large VLM-Based Re-ranking · ACM Multimedia 2025

Methods — techniques the papers use, named apart from their topics

large vision-language model re-ranking · 1.7contrastive learning · 1.7
YearPublicationVenuePosition
2025 Enhancing Endoscopic Image Retrieval via Self-Supervised Learning and Large VLM-Based Re-ranking
abstract
Medical image retrieval is essential for clinical diagnosis and medical education, yet remains highly challenging in endoscopic imaging due to limited annotated data, the lack of domain-specific pretrained models, and subtle visual similarities across anatomical regions. In this work, we utilize self-supervised contrastive learning to pretrain a strong image encoder tailored for endoscopic data, which serves as the backbone for downstream retrieval tasks. For text-to-image retrieval, we adopt a multi-modal contrastive learning approach that aligns textual and visual representations based on this pretrained backbone. To further enhance retrieval performance, we propose a novel re-ranking module that leverages the reasoning capabilities of large vision-language models (LVLMs), such as GPT-4o and Gemini. We also provide a comparative analysis of various retrieval strategies, offering insights into their effectiveness in clinical scenarios. Our method achieves top-2 in text-to-image and top-5 in image-to-image retrieval at the ENTRep Challenge 2025, demonstrating its potential value for endoscopic image retrieval. Source code is available at https://github.com/ELO-Lab/ENTRep-LDSF.
Linh Ly, Duy Khanh Ho, Ngoc Hoang Luong
ACM Multimedia2