VLDB 2026 Research / reviewers in the wild / expert
Jeong Ryong Lee
dblp:300/5592
· DBLP profile ↗
5ranked-venue papers
3as first author
5since 2021 · last 2026
0000-0001-7251-2126ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Vision and language · 40% Trustworthy machine learning · 23% Transfer learning and domain adaptation · 13% |
Topics — the 9 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language
cross-modal alignment |
0.9 | 1 | 2025 | Diffusion Bridge: Leveraging Diffusion Model to Reduce the Modality Gap Between Text and Vision for Zero-Shot Image Captioning · CVPR 2025 |
Computer vision › Vision and language
image captioning |
0.9 | 1 | 2025 | Diffusion Bridge: Leveraging Diffusion Model to Reduce the Modality Gap Between Text and Vision for Zero-Shot Image Captioning · CVPR 2025 |
Machine learning › Transfer learning and domain adaptation › domain alignment
modality gap reduction |
0.9 | 1 | 2025 | Diffusion Bridge: Leveraging Diffusion Model to Reduce the Modality Gap Between Text and Vision for Zero-Shot Image Captioning · CVPR 2025 |
Computer vision › Vision and language › image captioning › low-shot image captioning
zero-shot image captioning |
0.9 | 1 | 2025 | Diffusion Bridge: Leveraging Diffusion Model to Reduce the Modality Gap Between Text and Vision for Zero-Shot Image Captioning · CVPR 2025 |
Natural language and speech › Language models and text generation
chain-of-thought reasoning |
0.8 | 1 | 2024 | Large Language Models Are Clinical Reasoners: Reasoning-Aware Diagnosis Framework with Prompt-Generated Rationales · AAAI 2024 |
Natural language and speech › Question answering and dialogue systems
medical diagnosis |
0.8 | 1 | 2024 | Large Language Models Are Clinical Reasoners: Reasoning-Aware Diagnosis Framework with Prompt-Generated Rationales · AAAI 2024 |
Machine learning › Trustworthy machine learning › interpretability › visual explanation
class activation map |
0.5 | 1 | 2021 | Relevance-CAM: Your Model Already Knows Where To Look · CVPR 2021 |
Machine learning › Trustworthy machine learning
interpretability |
0.5 | 1 | 2021 | Relevance-CAM: Your Model Already Knows Where To Look · CVPR 2021 |
Machine learning › Trustworthy machine learning › interpretability › visual explanation
saliency map |
0.5 | 1 | 2021 | Relevance-CAM: Your Model Already Knows Where To Look · CVPR 2021 |
Methods — techniques the papers use, named apart from their topics
reverse diffusion · 0.9denoising diffusion probabilistic model · 0.9prompt-based learning · 0.8large language model · 0.8layer-wise relevance propagation · 0.5
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | HP-GAN: Harnessing pretrained networks for GAN improvement with FakeTwins and discriminator consistency
Geonhui Son, Jeong Ryong Lee, Dosik Hwang |
Neural Networks | 2 |
| 2026 | FeDi: Feature disentanglement for self-supervised learning
Jeong Ryong Lee, Geonhui Son, Dosik Hwang |
Pattern Recognit. | 1 |
| 2025 | Diffusion Bridge: Leveraging Diffusion Model to Reduce the Modality Gap Between Text and Vision for Zero-Shot Image CaptioningabstractThe modality gap between vision and text embeddings in CLIP presents a significant challenge for zero-shot image captioning, limiting effective cross-modal representation. Traditional approaches, such as noise injection and memory-based similarity matching, attempt to address this gap, yet these methods either rely on indirect alignment or relatively naive solutions with heavy computation. Diffusion Bridge introduces a novel approach to directly reduce this modality gap by leveraging Denoising Diffusion Probabilistic Models (DDPM), trained exclusively on text embeddings to model their distribution. Our approach is motivated by the observation that, while paired vision and text embeddings are relatively close, a modality gap still exists due to stable regions created by the contrastive loss. This gap can be interpreted as noise in cross-modal mappings, which we approximate as Gaussian noise. To bridge this gap, we employ a reverse diffusion process, where image embeddings are strategically introduced at an intermediate step in the reverse process, allowing them to be refined progressively toward the text embedding distribution. This process transforms vision embeddings into text-like representations closely aligned with paired text embeddings, effectively minimizing discrepancies between modalities. Experimental results demonstrate that these text-like vision embeddings significantly enhance alignment with their paired text embeddings, leading to improved zero-shot captioning performance on MSCOCO and Flickr30K. Diffusion Bridge achieves competitive results without reliance on memory banks or entity-driven methods, offering a novel pathway for cross-modal alignment and opening new possibilities for the application of diffusion models in multi-modal tasks. The source code is available at: https://github.com/mongeoroo/diffusion-bridge Jeong Ryong Lee, Yejee Shin, Geonhui Son, Dosik Hwang |
CVPR | 1 |
| 2024 | Large Language Models Are Clinical Reasoners: Reasoning-Aware Diagnosis Framework with Prompt-Generated RationalesabstractMachine reasoning has made great progress in recent years owing to large language models (LLMs). In the clinical domain, however, most NLP-driven projects mainly focus on clinical classification or reading comprehension, and under-explore clinical reasoning for disease diagnosis due to the expensive rationale annotation with clinicians. In this work, we present a "reasoning-aware" diagnosis framework that rationalizes the diagnostic process via prompt-based learning in a time- and labor-efficient manner, and learns to reason over the prompt-generated rationales. Specifically, we address the clinical reasoning for disease diagnosis, where the LLM generates diagnostic rationales providing its insight on presented patient data and the reasoning path towards the diagnosis, namely Clinical Chain-of-Thought (Clinical CoT). We empirically demonstrate LLMs/LMs' ability of clinical reasoning via extensive experiments and analyses on both rationale generation and disease diagnosis in various settings. We further propose a novel set of criteria for evaluating machine-generated rationales' potential for real-world clinical settings, facilitating and benefiting future research in this area. Taeyoon Kwon, Kai Tzu-iunn Ong, Dongjin Kang, Seungjun Moon, Jeong Ryong Lee, Dosik Hwang, Beomseok Sohn, Yongsik Sim, Dongha Lee 0003, Jinyoung Yeo |
AAAI | 5 |
| 2021 | Relevance-CAM: Your Model Already Knows Where To LookabstractWith increasing fields of application for neural networks and the development of neural networks, the ability to explain deep learning models is also becoming increasingly important. Especially, prior to practical applications, it is crucial to analyze a model’s inference and the process of generating the results. A common explanation method is Class Activation Mapping(CAM) based method where it is often used to understand the last layer of the convolutional neural networks popular in the field of Computer Vision. In this paper, we propose a novel CAM method named Relevance-weighted Class Activation Mapping(Relevance-CAM) that utilizes Layer-wise Relevance Propagation to obtain the weighting components. This allows the explanation map to be faithful and robust to the shattered gradient problem, a shared problem of the gradient based CAM methods that causes noisy saliency maps for intermediate layers. Therefore, our proposed method can better explain a model by correctly analyzing the intermediate layers as well as the last convolutional layer. In this paper, we visualize how each layer of the popular image processing models extracts class specific features using Relevance-CAM, evaluate the localization ability, and show why the gradient based CAM cannot be used to explain the intermediate layers, proven by experimenting the weighting component. Relevance-CAM outperforms other CAM-based methods in recognition and localization evaluation in layers of any depth. The source code is available at: https://github.com/mongeoroo/Relevance-CAM Jeong Ryong Lee, Sewon Kim, Inyong Park, Taejoon Eo, Dosik Hwang |
CVPR | 1 |