VLDB 2026 Research / reviewers in the wild / expert
Sanghwan Kim
dblp:60/10274
· DBLP profile ↗
6ranked-venue papers
4as first author
6since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 4 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Learning paradigms · 25% Vision and language · 24% Efficient and distributed learning · 22% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% |
Topics — the 16 heaviest of 16, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Vision and language › cross-modal alignment › image-text alignment
fine-grained image-text alignment |
0.9 | 1 | 2025 | FLAIR: VLM with Fine-grained Language-informed Image Representations · CVPR 2025 |
Machine learning › Efficient and distributed learning › model compression › knowledge distillation
self-distillation |
0.9 | 1 | 2025 | COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training · CVPR 2025 |
Computer vision › Vision and language
vision-language pretraining |
0.9 | 1 | 2025 | COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-training · CVPR 2025 |
Computer vision › Vision and language › multimodal representation
vision-language representation learning |
0.9 | 1 | 2025 | FLAIR: VLM with Fine-grained Language-informed Image Representations · CVPR 2025 |
Information retrieval
cross-modal retrieval |
0.9 | 1 | 2025 | FLAIR: VLM with Fine-grained Language-informed Image Representations · CVPR 2025 |
Information retrieval › image retrieval › content-based image retrieval
fine-grained image retrieval |
0.9 | 1 | 2025 | FLAIR: VLM with Fine-grained Language-informed Image Representations · CVPR 2025 |
Computer vision › Video understanding and tracking
action anticipation |
0.8 | 1 | 2024 | PALM: Predicting Actions through Language Models · ECCV (82) 2024 |
Machine learning › Generative modeling
diffusion model |
0.8 | 1 | 2024 | Distilling ODE Solvers of Diffusion Models into Smaller Steps · CVPR 2024 |
Machine learning › Generative modeling › diffusion model
diffusion sampling |
0.8 | 1 | 2024 | Distilling ODE Solvers of Diffusion Models into Smaller Steps · CVPR 2024 |
Machine learning › Efficient and distributed learning › model compression
knowledge distillation |
0.8 | 1 | 2024 | Distilling ODE Solvers of Diffusion Models into Smaller Steps · CVPR 2024 |
Natural language and speech › Language models and text generation
large language model reasoning |
0.8 | 1 | 2024 | PALM: Predicting Actions through Language Models · ECCV (82) 2024 |
Machine learning › Efficient and distributed learning
model compression |
0.8 | 1 | 2024 | Distilling ODE Solvers of Diffusion Models into Smaller Steps · CVPR 2024 |
Machine learning › Learning paradigms › continual learning
class-incremental learning |
0.7 | 1 | 2023 | Achieving a Better Stability-Plasticity Trade-off via Auxiliary Networks in Continual Learning · CVPR 2023 |
Machine learning › Learning paradigms
continual learning |
0.7 | 1 | 2023 | Achieving a Better Stability-Plasticity Trade-off via Auxiliary Networks in Continual Learning · CVPR 2023 |
Machine learning › Learning paradigms › continual learning
stability-plasticity trade-off |
0.7 | 1 | 2023 | Achieving a Better Stability-Plasticity Trade-off via Auxiliary Networks in Continual Learning · CVPR 2023 |
Machine learning › Learning paradigms › continual learning
task incremental learning |
0.7 | 1 | 2023 | Achieving a Better Stability-Plasticity Trade-off via Auxiliary Networks in Continual Learning · CVPR 2023 |
Methods — techniques the papers use, named apart from their topics
contrastive learning · 2.6attention pooling · 1.7text-cropping · 0.9cross-attention · 0.9prompting · 0.8ordinary differential equation solver · 0.8language model · 0.8knowledge distillation · 0.8regularization · 0.7auxiliary network · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | COSMOS: Cross-Modality Self-Distillation for Vision Language Pre-trainingabstractVision-Language Models (VLMs) trained with contrastive loss have achieved significant advancements in various vision and language tasks. However, the global nature of the contrastive loss makes VLMs focus predominantly on foreground objects, neglecting other crucial information in the image, which limits their effectiveness in downstream tasks. To address these challenges, we propose COSMOS: CrOSs-MOdality Self-distillation for vision-language pre-training that integrates a novel text-cropping strategy and cross-attention module into a self-supervised learning framework. We create global and local views of images and texts (i.e., multi-modal augmentations), which are essential for self-distillation in VLMs. We further introduce a cross-attention module, enabling COSMOS to learn comprehensive cross-modal representations optimized via a cross-modality self-distillation loss. COSMOS consistently outperforms previous strong baselines on various zero-shot downstream tasks, including retrieval, classification, and semantic segmentation. Additionally, it surpasses CLIP-based models trained on larger datasets in visual perception and contextual understanding tasks. Code is available at https://github.com/ExplainableML/cosmos. Sanghwan Kim, Mariana-Iuliana Georgescu, Stephan Alaniz, Zeynep Akata |
CVPR | 1 |
| 2025 | FLAIR: VLM with Fine-grained Language-informed Image RepresentationsabstractCLIP has shown impressive results in aligning images and texts at scale. However, its ability to capture detailed visual features remains limited because CLIP matches images and texts at a global level. To address this issue, we propose FLAIR, Fine-grained Language-informed Image Representations, an approach that utilizes long and detailed image descriptions to learn localized image embeddings. By sampling diverse sub-captions that describe fine-grained details about an image, we train our vision-language model to produce not only global embeddings but also text-specific image representations. Our model introduces text-conditioned attention pooling on top of local image tokens to produce fine-grained image representations that excel at retrieving detailed image content. We achieve state-of-the-art performance on both, existing multimodal retrieval benchmarks, as well as, our newly introduced fine-grained retrieval task which evaluates vision-language models’ ability to retrieve partial image content. Furthermore, our experiments demonstrate the effectiveness of FLAIR trained on 30M image-text pairs in capturing fine-grained visual information, including zero-shot semantic segmentation, outperforming models trained on billions of pairs. Code is available at https://github.com/ExplainableML/flair. Sanghwan Kim, Mariana-Iuliana Georgescu, Zeynep Akata, Stephan Alaniz |
CVPR | 2 |
| 2024 | Distilling ODE Solvers of Diffusion Models into Smaller StepsabstractDiffusion models have recently gained prominence as a novel category of generative models. Despite their success, these models face a notable drawback in terms of slow sampling speeds, requiring a high number of function evaluations (NFE) in the order of hundreds or thousands. In response, both learning-free and learning-based sampling strategies have been explored to expedite the sampling process. Learning-free sampling employs various ordinary differential equation (ODE) solvers based on the formulation of diffusion ODEs. However, it encounters challenges in faithfully tracking the true sampling trajectory, particularly for small NFE. Conversely, learning-based sampling methods, such as knowledge distillation, demand extensive additional training, limiting their practical applicability. To overcome these limitations, we introduce Distilled-ODE solvers (D-ODE solvers), a straightforward distillation approach grounded in ODE solver formulations. Our method seamlessly integrates the strengths of both learning-free and learning-based sampling. D-ODE solvers are constructed by introducing a single parameter adjustment to existing ODE solvers. Furthermore, we optimize D-ODE solvers with smaller steps using knowledge distillation from ODE solvers with larger steps across a batch of samples. Comprehensive experiments demonstrate the superior performance of D- ODE solvers compared to existing ODE solvers, including DDIM, PNDM, DPM-Solver, DEIS, and EDM, particularly in scenarios with fewer NFE. Notably, our method incurs negligible computational overhead compared to previous distillation techniques, facilitating straightforward and rapid integration with existing samplers. Qualitative analysis reveals that D-ODE solvers not only enhance image quality but also faithfully follow the target ODE trajectory. Sanghwan Kim, Hao Tang 0005, Fisher Yu 0001 |
CVPR | 1 |
| 2024 | PALM: Predicting Actions through Language Models
Sanghwan Kim, Daoji Huang, Yongqin Xian, Otmar Hilliges, Luc Van Gool, Xi Wang 0021 |
ECCV (82) | 1 |
| 2024 | Extracting lung cancer staging descriptors from pathology reports: A generative language model approachabstractBACKGROUND: In oncology, electronic health records contain textual key information for the diagnosis, staging, and treatment planning of patients with cancer. However, text data processing requires a lot of time and effort, which limits the utilization of these data. Recent advances in natural language processing (NLP) technology, including large language models, can be applied to cancer research. Particularly, extracting the information required for the pathological stage from surgical pathology reports can be utilized to update cancer staging according to the latest cancer staging guidelines. OBJECTIVES: This study has two main objectives. The first objective is to evaluate the performance of extracting information from text-based surgical pathology reports and determining pathological stages based on the extracted information using fine-tuned generative language models (GLMs) for patients with lung cancer. The second objective is to determine the feasibility of utilizing relatively small GLMs for information extraction in a resource-constrained computing environment. METHODS: Lung cancer surgical pathology reports were collected from the Common Data Model database of Seoul National University Bundang Hospital (SNUBH), a tertiary hospital in Korea. We selected 42 descriptors necessary for tumor-node (TN) classification based on these reports and created a gold standard with validation by two clinical experts. The pathology reports and gold standard were used to generate prompt-response pairs for training and evaluating GLMs which then were used to extract information required for staging from pathology reports. RESULTS: We evaluated the information extraction performance of six trained models as well as their performance in TN classification using the extracted information. The Deductive Mistral-7B model, which was pre-trained with the deductive dataset, showed the best performance overall, with an exact match ratio of 92.24% in the information extraction problem and an accuracy of 0.9876 (predicting T and N classification concurrently) in classification. CONCLUSION: This study demonstrated that training GLMs with deductive datasets can improve information extraction performance, and GLMs with a relatively small number of parameters at approximately seven billion can achieve high performance in this problem. The proposed GLM-based information extraction method is expected to be useful in clinical decision-making support, lung cancer staging and research. Hyeongmin Cho, Sooyoung Yoo, Borham Kim, Sowon Jang, Leonard Sunwoo, Sanghwan Kim, Donghyoung Lee, Sejin Nam, Jin-Haeng Chung |
J. Biomed. Informatics | 6 |
| 2023 | Achieving a Better Stability-Plasticity Trade-off via Auxiliary Networks in Continual LearningabstractIn contrast to the natural capabilities of humans to learn new tasks in a sequential fashion, neural networks are known to suffer from catastrophic forgetting, where the model's performances on old tasks drop dramatically after being optimized for a new task. Since then, the continual learning (CL) community has proposed several solutions aiming to equip the neural network with the ability to learn the current task (plasticity) while still achieving high accuracy on the previous tasks (stability). Despite remarkable improvements, the plasticity-stability trade-off is still far from being solved and its underlying mechanism is poorly understood. In this work, we propose Auxiliary Network Continual Learning (ANCL), a novel method that applies an additional auxiliary network which promotes plasticity to the continually learned model which mainly focuses on stability. More concretely, the proposed framework materializes in a regularizer that naturally interpolates between plasticity and stability, surpassing strong baselines on task incremental and class incremental scenarios. Through extensive analyses on ANCL solutions, we identify some essential principles beneath the stability-plasticity tradeoff. The code implementation of our work is available at https://github.com/kim-sanghwan/ANCL. Sanghwan Kim, Lorenzo Noci, Antonio Orvieto, Thomas Hofmann 0001 |
CVPR | 1 |