VLDB 2026 Research / reviewers in the wild / expert
Jae-Myung Kim
dblp:51/1888 · also Jae Myung Kim
· DBLP profile ↗
13ranked-venue papers
3as first author
6since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 3Computer networks · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | DataDream: Few-Shot Guided Dataset Generation
Jae-Myung Kim, Jessica Bader, Stephan Alaniz, Cordelia Schmid, Zeynep Akata |
ECCV (71) | 1 |
| 2024 | Improving Intervention Efficacy via Concept Realignment in Concept Bottleneck Models
Nishad Singhi, Jae-Myung Kim, Karsten Roth, Zeynep Akata |
ECCV (26) | 2 |
| 2023 | Bridging the Gap Between Model Explanations in Partially Annotated Multi-Label ClassificationabstractDue to the expensive costs of collecting labels in multi-label classification datasets, partially annotated multi-label classification has become an emerging field in computer vision. One baseline approach to this task is to assume unobserved labels as negative labels, but this assumption induces label noise as a form of false negative. To understand the negative impact caused by false negative labels, we study how these labels affect the model's explanation. We observe that the explanation of two models, trained with full and partial labels each, highlights similar regions but with different scaling, where the latter tends to have lower attribution scores. Based on these findings, we propose to boost the attribution scores of the model trained with partial labels to make its explanation resemble that of the model trained with full labels. Even with the conceptually simple approach, the multi-label classification performance improves by a large margin in three different datasets on a single positive label setting and one on a large-scale partial label setting. Code is available at https://github.com/youngwk/BridgeGapExplanationPAMC. Youngwook Kim 0005, Jae-Myung Kim, Jieun Jeong, Cordelia Schmid, Zeynep Akata, Jungwoo Lee 0001 |
CVPR | 2 |
| 2023 | Waffling around for Performance: Visual Classification with Random Words and Broad ConceptsabstractThe visual classification performance of vision-language models such as CLIP has been shown to benefit from additional semantic knowledge from large language models (LLMs) such as GPT-3. In particular, averaging over LLM-generated class descriptors, e.g. "waffle, which has a round shape", can notably improve generalization performance. In this work, we critically study this behavior and propose WaffleCLIP, a framework for zero-shot visual classification which simply replaces LLM-generated descriptors with random character and word descriptors. Without querying external models, we achieve comparable performance gains on a large number of visual classification tasks. This allows WaffleCLIP to both serve as a low-cost alternative, as well as a sanity check for any future LLM-based vision-language model extensions. We conduct an extensive experimental study on the impact and shortcomings of additional semantics introduced with LLM-generated descriptors, and showcase how - if available - semantic context is better leveraged by querying LLMs for high-level concepts, which we show can be done to jointly resolve potential class name ambiguities. Code is available here: https://github.com/ExplainableML/WaffleCLIP. Karsten Roth, Jae-Myung Kim, A. Sophia Koepke, Oriol Vinyals, Cordelia Schmid, Zeynep Akata |
ICCV | 2 |
| 2022 | Large Loss Matters in Weakly Supervised Multi-Label ClassificationabstractWeakly supervised multi-label classification (WSML) task, which is to learn a multi-label classification using partially observed labels per image, is becoming increasingly important due to its huge annotation cost. In this work, we first regard unobserved labels as negative labels, casting the WSML task into noisy multi-label classification. From this point of view, we empirically observe that memorization effect, which was first discovered in a noisy multi-class setting, also occurs in a multi-label setting. That is, the model first learns the representation of clean labels, and then starts memorizing noisy labels. Based on this finding, we propose novel methods for WSML which reject or correct the large loss samples to prevent model from memorizing the noisy label. Without heavy and complex components, our proposed methods outperform previous state-of-the-art WSML methods on several partial label settings including Pascal VOC 2012, MS COCO, NUSWIDE, CUB, and OpenImages V3 datasets. Various analysis also show that our methodology actually works well, validating that treating large loss properly matters in a weakly supervised multi-label classification. Our code is available at https://github.com/snucml/LargeLossMatters. Youngwook Kim 0005, Jae-Myung Kim, Zeynep Akata, Jungwoo Lee 0001 |
CVPR | 2 |
| 2021 | Keep CALM and Improve Visual Feature AttributionabstractThe class activation mapping, or CAM, has been the cornerstone of feature attribution methods for multiple vision tasks. Its simplicity and effectiveness have led to wide applications in the explanation of visual predictions and weakly-supervised localization tasks. However, CAM has its own shortcomings. The computation of attribution maps relies on ad-hoc calibration steps that are not part of the training computational graph, making it difficult for us to understand the real meaning of the attribution values. In this paper, we improve CAM by explicitly incorporating a la-tent variable encoding the location of the cue for recognition in the formulation, thereby subsuming the attribution map into the training computational graph. The resulting model, class activation latent mapping, or CALM, is trained with the expectation-maximization algorithm. Our experiments show that CALM identifies discriminative attributes for image classifiers more accurately than CAM and other visual attribution baselines. CALM also shows performance improvements over prior arts on the weakly-supervised object localization benchmarks. Our code is available at https://github.com/naver-ai/calm. Jae-Myung Kim, Junsuk Choe, Zeynep Akata, Seong Joon Oh |
ICCV | 1 |
| 2020 | REST: Performance Improvement of a Black Box Model via RL-Based Spatial Transformation
Jae-Myung Kim, Chanwoo Park, Jungwoo Lee 0001 |
AAAI | 1 |
| 2012 | Reducing cache misses in hash join probing phase by pre-sorting strategy (abstract only)abstractRecently, several studies on multi-core cache-aware hash join have been carried out [Kim09VLDB, Blanas11SIGMOD]. In particular, the work of Blanas has shown that rather simple no-partitioning hash join can outperform the work of Kim. Meanwhile, the simple but best performing hash join of Blanas still experiences severe cache misses in probing phase. Because the key values of tuples in outer relation are not sorted or clustered, each outer record has different hashed key value and thus accesses the different hash bucket. Since the size of hash table of inner table is usually much larger than that of the CPU cache, it is highly probable that the reference to hash bucket of inner table by each outer record would encounter cache miss. To reduce the cache misses in hash join probing phase, we propose a new join algorithm, Sorted Probing (in short, SP), which pre-sorts the hashed key values of outer table of hash join so that the access to the hash bucket of inner table has strong temporal locality, thus minimizing the cache misses during the probing phase. As an optimization technique of sorting, we used the cache-aware AlphaSort technique, which extracts the key from each record of data set to be sorted and its pointer, and then sorts the pairs of (key, rec_ptr). For performance evaluation, we used two hash join algorithms from Blanas' work, no partitioning(NP) and independent partitioning(IP) in a standard C++ program, provided by Blanas. Also, we implemented the AlphaSort and added it before each probing phase of NP and IP, and we call each algorithm as NP+SP and IP+SP. For syntactic workload, IP+SP outperforms all other algorithms: IP+SP is faster than other altorithms up to 30%. Gi-Hwan Oh, Jae-Myung Kim, Woon-Hak Kang, Sang-Won Lee 0001 |
SIGMOD Conference | 2 |
| 2009 | BAR: bitmap-based association rule: an implementation and its optimizationsabstractThe association rule mining, one of the most popular data mining techniques, is to find the frequent itemsets which occur commonly in transaction database. Of the various association algorithms, the Apriori is the most popular one, and its implementation technique to improve the performance has been continuously developed during the past decade. In this paper, we propose a bitmap-based association rule technique, called BAR, in order to drastically improve the performance of the Apriori algorithm. Compared to the latest Apriori implementation, our approach can improve the performance by nearly up to two orders of magnitude. This gain comes mainly from the following characteristics of BAR: 1) bitmap based implementation paradigm, 2) reduction of redundant bitmap-AND operations, and 3) an efficient implementation of bitmap-AND and bit-counting operation by exploiting the advanced CPU technology, including SIMD and SW prefetching. We will describe the basic concept of BAR approach and its optimization techniques, and will show, through experimental results, how each of the above characteristics of BAR can contribute the performance improvement. Sung-Tan Kim, Jae-Myung Kim, Sang-Won Lee 0001 |
MoMM | 2 |
| 2008 | A case for flash memory ssd in enterprise database applicationsabstractDue to its superiority such as low access latency, low energy consumption, light weight, and shock resistance, the success of flash memory as a storage alternative for mobile computing devices has been steadily expanded into personal computer and enterprise server markets with ever increasing capacity of its storage. However, since flash memory exhibits poor performance for small-to-moderate sized writes requested in a random order, existing database systems may not be able to take full advantage of flash memory without elaborate flash-aware data structures and algorithms. The objective of this work is to understand the applicability and potential impact that flash memory SSD (Solid State Drive) has for certain type of storage spaces of a database server where sequential writes and random reads are prevalent. We show empirically that up to more than an order of magnitude improvement can be achieved in transaction processing by replacing magnetic disk with flash memory SSD for transaction log, rollback segments, and temporary table spaces. Sang-Won Lee 0001, Bongki Moon, Chanik Park, Jae-Myung Kim |
SIGMOD Conference | 4 |
| 2007 | Implementation of Bitmap Based Incognito and Performance Evaluation
Hyun-Ho Kang, Jae-Myung Kim, Gap-Joo Na, Sang-Won Lee 0001 |
DASFAA | 2 |
| 2007 | Research issues in next generation DBMS for mobile platformsabstractRecently, flash memory(in particular, NAND) is being rapidly deployed as data storage for mobile platforms such as PDAs, MP3 players, mobile phones and digital cameras, mainly because of its many advantages over its competitor, hard disk, including its low electronic power, non-volatile storage, high performance, physical stability, smaller size, light weight, and portability. Considering its rapid technical improvement both in capacity and speed, it will have a competitive advantage over its rivalry minidrive (i.e. a small size hard disk) under 100 Gbytes within a few years, As the applications in next generation mobile platforms become large, complex, and more data-oriented, they requires the database technology, because the file interface is too complex to manage their complicated data requirements. However, flash memory, compared to hard disk, has a few unique characteristics, and thus the traditional disk-based database technology does not seem to go well with flash memory. Therefore, we need to revisit almost every aspect of DBMS implementation techniques from the perspectives of flash memory. In this paper, we introduce the technical characteristics of flash memory, which we think might have huge impact on database performance to database community that are 1) no-overwrite (erase-before-write paradigm), 2) asymmetric read and write speed, and 3) no seek or rotation time. These small differences necessitate us to revisit all the major DBMS modules which have evolved over the several decades. Based on the characteristics, we identify several key issues in implementing major DBMS modules, and suggest alternative approaches to solve the issues. The topics covered in this article are neither comprehensive nor in-depth, but the main goal of this article is just to issue that a practical and urgent research topic is ahead and it poses us many challenges and opportunities. Sang-Won Lee 0001, Gap-Joo Na, Jae-Myung Kim, Joo-Hyung Oh |
Mobile HCI | 3 |
| 2006 | A Novel Rekey Management Scheme in Digital Broadcasting Network
Han-Seung Koo, Il-Kyoo Lee, Jae-Myung Kim, Sung-Woong Ra |
APNOMS | 3 |