EDBT 2026 Demo / reviewers in the wild / expert
Dae-Shik Kim
dblp:25/2348 · also Daeshik Kim
· DBLP profile ↗
34ranked-venue papers
0as first author
19since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Systems, architecture and hardware · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AIPO: Automatic Instruction Prompt Optimization by model itself with "Gradient Ascent"
Kyeonghye Park, Dae-Shik Kim |
Comput. Speech Lang. | 2 |
| 2025 | MMR: A Large-scale Benchmark Dataset for Multi-target and Multi-granularity Reasoning SegmentationabstractThe fusion of Large Language Models (LLMs) with vision models is pioneering new possibilities in user-interactive vision-language tasks. A notable application is reasoning segmentation, where models generate pixel-level segmentation masks by comprehending implicit meanings in human instructions. However, seamless human-AI interaction demands more than just object-level recognition; it requires understanding both objects and the functions of their detailed parts, particularly in multi-target scenarios. For example, when instructing a robot to \textit{“turn on the TV"}, there could be various ways to accomplish this command. Recognizing multiple objects capable of turning on the TV, such as the TV itself or a remote control (multi-target), provides more flexible options and aids in finding the optimized scenario. Furthermore, understanding specific parts of these objects, like the TV's button or the remote's button (part-level), is important for completing the action. Unfortunately, current reasoning segmentation datasets predominantly focus on a single target object-level reasoning, which limits the detailed recognition of an object's parts in multi-target contexts. To address this gap, we construct a large-scale dataset called Multi-target and Multi-granularity Reasoning (MMR). MMR comprises 194K complex and implicit instructions that consider multi-target, object-level, and part-level aspects, based on pre-existing image-mask sets. This dataset supports diverse and context-aware interactions by hierarchically providing object and part information. Moreover, we propose a straightforward yet effective framework for multi-target, object-level, and part-level reasoning segmentation. Experimental results on MMR show that the proposed method can reason effectively in multi-target and multi-granularity scenarios, while the existing reasoning segmentation model still has room for improvement. The dataset is available at \url{https://github.com/jdg900/MMR}. Donggon Jang, Yucheol Cho, Suin Lee, Taehyeon Kim 0005, Dae-Shik Kim |
ICLR | 5 |
| 2025 | TexTailor: Customized Text-aligned Texturing via Effective ResamplingabstractWe present TexTailor, a novel method for generating consistent object textures from textual descriptions. Existing text-to-texture synthesis approaches utilize depth-aware diffusion models to progressively generate images and synthesize textures across predefined multiple viewpoints. However, these approaches lead to a gradual shift in texture properties across viewpoints due to (1) insufficient integration of previously synthesized textures at each viewpoint during the diffusion process and (2) the autoregressive nature of the texture synthesis process. Moreover, the predefined selection of camera positions, which does not account for the object's geometry, limits the effective use of texture information synthesized from different viewpoints, ultimately degrading overall texture consistency. In TexTailor, we address these issues by (1) applying a resampling scheme that repeatedly integrates information from previously synthesized textures within the diffusion process, and (2) fine-tuning a depth-aware diffusion model on these resampled textures. During this process, we observed that using only a few training images restricts the model's original ability to generate high-fidelity images aligned with the conditioning, and therefore propose an performance preservation loss to mitigate this issue. Additionally, we improve the synthesis of view-consistent textures by adaptively adjusting camera positions based on the object's geometry. Experiments on a subset of the Objaverse dataset and the ShapeNet car dataset demonstrate that TexTailor outperforms state-of-the-art methods in synthesizing view-consistent textures. Suin Lee, Dae-Shik Kim |
ICLR | 2 |
| 2025 | Interpretable fMRI Captioning via Contrastive Learning
Vyacheslav Shen, Kassymzhomart Kunanbayev, Donggon Jang, Dae-Shik Kim |
MICCAI (14) | 4 |
| 2025 | Generality-aware self-supervised transformer for multivariate time series anomaly detection
Yucheol Cho, Jae-Hyeok Lee 0001, Gyeongdo Ham, Donggon Jang, Dae-Shik Kim |
Appl. Intell. | 5 |
| 2025 | Teach sample-specific knowledge: Separated distillation based on samples
Seonghak Kim, Gyeongdo Ham, Suin Lee, Dae-Shik Kim |
Eng. Appl. Artif. Intell. | 4 |
| 2024 | Training ViT with Limited Data for Alzheimer's Disease Classification: An Empirical Study
Kassymzhomart Kunanbayev, Vyacheslav Shen, Dae-Shik Kim |
MICCAI (12) | 3 |
| 2024 | Robust Unsupervised Domain Adaptation through Negative-View RegularizationabstractIn the realm of Unsupervised Domain Adaptation (UDA), Vision Transformers (ViTs) have recently demonstrated remarkable adaptability surpassing that of traditional Convolutional Neural Networks (CNNs). Nevertheless, the patch-based structure of ViTs heavily relies on local features within image patches, potentially leading to reduced robustness when confronted with out-of-distribution (OOD) samples. To address this concern, we introduce a novel regularizer tailored specifically for UDA. By leveraging negative views, i.e. target-domain samples applied by negative augmentations, we make the learning process more intricate, thereby preventing models from taking shortcuts in spatial context recognition. We present a novel loss function, rooted in contrastive principles, to effectively distinguish between the negative views and original target samples. By integrating this novel regularizer with existing UDA methodologies, we guide ViTs to prioritize context relationships among local patches, thereby enhancing the robustness of ViTs. Our proposed Negative View-based Contrastive (NVC) regularizer substantially boosts the performance of baseline UDA methods across diverse benchmark datasets. Furthermore, we release new dataset, Retail-71, comprising 71 classes of images commonly encountered in retail stores. Through comprehensive experimentation, we showcase the effectiveness of our approach on traditional benchmarks as well as the novel retail domain. These results substantiate the robust adaptation capabilities of our proposed method. Our method is implemented at our repository. Joonhyeok Jang, Sunhyeok Lee, Seonghak Kim, Jung-Un Kim, Seonghyun Kim, Dae-Shik Kim |
WACV | 6 |
| 2024 | Difficulty level-based knowledge distillation
Gyeongdo Ham, Yucheol Cho, Jae-Hyeok Lee 0001, Minchan Kang, Gyuwon Choi, Dae-Shik Kim |
Neurocomputing | 6 |
| 2024 | Adaptive class token knowledge distillation for efficient vision transformer
Minchan Kang, Sanghyeok Son, Dae-Shik Kim |
Knowl. Based Syst. | 3 |
| 2024 | Maximizing discrimination capability of knowledge distillation with energy function
Seonghak Kim, Gyeongdo Ham, Suin Lee, Donggon Jang, Dae-Shik Kim |
Knowl. Based Syst. | 5 |
| 2024 | Energy-Based Domain Adaptation Without Intermediate Domain Dataset for Foggy Scene SegmentationabstractRobust segmentation performance under dense fog is crucial for autonomous driving, but collecting labeled real foggy scene datasets is burdensome in the real world. To this end, existing methods have adapted models trained on labeled clear weather images to the unlabeled real foggy domain. However, these approaches require intermediate domain datasets (e.g. synthetic fog) and involve multi-stage training, making them cumbersome and less practical for real-world applications. In addition, the issue of overconfident pseudo-labels by a confidence score remains less explored in self-training for foggy scene adaptation. To resolve these issues, we propose a new framework, named DAEN, which Directly Adapts without additional datasets or multi-stage training and leverages an ENergy score in self-training. Notably, we integrate a High-order Style Matching (HSM) module into the network to match high-order statistics between clear weather features and real foggy features. HSM enables the network to implicitly learn complex fog distributions without relying on intermediate domain datasets or multi-stage training. Furthermore, we introduce Energy Score-based Pseudo-Labeling (ESPL) to mitigate the overconfidence issue of the confidence score in self-training. ESPL generates more reliable pseudo-labels through a pixel-wise energy score, thereby alleviating bias and preventing the model from assigning pseudo-labels exclusively to head classes. Extensive experiments demonstrate that DAEN achieves state-of-the-art performance on three real foggy scene datasets and exhibits a generalization ability to other adverse weather conditions. Code is available at https://github.com/jdg900/daen. Donggon Jang, Sunhyeok Lee, Gyuwon Choi, Yejin Lee 0013, Sanghyeok Son, Dae-Shik Kim |
IEEE Trans. Image Process. | 6 |
| 2024 | Robustness-Reinforced Knowledge Distillation With Correlation Distance and Network PruningabstractThe improvement in the performance of efficient and lightweight models (i.e., the student model) is achieved through knowledge distillation (KD), which involves transferring knowledge from more complex models (i.e., the teacher model). However, most existing KD techniques rely on Kullback-Leibler (KL) divergence, which has certain limitations. First, if the teacher distribution has high entropy, the KL divergence's mode-averaging nature hinders the transfer of sufficient target information. Second, when the teacher distribution has low entropy, the KL divergence tends to excessively focus on specific modes, which fails to convey an abundant amount of valuable knowledge to the student. Consequently, when dealing with datasets that contain numerous confounding or challenging samples, student models may struggle to acquire sufficient knowledge, resulting in subpar performance. Furthermore, in previous KD approaches, we observed that data augmentation, a technique aimed at enhancing a model's generalization, can have an adverse impact. Therefore, we propose a Robustness-Reinforced Knowledge Distillation (R2KD) that leverages correlation distance and network pruning. This approach enables KD to effectively incorporate data augmentation for performance improvement. Extensive experiments on various datasets, including CIFAR-100, FGVR, TinyImagenet, and ImageNet, demonstrate our method's superiority over current state-of-the-art methods. Seonghak Kim, Gyeongdo Ham, Yucheol Cho, Dae-Shik Kim |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2023 | ICE-NeRF: Interactive Color Editing of NeRFs via Decomposition-Aware Weight OptimizationabstractNeural Radiance Fields (NeRFs) have gained considerable attention for their high-quality results in 3D scene reconstruction and rendering. Recently, there have been active studies on various tasks such as novel view synthesis and scene editing. However, editing NeRFs is challenging as accurately decomposing the desired area of 3D space and ensuring the consistency of edited results from different angles is difficult. In this paper, we propose ICE-NeRF, an Interactive Color Editing framework that performs color editing by taking a pre-trained NeRF and a rough user mask as input. Our proposed method performs the entire color editing process in only under a minute using a partial fine-tuning approach. To perform effective color editing, we address two issues: (1) the entanglement of the implicit representation that causes unwanted color changes in undesired areas when learning weights, and (2) the loss of multi-view consistency when fine-tuning for a single or a few views. To address these issues, we introduce two techniques: Activation Field-based Regularization (AFR) and Single-mask Multi-view Rendering (SMR). The AFR performs weight regularization during fine-tuning based on the assumption that not all weights have an equal impact on the desired area. The SMR maps the 2D mask to 3D space through inverse projection and renders it from other views to generate multi-view masks. ICE-NeRF not only enables well-decomposed, multi-view consistent color editing but also significantly reduces processing time compared to existing methods. Jae-Hyeok Lee 0001, Dae-Shik Kim |
ICCV | 2 |
| 2023 | Burst super-resolution with adaptive feature refinement and enhanced group up-sampling
Minchan Kang, Woo Jin Jeong, Sanghyeok Son, Gyeongdo Ham, Dae-Shik Kim |
Appl. Intell. | 5 |
| 2023 | Ambiguity-aware robust teacher (ART): Enhanced self-knowledge distillation framework with pruned teacher networkabstractSelf-knowledge distillation (self-KD) methods, which use a student model itself as the teacher model instead of a large and complex teacher model, are currently a subject of active study. Since most previous self-KD approaches relied on the knowledge of a single teacher model, if the teacher model incorrectly predicted confusing samples, poor-quality knowledge was transferred to the student model. Unfortunately, natural images are often ambiguous for teacher models due to multiple objects, mislabeling, or low quality. In this paper, we propose a novel knowledge distillation framework named ambiguity-aware robust teacher knowledge distillation (ART-KD) that provides refined knowledge, that reflects the ambiguity of the samples with network pruning. Since the pruned teacher model is simply obtained by copying and pruning the teacher model, re-training process is unnecessary in ART-KD. The key insight of ART-KD lies in the predictions of a teacher model and pruned teacher model for ambiguous samples providing different distributions with low similarity. From these two distributions, we obtain a joint distribution considering the ambiguity of the samples as teacher’s knowledge for distillation. We comprehensively evaluate our method on public classification benchmarks, as well as more challenging benchmarks for fine-grained visual recognition (FGVR), achieving much superior performance to state-of-the-art counterparts. Yucheol Cho, Gyeongdo Ham, Jae-Hyeok Lee 0001, Dae-Shik Kim |
Pattern Recognit. | 4 |
| 2022 | Learning Color Representations for Low-Light Image EnhancementabstractColor conveys important information about the visible world. However, under low-light conditions, both pixel intensity, as well as true color distribution, can be significantly shifted. Moreover, most of such distortions are non-recoverable due to inverse problems. In the present study, we utilized recent advancements in learning-based methods for low-light image enhancement. However, while most "deep learning" methods aim to restore high-level and object-oriented visual information, we hypothesized that learning-based methods can also be used for restoring color-based information. To address this question, we propose a novel color representation learning method for low-light image enhancement. More specifically, we used a channel-aware residual network and a differentiable intensity histogram to capture color features. Experimental results using synthetic and natural datasets suggest that the proposed learning scheme achieves state-of-the-art performance. We conclude from our study that inter-channel dependency and color distribution matching are crucial factors for learning color representations under low-light conditions. Bomi Kim, Sunhyeok Lee, Nahyun Kim, Donggon Jang, Dae-Shik Kim |
WACV | 5 |
| 2021 | Unsupervised Image Denoising with Frequency Domain Knowledge
Nahyun Kim, Donggon Jang, Sunhyeok Lee, Bomi Kim, Dae-Shik Kim |
BMVC | 5 |
| 2021 | Cross-Active Connection for Image-Text Multimodal Feature Fusion
JungHyuk Im, Wooyeong Cho, Dae-Shik Kim |
NLDB | 3 |
| 2020 | MHSAN: Multi-Head Self-Attention Network for Visual Semantic EmbeddingabstractVisual-semantic embedding enables various tasks such as image-text retrieval, image captioning, and visual question answering. The key to successful visual-semantic embedding is to express visual and textual data properly by accounting for their intricate relationship. While previous studies have achieved much advance by encoding the visual and textual data into a joint space where similar concepts are closely located, they often represent data by a single vector ignoring the presence of multiple important components in an image or text. Thus, in addition to the joint embedding space, we propose a novel multi-head self-attention network to capture various components of visual and textual data by attending to important parts in data. Our approach achieves the new state-of-the-art results in image-text retrieval tasks on MS-COCO and Flicker30K datasets. Through the visualization of the attention maps that capture distinct semantic components at multiple positions in the image and the text, we demonstrate that our method achieves an effective and interpretable visual-semantic joint space. Geondo Park, Chihye Han, Dae-Shik Kim, Wonjun Yoon |
WACV | 3 |
| 2019 | Progressive Face Super-Resolution via Attention to Facial Landmark
Deokyun Kim, Minseon Kim, Gihyun Kwon, Dae-Shik Kim |
BMVC | 4 |
| 2019 | Representation of white- and black-box adversarial examples in deep neural networks and humans: A functional magnetic resonance imaging studyabstractThe recent success of brain-inspired deep neural networks (DNNs) in solving complex, high-level visual tasks has led to rising expectations for their potential to match the human visual system. However, DNNs exhibit idiosyncrasies that suggest their visual representation and processing might be substantially different from human vision. One limitation of DNNs is that they are vulnerable to adversarial examples, input images on which subtle, carefully designed noises are added to fool a machine classifier. The robustness of the human visual system against adversarial examples is potentially of great importance as it could uncover a key mechanistic feature that machine vision is yet to incorporate. In this study, we compare the visual representations of white- and black-box adversarial examples in DNNs and humans by leveraging functional magnetic resonance imaging (fMRI). We find a small but significant difference in representation patterns for different (i.e. white- versus black-box) types of adversarial examples for both humans and DNNs. However, human performance on categorical judgment is not degraded by noise regardless of the type unlike DNN. These results suggest that adversarial examples may be differentially represented in the human visual system, but unable to affect the perceptual experience. Chihye Han, Wonjun Yoon, Gihyun Kwon, Dae-Shik Kim, Seungkyu Nam |
IJCNN | 4 |
| 2019 | Generation of 3D Brain MRI Using Auto-Encoding Generative Adversarial Networks
Gihyun Kwon, Chihye Han, Dae-Shik Kim |
MICCAI (3) | 3 |
| 2019 | Latent Question Interpretation Through Variational AdaptationabstractMost artificial neural network models for question-answering rely on complex attention mechanisms. These techniques demonstrate high performance on existing datasets; however, they are limited in their ability to capture natural language variability, and to generate diverse relevant answers. To address this limitation, we propose a model that learns multiple interpretations of a given question. This diversity is ensured by our interpretation policy module which automatically adapts the parameters of a question-answering model with respect to a discrete latent variable. This variable follows the distribution of interpretations learned by the interpretation policy through a semi-supervised variational inference framework. To boost the performance further, the resulting policy is fine-tuned using the rewards from the answer accuracy with a policy gradient. We demonstrate the relevance and efficiency of our model through a large panel of experiments. Qualitative results, in particular, underline the ability of the proposed architecture to discover multiple interpretations of a question. When tested using the Stanford Question Answering Dataset 1.1, our model outperforms the baseline methods in finding multiple and diverse answers. To assess our strategy from a human standpoint, we also conduct a large-scale user study. This study highlights the ability of our network to produce diverse and coherent answers compared to existing approaches. Our Pytorch implementation is available as open source.11github.com/parshakova/APIP. Tetiana Parshakova, François Rameau, Andriy Serdega, In-So Kweon, Dae-Shik Kim |
IEEE ACM Trans. Audio Speech Lang. Process. | 5 |
| 2018 | End-to-End Speech Command Recognition with Capsule Network
Jaesung Bae, Dae-Shik Kim |
INTERSPEECH | 2 |
| 2018 | Competitive Hyperparameter Balancing on Spiking Neural Network for a Fast, Accurate and Energy-Efficient Inference
Dae-Shik Kim |
ISNN | 2 |
| 2014 | Quantification and reduction of visual load during BCI operationabstractOperating a brain-actuated vehicle in real-world environments requires much of our visual attention. However, a typical brain-computer interface (BCI) sends the feedback information about the current status of the user's brain also via the visual channel. As a result, users have to split their visual attention into two: one for the surroundings and the other for the visual BCI feedback. Therefore, we recently developed a tactile stimulation system that successfully replaced the conventional visual feedback. Here we employ the multiple object tracking experiments to quantify the visual load added by the visual feedback. The result show that the additional visual load is almost eliminated, and the true negative rate of the BCI operation (intentional non-control) is improved when the visual feedback is replaced by the tactile feedback. Kiuk Gwak, Robert Leeb, José del R. Millán, Dae-Shik Kim |
SMC | 4 |
| 2013 | Planning as Inference in a Hierarchical Predictive Memory
Hansol Choi, Dae-Shik Kim |
ICONIP (1) | 2 |
| 2013 | Multiple Kernel Learning with Hierarchical Feature Representations
Juhyeon Lee, Jae Hyun Lim 0001, Hyungwon Choi, Dae-Shik Kim |
ICONIP (3) | 4 |
| 2012 | Apparent Volitional Behavior Selection Based on Memory Predictions
Jun-Cheol Park, Jae Hyeon Yoo, Juhyeon Lee, Dae-Shik Kim |
ICONIP (1) | 4 |
| 2012 | Reward hierarchical temporal memoryabstractIn humans and animals, reward prediction error encoded by dopamine systems is thought to be important in the temporal difference learning class of reinforcement learning (RL). With RL algorithms, many brain models have described the function of dopamine and related areas, including the basal ganglia and frontal cortex. In spite of this importance, how the reward prediction error itself is computed is not understood well, including the problem of how the current states are assigned to a memorized states and how the values of the states are memorized. In this paper, we describe a neocortical model for memorizing state space and computing reward prediction error, known as `reward hierarchical temporal memory' (rHTM). In this model, the temporal relationships among events are hierarchically stored. Using this memory, rHTM computes reward prediction errors by associating the memorized sequences to rewards and inhibits the predicted reward. In a simulation, our model behaved similarly to dopaminergic neurons. We suggest that our model can provide a hypothetical framework of interaction between cortex and dopamine neurons. Hansol Choi, Jun-Cheol Park, Jae Hyun Lim 0001, Jae Young Jun, Dae-Shik Kim |
IJCNN | 5 |
| 2012 | Learning spatio-temporally invariant representations from videoabstractLearning invariant representations of environments through experience has been important area of research both in the field of machine learning as well as in computational neuroscience. In the present study, we propose a novel unsupervised method for the discovery of invariants from a single video input based on the learning of the spatio-temporal relationship of inputs. In an experiment, we tested the learning of spatio-temporal invariant features from a single video that involves rotational movements of faces of several subjects. From the results of this experiment, we demonstrate that the proposed system for the learning of invariants based on spatio-temporal continuity can be used as a compelling unsupervised method for learning invariants from an input that includes temporal information. Jae Hyun Lim 0001, Hansol Choi, Jun-Cheol Park, Jae Young Jun, Dae-Shik Kim |
IJCNN | 5 |
| 2001 | Magnetic resonance imaging of brain function and neurochemistryabstractIn the past decade, magnetic resonance imaging (MRI) research has been focused on the acquisition of physiological and biochemical information noninvasively. Probably the most notable accomplishment in this general effort has been the introduction of the MR approaches to map brain function. This capability, often referred to as functional magnetic resonance imaging, or fMRI, is based on the sensitivity of MR signals to secondary metabolic and hemodynamic responses that accompany increased neuronal activity. Despite this indirect link to neurotransmission, recent studies demonstrate that under appropriate conditions, these fMRI maps have accuracy at the scale of submillimeter neuronal organizations such as the orientation columns of the visual cortex, and are directly proportional in magnitude to electrical signals generated by the neurons. High magnetic fields have been critical in achieving such specificity in functional maps because they provide advantages through increased signal-to-noise ratio, diminishing blood-related contributions to mapping signals, and enhanced sensitivity to microvasculature. Equally important is MR spectroscopy studies, which, at high magnetic fields, provide for the first time the opportunity to measure local metabolic correlates of human brain function and neurotransmission rates. Together, these MR methods provide a complementary set of approaches for probing important aspects of the nervous system. Kâmil Ugurbil, Dae-Shik Kim, Timothy Q. Duong, Xiaoping Hu 0001, Seiji Ogawa, Rolf Gruetter, Wei Chen 0086, Seong-Gi Kim, Xiao-Hung Zhu, Essa Yacoub, Pierre-François van de Moortele, Amir Shmuel, Josef Pfeuffer, Hellmut Merkle, Peter Andersen, Gregor Adriany |
Proc. IEEE | 2 |
| 1990 | Computation of large-scale constrained matrix problems: the splitting equilibration algorithmabstractThe authors introduce a general parallelizable computational method called the splitting equilibration algorithm for solving the entire class of constrained matrix problems. The empirical performance of the algorithm is investigated on the largest quadratic constrained matrix problems reported to date using the IBM 3090-600E at the Cornell National Supercomputer Facility in a serial and in a parallel environment. The goals are to compare the relative efficiency of the splitting equilibration algorithm to both the earlier equilibration algorithm and the much-cited Bachem and Korte algorithm, (1978) and to investigate the speedups obtained with parallelization of the splitting equilibration algorithm.> Anna Nagurney, Alexander Eydeland, Dae-Shik Kim |
SC | 3 |