Taehwan Kim 0013

dblp:86/3976-13 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
5since 2021 · last 2025
0000-0002-6571-4632ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Generative modeling · 100%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
1.422024
Grid Diffusion Models for Text-to-Video Generation · CVPR 2024
Generating Realistic Images from In-the-wild Sounds · ICCV 2023
Machine learning › Generative modeling › video generation
text-to-video generation
0.812024
Grid Diffusion Models for Text-to-Video Generation · CVPR 2024
Machine learning › Generative modeling › cross-modal generation
audio-to-image generation
0.712023
Generating Realistic Images from In-the-wild Sounds · ICCV 2023
Machine learning › Generative modeling
cross-modal generation
0.712023
Generating Realistic Images from In-the-wild Sounds · ICCV 2023
Machine learning › Generative modeling › diffusion model
text-to-image generation
0.712023
Generating Realistic Images from In-the-wild Sounds · ICCV 2023

Methods — techniques the papers use, named apart from their topics

grid diffusion · 0.83d u-net · 0.8sentence attention · 0.7diffusion model · 0.7audio captioning · 0.7audio attention · 0.7CLIPScore · 0.7AudioCLIP · 0.7
YearPublicationVenuePosition
2025 DiffListener: Discrete Diffusion Model for Listener Generation
abstract
The listener head generation (LHG) task aims to generate natural nonverbal listener responses based on the speaker’s multimodal cues. While prior work either rely on limited modalities (e.g. audio and facial information) or employ autoregressive approaches which have limitations such as accumulating prediction errors. To address these limitations, we propose DiffListener, a discrete diffusion based approach for non-autoregressive listener head generation. Our model takes the speaker’s facial information, audio, and text as inputs, additionally incorporating facial differential information to represent the temporal dynamics of expressions and movements. With this explicit modeling of facial dynamics, DiffListener can generate coherent reaction sequences in a non-autoregressive manner. Through comprehensive experiments, DiffListener demonstrates state-of-the-art performance in both quantitative and qualitative evaluations. The user study shows that DiffListener generates natural context-aware listener reactions that are well synchronized with the speaker. The code and demo videos are available in https://siyeoljung.github.io/DiffListener.
Siyeol Jung, Taehwan Kim 0013
ICASSP2
2024 Grid Diffusion Models for Text-to-Video Generation
abstract
Recent advances in the diffusion models have significantly improved text-to-image generation. However, generating videos from text is a more challenging task than generating images from text, due to the much larger dataset and higher computational cost required. Most existing video generation methods use either a 3D U-Net architecture that considers the temporal dimension or autoregressive generation. These methods require large datasets and are limited in terms of computational costs compared to text-to-image generation. To tackle these challenges, we propose a simple but effective novel grid diffusion for text-to-video generation without temporal dimension in architecture and a large text-video paired dataset. We can generate a high-quality video using a fixed amount of GPU memory regardless of the number of frames by representing the video as a grid image. Additionally, since our method reduces the dimensions of the video to the dimensions of the image, various image-based methods can be applied to videos, such as textguided video manipulation from image manipulation. Our proposed method outperforms the existing methods in both quantitative and qualitative evaluations, demonstrating the suitability of our model for real-world video generation.
Taegyeong Lee, Soyeong Kwon, Taehwan Kim 0013
CVPR3
2023 Effective Slogan Generation with Noise Perturbation
abstract
Slogans play a crucial role in building the brand's identity of the firm. A slogan is expected to reflect firm's vision and the brand's value propositions in memorable and likeable ways. Automating the generation of slogans with such characteristics is challenging. Previous studies developed and tested slogan generation with syntactic control and summarization models which are not capable of generating distinctive slogans. We introduce a novel approach that leverages pre-trained transformer T5 model with noise perturbation on newly proposed 1:N matching pair dataset. This approach serves as a contributing factor in generating distinctive and coherent slogans. Furthermore, the proposed approach incorporates descriptions about the firm and brand into the generation of slogans. We evaluate generated slogans based on ROUGE-1, ROUGE-L and Cosine Similarity metrics and also assess them with human subjects in terms of slogan's distinctiveness, coherence, and fluency. The results demonstrate that our approach yields better performance than baseline models and other transformer-based models.
MinChung Kim, Taehwan Kim 0013
CIKM3
2023 Generating Realistic Images from In-the-wild Sounds
abstract
Representing wild sounds as images is an important but challenging task due to the lack of paired datasets between sound and images and the significant differences in the characteristics of these two modalities. Previous studies have focused on generating images from sound in limited categories or music. In this paper, we propose a novel approach to generate images from in-the-wild sounds. First, we convert sound into text using audio captioning. Second, we propose audio attention and sentence attention to represent the rich characteristics of sound and visualize the sound. Lastly, we propose a direct sound optimization with CLIPscore and AudioCLIP and generate images with a diffusion-based model. In experiments, it shows that our model is able to generate high quality images from wild sounds and outperforms baselines in both quantitative and qualitative evaluations on wild audio datasets.
Taegyeong Lee, Jeonghun Kang, Hyeonyu Kim, Taehwan Kim 0013
ICCV4
2022 An Empirical Study on How People Perceive AI-generated Music
abstract
Music creation is difficult because one must express one's creativity while following strict rules. The advancement of deep learning technologies has diversified the methods to automate complex processes and express creativity in music composition. However, prior research has not paid much attention to exploring the audiences' subjective satisfaction to improve music generation models. In this paper, we evaluate human satisfaction with the state-of-the-art automatic symbolic music generation models using deep learning. In doing so, we define a taxonomy for music generation models and suggest nine subjective evaluation metrics. Through an evaluation study, we obtained more than 700 evaluations from 100 participants, using the suggested metrics. Our evaluation study reveals that the token representation method and models' characteristics affect subjective satisfaction. Through our qualitative analysis, we deepen our understanding of AI-generated music and suggested evaluation metrics. Lastly, we present lessons learned and discuss future research directions of deep learning models for music creation.
Hyeshin Chu, Joohee Kim, Seongouk Kim, Hongkyu Lim, Hyunwook Lee, Seungmin Jin, Jongeun Lee, Taehwan Kim 0013, Sungahn Ko
CIKM8