VLDB 2026 Research / reviewers in the wild / expert
Ilhwan Kim
dblp:16/2774
· DBLP profile ↗
11ranked-venue papers
2as first author
5since 2021 · last 2026
0009-0004-2739-8001ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 5 · 3 since 2021Systems, architecture and hardware · 2Human-computer interaction and ubiquitous computing · 2Applied, interdisciplinary, general and emerging computing · 2Computer networks · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | NAPP: Noise-Adaptive Prototype Perturbation for Few-Shot LearningabstractFew-shot learning aims to generalize deep models to novel categories with only a handful of labeled examples, but existing methods remain vulnerable to task-irrelevant noise, unstable prototype estimation, and limited adaptability under domain shift. To address these issues, we propose the Noise-Adaptive Prototype Perturbation Network (NAPP), a framework that enhances robustness and generalization for few-shot learning. NAPP introduces three key innovations: (1) a Noise Cancellation Mechanism embedded in Vision Transformer self-attention layers that dynamically suppresses spurious, task-irrelevant features. (2) a Mix-Perturbation Module that perturbs class prototypes through augmented feature combinations, producing more stable and transferable prototype representations. (3) an Adaptive Noise-Conditioned Meta-Learning scheme that finetunes less than 0.02% of noise-related parameters at metatest time, enabling efficient and rapid adaptation to unseen classes without eroding pretrained knowledge. Extensive experiments demonstrate that NAPP achieves competitive and superior performance compared to state-of-the-art fewshot classification methods across both in-domain and challenging cross-domain benchmarks. Ilhwan Kim, Sangwoo Yun, Seongsu Kim, Joonki Paik |
WACV | 1 |
| 2024 | Faces that Speak: Jointly Synthesising Talking Face and Speech from TextabstractThe goal of this work is to simultaneously generate natural talking faces and speech outputs from text. We achieve this by integrating Talking Face Generation (TFG) and Text-to-Speech (TTS) systems into a unified framework. We address the main challenges of each task: (1) generating a range of head poses representative of real-world scenarios, and (2) ensuring voice consistency despite variations in facial motion for the same identity. To tackle these issues, we introduce a motion sampler based on conditional flow matching, which is capable of high-quality motion code generation in an efficient way. Moreover, we introduce a novel conditioning method for the TTS system, which utilises motion-removed features from the TFG model to yield uniform speech outputs. Our extensive experiments demonstrate that our method effectively creates natural-looking talking faces and speech that accurately match the input text. To our knowledge, this is the first effort to build a multimodal synthesis system that can generalise to unseen identities. Youngjoon Jang 0001, Junseok Ahn, Doyeop Kwak, Hongsun Yang, Yooncheol Ju, Ilhwan Kim, Byeong-Yeol Kim, Joon Son Chung |
CVPR | 7 |
| 2023 | CROSSSPEECH: Speaker-Independent Acoustic Representation for Cross-Lingual Speech SynthesisabstractWhile recent text-to-speech (TTS) systems have made remarkable strides toward human-level quality, the performance of cross-lingual TTS lags behind that of intra-lingual TTS. This gap is mainly rooted from the speaker-language entanglement problem in cross-lingual TTS. In this paper, we propose CrossSpeech which improves the quality of cross-lingual speech by effectively disentangling speaker and language information in the level of acoustic feature space. Specifically, CrossSpeech decomposes the speech generation pipeline into the speaker-independent generator (SIG) and speaker-dependent generator (SDG). The SIG produces the speaker-independent acoustic representation which is not biased to specific speaker distributions. On the other hand, the SDG models speaker-dependent speech variation that characterizes speaker attributes. By handling each information separately, CrossSpeech can obtain disentangled speaker and language representations. From the experiments, we verify that CrossSpeech achieves significant improvements in cross-lingual TTS, especially in terms of speaker similarity to the target speaker. Hongsun Yang, Yooncheol Ju, Ilhwan Kim, Byeong-Yeol Kim |
ICASSP | 4 |
| 2023 | FACTSpeech: Speaking a Foreign Language Pronunciation Using Only Your Native Characters
Hongsun Yang, Yooncheol Ju, Ilhwan Kim, Byeong-Yeol Kim, Shukjae Choi, Hyung Yong Kim |
INTERSPEECH | 4 |
| 2022 | TriniTTS: Pitch-controllable End-to-end TTS without External Aligner
Yooncheol Ju, Ilhwan Kim, Hongsun Yang, Byeong-Yeol Kim, Soumi Maiti, Shinji Watanabe 0001 |
INTERSPEECH | 2 |
| 2015 | Applying GPGPU to recurrent neural network language model based fast network search in the real-time LVCSRabstractRecurrent Neural Network Language Models (RNNLMs) have started to be used in various fields of speech recognition due to their outstanding performance. However, the high computational complexity of RNNLMs has been a hurdle in applying the RNNLM to a real-time Large Vocabulary Continuous Speech Recognition (LVCSR). In order to accelerate the speed of RNNLM-based network searches during decoding, we apply the General Purpose Graphic Processing Units (GPGPUs). This paper proposes a novel method of applying GPGPUs to RNNLM-based graph traversals. We have achieved our goal by reducing redundant computations on CPUs and amount of transfer between GPGPUs and CPUs. The proposed approach was evaluated on both WSJ corpus and in-house data. Experiments shows that the proposed approach achieves the real-time speed in various circumstances while maintaining the Word Error Rate (WER) to be relatively 10% lower than that of n-gram models. Kyungmin Lee, Chiyoun Park, Ilhwan Kim, Namhoon Kim |
INTERSPEECH | 3 |
| 2013 | Development of Gazing Interface for Operation of a Page Turner MachineabstractThis paper describes a man machine interface based on gazing input without user's body restrictions. Gazing points forward to the side directions were estimated by the relative distance of an iris and the outer corner of left and right eyes. Whether or not the user gazes the camera was recognized by the intersection of the two perpendicular bisectors derived by the position of a cornea and the eyelid edges. The suggested recognition system was applied to the operations of the page turner machine. It could judge the user's states of both reading a book and operating the machine by the special feature points on an eye. Nobuaki Nakazawa, Ilhwan Kim, Hiroki Murakawa, Toshikazu Matsui |
SMC | 2 |
| 2012 | Hand-free interface based on facial orientationsabstractThis paper suggests a hand-free interface based on the facial orientations. The operator's face was observed by the USB camera and the changes in the darkness area of the both nostrils were utilized for recognition of the face orientations. When the operator faced to up and downward, the darkness areas of both nostrils were increased and decreased, respectively. On the other hand, the difference between two nostril areas could be caused in cases where the face was turn to the side. Here, these characteristics were reflected to the recognition of the face orientations. The facial orientations were applied to the auto-wheelchair operations, instead of a joystick device. Nobuaki Nakazawa, Takashi Mori, Ilhwan Kim, Toshikazu Matsui, Kou Yamada, Aya Maeda |
IECON | 3 |
| 2012 | Development of an intuitive interface based on facial orientations and gazing actions for auto-wheel chair operationabstractThis paper suggests an intuitive interface based on the facial orientations and gazing actions for auto-wheel chair operation. The real-time image of the operator's face was taken to the computer through the USB camera to observe the operational intention. The changes in the darkness area of the both nostrils were utilized for recognition of the face orientations. When the operator faced to up and downward, the darkness areas of both nostrils were increased and decreased, respectively. On the other hand, the difference between two nostril areas could be caused in cases where the face was turn to the side. Here, these characteristics were utilized for the recognition of the face orientations. Moreover, gazing actions was recognized by the curve ratio of the operator's eye lines. Only when the operator gazed to the control computer, the facial orientation was reflected to operate the auto-wheelchair, instead of a joystick interface. Nobuaki Nakazawa, Ilhwan Kim, Takashi Mori, Hiroki Murakawa, Motohiro Kano, Aya Maeda, Toshikazu Matsui, Kou Yamada |
RO-MAN | 2 |
| 2006 | Mixing Heterogeneous Address Spaces in a Single Edge Network
Ilhwan Kim, Heon Young Yeom |
APNOMS | 1 |
| 1996 | IP Multiplexing by Transparent Port-Address Translator
Heon Young Yeom, Jungsoo Ha, Ilhwan Kim |
LISA | 3 |