EDBT 2026 Demo / reviewers in the wild / expert
Joanna Hong
dblp:255/6341
· DBLP profile ↗
15ranked-venue papers
6as first author
11since 2021 · last 2024
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 12 · 5 first-author · 8 since 2021Artificial intelligence and machine learning · 10 · 4 first-author · 9 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Let's Go Real Talk: Spoken Dialogue Model for Face-to-Face ConversationabstractIn this paper, we introduce a novel Face-to-Face spoken dialogue model.It processes audio-visual speech from user input and generates audio-visual speech as the response, marking the initial step towards creating an avatar chatbot system without relying on intermediate text.To this end, we newly introduce MultiDialog, the first large-scale multimodal (i.e., audio and visual) spoken dialogue corpus containing 340 hours of approximately 9,000 dialogues, recorded based on the open domain dialogue dataset, TopicalChat.The MultiDialog contains parallel audio-visual recordings of conversation partners acting according to the given script with emotion annotations, which we expect to open up research opportunities in multimodal synthesis.Our Face-to-Face spoken dialogue model incorporates a textually pretrained large language model and adapts it into the audio-visual spoken dialogue domain by incorporating speech-text joint pretraining.Through extensive experiments, we validate the effectiveness of our model in facilitating a face-to-face conversation. Se Jin Park, Chae Won Kim, Hyeongseop Rha, Minsu Kim 0001, Joanna Hong, Jeong Hun Yeo, Yong Man Ro |
ACL (1) | 5 |
| 2023 | Watch or Listen: Robust Audio-Visual Speech Recognition with Visual Corruption Modeling and Reliability ScoringabstractThis paper deals with Audio-Visual Speech Recognition (AVSR) under multimodal input corruption situations where audio inputs and visual inputs are both corrupted, which is not well addressed in previous research directions. Previous studies have focused on how to complement the corrupted audio inputs with the clean visual inputs with the assumption of the availability of clean visual inputs. However, in real life, clean visual inputs are not always accessible and can even be corrupted by occluded lip regions or noises. Thus, we firstly analyze that the previous AVSR models are not indeed robust to the corruption of multimodal input streams, the audio and the visual inputs, compared to uni-modal models. Then, we design multimodal input corruption modeling to develop robust AVSR models. Lastly, we propose a novel AVSR framework, namely Audio-Visual Reliability Scoring module (AV-RelScore), that is robust to the corrupted multimodal inputs. The AV-RelScore can determine which input modal stream is reliable or not for the prediction and also can exploit the more reliable streams in prediction. The effectiveness of the proposed method is evaluated with comprehensive experiments on popular benchmark databases, LRS2 and LRS3. We also show that the reliability scores obtained by AV-RelScore well reflect the degree of corruption and make the proposed model focus on the reliable multimodal representations. Joanna Hong, Minsu Kim 0001, Jeongsoo Choi, Yong Man Ro |
CVPR | 1 |
| 2023 | Lip-to-Speech Synthesis in the Wild with Multi-Task LearningabstractRecent studies have shown impressive performance in Lip-to-speech synthesis that aims to reconstruct speech from visual information alone. However, they have been suffering from synthesizing accurate speech in the wild, due to insufficient supervision for guiding the model to infer the correct content. Distinct from the previous methods, in this paper, we develop a powerful Lip2Speech method that can reconstruct speech with correct contents from the input lip movements, even in a wild environment. To this end, we design multitask learning that guides the model using multimodal supervision, i.e. text and audio, to complement the insufficient word representations of acoustic feature reconstruction loss. Thus, the proposed framework brings the advantage of synthesizing speech containing the right content of multiple speakers with unconstrained sentences. We verify the effectiveness of the proposed method using LRS2, LRS3, and LRW datasets. Minsu Kim 0001, Joanna Hong, Yong Man Ro |
ICASSP | 2 |
| 2023 | DiffV2S: Diffusion-based Video-to-Speech Synthesis with Vision-guided Speaker EmbeddingabstractRecent research has demonstrated impressive results in video-to-speech synthesis which involves reconstructing speech solely from visual input. However, previous works have struggled to accurately synthesize speech due to a lack of sufficient guidance for the model to infer the correct content with the appropriate sound. To resolve the issue, they have adopted an extra speaker embedding as a speaking style guidance from a reference auditory information. Nevertheless, it is not always possible to obtain the audio information from the corresponding video input, especially during the inference time. In this paper, we present a novel vision-guided speaker embedding extractor using a self-supervised pretrained model and prompt tuning technique. In doing so, the rich speaker embedding information can be produced solely from input visual information, and the extra audio information is not necessary during the inference time. Using the extracted vision-guided speaker embedding representations, we further develop a diffusion-based video-to-speech synthesis model, so called DiffV2S, conditioned on those speaker embeddings and the visual representation extracted from the input video. The proposed DiffV2S not only maintains phoneme details contained in the input video frames, but also creates a highly intelligible mel-spectrogram in which the speaker identities of the multiple speakers are all preserved. Our experimental results show that DiffV2S achieves the state-of-the-art performance compared to the previous video-to-speech synthesis technique. Jeongsoo Choi, Joanna Hong, Yong Man Ro |
ICCV | 2 |
| 2022 | SyncTalkFace: Talking Face Generation with Precise Lip-Syncing via Audio-Lip MemoryabstractThe challenge of talking face generation from speech lies in aligning two different modal information, audio and video, such that the mouth region corresponds to input audio. Previous methods either exploit audio-visual representation learning or leverage intermediate structural information such as landmarks and 3D models. However, they struggle to synthesize fine details of the lips varying at the phoneme level as they do not sufficiently provide visual information of the lips at the video synthesis step. To overcome this limitation, our work proposes Audio-Lip Memory that brings in visual information of the mouth region corresponding to input audio and enforces fine-grained audio-visual coherence. It stores lip motion features from sequential ground truth images in the value memory and aligns them with corresponding audio features so that they can be retrieved using audio input at inference time. Therefore, using the retrieved lip motion features as visual hints, it can easily correlate audio with visual dynamics in the synthesis step. By analyzing the memory, we demonstrate that unique lip features are stored in each memory slot at the phoneme level, capturing subtle lip motion based on memory addressing. In addition, we introduce visual-visual synchronization loss which can enhance lip-syncing performance when used along with audio-visual synchronization loss in our model. Extensive experiments are performed to verify that our method generates high-quality video with mouth shapes that best align with the input audio, outperforming previous state-of-the-art methods. Se Jin Park, Minsu Kim 0001, Joanna Hong, Jeongsoo Choi, Yong Man Ro |
AAAI | 3 |
| 2022 | VisageSynTalk: Unseen Speaker Video-to-Speech Synthesis via Speech-Visage Feature Selection
Joanna Hong, Minsu Kim 0001, Yong Man Ro |
ECCV (36) | 1 |
| 2022 | Visual Context-driven Audio Feature Enhancement for Robust End-to-End Audio-Visual Speech RecognitionabstractThis paper focuses on designing a noise-robust end-to-end Audio-Visual Speech Recognition (AVSR) system.To this end, we propose Visual Context-driven Audio Feature Enhancement module (V-CAFE) to enhance the input noisy audio speech with a help of audio-visual correspondence.The proposed V-CAFE is designed to capture the transition of lip movements, namely visual context and to generate a noise reduction mask by considering the obtained visual context.Through context-dependent modeling, the ambiguity in viseme-to-phoneme mapping can be refined for mask generation.The noisy representations are masked out with the noise reduction mask resulting in enhanced audio features.The enhanced audio features are fused with the visual features and taken to an encoder-decoder model composed of Conformer and Transformer for speech recognition.We show the proposed end-to-end AVSR with the V-CAFE can further improve the noise-robustness of AVSR.The effectiveness of the proposed method is evaluated in noisy speech recognition and overlapped speech recognition experiments using the two largest audio-visual datasets, LRS2 and LRS3. Joanna Hong, Minsu Kim 0001, Daehun Yoo, Yong Man Ro |
INTERSPEECH | 1 |
| 2022 | CroMM-VSR: Cross-Modal Memory Augmented Visual Speech RecognitionabstractVisual Speech Recognition (VSR) is a task that recognizes speech from external appearances of the face (${\it i}.{\it e}.$, lips) into text. Since the information from the visual lip movements is not sufficient to fully represent the speech, VSR is considered as one of the challenging problems. One possible way to resolve this problem is additionally utilizing audio which contains rich information for speech recognition. However, the audio information could not be always available such as in crowded situations. Thus, it is necessary to find a way that successfully provides enough information for speech recognition with visual inputs only. In this paper, we alleviate the information insufficiency of visual lip movement by proposing a cross-modal memory augmented VSR with Visual-Audio Memory (VAM). The proposed framework tries to utilize the complementary information of audio even when the audio inputs are not provided at the inference time. Concretely, the proposed VAM learns to imprint audio features of short clip-level into a memory network using the corresponding visual features. To this end, the VAM contains two memories, lip-video key and audio value. We guide the audio value memory to imprint the audio feature and the lip-video key memory to memorize the location of the imprinted audio. By doing this, the VAM can exploit rich audio information by accessing the memory using visual inputs only. Experimental results show that the proposed method achieves state-of-the-art performance on both word- and sentence-level VSR. In addition, we verify the learned representations inside the VAM contain meaningful information for VSR. Minsu Kim 0001, Joanna Hong, Se Jin Park, Yong Man Ro |
IEEE Trans. Multim. | 2 |
| 2021 | Multi-modality Associative Bridging through Memory: Speech Sound Recollected from Face VideoabstractIn this paper, we introduce a novel audio-visual multi-modal bridging framework that can utilize both audio and visual information, even with uni-modal inputs. We exploit a memory network that stores source (i.e., visual) and target (i.e., audio) modal representations, where source modal representation is what we are given, and target modal representations are what we want to obtain from the memory network. We then construct an associative bridge between source and target memories that considers the inter-relationship between the two memories. By learning the interrelationship through the associative bridge, the proposed bridging framework is able to obtain the target modal representations inside the memory network, even with the source modal input only, and it provides rich information for its downstream tasks. We apply the proposed framework to two tasks: lip reading and speech reconstruction from silent video. Through the proposed associative bridge and modality-specific memories, each task knowledge is enriched with the recalled audio context, achieving state-of-the-art performance. We also verify that the associative bridge properly relates the source and target memories. Minsu Kim 0001, Joanna Hong, Se Jin Park, Yong Man Ro |
ICCV | 2 |
| 2021 | Lip to Speech Synthesis with Visual Context Attentional GANabstractIn this paper, we propose a novel lip-to-speech generative adversarial network, Visual Context Attentional GAN (VCA-GAN), which can jointly model local and global lip movements during speech synthesis. Specifically, the proposed VCA-GAN synthesizes the speech from local lip visual features by finding a mapping function of viseme-to-phoneme, while global visual context is embedded into the intermediate layers of the generator to clarify the ambiguity in the mapping induced by homophene. To achieve this, a visual context attention module is proposed where it encodes global representations from the local visual features, and provides the desired global visual context corresponding to the given coarse speech representation to the generator through audio-visual attention. In addition to the explicit modelling of local and global visual representations, synchronization learning is introduced as a form of contrastive learning that guides the generator to synthesize a speech in sync with the given input lip movements. Extensive experiments demonstrate that the proposed VCA-GAN outperforms existing state-of-the-art and is able to effectively synthesize the speech from multi-speaker that has been barely handled in the previous works. Minsu Kim 0001, Joanna Hong, Yong Man Ro |
NeurIPS | 2 |
| 2021 | Speech Reconstruction With Reminiscent Sound Via Visual Voice MemoryabstractThe goal of this work is to reconstruct speech from silent video, in both speaker dependent and speaker independent ways. Unlike previous works that have been mostly restricted to a speaker dependent setting, we propose Visual Voice memory to restore essential auditory information to generate proper speech from different speakers and even unseen speakers. The proposed memory takes additional auditory information that corresponds to the input face movements and stores the auditory contexts that can be recalled by the given input visual features. Specifically, the Visual Voice memory contains value and key memory slots, where value memory slots are for saving the audio features, and key memory slots are for storing the visual features in the same location of the saved audio features. Guiding each memory to properly save each feature, the model can adequately produce the speech through auxiliary information of audio. Hence, our method employs both video and audio information during training time, but does not require any additional auditory input in the inference time. Our key contributions are: (1) proposing the Visual Voice memory that brings rich information of audio that complements the visual features, thus producing high-quality speech from silent video, and (2) enabling multi-speaker and speaker independent training by memorizing auditory features and the corresponding visual features. We validate the proposed framework on GRID and Lip2Wav datasets and show that our method surpasses the performance of previous works. Moreover, we experiment on both multi-speaker and speaker independent settings and verify the effectiveness of the Visual Voice memory. We also demonstrate that the Visual Voice memory contains meaningful information to reconstruct speech. Joanna Hong, Minsu Kim 0001, Se Jin Park, Yong Man Ro |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2020 | Comprehensive Facial Expression Synthesis Using Human-Interpretable LanguageabstractRecent advances in facial expression synthesis have shown promising results using diverse expression representations including facial action units. Facial action units for an elaborate facial expression synthesis need to be intuitively represented for human comprehension, not a numeric categorization of facial action units. To address this issue, we utilize human-friendly approach: use of natural language where language helps human grasp conceptual contexts. In this paper, therefore, we propose a new facial expression synthesis model from language-based facial expression description. Our method can synthesize the facial image with detailed expressions. In addition, effectively embedding language features on facial features, our method can control individual word to handle each part of facial movement. Extensive qualitative and quantitative evaluations were conducted to verify the effectiveness of the natural language. Joanna Hong, Jung Uk Kim, Sangmin Lee 0001, Yong Man Ro |
ICIP | 1 |
| 2020 | Learning Style Correlation for Elaborate Few-Shot ClassificationabstractFew-shot classification is defined as a task where the network aims to classify unseen classes given only a few samples. Recent approaches, especially metric-based methods, have great progress in few-shot classification. However, the existing metric-based methods have a limitation in deploying discriminative features for elaborate comparison. They usually extract features from the embedding network without direct consideration of the relationship between support and query sets. To address the relationship, we propose a novel architecture, Style Correlated Module (SCM) to learn style correlation between support and query sets for few-shot classification. The proposed module leads support and query feature maps to focus on significant style correlated features and encourage the metric network to conduct an elaborate comparison. Furthermore, the proposed module can be generally applied to the existing metric-based approaches by adding the SCM behind the embedding network. We evaluate our proposed method with comprehensive experiments on two publicly available datasets and demonstrate its effectiveness with comparable results. Minsu Kim 0001, Jung Uk Kim, Hong Joo Lee 0001, Sangmin Lee 0001, Joanna Hong, Yong Man Ro |
ICIP | 6 |
| 2020 | Unsupervised Disentangling of Viewpoint and Residues Variations by Substituting Representations for Robust Face RecognitionabstractIt is well-known that identity-unrelated variations (e.g., viewpoint or illumination) degrade the performances of face recognition methods. In order to handle this challenge, a robust method for disentangling the identity and view representations has drawn an attention in the machine learning area. However, existing methods learn discriminative features which require a manual supervision of such factors of variations. In this paper, we propose a novel disentangling framework through modeling three representations of identity, viewpoint, and residues (i.e., identity and pose unrelated) which do not require supervision of the variations. By jointly modeling the three representations, we enhance the disentanglement of each representation and achieve robust face recognition performance. Further, the learned viewpoint representation can be utilized for pose estimation or editing of a posed facial image. Extensive quantitative and qualitative evaluations verify the effectiveness of our proposed method which disentangles identity, viewpoint, and residues of facial images. Minsu Kim 0001, Joanna Hong, Hong Joo Lee 0001, Yong Man Ro |
ICPR | 2 |
| 2020 | Face Tells Detailed Expression: Generating Comprehensive Facial Expression Sentence Through Facial Action Units
Joanna Hong, Hong Joo Lee 0001, Yelin Kim, Yong Man Ro |
MMM (2) | 1 |