Inchul Hwang

dblp:151/4034 · DBLP profile ↗
← Back
12ranked-venue papers
2as first author
5since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 Counterfactual Activation Editing for Post-hoc Prosody and Mispronunciation Correction in TTS Models
Kyowoon Lee, Artyom Stitsyuk, Gunu Jho, Inchul Hwang, Jaesik Choi
INTERSPEECH4
2024 Investigating Disentanglement in a Phoneme-Level Speech Codec for Prosody Modeling
abstract
Most of the prevalent approaches in speech prosody modeling rely on learning global style representations or a continuous latent space which encode and transfer the attributes of reference speech. However, recent work on neural codecs which are based on Residual Vector Quantization (RVQ) already shows great potential offering distinct advantages. We investigate the prosody modeling capabilities of the discrete space of such an RVQ-VAE model, modifying it to operate on the phoneme-level. We condition both the encoder and decoder of the model on linguistic representations and apply a global speaker embedding in order to factor out both phonetic and speaker information. We conduct an extensive set of investigations based on subjective experiments and objective measures to show that the phoneme-level discrete latent representations obtained this way achieve a high degree of disentanglement, capturing fine-grained prosodic information that is robust and transferable. The latent space turns out to have interpretable structure with its principal components corresponding to pitch and energy.
Sotirios Karapiperis, Nikolaos Ellinas, Alexandra Vioni, Junkwang Oh, Gunu Jho, Inchul Hwang, Spyros Raptis
SLT6
2023 Investigating Content-Aware Neural Text-to-Speech MOS Prediction Using Prosodic and Linguistic Features
abstract
Current state-of-the-art methods for automatic synthetic speech evaluation are based on MOS prediction neural models. Such MOS prediction models include MOSNet and LDNet that use spectral features as input, and SSL-MOS that relies on a pretrained selfsupervised learning model that directly uses the speech signal as input. In modern high-quality neural TTS systems, prosodic appropriateness with regard to the spoken content is a decisive factor for speech naturalness. For this reason, we propose to include prosodic and linguistic features as additional inputs in MOS prediction systems, and evaluate their impact on the prediction outcome. We consider phoneme-level F0 and duration features as prosodic inputs, as well as Tacotron encoder outputs, POS tags and BERT embeddings as higher-level linguistic inputs. All MOS prediction systems are trained on SOMOS, a neural TTS-only dataset with crowdsourced naturalness MOS evaluations. Results show that the proposed additional features are beneficial in the MOS prediction task, by improving the predicted MOS scores’ correlation with the ground truths, both at utterance-level and system-level predictions.
Alexandra Vioni, Georgia Maniati, Nikolaos Ellinas, June Sig Sung, Inchul Hwang, Aimilios Chalamandaris, Pirros Tsiakoulis
ICASSP5
2023 Generating Multilingual Gender-Ambiguous Text-to-Speech Voices
Konstantinos Markopoulos, Georgia Maniati, Georgios Vamvoukakis, Nikolaos Ellinas, Georgios Vardaxoglou, Panos Kakoulidis, Junkwang Oh, Gunu Jho, Inchul Hwang, Aimilios Chalamandaris, Pirros Tsiakoulis, Spyros Raptis
INTERSPEECH9
2021 Task Aware Multi-Task Learning for Speech to Text Tasks
abstract
In general, the direct Speech-to-text translation (ST) is jointly trained with Automatic Speech Recognition (ASR), and Machine Translation (MT) tasks. However, the issues with the current joint learning strategies inhibit the knowledge transfer across these tasks. We propose a task modulation network which allows the model to learn task specific features, while learning the shared features simultaneously. This proposed approach removes the need for separate finetuning step resulting in a single model which performs all these tasks. This single model achieves a performance of 28.64 BLEU score on ST MuST-C English-German, WER of 11.61% on ASR TEDLium v3, 23.35 BLEU score on MT WMT’15 English-German task. This sets a new state-of-the-art performance (SOTA) on the ST task while outperforming the existing end-to-end ASR systems.
Sathish Reddy Indurthi, Mohd Abbas Zaidi, Nikhil Kumar Lakumarapu, Beomseok Lee, Hyojung Han 0003, Seokchan Ahn, Sangha Kim 0002, Chanwoo Kim 0001, Inchul Hwang
ICASSP9
2019 Paraphrase Generation Based on VAE and Pointer-Generator Networks
abstract
Paraphrase generation is a challenging task that involves expressing the meaning of a sentence using synonyms or different phrases, either to achieve variations or a certain stylistic response. Most previous sequence-to-sequence (Seq2Seq) models focus on either generating variations or preserving the content. We mainly address the issue of preserving the content in a sentence while generating diverse paraphrases. In this paper, we propose a novel approach for paraphrase generation using variational autoencoder (VAE) and Pointer Generator Network (PGN). The proposed model uses a copy mechanism to control the content transfer, a VAE to introduce variations and a training technique to restrict the gradient flow for efficient learning. Our evaluations on QUORA and MS COCO datasets show that our model outperforms the state-of-the-art approaches and the generated paraphrases are highly diverse as well as consistent with their original meaning.
Lohith Ravuru, Hyungtak Choi, Siddarth K. M., Hojung Lee, Inchul Hwang
ASRU5
2019 Deep Reinforcement Learning for Chatbots Using Clustered Actions and Human-Likeness Rewards
abstract
Training chatbots using the reinforcement learning paradigm is challenging due to high-dimensional states, infinite action spaces and the difficulty in specifying the reward function. We address such problems using clustered actions instead of infinite actions, and a simple but promising reward function based on human-likeness scores derived from human-human dialogue data. We train Deep Reinforcement Learning (DRL) agents using chitchat data in raw text—without any manual annotations. Experimental results using different splits of training data report the following. First, that our agents learn reasonable policies in the environments they get familiarised with, but their performance drops substantially when they are exposed to a test set of unseen dialogues. Second, that the choice of sentence embedding size between 100 and 300 dimensions is not significantly different on test data. Third, that our proposed human-likeness rewards are reasonable for training chatbots as long as they use lengthy dialogue histories of ≥10 sentences.
Heriberto Cuayáhuitl, Seonghan Ryu, Sungja Choi, Inchul Hwang, Jihie Kim
IJCNN5
2019 VAE-PGN based Abstractive Model in Multi-stage Architecture for Text Summarization
abstract
This paper describes our submission to the TL;DR challenge.Neural abstractive summarization models have been successful in generating fluent and consistent summaries with advancements like the copy (Pointer-generator) and coverage mechanisms.However, these models suffer from their extractive nature as they learn to copy words from the source text.In this paper, we propose a novel abstractive model based on Variational Autoencoder (VAE) to address this issue.We also propose a Unified Summarization Framework for the generation of summaries.Our model eliminates non-critical information at a sentencelevel with an extractive summarization module and generates the summary word by word using an abstractive summarization module.To implement our framework, we combine submodules with state-of-the-art techniques including Pointer-Generator Network (PGN) and BERT while also using our new VAE-PGN abstractive model.We evaluate our model on the benchmark Reddit corpus as part of the TL;DR challenge and show that our model outperforms the baseline in ROUGE score while generating diverse summaries.
Hyungtak Choi, Lohith Ravuru, Tomasz Dryjanski, Sunghan Rye, Hojung Lee, Inchul Hwang
INLG7
2019 Ensemble-based deep reinforcement learning for chatbots
Heriberto Cuayáhuitl, Seonghan Ryu, Yongjin Cho, Sungja Choi, Sathish Reddy Indurthi, Seunghak Yu, Hyungtak Choi, Inchul Hwang, Jihie Kim
Neurocomputing9
2018 Self-Learning Architecture for Natural Language Generation
abstract
In this paper, we propose a self-learning architecture for generating natural language templates for conversational assistants.Generating templates to cover all the combinations of slots in an intent is time consuming and labor-intensive.We examine three different models based on our proposed architecture -Rule-based model, Sequence-to-Sequence (Seq2Seq) model and Semantically Conditioned LSTM (SC-LSTM) model for the IoT domain -to reduce the human labor required for template generation.We demonstrate the feasibility of template generation for the IoT domain using our self-learning architecture.In both automatic and human evaluation, the self-learning architecture outperforms previous works trained with a fully human-labeled dataset.This is promising for commercial conversational assistant solutions.
Hyungtak Choi, Siddarth K. M., Haehun Yang, Heesik Jeon, Inchul Hwang, Jihie Kim
INLG5
2015 Design and implementation of cloud offloading framework among devices for web applications
abstract
Nowadays, cloud computing has become a common computing infrastructure. As the computing paradigm has been shifted to cloud computing, devices can utilize computing resource any-where/any-time/any-device. Many research papers in mobile cloud called ‘cloud offloading’ which migrates a device's workload to a server or to the other device have been proposed. However, previous cloud offloading methods are mainly focusing on the cloud offloading between a device and a server. Furthermore, these proposed methods are difficult to be commercialized because the proposed methods were very complex - difficulty of partitioning application tasks and maintaining execution status sync between a device and a server in the cloud. In this paper, I proposed the framework for cloud offloading based on the web application standard - HTML5 specification - for web applications among devices. In HTML5 specification, there is the method for the parallel execution of the task named ‘Web Worker’ and the method for the communication among browsers name ‘WebRTC’. Utilizing the property of this specification, I proposed a seamless method to do the cloud offloading for parallelized tasks of the web applications among devices. Based on proposed method, a device can seamlessly migrate a part of web application workload with the Web Worker to other devices with a little modification of web applications. As a result, I can successfully build the environment where a device which has a HTML5 browser such as a mobile phone and a smart TV can share the workload among them in various situations - out-of-battery, good network connection.
Inchul Hwang
CCNC1
2015 Adaptive Computational Workload Offloading Method for Web Applications
Inchul Hwang
ICCSA (1)1