EDBT 2026 Demo / reviewers in the wild / expert
June Sig Sung
dblp:47/9229
· DBLP profile ↗
19ranked-venue papers
4as first author
10since 2021 · last 2023
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 19 · 4 first-author · 10 since 2021Artificial intelligence and machine learning · 14 · 4 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Investigating Content-Aware Neural Text-to-Speech MOS Prediction Using Prosodic and Linguistic FeaturesabstractCurrent state-of-the-art methods for automatic synthetic speech evaluation are based on MOS prediction neural models. Such MOS prediction models include MOSNet and LDNet that use spectral features as input, and SSL-MOS that relies on a pretrained selfsupervised learning model that directly uses the speech signal as input. In modern high-quality neural TTS systems, prosodic appropriateness with regard to the spoken content is a decisive factor for speech naturalness. For this reason, we propose to include prosodic and linguistic features as additional inputs in MOS prediction systems, and evaluate their impact on the prediction outcome. We consider phoneme-level F0 and duration features as prosodic inputs, as well as Tacotron encoder outputs, POS tags and BERT embeddings as higher-level linguistic inputs. All MOS prediction systems are trained on SOMOS, a neural TTS-only dataset with crowdsourced naturalness MOS evaluations. Results show that the proposed additional features are beneficial in the MOS prediction task, by improving the predicted MOS scores’ correlation with the ground truths, both at utterance-level and system-level predictions. Alexandra Vioni, Georgia Maniati, Nikolaos Ellinas, June Sig Sung, Inchul Hwang, Aimilios Chalamandaris, Pirros Tsiakoulis |
ICASSP | 4 |
| 2023 | Controllable speech synthesis by learning discrete phoneme-level prosodic representations
Nikolaos Ellinas, Myrsini Christidou, Alexandra Vioni, June Sig Sung, Aimilios Chalamandaris, Pirros Tsiakoulis, Paris Mastorocostas |
Speech Commun. | 4 |
| 2022 | Karaoker: Alignment-free singing voice synthesis with speech training dataabstractExisting singing voice synthesis models (SVS) are usually trained on singing data and depend on either error-prone time-alignment and duration features or explicit music score information. In this paper, we propose Karaoker, a multispeaker Tacotron-based model conditioned on voice characteristic features that is trained exclusively on spoken data without requiring time-alignments. Karaoker synthesizes singing voice and transfers style following a multi-dimensional template extracted from a source waveform of an unseen singer/speaker. The model is jointly conditioned with a single deep convolutional encoder on continuous data including pitch, intensity, harmonicity, formants, cepstral peak prominence and octaves. We extend the text-to-speech training objective with feature reconstruction, classification and speaker identification tasks that guide the model to an accurate result. In addition to multitasking, we also employ a Wasserstein GAN training scheme as well as new losses on the acoustic model's output to further refine the quality of the model. Panos Kakoulidis, Nikolaos Ellinas, Georgios Vamvoukakis, Konstantinos Markopoulos, June Sig Sung, Gunu Jho, Pirros Tsiakoulis, Aimilios Chalamandaris |
INTERSPEECH | 5 |
| 2022 | Self supervised learning for robust voice cloningabstractVoice cloning is a difficult task which requires robust and informative features incorporated in a high quality TTS system in order to effectively copy an unseen speaker's voice.In our work, we utilize features learned in a self-supervised framework via the Bootstrap Your Own Latent (BYOL) method, which is shown to produce high quality speech representations when specific audio augmentations are applied to the vanilla algorithm.We further extend the augmentations in the training procedure to aid the resulting features to capture the speaker identity and to make them robust to noise and acoustic conditions.The learned features are used as pre-trained utterance-level embeddings and as inputs to a Non-Attentive Tacotron based architecture, aiming to achieve multispeaker speech synthesis without utilizing additional speaker features.This method enables us to train our model in an unlabeled multispeaker dataset as well as use unseen speaker embeddings to copy a speaker's voice.Subjective and objective evaluations are used to validate the proposed model, as well as the robustness to the acoustic conditions of the target utterance. Konstantinos Klapsas, Nikolaos Ellinas, Karolos Nikitaras, Georgios Vamvoukakis, Panos Kakoulidis, Konstantinos Markopoulos, Spyros Raptis, June Sig Sung, Gunu Jho, Aimilios Chalamandaris, Pirros Tsiakoulis |
INTERSPEECH | 8 |
| 2022 | SOMOS: The Samsung Open MOS Dataset for the Evaluation of Neural Text-to-Speech SynthesisabstractIn this work, we present the SOMOS dataset, the first large-scale mean opinion scores (MOS) dataset consisting of solely neural text-to-speech (TTS) samples. It can be employed to train automatic MOS prediction systems focused on the assessment of modern synthesizers, and can stimulate advancements in acoustic model evaluation. It consists of 20K synthetic utterances of the LJ Speech voice, a public domain speech dataset which is a common benchmark for building neural acoustic models and vocoders. Utterances are generated from 200 TTS systems including vanilla neural acoustic models as well as models which allow prosodic variations. An LPCNet vocoder is used for all systems, so that the samples' variation depends only on the acoustic models. The synthesized utterances provide balanced and adequate domain and length coverage. We collect MOS naturalness evaluations on 3 English Amazon Mechanical Turk locales and share practices leading to reliable crowdsourced annotations for this task. We provide baseline results of state-of-the-art MOS prediction models on the SOMOS dataset and show the limitations that such models face when assigned to evaluate TTS utterances. Georgia Maniati, Alexandra Vioni, Nikolaos Ellinas, Karolos Nikitaras, Konstantinos Klapsas, June Sig Sung, Gunu Jho, Aimilios Chalamandaris, Pirros Tsiakoulis |
INTERSPEECH | 6 |
| 2022 | Fine-grained Noise Control for Multispeaker Speech SynthesisabstractA text-to-speech (TTS) model typically factorizes speech attributes such as content, speaker and prosody into disentangled representations.Recent works aim to additionally model the acoustic conditions explicitly, in order to disentangle the primary speech factors, i.e. linguistic content, prosody and timbre from any residual factors, such as recording conditions and background noise.This paper proposes unsupervised, interpretable and fine-grained noise and prosody modeling.We incorporate adversarial training, representation bottleneck and utterance-to-frame modeling in order to learn frame-level noise representations.To the same end, we perform fine-grained prosody modeling via a Fully Hierarchical Variational AutoEncoder (FVAE) which additionally results in more expressive speech synthesis.Experimental results support our claims and ablation studies verify the importance of each proposed component.Audio samples are available in our demo page 1 . Karolos Nikitaras, Georgios Vamvoukakis, Nikolaos Ellinas, Konstantinos Klapsas, Konstantinos Markopoulos, Spyros Raptis, June Sig Sung, Gunu Jho, Aimilios Chalamandaris, Pirros Tsiakoulis |
INTERSPEECH | 7 |
| 2022 | Bunched LPCNet2: Efficient Neural Vocoders Covering Devices from Cloud to EdgeabstractText-to-Speech (TTS) services that run on edge devices have many advantages compared to cloud TTS, e.g., latency and privacy issues.However, neural vocoders with a low complexity and small model footprint inevitably generate annoying sounds.This study proposes a Bunched LPCNet2, an improved LPCNet architecture that provides highly efficient performance in high-quality for cloud servers and in a lowcomplexity for low-resource edge devices.Single logistic distribution achieves computational efficiency, and insightful tricks reduce the model footprint while maintaining speech quality.A DualRate architecture, which generates a lower sampling rate from a prosody model, is also proposed to reduce maintenance costs.The experiments demonstrate that Bunched LPCNet2 generates satisfactory speech quality with a model footprint of 1.1MB while operating faster than real-time on a RPi 3B.Our audio samples are available at https://srtts.github.io/bunchedLPCNet2. Kihyun Choo, Anton V. Porov, Konstantin Osipov, June Sig Sung |
INTERSPEECH | 6 |
| 2021 | Vibrato Learning in Multi-Singer Singing Voice SynthesisabstractDecent vibratos are a trait of good vocal training, often associated with perceived level of singing skill. In this paper we present a system for multi-singer singing voice synthesis, which is capable of producing high quality singing with convincing, controllable vibratos, and can also synthesize natural singing voices for target speakers with only speech data. This is enabled by using a unified speech-and-singing acoustic model that not only bridges the modality gap but also helps make best use of both types of data. The acoustic model exposes the full F0 contour therefore allowing explicitly modelling of F0 characteristics specific to singing voice. We observe that short-time Fourier transform of the F0 contour sparsely encode vibrato characteristics, and derive a learning objective therefrom for improved vibrato production. Control of the synthesized vibrato extent is possible by wiring in a supervised “extent” neuron and expose it to outer system. Experimental results confirm the effectiveness of proposed objective in producing good vibratos and improving overall perceived singing voice quality. Ruolan Liu, Xue Wen 0002, Chunhui Lu, Liming Song, June Sig Sung |
ASRU | 5 |
| 2021 | Prosodic Clustering for Phoneme-Level Prosody Control in End-to-End Speech SynthesisabstractThis paper presents a method for controlling the prosody at the phoneme level in an autoregressive attention-based text-to-speech system. Instead of learning latent prosodic features with a variational framework as is commonly done, we directly extract phoneme-level F0 and duration features from the speech data in the training set. Each prosodic feature is discretized using unsupervised clustering in order to produce a sequence of prosodic labels for each utterance. This sequence is used in parallel to the phoneme sequence in order to condition the decoder with the utilization of a prosodic encoder and a corresponding attention module. Experimental results show that the proposed method retains the high quality of generated speech, while allowing phoneme-level control of F0 and duration. By replacing the F0 cluster centroids with musical notes, the model can also provide control over the note and octave within the range of the speaker. Alexandra Vioni, Myrsini Christidou, Nikolaos Ellinas, Georgios Vamvoukakis, Panos Kakoulidis, June Sig Sung, Hyoungmin Park, Aimilios Chalamandaris, Pirros Tsiakoulis |
ICASSP | 7 |
| 2021 | Cross-Lingual Low Resource Speaker Adaptation Using Phonological FeaturesabstractThe idea of using phonological features instead of phonemes as input to sequence-to-sequence TTS has been recently proposed for zero-shot multilingual speech synthesis. This approach is useful for code-switching, as it facilitates the seamless uttering of foreign text embedded in a stream of native text. In our work, we train a language-agnostic multispeaker model conditioned on a set of phonologically derived features common across different languages, with the goal of achieving cross-lingual speaker adaptation. We first experiment with the effect of language phonological similarity on cross-lingual TTS of several source-target language combinations. Subsequently, we fine-tune the model with very limited data of a new speaker's voice in either a seen or an unseen language, and achieve synthetic speech of equal quality, while preserving the target speaker's identity. With as few as 32 and 8 utterances of target speaker data, we obtain high speaker similarity scores and naturalness comparable to the corresponding literature. In the extreme case of only 2 available adaptation utterances, we find that our model behaves as a few-shot learner, as the performance is similar in both the seen and unseen adaptation language scenarios. Georgia Maniati, Nikolaos Ellinas, Konstantinos Markopoulos, Georgios Vamvoukakis, June Sig Sung, Hyoungmin Park, Aimilios Chalamandaris, Pirros Tsiakoulis |
Interspeech | 5 |
| 2020 | High Quality Streaming Speech Synthesis with Low, Sentence-Length-Independent LatencyabstractThis paper presents an end-to-end text-to-speech system with low latency on a CPU, suitable for real-time applications. The system is composed of an autoregressive attention-based sequence-to-sequence acoustic model and the LPCNet vocoder for waveform generation. An acoustic model architecture that adopts modules from both the Tacotron 1 and 2 models is proposed, while stability is ensured by using a recently proposed purely location-based attention mechanism, suitable for arbitrary sentence length generation. During inference, the decoder is unrolled and acoustic feature generation is performed in a streaming manner, allowing for a nearly constant latency which is independent from the sentence length. Experimental results show that the acoustic model can produce feature sequences with minimal latency about 31 times faster than real-time on a computer CPU and 6.5 times on a mobile CPU, enabling it to meet the conditions required for real-time applications on both devices. The full end-to-end system can generate almost natural quality speech, which is verified by listening tests. Nikolaos Ellinas, Georgios Vamvoukakis, Konstantinos Markopoulos, Aimilios Chalamandaris, Georgia Maniati, Panos Kakoulidis, Spyros Raptis, June Sig Sung, Hyoungmin Park, Pirros Tsiakoulis |
INTERSPEECH | 8 |
| 2013 | Factored maximum likelihood kernelized regression for HMM-based singing voice synthesis
June Sig Sung, Doo Hwa Hong, Hyun Woo Koo, Nam Soo Kim |
INTERSPEECH | 1 |
| 2012 | Artificial stereo data generation for speech feature mappingabstractFeature mapping technique is widely used to eliminate the mismatch between the training and test conditions of speech recognition. In the feature mapping, a target (mismatched) feature vector sequence is mapped closer to the corresponding reference (matched) feature vector stream. The training of the mapping system is usually carried out based on a set of stereo data which consists of simultaneous recordings obtained in both the reference and target conditions. In this paper, we propose a novel approach to blind parameter estimation which does not require the reference feature vectors. The proposed approach is motivated by the hidden Markov model (HMM)-based speech synthesis algorithm. Chang Woo Han, Tae Gyoon Kang, Shin Jae Kang, June Sig Sung, Nam Soo Kim |
ICASSP | 4 |
| 2012 | Factored MLLR Adaptation Algorithm for HMM-based Expressive TTS
June Sig Sung, Doo Hwa Hong, Hyun Woo Koo, Nam Soo Kim |
INTERSPEECH | 1 |
| 2011 | Decision Tree-Based Clustering with Outlier Detection for HMM-Based Speech Synthesis
Kyung Hwan Oh, June Sig Sung, Doo Hwa Hong, Nam Soo Kim |
INTERSPEECH | 2 |
| 2011 | Factored MLLR Adaptation for Singing Voice GenerationabstractIn our previous study, we proposed factored MLLR (FMLLR) where each MLLR parameter is defined as a function of a control vector. We presented a method to train the FMLLR parameters based on a general framework of the expectationmaximization (EM) algorithm. In this paper, we extend the FMLLR structure from diagonal to unrestricted full matrix with a sophisticated algorithm for the training of relevant parameters. In the experiments on artificial generation of singing voice, we evaluate the performance of the FMLLR technique with two matrix structures and also compare with other approaches to parameter adaptation in HMM-based speech synthesis. Index Terms: Parameter adaptation, MLLR, HRHSMM, factored MLLR June Sig Sung, Doo Hwa Hong, Shin Jae Kang, Nam Soo Kim |
INTERSPEECH | 1 |
| 2011 | Factored MLLR AdaptationabstractOne of the most popular approaches to parameter adaptation in hidden Markov model (HMM) based systems is the maximum likelihood linear regression (MLLR) technique. In this letter, we extend MLLR to factored MLLR (FMLLR) in which the MLLR parameters depend on a continuous-valued control vector. Since it is practically impossible to estimate the MLLR parameters for each control vector separately, we propose a compact parametric form of the MLLR parameters. In the proposed approach, each MLLR parameter is represented as an inner product between a regression vector and transformed control vector. We present an algorithm to train the FMLLR parameters based on a general framework of the expectation-maximization (EM) algorithm. The proposed approach is applied to adapt the HMM parameters obtained from a database of reading-style speech to singing-style voices while treating the pitches and durations extracted from the musical notes as the control vectors. This enables to efficiently construct a singing voice synthesizer with only a small amount of singing data. Nam Soo Kim, June Sig Sung, Doo Hwa Hong |
IEEE Signal Process. Lett. | 2 |
| 2010 | Excitation modeling based on waveform interpolation for HMM-based speech synthesisabstractIt is generally known that a well-designed excitation produces high quality signals in hidden Markov model (HMM)-based speech synthesis systems. This paper proposes a novel tech- niques for generating excitation based on the waveform inter- polation (WI). For modeling WI parameters, we implemented statistical method like principal component analysis (PCA). The parameters of the proposed excitation modeling techniques can be easily combined with the conventional speech synthesis sys- tem under the HMM framework. From a number of experi- ments, the proposed method has been found to generate more naturally sounding speech. Index Terms: HMM-based speech synthesis, Waveform Inter- polation, Principal Component Analysis In this paper, we propose a novel approach to excitation modeling under the waveform interpolation (WI) framework. For parameterizing the excitation generation model, a charac- teristic waveform (CW) is extracted from each frame of LP residual signals. To derive a compact representation of each CW, we apply principal component analysis (PCA) to a collec- tion of the extracted CW's. Once PCA is done, each CW can be compactly approximated as a linear combination of a few PCA basis vectors. The statistical distribution of the linear com- bination coefficients and their dynamics can be efficiently de- scribed by means of HMM's for which the relevant parameters are estimated by following the conventional HMM training pro- cedure. Given a sentence we want to synthesize, the sequence of CW's can be generated from the trained HMM's according to the maximum likelihood (ML) criterion. The WI algorithm enables a smooth transition between adjacent CW's resulting in a more natural excitation signal. The major advantages of the proposed technique are twofold. First, instead of using a fixed set of waveforms such as the impulse train and the ran- dom noise, the proposed method finds CWs which represents the excitation waveforms from the various kinds of modeling in frequency domain. Second, the WI approach lets the excita- tion signal evolve smoothly, which may reduce the audible arti- facts of the synthesized speech. From a number of experiments on speech synthesis, it has been demonstrated that the propose technique enhances the quality of the synthesized speech. June Sig Sung, Doo Hwa Hong, Kyung Hwan Oh, Nam Soo Kim |
INTERSPEECH | 1 |
| 2007 | Speech reinforcement based on partial specific loudness
Jong Won Shin, Woohyung Lim, June Sig Sung, Nam Soo Kim |
INTERSPEECH | 3 |