VLDB 2026 Research / reviewers in the wild / expert
Milos Cernak
dblp:60/2378
· DBLP profile ↗
54ranked-venue papers
17as first author
17since 2021 · last 2025
0000-0002-5569-9491ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 48 · 13 first-author · 17 since 2021Artificial intelligence and machine learning · 35 · 11 first-author · 10 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Semi-intrusive audio evaluation: Casting non-intrusive assessment as a multi-modal text prediction taskabstractHuman perception has the unique ability to focus on specific events in a mixture of signals – a challenging task for existing non-intrusive assessment methods. In this work, we introduce semi-intrusive assessment that emulates human attention by framing audio assessment as a text-prediction task with audio-text inputs. To this end, we extend the multi-modal PENGI model through instruction fine-tuning for MOS and SNR estimation. For MOS, our approach achieves absolute Pearson correlation gains of 0.06 and 0.20 over the re-trained MOSRA model and the pre-trained PAM model, respectively. We further propose a novel SNR estimator that can focus on a specific audio source in a mixture, outperforming a random baseline and the fixed-prompt counterpart. Our findings suggest that semi-intrusive assessment can effectively capture humanlike selective listening capabilities. Samples are available at https://jozefcoldenhoff.github.io/semi-intrusive-assessment. Jozef Coldenhoff, Milos Cernak |
ICASSP | 2 |
| 2025 | OpenACE: An Open Benchmark for Evaluating Audio Coding PerformanceabstractAudio and speech coding lack unified evaluation and open-source testing. Many candidate systems were evaluated on proprietary, non-reproducible, or small data, and machine learning-based codecs are often tested on datasets with similar distributions as trained on, which is unfairly compared to digital signal processing-based codecs that usually work well with unseen data. This paper presents a full-band audio and speech coding quality benchmark with more variable content types, including traditional open test vectors. An example use case of audio coding quality assessment is presented with open-source Opus, 3GPP’s EVS, and recent ETSI’s LC3 with LC3+ used in Bluetooth LE Audio profiles. Besides, quality variations of emotional speech encoding at 16 kbps are shown. The proposed open-source benchmark contributes to audio and speech coding democratization and is available at https://github.com/JozefColdenhoff/OpenACE. Jozef Coldenhoff, Niclas Granqvist, Milos Cernak |
ICASSP | 3 |
| 2025 | Model as Loss: A Self-Consistent Training Paradigm
Saisamarth Rajesh Phaye, Milos Cernak, Andrew Harper |
INTERSPEECH | 2 |
| 2025 | DeepFilterGAN: A Full-band Real-time Speech Enhancement System with GAN-based Stochastic RegenerationabstractIn this work, we propose a full-band real-time speech enhancement system with GAN-based stochastic regeneration. Predictive models focus on estimating the mean of the target distribution, whereas generative models aim to learn the full distribution. This behavior of predictive models may lead to over-suppression, i.e. the removal of speech content. In the literature, it was shown that combining a predictive model with a generative one within the stochastic regeneration framework can reduce the distortion in the output. We use this framework to obtain a real-time speech enhancement system. With 3.58M parameters and a low latency, our system is designed for real-time streaming with a lightweight architecture. Experiments show that our system improves over the first stage in terms of NISQA-MOS metric. Finally, through an ablation study, we show the importance of noisy conditioning in our system. We participated in 2025 Urgent Challenge with our model and later made further improvements. Sanberk Serbest, Tijana Stojkovic, Milos Cernak, Andrew Harper |
INTERSPEECH | 3 |
| 2024 | Multi-Channel Mosra: Mean Opinion Score and Room Acoustics Estimation Using Simulated Data and A Teacher ModelabstractPrevious methods for predicting room acoustic parameters and speech quality metrics have focused on the single-channel case, where room acoustics and Mean Opinion Score (MOS) are predicted for a single recording device. However, quality-based device selection for rooms with multiple recording devices may benefit from a multi-channel approach where the descriptive metrics are predicted for multiple devices in parallel. Following our hypothesis that a model may benefit from multi-channel training, we develop a multi-channel model for joint MOS and room acoustics prediction (MOSRA) for five channels in parallel. The lack of multi-channel audio data with ground truth labels necessitated the creation of simulated data using an acoustic simulator with room acoustic labels extracted from the generated impulse responses and labels for MOS generated in a student-teacher setup using a wav2vec2-based MOS prediction model. Our experiments show that the multi-channel model improves the prediction of the direct-to-reverberation ratio, clarity, and speech transmission index over the single-channel model with roughly 5× less computation while suffering minimal losses in the performance of the other metrics. Jozef Coldenhoff, Andrew Harper, Paul Kendrick, Tijana Stojkovic, Milos Cernak |
ICASSP | 5 |
| 2024 | On Real-Time Multi-Stage Speech Enhancement SystemsabstractRecently, multi-stage systems have stood out among deep learning-based speech enhancement methods. However, these systems are always high in complexity, requiring millions of parameters and powerful computational resources, which limits their application for real-time processing in low-power devices. Besides, the contribution of various influencing factors to the success of multi-stage systems remains unclear, which presents challenges to reduce the size of these systems. In this paper, we extensively investigate a lightweight two-stage network with only 560k total parameters. It consists of a Mel-scale magnitude masking model in the first stage and a complex spectrum mapping model in the second stage. We first provide a consolidated view of the roles of gain power factor, post-filter, and training labels for the Mel-scale masking model. Then, we explore several training schemes for the two-stage network and provide some insights into the superiority of the two-stage network. We show that the proposed two-stage network trained by an optimal scheme achieves a performance similar to a four times larger open source model DeepFilterNet2 [1]. Lingjun Meng, Jozef Coldenhoff, Paul Kendrick, Tijana Stojkovic, Andrew Harper, Kiril Ratmanski, Milos Cernak |
ICASSP | 7 |
| 2023 | Efficient Speech Quality Assessment Using Self-Supervised Framewise EmbeddingsabstractAutomatic speech quality assessment is essential for audio researchers, developers, speech and language pathologists, and system quality engineers. The current state-of-the-art systems are based on framewise speech features (hand-engineered or learnable) combined with time dependency modeling. This paper proposes an efficient system with results comparable to the best performing model in the ConferencingSpeech 2022 challenge. Our proposed system is characterized by a smaller number of parameters (40-60x), fewer FLOPS (100x), lower memory consumption (10-15x), and lower latency (30x). Speech quality practitioners can therefore iterate much faster, deploy the system on resource-limited hardware, and, overall, the proposed system contributes to sustainable machine learning. The paper also concludes that framewise embeddings outperform utterance-level embeddings and that multi-task training with acoustic conditions modeling does not degrade speech quality prediction while providing better interpretation. Karl El Hajal, Zihan Wu 0009, Neil Scheidwasser-Clow, Gasser Elbanna, Milos Cernak |
ICASSP | 5 |
| 2023 | Personalized Task Load Prediction in Speech CommunicationabstractEstimating the quality of remote speech communication is a complex task influenced by the speaker, transmission channel, and listener. For example, the degradation of transmission quality can increase listeners’ cognitive load, which can influence the overall perceived quality of the conversation. This paper presents a framework that isolates quality-dependent changes and controls most outside influencing factors like personal preference in a simulated conversational environment. The performed statistical analysis finds significant relationships between stimulus quality and the listener’s valence and personality (agreeableness and openness) and, similarly, between the perceived task load during the listening task and the listener’s personality and frustration intolerance. The machine learning model of the task load prediction improves the correlation coefficients from 0.48 to 0.76 when listeners’ individuality is considered. The proposed evaluation framework and results pave the way for personalized audio quality assessment that includes speakers’ and listeners’ individuality beyond conventional channel modeling. Robert P. Spang, Karl El Hajal, Sebastian Möller 0001, Milos Cernak |
ICASSP | 4 |
| 2023 | Speaker Embeddings as Individuality Proxy for Voice Stress Detection
Zihan Wu 0009, Neil Scheidwasser-Clow, Karl El Hajal, Milos Cernak |
INTERSPEECH | 4 |
| 2023 | ALO-VC: Any-to-any Low-latency One-shot Voice Conversion
Damien Ronssin, Milos Cernak |
INTERSPEECH | 3 |
| 2022 | SERAB: A Multi-Lingual Benchmark for Speech Emotion RecognitionabstractRecent developments in speech emotion recognition (SER) often leverage deep neural networks (DNNs). Comparing and benchmarking different DNN models can often be tedious due to the use of different datasets and evaluation protocols. To facilitate the process, here, we present the Speech Emotion Recognition Adaptation Benchmark (SERAB), a framework for evaluating the performance and generalization capacity of different approaches for utterance-level SER. The benchmark is composed of nine datasets for SER in six languages. Since the datasets have different sizes and numbers of emotional classes, the proposed setup is particularly suitable for estimating the generalization capacity of pre-trained DNN-based feature extractors. We used the proposed framework to evaluate a selection of standard hand-crafted feature sets and state-of-the-art DNN representations. The results highlight that using only a subset of the data included in SERAB can result in biased evaluation, while compliance with the proposed protocol can circumvent this issue. Neil Scheidwasser-Clow, Mikolaj Kegler, Pierre Beckmann, Milos Cernak |
ICASSP | 4 |
| 2022 | PEAF: Learnable Power Efficient Analog Acoustic Features for Audio RecognitionabstractAt the end of Moore's law, new computing paradigms are required to prolong the battery life of wearable and IoT smart audio devices. Theoretical analysis and physical validation have shown that analog signal processing (ASP) can be more power-efficient than its digital counterpart in the realm of lowto-medium signal-to-noise ratio applications. In addition, ASP allows a direct interface with an analog microphone without a power-hungry analog-to-digital converter. Here, we present power-efficient analog acoustic features (PEAF) that are validated by fabricated CMOS chips for running audio recognition. Linear, non-linear, and learnable PEAF variants are evaluated on two speech processing tasks that are demanded in many battery-operated devices: wake word detection (WWD) and keyword spotting (KWS). Compared to digital acoustic features, higher power efficiency with competitive classification accuracy can be obtained. A novel theoretical framework based on information theory is established to analyze the information flow in each individual stage of the feature extraction pipeline. The analysis identifies the information bottleneck and helps improve the KWS accuracy by up to 7%. This work may pave the way to building more power-efficient smart audio devices with best-in-class inference performance. Boris Bergsma, Minhao Yang, Milos Cernak |
INTERSPEECH | 3 |
| 2022 | Hybrid Handcrafted and Learnable Audio Representation for Analysis of Speech Under Cognitive and Physical LoadabstractAs a neurophysiological response to threat or adverse conditions, stress can affect cognition, emotion and behaviour with potentially detrimental effects on health in the case of sustained exposure. Since the affective content of speech is inherently modulated by an individual's physical and mental state, a substantial body of research has been devoted to the study of paralinguistic correlates of stress-inducing task load. Historically, voice stress analysis (VSA) has been conducted using conventional digital signal processing (DSP) techniques. Despite the development of modern methods based on deep neural networks (DNNs), accurately detecting stress in speech remains difficult due to the wide variety of stressors and considerable variability in the individual stress perception. To that end, we introduce a set of five datasets for task load detection in speech. The voice recordings were collected as either cognitive or physical stress was induced in the cohort of volunteers, with a cumulative number of more than a hundred speakers. We used the datasets to design and evaluate a novel self-supervised audio representation that leverages the effectiveness of handcrafted features (DSP-based) and the complexity of data-driven DNN representations. Notably, the proposed approach outperformed both extensive handcrafted feature sets and novel DNN-based audio representation learning approaches. Gasser Elbanna, Alice Biryukov, Neil Scheidwasser-Clow, Lara Orlandic, Pablo Mainar, Mikolaj Kegler, Pierre Beckmann, Milos Cernak |
INTERSPEECH | 8 |
| 2022 | MOSRA: Joint Mean Opinion Score and Room Acoustics Speech Quality AssessmentabstractThe acoustic environment can degrade speech quality during communication (e.g., video call, remote presentation, outside voice recording), and its impact is often unknown.Objective metrics for speech quality have proven challenging to develop given the multi-dimensionality of factors that affect speech quality and the difficulty of collecting labeled data.Hypothesizing the impact of acoustics on speech quality, this paper presents MOSRA: a non-intrusive multi-dimensional speech quality metric that can predict room acoustics parameters (SNR, STI, T60, DRR, and C50) alongside the overall mean opinion score (MOS) for speech quality.By explicitly optimizing the model to learn these room acoustics parameters, we can extract more informative features and improve the generalization for the MOS task when the training data is limited.Furthermore, we also show that this joint training method enhances the blind estimation of room acoustics, improving the performance of current state-of-the-art models.An additional side-effect of this joint prediction is the improvement in the explainability of the predictions, which is a valuable feature for many applications. Karl El Hajal, Milos Cernak, Pablo Mainar |
INTERSPEECH | 2 |
| 2022 | Application for Real-time Personalized Speaker Extraction
Damien Ronssin, Milos Cernak |
INTERSPEECH | 2 |
| 2021 | AC-VC: Non-Parallel Low Latency Phonetic Posteriorgrams Based Voice ConversionabstractThis paper presents AC-VC (Almost Causal Voice Conversion), a phonetic posteriorgrams based voice conversion system that can perform any-to-many voice conversion while having only 57.5 ms future look-ahead. The complete system is composed of three neural networks trained separately with non-parallel data. While most of the current voice conversion systems focus primarily on quality irrespective of algorithmic latency, this work elaborates on designing a method using a minimal amount of future context thus allowing a future real-time implementation. According to a subjective listening test organized in this work, the proposed AC-VC system achieves parity with the non-causal ASR-TTS baseline of the Voice Conversion Challenge 2020 in naturalness with a MOS of 3.5. In contrast, the results indicate that missing future context impacts speaker similarity. Obtained similarity percentage of 65% is lower than the similarity of current best voice conversion systems. Damien Ronssin, Milos Cernak |
ASRU | 2 |
| 2021 | Non-Intrusive Speech Quality Assessment with Transfer Learning and Subject-Specific ScalingabstractIn communication systems, it is crucial to estimate the perceived quality of audio and speech. The industrial standards for many years have been PESQ, 3QUEST, and POLQA, which are intrusive methods. This restricts the possibilities of using these metrics in real-world conditions, where we might not have access to the clean reference signal. In this work, we develop a new non-intrusive metric based on crowd-sourced data. We build a new speech dataset by combining publicly available speech, noises, and reverberations. Then we follow the ITU P.808 recommendation to label the dataset with mean opinion scores (MOS). Finally, we train a deep neural network to estimate the MOS from the speech data in a non-intrusive way. We propose two novelties in our work. First, we explore transfer learning by pre-training a model using a larger set of POLQA scores and finetuning with the smaller (and thus cheaper) human-labeled set. Secondly, we perform a subject-specific scaling in the MOS scores to adjust for their different subjective scales. Our model yields better accuracy than PESQ, POLQA, and other non-intrusive methods when evaluated on the independent VCTK test set. We also report misleading POLQA scores for reverberant speech. Natalia Nessler, Milos Cernak, Paolo Prandoni, Pablo Mainar |
Interspeech | 2 |
| 2020 | A Bin Encoding Training of a Spiking Neural Network Based Voice Activity DetectionabstractAdvances of deep learning for Artificial Neural Networks (ANNs) have led to significant improvements in the performance of digital signal processing systems implemented on digital chips. Although recent progress in low-power chips is remarkable, neuromorphic chips that run Spiking Neural Networks (SNNs) based applications offer an even lower power consumption, as a consequence of the ensuing sparse spikebased coding scheme. In this work, we develop a SNN-based Voice Activity Detection (VAD) system that belongs to the building blocks of any audio and speech processing system. We propose to use the bin encoding, a novel method to convert log mel filterbank bins of single-time frames into spike patterns. We integrate the proposed scheme in a bilayer spiking architecture which was evaluated on the QUT-NOISE-TIMIT corpus. Our approach shows that SNNs enable an ultra low-0power implementation of a VAD classifier that consumes only 3.8 μW, while achieving state-of-the-art performance. The code is freely available on Code Ocean [1]. Giorgia Dellaferrera, Flavio Martinelli, Milos Cernak |
ICASSP | 3 |
| 2020 | Spiking Neural Networks Trained With Backpropagation for Low Power Neuromorphic Implementation of Voice Activity DetectionabstractRecent advances in Voice Activity Detection (VAD) are driven by artificial and Recurrent Neural Networks (RNNs), however, using a VAD system in battery-operated devices requires further power efficiency. This can be achieved by neuromorphic hardware, which enables Spiking Neural Networks (SNNs) to perform inference at very low energy consumption. Spiking networks are characterized by their ability to process information efficiently, in a sparse cascade of binary events in time called spikes. However, a big performance gap separates artificial from spiking networks, mostly due to a lack of powerful SNN training algorithms. To overcome this problem we exploit an SNN model that can be recast into a recurrent network and trained with known deep learning techniques. We describe a training procedure that achieves low spiking activity and apply pruning algorithms to remove up to 85% of the network connections with no performance loss. The model competes with state-of-the-art performance at a fraction of the power consumption comparing to other methods. Flavio Martinelli, Giorgia Dellaferrera, Pablo Mainar, Milos Cernak |
ICASSP | 4 |
| 2020 | Deep Speech Inpainting of Time-Frequency MasksabstractTransient loud intrusions, often occurring in noisy environments, can completely overpower speech signal and lead to an inevitable loss of information. While existing algorithms for noise suppression can yield impressive results, their efficacy remains limited for very low signal-to-noise ratios or when parts of the signal are missing. To address these limitations, here we propose an end-to-end framework for speech inpainting, the context-based retrieval of missing or severely distorted parts of time-frequency representation of speech. The framework is based on a convolutional U-Net trained via deep feature losses, obtained using speechVGG, a deep speech feature extractor pre-trained on an auxiliary word classification task. Our evaluation results demonstrate that the proposed framework can recover large portions of missing or distorted time-frequency representation of speech, up to 400 ms and 3.2 kHz in bandwidth. In particular, our approach provided a substantial increase in STOI & PESQ objective metrics of the initially corrupted speech samples. Notably, using deep feature losses to train the framework led to the best results, as compared to conventional approaches. Mikolaj Kegler, Pierre Beckmann, Milos Cernak |
INTERSPEECH | 3 |
| 2019 | Phone-Attribute Posteriors to Evaluate the Speech of Cochlear Implant UsersabstractPeople with pre- and postlingual onset of deafness, i.e, age of occurrence of hearing loss, often present speech production\nproblems even after hearing rehabilitation by cochlear implantation. In this paper, the speech of 20 prelinguals (aged between 18 to 71 years old), 20 postlinguals (aged between 33 to 78 years old) and 20 healthy control (aged between 31 to 62 years old) German native speakers are analyzed considering phone-attribute features extracted with pre-trained Deep Neural Networks. Speech signals are analyzed with reference to the manner of articulation of consonants according to 5 groups: nasals, sibilants, fricatives, voiced-stops, and voiceless-stops. According to the results, it is possible to detect alterations in the consonant production of CI users when compared with healthy speakers. A comprehensive evaluation of speech changes of CI users will help in the rehabilitation after deafening. Tomás Arias-Vergara, Juan Rafael Orozco-Arroyave, Milos Cernak, Sandra Gollwitzer, Maria Schuster, Elmar Nöth |
INTERSPEECH | 3 |
| 2019 | Evaluating Audiovisual Source Separation in the Context of Video ConferencingabstractSource separation involving mono-channel audio is a challenging problem, in particular for speech separation where source contributions overlap both in time and frequency. This task is of high interest for applications such as video conferencing. Recent progress in machine learning has shown that the combination of visual cues, coming from the video, can increase the source separation performance. Starting from a recently designed deep neural network, we assess its ability and robustness to separate the visible speakers’ speech from other interfering speeches or signals. We test it for different configuration of video recordings where the speaker’s face may not be fully visible. We also asses the performance of the network with respect to different sets of visual features from the speakers’ faces. Berkay Inan, Milos Cernak, Helmut Grabner, Helena Peic Tukuljac, Rodrigo C. G. Pena, Benjamin Ricaud |
INTERSPEECH | 2 |
| 2019 | Open-Vocabulary Keyword Spotting with Audio and Text EmbeddingsabstractKeyword Spotting (KWS) systems allow detecting a set of spoken (pre-defined) keywords. Open-vocabulary KWS systems search for the keywords in the set of word hypotheses generated by an automatic speech recognition (ASR) system which is computationally expensive and, therefore, often implemented as a cloud-based service. Besides, KWS systems could use also word classification algorithms that do not allow easily changing the set of words to be recognized, as the classes have to be defined a priori, even before training the system. In this paper, we propose the implementation of an open-vocabulary ASR-free KWS system based on speech and text encoders that allow matching the computed embeddings in order to spot whether a keyword has been uttered. This approach would allow choosing the set of keywords a posteriori while requiring low computational power. The experiments, performed on two different datasets, show that our method is competitive with other state of the art KWS systems while allowing for a flexibility of configuration and being computationally efficient. Niccolò Sacchi, Alexandre Nanchen, Martin Jaggi, Milos Cernak |
INTERSPEECH | 4 |
| 2019 | End-to-End Accented Speech Recognition
Thibault Viglino, Petr Motlícek, Milos Cernak |
INTERSPEECH | 3 |
| 2018 | Nasal Speech Sounds Detection Using Connectionist Temporal ClassificationabstractPhone attributes, known also as distinctive or phonological features, belong to important classification of the speech sounds used in automatic speech processing. Training of conventional phone attribute detectors (classifiers), either based on acoustic measurements or deep learning approaches, requires decent phone boundary segmentation. This paper proposes a solution to train a phone attribute detector without phone alignment using an end-to-end phone attribute modeling based on the connectionist temporal classification. Experiments, performed for the nasal phone attribute on the LibriSpeech database, confirm that the proposed system outperforms conventional deep neural network detector, trained even on the same training data. Further improvements are observed with more training data. Conventional complex system that consists of feature extraction, phone force-alignment and deep neural network training is replaced by a more simpler Python package based on PyTorch, released as open-source. Milos Cernak, Sibo Tong |
ICASSP | 1 |
| 2017 | On the impact of non-modal phonation on phonological featuresabstractDifferent modes of vibration of the vocal folds contribute significantly to the voice quality. The neutral mode phonation, often used in a modal voice, is one against which the other modes can be contrastively described, also called non-modal phonations. This paper investigates the impact of non-modal phonation on phonological posteriors, the probabilities of phonological features inferred from the speech signal using a deep learning approach. Five different non-modal phonations are considered: falsetto, creaky, harshness, tense and breathiness. The impact of such non-modal phonation on phonological features, the Sound Patterns of English (SPE), is investigated in both speech analysis and synthesis tasks. We found that breathy and tense phonation impact the SPE features less, creaky phonation impacts the features moderately, and harsh and falsetto phonation impact the phonological features the most. We also report invariant and the most different SPE features impacted by non-modal phonation. Milos Cernak, Elmar Nöth, Frank Rudzicz, Heidi Christensen, Juan Rafael Orozco-Arroyave, Raman Arora, Tobias Bocklet, Hamid R. Chinaei, Julius Hannink, Phani S. Nidadavolu, Juan Camilo Vásquez-Correa, Maria Yancheva, Alyssa Vann, Nikolai Vogler |
ICASSP | 1 |
| 2017 | Multi-view representation learning via gcca for multimodal analysis of Parkinson's diseaseabstractInformation from different bio-signals such as speech, handwriting, and gait have been used to monitor the state of Parkinson's disease (PD) patients, however, all the multimodal bio-signals may not always be available. We propose a method based on multi-view representation learning via generalized canonical correlation analysis (GCCA) for learning a representation of features extracted from handwriting and gait that can be used as a complement to speech-based features. Three different problems are addressed: classification of PD patients vs. healthy controls, prediction of the neurological state of PD patients according to the UPDRS score, and the prediction of a modified version of the Frenchay dysarthria assessment (m-FDA). According to the results, the proposed approach is suitable to improve the results in the addressed problems, specially in the prediction of the UPDRS, and m-FDA scores. Juan Camilo Vásquez-Correa, Juan Rafael Orozco-Arroyave, Raman Arora, Elmar Nöth, Najim Dehak, Heidi Christensen, Frank Rudzicz, Tobias Bocklet, Milos Cernak, Hamid R. Chinaei, Julius Hannink, Phani S. Nidadavolu, Maria Yancheva, Alyssa Vann, Nikolai Vogler |
ICASSP | 9 |
| 2017 | Bob Speaks Kaldi
Milos Cernak, Alain Komaty, Amir Mohammadi, André Anjos, Sébastien Marcel |
INTERSPEECH | 1 |
| 2017 | Speech vocoding for laboratory phonologyabstractUsing phonological speech vocoding, we propose a platform for exploring relations between phonology and speech processing, and in broader terms, for exploring relations between the abstract and physical structures of a speech signal. Our goal is to make a step towards bridging phonology and speech processing and to contribute to the program of Laboratory Phonology. We show three application examples for laboratory phonology: compositional phonological speech modelling, a comparison of phonological systems and an experimental phonological parametric text-to-speech (TTS) system. The featural representations of the following three phonological systems are considered in this work: (i) Government Phonology (GP), (ii) the Sound Pattern of English (SPE), and (iii) the extended SPE (eSPE). Comparing GP- and eSPE-based vocoded speech, we conclude that the latter achieves slightly better results than the former. However, GP – the most compact phonological speech representation – performs comparably to the systems with a higher number of phonological features. The parametric TTS based on phonological speech representation, and trained from an unlabelled audiobook in an unsupervised manner, achieves intelligibility of 85% of the state-of-the-art parametric speech synthesis. We envision that the presented approach paves the way for researchers in both fields to form meaningful hypotheses that are explicitly testable using the concepts developed and exemplified in this paper. On the one hand, laboratory phonologists might test the applied concepts of their theoretical models, and on the other hand, the speech processing community may utilize the concepts developed for the theoretical phonological models for improvements of the current state-of-the-art applications. Milos Cernak, Stefan Benus, Alexandros Lazaridis |
Comput. Speech Lang. | 1 |
| 2017 | Characterisation of voice quality of Parkinson's disease using differential phonological posterior features
Milos Cernak, Juan Rafael Orozco-Arroyave, Frank Rudzicz, Heidi Christensen, Juan Camilo Vásquez-Correa, Elmar Nöth |
Comput. Speech Lang. | 1 |
| 2017 | Perceptual Information Loss due to Impaired Speech ProductionabstractPhonological classes define articulatory-free and articulatory-bound phone attributes. Deep neural network is used to estimate the probability of phonological classes from the speech signal. In theory, a unique combination of phone attributes form a phoneme identity. Probabilistic inference of phonological classes thus enables estimation of their compositional phoneme probabilities. A novel information theoretic framework is devised to quantify the information conveyed by each phone attribute, and assess the speech production quality for perception of phonemes. As a use case, we hypothesize that disruption in speech production leads to information loss in phone attributes, and thus confusion in phoneme identification. We quantify the amount of information loss due to dysarthric articulation recorded in the TORGO database. A novel information measure is formulated to evaluate the deviation from an ideal phone attribute production leading us to distinguish healthy production from pathological speech. Afsaneh Asaei, Milos Cernak, Hervé Bourlard |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2016 | Phonetic and Phonological Posterior Search Space Hashing Exploiting Class-Specific Sparsity StructuresabstractThis paper shows that exemplar-based speech processing using class-conditional posterior probabilities admits a highly effective search strategy relying on posteriors' intrinsic sparsity structures. The posterior probabilities are estimated for phonetic and phonological classes using deep neural network (DNN) computational framework. Exploiting the class-specific sparsity leads to a simple quantized posterior hashing procedure to reduce the search space of posterior exemplars. To that end, small number of quantized posteriors are regarded as representatives of the posterior space and used as hash keys to index subsets of neighboring exemplars. The $k$ nearest neighbor ($k$NN) method is applied for posterior based classification problems. The phonetic posterior probabilities are used as exemplars for phonetic classification whereas the phonological posteriors are used as exemplars for automatic prosodic event detection. Experimental results demonstrate that posterior hashing improves the efficiency of $k$NN classification drastically. This work encourages the use of posteriors as discriminative exemplars appropriate for large scale speech classification tasks. Afsaneh Asaei, Gil Luyet, Milos Cernak, Hervé Bourlard |
INTERSPEECH | 3 |
| 2016 | Sound Pattern Matching for Automatic Prosodic Event DetectionabstractLIDIAP Milos Cernak, Afsaneh Asaei, Pierre-Edouard Honnet, Philip N. Garner, Hervé Bourlard |
INTERSPEECH | 1 |
| 2016 | PhonVoc: A Phonetic and Phonological Vocoding ToolkitabstractWe present the PhonVoc toolkit, a cascaded deep neural network (DNN) composed of speech analyser and synthesizer that use a shared phonetic and/or phonological speech representation.The free toolkit is distributed as open-source software under a BSD 3-Clause License, available at https://github.com/idiap/phonvoc with the pre-trained US English analysis and synthesis DNNs, and thus it is ready for immediate use.In a broader context, the toolkit implements training and testing of the analysis by synthesis heuristic model.It is thus designed for the wider speech community working in acoustic phonetics, laboratory phonology, and parametric speech coding.The toolkit interprets the phonetic posterior probabilities as a sequential scheme, whereas the phonological posterior-class probabilities are considered as a parallel via K different phonological classes.A case study is presented on a LibriSpeech database and a LibriVox US English native female speaker.The phonetic and phonological vocoding yield comparable performance, improving speech quality by merging the phonetic and phonological speech representation. Milos Cernak, Philip N. Garner |
INTERSPEECH | 1 |
| 2016 | Probabilistic Amplitude Demodulation Features in Speech Synthesis for Improving ProsodyabstractAbstract Amplitude demodulation (AM) is a signal decomposition technique by which a signal can be decomposed to a product of two signals, i.e, a quickly varying carrier and a slowly varying modulator. In this work, the probabilistic amplitude demodulation (PAD) features are used to improve prosody in speech synthesis. The PAD is applied iteratively for generating syllable and stress amplitude modulations in a cascade manner. The PAD features are used as a secondary input scheme along with the standard text-based input features in statistical parametric speech syn- thesis. Specifically, deep neural network (DNN)-based speech synthesis is used to evaluate the importance of these features. Objective evaluation has shown that the proposed system using the PAD features has improved mainly prosody modelling; it outperforms the baseline system by approximately 5% in terms of relative reduction in root mean square error (RMSE) of the fundamental frequency (F0). The significance of this improvement is validated by subjective evaluation of the overall speech quality, achieving 38.6% over 19.5% preference score in respect to the baseline system, in an ABX test. Alexandros Lazaridis, Milos Cernak, Philip N. Garner |
INTERSPEECH | 2 |
| 2016 | HMM-Based Non-Native Accent Assessment Using Posterior FeaturesabstractAutomatic non-native accent assessment has potential benefits in language learning and speech technologies. The three fundamental challenges in automatic accent assessment are to characterize, model and assess individual variation in speech of the non-native speaker. In our recent work, accentedness score was automatically obtained by comparing two phone probability sequences obtained through instances of non-native and native speech. Although automatic accentedness ratings of the approach correlated well with human accent ratings, the approach is critically constrained because of the requirement of native speech instance. In this paper, we build on the previous work and obtain the native latent symbol probability sequence through the word hypothesis modeled as a hidden Markov model (HMM). The latent symbols are either context-independent phonemes or clustered context-dependent phonemes. The advantage of the proposed approach is that it requires just reference text transcription instead of native speech recordings. Using the HMMs trained on an auxiliary native speech corpus, the proposed approach achieves a correlation of 0.68 with human accent ratings on the ISLE corpus. This is further interesting considering that the approach does not use any non-native data and human accent ratings at any stage of the system development. Ramya Rasipuram, Milos Cernak, Mathew Magimai-Doss |
INTERSPEECH | 2 |
| 2016 | On structured sparsity of phonological posteriors for linguistic parsing
Milos Cernak, Afsaneh Asaei, Hervé Bourlard |
Speech Commun. | 1 |
| 2016 | Composition of Deep and Spiking Neural Networks for Very Low Bit Rate Speech CodingabstractMost current very low bit rate (VLBR) speech coding systems use hidden Markov model (HMM) based speech recognition and synthesis techniques. This allows transmission of information (such as phonemes) segment by segment; this decreases the bit rate. However, an encoder based on a phoneme speech recognition may create bursts of segmental errors; these would be further propagated to any suprasegmental (such as syllable) information coding. Together with the errors of voicing detection in pitch parametrization, HMM-based speech coding leads to speech discontinuities and unnatural speech sound artifacts. In this paper, we propose a novel VLBR speech coding framework based on neural networks (NNs) for end-to-end speech analysis and synthesis without HMMs. The speech coding framework relies on a phonological (subphonetic) representation of speech. It is designed as a composition of deep and spiking NNs: a bank of phonological analyzers at the transmitter, and a phonological synthesizer at the receiver. These are both realized as deep NNs, along with a spiking NN as an incremental and robust encoder of syllable boundaries for coding of continuous fundamental frequency (F0). A combination of phonological features defines much more sound patterns than phonetic features defined by HMM-based speech coders; this finer analysis/synthesis code contributes to smoother encoded speech. Listeners significantly prefer the NN-based approach due to fewer discontinuities and speech artifacts of the encoded speech. A single forward pass is required during the speech encoding and decoding. The proposed VLBR speech coding operates at a bit rate of approximately 360 bits/s. Milos Cernak, Alexandros Lazaridis, Afsaneh Asaei, Philip N. Garner |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2015 | Phonological vocoding using artificial neural networksabstractWe investigate a vocoder based on artificial neural networks using a phonological speech representation. Speech decomposition is based on the phonological encoders, realised as neural network classifiers, that are trained for a particular language. The speech reconstruction process involves using a Deep Neural Network (DNN) to map phonological features posteriors to speech parameters - line spectra and glottal signal parameters - followed by LPC resynthesis. This DNN is trained on a target voice without transcriptions, in a semi-supervised manner. Both encoder and decoder are based on neural networks and thus the vocoding is achieved using a simple fast forward pass. An experiment with French vocoding and a target male voice trained on 21 hour long audio book is presented. An application of the phonological vocoder to low bit rate speech coding is shown, where transmitted phonological posteriors are pruned and quantized. The vocoder with scalar quantization operates at 1 kbps, with potential for lower bit-rate. Milos Cernak, Blaise Potard, Philip N. Garner |
ICASSP | 1 |
| 2015 | On compressibility of neural network phonological features for low bit rate speech codingabstractPhonological features extracted by neural network have shown interesting potential for low bit rate speech vocoding. The span of phonological features is wider than the span of phonetic features, and thus fewer frames need to be transmitted. Moreover, the binary nature of phonological features enables a higher compression ratio at minor quality cost. In this paper, we study the compressibility and structured sparsity of the phonological features. We propose a compressive sampling framework for speech coding and sparse reconstruction for decoding prior to synthesis. Compressive sampling is found to be a principled way for compression in contrast to the conventional pruning approach; it leads to $50$\\% reduction in the bit-rate for better or equal quality of the decoded speech. Furthermore, exploiting the structured sparsity and binary characteristic of these features have shown to enable very low bit-rate coding at 700 bps with negligible quality loss; this coding scheme imposes no latency. If we consider a latency of $256$~ms for supra-segmental structures, the rate of $250-350$~bps is achieved. Afsaneh Asaei, Milos Cernak, Hervé Bourlard |
INTERSPEECH | 2 |
| 2015 | An empirical model of emphatic word detectionabstractLIDIAP Milos Cernak, Pierre-Edouard Honnet |
INTERSPEECH | 1 |
| 2015 | Neuromorphic based oscillatory device for incremental syllable boundary detectionabstractLIDIAP Alexandre Hyafil, Milos Cernak |
INTERSPEECH | 2 |
| 2015 | Automatic accentedness evaluation of non-native speech using phonetic and sub-phonetic posterior probabilitiesabstractAutomatic evaluation of non-native speech accentedness has potential implications for not only language learning and accent identification systems but also for speaker and speech recognition systems. From the perspective of speech production, the two primary factors influencing the accentedness are the phonetic and prosodic structure. In this paper, we propose an approach for automatic accentedness evaluation based on comparison of instances of native and non-native speakers at the acoustic-phonetic level. Specifically, the proposed approach measures accentedness by comparing phone class conditional probability sequences corresponding to the instances of native and non-native speakers, respectively. We evaluate the proposed approach on the EMIME bilingual and EMIME Mandarin bilingual corpora, which contains English speech from native English speakers and various non-native English speakers, namely Finnish, German and Mandarin. We also investigate the influence of the granularity of the phonetic unit representation on the performance of the proposed accentedness measure. Our results indicate that the accentedness ratings by the proposed approach correlate consistently with the human ratings of accentedness. In addition, our studies show that the granularity of the phonetic unit representation that yields the best correlation with the human accentedness ratings varies with respect to the native language of the non-native speakers. Ramya Rasipuram, Milos Cernak, Alexandre Nanchen, Mathew Magimai-Doss |
INTERSPEECH | 2 |
| 2015 | Incremental Syllable-Context Phonetic VocodingabstractCurrent very low bit rate speech coders are, due to complexity limitations, designed to work off-line. This paper investigates incremental speech coding that operates real-time and incrementally (i.e., encoded speech depends only on already-uttered speech without the need of future speech information). Since human speech communication is asynchronous (i.e., different information flows being simultaneously processed), we hypothesized that such an incremental speech coder should also operate asynchronously. To accomplish this task, we describe speech coding that reflects the human cortical temporal sampling that packages information into units of different temporal granularity, such as phonemes and syllables, in parallel. More specifically, a phonetic vocoder-cascaded speech recognition and synthesis systems-extended with syllable-based information transmission mechanisms is investigated. There are two main aspects evaluated in this work, the synchronous and asynchronous coding. Synchronous coding refers to the case when the phonetic vocoder and speech generation process depend on the syllable boundaries during encoding and decoding respectively. On the other hand, asynchronous coding refers to the case when the phonetic encoding and speech generation processes are done independently of the syllable boundaries. Our experiments confirmed that the asynchronous incremental speech coding performs better, in terms of intelligibility and overall speech quality, mainly due to better alignment of the segmental and prosodic information. The proposed vocoding operates at an uncompressed bit rate of 213 bits/sec and achieves an average communication delay of 243 ms. Milos Cernak, Philip N. Garner, Alexandros Lazaridis, Petr Motlícek, Xingyu Na |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2014 | Stress and accent transmission in HMM-based syllable-context very low bit rate speech codingabstractLIDIAP Milos Cernak, Alexandros Lazaridis, Philip N. Garner, Petr Motlícek |
INTERSPEECH | 1 |
| 2014 | Development of bilingual ASR system for MediaParl corpusabstractThe development of an Automatic Speech Recognition (ASR) system for the bilingual MediaParl corpus is challenging for several reasons: (1) reverberant recordings, (2) accented speech, and (3) no prior information about the language. In that context, we employ frequency domain linear prediction-based (FDLP) features to reduce the effect of reverberation, exploit bilingual deep neural networks applied in Tandem and hybrid acoustic modeling approaches to significantly improve ASR for accented speech and develop a fully bilingual ASR system using entropy-based decoding-graph selection. Our experiments indicate that the proposed bilingual ASR system performs similar to a language-specific ASR system if approximately five seconds of speech are available. Petr Motlícek, David Imseng, Milos Cernak, Namhoon Kim |
INTERSPEECH | 3 |
| 2013 | Automatic Staging of Audio with EmotionsabstractCurrent day text-to-speech technologies are mature enough to be acceptable in quality for the users. There is still a large gap between a synthesised speech and a real human speech due to lack of expressions and emotions. Geneemo is a technology for automatic addition of emotions and expressions to any audio. The process of staging the text is to dramatize it. The text is enriched and transformed into a performance. Similarly, "staging the audio" refers to extending text dramatisation to audio by enriching emotionally neutral audio content into a natural human speech with real expressions. The audio can be generated by any text-to-speech technology. The aim of the project is to make human computer interactions as natural as possible with expressive speech. This also opens up a portfolio of applications replacing real human voices. Lakshmi Babu Saheer, Milos Cernak |
ACII | 2 |
| 2013 | On the (UN)importance of the contextual factors in HMM-based speech synthesis and codingabstractThis paper presents an evaluation of the contextual factors of HMMbased speech synthesis and coding systems.Two experimental setups are proposed that are based on successive context addition from phonetic to full-context.The aim was to investigate the impact of the individual contextual factors on the speech quality.In that sense important and unimportant (i.e., not having significant impact on speech quality, also called weak) contextual factors were identified.The results imply that in speech coding the improvement in quality can be achieved just with reconstruction of syllable contexts.The sentence and utterance contexts are unimportant on the decoder side, and it is not necessary to deal with them.Although in speech coding the wider context was not necessary, in speech synthesis current syllable and utterance contexts are more important over others (previous and next word/phrase contexts). Milos Cernak, Petr Motlícek, Philip N. Garner |
ICASSP | 1 |
| 2013 | Syllable-based pitch encoding for low bit rate speech coding with recognition/synthesis architectureabstractCurrent HMM-based low bit rate speech coding systems work with phonetic vocoders.Pitch contour coding (on frame or phoneme level) is usually fairly orthogonal to other speech coding parameters.We make an assumption in our work that the speech signal contains supra-segmental cues.Hence, we present encoding of the pitch on the syllable level, used in the framework of a recognition/synthesis speech coder with phonetic vocoder.The results imply that high accuracy pitch contour reconstruction with negligible speech quality degradation is possible.The proposed pitch encoding technique operates on 30-35 bits per second. Milos Cernak, Xingyu Na, Philip N. Garner |
INTERSPEECH | 1 |
| 2013 | A Simple Continuous Pitch Estimation AlgorithmabstractRecent work in text to speech synthesis has pointed to the benefit of using a continuous pitch estimate; that is, one that records pitch even when voicing is not present. Such an approach typically requires interpolation. The purpose of this letter is to show that a continuous pitch estimation is available from a combination of otherwise well known techniques. Further, in the case of an autocorrelation based estimate, the continuous requirement negates the need for other heuristics to correct for common errors. An algorithm is suggested, illustrated, and demonstrated using a parametric vocoder. Philip N. Garner, Milos Cernak, Petr Motlícek |
IEEE Signal Process. Lett. | 2 |
| 2012 | Robust triphone mapping for acoustic modelingabstractIn this paper we revisit the recently proposed triphone mapping as an alternative to decision tree state clustering. We generalize triphone mapping to Kullback-Leibler based hidden Markov models for acoustic modeling and propose a modified training procedure for the Gaussian mixture model based acoustic modeling. We compare the triphone mapping to decision tree state clustering on the Wall Street Journal task as well as in the context of an under-resourced language by using Greek data from the SpeechDat(II) corpus. Experiments reveal that triphone mapping has the best overall performance and is robust against varying the acoustic modeling technique as well as variable amounts of training data. Milos Cernak, David Imseng, Hervé Bourlard |
INTERSPEECH | 1 |
| 2011 | Effective Triphone Mapping for Acoustic Modeling in Speech RecognitionabstractThis paper presents effective triphone mapping for acoustic models training in automatic speech recognition, which allows the synthesis of unseen triphones. The description of this data-driven model clustering, including experiments performed using 350 hours of a Slovak audio database of mixed read and spontaneous speech, are presented. The proposed technique is compared with treebased state tying, and it is shown that for bigger acoustic models, at a size of 4000 states and more, a triphone mapped HMM system achieves better performance than a tree-based state tying system. The main gain in performance is due to latent application of triphone mapping on monophones with multiple Gaussian pdfs, so the cloned triphones are initialized better than with single Gaussians monophones. Absolute decrease of word error rate was 0.46% (5.73% relatively) for models with 7500 states, and decreased to 0.4% (5.17% relatively) gain at 11500 states. Sakhia Darjaa, Milos Cernak, Marián Trnka, Milan Rusko, Róbert Sabo |
INTERSPEECH | 2 |
| 2006 | Unit Selection Speech Synthesis in NoiseabstractThe paper presents an approach to unit selection speech synthesis in noise. The approach is based on a modification of the speech synthesis method originally published in [1], where the distance of a candidate unit from its cluster center is used as the unit selection cost. We found out that using an additional measure evaluating intelligibility for the unit cost may improve the overall understandability of speech in noise. The measure we have chosen for prediction of speech intelligibility in noise is Speech Intelligibility Index (SII). While the calculation of the SII value for each unit in the speech corpus was made off-line, a pink noise was used as a representative noise for the calculation. Listening tests imply that such a simple modification of the unit cost in unit selection synthesis can improve understandability of speech delivered under poor channel conditions. Milos Cernak |
ICASSP (1) | 1 |
| 2005 | TTSBOX: a MATLAB toolbox for teaching text-to-speech synthesisabstractThis paper presents a new toolbox for teaching TTS synthesis. TTSBOX performs the synthesis of Genglish (for "generic English"), an imaginary language obtained by replacing English words by generic words. Genglish therefore has a rather limited lexicon, but its pronunciation maintains most of the problems encountered in natural languages. TTSBOX uses simple data-driven techniques (bigrams, CARTs, NUUs) while trying to keep the code minimal, so as to keep it readable for students with reasonable MATLAB practice. TTSBOX was designed with the hope that it can help to increase the personal involvement of undergraduate and graduate students in their TTS courses. Thierry Dutoit, Milos Cernak |
ICASSP (5) | 2 |