Yannis Pantazis

dblp:10/7000 · DBLP profile ↗
← Back
33ranked-venue papers
14as first author
11since 2021 · last 2025
0000-0002-2009-7562ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 20 · 9 first-author · 4 since 2021Artificial intelligence and machine learning · 17 · 5 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 3 first-author · 1 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Dynamic Speech Generation to Enhance Intelligibility in Noisy Environments
abstract
Synthetic speech is occasionally required to be enhanced by a fixed amount before being presented to the listener. However, this approach neglects the diverse types and levels of noise, potentially resulting in either unintelligible or unpleasant speech. This paper proposes a dynamic speech enhancement approach, which generates speech while accounting for the presence of noise. Our methodology extends the WaveRNN vocoder by conditioning not only the speech mel-spectrograms but also the spectrogram from the background noise. In training, the target speech is tilted according to the listeners’ preference and we further introduce an objective metric that aims at maximizing the spectro-temporal regions where target energy exceeds the masker energy. The generated speech was tested in speech-shaped noise at various noise levels. Our evaluation results showed that (a) in more adverse conditions, both a fixed-amount post-enhanced baseline system and the suggested dynamically-enhanced system performed equally well in terms of intelligibility and preference and (b) in the least noisy conditions, where the intelligibility scores of all models were nearly identical, listeners preferred the dynamically-enhanced speech. Our findings demonstrate and reinforce the benefits of using dynamic speech enhancement techniques in noisy environments.
Olympia Simantiraki, Maria E. Markaki, Yannis Pantazis
ICASSP3
2025 Classifier-driven generative adversarial networks for enhanced antimicrobial peptide design
abstract
The development of antimicrobial peptides (AMPs) presents a promising approach to addressing antibiotic-resistant pathogens. Computational methods, such as Feedback Generative Adversarial Networks (FBGANs), have demonstrated strong performance in optimizing AMP design. FBGAN operates as a classifier-guided Generative Adversarial Network (GAN), refining training data by replacing them with the classifier's most accurate predictions based on a predefined threshold. However, this method may introduce bias and constrain the diversity and quality of the generated peptides. To address these limitations, we propose a novel classifier-driven GAN (cdGAN) framework that seamlessly integrates classifier predictions into the generative model's loss function. This enables an adaptive, end-to-end learning process that enhances AMP generation without requiring explicit data modifications. By embedding classifier guidance within the loss computation, cdGAN dynamically optimizes both peptide diversity and functionality. Comparative studies indicate that cdGAN outperforms conventional guided-GAN architectures, such as Conditional GANs and Auxiliary Classifier GANs, while achieving performance comparable to or exceeding established AMP design methods. Additionally, cdGAN's flexible architecture allows for the simultaneous optimization of multiple peptide attributes. To demonstrate this capability, we introduce a multi-task classifier based on the Evolutionary Scale Modeling 2 (ESM2) model, enabling cdGAN to assess both antimicrobial activity and peptide structural properties in parallel. This enhancement improves the likelihood of generating viable therapeutic candidates with enhanced antimicrobial effectiveness and reduced toxicity.
Michaela Areti Zervou, Effrosyni Doutsi, Yannis Pantazis, Panagiotis Tsakalides
Briefings Bioinform.3
2024 Multitask Classification of Antimicrobial Peptides for Simultaneous Assessment of Antimicrobial Property and Structural Fold
abstract
Antimicrobial peptides (AMPs) play a significant role in guiding drug design, advancing targeted therapies, and cancer treatment research. The function of peptides is highly associated with their three-dimensional structure. AMPs particularly favor alpha-helical structures, or alpha-folds, due to their ability to disrupt the protective layers that surround cells effectively and their structural stability. Existing classifiers mainly identify AMPs but overlook their structural fold which can provide valuable insights into their function. To address this limitation, we introduce an innovative multitask classifier that recognizes AMPs and predicts their alphahelical folds simultaneously. Our approach employs k-mers and Transformer networks for efficient, accurate multitask classification. Results on the datasets indicate comparable performance compared to single-task methods in half the time and complexity.
Michaela Areti Zervou, Effrosyni Doutsi, Yannis Pantazis, Panagiotis Tsakalides
ICASSP3
2023 Function-space regularized Rényi divergences
Jeremiah Birrell, Yannis Pantazis, Paul Dupuis, Luc Rey-Bellet, Markos A. Katsoulakis
ICLR2
2023 The effect of masking noise on listeners' spectral tilt preferences
Olympia Simantiraki, Yannis Pantazis, Martin Cooke
INTERSPEECH2
2023 Learning biologically-interpretable latent representations for gene expression data
abstract
Abstract Molecular gene-expression datasets consist of samples with tens of thousands of measured quantities (i.e., high dimensional data). However, lower-dimensional representations that retain the useful biological information do exist. We present a novel algorithm for such dimensionality reduction called Pathway Activity Score Learning (PASL). The major novelty of PASL is that the constructed features directly correspond to known molecular pathways (genesets in general) and can be interpreted as pathway activity scores . Hence, unlike PCA and similar methods, PASL’s latent space has a fairly straightforward biological interpretation. PASL is shown to outperform in predictive performance the state-of-the-art method (PLIER) on two collections of breast cancer and leukemia gene expression datasets. PASL is also trained on a large corpus of 50000 gene expression samples to construct a universal dictionary of features across different tissues and pathologies. The dictionary validated on 35643 held-out samples for reconstruction error. It is then applied on 165 held-out datasets spanning a diverse range of diseases. The AutoML tool JADBio is employed to show that the predictive information in the PASL-created feature space is retained after the transformation. The code is available at https://github.com/mensxmachina/PASL .
Ioulia Karagiannaki, Krystallia Gourlia, Vincenzo Lagani, Yannis Pantazis, Ioannis Tsamardinos
Mach. Learn.4
2023 Cumulant GAN
abstract
In this article, we propose a novel loss function for training generative adversarial networks (GANs) aiming toward deeper theoretical understanding as well as improved stability and performance for the underlying optimization problem. The new loss function is based on cumulant generating functions (CGFs) giving rise to Cumulant GAN. Relying on a recently derived variational formula, we show that the corresponding optimization problem is equivalent to Rényi divergence minimization, thus offering a (partially) unified perspective of GAN losses: the Rényi family encompasses Kullback–Leibler divergence (KLD), reverse KLD, Hellinger distance, and$\chi ^{2}$-divergence. Wasserstein GAN is also a member of cumulant GAN. In terms of stability, we rigorously prove the linear convergence of cumulant GAN to the Nash equilibrium for a linear discriminator, Gaussian distributions, and the standard gradient descent ascent algorithm. Finally, we experimentally demonstrate that image generation is more robust relative to Wasserstein GAN and it is substantially improved in terms of both inception score (IS) and Fréchet inception distance (FID) when both weaker and stronger discriminators are considered.
Yannis Pantazis, Dipjyoti Paul, Michail Fasoulakis, Yannis Stylianou, Markos A. Katsoulakis
IEEE Trans. Neural Networks Learn. Syst.1
2022 Forward Looking Best-Response Multiplicative Weights Update Methods for Bilinear Zero-sum Games
abstract
Our work focuses on extra gradient learning algorithms for finding Nash equilibria in bilinear zero-sum games. The proposed method, which can be formally considered as a variant of Optimistic Mirror Descent (Mertikopoulos et al., 2019), uses a large learning rate for the intermediate gradient step which essentially leads to computing (approximate) best response strategies against the profile of the previous iteration. Although counter-intuitive at first sight due to the irrationally large, for an iterative algorithm, intermediate learning step, we prove that the method guarantees last-iterate convergence to an equilibrium. Particularly, we show that the algorithm reaches first an $\eta^{1/\rho}$-approximate Nash equilibrium, with $\rho > 1$, by decreasing the Kullback-Leibler divergence of each iterate by at least $\Omega(\eta^{1+\frac{1}{\rho}})$, for sufficiently small learning rate $\eta$, until the method becomes a contracting map, and converges to the exact equilibrium. Furthermore, we perform experimental comparisons with the optimistic variant of the multiplicative weights update method, by Daskalakis and Panageas (2019) and show that our algorithm has significant practical potential since it offers substantial gains in terms of accelerated convergence.
Michail Fasoulakis, Evangelos Markakis 0001, Yannis Pantazis, Konstantinos Varsos 0001
AISTATS3
2022 (f, Gamma)-Divergences: Interpolating between f-Divergences and Integral Probability Metrics
abstract
We develop a rigorous and general framework for constructing information-theoretic divergences that subsume both $f$-divergences and integral probability metrics (IPMs), such as the $1$-Wasserstein distance. We prove under which assumptions these divergences, hereafter referred to as $(f,\Gamma)$-divergences, provide a notion of `distance' between probability measures and show that they can be expressed as a two-stage mass-redistribution/mass-transport process. The $(f,\Gamma)$-divergences inherit features from IPMs, such as the ability to compare distributions which are not absolutely continuous, as well as from $f$-divergences, namely the strict concavity of their variational representations and the ability to control heavy-tailed distributions for particular choices of $f$. When combined, these features establish a divergence with improved properties for estimation, statistical learning, and uncertainty quantification applications. Using statistical learning as an example, we demonstrate their advantage in training generative adversarial networks (GANs) for heavy-tailed, not-absolutely continuous sample distributions. We also show improved performance and stability over gradient-penalized Wasserstein GAN in image generation.
Jeremiah Birrell, Paul Dupuis, Markos A. Katsoulakis, Yannis Pantazis, Luc Rey-Bellet
J. Mach. Learn. Res.4
2022 Optimizing Variational Representations of Divergences and Accelerating Their Statistical Estimation
abstract
Variational representations of divergences and distances between high-dimensional probability distributions offer significant theoretical insights and practical advantages in numerous research areas. Recently, they have gained popularity in machine learning as a tractable and scalable approach for training probabilistic models and for statistically differentiating between data distributions. Their advantages include: 1) They can be estimated from data as statistical averages. 2) Such representations can leverage the ability of neural networks to efficiently approximate optimal solutions in function spaces. However, a systematic and practical approach to improving the tightness of such variational formulas, and accordingly accelerate statistical learning and estimation from data, is currently lacking. Here we develop such a methodology for building new, tighter variational representations of divergences. Our approach relies on improved objective functionals constructed via an auxiliary optimization problem. Furthermore, the calculation of the functional Hessian of objective functionals unveils the local curvature differences around the common optimal variational solution; this quantifies and orders the tightness gains between different variational representations. Finally, numerical simulations utilizing neural network optimization demonstrate that tighter representations can result in significantly faster learning and more accurate estimation of divergences in both synthetic and real datasets (of more than 1000 dimensions), often accelerated by nearly an order of magnitude.
Jeremiah Birrell, Markos A. Katsoulakis, Yannis Pantazis
IEEE Trans. Inf. Theory3
2021 A Universal Multi-Speaker Multi-Style Text-to-Speech via Disentangled Representation Learning Based on Rényi Divergence Minimization
Dipjyoti Paul, Sankar Mukherjee, Yannis Pantazis, Yannis Stylianou
Interspeech3
2020 Latent Feature Representations for Human Gene Expression Data Improve Phenotypic Predictions
abstract
High-throughput technologies such as microarrays and RNA-sequencing (RNA-seq) allow to precisely quantify transcriptomic profiles, generating datasets that are inevitably high-dimensional. In this work, we investigate whether the whole human transcriptome can be represented in a compressed, low dimensional latent space without loosing relevant information. We thus constructed low-dimensional latent feature spaces of the human genome, by utilizing three dimensionality reduction approaches and a diverse set of curated datasets. We applied standard Principal Component Analysis (PCA), kernel PCA and Autoencoder Neural Networks on 1360 datasets from four different measurement technologies. The latent feature spaces are tested for their ability to (a) reconstruct the original data and (b) improve predictive performance on validation datasets not used during the creation of the feature space. While linear techniques show better reconstruction performance, nonlinear approaches, particularly, neural-based models seem to be able to capture non-additive interaction effects, and thus enjoy stronger predictive capabilities. Despite the limited sample size of each dataset and the biological / technological heterogeneity across studies, our results show that low dimensional representations of the human transcriptome can be achieved by integrating hundreds of datasets. The created space is two to three orders of magnitude smaller compared to the raw data, offering the ability of capturing a large portion of the original data variability and eventually reducing computational time for downstream analyses.
Yannis Pantazis, Christos Tselas, Kleanthi Lakiotaki, Vincenzo Lagani, Ioannis Tsamardinos
BIBM1
2020 Pathway Activity Score Learning for Dimensionality Reduction of Gene Expression Data
abstract
Abstract Molecular gene-expression datasets consist of samples with tens of thousands of measured quantities (e.g., high dimensional data). However, there exist lower-dimensional representations that retain the useful information. We present a novel algorithm for such dimensionality reduction called Pathway Activity Score Learning (PASL). The major novelty of PASL is that the constructed features directly correspond to known molecular pathways and can be interpreted as pathway activity scores. Hence, unlike PCA and similar methods, PASL’s latent space has a relatively straight-forward biological interpretation. As a use-case, PASL is applied on two collections of breast cancer and leukemia gene expression datasets. We show that PASL does retain the predictive information for disease classification on new, unseen datasets, as well as outperforming PLIER, a recently proposed competitive method. We also show that differential activation pathway analysis provides complementary information to standard gene set enrichment analysis. The code is available at https://github.com/mensxmachina/PASL .
Ioulia Karagiannaki, Yannis Pantazis, Ekaterini Chatzaki, Ioannis Tsamardinos
DS2
2020 Effects of Spectral Tilt on Listeners' Preferences And Intelligibility
abstract
High intelligibility can be achieved when listening to synthetic or artificially-produced speech under adverse conditions. But can listener preferences reveal any extra information when intelligibility is at ceiling? This paper describes a real-time speech modification technique which allows evaluation of the impact of individual speech properties on listeners' preferences. The current study investigates spectral tilt, a feature which also changes naturally in different speaking styles such as Lombard speech. In the listening experiment, participants were asked to adjust the spectral tilt in masked conditions; subsequently, intelligibility was assessed. Listeners preferred flatter spectral tilts as SNR decreased. Additionally, tilt preferences were evident even while intelligibility was at or close to ceiling levels, suggesting that preferences provide information over and above that measured in traditional intelligibility-based studies. On the basis of these findings a method for probabilistic modelling of listener preferences, useful in the development of speech enrichment algorithms, is proposed.
Olympia Simantiraki, Martin Cooke, Yannis Pantazis
ICASSP3
2020 Speaker Conditional WaveRNN: Towards Universal Neural Vocoder for Unseen Speaker and Recording Conditions
abstract
Recent advancements in deep learning led to human-level performance in single-speaker speech synthesis. However, there are still limitations in terms of speech quality when generalizing those systems into multiple-speaker models especially for unseen speakers and unseen recording qualities. For instance, conventional neural vocoders are adjusted to the training speaker and have poor generalization capabilities to unseen speakers. In this work, we propose a variant of WaveRNN, referred to as speaker conditional WaveRNN (SC-WaveRNN). We target towards the development of an efficient universal vocoder even for unseen speakers and recording conditions. In contrast to standard WaveRNN, SC-WaveRNN exploits additional information given in the form of speaker embeddings. Using publicly-available data for training, SC-WaveRNN achieves significantly better performance over baseline WaveRNN on both subjective and objective metrics. In MOS, SC-WaveRNN achieves an improvement of about 23% for seen speaker and seen recording condition and up to 95% for unseen speaker and unseen condition. Finally, we extend our work by implementing a multi-speaker text-to-speech (TTS) synthesis similar to zero-shot speaker adaptation. In terms of performance, our system has been preferred over the baseline TTS system by 60% over 15.5% and by 60.9% over 32.6%, for seen and unseen speakers, respectively.
Dipjyoti Paul, Yannis Pantazis, Yannis Stylianou
INTERSPEECH2
2020 Enhancing Speech Intelligibility in Text-To-Speech Synthesis Using Speaking Style Conversion
abstract
The increased adoption of digital assistants makes text-to-speech (TTS) synthesis systems an indispensable feature of modern mobile devices. It is hence desirable to build a system capable of generating highly intelligible speech in the presence of noise. Past studies have investigated style conversion in TTS synthesis, yet degraded synthesized quality often leads to worse intelligibility. To overcome such limitations, we proposed a novel transfer learning approach using Tacotron and WaveRNN based TTS synthesis. The proposed speech system exploits two modification strategies: (a) Lombard speaking style data and (b) Spectral Shaping and Dynamic Range Compression (SSDRC) which has been shown to provide high intelligibility gains by redistributing the signal energy on the time-frequency domain. We refer to this extension as Lombard-SSDRC TTS system. Intelligibility enhancement as quantified by the Intelligibility in Bits (SIIB-Gauss) measure shows that the proposed Lombard-SSDRC TTS system shows significant relative improvement between 110% and 130% in speech-shaped noise (SSN), and 47% to 140% in competing-speaker noise (CSN) against the state-of-the-art TTS approach. Additional subjective evaluation shows that Lombard-SSDRC TTS successfully increases the speech intelligibility with relative improvement of 455% for SSN and 104% for CSN in median keyword correction rate compared to the baseline TTS method.
Dipjyoti Paul, P. V. Muhammed Shifas, Yannis Pantazis, Yannis Stylianou
INTERSPEECH3
2019 Towards a Robust and Accurate Screening Tool for Dyslexia with Data Augmentation using GANs
abstract
Eye movements during text reading can provide insights about reading disorders. We developed the DysLexML, a screening tool for developmental dyslexia, based on various ML algorithms that analyze gaze points recorded via eye-tracking during silent reading of children. We comparatively evaluated its performance using measurements collected from two systematic field studies with 221 participants in total. This work presents DysLexML and its performance. It identifies the features with prominent predictive power and performs dimensionality reduction. Specifically, it achieves its best performance using linear SVM, with an accuracy of 97% and 84% respectively, using a small feature set. We show that DysLexML is also robust in the presence of noise. These encouraging results set the basis for developing screening tools in less controlled, larger-scale environments, with inexpensive eye-trackers, potentially reaching a larger population for early intervention. Unlike other related studies, DysLexML achieves the aforementioned performance by employing only a small number of selected features, that have been identified with prominent predictive power. Finally, we developed a new data augmentation/substitution technique based on GANs for generating synthetic data similar to the original distributions.
Thomais Asvestopoulou, Victoria Manousaki, Antonis Psistakis, Erjona Nikolli, Vassilios Andreadakis, Ioannis M. Aslanides, Yannis Pantazis, Ioannis Smyrnakis, Maria Papadopouli
BIBE7
2019 Speech Enhancement for Noise-Robust Speech Synthesis Using Wasserstein GAN
Nagaraj Adiga, Yannis Pantazis, Vassilis Tsiaras, Yannis Stylianou
INTERSPEECH2
2019 Non-Parallel Voice Conversion Using Weighted Generative Adversarial Networks
abstract
In this paper, we suggest a novel way to train GenerativeAdversarial Network (GAN) for the purpose of non-parallel,many-to-many voice conversion. The goal of voice conversion(VC) is to transform speech from a source speaker to that of atarget speaker without changing the phonetic contents. Basedon ideas from Game Theory, we suggest to multiply the gradi-ent of the Generator with suitable weights. Weights are calcu-lated so that they increase the power of fake samples that foolthe Discriminator resulting in a stronger Generator. Motivatedby a recently presented GAN based approach for VC, StarGAN-VC, we suggest a variation to StarGAN, referred to as WeightedStarGAN (WeStarGAN). The experiments are conducted onstandard CMU ARCTIC database. WeStarGAN-VC approachachieves significantly better relative performance and is clearlypreferred over recently proposed StarGAN-VC method in termsof speech subjective quality and speaker similarity with 75% and 65%preference scores, respectively.
Dipjyoti Paul, Yannis Pantazis, Yannis Stylianou
INTERSPEECH2
2019 A unified approach for sparse dynamical system inference from temporal measurements
abstract
MOTIVATION: Temporal variations in biological systems and more generally in natural sciences are typically modeled as a set of ordinary, partial or stochastic differential or difference equations. Algorithms for learning the structure and the parameters of a dynamical system are distinguished based on whether time is discrete or continuous, observations are time-series or time-course and whether the system is deterministic or stochastic, however, there is no approach able to handle the various types of dynamical systems simultaneously. RESULTS: In this paper, we present a unified approach to infer both the structure and the parameters of non-linear dynamical systems of any type under the restriction of being linear with respect to the unknown parameters. Our approach, which is named Unified Sparse Dynamics Learning (USDL), constitutes of two steps. First, an atemporal system of equations is derived through the application of the weak formulation. Then, assuming a sparse representation for the dynamical system, we show that the inference problem can be expressed as a sparse signal recovery problem, allowing the application of an extensive body of algorithms and theoretical results. Results on simulated data demonstrate the efficacy and superiority of the USDL algorithm under multiple interventions and/or stochasticity. Additionally, USDL's accuracy significantly correlates with theoretical metrics such as the exact recovery coefficient. On real single-cell data, the proposed approach is able to induce high-confidence subgraphs of the signaling pathway. AVAILABILITY AND IMPLEMENTATION: Source code is available at Bioinformatics online. USDL algorithm has been also integrated in SCENERY (http://scenery.csd.uoc.gr/); an online tool for single-cell mass cytometry analytics. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Yannis Pantazis, Ioannis Tsamardinos
Bioinform.1
2013 Parametric sensitivity analysis for biochemical reaction networks based on pathwise information theory
abstract
BACKGROUND: Stochastic modeling and simulation provide powerful predictive methods for the intrinsic understanding of fundamental mechanisms in complex biochemical networks. Typically, such mathematical models involve networks of coupled jump stochastic processes with a large number of parameters that need to be suitably calibrated against experimental data. In this direction, the parameter sensitivity analysis of reaction networks is an essential mathematical and computational tool, yielding information regarding the robustness and the identifiability of model parameters. However, existing sensitivity analysis approaches such as variants of the finite difference method can have an overwhelming computational cost in models with a high-dimensional parameter space. RESULTS: We develop a sensitivity analysis methodology suitable for complex stochastic reaction networks with a large number of parameters. The proposed approach is based on Information Theory methods and relies on the quantification of information loss due to parameter perturbations between time-series distributions. For this reason, we need to work on path-space, i.e., the set consisting of all stochastic trajectories, hence the proposed approach is referred to as "pathwise". The pathwise sensitivity analysis method is realized by employing the rigorously-derived Relative Entropy Rate, which is directly computable from the propensity functions. A key aspect of the method is that an associated pathwise Fisher Information Matrix (FIM) is defined, which in turn constitutes a gradient-free approach to quantifying parameter sensitivities. The structure of the FIM turns out to be block-diagonal, revealing hidden parameter dependencies and sensitivities in reaction networks. CONCLUSIONS: As a gradient-free method, the proposed sensitivity analysis provides a significant advantage when dealing with complex stochastic systems with a large number of parameters. In addition, the knowledge of the structure of the FIM can allow to efficiently address questions on parameter identifiability, estimation and robustness. The proposed method is tested and validated on three biochemical systems, namely: (a) a protein production/degradation model where explicit solutions are available, permitting a careful assessment of the method, (b) the p53 reaction network where quasi-steady stochastic oscillations of the concentrations are observed, and for which continuum approximations (e.g. mean field, stochastic Langevin, etc.) break down due to persistent oscillations between high and low populations, and (c) an Epidermal Growth Factor Receptor model which is an example of a high-dimensional stochastic reaction network with more than 200 reactions and a corresponding number of parameters.
Yannis Pantazis, Markos A. Katsoulakis, Dionisios G. Vlachos
BMC Bioinform.1
2012 An extension of the adaptive Quasi-Harmonic Model
abstract
In this paper, we present an extension of a recently developed AM-FM decomposition algorithm, which will be referred to as the extended adaptive Quasi-Harmonic Model (eaQHM). It was previously shown that the adaptive Quasi-Harmonic Model (aQHM) [1] is an efficient AM-FM decomposition algorithm with applications in speech analysis. In this paper, we show that a simple extension of the aQHM algorithm to include not only frequency but also amplitude adaptation results in higher performance in terms of Signal-to-Reconstruction-Error Ratio (SRER). To support our hypothesis, eaQHM is tested both on synthetic signals and on a subset of the ARCTIC database of speech. Overall, compared with aQHM, eaQHM improves the SRER by more than 2 dB, on average.
George P. Kafentzis, Yannis Pantazis, Olivier Rosec, Yannis Stylianou
ICASSP2
2011 Adaptive AM-FM Signal Decomposition With Application to Speech Analysis
abstract
In this paper, we present an iterative method for the accurate estimation of amplitude and frequency modulations (AM-FM) in time-varying multi-component quasi-periodic signals such as voiced speech. Based on a deterministic plus noise representation of speech initially suggested by Laroche (“HNM: A simple, efficient harmonic plus noise model for speech,” Proc. WASPAA, Oct., 1993, pp. 169-172), and focusing on the deterministic representation, we reveal the properties of the model showing that such a representation is equivalent to a time-varying quasi-harmonic representation of voiced speech. Next, we show how this representation can be used for the estimation of amplitude and frequency modulations and provide the conditions under which such an estimation is valid. Finally, we suggest an adaptive algorithm for nonparametric estimation of AM-FM components in voiced speech. Based on the estimated amplitude and frequency components, a high-resolution time-frequency representation is obtained. The suggested approach was evaluated on synthetic AM-FM signals, while using the estimated AM-FM information, speech signal reconstruction was performed, resulting in a high signal-to-reconstruction error ratio (around 30 dB).
Yannis Pantazis, Olivier Rosec, Yannis Stylianou
IEEE Trans. Speech Audio Process.1
2010 On the robustness of the Quasi-Harmonic model of speech
abstract
In this paper we discuss the robustness of the Quasi-Harmonic model, QHM, previously suggested for speech analysis and AM-FM decomposition of speech. Assuming a frame by frame analysis, QHM suggests an iterative estimator for the actual frequencies of the speech components at the center of analysis window. In this paper, we show that this is a biased estimator and then, we compute analytically and numerically the bias of the estimator showing its dependence on the type and length of the analysis window. Moreover, we analyze the robustness of the QHM estimator in white Gaussian noise, showing that the suggested iterative estimator asymptotically attains the corresponding Cramer-Rao lower bound even in adverse noisy conditions. Examples of synthetic signals are provided to support our analysis.
Yannis Pantazis, Olivier Rosec, Yannis Stylianou
ICASSP1
2010 Analysis/synthesis of speech based on an adaptive quasi-harmonic plus noise model
abstract
Decomposition of speech into a deterministic part and a stochastic part is a typical modeling. Usually, the deterministic part in voiced speech is modeled as a sum of time-varying sinusoids while the stochastic part is modeled as modulated noise. The estimation of sinusoidal parameters assumes that locally speech is a stationary signal. However, this is not true leading to biased amplitude and phase estimation. In this paper, we develop a scheme for speech analysis and synthesis which is able to deal with locally nonstationary frames. Thus, deterministic part it modeled using an adaptive quasi-harmonic model while stochastic part is modeled as time-modulated and frequency-modulated noise. Results show that the reconstructed signal is almost indistinguishable from the original.
Yannis Pantazis, Georgios Tzedakis, Olivier Rosec, Yannis Stylianou
ICASSP1
2010 Fast least-squares solution for sinusoidal, harmonic and quasi-harmonic models
Georgios Tzedakis, Yannis Pantazis, Olivier Rosec, Yannis Stylianou
INTERSPEECH2
2010 Iterative Estimation of Sinusoidal Signal Parameters
abstract
While the problem of estimating the amplitudes of sinusoidal components in signals, given an estimation of their frequencies, is linear and tractable, it is biased due to the unavoidable, in practice, errors in the estimation of frequencies. These errors are of great concern for processing signals with many sinusoidal like components as is the case of speech and audio. In this letter, we suggest using a time-varying sinusoidal representation which is able to iteratively correct frequency estimation errors. Then the corresponding amplitudes are computed through Least Squares. Experiments conducted on synthetic and speech signals show the suggested model's effectiveness in correcting frequency estimation errors and robustness in additive noise conditions.
Yannis Pantazis, Olivier Rosec, Yannis Stylianou
IEEE Signal Process. Lett.1
2010 Reply to "Comments on 'Iterative Estimation of Sinusoidal Signal Parameters'"
abstract
The method proposed in , referred to as iterative Quasi-Harmonic Model (iQHM), is an iterative frequency estimation technique. It is based on a linearized version of the frequency error between the true frequency and the initially provided frequency, while the estimation of its unknown parameters is performed by Least Squares (LS) method. In , a relationship between iQHM and Gauss-Newton (GN) method was presented. More specifically it was claimed that iQHM is actually equivalent to an approximate Gauss-Newton method (AGN). In this correspondence, we show that iQHM is actually equivalent to a sequential version of the GN method.
Yannis Pantazis, Olivier Rosec, Yannis Stylianou
IEEE Signal Process. Lett.1
2009 Chirp rate estimation of speech based on a time-varying quasi-harmonic model
abstract
The speech signal is usually considered as stationary during short analysis time intervals. Though this assumption may be sufficient in some applications, it is not valid for high-resolution speech analysis and in applications such as speech transformation and objective voice function assessment for detection of voice disorders. In speech, there are non stationary components, for instance time-varying amplitudes and frequencies, which may change quickly over short time intervals. In this paper, a previously suggested time-varying quasi-harmonic model is extended in order or to estimate the chirp rate for each sinusoidal component, thus successfully tracking fast variations in frequency and amplitude. The parameters of the model are estimated through linear Least Squares and the model accuracy is evaluated on synthetic chirp signals. Experiments on speech signals indicate that the new model is able to efficiently estimate the signal component chirp rates, providing means to develop more accurate speech models for high-quality speech transformations.
Yannis Pantazis, Olivier Rosec, Yannis Stylianou
ICASSP1
2009 AM-FM estimation for speech based on a time-varying sinusoidal model
Yannis Pantazis, Olivier Rosec, Yannis Stylianou
INTERSPEECH1
2008 Improving the modeling of the noise part in the harmonic plus noise model of speech
abstract
Harmonic + noise model (HNM) is a hybrid model of speech with a harmonic component and a noise component. While the harmonic part describes efficiently the periodicities in speech signals (voiced parts), modeling of the noise part introduces artifacts primarily because of the specific time-domain characteristics of noise in voiced speech. In this paper, we concentrated on the modeling of noise in voiced frames. To model the temporal characteristics of noise, we study three time envelopes in the context of HNM; triangular envelope, Hilbert envelope and energy envelope. Listening tests showed a clear preference for the Energy envelope and Hilbert envelope for male voices and to a lesser extent the same conclusions can be drawn for female voices.
Yannis Pantazis, Yannis Stylianou
ICASSP1
2008 On the properties of a time-varying quasi-harmonic model of speech
Yannis Pantazis, Olivier Rosec, Yannis Stylianou
INTERSPEECH1
2005 Discontinuity detection in concatenated speech synthesis based on nonlinear speech analysis
abstract
An objective distance measure which is able to predict audible discontinuity in concatenated speech synthesis systems is very important. Previous works were primarily based on features estimated by linear and/or stationary models of speech. In this paper, we introduce two nonlinear approaches for the detection of discontinuity. The first method is based on a nonlinear harmonic model of speech while the second method is based on the demodulation of speech in an amplitude and a frequency component using the Teager energy operator. Fisher’s linear discriminant was used for the separation of signals with audible discontinuity from those perceived as continuous. When we combined the two methods using Fisher’s linear discriminant a detection rate of 56.5 % was achieved which is an 90 % improvement over previously published results on the same database. 1.
Yannis Pantazis, Yannis Stylianou, Esther Klabbers
INTERSPEECH1