VLDB 2026 Research / reviewers in the wild / expert
Yannis Agiomyrgiannakis
dblp:08/8055
· DBLP profile ↗
21ranked-venue papers
14as first author
0since 2021 · last 2018
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 19 · 12 first-authorArtificial intelligence and machine learning · 8 · 5 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
2 papers |
Audio and music processing · 100% |
Topics — the 1 heaviest of 2, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Audio and music processing
speech coding |
0.2 | 2 | 2009 | Wrapped Gaussian Mixture Models for Modeling and High-Rate Quantization of Phase Data of Speech · IEEE Trans. Speech Audio Process. 2009 Conditional Vector Quantization for Speech Coding · IEEE Trans. Speech Audio Process. 2007 |
Methods — techniques the papers use, named apart from their topics
scalar quantization · 0.1expectation-maximization · 0.1vector quantization · 0.1divide-and-conquer · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2018 | B-Spline Pdf: A Generalization of Histograms to Continuous Density Models for Generative Audio NetworksabstractMany modern neural networks use histograms to efficiently model continuous random variables. This implies that the parametric space of the multinomial distribution is easier for training large neural networks. In applications like generative audio networks, this approach introduces audible quantization noise to the generated signal. This work presents a novel probability density function (PDF), referred to as B-Spline PDF, that is a direct generalization of histograms to continuous densities while retaining the multinomial parameter space. The latter uses k-th order B-Splines to ensure continuity up to the (k-1)-th order derivative. B-Spline PDF is amenable for neural network training via closed-form gradients that are easy and fast to compute. For other applications, one may use a novel algorithm, referred to as the Expectation algorithm, to efficiently estimate the model parameters. Further, a novel sample generation algorithm is derived that is fast and simple. The theoretical results, coupled with illustrative examples, suggest that B-Spline PDF may directly replace histograms in many related applications. Yannis Agiomyrgiannakis |
ICASSP | 1 |
| 2018 | Natural TTS Synthesis by Conditioning Wavenet on MEL Spectrogram PredictionsabstractThis paper describes Tacotron 2, a neural network architecture for speech synthesis directly from text. The system is composed of a recurrent sequence-to-sequence feature prediction network that maps character embeddings to mel-scale spectrograms, followed by a modified WaveNet model acting as a vocoder to synthesize time-domain waveforms from those spectrograms. Our model achieves a mean opinion score (MOS) of 4.53 comparable to a MOS of 4.58 for professionally recorded speech. To validate our design choices, we present ablation studies of key components of our system and evaluate the impact of using mel spectrograms as the conditioning input to WaveNet instead of linguistic, duration, and F0 features. We further show that using this compact acoustic intermediate representation allows for a significant reduction in the size of the WaveNet architecture. Jonathan Shen, Ruoming Pang, Ron J. Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Yu Zhang 0033, Yuxuan Wang 0002, R. J. Skerry-Ryan, Rif A. Saurous, Yannis Agiomyrgiannakis |
ICASSP | 12 |
| 2017 | Google's Next-Generation Real-Time Unit-Selection Synthesizer Using Sequence-to-Sequence LSTM-Based Autoencoders
Vincent Wan, Yannis Agiomyrgiannakis, Hanna Silén, Jakub Vit |
INTERSPEECH | 2 |
| 2017 | Tacotron: Towards End-to-End Speech SynthesisabstractA text-to-speech synthesis system typically consists of multiple stages, such as a text analysis frontend, an acoustic model and an audio synthesis module.Building these components often requires extensive domain expertise and may contain brittle design choices.In this paper, we present Tacotron, an end-to-end generative text-to-speech model that synthesizes speech directly from characters.Given pairs, the model can be trained completely from scratch with random initialization.We present several key techniques to make the sequence-tosequence framework perform well for this challenging task.Tacotron achieves a 3.82 subjective 5-scale mean opinion score on US English, outperforming a production parametric system in terms of naturalness.In addition, since Tacotron generates speech at the frame level, it's substantially faster than sample-level autoregressive methods. Yuxuan Wang 0002, R. J. Skerry-Ryan, Daisy Stanton, Ron J. Weiss, Navdeep Jaitly, Zongheng Yang, Samy Bengio, Quoc V. Le, Yannis Agiomyrgiannakis, Rob Clark, Rif A. Saurous |
INTERSPEECH | 12 |
| 2016 | The matching-minimization algorithm, the INCA algorithm and a mathematical framework for voice conversion with unaligned corporaabstractThis paper presents a mathematical framework that is suitable for voice conversion and adaptation in speech processing. Voice conversion is formulated as a search for the optimal correspondances between a set of source-speaker spectra and a set of target-speaker spectra under a transform that compensates speaker differences. It is possible to simultaneously recover a bi-directional mapping between two sets of vectors that is a parametric mapping (a transform) in one direction and a non-parametric mapping (correspondences) in the reverse direction. An algorithm referred to as Matching-Minimization (MM) is formally derived with proven convergence and an optimal closed-form solution for each step. The algorithm is closely related to the asymmetric-1 variant of the well-known INCA algorithm [1] for which we also provide a proof within the same framework. The differences between MM and INCA are delineated both theoretically and experimentally. MM outperforms INCA in all scenarios. Like INCA, MM does not require parallel corpora. Unlike INCA, MM is suitable when only a few adaptation data are available. Yannis Agiomyrgiannakis |
ICASSP | 1 |
| 2016 | Voice Morphing that improves TTS quality using an optimal dynamic frequency warping-and-weighting transformabstractDynamic Frequency Warping (DFW) is widely used to align spectra of different speakers. It has long been argued that frequency warping captures inter-speaker differences but DFW practice always involves a tricky preprocessing part to remove spectral tilt. The DFW residual is successfully used in Voice Morphing to improve the quality and the similarity of synthesized speech but the estimation of the DFW residual remains largely heuristic and sub-optimal. This paper presents a dynamic programming algorithm that simultaneously estimates the Optimal Frequency Warping and Weighting transform (ODFWW) and therefore needs no preprocessing step and fine-tuning while source/target-speaker data are matched using the Matching-Minimization algorithm [1]. The transform is used to morph the output of a state-of-the-art Vocaine-based [2] TTS synthesizer in order to generate different voices in runtime with only +8% computational overhead. Some morphed TTS voices exhibit significantly higher quality than the original one as morphing seems to "correct" the voice characteristics of the TTS voice. Yannis Agiomyrgiannakis, Zoi Roupakia |
ICASSP | 1 |
| 2016 | Fast, Compact, and High Quality LSTM-RNN Based Statistical Parametric Speech Synthesizers for Mobile DevicesabstractAcoustic models based on long short-term memory recurrent neural networks (LSTM-RNNs) were applied to statistical parametric speech synthesis (SPSS) and showed significant improvements in naturalness and latency over those based on hidden Markov models (HMMs). This paper describes further optimizations of LSTM-RNN-based SPSS for deployment on mobile devices; weight quantization, multi-frame inference, and robust inference using an ε-contaminated Gaussian loss function. Experimental results in subjective listening tests show that these optimizations can make LSTM-RNN-based SPSS comparable to HMM-based SPSS in runtime speed while maintaining naturalness. Evaluations between LSTM-RNN- based SPSS and HMM-driven unit selection speech synthesis are also presented. Heiga Zen, Yannis Agiomyrgiannakis, Niels Egberts, Fergus Henderson, Przemyslaw Szczepaniak |
INTERSPEECH | 2 |
| 2015 | Vocaine the vocoder and applications in speech synthesisabstractVocoders received renewed attention recently as basic components in speech synthesis applications such as voice transformation, voice conversion and statistical parametric speech synthesis. This paper presents a new vocoder synthesizer, referred to as Vocaine, that features a novel Amplitude Modulated-Frequency Modulated (AM-FM) speech model, a new way to synthesize non-stationary sinusoids using quadratic phase splines and a super fast cosine generator. Extensive evaluations are made against several state-of-the-art methods in Copy-Synthesis and Text-To-Speech synthesis experiments. Vocaine matches or outperforms STRAIGHT in Copy-Synthesis experiments and outperforms our baseline real-time optimized Mixed-Excitation vocoder with the same computational cost. We report that Vocaine considerably improves our statistical TTS synthesizers and that our new statistical parametric synthesizer [1] matched the quality of our mature production Unit-Selection system with uncompressed waveforms. Yannis Agiomyrgiannakis |
ICASSP | 1 |
| 2014 | A frequency-weighted post-filtering transform for compensation of the over-smoothing effect in HMM-based speech synthesisabstractOver-smoothing is one of the major sources of quality degradation in statistical parametric speech synthesis. Many methods have been proposed to compensate over-smoothing with the speech parameter generation algorithm considering Global Variance (GV) being one of the most successfull. This paper models over-smoothing as a radial relocation of poles and zeros of the spectral envelope towards the origin of the z-plane and uses radial scaling to enhance spectral peaks and to deepen spectral valeys. The radial scaling technique is improved by introducing over-emphasis, spectral-tilt compensation and frequency weighting. Listening test results indicate that the proposed method is 11%-13% more preferable than GV while it has less algorithmic delay (only 5 ms) and computational complexity. Florian Eyben, Yannis Agiomyrgiannakis |
ICASSP | 2 |
| 2011 | ON the recovery of time-varying spectral envelope information from AQHM-derived spectraabstractSpectral envelopes of speech signals are typically obtained by making stationarity assumptions about the signal which are not always valid. The Adaptive Quasi-Harmonic Model (AQHM), a non-stationary signal model, is capable of capturing the time-varying quasi harmonics in voiced speech. This paper suggests the use of AQHM in a multi-layer scheme which results in a high-resolution time-frequency representation of speech. This representation is then used for the recovery of the evolving spectral envelope and thus, a time-frequency spectral envelope estimation algorithm is introduced related to the Papoulis-Gerchberg algorithm for data extrapolation. Results on voiced speech sounds show that the estimated spectral envelopes are smoother than those estimated by state-of-the-art spectral envelope estimators, while maintaining the important spectral details of the speech spectrum. Yannis Agiomyrgiannakis, Yannis Stylianou |
ICASSP | 1 |
| 2009 | ARX-LF-based source-filter methods for voice modification and transformationabstractTwo ARX-LF-based source/filters models for speech signals are presented. A robust glottal inversion technique is used to deconvolve the signal into an excitation component and a filter component. The excitation component is further decomposed into an LF part and a residual part. The first model, referred to as the LF-vocoder, is a high quality vocoder that replaces the residual part with modulated noise. The second model uses a sinusoidal harmonic representation of the residual signal. The latter does not degrade the signal during analysis/synthesis and provides higher quality for small modification factors, while the former has the advantage of being a compact, fully parametric representation that is suitable for low-bit-rate speech coding as well as parametric speech synthesis applications. Yannis Agiomyrgiannakis, Olivier Rosec |
ICASSP | 1 |
| 2009 | Wrapped Gaussian Mixture Models for Modeling and High-Rate Quantization of Phase Data of SpeechabstractThe harmonic representation of speech signals has found many applications in speech processing. This paper presents a novel statistical approach to model the behavior of harmonic phases. Phase information is decomposed into three parts: a minimum phase part, a translation term, and a residual term referred to as dispersion phase. Dispersion phases are modeled by wrapped Gaussian mixture models (WGMMs) using an expectation-maximization algorithm suitable for circular vector data. A multivariate WGMM-based phase quantizer is then proposed and constructed using novel scalar quantizers for circular random variables. The proposed phase modeling and quantization scheme is evaluated in the context of a narrowband harmonic representation of speech. Results indicate that it is possible to construct a variable-rate harmonic codec that is equivalent to iLBC at approximately 13 kbps. Yannis Agiomyrgiannakis, Yannis Stylianou |
IEEE Trans. Speech Audio Process. | 1 |
| 2008 | Towards flexible speech coding for speech synthesis: an LF + modulated noise vocoder
Yannis Agiomyrgiannakis, Olivier Rosec |
INTERSPEECH | 1 |
| 2007 | Stochastic Modeling and Quantization of Harmonic Phases in Speech using Wrapped Gaussian Mixture ModelsabstractHarmonic sinusoidal representations of speech have proven to be useful in many speech processing tasks. This work focuses on the phase spectra of the harmonics and provides a methodology to analyze and subsequently to model the statistics of the harmonic phases. To do so, we propose the use of a wrapped Gaussian mixture model (WGMM), a model suitable for random variables that belong to circular spaces, and provide an expectation-maximization algorithm for training. The WGMM is then used to construct a phase quantizer. The quantizer is employed in a prototype variable rate narrow-band VoIP sinusoidal codec that is equivalent to iLBC in terms of PESQ-MOS, at ~13 kbps. Yannis Agiomyrgiannakis, Yannis Stylianou |
ICASSP (4) | 1 |
| 2007 | Conditional Vector Quantization for Voice ConversionabstractVoice conversion methods have the objective of transforming speech spoken by a particular source speaker, so that it sounds as if spoken by a different target speaker. The majority of voice conversion methods is based on transforming the short-time spectral envelope of the source speaker, based on derived correspondences between the source and target vectors using training speech data from both speakers. These correspondences are usually obtained by segmenting the spectral vectors of one or both speakers into clusters, using soft (GMM-based) or hard (VQ-based) clustering. Here, we propose that voice conversion performance can be improved by taking advantage of the fact that often the relationship between the source and target vectors is one-to-many. In order to illustrate this, we propose that a VQ approach namely constrained vector quantization (CVQ), can be used for voice conversion. Results indicate that indeed such a relationship between the source and target data exists and can be exploited by following a CVQ-based function for voice conversion. Athanasios Mouchtaris, Yannis Agiomyrgiannakis, Yannis Stylianou |
ICASSP (4) | 2 |
| 2007 | The harmonic model codec (HMC) framework for voIP
Yannis Agiomyrgiannakis, Yannis Stylianou |
INTERSPEECH | 1 |
| 2007 | Bit-erasure channel decoding for GMM-based multiple description coding
Yannis Agiomyrgiannakis, Yannis Stylianou |
INTERSPEECH | 1 |
| 2007 | Conditional Vector Quantization for Speech CodingabstractIn many speech-coding-related problems, there is available information and lost information that must be recovered. When there is significant correlation between the available and the lost information source, coding with side information (CSI) can be used to benefit from the mutual information between the two sources. In this paper, we consider CSI as a special VQ problem which will be referred to as conditional vector quantization (CVQ). A fast two-step divide-and-conquer solution is proposed. CVQ is then used in two applications: the recovery of highband (4-8 kHz) spectral envelopes for speech spectrum expansion and the recovery of lost narrowband spectral envelopes for voice over IP. Comparisons with alternative approaches like estimation and simple VQ-based schemes show that CVQ provides significant distortion reductions at very low bit rates. Subjective evaluations indicate that CVQ provides noticeable perceptual improvements over the alternative approaches Yannis Agiomyrgiannakis, Yannis Stylianou |
IEEE Trans. Speech Audio Process. | 1 |
| 2006 | Fast Analysis/Synthesis of Harmonic SignalsabstractHarmonic models are commonly used in signal processing. The analysis of harmonic signals requires the solution of a symmetric Toeplitz system of equations. Levinson-based Toeplitz solvers have a O(n 2) complexity. This paper proposes an O(n) algorithm by encoding the inverse matrices required for the solution of the linear system to a few parameters in order to obtain an approximate solution for the harmonic model. For speech related applications, the proposed algorithm is 2-30 times faster than the Levinson algorithm, while degradation is minimal and memory requirements are very low Miltiadis Vasilakis, Yannis Agiomyrgiannakis, Yannis Stylianou |
ICASSP (3) | 2 |
| 2005 | Coding with Side Information Techniques for LSF Reconstruction in Voice Over IPabstractIn applications like VoIP, speech codecs have to deal with excessive packet losses, caused by network errors and/or delays. In this paper, a new method for the reconstruction of lost speech spectral envelopes is presented, which is based on a statistical estimation function. We suggest the usage of a minimal "corrective" bitstream and propose coding with side information (CSI) techniques for an efficient forward error correction (FEC) strategy. The proposed methods are tested on multiple scenarios of missing frames. Objective results indicate that with only 4 bits per lost frame, a spectral distortion reduction of 0.77-1.14 dB is achieved, compared to results obtained by current state-of-the-art estimation methods. Compared to "predictive" estimation methods, the use of the jitter buffer as side information and 4 bits per lost frame provide a 42% reduction of spectral distortion for single packet losses, and a 32% reduction for double packet losses. Subjective results indicate that the corrected speech has fewer artifacts. Yannis Agiomyrgiannakis, Yannis Stylianou |
ICASSP (1) | 1 |
| 2004 | Combined estimation/coding of highband spectral envelopes for speech spectrum expansionabstractThe paper addresses the problem of expanding the bandwidth of narrowband speech signals, focusing on the estimation of highband spectral envelopes. It is well known that there is not enough mutual information between the two bands. We show that this happens because narrowband spectral envelopes have a one-to-many relationship with highband spectral envelopes. A combined estimation/coding scheme for the missing spectral envelope is proposed, which employs this relationship to produce a high quality highband reconstruction, provided that there is an appropriate excitation. Subjective tests using the TIMIT database indicate that 134 bits/sec for the highband spectral envelope are adequate for a DCR (degradation category rating) score of 4.41. This is an improvement of 22.8% over a typical estimation of highband envelopes using the usual mapping functions, in terms of DCR score. Yannis Agiomyrgiannakis, Yannis Stylianou |
ICASSP (1) | 1 |