VLDB 2026 Research / reviewers in the wild / expert
Christer Gobl
dblp:00/4657
· DBLP profile ↗
49ranked-venue papers
8as first author
5since 2021 · last 2023
0000-0002-5958-3891ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 44 · 8 first-author · 5 since 2021Artificial intelligence and machine learning · 39 · 6 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Towards Dialect-inclusive Recognition in a Low-resource Language: Are Balanced Corpora the Answer?
Liam Lonergan, Neasa Ní Chiaráin, Christer Gobl, Ailbhe Ní Chasaide |
INTERSPEECH | 4 |
| 2023 | A System for Generating Voice Source Signals that Implements the Transformed LF-model Parameter ControlabstractThis paper describes a system which fully implements the \ntransformed LF glottal flow model, incorporating also the \noften-overlooked k-factors. A problem with the original \nproposal is that the global waveshape parameter Rd, central to \nthe transformed model and used to predict some of the other \nparameters, cannot be directly controlled. Instead, a stylisation \nof the model was used to indirectly control Rd. However, this \napproach can yield substantial errors in the Rd value of the \npulse, thus undermining the usefulness of the model. To \novercome this problem, an iterative algorithm is presented, \nwhich ensures that for a given Rd input value, a pulse with that \nRd value will be produced. Using this new Rd control, it \ntranspired that some of the original parameter predictions give \nrise to combinations of values that are incompatible with the LF \nmodel. Modifications to the original predictions of the Rk \nparameter were incorporated to ensure model conformity. Christer Gobl |
INTERSPEECH | 2 |
| 2022 | Cross-dialect lexicon optimisation for an endangered language ASR system: the case of IrishabstractLexicon optimisation strategies, addressing the problem of dialect divergence, are tested in an ASR system for Irish. As in many endangered languages, Irish has no spoken standard, but rather, three very different dialects of Ulster (Ul), Connaught (Co) and Munster (Mu). Furthermore, the complex sound system and ancient, opaque writing system result in sound-to-grapheme mappings that differ considerably across dialects. A hybrid ASR system was trained on (predominantly) native speaker speech data, balanced across the dialects. Experiment 1 tested whether a Global lexicon, which captures dialect variant forms with relatively abstract representations, can perform as well as a Multi-dialect lexicon containing all dialect variants. Three dialect-specific lexicons were also included in the tests. The Global lexicon did yield the best performance and experiment 2 tested whether further reductions to its phoneset might further enhance its performance. These included (i) merging a Tense-Lax contrast among coronal sonorants, not common to all dialects, and (ii) merging the contrast of voiceless-voiced sonorants, as the voiceless member is relatively infrequent. Results showed but a slight enhancement and only for Mu dialect, which is the one most aligned to the phoneset reduction. Liam Lonergan, Neasa Ní Chiaráin, Christer Gobl, Ailbhe Ní Chasaide |
INTERSPEECH | 4 |
| 2022 | Contribution of the glottal flow residual in affect-related voice transformationabstractThis paper explores the contribution of the glottal flow residual in affect-related voice transformation.This signal, which is defined as the difference between the output of the inverse filter estimating the glottal flow signal and the modelled source signal, was analysed using multiple regression analysis.Results show that the strength of the residual varies as a function of the source parameters and this variation is frequency dependent: low frequency energy in the residual is mainly determined by the glottal excitation strength, whereas mid to high frequencies are more influenced by the glottal pulse shape.A method for modelling the residual is presented, which enables modifications based on the changes in source parameters used for voice transformation.This method makes it possible to use the residual as part of the voice source signal when transforming the voice quality in expressive speech synthesis.The result of a listening test, involving the transformation of a neutral voice to an angry or a sad voice, shows that including the glottal flow residual can improve the perceived naturalness of the synthesis.However, the fact that the transformed utterances are still relatively degraded indicates that other factors also need to be considered. Christer Gobl |
INTERSPEECH | 2 |
| 2021 | The LF Model in the Frequency Domain for Glottal Airflow Modelling Without Aliasing DistortionabstractMany of the commonly used voice source models are based on piecewise elementary functions defined in the time domain. \nThe discrete-time implementation of such models generally causes aliasing distortion, which make them less useful for certain applications. This paper presents a method which eliminates this distortion. The key component of the proposed method is the frequency domain description of the source model. By deploying the Laplace transform and phasor arithmetic, closed-form expressions of the source model spectrum can be derived. This facilitates the calculation of the spectrum directly from the model parameters, which in turn makes it possible to obtain the ideal discrete spectrum of the model given the sampling frequency used. This discrete spectrum is entirely free of aliasing distortion, and the inverse discrete Fourier transform is used to compute the sampled glottal flow pulse. The proposed method was applied to the widely used LF model, and the complete Laplace transform of the model is \npresented. Also included are closed-form expressions of the amplitude spectrum and the phase spectrum for the calculation of the LF model spectrum. Christer Gobl |
Interspeech | 1 |
| 2019 | Time to Frequency Domain Mapping of the Voice Source: The Influence of Open Quotient and Glottal Skew on the Low End of the Source Spectrum
Christer Gobl, Ailbhe Ní Chasaide |
INTERSPEECH | 1 |
| 2019 | The Role of Voice Quality in the Perception of Prominence in Synthetic SpeechabstractThis paper explores how prominence can be modelled in speech synthesis through voice quality variation. Synthetic \nutterances varying in voice quality (breathy, modal, tense) were generated using a glottal source model where the global \nwaveshape parameter Rd was the main control parameter and f0 was not varied. A manipulation task perception experiment was conducted to establish perceptually salient Rd values in the signalling of focus. The participants were presented with mini-dialogues designed to elicit narrow focus (with different focal syllable locations) and were asked to manipulate an unknown parameter in the synthetic utterances to produce a natural response. The results showed that participants manipulated Rd not only in focal syllables, but also in the pre- and postfocal material. The direction of Rd manipulation in the focal syllables was the same across the three voice qualities – towards decreased Rd values (tenser phonation). The magnitude of the decrease in Rd was significantly less for tense voice compared to breathy and modal voice, but did not vary with the location of the focal syllable in the utterance. Overall, the results suggest that Rd is effective as a control parameter for modelling prominence in synthetic speech. Andy Murphy, Irena Yanushevskaya, Ailbhe Ní Chasaide, Christer Gobl |
INTERSPEECH | 4 |
| 2018 | On the Relationship between Glottal Pulse Shape and Its Spectrum: Correlations of Open Quotient, Pulse Skew and Peak Flow with Source Harmonic Amplitudes
Christer Gobl, Andy Murphy, Irena Yanushevskaya, Ailbhe Ní Chasaide |
INTERSPEECH | 1 |
| 2018 | Voice Source Contribution to Prominence Perception: Rd Implementation
Andy Murphy, Irena Yanushevskaya, Ailbhe Ní Chasaide, Christer Gobl |
INTERSPEECH | 4 |
| 2017 | The ABAIR Initiative: Bringing Spoken Irish into the Digital SpaceabstractThe processes of language demise take hold when a language ceases to belong to the mainstream of life’s activities. Digital communication technology increasingly pervades all aspects of modern life. Languages not digitally ‘available’ are ever more marginalised, whereas a digital presence often yields unexpected opportunities to integrate the language into the mainstream. The ABAIR initiative embraces three central aspects of speech technology development for Irish (Gaelic): the provision of technology-oriented linguistic-phonetic resources; the building and perfecting of core speech technologies; and the development of technology applications, which exploit both the technologies and the linguistic resources. The latter enable the public, learners, and those with disabilities to integrate Irish into their day-to-day usage. This paper outlines some of the specific linguistic and sociolinguistic challenges and the approaches adopted to address them. Although machine-learning approaches are helping to speed up the process of technology provision, the ABAIR experience highlights how phonetic-linguistic resources are also crucial to the development process. For the endangered language, linguistic resources are central to many applications that impact on language usage. The sociolinguistic context and the needs of potential end users should be central considerations in setting research priorities and deciding on methods. Ailbhe Ní Chasaide, Neasa Ní Chiaráin, Christoph Wendler, Harald Berthelsen, Andy Murphy, Christer Gobl |
INTERSPEECH | 6 |
| 2017 | Voice-to-Affect Mapping: Inferences on Language Voice Baseline SettingsabstractModulations of the voice convey affect, and the precise mapping of voice-to-affect may vary for different languages. However, affect-related modulations occur relative to the baseline affect-neutral voice, which tends to differ from language to language. Little is known about the characteristic long-term voice settings for different languages, and how they influence the use of voice quality to signal affect. In this paper, data from a voice-to-affect perception test involving Russian, English, Spanish and Japanese subjects is re-examined to glean insights concerning likely baseline settings in these languages. The test used synthetic stimuli with different voice qualities (modelled on a male voice), with or without extreme f0 contours as might be associated with affect. Cross-language differences in affect ratings for modal and tense voice suggest that the baseline in Spanish and Japanese is inherently tenser than in Russian and English, and that as a corollary, tense voice serves as a more potent cue to high-activation affects in the latter languages. A relatively tenser baseline in Japanese and Spanish is further suggested by the fact that tense voice can be associated with intimate, a low activation state, just as readily as with the high-activation state interested. Ailbhe Ní Chasaide, Irena Yanushevskaya, Christer Gobl |
INTERSPEECH | 3 |
| 2017 | Reshaping the Transformed LF Model: Generating the Glottal Source from the Waveshape Parameter RdabstractPrecise specification of the voice source would facilitate better modelling of expressive nuances in human spoken interaction. This paper focuses on the transformed version of the widely used LF voice source model, and proposes an algorithm which makes it possible to use the wave shape parameter Rd to directly control the LF pulse, for more effective analysis and synthesis of voice modulations. The Rd parameter, capturing much of the natural covariation between glottal parameters, is central to the transformed LF model. It is used to predict the standard R-parameters, which in turn are used to synthesise the LF waveform. However, the LF pulse that results from these predictions may have an Rd value noticeably different from the specified Rd, yielding undesirable artefacts, particularly when the model is used for detailed analysis and syn-thesis of non-modal voice. A further limitation is that only a subset of possible Rd values can be used, to avoid conflicting LF parameter settings. To eliminate these problems, a new iterative algorithm was developed based on the Newton-Raphson method for two variables, but modified to include constraints. This ensures that the correct Rd is always obtained and that the algorithm converges for effectively all permissible Rd values. Christer Gobl |
INTERSPEECH | 1 |
| 2017 | Rd as a Control Parameter to Explore Affective Correlates of the Tense-Lax ContinuumabstractThis study uses the Rd glottal waveshape parameter to simulate the phonatory tense-lax continuum and to explore its affective correlates in terms of activation and valence. Based on a natural utterance which was inverse filtered and source-parameterised, a range of synthesized stimuli varying along the tense-lax continuum were generated using Rd as a control parameter. Two additional stimuli were included, which were versions of the most lax stimuli with additional creak (lax-creaky voice). In a listening test, participants chose an emotion from a set of affective labels and indicated its perceived strength. They also indicated the naturalness of the stimulus and their confidence in their judgment. Results showed that stimuli at the tense end of the range were most frequently associated with angry, at the lax end of the range the association was with sad, and in the intermediate range, the association was with content. Results also indicate, as was found in our earlier work, that a particular stimulus can be associated with more than one affect. Overall these results show that Rd can be used as a single control parameter to generate variation along the tense-lax continuum of phonation. Andy Murphy, Irena Yanushevskaya, Ailbhe Ní Chasaide, Christer Gobl |
INTERSPEECH | 4 |
| 2017 | Cross-Speaker Variation in Voice Source Correlates of Focus and DeaccentuationabstractThis paper describes cross-speaker variation in the voice source correlates of focal accentuation and deaccentuation. A set of utterances with varied narrow focus placement as well as broad focus and deaccented renditions were produced by six speakers of English. These were manually inverse filtered and parameterized on a pulse-by-pulse basis using the LF source model. Z-normalized F0, EE, OQ and RD parameters (selected through correlation and factor analysis) were used to generate speaker specific baseline voice profiles and to explore cross-speaker variation in focal and non-focal (post- and prefocal) syllables. As expected, source parameter values were found to differ in the focal and postfocal portions of the utterance. For four of the six speakers the measures revealed a trend of tenser phonation on the focal syllable (an increase in EE and F0 and typically, a decrease in OQ and RD) as well as increased laxness in the postfocal part of the utterance. For two of the speakers, however, the measurements showed a different trend. These speakers had very high F0 and often high EE on the focal accent. In these cases, RD and OQ values tended to be raised rather than lowered. The possible reasons for these differences are discussed. Irena Yanushevskaya, Ailbhe Ní Chasaide, Christer Gobl |
INTERSPEECH | 3 |
| 2016 | Perceptual Salience of Voice Source Parameters in Signaling Focal Prominenceabstractpaper describes listening tests investigating the perceptual role of voice source parameters (other than F0) in signaling focal prominence. Synthesized stimuli were constructed on the basis of an inverse filtered utterance ‘We were away a year ago’. Voice source parameters were manipulated in the two potentially accentable syllables WAY and YEAR (in terms of the absolute magnitude and alignment of peaks) and to provide source deaccentuation of post-focal material. Participants in the first listening test were asked to decide whether the syllable WAY, YEAR or neither was deemed the most prominent: judgments on the degree of prominence and naturalness were also indicated on a continuous visual analogue scale. In the second test listeners indicated the degree of prominence for every syllable in the phrase. For WAY, voice source manipulations can cue focal accentuation, and both the magnitude of the source manipulation of the syllable and the presence of source deaccentuation contribute to the effect. However, for YEAR, listeners’ perception of focal accentuation tended to show relatively minor increases in perceived prominence regardless of the source manipulations involved. It therefore appears that the source expression of focus is sensitive to the location of focus in the intonational phrase. Irena Yanushevskaya, Andy Murphy, Christer Gobl, Ailbhe Ní Chasaide |
INTERSPEECH | 3 |
| 2015 | The relationship between voice source parameters and the maxima dispersion quotient (MDQ)abstractThis study examines the relationship between the Maxima Dispersion Quotient (MDQ), a recently proposed measure of the tense-lax dimension of voice quality, and voice source parameters manually measured from glottal flow data, where there is linguistically and paralinguistically determined voice source modulation.MDQ was found to correlate most closely to the open quotient (OQ) and the RD parameter.The paralinguistically varying data, which involved more extensive voice source modulation than the linguistic, also showed a higher degree of correlation with these parameters, and higher correlations overall with a range of voice source parameters.The high correlation with OQ and RD found in these analyses would suggest that MDQ can be a useful, additional parameter for the analysis of glottal source dynamics. Christer Gobl, Irena Yanushevskaya, Ailbhe Ní Chasaide |
INTERSPEECH | 1 |
| 2015 | Pitch declination and reset as a function of utterance duration in conversational speech dataabstractThis paper describes the declination trends of f 0 in conversational speech data.A 10-minute dialogue interaction from a corpus of spontaneous speech was annotated to identify intersilence units (ISU) and turns.Detailed annotation of the ISUs was conducted in terms of communicative types and pitch patterns.f 0 declination was measured by (1) fitting a regression line to f 0 trajectories and (2) by fitting additional regression lines to the data points below and above the original (central) regression line.The slope of declination as well as the height of ISU/turn-initial f 0 peak were examined as a function of the duration of the ISU or turn.The results suggest that declination is indeed present in conversational speech data, at the level of both the ISU and the turn (73% of the analysed ISUs exhibited negative f 0 declination slope).There is a tendency for the steepness of the slope to decrease and the height of ISturn-initial f 0 peak to increase as the duration of the ISU or turn increases.The results are discussed in the context of Projection and Reaction theories and of Hard vs. Soft preplanning of speech production.The findings are of potential interest for the development of human-machine dialogue systems. Céline De Looze, Irena Yanushevskaya, Andy Murphy, Eoghan O'Connor, Christer Gobl |
INTERSPEECH | 5 |
| 2014 | Data-driven detection and analysis of the patterns of creaky voice
Thomas Drugman, John Kane 0002, Christer Gobl |
Comput. Speech Lang. | 3 |
| 2014 | Phonetic feature extraction for context-sensitive glottal source processing
John Kane 0002, Matthew P. Aylett, Irena Yanushevskaya, Christer Gobl |
Speech Commun. | 4 |
| 2013 | Prediction of creaky voice from contextual factorsabstractCreaky voice, also referred to as vocal fry, is a voice quality frequently produced in many languages, in both read and conversational speech. In order to enhance the naturalness of speech synthesisers, these latter should be able to generate speech in all its expressive diversity. This includes a proper use of creaky voice. The goal of this paper is two-fold. Firstly we analyse how contextual factors can be informative for the prediction of creaky use. It is observed that a few contextual factors related to speech production preceding a silence or a pause are of particular interest. This study validates that creaky voice plays a crucial syntactic role, allowing for a better structuring of phrases. In a second experiment, we investigate the prediction of creakiness from contextual factors based on HMMs. Four methods are compared on a US English and a Finnish speaker. It is shown that the best prediction technique achieves a promising performance comparable to what is carried out with the creaky detection algorithm on which HMMs were trained. Thomas Drugman, John Kane 0002, Tuomo Raitio, Christer Gobl |
ICASSP | 4 |
| 2013 | Speaker and language independent voice quality classification applied to unlabelled corpora of expressive speechabstractVoice quality plays a pivotal role in speech style variation. Therefore, control and analysis of voice quality is critical for many areas of speech technology. Until now, most work has focused on small purpose built corpora. In this paper we apply state-of-the-art voice quality analysis to large speech corpora built for expressive speech synthesis. A fuzzy-input fuzzy-output support vector machine classifier is trained and validated using features extracted from these corpora. We then apply this classifier to freely available audiobook data and demonstrate a clustering of the voice qualities that approximates the performance of human perceptual ratings. The ability to detect voice quality variation in these widely available unlabelled audiobook corpora means that the proposed method may be used as a valuable resource in expressive speech synthesis. John Kane 0002, Stefan Scherer, Matthew P. Aylett, Louis-Philippe Morency, Christer Gobl |
ICASSP | 5 |
| 2013 | The voice prominence hypothesis: the interplay of F0 and voice source features in accentuationabstractThis paper explores the interplay of source correlates of accentuation, examining a hypothesis (the Voice Prominence Hypothesis) that different source parameters are involved and may serve as equivalent. It predicts that where accentuation is not marked by pitch salience there will be more extensive changes in other source parameters. This follows our assumption that prosodic entities such as accentuation, focus, declination, etc. involve adjustments to the entire voice source and not simply to F0. Twelve 3-accent sentences of Connemara Irish (declaratives, WH questions and Yes/No questions) were analysed. These are typically produced and transcribed as H* H* H*L. Of particular interest were the second accents: although they are heard as accented, there are no particular pitch excursions that would account for their salience. Inverse filtering and subsequent source parameterisation was carried out to yield measures for a range of source parameters. Results support the voice prominence hypothesis: as predicted, the most striking source adjustments were found in the second accent. Even where there is substantial pitch movement (final accent), parameters other than F0 appear to be contributing to the salience of the accented syllable. The precise source changes associated with accentuation varied across sentence types and within the prosodic phrase. Index Terms: voice source, accentuation, prominence. Ailbhe Ní Chasaide, Irena Yanushevskaya, John Kane 0002, Christer Gobl |
INTERSPEECH | 4 |
| 2013 | A comparative study of glottal open quotient estimation techniquesabstractThe robust and efficient extraction of features related to the glottal excitation source has become increasingly important for speech technology. The glottal open quotient (OQ) is one rel-evant measurement which is known to significantly vary with changes in voice quality on a breathy to tense continuum. The extraction of OQ, however, is hampered in the time-domain by the difficulty in consistently locating the point of glottal open-ing as well the computational load of its measurement. De-termining OQ correlates in the frequency domain is an attrac-tive alternative, however the lower frequencies of glottal source spectrum are also affected by other aspects of the glottal pulse shape thereby precluding closed-form solutions and straightfor-ward mappings. The present study provides a comparison of three OQ estimation methods and shows a new method based on spectral features and artificial neural networks to outperform existing methods in terms of discrimination of voice quality, lower error values on a large volume of speech data and dra-matically reduced computation time. John Kane 0002, Stefan Scherer, Louis-Philippe Morency, Christer Gobl |
INTERSPEECH | 4 |
| 2013 | Using phonetic feature extraction to determine optimal speech regions for maximising the effectiveness of glottal source analysisabstractParameterisation of the glottal source has become increasingly useful for speech technology. For many applications it may \nbe desirable to restrict the glottal source feature data to only speech regions where it can be reliably extracted. In this paper we exploit the previously proposed set of binary phonetic feature extractors to help determine optimal regions for glottal source analysis. Besides validation of the phonetic feature extractors, we also quantitatively assess their usefulness for improving voice quality classification and find highly significant reductions in error rates in particular when nasals and fricative regions are excluded. John Kane 0002, Irena Yanushevskaya, John Dalton, Christer Gobl, Ailbhe Ní Chasaide |
INTERSPEECH | 4 |
| 2013 | HMM-based synthesis of creaky voiceabstractCreaky voice, also referred to as vocal fry, is a voice quality frequently produced in many languages, in both read and conversational speech. To enhance the naturalness of speech synthesis, these latter should be able to generate speech in all its expressive diversity, including creaky voice. The present study looks to exploit our recent developments, including creaky voice detection, prediction of creaky voice from context, and rendering of the creaky excitation, into a fully functioning and automatic HMM-based synthesis system. HMM-based synthetic creaky voices are built and evaluated in subjective listening tests, which show that the best synthetic creaky voices are rated more natural and more creaky compared to a conventional voice. A noncreaky voice is also successfully transformed to use creak by modifying the F0 contour and excitation of the predicted creaky parts. The transformed voice is rated equal in terms of naturalness and clearly more creaky compared to the original voice. Index Terms: speech synthesis, creaky voice, contextual factors, F0 estimation, excitation modeling Tuomo Raitio, John Kane 0002, Thomas Drugman, Christer Gobl |
INTERSPEECH | 4 |
| 2013 | Improved automatic detection of creak
John Kane 0002, Thomas Drugman, Christer Gobl |
Comput. Speech Lang. | 3 |
| 2013 | Investigating fuzzy-input fuzzy-output support vector machines for robust voice quality classification
Stefan Scherer, John Kane 0002, Christer Gobl, Friedhelm Schwenker |
Comput. Speech Lang. | 3 |
| 2013 | Evaluation of glottal closure instant detection in a range of voice qualities
John Kane 0002, Christer Gobl |
Speech Commun. | 2 |
| 2013 | Automating manual user strategies for precise voice source analysis
John Kane 0002, Christer Gobl |
Speech Commun. | 2 |
| 2013 | Wavelet Maxima Dispersion for Breathy to Tense Voice DiscriminationabstractThis paper proposes a new parameter, the Maxima Dispersion Quotient (MDQ), for differentiating breathy to tense voice. Maxima derived following wavelet decomposition are often used for detecting edges in image processing, where locations of these maxima organize in the vicinity of the edge location. Similarly for tense voice, which typically displays sharp glottal closing characteristics, maxima following wavelet analysis are organized in the vicinity of the glottal closure instant (GCI). Contrastingly, as the phonation type tends away from tense voice towards a breathier phonation it is observed that the maxima become increasingly dispersed. The MDQ parameter is designed to quantify the extent of this dispersion and is shown to compare favorably to existing voice quality parameters, particularly for the analysis of continuous speech. Also, classification experiments reveal a significant improvement in the detection of the voice qualities when MDQ is included as an input to the classifier. Finally, MDQ is shown to be robust to additive noise down to a Signal-to-Noise Ratio of 10 dB. John Kane 0002, Christer Gobl |
IEEE Trans. Speech Audio Process. | 2 |
| 2012 | Detecting a targeted voice style in an audiobook using voice quality featuresabstractAudiobooks are known to contain a variety of expressive speaking styles that occur as a result of the narrator mimicking a character in a story, or expressing affect. An accurate modeling of this variety is essential for the purposes of speech synthesis from an audiobook. Voice quality differences are important features characterizing these different speaking styles, which are realized on a gradient and are often difficult to predict from the text. The present study uses a parameter characterizing breathy to tense voice qualities using features of the wavelet transform, and a measure for identifying creaky segments in an utterance. Based on these features, a combination of supervised and unsupervised classification is used to detect the regions in an audiobook, where the speaker changes his regular voice quality to a particular voice style. The target voice style candidates are selected based on the agreement of the supervised classifier ensemble output, and evaluated in a listening test. Éva Székely, John Kane 0002, Stefan Scherer, Christer Gobl, Julie Carson-Berndsen |
ICASSP | 4 |
| 2012 | Modeling the Creaky Excitation for Parametric Speech SynthesisabstractIn order to produce natural sounding output, corpus-based speech synthesis systems need to be able to properly model the acoustic variability in the corpus. Creaky voice is a voice quality frequently produced in many languages, in both read and conversational speech settings. However, the creaky excitation displays different acoustic characteristics than modal excitations and is, hence, not suitably modelled by standard vocoders. This study presents an analysis of the creaky excitation which is used to derive an extension of the Deterministic plus Stochastic Model of the residual signal. This proposed model is designed to appropriately model creaky voice and is integrated into a vocoder for parametric speech synthesis. Copy-synthesis versions of short speech segments containing creaky voice were used in a subjective listening test which revealed clearly better rendering of the voice quality than a standard vocoder. Index Terms: Voice quality, speech synthesis, creak, vocal fry 1. Thomas Drugman, John Kane 0002, Christer Gobl |
INTERSPEECH | 3 |
| 2012 | Resonator-based creaky voice detectionabstractCreaky voice is used by speakers for a variety of interactive, expressive and stylistic reasons. As a result the accurate detection of creaky regions in speech can yield important information not captured within the propositional content of spoken utterances. Hence, we describe a new method for automatically detecting creaky regions following the observation that secondary peaks occur in the linear prediction residual signal. The proposed approach was shown to significantly outperform the state-of-theart in an objective evaluation on a range of speech databases. Index Terms: Voice quality, glottal source, creak, vocal fry 1. Thomas Drugman, John Kane 0002, Christer Gobl |
INTERSPEECH | 3 |
| 2011 | Evaluation of Glottal Epoch Detection Algorithms on Different Voice TypesabstractAccording to the source-filter model of speech production, speech can be represented by passing the excitation signal through the vocal tract filter. The epoch or instant of maximum excitation corresponds to the glottal closure instant. Several speech processing applications require robust epoch detection but this can be a difficult task. Although state-of-the-art epoch estimation methods can produce reliable results, they are generally evaluated using speech recorded with a neutral voice quality (modal voice). This paper reviews and evaluates six popular algorithms for the calculation of glottal closure instants on speech spoken with modal voice and seven additional voice qualities. Results show that the performance of each method is affected by the voice type and that some methods perform better than others for each voice quality. Index Terms: GCI, epoch detection, glottal source João P. Cabral, John Kane 0002, Christer Gobl, Julie Carson-Berndsen |
INTERSPEECH | 3 |
| 2011 | Identifying Regions of Non-Modal Phonation Using Features of the Wavelet TransformabstractThe present study proposes a new parameter for identifying breathy to tense voice qualities in a given speech segment using measurements from the wavelet transform. Techniques that can deliver robust information on the voice quality of a speech segment are desirable as they can help tune analysis strategies as well as provide automatic voice quality annotation in large corpora. The method described here involves wavelet-based decomposition of the speech signal into octave bands and then fitting a regression line to the maximum amplitudes at the different scales. The slope coefficient is then evaluated in terms of its ability to differentiate voice qualities compared to other parameters in the literature. The new parameter (named here Peak Slope) was shown to have robustness to babble noise added with signal to noise ratios as low as 10 dB. Furthermore, the proposed parameter was shown to provide better differentiation of breathy to tense voice qualities in both vowels and running speech. Index Terms: Voice quality, glottal source, wavelets 1. John Kane 0002, Christer Gobl |
INTERSPEECH | 2 |
| 2010 | A spectral LF model based approach to voice source parameterisationabstractThis paper presents a new method of extracting LF model based parameters using a spectral model matching approach. Strategies are described for overcoming some of the known difficulties of this type of approach, in particular high frequency noise. The new method performed well compared to a typical time based method particularly in terms of robustness against distortions introduced by the recording system and in terms of the ability of parameters extracted in this manner to differentiate three discrete voice qualities. Results from this study are very promising for the new method and offer a way of extracting a set of non-redundant spectral parameters that may be very useful in both recognition and synthesis systems. Index Terms: LF model, voice source, parameterisation, robustness, classification. John Kane 0002, Mark Kane, Christer Gobl |
INTERSPEECH | 3 |
| 2010 | An exploration of voice source correlates of focusabstractThis pilot study explores how the voice source parameters vary in focally accented syllables. It examines the dynamics of the voice source parameters in an all-voiced short declarative utterance in which the focus placement was varied. The voice Irena Yanushevskaya, Christer Gobl, John Kane 0002, Ailbhe Ní Chasaide |
INTERSPEECH | 2 |
| 2009 | Perceived loudness and voice quality in affect cueingabstractThe paper describes an auditory experiment aimed at testing whether the intrinsic loudness of a stimulus with a given voice quality influences the way in which it signals affect. Synthesised voice quality stimuli in which intrinsic loudness \nwas systematically manipulated were presented to listeners to test the effect of this manipulation on the affective colouring \nof the stimuli. The results showed that even when devoid of intrinsic loudness variation, non-modal voice quality stimuli \nwere capable of communicating affect. However, changing the loudness of a non-modal voice quality stimulus towards its \nintrinsic loudness resulted in the increase of affective ratings. Irena Yanushevskaya, Christer Gobl, Ailbhe Ní Chasaide |
INTERSPEECH | 2 |
| 2008 | Cross-dialect Irish prosody: linguistic constraints on Fujisaki modellingabstractWe describe here our approach to quantifying cross-dialect\ndifferences in Irish Gaelic, using the Fujisaki model. The\nbasic principle is that the way in which the modelling is\ncarried out respects a parallel linguistic (AM) analysis. The\naims are: (1) to ensure that our modelling strategies permit a\nreliable cross-dialect comparison, (2) that the model-derived\nmeasurements can be related to meaningful linguistic\ndimensions and (3) that the analysis forms the basis for multidialect\nsynthesis. Maria O'Reilly, Ailbhe Ní Chasaide, Christer Gobl |
INTERSPEECH | 3 |
| 2008 | Cross-language study of vocal correlates of affective states
Irena Yanushevskaya, Ailbhe Ní Chasaide, Christer Gobl |
INTERSPEECH | 3 |
| 2007 | Time- and Amplitude-Based Voice Source Correlates of Emotional Portrayals
Irena Yanushevskaya, Michelle Tooher, Christer Gobl, Ailbhe Ní Chasaide |
ACII | 3 |
| 2006 | Speech technology for minority languages: the case of Irish (gaelic)abstractAbstract?Unit selection is a data-driven approach to speech\nsynthesis that concatenates pieces of recorded speech from a\nlarge database in order to create novel sentences. Many corpora\nare available in the English language, including the Arctic\ndatabase [1], which allows a user to create small, reliable speech\nsynthesisers using only a small set of recorded sentences. Such\nresources for minority languages are scarce however, despite their\nincreasing importance for the survival of such languages. This\npaper describes the current research in creating efficient Irish\nlanguage corpora for speech synthesis. Corpus design techniques\nare discussed, in particular, two methods of data reduction that\nare applied to an aligned spoken corpus of Irish in order to\ncreate smaller, more efficient speech corpora. Ailbhe Ní Chasaide, John Wogan, Brian Ó Raghallaigh, Áine Ní Bhriain, Eric Zoerner, Harald Berthelsen, Christer Gobl |
INTERSPEECH | 7 |
| 2006 | Modelling aspiration noise during phonation using the LF voice source model
Christer Gobl |
INTERSPEECH | 1 |
| 2005 | Voice quality and f0 cues for affect expression: implications for synthesisabstractSynthesised stimuli were used to investigate how two notionally\nseparable dimensions of tone-of-voice ? voice quality and\nfundamental frequency ? are involved in the expression of\naffect. Listeners were presented with three series of stimuli:\n(1) stimuli exemplifying different voice qualities, (2) stimuli\nall with modal voice quality but with different affect-related f0\ncontours, and (3) stimuli incorporating variation in both voice\nquality and affect-related f0 contours. A total of 15 stimuli\nwere rated for 12 different affective attributes. Voice quality\ndifferentiation appears to account for the highest affect ratings\noverall, as indicated by the scores obtained for stimuli series\n(1) and (3). The relatively weaker affect signalling of stimuli\ndifferentiated by f0 alone corroborates findings in [2]. It also\nsuggests that for the generation of expressive, affectively\ncoloured speech synthesis, it is not sufficient to manipulate\nonly f0; we also need to capture the voice quality dimension\nof the voice source. Irena Yanushevskaya, Christer Gobl, Ailbhe Ní Chasaide |
INTERSPEECH | 2 |
| 2004 | Decomposing linguistic and affective components of phonatory qualityabstractThis paper is concerned with the role of phonatory quality in signalling affect. An overview of perception experiments is presented, which used synthetic stimuli with different phonatory qualities and f 0 contours in order to explore the mapping of voice quality to affect as well as the way in which voice quality combines with f 0. Results highlight the need for these phonetic parameters to be considered together. To identify the phonatory correlates of affect, we also need to understand the substrate of voice source variation, due to the linguistic content of utterances (prosodic and segmental) as well as to speaker specific characteristics. Illustrations of the former type of variation are presented, based on source parameterisation of inverse filtered data. A holistic analytic approach is advocated, which incorporates the main phonetic dimensions (voice quality, f 0 and temporal parameters) and which integrates the affective dimension with the more linguistic dimension of prosody. Ailbhe Ní Chasaide, Christer Gobl |
INTERSPEECH | 2 |
| 2003 | The role of voice quality in communicating emotion, mood and attitude
Christer Gobl, Ailbhe Ní Chasaide |
Speech Commun. | 1 |
| 1992 | Acoustic characteristics of voice quality
Christer Gobl, Ailbhe Ní Chasaide |
Speech Commun. | 1 |
| 1990 | Linguistic and paralinguistic variation in the voice source
Ailbhe Ní Chasaide, Christer Gobl |
ICSLP | 2 |
| 1989 | Voice source rules for text-to-speech synthesisabstractVoice source parameters are generally derived from time-domain inverse filtering. An effective system of parameterization is the LF model, which has been adopted for the authors' studies of voice source variations in connected speech. The effects of variations of the different LF parameters on the voice source spectrum are demonstrated. Dynamic variations within a linguistic frame and with respect to different speaker types are exemplified. The LF-model has been implemented in a text-to-speech system, and work on development of rules is progressing.> Rolf Carlson, Gunnar Fant, Christer Gobl, Björn Granström, Inger Karlsson, Qiguang Lin |
ICASSP | 3 |