Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Jan P. H. van Santen

dblp:15/6483 · DBLP profile ↗
← Back
69ranked-venue papers
15as first author
0since 2021 · last 2019
0000-0003-3245-5301ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 58 · 13 first-authorArtificial intelligence and machine learning · 55 · 12 first-authorTheory of computation · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
1 paper
Audio and music processing · 100%
Artificial intelligence
2 papers
Speech recognition and synthesis · 82% Efficient and distributed learning · 18%

Topics — the 6 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Audio and music processing › speech synthesis
concatenative speech synthesis
0.112007
The Contribution of Various Sources of Spectral Mismatch to Audible Discontinuities in a Diphone Database · IEEE Trans. Speech Audio Process. 2007
Audio and music processing
speech synthesis
0.112007
The Contribution of Various Sources of Spectral Mismatch to Audible Discontinuities in a Diphone Database · IEEE Trans. Speech Audio Process. 2007
Natural language and speech › Speech recognition and synthesis › speech analysis
formant tracking
0.112005
Formant Tracking Using Context-Dependent Phonemic Information · IEEE Trans. Speech Audio Process. 2005
Natural language and speech › Speech recognition and synthesis
automatic speech recognition
0.012005
Formant Tracking Using Context-Dependent Phonemic Information · IEEE Trans. Speech Audio Process. 2005
Machine learning › Efficient and distributed learning
data selection
0.011996
Selecting Training Inputs via Greedy Rank Covering · SODA 1996
Graph algorithms and graph theory
covering algorithm
0.011996
Selecting Training Inputs via Greedy Rank Covering · SODA 1996

Methods — techniques the papers use, named apart from their topics

pairwise comparison · 0.1formant resynthesis · 0.1viterbi search · 0.1linear prediction · 0.1hidden markov model · 0.1greedy rank covering · 0.0
YearPublicationVenuePosition
2019 Improving ASR Systems for Children with Autism and Language Impairment Using Domain-Focused DNN Transfer Techniques
abstract
grade performs best at a 26.21% WER.
Robert Gale, Liu Chen, Jill Dolata, Jan P. H. van Santen, Meysam Asgari
INTERSPEECH4
2018 Articulation-to-Speech Synthesis Using Articulatory Flesh Point Sensors' Orientation Information
Beiming Cao, Myung Jong Kim, Jun R. Wang, Jan P. H. van Santen, Ted Mau, Jun Wang 0037
INTERSPEECH4
2017 Automatic Scoring of a Nonword Repetition Test
abstract
In this study, we explore the feasibility of speech-based techniques to automatically evaluate a nonword repetition (NWR) test. NWR tests, a useful marker for detecting language impairment, require repetition of pronounceable nonwords, such as "D OY F", presented aurally by an examiner or via a recording. Our proposed method leverages ASR techniques to first transcribe verbal responses. Second, it applies machine learning techniques to ASR output for predicting gold standard scores provided by speech and language pathologists. Our experimental results for a sample of 101 children (42 with autism spectrum disorders, or ASD; 18 with specific language impairment, or SLI; and 41 typically developed, or TD) show that the proposed approach is successful in predicting scores on this test, with averaged product-moment correlations of 0.74 and mean absolute error of 0.06 (on a observed score range from 0.34 to 0.97) between observed and predicted ratings.
Meysam Asgari, Jan P. H. van Santen, Katina Papadakis
ICMLA2
2017 Integrating Articulatory Information in Deep Learning-Based Text-to-Speech Synthesis
Beiming Cao, Myung Jong Kim, Jan P. H. van Santen, Ted Mau, Jun Wang 0037
INTERSPEECH3
2015 Speaker intonation adaptation for transforming text-to-speech synthesis speaker identity
abstract
In this study, we propose a new intonation adaptation method to transform the perceived identity of a Text-To-Speech system to that of a target speaker with a small amount of training data. In the proposed method, during training we fit parametrized accent and phrase curves to parallel recordings of the target speaker F0 curves, and estimate the parameters of a mapping between the corresponding parameter spaces. During test, we fit the accent and phrase curves to the source utterances, apply the mapping, and create an F0 contour from the mapped accent and phrase curves. We compare the proposed method with a baseline adaptation method in which the source F0 contour is transformed linearly such that the per-utterance mean and variance of the target F0 contour is left unaltered. Perceptual tests showed that the proposed method was better than the baseline method in two subjective tests that assess similarity to the target speaker and speech quality, respectively.
Mahsa Sadat Elyasi Langarani, Jan P. H. van Santen
ASRU2
2015 Data-driven foot-based intonation generator for text-to-speech synthesis
abstract
We propose a method for generating F0 contours for text-tospeech synthesis. Training speech is automatically annotated in terms of feet, with features indicating start and end times of syllables, foot position, and foot length. During training, we fit a foot-based superpositional intonation model comprising accent curves and phrase curves. During synthesis, the method searches for stored, fitted accent curves associated with feet that optimally match to-be-synthesized feet in the feature space, while minimizing differences between successive accent curve heights. We tested the proposed method against the HMMbased Speech Synthesis System (HTS) by imposing contours generated by these two methods onto natural speech, and obtaining quality ratings. Test sets varied in how well they were covered by the training data. Contours generated by the proposed method were preferred over HTS-generated contours, especially for poorly-covered test items. To test the new method’s usefulness for processing marked-up text input, we compared its ability to convey contrastive stress with that of natural speech recordings, and found no difference. We conclude that the new method holds promise for generating comparatively highquality F0 contours, especially when training data are sparse and when mark-up is required.
Mahsa Sadat Elyasi Langarani, Jan P. H. van Santen, Seyed Hamidreza Mohammadi, Alexander Kain
INTERSPEECH2
2014 Automatic measurement of affective valence and arousal in speech
abstract
Methods are proposed for measuring affective valence and arousal in speech. The methods apply support vector regression to prosodic and text features to predict human valence and arousal ratings of three stimulus types: speech, delexicalized speech, and text transcripts. Text features are extracted from transcripts via a lookup table listing per-word valence and arousal values and computing per-utterance statistics from the per-word values. Prediction of arousal ratings of delexicalized speech and of speech from prosodic features was successful, with accuracy levels not far from limits set by the reliability of the human ratings. Prediction of valence for these stimulus types as well as prediction of both dimensions for text stimuli proved more difficult, even though the corresponding human ratings were as reliable. Text based features did add, however, to the accuracy of prediction of valence for speech stimuli. We conclude that arousal of speech can be measured reliably, but not valence, and that improving the latter requires better lexical features.
Meysam Asgari, Géza Kiss, Jan P. H. van Santen, Izhak Shafran, Xubo Song
ICASSP3
2014 A novel pitch decomposition method for the generalized linear alignment model
abstract
Superpositional models of intonation typically propose decomposing fundamental frequency (F0) contours into phrase curves and accent curves, aligned with phrases and left-headed feet, respectively. Extracting these component curves from F0contours without making undue assumptions is challenging. We propose a novel method for decomposing pitch curves, based on the assumption that accent curves can be described by combining skewed normal distributions and sigmoid functions. In contrast to an earlier pitch decomposition algorithm (“PRISM”), this allows for simple joint optimization of phrase and accent curve parameters, using fewer parameters. The proposed method was evaluated on three speech corpora containing: (1) synthetically generated pitch curves, (2) all-sonorant utterances, and (3) utterances containing both sonorant and non-sonorant speech sounds. The root weighted mean squared error is small, and, on the corpus for which comparable data are available, is significantly smaller than for PRISM.
Mahsa Sadat Elyasi Langarani, Esther Klabbers, Jan P. H. van Santen
ICASSP3
2014 Modeling fundamental frequency dynamics in hypokinetic dysarthria
abstract
Hypokinetic dysarthria (Hd), which often accompanies Parkinson's Disease (PD), is characterized by hypernasality and by compromised phonation, prosody, and articulation. This paper proposes automated methods for detection of Hd. Whereas most such studies focus on measures of phonation, this paper focuses on prosody, specifically on fundamental frequency (F0) dynamics. Prosody in Hd is clinically described as involving monopitch, which has been confirmed in numerous studies reporting reduced within-utterance pitch variability. We show that a new measure of F0 dynamics, based on a superpositional pitch model that decomposes the F0 contour into a declining phrase curve and (generally, single-peaked) accent curves, performs more accurate Hd vs. Control classification than simpler versions of the model or than conventional variability statistics.
Mahsa Sadat Elyasi Langarani, Jan P. H. van Santen
SLT2
2014 Computational analysis of trajectories of linguistic development in autism
abstract
Deficits in semantic and pragmatic expression are among the hallmark linguistic features of autism. Recent work in deriving computational correlates of clinical spoken language measures has demonstrated the utility of automated linguistic analysis for characterizing the language of children with autism. Most of this research, however, has focused either on young children still acquiring language or on small populations covering a wide age range. In this paper, we extract numerous linguistic features from narratives produced by two groups of children with and without autism from two narrow age ranges. We find that although many differences between diagnostic groups remain constant with age, certain pragmatic measures, particularly the ability to remain on topic and avoid digressions, seem to improve. These results confirm findings reported in the psychology literature while underscoring the need for careful consideration of the age range of the population under investigation when performing clinically oriented computational analysis of spoken language.
Emily Tucker Prud'hommeaux, Eric Morley, Masoud Rouhizadeh, Laura Silverman, Jan P. H. van Santen, Brian Roark, Richard Sproat, Sarah Kauper, Rachel DeLaHunta
SLT5
2013 Estimating speaker-specific intonation patterns using the linear alignment model
Géza Kiss, Jan P. H. van Santen
INTERSPEECH2
2013 Distributional semantic models for the evaluation of disordered language
Masoud Rouhizadeh, Emily Tucker Prud'hommeaux, Brian Roark, Jan P. H. van Santen
HLT-NAACL4
2012 Quantitative Analysis of Pitch in Speech of Children with Neurodevelopmental Disorders
abstract
We analyzed the prosody of children with Autism Spectrum Disorder, Developmental Language Disorder, and typical development in conversational speech, using the CSLU ADOS speech corpus. We found several significant differences in the pitch characteristics of these diagnostic groups, and report automatic classification utilizing these features that are well above chance level. We show that the choice of pitch tracker, its parameters, and the pitch correction method can substantially affect the results, thus the scientific relevance of studies on prosody, and may be one of the reasons for conflicting findings.
Géza Kiss, Jan P. H. van Santen, Emily Tucker Prud'hommeaux, Lois M. Black
INTERSPEECH2
2012 Interactions Between Turn-taking Gaps, Disfluencies and Social Obligation
abstract
Speakers strive to minimize inter-turn gaps when engaged in a dialogue. However, little work has addressed what impact this might have on the fluency of the following speech. In this paper we explore whether there are interactions between turn-taking gaps and turn-initial disfluencies and if the social pressure to respond to questions plays a role in that interaction. Our results indicate that child speakers are more likely to become disfluent both after a question and as the gap length increases, and that the two interact to further increase the likelihood. We also compared the speech of children with Typical Development (TD) to those with Autism Spectrum Disorder (ASD) or Developmental Language Disorder (DLD), where we found that those with ASD were less likely to become disfluent after a question. This finding suggests that the trade-off between timing and disfluencies is driven by social obligation, and that speakers are willing to tolerate disfluencies so as to maintain a short delay. Index Terms: turn-taking gaps, disfluencies, social pressure, language impairments
Rebecca Lunsford, Peter A. Heeman, Jan P. H. van Santen
INTERSPEECH3
2012 Making Conversational Vowels More Clear
abstract
Previously, it has been shown that using clear speech short-term spectra improves the intelligibility of conversational speech. In this paper, a speech transformation method is used to map the spectral features of conversational speech to resemble clear speech. A joint-density Gaussian mixture model is used as the mapping function. The transformation is studied in both the formant frequency and the line spectral frequency domains. Listening test results show that in noisier environments, the transformed speech signal improves vowel intelligibility significantly compared to the original conversational speech. There is also an increase in vowel intelligibility in less noisy environments, but the increase is not statistically significant. The significance tests are performed by Tukey comparison and planned one-tail t-test methods. Index Terms: conversational speech, clear speech, speech transformation, intelligibility
Seyed Hamidreza Mohammadi, Alexander Kain, Jan P. H. van Santen
INTERSPEECH3
2012 Synthetic F0 Can Effectively Convey Speaker ID in Delexicalized Speech
abstract
We investigate the extent to which F0 can convey speaker ID in the absence of spectral, segmental, and durational information. We propose two methods of F0 synthesis based on the Linear Alignment Model (LAM) [2]: one parametric, the other corpusbased. Through a perceptual experiment, we show that F0 alone is able to convey information about speaker ID. We find that F0 synthesized with either LAM-based method conveys speaker ID almost as effectively as natural F0. Index Terms: F0, prosody, speech synthesis, speaker identity, recombinant synthesis
Eric Morley, Esther Klabbers, Jan P. H. van Santen, Alexander Kain, Seyed Hamidreza Mohammadi
INTERSPEECH3
2011 F0 range and peak alignment across speakers and emotions
abstract
We present an analysis of F0range and peak alignment in emotional speech from a heterogeneous group of speakers varying in age and gender. Both speaker and emotion had a strong effect on F0range. Despite these large changes in the F0trajectory, peak alignment was remarkably stable. Using the Linear Alignment Model (LAM), we show that the effects on alignment of emotion and speaker differences, al though statistically significant, are small. This stability results in a conclusion that peak alignment, unlike F0range, does not appear to carry much information about speaker identity or emotional state. The LAM is effective in that it explains 42% of the variance in peak location on average, and furthermore it predicts the time of F0peaks with an average RMS error of 12ms.
Eric Morley, Jan P. H. van Santen, Esther Klabbers, Alexander Kain
ICASSP2
2010 Identifying informative features for ERP speller systems based on RSVP paradigm
Tian Lan 0007, Deniz Erdogmus, Lois M. Black, Jan P. H. van Santen
ESANN4
2010 Frequency-domain delexicalization using surrogate vowels
abstract
We propose a delexicalization algorithm that renders the lexical content of an utterance unintelligible, while preserving important acoustic prosodic cues, as well as naturalness and speaker identity. This is achieved by replacing voiced regions by spectral slices from a surrogate vowel, and by averaging the magnitude spectrum during unvoiced regions. Perceptual tests were carried out comparing sentences that were either unprocessed or delexicalized, using a baseline or the proposed method. An intelligibility test resulted in a keyword recall rate of 92% for the unprocessed sentences, and near complete unintelligibility for both delexicalization methods. Affect recognition was at 65% for unprocessed sentences, and 46% and 49% for the baseline and the proposed method, respectively. Preference tests showed that the proposed method preserved drastically more speaker identity, and sounded more natural than the baseline.
Alexander Kain, Jan P. H. van Santen
INTERSPEECH2
2010 Automated vocal emotion recognition using phoneme class specific features
Géza Kiss, Jan P. H. van Santen
INTERSPEECH2
2010 Evaluation of speaker mimic technology for personalizing SGD voices
abstract
In this paper, we demonstrate the use of state-of-the-art speech technology to transform speech from a source speaker to mimic a particular target speaker with the intention of providng personalized voices to users of Speech Generating Devices (SGDs). This speaker mimicry (SM) capability allows us to use highquality acoustic inventories from professional speakers and transform them to a different target speaker using a very limited set of sentences from that speaker. This technology targets future SGD users who still have a limited vocabulary or available previous recordings. The results of a perceptual study show that listeners can identify which SM voices most resemble their respective target voices. 1
Esther Klabbers, Alexander Kain, Jan P. H. van Santen
INTERSPEECH3
2010 Autism and Interactional Aspects of Dialogue
Peter A. Heeman, Rebecca Lunsford, Ethan Selfridge, Lois M. Black, Jan P. H. van Santen
SIGDIAL Conference5
2009 Using speech transformation to increase speech intelligibility for the hearing- and speaking-impaired
abstract
We present two speech transformation approaches designed to increase the intelligibility of speech. The first approach is used in the context of increasing the intelligibility of conversationally spoken speech for hearing-impaired listeners. An initial experiment showed that a relatively simple mapping function can map spectral features of conversationally spoken speech closer to context-equivalent spectral features of clearly spoken speech. The second approach aims to increase the intelligibility of speaking-impaired individuals by the general population. Results of listening tests indicated that although an intelligibility increase was not achieved, listeners preferred the transformed speech of the proposed system over that of an alternative system.
Alexander Kain, Jan P. H. van Santen
ICASSP2
2009 Perceptual cost function for cross-fading based concatenation
abstract
In earlier research, we applied a linear weighted cross-fading function to ensure smooth concatenation. However, this can cause unnaturally shaped spectral trajectories. We propose context-sensitive cross-fading. To train this system, a perceptually validated cost function is needed, which is the focus of this paper. A corpus was designed to generate a variety of formant trajectory shapes. A perceptual experiment was performed and a multiple linear regression model was applied to predict perceptual quality ratings from various distances between cross-faded and natural trajectories. Results show that perceptual quality could be predicted well from the proposed distance measures. Index Terms: perceptual score, formant frequency, crossfading function, concatenation errors
Qi Miao, Alexander Kain, Jan P. H. van Santen
INTERSPEECH3
2009 Integrating phrasing and intonation modelling using syntactic and morphosyntactic information
Francisco Campillo Díaz, Jan P. H. van Santen, Eduardo Rodríguez Banga
Speech Commun.2
2009 Automated assessment of prosody production
Jan P. H. van Santen, Emily Tucker Prud'hommeaux, Lois M. Black
Speech Commun.1
2007 Hybridizing conversational and clear speech
abstract
“Clear ” (CLR) speech is a speaking style that speakers adopt to be understood correctly in a difficult communication environment. Studies have shown that CLR speech, as opposed to “conversational” (CNV) speech, has significantly higher intelligibility in various conditions. While many differences in acoustic features have been identified, it is not known which individual feature or combinations of features cause the higher intelligibility of CLR speech. The objectives of the current study are to examine whether it is possible to improve speech intelligibility by approximating CLR speech features and to determine which acoustic features contribute to intelligibility. Our approach creates speech samples that combine acoustic features of CNV and CLR speech, using a hybridization algorithm. Results with normalhearing listeners showed significant sentence-level intelligibility improvements of 11–23 % over CNV speech when replacing certain acoustic features with those from CLR speech.
Akiko Amano-Kusumoto, Alexander Kain, John-Paul Hosom, Jan P. H. van Santen
INTERSPEECH4
2007 Dual-channel acoustic detection of nasalization states
abstract
Automatic detection of different oral-nasal configurations during speech is useful for understanding normal nasalization and assessing certain speech disorders. We propose an algorithm to extract nasalization features from dual-channel acoustic signals that are acquired by a simple two-microphone setup. The feature is based on a dual-channel acoustic model and the associated analysis method. We successfully test this feature in speaker-dependent and speaker-independent tasks by comparing it with the conventional single-channel MFCC feature. The proposed feature uniformly performs better in both tasks. Index Terms: speech production, nasalization, speech pathology, velopharyngeal function, nasal resonance
Xiaochuan Niu, Jan P. H. van Santen
INTERSPEECH2
2007 Improving the intelligibility of dysarthric speech
Alexander Kain, John-Paul Hosom, Xiaochuan Niu, Jan P. H. van Santen, Melanie Fried-Oken, Janice Staehely
Speech Commun.4
2007 The Contribution of Various Sources of Spectral Mismatch to Audible Discontinuities in a Diphone Database
abstract
One of the major problems in concatenative synthesis is the occurrence of audible discontinuities between two successive concatenative units. Several studies have attempted to discover objective distance measures that predict the audibility of these discontinuities. In this paper, we investigate mid-vowel joins for three vowels with a range of post-vocalic consonant contexts typical for diphone databases. A first perceptual experiment uses a pairwise comparison procedure to find two subsets of unit combinations: Those with versus without audible discontinuities. A second perceptual experiment uses these two subsets in a procedure where formant resynthesis is used to manipulate three sources of discontinuity separately: formant frequencies, formant bandwidths, and overall energy. Results show mismatch in formant frequencies provides the largest contribution to audible discontinuity, followed by mismatch in overall energy
Esther Klabbers, Jan P. H. van Santen, Alexander Kain
IEEE Trans. Speech Audio Process.2
2006 A model for the f0 reset in corpus-based intonation approaches
Francisco Campillo Díaz, Jan P. H. van Santen, Eduardo Rodríguez Banga
INTERSPEECH2
2006 A noninvasive, low-cost device to study the velopharyngeal port during speech and some preliminary results
abstract
It is desirable to monitor the status of the velopharyngeal port during speech. This paper reports on the design and usage of a noninvasive device that measures the static nasal airflow from the nose during speech. The signal can be recorded directly from a generic sound card. Neither does the usage of this device interfere with the articulatory process of speech, nor does it introduce distortions to the simultaneously recorded acoustic signal. We successfully tested the device in an experiment to analyze the velopharyngeal status during normal speech. Index Terms: speech production, nasal airflow, speech pathology, velopharyngeal function
Xiaochuan Niu, Alexander Kain, Jan P. H. van Santen
INTERSPEECH3
2005 Estimation of the acoustic properties of the nasal tract during the production of nasalized vowels
abstract
Accurate estimation of velar movements is useful for automatic speech recognition, speech enhancement, and diagnosis of certain speech disorders. This paper reports on initial results of a project on estimation of velar movements, for two-microphone setups where the microphones are differentially positioned to pick up nasal and oral speech output. Toward this goal, we propose a method that allows detailed estimation of the acoustic properties of the nasal tract in the simplified condition where the two microphone signals exhibit complete source separation. We successfully test the method against synthetic speech, generated by an articulatory synthesizer in which the acoustic properties of the simulated nasal tract are known. 1.
Xiaochuan Niu, Alexander Kain, Jan P. H. van Santen
INTERSPEECH3
2005 Synthesis of prosody using multi-level unit sequences
Jan P. H. van Santen, Alexander Kain, Esther Klabbers, Taniya Mishra
Speech Commun.1
2005 Duration and spectral balance of intervocalic consonants: A case for efficient communication
R. J. J. H. van Son, Jan P. H. van Santen
Speech Commun.2
2005 Formant Tracking Using Context-Dependent Phonemic Information
abstract
A new formant-tracking algorithm using phoneme information is proposed. Conventional formant-tracking algorithms obtain formant tracks by analyzing the acoustic speech signal using continuity constraints without any additional information. The formant-tracking error rate of the conventional methods is reportedly in the range of 10%-20%. In this paper, we show that if text or phoneme transcription of speech utterances is available, the error rate can be significantly reduced. The basic idea behind this approach is that given the phoneme identity, formant-tracking algorithms can have a better clue of where to look for formants. The algorithm consists of three phases: 1) analysis, 2) segmentation and alignment, and 3) formant tracking by the Viterbi searching algorithm. In the analysis phase, formant candidates are obtained for each analysis frame by solving the linear prediction polynomial. In the segmentation and alignment phase, the text corresponding to the input speech utterance is converted into a sequence of phoneme symbols. Then, the phoneme sequence is time aligned with the speech utterance. A hidden Markov model (HMM) based automatic segmentation algorithm is used for forced-time alignment. For each phoneme segment, nominal formant frequencies are assigned at the center of each phoneme segment. Then nominal formant tracks for the entire utterance are obtained by interpolating the nominal formant frequencies. In order to compensate for the coarticulation effect, different interpolation methods are used depending on the phonemic context. The interpolation process makes the formant-tracking algorithm robust to possible segmentation errors made by the HMM-based segmentation algorithm. As a result, the proposed formant-tracking algorithm does not require highly accurate alignment/segmentation. Finally, a set of formants is chosen from the formant candidates in such a way that the resulting formant tracks come close to the nominal formant tracks while satisfying the continuity constraints. The algorithm is tested using natural speech utterances and the performance is compared against formant tracks obtained by the conventional method using continuity constraints only. The new algorithm significantly reduces the formant-tracking error rate (5.03% for male and 3.73% for female) over the conventional formant-tracking algorithm (13.00% for male and 15.82% for female).
Jan P. H. van Santen, Bernd Möbius, Joseph P. Olive
IEEE Trans. Speech Audio Process.2
2004 Including dynamic and phonetic information in voice conversion systems
abstract
Voice Conversion (VC) systems modify a speaker voice (source speaker) to be perceived as if another speaker (target speaker) had uttered it. Previous published VC approaches using Gaussian Mixture Models [1] performs the conversion in a frame-by-frame basis using only spectral information. In this paper, two new approaches are studied in order to extend the GMM-based VC systems. First, dynamic information is used to build the speaker acoustic model. So, the transformation is carried out according to sequences of frames. Then, phonetic information is introduced in the training of the VC system. Objective and perceptual results compare the performance of the proposed systems.
Antonio Bonafonte, Alexander Kain, Jan P. H. van Santen, Helenca Duxans
INTERSPEECH3
2003 Intelligibility of modifications to dysarthric speech
abstract
Dysarthria is a motor speech impairment affecting millions of people. Dysarthric speech can be far less intelligible than that of non-dysarthric speakers, causing significant communication difficulties. The goal of our work is to understand the effect that certain modifications have on the intelligibility of dysarthric speech. These modifications are designed to identify aspects of the speech signal or signal processing that may be especially relevant to the effectiveness of a system that transforms dysarthric speech to improve its intelligibility. A result of this study is that dysarthric speech can, in the best case, be modified only at the short-term spectral level to improve intelligibility from 68% to 87%. A baseline transformation system using standard technology, however, does not show improvement in intelligibility. Prosody also has a significant (p<0.05) effect on intelligibility.
John-Paul Hosom, Alexander Kain, Taniya Mishra, Jan P. H. van Santen, Melanie Fried-Oken, Janice Staehely
ICASSP (1)4
2003 A speech model of acoustic inventories based on asynchronous interpolation
abstract
We propose a speech model that describes acoustic inventories of concatenative synthesizers. The model has the following characteristics: (i) very compact representations and thus high compression ratios are possible, (ii) re-synthezised speech is free of concatenation errors, (iii) the degree of articulation can be controlled explicitly, and (iv) voice transformation is feasible with relatively few additional recordings of a target speaker. The model represents a speech unit as a synthesis of several types of features, each of which has been computed using non-linear, asynchronous interpolation of neighboring basis vectors associated with known phonemic identities. During analysis, basis vectors and transition weights are estimated under a strict diphone assumption using a dynamic time warping approach. During synthesis, the estimated transition weight values are modified to produce changes in duration and articulation effort.
Alexander Kain, Jan P. H. van Santen
INTERSPEECH2
2003 Control and prediction of the impact of pitch modification on synthetic speech quality
abstract
In order to use speech synthesis to generate highly expressive speech convincingly, the problem of poor prosody (both prediction and generation) needs to be overcome. In this paper we will show that with a simple annotation scheme using the notion of foot structure, we can more accurately predict the shape of local pitch contours. The assumption is that with a better selection mechanism we can reduce the amount of pitch modification required, thereby reducing speech degradation. In addition, we present a perceptual experiment that investigates the degradation introduced by pitch modification using the OGIresLPC algorithm. We correlated the weighted perceptual score with different pitch and delta pitch distances. The best combination of distance measures is able to explain 63% of the variance in the perceptual scores. Decreasing the pitch is shown to have a higher impact on perception than increasing the pitch.
Esther Klabbers, Jan P. H. van Santen
INTERSPEECH2
2003 Detection of list-type sentences
abstract
In this paper, we explore a text type based scheme of text analysis, through the specific problem of detecting the list text type. This is important because TTS systems that can generate the very distinct F0 contour of lists sound more natural. The presented list detection algorithm uses part-of-speech tags as input, and detects lists by computing the alignment costs of clauses in a sentence. The algorithm detects lists with 80 % accuracy. 1.
Taniya Mishra, Esther Klabbers, Jan P. H. van Santen
INTERSPEECH3
2003 Applications of computer generated expressive speech for communication disorders
abstract
This paper focuses on generation of expressive speech, specifically speech displaying vocal affect. Generating speech with vocal affect is important for diagnosis, research, and remediation for children with autism and developmental language disorders. However, because vocal affect involves many acoustic factors working together in complex ways, it is unlikely that we will be able to generate compelling vocal affect with traditional diphone synthesis. Instead, methods are needed that preserve as much of the original signals as possible. We describe an approach to concatenative synthesis that attempts to combine the naturalness of unit selection based synthesis with the ability of diphone based synthesis to handle unrestricted input domains. 1.
Jan P. H. van Santen, Lois M. Black, Gilead Cohen, Alexander Kain, Esther Klabbers, Taniya Mishra, Jacques de Villiers, Xiaochuan Niu
INTERSPEECH1
2001 The ISCA special interest group on speech synthesis
abstract
This paper describes the constitution and activities of the ISCA Speech Synthesis Special Interest Group, SynSIG. It summarises past achievements and suggests ways in which future development could be maintained. The aims of the Special Interest Group on Speech Synthesis are to promote the study and diffusion of knowledge about speech synthesis in general, in a number of ways including: dedicated web pages, a mailing list, a bibliographic database, organisation of workshops on specific themes, exchange of students, and helping to co-ordinate sessions on speech synthesis in international conferences and workshops. The international and multi-disciplinary nature of the SIG also provides a means for diffusing information both to and from the different research communities involved in the synthesis of various languages. 1.
Nick Campbell 0001, Wolfgang Hess, Bernd Möbius, Jan P. H. van Santen
INTERSPEECH4
2001 Obituary: M. W. Macon 1969-2001
Mari Ostendorf, Jan P. H. van Santen, Mark A. Clements
Comput. Speech Lang.2
2000 Predicting segmental durations for Dutch using the sums-of-products approach
abstract
This paper presents the results of a duration study performed for Dutch using the sums-of-products approach [5]. With a relatively small corpus of 297 sentences, a duration model could be constructed with an RMSE of 27 ms, which compares well to similar models for English, French and German. In an evaluation study the predicted durations of the duration model were compared to those predicted by a rule-based duration model. 1.
Esther Klabbers, Jan P. H. van Santen
INTERSPEECH2
2000 When will synthetic speech sound human: role of rules and data
abstract
Text-to-speech synthesis research has moved away from building general purpose systems based on an understanding of human language and speech production towards building systems based on statistical algorithms applied to large text and speech corpora, and, recently, towards building such systems for specific domains. Despite substantial progress, the overall quality of even the best systems is often still inadequate for broad user acceptance in applications that cannot also be handled with simple phrase splicing. This tutorial paper analyzes which problems must be addressed to achieve the goal of generating naturalsounding speech in limited domains in a cost-effective way, and the roles of data and rules as we work towards solutions. 1.
Jan P. H. van Santen, Michael W. Macon, Andrew Cronk, John-Paul Hosom, Alexander Kain, Vincent Pagel, Johan Wouters
INTERSPEECH1
2000 Japanese intonation synthesis using superposition and linear alignment models
Jennifer J. Venditti, Jan P. H. van Santen
INTERSPEECH2
1999 An efficient speaker adaptation method for TTS duration model
abstract
In this paper we present a novel language-independent probabilistic model for automatic grapheme-to-phoneme and phoneme-to-grapheme conversion of words. In a fully unsupervised training procedure, two processes are applied; the transformation rules, which usually fail to provide the correct symbols, are eliminated, and new variable-length string transformation rules are defined improving the string transformation accuracy in the training data. In an iterative process the probabilistic transformation rules are updated in the direction of reducing the error rate of the transformed symbols. Long-term dependencies are defined automatically. Training and testing of the model was carried out on lexicon and natural language corpora of six European Languages. Accurate generalisations have been achieved in all experiments for both transformation directions using a relative small number of defined rules in the training procedure. It is demonstrated that the variable-length probabilistic rules are sufficiently effective for describing bi-directional transcription.
Wentao Gu, Chilin Shih, Jan P. H. van Santen
EUROSPEECH3
1999 Formant tracking using segmental phonemic information
abstract
A new formant tracking algorithm using phoneme dependent nominal formant values is tested. The algorithm consists of three phases: (1) analysis, (2) segmentation, and (3) formant tracking. In the analysis phase, formant candidates are obtained by solving for the roots of the linear prediction polynomial. In the segmentation phase, the input text is converted into a sequence of phonemic symbols. Then the sequence is time aligned with the speech utterance. Finally, a set of formant candidates that are close to the nominal formant estimates while satisfying the continuity constraints are chosen. The new algorithm significantly reduces the formant tracking error rate (3.62%) over a formant tracking algorithmusing only continuity constraints (13.04%). We will also discuss how to further reduce the tracking error rate. INTRODUCTION In the Bell Labs' Text-To-Speech (TTS) system [1], a limited number of acoustic units is stored in the inventory table. Therefore, it is important to be able to...
Jan P. H. van Santen, Bernd Möbius, Joseph P. Olive
EUROSPEECH2
1999 High-accuracy automatic segmentation
abstract
We propose a system for automatically determining boundaries between phonetic segments in a speechwave given a phonetic transcription: automatic segmentation. The system uses edge detectors that are applied to various speech representations; both are optimized for each diphone or diphone class. Output from these detectors, which contains spuriously detected edges, is then combined with alternative pronunciations generated via rules from the canonical pronunciation. The #nal output is generated with lowest-cost path algorithms applied to #nite state transducers. 1. INTRODUCTION Automatic segmentation is critical both for speech research and for speech technologies that rely on segmented speech corpora for training or construction purposes. For example, in text-to-speech synthesis #TTS# segmented corpora are used for the construction of intonation, duration, and synthesis components #4#. The standard approach to automated segmentation is to adapt an automatic speech recognition #ASR#...
Jan P. H. van Santen, Richard Sproat
EUROSPEECH1
1998 Efficient adaptation of TTS duration model to new speakers
abstract
This paper discusses a methodology using a minimal set of sentences to adapt an existing TTS duration model to capture interspeaker variations. The assumption is that the original duration database contains information of both language-specific and speaker-specific duration characteristics. In training a duration model for a new speaker, only the speaker-specific information needs to be modeled, therefore the size of the training data can be reduced drastically. Results from several experiments are compared and discussed.
Chilin Shih, Wentao Gu, Jan P. H. van Santen
ICSLP3
1998 Automatic ambiguity detection
abstract
Most work on sense disambiguation presumes that one knows beforehand -e.g. from a thesaurus -a set of polysemous terms.But published lists invariably give only partial coverage.For example, the English word tan has several obvious senses, but one may overlook the abbreviation for tangent.In this paper, we present an algorithm for identifying interesting polysemous terms and measuring their degree of polysemy, given an unlabeled corpus.The algorithm involves: (i) collecting all terms within a -term window of the target term; (ii) computing the inter-term distances of the contextual terms, and reducing the multi-dimensional distance space to two dimensions using standard methods; (iii) converting the two-dimensional representation into radial coordinates and using isotonic/antitonic regression to compute the degree to which the distribution deviates from a single-peak model.The amount of deviation is the proposed polysemy index.
Richard Sproat, Jan P. H. van Santen
ICSLP2
1998 Modeling vowel duration for Japanese text-to-speech synthesis
Jennifer J. Venditti, Jan P. H. van Santen
ICSLP2
1998 The use of large text corpora for evaluating text-to-speech systems
Louis C. W. Pols, Jan P. H. van Santen, Masanobu Abe, Dan Kahn, Eric Keller
LREC2
1997 The bell labs German text-to-speech system: an overview
abstract
In this paper we present an overview of the German version of the Bell Labs text-to-speech system, a high-quality concatenative synthesis system with extensive text analysis capabilities. We discuss problems of text analysis, and our solutions to these problems, including: the integration of text normalization tasks into linguistic text analysis; the capability to morphologically analyze compounds and unseen words; name analysis and pronunciation. We briefly describe the prosodic components of the text-to-speech system and their underlying duration and intonation models. Finally, the phonetically motivated structure of the acoustic inventory is presented.
Bernd Möbius, Richard Sproat, Jan P. H. van Santen, Joseph P. Olive
EUROSPEECH3
1997 Bell laboratories Russian text-to-speech system
abstract
This paper describes the Bell Labs Russian text-to-speech system, a concatenative system with extensive text-analysis capabilities. The construction of Russian-specific modules will be discussed, including the text-analysis module, the acoustic inventory, the duration module, and the intonation module.
Elena Pavlova, Yuri Pavlov, Richard Sproat, Chilin Shih, Jan P. H. van Santen
EUROSPEECH5
1997 Prosodic modelling in text-to-speech synthesis
Jan P. H. van Santen
EUROSPEECH1
1997 Combinatorial issues in text-to-speech synthesis
abstract
Enhanced storage capacities and new learning algorithms have increased the role of text and speech training data bases in the construction of text-to-speech systems. It has become apparent, however, that not always learning algorithms are available that have strong generalization capabilities – the ability to generalize from cases seen in the training data base to new cases encountered during TTS operation. This makes it important to measure and understand the degree of coverage of the input domain of a text-to-speech system (usually, the entire language) by a given training data base. The goal of this paper is to investigate the feasibility of coverage in several domains of interest for TTS. It is shown that, as a result of the combinatorics of language, coverage is typically quite disappointing. This puts a premium on the generalization capability of learning algorithms.
Jan P. H. van Santen
EUROSPEECH1
1997 Methods for optimal text selection
abstract
Construction of both text-to-speech synthesis (TTS) and automatic speech recognition (ASR) systems involves usage of speech data bases. These data bases usually consist of read text, which means that one has significant control over the content of the data base. Here we address how one can take advantage of this control, by discussing a number of variants of "greedy" text selection methods and showing their application in a variety of examples. 1. INTRODUCTION Both automatic speech recognition (ASR) systems and text to speech (TTS) systems have components that are trained on text---typically read text. Surprisingly often, training text is selected without giving much thought to optimality of the selected text. For limited domain situations, it may very well suffice to select randomly a subset from the domain for training purposes. In many ASR applications, and certainly in most TTS applications, however, the domain is open. And, as discussed at length in [7], in open domain situation...
Jan P. H. van Santen, Adam L. Buchsbaum
EUROSPEECH1
1997 Multi-lingual duration modeling
abstract
Controlling timing in text-to-speech synthesis systems is complicated, because there are many contextual factors that affect timing; moreover, which factors matter and what their precise effects are varies among languages. We describe here a language-independent approach for duration control. At run time, a language-independent timing module accesses languagespecific tables. These tables specify which sub-classes of the feature space (i.e., all combinations of context and phone identity) are homogeneous in the specific sense that the same factors have similar effects on the cases in a sub-class. Within a sub-class, durations are modeled by simple arithmetic models such as multiplicative, additive, or – more generally – sums-ofproducts models. Exploratory statistical methods (supervised) and parameter estimation techniques (unsupervised) are used for
Jan P. H. van Santen, Chilin Shih, Bernd Möbius, Evelyne Tzoukermann, Michael Tanenblatt
EUROSPEECH1
1997 Strong interaction between factors influencing consonant duration
R. J. J. H. van Son, Jan P. H. van Santen
EUROSPEECH2
1996 Modeling segmental duration in German text-to-speech synthesis
Bernd Möbius, Jan P. H. van Santen
ICSLP2
1996 Selecting Training Inputs via Greedy Rank Covering
Adam L. Buchsbaum, Jan P. H. van Santen
SODA2
1994 Segmental effects on timing and height of pitch contours
Jan P. H. van Santen, Julia Hirschberg
ICSLP1
1994 Assignment of segmental duration in text-to-speech synthesis
Jan P. H. van Santen
Comput. Speech Lang.1
1993 Timing in text-to-speech systems
Jan P. H. van Santen
EUROSPEECH1
1993 Perceptual experiments for diagnostic testing of text-to-speech systems
Jan P. H. van Santen
Comput. Speech Lang.1
1992 Diagnostic perceptual experiments for text-to-speech system evaluation
Jan P. H. van Santen
ICSLP1
1992 Contextual effects on vowel duration
Jan P. H. van Santen
Speech Commun.1