VLDB 2026 Research / reviewers in the wild / expert
Raul Fernandez
dblp:72/1615
· DBLP profile ↗
37ranked-venue papers
15as first author
8since 2021 · last 2024
0009-0009-7650-193XORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 31 · 15 first-author · 7 since 2021Artificial intelligence and machine learning · 17 · 7 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Speak While You Think: Streaming Speech Synthesis During Text GenerationabstractLarge Language Models (LLMs) demonstrate impressive capabilities, yet interaction with these models is mostly facilitated through text. Using Text-To-Speech to synthesize LLM outputs typically results in notable latency, which is impractical for fluent voice conversations. We propose LLM2Speech, an architecture to synthesize speech while text is being generated by an LLM which yields significant latency reduction. LLM2Speech mimics the predictions of a non-streaming teacher model while limiting the exposure to future context in order to enable streaming. It exploits the hidden embeddings of the LLM, a by-product of the text generation that contains informative semantic context. Experimental results show that LLM2Speech maintains the teacher’s quality while reducing the latency to enable natural conversations. Avihu Dekel, Slava Shechtman, Raul Fernandez, David Haws, Zvi Kons, Ron Hoory |
ICASSP | 3 |
| 2024 | Exploring the Benefits of Tokenization of Discrete Acoustic UnitsabstractTokenization algorithms that merge the units of a base vocabulary into larger, variable-rate units have become standard in natural language processing tasks.This idea, however, has been mostly overlooked when the vocabulary consists of phonemes or Discrete Acoustic Units (DAUs), an audio-based representation that is playing an increasingly important role due to the success of discrete language-modeling techniques.In this paper, we showcase the advantages of tokenization of phonetic units and of DAUs on three prediction tasks: grapheme-to-phoneme, grapheme-to-DAUs, and unsupervised speech generation using DAU language modeling.We demonstrate that tokenization yields significant improvements in terms of performance, as well as training and inference speed, across all three tasks.We also offer theoretical insights to provide some explanation for the superior performance observed. Avihu Dekel, Raul Fernandez |
INTERSPEECH | 2 |
| 2024 | Creating an African American-Sounding TTS: Guidelines, Technical Challenges, and Surprising EvaluationsabstractRepresentations of AI agents in user interfaces and robotics are predominantly White, not only in terms of facial and skin features, but also in the synthetic voices they use. In this paper we explore some unexpected challenges in the representation of race we found in the process of developing an U.S. English Text-to-Speech (TTS) system aimed to sound like an educated, professional, regional accent-free African American woman. The paper starts by presenting the results of focus groups with African American IT professionals where guidelines and challenges for the creation of a representative and appropriate TTS system were discussed and gathered, followed by a discussion about some of the technical difficulties faced by the TTS system developers. We then describe two studies with U.S. English speakers where the participants were not able to attribute the correct race to the African American TTS voice while overwhelmingly correctly recognizing the race of a White TTS system of similar quality. A focus group with African American IT workers not only confirmed the representativeness of the African American voice we built, but also suggested that the surprising recognition results may have been caused by the inability or the latent prejudice from non-African Americans to associate educated, non-vernacular, professionally-sounding voices to African American people. Claudio S. Pinhanez, Raul Fernandez, Marcelo Grave, Julio Nogima, Ron Hoory |
IUI | 2 |
| 2023 | A Neural TTS System with Parallel Prosody Transfer from Unseen SpeakersabstractModern neural TTS systems are capable of generating natural and expressive speech when provided with sufficient amounts of training data. Such systems can be equipped with prosody-control functionality, allowing for more direct shaping of the speech output at inference time. In some TTS applications, it may be desirable to have an option that guides the TTS system with an ad-hoc speech recording exemplar to impose an implicit fine-grained, user-preferred prosodic realization for certain input prompts. In this work we present a first-of-its-kind neural TTS system equipped with such functionality to transfer the prosody from a parallel text recording from an unseen speaker. We demonstrate that the proposed system can precisely transfer the speech prosody from novel speakers to various trained TTS voices with no quality degradation, while preserving the target TTS speakers' identity, as evaluated by a set of subjective listening experiments. Slava Shechtman, Raul Fernandez |
INTERSPEECH | 2 |
| 2022 | Transplantation of Conversational Speaking Style with Interjections in Sequence-to-Sequence Speech SynthesisabstractSequence-to-Sequence Text-to-Speech architectures that directly generate low level acoustic features from phonetic sequences are known to produce natural and expressive speech when provided with adequate amounts of training data.Such systems can learn and transfer desired speaking styles from one seen speaker to another (in multi-style multi-speaker settings), which is highly desirable for creating scalable and customizable Human-Computer Interaction systems.In this work we explore one-to-many style transfer from a dedicated single-speaker conversational corpus with style nuances and interjections.We elaborate on the corpus design and explore the feasibility of such style transfer when assisted with Voice-Conversion-based data augmentation.In a set of subjective listening experiments, this approach resulted in high-fidelity style transfer with no quality degradation.However, a certain voice persona shift was observed, requiring further improvements in voice conversion. Raul Fernandez, David Haws, Guy Lorberbom, Slava Shechtman, Alexander Sorin |
INTERSPEECH | 1 |
| 2021 | Stable Checkpoint Selection and Evaluation in Sequence to Sequence Speech SynthesisabstractAutoregressive Attentive Sequence-to-Sequence (S2S) speech synthesis is considered state-of-the-art in terms of speech quality and naturalness, as evaluated on a finite set of testing utterances. However, it can occasionally suffer from stability issues at inference time, such as local intelligibility problems or utterance incompletion. Frequently, a model’s stability varies from one checkpoint to another, even after the training loss shows signs of convergence, making the selection of a stable model a tedious and time-consuming task. In this work we propose a novel stability metric designed for automatic checkpoint selection based on incomplete utterance counts within a validation set. The metric is based solely on attention matrix analysis in inference mode and requires no ground-truth output targets. The proposal runs 125 times faster than real-time on a GPU (TeslaK80), allowing convenient incorporation during training to filter out unstable checkpoints, and we demonstrate, via objective and perceptual metrics, its effectiveness in selecting a robust model that attains a good trade-off between stability and quality. Slava Shechtman, David Haws, Raul Fernandez |
ICASSP | 3 |
| 2021 | Synthesis of Expressive Speaking Styles with Limited Training Data in a Multi-Speaker, Prosody-Controllable Sequence-to-Sequence Architecture
Slava Shechtman, Raul Fernandez, Alexander Sorin, David Haws |
Interspeech | 2 |
| 2021 | Supervised and unsupervised approaches for controlling narrow lexical focus in sequence-to-sequence speech synthesisabstractAlthough Sequence-to-Sequence (S2S) architectures have become state-of-the-art in speech synthesis, capable of generating outputs that approach the perceptual quality of natural samples, they are limited by a lack of flexibility when it comes to controlling the output. In this work we present a framework capable of controlling the prosodic output via a set of concise, interpretable, disentangled parameters. We apply this framework to the realization of emphatic lexical focus, proposing a variety of architectures designed to exploit different levels of supervision based on the availability of labeled resources. We evaluate these approaches via listening tests that demonstrate we are able to successfully realize controllable focus while maintaining the same, or higher, naturalness over an established baseline, and we explore how the different approaches compare when synthesizing in a target voice with or without labeled data. Slava Shechtman, Raul Fernandez, David Haws |
SLT | 2 |
| 2018 | Measuring the Effect of Linguistic Resources on Prosody Modeling for Speech SynthesisabstractThe generation of natural and expressive prosodic contours is an important component of a text-to-speech (TTS) system which, in most classical architectures, relies on the existence of a text-analysis processor that can extract prosody-predictive features and pass them to a statistical learning model. These features can range from basic properties of the input string to rich high-level features which may not be always available when developing a TTS system in a new language with sparse computational resources. In this work we investigate how the prosody model of a speech-synthesis system performs as a function of different predictive feature sets that assume access to a certain amount of rich resources. We investigate, using objective metrics, the effect of relaxing the assumptions on input representations for prosody prediction for 5 languages, and evaluate the perceptual implications for US English. Andrew Rosenberg, Raul Fernandez, Bhuvana Ramabhadran |
ICASSP | 2 |
| 2018 | Data Augmentation Improves Recognition of Foreign Accented Speech
Takashi Fukuda, Raul Fernandez, Andrew Rosenberg, Samuel Thomas 0001, Bhuvana Ramabhadran, Alexander Sorin, Gakuto Kurata |
INTERSPEECH | 2 |
| 2018 | Comparing Prosodic Frameworks: Investigating the Acoustic-Symbolic Relationship in ToBI and RaPabstractToBI is the dominant tool for symbolically describing prosodic content in American English speech material. This is due to its descriptive power and its theoretical grounding, but also to the amount of available annotated data. Recently, a modest amount of material annotated with the Rhythm and Pitch (RaP) framework was released publicly. In this paper, we investigate the acoustic-symbolic relationship under these two systems. We present experiments looking at this relationship in both directions. From acoustic to symbolic, we compare the automatic prediction of prosodic prominence as defined under the two systems. From symbolic to acoustic, we examine the utility of these annotation standards to correctly prescribe the acoustics of a given utterance from their symbolic sequences. We find RaP to be promising, showing a somewhat stronger acoustic-symbolic relationship than ToBI given a comparable amount of data for some aspects of these tasks. While with more annotated data ToBI results are stronger, it remains to be shown whether RaP performance can scale up. Raul Fernandez, Andrew Rosenberg |
SLT | 1 |
| 2017 | Voice-transformation-based data augmentation for prosodic classificationabstractIn this work we explore data-augmentation techniques for the task of improving the performance of a supervised recurrent-neural-network classifier tasked with predicting prosodic-boundary and pitch-accent labels. The technique is based on applying voice transformations to the training data that modify the pitch baseline and range, as well as the vocal-tract and vocal-source characteristics of the speakers to generate further training examples. We demonstrate the validity of the approach by improving performance when the amount of base labeled examples is small (showing reductions in the range of 7%–12% for reduced-data conditions) as well as in terms of its generalization to speakers unseen in the training set (showing a relative reduction in the error rate of 8.74% and 4.75%, on the average, for boundaries and accent tasks respectively, in leave-one-speaker-out validation). Raul Fernandez, Andrew Rosenberg, Alexander Sorin, Bhuvana Ramabhadran, Ron Hoory |
ICASSP | 1 |
| 2017 | Weakly-Supervised Phrase Assignment from Text in a Speech-Synthesis System Using Noisy Labels
Asaf Rendel, Raul Fernandez, Zvi Kons, Andrew Rosenberg, Ron Hoory, Bhuvana Ramabhadran |
INTERSPEECH | 2 |
| 2016 | Using continuous lexical embeddings to improve symbolic-prosody prediction in a text-to-speech front-endabstractThe prediction of symbolic prosodic categories from text is an important, but challenging, natural-language processing task given the various ways in which an input can be realized, and the fact that knowledge about what features determine this realization is incomplete or inaccessible to the model. In this work, we look at augmenting baseline features with lexical representations that are derived from text, providing continuous embeddings of the lexicon in a lower-dimensional space. Although learned in an unsupervised fashion, such features capture semantic and syntactic properties that make them amenable for prosody prediction. We deploy various embedding models on prominence- and phrase-break prediction tasks, showing substantial gains, particularly for prominence prediction. Asaf Rendel, Raul Fernandez, Ron Hoory, Bhuvana Ramabhadran |
ICASSP | 2 |
| 2015 | Using deep bidirectional recurrent neural networks for prosodic-target prediction in a unit-selection text-to-speech system
Raul Fernandez, Asaf Rendel, Bhuvana Ramabhadran, Ron Hoory |
INTERSPEECH | 1 |
| 2015 | Modeling phrasing and prominence using deep recurrent learning
Andrew Rosenberg, Raul Fernandez, Bhuvana Ramabhadran |
INTERSPEECH | 2 |
| 2014 | Exploiting vocal-source features to improve ASR accuracy for low-resource languages
Raul Fernandez, Jia Cui, Andrew Rosenberg, Bhuvana Ramabhadran |
INTERSPEECH | 1 |
| 2014 | Prosody contour prediction with long short-term memory, bi-directional, deep recurrent neural networksabstractDeep Neural Networks (DNNs) have been shown to provide state-of-the-art performance over other baseline models in the task of predicting prosodic targets from text in a speechsynthesis system. However, prosody prediction can be affected by an interaction of short- and long-term contextual factors that a static model that depends on a fixed-size context window can fail to properly capture. In this work, we look at a recurrent formulation of neural networks (RNNs) that are deep in time and can store state information from an arbitrarily large input history when making a prediction. We show that RNNs provide improved performance over DNNs of comparable size in terms of various objective metrics for a variety of prosodic streams (notably, a relative reduction of about 6% in F0 mean-square error accompanied by a relative increase of about 14% in F0 variance), as well as in terms of perceptual quality assessed through mean-opinion-score listening tests. Index Terms: speech synthesis,text-to-speech, prosody prediction, recurrent neural networks, deep learning Raul Fernandez, Asaf Rendel, Bhuvana Ramabhadran, Ron Hoory |
INTERSPEECH | 1 |
| 2013 | F0 contour prediction with a deep belief network-Gaussian process hybrid modelabstractIn this work we look at using non-parametric, exemplar-based regression for the prediction of prosodic contour targets from textual features in a speech synthesis system. We investigate the performance of Gaussian Process regression on this task when the covariance kernel operates on a variety of input feature spaces. In particular, we consider non-linear features extracted via Deep Belief Networks. We motivate the use of this hybrid model by considering the initial deep-layer model as a feature extractor that can summarize high-level structure from the raw inputs to improve the regression of an exemplar-based model in the second part of the approach. By looking at both objective metrics and perceptual listening tests, we evaluate these proposals against each other, and against the standard clustering-tree techniques implemented in parametric synthesis for the prediction of prosodic targets. Raul Fernandez, Asaf Rendel, Bhuvana Ramabhadran, Ron Hoory |
ICASSP | 1 |
| 2012 | Prediction of F0 contours from symbolic and numerical variables using continuous conditional random fieldsabstractRegression of continuous-valued variables as a function of both categorical and continuous predictors arises in some areas of speech processing, such as when predicting prosodic targets in a text-to-speech system. In this work we investigate the use of Continuous Conditional Random Fields (CCRF) to conditionally predict F0 targets from a series of s symbolic and numerical predictive features derived from text. We derive the training equations for the model using a Least-Squares-Error criterion within a supervised framework, and evaluate the proposed system using this objective criterion against other baseline models that can handle mixed inputs, such as regression trees and ensemble of regression trees. Raul Fernandez, Steve Minnis, Bhuvana Ramabhadran |
ICASSP | 1 |
| 2012 | Phrase Boundary Assignment from Text in Multiple DomainsabstractDetecting and modeling proper phrasing from an input text string is an important aspect when producing synthesis that sounds intelligible and natural. Knowledge of proper phrase structure influences, e.g., the placement and length of pauses, and the realization of phrase-final boundary contours, both of which can have an effect in a listener’s percepts ranging from naturalness to semantic interpretation. In this work, we look at modeling the occurrence, and types, of phrase breaks from purely textual features, paying close attention to how the performance of the systems generalizes inand out-of-domain for corpora of various types (such as broadcast news, spontaneous speech, and synthesis databases), and as a function of various subsets of syntactical and lexical features investigated. Andrew Rosenberg, Raul Fernandez, Bhuvana Ramabhadran |
INTERSPEECH | 2 |
| 2011 | Exploiting active-learning strategies for annotating prosodic events with limited labeled dataabstractMany applications of spoken-language systems can benefit from having access to annotations of prosodic events. Unfortunately, obtaining human annotations of these events, even sensible amounts to train a supervised system, can become a laborious and costly effort. Given these constraints, this task serves as a good case study for approaches that judiciously guide the selection of data in order to maximize the gain from the human-labeling process or which minimize the size of the training set. To address this, we explore active learning techniques with the objective of reducing the amount of human-annotated data needed to attain a given level of performance. We review strategies that can be used to guide the selection of sequences by combining the output of a classifier and information about the structure of the data into a criterion that can be used during the learning process to query the label of data points that are both informative and representative of the task, and show that for most of the cases considered, active selection strategies when labeling pitch accents and prosodic boundaries are as good as or exceed the performance of random data selection. Raul Fernandez, Bhuvana Ramabhadran |
ICASSP | 1 |
| 2011 | "What is... Dengue Fever?" - Modeling and Predicting Pronunciation Errors in a Text-to-Speech SystemabstractWe propose a system to predict baseform-generation errors in a text-to-speech (TTS) front-end, and aid in the process of cus-tomizing the synthesis engine to a novel application with a large, open-ended vocabulary. We motivate the use of the sys-tem by using data collected during the deployment of the IBM TTS engine in the Watson Deep Question-Answering system customized to play a game of Jeopardy!. We propose a set of features derived from a lexeme’s orthography and candidate baseform, and use a variety of learning schemes and data sam-pling algorithms to address the issue of skewed class priors in the training data. We show that 1) these different approaches provide complementary information that can then be exploited by fusion schemes to improve on the baseline performances, and 2) it is possible to use these techniques to retrieve a list of likely incorrect lexemes so as to reduce the number of tokens that must be vetted before finding and fixing an error. Index Terms: front-end error modeling, speech synthesis 1. Andrew Rosenberg, Raul Fernandez, Bhuvana Ramabhadran |
INTERSPEECH | 2 |
| 2011 | Recognizing affect from speech prosody using hierarchical graphical models
Raul Fernandez, Rosalind W. Picard |
Speech Commun. | 1 |
| 2010 | An autoencoder neural-network based low-dimensionality approach to excitation modeling for HMM-based text-to-speechabstractHMM-TTS synthesis is a popular approach toward flexible, low-footprint, data driven systems that produce highly intelligible speech. In spite of these strengths, speech generated by these systems exhibit some degradation in quality, attributable to an inadequacy in modeling the excitation signal that drives the parametric models of the vocal tract. This paper proposes a novel method for modeling the excitation as a low-dimensional set of coefficients, based on a non-linear map learned through an autoencoder. Through analysis-and-resynthesis experiments, and a formal listening test, we show that this model produces speech of higher perceptual quality compared to conventional pulse-excited speech signals at the p <; 0.01 significance level. Srikanth Vishnubhotla, Raul Fernandez, Bhuvana Ramabhadran |
ICASSP | 2 |
| 2010 | Discriminative training and unsupervised adaptation for labeling prosodic events with limited training data
Raul Fernandez, Bhuvana Ramabhadran |
INTERSPEECH | 1 |
| 2010 | Efficient peerGroup management in JXTA-Overlay P2P system for developing groupware tools
Fatos Xhafa, Leonard Barolli, Santi Caballé, Raul Fernandez |
J. Supercomput. | 4 |
| 2008 | Efficient Peer Selection in P2P JXTA-Based PlatformsabstractP2P systems are nowadays being used not only for file sharing but also for developing large-scale distributed applications. As an emerging paradigm for distributed computing, P2P systems are raising important issues as many novel aspects have to be dealt with in such systems. One key issue in P2P distributed computing is the efficient discovery and selection of peers, which is needed for many purposes such as efficient allocation of jobs to peers, load balancing, efficient file transfer, etc. Existing P2P distributed applications use ad hoc ways to discover and select peers, usually without any performance guarantee. In this paper we address the problem of the efficient peer selection in P2P distributed platforms. To this end, we have developed a P2P distributed platform using Sun's JXTA technology, which is endowed with resource brokerage strategies to efficiently select peers using four selection models: (a) economic scheduling model; (b) priced-based model; (c) peer-priority selection model;and, (d) random selection model. These different models are aimed to match different needs of P2P applications. Next, we have deployed the P2P JXTA platform in a real network using nodes of the PlaneLab - a planetary scale distributed infrastructure- and have experimentally evaluated the performance of the peer selection models by using a distributed application for processing large size log files of a virtual campus, which requires both efficient file transmission and processing in P2P nodes. The results of our work showed the need to develop, implement and evaluate appropriate models for efficient peer selection to match the different requirements of large-scale P2P distributed applications. Although we have used a concrete technology such as JXTA, our approach is applicable in a more general context of P2P and grid computing domain. Finally, our approach to peer selection through brokerage services is very flexible allowing extensions with other models. Fatos Xhafa, Thanasis Daradoumis, Leonard Barolli, Raul Fernandez, Santi Caballé, Vladi Kolici |
AINA | 4 |
| 2008 | Extending JXTA Protocols for P2P File Sharing SystemsabstractFile sharing is among the most important features of the today's Internet-based applications. Most of such applications are server-based approaches inheriting thus the disadvantages of centralized systems. Advances in P2P systems are allowing to share huge quantities of data and files in a distributed way. In this paper, we present extensions of JXTA protocols to support filesharing in P2P systems with the aim to overcome limitations of server-mediated approaches. Our proposal is validated in practice by deploying a P2P file sharing system in a real P2P network. The empirical study revealed the benefits and drawbacks of using JXTA protocols for P2P file sharing systems. Fatos Xhafa, Leonard Barolli, Raul Fernandez, Thanasis Daradoumis, Santi Caballé, Vladi Kolici |
CISIS | 3 |
| 2007 | Database Mining for Flexible Concatenative Text-to-SpeechabstractIn this paper we explore mining a concatenative text-to-speech database to exploit subtle, naturally-occurring stylistic and contextual variability for runtime synthesis. By making a desired style or context known to the search during synthesis, the cost function can be biased toward finding units which satisfy these additional criteria. Having the ability to bias the output of the synthesizer towards a particular voice quality, or other characteristic such as speaking rate, increases its flexibility and potential value. In this paper we illustrate the approach to synthesizing subtle speech variation by focusing on three aspects: prosodic structure (phrase-finalness), prosodic prominence (prosodic accent), and voice quality (breathiness). Target values for the first two of these are automatically generated, while the target value for breathiness is specified by the user. We present results which indicate the value of distinguishing our data along these dimensions, and discuss possible improvements and new uses in the future. Ellen Eide, Raul Fernandez |
ICASSP (4) | 2 |
| 2006 | The IBM expressive text-to-speech synthesis system for American EnglishabstractExpressive text-to-speech (TTS) synthesis should contribute to the pleasantness, intelligibility, and speed of speech-based human-machine interactions which use TTS. We describe a TTS engine which can be directed, via text markup, to use a variety of expressive styles, here, questioning, contrastive emphasis, and conveying good and bad news. Differences in these styles lead us to investigate two approaches for expressive TTS, a "corpus-driven" and a "prosodic-phonology" approach. Each speaker records 11 h (excluding silences) of "neutral" sentences. In the corpus-driven approach, the speaker also records 1-h corpora in each expressive style; these segments are tagged by style for use during search, and decision trees for determining f0contours and timing are trained separately for each of the neutral and expressive corpora. In the prosodic-phonology approach, rules translating certain expressive markup elements to tones and break indices (ToBI) are manually determined, and the ToBI elements are used in single f0and duration trees for all expressions. Tests show that listeners identify synthesis in particular styles ranging from 70% correctly for "conveying bad news" to 85% for "yes-no questions". Further improvements are demonstrated through the use of speaker-pooled f0and duration models John F. Pitrelli, Raimo Bakis, Ellen Eide, Raul Fernandez, Wael Hamza, Michael Picheny |
IEEE Trans. Speech Audio Process. | 4 |
| 2005 | Upending the Uncanny Valley
David Hanson, Andrew Olney, Steve Prilliman, Eric Mathews, Marge Zielke, Derek Hammons, Raul Fernandez, Harry E. Stephanou |
AAAI | 7 |
| 2005 | Classical and novel discriminant features for affect recognition from speechabstractThis paper investigates the performance and relevance of a set of acoustic features for the task of automatic recognition of affect from speech using machine learning techniques. Eighty seven novel and classical features related to loudness, intonation, and voice quality, are examined. Using feature selection, the results yield a performance level of 49.4% recognition rate (compared to a human performance rate of 60.4% and a chance level of 20%), while the relevance results show that the more exploratory and novel subset of these features outrank the more classical features in the recognition task. In the active research area of recognition of affect from speech it is of particular interest to obtain acoustic features that provide results closer to those of human recognition abilities. While many now “classic” features have been proposed in the literature, their performance has still fallen short of human recognition, suggesting the need to continue a search for novel features and methods. This paper briefly highlights results from an extensive investigation developing new features, and comparing them side-by-side with classical ones using machine learning techniques. (See [1] for many details omitted in this paper.) Algorithms and features associated with modeling loudness, intonation, and voice quality are highlighted in § 2, 3 and 4 respectively, and results of the experiments in §5 with some concluding remarks in § 6. Raul Fernandez, Rosalind W. Picard |
INTERSPEECH | 1 |
| 2005 | Toward multiple-language TTS: experiments in English and MandarinabstractText-to-speech systems have dramatically improved in recent years through the use of corpus-based concatenative approaches, and we are beginning to see an interest in endowing them with the ability to handle more than the native language for which they have been developed. In this paper we present ongoing work at IBM in text-to-speech systems that can produce high-quality synthesis in more than one language. We illustrate the discussion with a case study in which two systems, originally developed to support English and Mandarin respectively, have been extended to support each other’s languages. We describe the challenges faced when adapting one system to a different target language, propose adaptation solutions, and present the results of perceptual tests carried out to evaluate how the approaches compare with the performance of the native systems. Raul Fernandez, Wei Zhang 0022, Ellen Eide, Raimo Bakis, Wael Hamza, Michael Picheny, John F. Pitrelli, Yong Qing, Zhiwei Shuang, Li Qin Shen |
INTERSPEECH | 1 |
| 2003 | Modeling drivers' speech under stress
Raul Fernandez, Rosalind W. Picard |
Speech Commun. | 1 |
| 2001 | Frustrating the user on purpose: a step toward building an affective computerabstractUsing a deliberately slow computer–game-interface to induce a state of hypothesised frustration in users, we collected physiological, video and behavioural data, and developed a strategy for coupling these data with real-world events. The effectiveness of our strategy was tested in a study with thirty six subjects, where the system was shown to reliably synchronise and gather data for affect analysis. A pattern-recognition strategy known as Hidden Markov Models was applied to each subject's physiological signals of skin conductivity and blood volume pressure in an effort to see if regimes of likely frustration could be automatically discriminated from regimes when frustration was much less likely. This pattern-recognition approach performed significantly better than random guessing at classifying the two regimes. Mouse-clicking behaviour was also synchronised to frustration-eliciting events and analysed, revealing four distinct patterns of clicking responses. We provide recommendations and guidelines for using physiology as a dependent measure for HCI experiments, especially when considering human emotions in the HCI equation. Jocelyn Scheirer, Raul Fernandez, Jonathan Klein, Rosalind W. Picard |
Interact. Comput. | 2 |
| 1998 | Signal processing for recognition of human frustrationabstractIn this work, inspired by the application of human-machine interaction and the potential use that human-computer interfaces can make of knowledge regarding the affective state of a user, we investigate the problem of sensing and recognizing typical affective experiences that arise when people communicate with computers. In particular, we address the problem of detecting "frustration" in human computer interfaces. By first sensing human biophysiological correlates of internal affective states, we proceed to stochastically model the biological time series with hidden Markov models to obtain user-dependent recognition systems that learn affective patterns from a set of training data. Labeling criteria to classify the data are discussed, and generalization of the results to a set of unobserved data is evaluated. Significant recognition results (greater than random) are reported for 21 of 24 subjects. Raul Fernandez, Rosalind W. Picard |
ICASSP | 1 |