Nigel G. Ward

dblp:53/9229 · DBLP profile ↗
← Back
33ranked-venue papers
25as first author
12since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 28 · 21 first-author · 9 since 2021Artificial intelligence and machine learning · 23 · 17 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 Which Prosodic Features Matter Most for Pragmatics?
abstract
We investigate which prosodic features matter most in conveying pragmatic functions. We use the problem of predicting human perceptions of pragmatic similarity among utterance pairs to evaluate the utility of prosodic features of different types. We find evidence that the duration-related features are most important, that pitch-related features are much less important and less adequate, and that complete modeling will require additional acoustic and prosodic features, including nasality and phonetic reduction. These findings can guide future basic research in prosody, and suggest how to improve speech synthesis evaluation, among other applications.
Nigel G. Ward, Divette Marco, Olac Fuentes
ICASSP1
2025 Towards Precision Characterization of Communication Disorders using Models of Perceived Pragmatic Similarity
abstract
The diagnosis and treatment of individuals with communication disorders offers many opportunities for the application of speech technology, but research so far has not adequately considered: the diversity of conditions, the challenges of limited data, and the role of pragmatic deficits. This paper explores how a general-purpose model of perceived pragmatic similarity may overcome these limitations. It shows that a simple model can capture utterance aspects that are relevant to diagnosis of autism and of specific language impairment, outlines how it might support several use cases for clinicians and clients, and analyzes its performance and limitations.
Nigel G. Ward, Andres Segura, Georgina Bugarini, Heike Lehnert-LeHouillier, Dancheng Liu, Jinjun Xiong, Olac Fuentes
ICASSP1
2025 Phonetic reduction is associated with positive assessment and other pragmatic functions
Nigel G. Ward, Raul O. Gomez, Carlos A. Ortega, Georgina Bugarini
Speech Commun.1
2024 A Collection of Pragmatic-Similarity Judgments over Spoken Dialog Utterances
abstract
Automatic measures of similarity between sentences or utterances are invaluable for training speech synthesizers, evaluating machine translation, and assessing learner productions. While there exist measures for semantic similarity and prosodic similarity, there are as yet none for pragmatic similarity. To enable the training of such measures, we developed the first collection of human judgments of pragmatic similarity between utterance pairs. 9 judges listened to 220 utterance pairs, each consisting of an utterance extracted from a recorded dialog and a re-enactment of that utterance under various conditions designed to create various degrees of similarity. Each pair was rated on a continuous scale. The average inter-judge correlation was 0.45. We make this data available at https://github.com/divettemarco/PragSim .
Nigel G. Ward, Divette Marco
LREC/COLING1
2024 Pragmatically similar utterance finder demonstration
Nigel G. Ward, Andres Segura
INTERSPEECH1
2024 Towards a General-Purpose Model of Perceived Pragmatic Similarity
Nigel G. Ward, Andres Segura, Alejandro Ceballos, Divette Marco
INTERSPEECH1
2023 Towards Cross-Language Prosody Transfer for Dialog
Jonathan E. Avila, Nigel G. Ward
INTERSPEECH2
2023 A dimensional model of interaction style variation in spoken dialog
Nigel G. Ward, Jonathan E. Avila
Speech Commun.1
2022 Comparison of Models for Detecting Off-Putting Speaking Styles
Diego Aguirre, Nigel G. Ward, Jonathan E. Avila, Heike Lehnert-LeHouillier
INTERSPEECH2
2022 On the Utility of Self-Supervised Models for Prosody-Related Tasks
abstract
Self-Supervised Learning (SSL) from speech data has produced models that have achieved remarkable performance in many tasks, and that are known to implicitly represent many aspects of information latently present in speech signals. However, relatively little is known about the suitability of such models for prosody-related tasks or the extent to which they encode prosodic information. We present a new evaluation framework, “SUPERB-prosody,” consisting of three prosody-related downstream tasks and two pseudo tasks. We find that 13 of the 15 SSL models outperformed the baseline on all the prosody-related tasks. We also show good performance on two pseudo tasks: prosody reconstruction and future prosody prediction. We further analyze the layerwise contributions of the SSL models. Overall we conclude that SSL speech models are highly effective for prosody-related tasks. We release our code11https://github.com/JSALT-2022-SSL/superb-prosody for the community to support further investigation of SSL models' utility for prosody.
Guan-Ting Lin, Chi-Luen Feng, Wei-Ping Huang, Yuan Tseng, Tzu-Han Lin, Chen-An Li, Hung-yi Lee, Nigel G. Ward
SLT8
2022 Spoken language interaction with robots: Recommendations for future research
abstract
With robotics rapidly advancing, more effective human–robot interaction is increasingly needed to realize the full potential of robots for society. While spoken language must be part of the solution, our ability to provide spoken language interaction capabilities is still very limited. In this article, based on the report of an interdisciplinary workshop convened by the National Science Foundation, we identify key scientific and engineering advances needed to enable effective spoken language interaction with robotics. We make 25 recommendations, involving eight general themes: putting human needs first, better modeling the social and interactive aspects of language, improving robustness, creating new methods for rapid adaptation, better integrating speech and language with other communication modalities, giving speech and language components access to rich representations of the robot’s current knowledge and state, making all components operate in real time, and improving research infrastructure and resources. Research and development that prioritizes these topics will, we believe, provide a solid foundation for the creation of speech-capable robots that are easy and effective for humans to work with.
Matthew Marge, Carol Y. Espy-Wilson, Nigel G. Ward, Abeer Alwan, Yoav Artzi, Mohit Bansal, Gilmer L. Blankenship, Joyce Y. Chai, Hal Daumé III, Debadeepta Dey, Mary P. Harper, Thomas Howard, Casey Kennington, Ivana Kruijff-Korbayová, Dinesh Manocha, Cynthia Matuszek, Ross Mead, Raymond J. Mooney, Roger K. Moore, Mari Ostendorf, Heather Pon-Barry, Alexander I. Rudnicky, Matthias Scheutz, Robert St. Amant, Stefanie Tellex, David R. Traum, Zhou Yu 0005
Comput. Speech Lang.3
2021 Towards Continuous Estimation of Dissatisfaction in Spoken Dialog
abstract
We collected a corpus of human-human taskoriented dialogs rich in dissatisfaction and built a model that used prosodic features to predict when the user was likely dissatisfied.For utterances this attained a F .25 score of 0.62, against a baseline of 0.39.Based on qualitative observations and failure analysis, we discuss likely ways to improve this result to make it have practical utility.
Nigel G. Ward, Jonathan E. Avila, Aaron M. Alarcon
SIGDIAL1
2019 Survey Talk: Prosody Research and Applications: The State of the Art
Nigel G. Ward
INTERSPEECH1
2018 Turn-Taking Predictions across Languages and Genres Using an LSTM Recurrent Neural Network
abstract
Going beyond turn-taking models built to solve specific tasks, such as predicting if a user will hold his/her turn after a pause, there is growing interest in more general models for turn taking that subsume many such tasks, and very good results have recently been obtained [1]. Here we present an improved recurrent network model that outperforms [1] and does so without requiring lexical annotation. Further, we show that this model can be trained for different languages with no modifications, providing good results in turn-taking prediction for English, Spanish, Japanese, Mandarin and French. We also show that our model performs well across genres, including task-oriented dialog and general conversation.
Nigel G. Ward, Diego Aguirre, Gerardo Cervantes, Olac Fuentes
SLT1
2018 Inferring stance in news broadcasts from prosodic-feature configurations
Nigel G. Ward, Jason C. Carlson, Olac Fuentes
Comput. Speech Lang.1
2017 Inferring Stance from Prosody
Nigel G. Ward, Jason C. Carlson, Olac Fuentes, Diego Castán, Elizabeth Shriberg, Andreas Tsiartas
INTERSPEECH1
2016 On the possibility of predicting gaze aversion to improve video-chat efficiency
abstract
A possible way to make video chat more efficient is to only send video frames that are likely to be looked at by the remote participant. Gaze in dialog is intimately tied to dialog states and behaviors, so prediction of such times should be possible. To investigate, we collected data on both participants in 6 video-chat sessions, totalling 65 minutes, and created a model to predict whether a participant will be looking at the screen 300 milliseconds in the future, based on prosodic and gaze information available at the other side. A simple predictor had a precision of 42% at the equal error rate. While this is probably not good enough to be useful, improved performance should be readily achievable.
Nigel G. Ward, Chelsey N. Jurado, Ricardo A. Garcia, Florencia A. Ramos
ETRA1
2016 Prediction and Generation of Backchannel Form for Attentive Listening Systems
Tatsuya Kawahara, Takashi Yamaguchi, Koji Inoue, Katsuya Takanashi, Nigel G. Ward
INTERSPEECH5
2015 A prosody-based vector-space model of dialog activity for information retrieval
Nigel G. Ward, Steven D. Werner, Fernando García 0001, Emilio Sanchis Arnal
Speech Commun.1
2013 Using dialog-activity similarity for spoken information retrieval
abstract
We want to enable users to locate desired information in spo-ken audio documents using not only the words, but also dia-log activities. Following previous research, we infer this infor-mation from prosodic features, however, instead of retrieval by matching to a predefined finite set of activities, we estimate sim-ilarity using a vector space representation. Utterances close in this vector space are frequently similar not only pragmatically, but also topically. Using this we implemented a dialog-based query-by-example function and built it into an interface for use in combination with normal lexical search. In an experiment searchers used the new feature and considered it helpful, but only for some search tasks. Index Terms: audio information retrieval, spoken content re-trieval, prosody, vector-space model, query by example
Nigel G. Ward, Steven D. Werner
INTERSPEECH1
2012 Towards Empirical Dialog-State Modeling and its Use in Language Modeling
abstract
Inspired by the goal of modeling the dialog state and the speaker’s mental state, moment by moment, we apply Principal Component Analysis to a vector of 76 prosodic features spanning 6 seconds of context. This gives a multidimensional representation of the current state. We find that word probabilities vary strongly with several of these dimensions, that the use of this information in a language model gives a 27 % reduction in perplexity, and that many of the dimensions do relate to aspects of mental state and dialog state.
Nigel G. Ward, Alejandro Vega
INTERSPEECH1
2012 A Bottom-Up Exploration of the Dimensions of Dialog State in Spoken Interaction
Nigel G. Ward, Alejandro Vega
SIGDIAL Conference1
2012 Prosodic and temporal features for language modeling for dialog
Nigel G. Ward, Alejandro Vega, Timo Baumann
Speech Commun.1
2011 Achieving rapport with turn-by-turn, user-responsive emotional coloring
Jaime C. Acosta, Nigel G. Ward
Speech Commun.2
2010 Dialog prediction for a general model of turn-taking
abstract
Today there are solutions for some specific turn-taking problems, but no general model. We show how turn-taking can be reduced to two more general problems, prediction and selection. We also discuss the value of predicting not only future speech/silence but also prosodic features, thereby handing not only turn-taking but “turn-shaping”. To illustrate how such predictions can be made, we trained a neural network predictor. This was adequate to support some specific turn-taking decisions and was modestly accurate overall. Index Terms: dialog system, predictive, prosody, endpointing, back-channeling, time, interaction control, dialog model
Nigel G. Ward, Olac Fuentes, Alejandro Vega
INTERSPEECH1
2009 Towards the use of inferred cognitive states in language modeling
abstract
In spoken dialog, speakers are simultaneously engaged in various mental processes, and it seems likely that the word that will be said next depends, to some extent, on the states of these mental processes. Further, these states can be inferred, to some extent, from properties of the speaker's voice as they change from moment to moment. As a illustration of how to apply these ideas in language modeling, we examine volume and speaking rate as predictors of the upcoming word. Combining the information which these provide with a trigram model gave a 2.6% improvement in perplexity.
Nigel G. Ward, Alejandro Vega
ASRU1
2009 Responding to user emotional state by adding emotional coloring to utterances
abstract
When people speak to each other, they share a rich set of nonverbal behaviors such as varying prosody in voice. These behaviors, sometimes interpreted as demonstrations of emotions, call for appropriate responses, but today’s spoken dialog systems lack the ability to do so. We collected a corpus of persuasive dialogs, specifically conversations about graduate school between a staff member and students, and had judges label all utterances with triples indicating the perceived emotions, using the three dimensions: activation, evaluation, and power. We found immediate response patterns, in which the staff member colored her utterances in response to the emotion shown by the student in the immediately previous utterance, and built a predictive model suitable for use in a dialog system to persuasively discuss graduate school with students. Index Terms: emotional responses, dimensional emotions, user
Jaime C. Acosta, Nigel G. Ward
INTERSPEECH2
2009 Using responsive prosodic variation to acknowledge the user's current state
abstract
Spoken dialog systems today do not vary the prosody of their utterances, although prosody is known to have many useful expressive functions. In a corpus of memory quizzes, we identify eleven dimensions of prosodic variation, each with its own expressive function. We identified the situations in which each was used, and developed rules for detecting these situations from the dialog context and the prosody of the interlocutor’s previous utterance. We implemented the resulting rules and had 21 users interact with two versions of the system. Overall they preferred the version in which the prosodic forms of the acknowledgments were chosen to be suitable for each specific context. This suggests that simple adjustments to system prosody based on local context can have value to users.
Nigel G. Ward, Rafael Escalante-Ruiz
INTERSPEECH1
2009 Estimating the potential of signal and interlocutor-track information for language modeling
abstract
Although today most language models treat language purely as word sequences, there is recurring interest in tapping new sources of information, such as disfluencies, prosody, the interlocutor’s dialog act, and the interlocutor’s recent words. In order to estimate the potential value of such sources of information, we extend Shannon’s guessing-game method for estimating entropy to work for spoken dialog. Four teams of two subjects each predicted the next word in a dialog using various amounts of context: one word, two words, all the words spoken so far, or the full dialog audio so far. The entropy benefit in the full-audio condition over the full text condition was substantial,.64 bits per word, greater than the.54 bit benefit of full text context over trigrams. This suggests that language models may be improved by use of the prosody of the speaker and context from the interlocutor. Index Terms: entropy, perplexity, Shannon’s guessing game, prediction, context, prosody
Nigel G. Ward, Benjamin H. Walker
INTERSPEECH1
2008 Modeling the effects on time-into-utterance on word probabilities
abstract
Most language models treat speech as simply sequences of words, ignoring the fact that words are also events in time. This paper reports an initial exploration of how word probabilities vary with time-into-utterance, and proposes a method for using this information to improve a language model. This is done by computing the ratio of the probability of the word at a specific time to its overall unigram probability, and using this ratio to adjust the n-gram probability. On casual dialogs from Switchboard this method gave a modest reduction in perplexity.
Nigel G. Ward, Alejandro Vega
INTERSPEECH1
2006 Prosodic feature generation for back-channel prediction
abstract
Using prosodic information to predict when back-channels are appropriate in spontaneous dialogs has become somewhat of a reference problem for automatic discovery techniques. Here we present experiments with two ideas: the use of features derived from randomly generated pitch and energy filters, and the use of instancebased learning, specifically the Locally Weighted Linear Regression (LWLR) algorithm. For the task of predicting possible backchannel locations in Iraqi Arabic [6], we obtain 22 % precision and 51 % recall, which is as good as that obtained using a laboriously developed and hand-tuned rule. 1.
Thamar Solorio, Olac Fuentes, Nigel G. Ward, Yaffa Al Bayyari
INTERSPEECH3
2006 A case study in the identification of prosodic cues to turn-taking: back-channeling in Arabic
abstract
Discovering and quantifying the prosodic signals that help manage turn-taking is difficult, in part because of the limitations of commonly used methods. This paper presents an integrated method that uses both perceptually-based analysis and quantitative analysis. The eight activities involved in the method — clarification of aims, problem formulation, corpus preparation, feature discovery, feature combination, hypothesis refinement, tuning, and evaluation — are illustrated using task of finding prosodic cues for back-channel feedback in Arabic.
Nigel G. Ward, Yaffa Al Bayyari
INTERSPEECH1
2005 Root causes of lost time and user stress in a simple dialog system
abstract
As a priority-setting exercise, we compared interactions between users and a simple spoken dialog system to interactions between users and a human operator. We observed usability events, places in which system behavior differed from human behavior, and for each we noted the impact, root causes, and prospects for improvement. We suggest some priority issues for research, involving not only such core areas as speech recognition and synthesis and language understanding and generation, but also less-studied topics such as adaptive or flexible timeouts, turn-taking and speaking rate.
Nigel G. Ward, Anais G. Rivera, Karen Ward, David G. Novick
INTERSPEECH1