EDBT 2026 Demo / reviewers in the wild / expert
Jens Edlund
dblp:10/3080
· DBLP profile ↗
51ranked-venue papers
12as first author
12since 2021 · last 2026
0000-0001-9327-9482ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 44 · 11 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 36 · 8 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Setting the Stage for Disfluency: Implications of Contextual Task Framing Effects for the Design of Listening Tasks
Ambika Kirkland, Jens Edlund |
LREC | 2 |
| 2026 | A Shoal of Voices: Parallel Read Speech from Professional Swedish Narrators
Christina Tånnander, Jim O'Regan, Jens Edlund |
LREC | 3 |
| 2026 | The use of variable length stimuli for assessing segmental distortion in TTS evaluation
Ayushi Pandey, Jens Edlund, Sébastien Le Maguer, Naomi Harte |
Comput. Speech Lang. | 2 |
| 2025 | Who knows best? Effects of speech disfluencies on incentivized decision-making
Ambika Kirkland, Jens Edlund |
INTERSPEECH | 2 |
| 2025 | Intrasentential English in Swedish TTS: perceived English-accentedness
Christina Tånnander, David House, Jonas Beskow, Jens Edlund |
INTERSPEECH | 4 |
| 2024 | Revisiting Three Text-to-Speech Synthesis Experiments with a Web-Based Audience Response SystemabstractIn order to investigate the strengths and weaknesses of Audience Response System (ARS) in text-to-speech synthesis (TTS) evaluations, we revisit three previously published TTS studies and perform an ARS-based evaluation on the stimuli used in each study. The experiments are performed with a participant pool of 39 respondents, using a web-based tool that emulates an ARS experiment. The results of the first experiment confirms that ARS is highly useful for evaluating long and continuous stimuli, particularly if we wish for a diagnostic result rather than a single overall metric, while the second and third experiments highlight weaknesses in ARS with unsuitable materials as well as the importance of framing and instruction when conducting ARS-based evaluation. Christina Tånnander, Jens Edlund, Joakim Gustafson |
LREC/COLING | 2 |
| 2024 | Assessing the impact of contextual framing on subjective TTS qualityabstractEdlund J, Tånnander C, LeMaguer S, Wagner P. Assessing the impact of contextual framing on subjective TTS quality. In: Proceedings of INTERSPEECH 2024. 2024: 1205--1209. Jens Edlund, Christina Tånnander, Sébastien Le Maguer, Petra Wagner |
INTERSPEECH | 1 |
| 2024 | Beyond graphemes and phonemes: continuous phonological features in neural text-to-speech synthesis
Christina Tånnander, Shivam Mehta, Jonas Beskow, Jens Edlund |
INTERSPEECH | 4 |
| 2023 | Crowdsource-based Validation of the Audio Cocktail as a Sound Browsing Tool
Per Fallgren, Jens Edlund |
INTERSPEECH | 2 |
| 2023 | Listener sensitivity to deviating obstruents in WaveNet
Ayushi Pandey, Jens Edlund, Sébastien Le Maguer, Naomi Harte |
INTERSPEECH | 2 |
| 2021 | Human-in-the-Loop Efficiency Analysis for Binary Classification in Edyson
Per Fallgren, Jens Edlund |
Interspeech | 2 |
| 2021 | Understanding acceptability of disordered speech through Audience Response Systems-based evaluationabstractWe explore the validity and reliability of an Audience Response Systems (ARS)-based measure of acceptability, applied to speech produced by children with speech sound disorder (SSD). We further explore how the suggested measure relates to an ARS-based measure of intelligibility. Finally, we explore potential differences between speech-language pathologists (SLPs), untrained adults, and children in their assessments. Fifty-three listeners participated in ARS-based assessments of acceptability and intelligibility: 19 SLPs, 18 untrained adults, and 16 children (aged 6-10 years). The listeners assessed speech samples collected from 14 children with SSD and 2 children with typical speech. The validity of the ARS-based acceptability measure was investigated through correlation analyses with reference to rating-based acceptability, as well as to measures of intelligibility and speech proficiency. The ARS-based acceptability measure correlated strongly with all assumed related measures. Listeners reacted more often to affected acceptability than to unintelligibility. Child listeners reacted less frequently to acceptability disruption than the SLPs and untrained adults. ARS-based assessment of acceptability provides a valid and reliable measure of acceptability. The ARS-based methodology captures an anticipated pattern where listeners react more often to disruptions of acceptability than to unintelligibility. Children appear less sensitive to traits signaling SSD than SLPs and other adults are; however, replication is required to establish this for sure. Sofia Strömbergsson, Jens Edlund, Anita McAllister, Tove Lagerberg |
Speech Commun. | 2 |
| 2020 | Augmented Prompt Selection for Evaluation of Spontaneous Speech SynthesisabstractBy definition, spontaneous speech is unscripted and created on the fly by the speaker. It is dramatically different from read speech, where the words are authored as text before they are spoken. Spontaneous speech is emergent and transient, whereas text read out loud is pre-planned. For this reason, it is unsuitable to evaluate the usability and appropriateness of spontaneous speech synthesis by having it read out written texts sampled from for example newspapers or books. Instead, we need to use transcriptions of speech as the target - something that is much less readily available. In this paper, we introduce Starmap, a tool allowing developers to select a varied, representative set of utterances from a spoken genre, to be used for evaluation of TTS for a given domain. The selection can be done from any speech recording, without the need for transcription. The tool uses interactive visualisation of prosodic features with t-SNE, along with a tree-based algorithm to guide the user through thousands of utterances and ensure coverage of a variety of prompts. A listening test has shown that with a selection of genre-specific utterances, it is possible to show significant differences across genres between two synthetic voices built from spontaneous speech. Éva Székely, Jens Edlund, Joakim Gustafson |
LREC | 2 |
| 2019 | How to Annotate 100 Hours in 45 MinutesabstractSpeech data found in the wild hold many advantages over artificially constructed speech corpora in terms of ecological validity and cultural worth. Perhaps most importantly, there is a lot of it. H ... Per Fallgren, Zofia Malisz, Jens Edlund |
INTERSPEECH | 3 |
| 2019 | Spot the Pleasant People! Navigating the Cocktail Party BuzzabstractWe present an experimental platform for making voice likability assessments that are decoupled from individual voices, and instead capture voice characteristics over groups of speakers. We employ methods that we have previously used for other purposes to create the Cocktail platform, where respondents navigate in a voice buzz made up of about 400 voices on a touch screen. They then choose the location where they find the voice buzz most pleasant. Since there is no image or message on the screen, the platform can be used by visually impaired people, who often need to rely on spoken text, on the same premises as seeing people. In this paper, we describe the platform and its motivation along with our analysis method. We conclude by presenting two experiments in which we verify that the platform behaves as expected: one simple sanity test, and one experiment with voices grouped according to their mean pitch variance. Christina Tånnander, Per Fallgren, Jens Edlund, Joakim Gusafsson |
INTERSPEECH | 3 |
| 2019 | The State of Speech in HCI: Trends, Themes and ChallengesabstractAbstract Speech interfaces are growing in popularity. Through a review of 99 research papers this work maps the trends, themes, findings and methods of empirical research on speech interfaces in the field of human–computer interaction (HCI). We find that studies are usability/theory-focused or explore wider system experiences, evaluating Wizard of Oz, prototypes or developed systems. Measuring task and interaction was common, as was using self-report questionnaires to measure concepts like usability and user attitudes. A thematic analysis of the research found that speech HCI work focuses on nine key topics: system speech production, design insight, modality comparison, experiences with interactive voice response systems, assistive technology and accessibility, user speech production, using speech technology for development, peoples’ experiences with intelligent personal assistants and how user memory affects speech interface interaction. From these insights we identify gaps and challenges in speech research, notably taking into account technological advancements, the need to develop theories of speech interface interaction, grow critical mass in this domain, increase design work and expand research from single to multiple user interaction contexts so as to reflect current use contexts. We also highlight the need to improve measure reliability, validity and consistency, in the wild deployment and reduce barriers to building fully functional speech interfaces for research. RESEARCH HIGHLIGHTS Most papers focused on usability/theory-based or wider system experience research with a focus on Wizard of Oz and developed systems Questionnaires on usability and user attitudes often used but few were reliable or validated Thematic analysis showed nine primary research topics Challenges identified in theoretical approaches and design guidelines, engaging with technological advances, multiple user and in the wild contexts, critical research mass and barriers to building speech interfaces Leigh Clark, Philip R. Doyle, Diego Garaialde, Emer Gilmartin, Stephan Schlögl, Jens Edlund, Matthew P. Aylett, João P. Cabral, Cosmin Munteanu, Justin Edwards, Benjamin R. Cowan |
Interact. Comput. | 6 |
| 2018 | Bringing Order to Chaos: A Non-Sequential Approach for Browsing Large Sets of Found Audio Data
Per Fallgren, Zofia Malisz, Jens Edlund |
LREC | 3 |
| 2017 | Approximating Phonotactic Input in Children's Linguistic Environments from Orthographic TranscriptsabstractChild-directed spoken data is the ideal source of support for claims about children’s linguistic environments. However, phonological transcriptions of child-directed speech are scarce,compared to sources like adult-directed speech or text data. Acquiring reliable descriptions of children’s phonological environments from more readily accessible sources would mean considerable savings of time and money. The first step towards this goal is to quantify the reliability of descriptions derived from such secondary sources. We investigate how phonological distributions vary across different modalities (spoken vs. written), and across the age of the intended audience (children vs. adults). Using a previously unseen collection of Swedish adult- and child-directed spoken and written data, we combine lexicon look-up and grapheme-to-phonemeconversion to approximate phonological characteristics. The analysis shows distributional differences across datasets both for single phonemes and for longer phoneme sequences. Some of these are predictably attributed to lexical and contextual characteristics of text vs. speech.The generated phonological transcriptions are remarkably reliable. The differences in phonological distributions between child-directed speech and secondary sources highlight a need for compensatory measures when relying on written data or onadult-directed spoken data, and/or for continued collection ofactual child-directed speech in research on children’s language environments. Sofia Strömbergsson, Jens Edlund, Jana Götze, Kristina N. Björkenstam |
INTERSPEECH | 2 |
| 2016 | Hidden Resources ― Strategies to Acquire and Exploit Potential Spoken Language Resources in National Archives
Jens Edlund, Joakim Gustafson |
LREC | 1 |
| 2015 | Communicative needs and respiratory constraintsabstractThis study investigates timing of communicative behaviour with respect to speaker’s respiratory cycle. The data is drawn from a corpus of multiparty conversations in Swedish. We find that while longer utterances (> 1 s) are tied, predictably, primarily to exhalation onset, shorter vocalisations are spread more uni- formly across the respiratory cycle. In addition, nods, which are free from any respiratory constraints, are most frequently found around exhalation offsets, where respiratory requirements for even a short utterance are not satisfied. We interpret the results to reflect the economy principle in speech production, whereby respiratory effort, associated primarily with starting a new respiratory cycle, is minimised within the scope of speaker’s communicative goals. Marcin Wlodarczak, Mattias Heldner, Jens Edlund |
INTERSPEECH | 3 |
| 2014 | Ranking severity of speech errors by their phonological impact in contextabstractChildren with speech disorders often present with systematic speech error patterns. In clinical assessments of speech disorders, evaluating the severity of the disorder is central. Current measures of severity have limited sensitivity to factors like the frequency of the target sounds in the child’s language and the degree of phonological diversity, which are factors that can be assumed to affect intelligibility. By constructing phonological filters to simulate eight speech error patterns often observed in children, and applying these filters to a phonologically transcribed corpus of 350K words, this study explores three quantitative measures of phonological impact: Percentage of Consonants Correct (PCC), edit distance, and degree of homonymy. These metrics were related to estimated ratings of severity collected from 34 practicing clinicians. The results show an expected high correlation between the PCC and edit distance metrics, but that none of the three metrics align with clinicians’ ratings. Although these results do not generate definite answers to what phonological factors contribute the most to (un)intelligibility, this study demonstrates a methodology that allows for large-scale investigations of the interplay between phonological errors and their impact on speech in context, within and across languages. Sofia Strömbergsson, Christina Tånnander, Jens Edlund |
INTERSPEECH | 3 |
| 2013 | Analysis of gaze and speech patterns in three-party quiz game interactionabstractIn order to understand and model the dynamics between interaction phenomena such as gaze and speech in face-to-face multiparty interaction between humans, we need large quantities of reliable, objective data of such interactions. To date, this type of data is in short supply. We present a data collection setup using automated, objective techniques in which we capture the gaze and speech patterns of triads deeply engaged in a high-stakes quiz game. The resulting corpus consists of five one-hour recordings, and is unique in that it makes use of three state-of-the-art gaze trackers (one per subject) in combination with a state-of-theart conical microphone array designed to capture roundtable meetings. Several video channels are also included. In this paper we present the obstacles we encountered and the possibilities afforded by a synchronised, reliable combination of large-scale multi-party speech and gaze data, and an overview of the first analyses of the data. Index Terms: multimodal corpus, multiparty dialogue, gaze patterns, multiparty gaze. Samer Al Moubayed, Jens Edlund, Joakim Gustafson |
INTERSPEECH | 2 |
| 2013 | Timing responses to questions in dialogueabstractQuestions and answers play an important role in spoken dialogue systems as well as in human-human interaction. A critical concern when responding to a question is the timing of the response. While human response times depend on a wide set of features, dialogue systems generally respond as soon as they can, that is, when the end of the question has been detected and the response is ready to be deployed. This paper presents an analysis of how different semantic and pragmatic features affect the response times to questions in two different data sets of spontaneous human-human dialogues: the Swedish Spontal Corpus and the US English Switchboard corpus. Our analysis shows that contextual features such as question type, response type, and conversation topic influence human response times. Based on these results, we propose that more sophisticated response timing can be achieved in spoken dialogue systems by using these features to automatically and deliberately target system response timing. Index Terms: speech prosody, spontaneous speech, question intonation, response times Sofia Strömbergsson, Anna Hjalmarsson, Jens Edlund, David House |
INTERSPEECH | 3 |
| 2012 | On the effect of the acoustic environment on the accuracy of perception of speaker orientation from auditory cues aloneabstractThe ability of people, and of machines, to determine the position of a sound source in a room is well studied. The related ability to determine the orientation of a directed sound source, on the other hand, is not, but the few studies there are show people to be surprisingly skilled at it. This has bearing for studies of face-to- face interaction and of embodied spoken dialogue systems, as sound source orientation of a speaker is connected to the head pose of the speaker, which is meaningful in a number of ways. The feature most often implicated for detection of sound source orientation is the inter-aural level difference - a feature which it is assumed is more easily exploited in anechoic chambers than in everyday surroundings. We expand here on our previous studies and compare detection of speaker orientation within and outside of the anechoic chamber. Our results show that listeners find the task easier, rather than harder, in everyday surroundings, which suggests that inter-aural level differences is not the only feature at play. Jens Edlund, Mattias Heldner, Joakim Gustafson |
INTERSPEECH | 1 |
| 2012 | On the Dynamics of Overlap in Multi-Party ConversationabstractOverlap, although short in duration, occurs frequently in multiparty conversation. We show that its duration is approximately log-normal, and inversely proportional to the number of simultaneously speaking parties. Using a simple model, we demonstrate that simultaneous talk tends to end simultaneously less frequently than in begins simultaneously, leading to an arrow of time in chronograms constructed from speech activity alone. The asymmetry is significant and discriminative. It appears to be due to dialog acts which do not carry propositional content, and those which are not brought to completion. Index Terms: multi-party conversation, overlap, turn-taking. 1. Kornel Laskowski, Mattias Heldner, Jens Edlund |
INTERSPEECH | 3 |
| 2012 | Gaze Patterns in Turn-TakingabstractOertel C, Wlodarczak M, Edlund J, Wagner P, Gustafson J. Gaze patterns in turn-taking. In: 13th Annual Conference of the International Speech Communication Association 2012 (INTERSPEECH 2012). Red Hook, NY: Curran; 2013: 2243-2246. Catharine Oertel, Marcin Wlodarczak, Jens Edlund, Petra Wagner, Joakim Gustafson |
INTERSPEECH | 3 |
| 2012 | Prosodic measurements and question types in the Spontal corpus of Swedish dialoguesabstractStudies of questions present strong evidence that there is no oneto-one relationship between intonation and interrogative mode. In this paper, we describe some aspects of prosodic variation in the Spontal corpus of 120 half-hour spontaneous dialogues in Swedish. The study is part of ongoing work aimed at extracting a database of 600 questions from the corpus, complete with categorization and prosodic descriptions. We report on coding and annotation of question typology and present results concerning some prosodic correlates related to question type for the 600 questions. A prosodically salient distinction was found between the two categories termed, in our typology, forward and backward looking questions. Index Terms: speech prosody, spontaneous speech, question intonation, interrogative intonation Sofia Strömbergsson, Jens Edlund, David House |
INTERSPEECH | 2 |
| 2012 | 3rd party observer gaze as a continuous measure of dialogue flow
Jens Edlund, Simon Alexanderson, Jonas Beskow, Lisa Gustavsson, Mattias Heldner, Anna Hjalmarsson, Petter Kallionen, Ellen Marklund |
LREC | 1 |
| 2012 | Taming Mona Lisa: Communicating gaze faithfully in 2D and 3D facial projectionsabstractThe perception of gaze plays a crucial role in human-human interaction. Gaze has been shown to matter for a number of aspects of communication and dialogue, especially for managing the flow of the dialogue and participant attention, for deictic referencing, and for the communication of attitude. When developing embodied conversational agents (ECAs) and talking heads, modeling and delivering accurate gaze targets is crucial. Traditionally, systems communicating through talking heads have been displayed to the human conversant using 2D displays, such as flat monitors. This approach introduces severe limitations for an accurate communication of gaze since 2D displays are associated with several powerful effects and illusions, most importantly the Mona Lisa gaze effect, where the gaze of the projected head appears to follow the observer regardless of viewing angle. We describe the Mona Lisa gaze effect and its consequences in the interaction loop, and propose a new approach for displaying talking heads using a 3D projection surface (a physical model of a human head) as an alternative to the traditional flat surface projection. We investigate and compare the accuracy of the perception of gaze direction and the Mona Lisa gaze effect in 2D and 3D projection surfaces in a five subject gaze perception experiment. The experiment confirms that a 3D projection surface completely eliminates the Mona Lisa gaze effect and delivers very accurate gaze direction that is independent of the observer's viewing angle. Based on the data collected in this experiment, we rephrase the formulation of the Mona Lisa gaze effect. The data, when reinterpreted, confirms the predictions of the new model for both 2D and 3D projection surfaces. Finally, we discuss the requirements on different spatially interactive systems in terms of gaze direction, and propose new applications and experiments for interaction in a human-ECA and a human-robot settings made possible by this technology. Samer Al Moubayed, Jens Edlund, Jonas Beskow |
ACM Trans. Interact. Intell. Syst. | 2 |
| 2011 | Syllabification of conversational speech using Bidirectional Long-Short-Term Memory Neural NetworksabstractSegmentation of speech signals is a crucial task in many types of speech analysis. We present a novel approach at segmentation on a syllable level, using a Bidirectional Long-Short-Term Memory Neural Network. It performs estimation of syllable nucleus positions based on regression of perceptually motivated input features to a smooth target function. Peak selection is performed to attain valid nuclei positions. Performance of the model is evaluated on the levels of both syllables and the vowel segments making up the syllable nuclei. The general applicability of the approach is illustrated by good results for two common databases-Switchboard and TIMIT-for both read and spontaneous speech, and a favourable comparison with other published results. Christian Landsiedel, Jens Edlund, Florian Eyben, Daniel Neiberg, Björn W. Schuller |
ICASSP | 2 |
| 2011 | A single-port non-parametric model of turn-taking in multi-party conversationabstractThe taking of turns to speak is an intrinsic property of conversation. It is expected that models of taking turns, providing a prior distribution over conversational form, can reduce the perplexity of what is attended to and processed by spoken dialogue systems. We propose a single-port model of multi-party turn-taking which allows conversants to behave independently but to condition their behavior on the past of the entire group. The model performs at least as well as an existing multi-port model on perplexity over subsequent speech activity. We quantify the effect of longer histories and more distant future horizons, and argue that the framework has the potential to inform the design and behavior of spoken dialogue systems. Kornel Laskowski, Jens Edlund, Mattias Heldner |
ICASSP | 2 |
| 2011 | Very Short Utterances and Timing in Turn-TakingabstractThis work explores the timing of very short utterances in conversations, as well as the effects of excluding intervals adjacent to such utterances from distributions of betweenspeaker interval durations. The results show that very short utterances are more precisely timed to the preceding utterance than longer utterances in terms of a smaller variance and a larger proportion of no-gap-no-overlaps. Excluding intervals adjacent to very short utterances furthermore results in measures of central tendency closer to zero (i.e. no-gap-no-overlaps) as well as larger variance (i.e. relatively longer gaps and overlaps). Index Terms: Human speech production, Prosody 1. Mattias Heldner, Jens Edlund, Anna Hjalmarsson, Kornel Laskowski |
INTERSPEECH | 2 |
| 2011 | Incremental Learning and Forgetting in Stochastic Turn-Taking ModelsabstractWe present a computational framework for stochastically modeling dyad interaction chronograms. The framework’s most novel feature is the capacity for incremental learning and forgetting. To showcase its flexibility, we design experiments answering four concrete questions about the systematics of spoken interaction. The results show that: (1) individuals are clearly affected by one another; (2) there is individual variation in interaction strategy; (3) strategies wander in time rather than converge; and (4) individuals exhibit similarity with their interlocutors. We expect the proposed framework to be capable of answering many such questions with little additional effort. Index Terms: interaction, chronogram modeling, turn-taking, incremental learning. Kornel Laskowski, Jens Edlund, Mattias Heldner |
INTERSPEECH | 2 |
| 2011 | The Mona Lisa Gaze Effect as an Objective Metric for Perceived Cospatiality
Jens Edlund, Samer Al Moubayed, Jonas Beskow |
IVA | 1 |
| 2010 | Pitch similarity in the vicinity of backchannelsabstractDynamic modeling of spoken dialogue seeks to capture how interlocutors change their speech over the course of a conversation. Much work has focused on how speakers adapt or entrain to different aspects of one another’s speaking style. In this paper we focus on local aspects of this adaptation. We investigate the relationship between backchannels and the interlocutor utterances that precede them with respect to pitch. We demonstrate that the pitch of backchannels is more similar to the immediately preceding utterance than non-backchannels. This inter-speaker pitch relationship captures the same distinctions as more cumbersome intra-speaker relations, and supports the intuition that, in terms of pitch, such similarity may be one of the mechanisms by which backchannels are rendered ’unobtrusive’. Mattias Heldner, Jens Edlund, Julia Hirschberg |
INTERSPEECH | 2 |
| 2010 | Spontal: A Swedish Spontaneous Dialogue Corpus of Audio, Video and Motion Capture
Jens Edlund, Jonas Beskow, Kjell Elenius, Kahl Hellmer, Sofia Strömbergsson, David House |
LREC | 1 |
| 2010 | A Snack Implementation and Tcl/Tk Interface to the Fundamental Frequency Variation Spectrum Algorithm
Kornel Laskowski, Jens Edlund |
LREC | 2 |
| 2010 | Spontal-N: A Corpus of Interactional Spoken Norwegian
Rein Ove Sikveland, Anton Öttl, Ingunn Amdal, Mirjam Ernestus, Torbjørn Svendsen, Jens Edlund |
LREC | 6 |
| 2009 | The MonAMI reminder: a spoken dialogue system for face-to-face interactionabstractWe describe the MonAMI Reminder, a multimodal spoken dialogue system which can assist elderly and disabled people in organising and initiating their daily activities. Based on deep interviews with potential users, we have designed a calendar and reminder application which uses an innovative mix of an embodied conversational agent, digital pen and paper, and the web to meet the needs of those users as well as the current constraints of speech technology. We also explore the use of head pose tracking for interaction and attention control in human-computer face-to-face interaction. Jonas Beskow, Jens Edlund, Björn Granström, Joakim Gustafson, Gabriel Skantze, Helena Tobiasson |
INTERSPEECH | 2 |
| 2009 | Pause and gap length in face-to-face interactionabstractIt has long been noted that conversational partners tend to exhibit increasingly similar pitch, intensity, and timing behavior over the course of a conversation. However, the metrics developed to measure this similarity to date have generally failed to capture the dynamic temporal aspects of this process. In this paper, we propose new approaches to measuring interlocutor similarity in spoken dialogue. define similarity in terms of convergence and synchrony and propose approaches to capture these, illustrating our techniques on gap and pause production in Swedish spontaneous dialogues. Jens Edlund, Mattias Heldner, Julia Hirschberg |
INTERSPEECH | 1 |
| 2009 | A general-purpose 32 ms prosodic vector for hidden Markov modelingabstractProsody plays a central role in conversation, making it impor-tant for speech technologies to model. Unfortunately, the ap-plication of standard modeling techniques to the acoustics of prosody has been hindered by difficulties in modeling intona-tion. In this work, we explore the suitability of the recently introduced fundamental frequency variation (FFV) spectrum as a candidate general representation of tone. Experiments on 4 tasks demonstrate that FFV features are complimentary to other acoustic measures of prosody and that hidden Markov models offer a suitable modeling paradigm. Proposed improvements yield a 35 % relative decrease in error on unseen data and simul-taneously reduce time complexity by a factor of five. The result-ing representation is sufficiently mature for general deployment in a broad range of automatic speech processing applications. 1. Kornel Laskowski, Mattias Heldner, Jens Edlund |
INTERSPEECH | 3 |
| 2008 | An instantaneous vector representation of delta pitch for speaker-change prediction in conversational dialogue systemsabstractAs spoken dialogue systems become deployed in increasingly complex domains, they face rising demands on the naturalness of interaction. We focus on system responsiveness, aiming to mimic human-like dialogue flow control by predicting speaker changes as observed in real human-human conversations. We derive an instantaneous vector representation of pitch variation and show that it is amenable to standard acoustic modeling techniques. Using a small amount of automatically labeled data, we train models which significantly outperform current state-of-the-art pause-only systems, and replicate to within 1% absolute the performance of our previously published hand-crafted baseline. The new system additionally offers scope for run-time control over the precision or recall of locations at which to speak. Kornel Laskowski, Jens Edlund, Mattias Heldner |
ICASSP | 2 |
| 2008 | Innovative interfaces in MonAMI: the reminderabstractThis demo paper presents an early version of the Reminder, a prototype ECA developed in the European project MonAMI, which aims at "mainstreaming accessibility in consumer goods and services, using advanced technologies to ensure equal access, independent living and participation for all". The Reminder helps users to plan activities and to remember what to do. The prototype merges mobile ECA technology with other, existing technologies: Google Calendar and a digital pen and paper. The solution allows users to continue using a paper calendar in the manner they are used to, whilst the ECA provides notifications on what has been written in the calendar. Users may ask questions such as "When was I supposed to meet Sara?" or "What's my schedule today?" Jonas Beskow, Jens Edlund, Teodore Gjermani, Björn Granström, Joakim Gustafson, Oskar Jonsson, Gabriel Skantze, Helena Tobiasson |
ICMI | 2 |
| 2008 | Towards human-like spoken dialogue systems
Jens Edlund, Joakim Gustafson, Mattias Heldner, Anna Hjalmarsson |
Speech Commun. | 1 |
| 2007 | Pushy versus meek - using avatars to influence turn-taking behaviourabstractThe flow of spoken interaction between human interlocutors is a widely studied topic. Amongst other things, studies have shown that we use a number of facial gestures to improve this flow – for example to control the taking of turns. This type of gestures ought to be useful in systems where an animated talking head is used, be they systems for computer mediated human-human dialogue or spoken dialogue systems, where the computer itself uses speech to interact with users. In this article, we show that a small set of simple interaction control gestures and a simple model of interaction can be used to influence users ’ behaviour in an unobtrusive manner. The results imply that such a model may improve the flow of computer mediated interaction between humans under adverse circumstances, such as network latency, or to create more human-like spoken human-computer interaction. 1. Jens Edlund, Jonas Beskow |
INTERSPEECH | 1 |
| 2006 | /nailon/ - software for online analysis of prosodyabstractThis paper presents /nailon/ – a software package for online real-time prosodic analysis that captures a number of prosodic features relevant for interaction control in spoken dialogue systems. The current implementation captures silence durations; voicing, intensity, and pitch; pseudo-syllable durations; and intonation patterns. The paper provides detailed information on how this is achieved. As an example application of /nailon/, we demonstrate how it is used to improve the efficiency of identifying relevant places at which a machine can legitimately begin to talk to a human interlocutor, as well as to shorten system response times. Index Terms: automatic extraction of prosodic features, dialogue systems, interaction control Jens Edlund, Mattias Heldner |
INTERSPEECH | 1 |
| 2006 | User responses to prosodic variation in fragmentary grounding utterances in dialogabstractIn a previous study we demonstrated that subjects could use prosodic features (primarily peak height and alignment) to make different interpretations of synthesized fragmentary grounding utterances. In the present study we test the hypothesis that subjects also change their behavior accordingly in a human-computer dialog setting. We report on an experiment in which subjects participate in a color-naming task in a Wizard-of-Oz controlled human-computer dialog in Swedish. The results show that two annotators were able to categorize the subjects ’ responses based on pragmatic meaning. Moreover, the subjects ’ response times differed significantly, depending on the prosodic features of the grounding fragment spoken by the system. Index terms: dialog systems, prosody, error handling 1. Gabriel Skantze, David House, Jens Edlund |
INTERSPEECH | 3 |
| 2005 | The effects of prosodic features on the interpretation of clarification ellipsesabstractIn this paper, the effects of prosodic features on the interpretation of elliptical clarification requests in dialogue are studied. An experiment is presented where subjects were asked to listen to ... Jens Edlund, David House, Gabriel Skantze |
INTERSPEECH | 1 |
| 2004 | Higgins - a spoken dialogue system for investigating error handling techniquesabstractIn this paper, an overview of the Higgins project and the research within the project is presented. The project incorporates studies of error handling for spoken dialogue systems on several levels, from processing to dialogue level. A domain in which a range of different error types can be studied has been chosen: pedestrian navigation and guiding. Several data collections within Higgins have been analysed along with data from Higgins' predecessor, the AdApt system. The error handling research issues in the project are presented in light of these analyses. Jens Edlund, Gabriel Skantze, Rolf Carlson |
INTERSPEECH | 1 |
| 2002 | Specification and realisation of multimodal output in dialogue systemsabstractWe present a high level formalism for specifying verbal and nonverbal output from a multimodal dialogue system. The output specification is XML-based and provides information about communicative functions of the output without detailing the realisation of these functions. The specification can be used to control an animated character that uses speech and gestures. We give examples from an implementation in a multimodal spoken dialogue system, and describe how facial gestures are implemented in a 3D-animated talking agent within this system. Jonas Beskow, Jens Edlund, Magnus Nordstrand |
INTERSPEECH | 2 |
| 2000 | Adapt - a multimodal conversational dialogue system in an apartment domainabstractA general overview of the AdApt project and the research that is performed within the project is presented. In this project various aspects of human-computer interaction in a multimodal conversational dialogue systems are investigated. The project will also include studies on the integration of user/system/dialogue dependent speech recognition and multimodal speech synthesis. A domain in which multimodal interaction is highly useful has been chosen, namely, finding available apartments in Stockholm. A Wizard-of-Oz data collection within this domain is also described. 1. Joakim Gustafson, Linda Bell, Jonas Beskow, Johan Boye, Rolf Carlson, Jens Edlund, Björn Granström, David House, Mats Wirén |
INTERSPEECH | 6 |