VLDB 2026 Research / reviewers in the wild / expert
Olov Engwall
dblp:38/6295
· DBLP profile ↗
42ranked-venue papers
16as first author
6since 2021 · last 2026
0000-0003-4532-014XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 32 · 14 first-author · 1 since 2021Artificial intelligence and machine learning · 31 · 13 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 8 · 2 first-author · 5 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Co-designing AI-mediated Paired Reading for Children in Swedish ClassroomsabstractReading proficiency is foundational for children’s future opportunities, yet many struggle to develop adequate literacy skills due to limited instructional resources and increasingly diverse classrooms. While one-to-one tutoring is highly effective, it is difficult to scale. We present AI paired reading for Swedish, co-designed iteratively with teachers and children to complement classroom practice and promote reading engagement. The system combines automatic speech recognition, text-to-speech, and large language models to enable children to read with a virtual peer. A mixed-methods classroom study with 43 second-grade pupils showed high engagement, sustained motivation and voluntary re-use. However, the study also revealed key design tensions around feedback clarity, difficulty adaptation, and the role of AI as a peer versus evaluator. We contribute empirically grounded design insights for AI-supported reading with children and discuss implications for integrating AI into classroom literacy practices. Olga Viberg, Patrik Larsson, Elias Hedlin, Olov Engwall |
IDC | 4 |
| 2024 | Conformity and Trust in Multi-party vs. Individual Human-Robot InteractionabstractIn this study, we explored how conformity and trust vary in adolescent students’ interactions with a social robot. Specifically, we compared how this was influenced by whether the participants had individual or multi-party interaction with robot and whether the robot was portrayed as an adult or a child through appearance and voice. Our experiment involved 75 Swedish middle school students participating in a card sorting game with the Furhat robot, where the objective was to discuss and reach an agreement on the card sequence. The data analysis focused firstly on the participants’ willingness to rearrange cards following the robot’s suggestions and secondly their post-session subjective trust in the robot’s advice. Results indicated that individuals interacting with the robot individually were more likely to conform to its suggestions than those interacting with it together with a peer. Individuals interacting alone with the robot also showed higher post-session trust levels than those in multi-party settings, indicating group size impacts robot trustworthiness perceptions. However, the robot’s perceived age did not affect the level of conformity. Exploratory analyses also showed that mutual understanding was lower in the multi-party setting, while the child robot condition improved user experience, highlighting the complex influence of group dynamics and robot portrayal on human-robot interactions in education. Alireza Mahmoudi Kamelabad, Olov Engwall, Gabriel Skantze |
IVA | 2 |
| 2022 | Shaping unbalanced multi-party interactions through adaptive robot backchannels
Ronald Cumbal, Daniel Alexander Kazzi, Vincent Winberg, Olov Engwall |
IVA | 4 |
| 2022 | Identification of Low-engaged Learners in Robot-led Second Language Conversations with AdultsabstractThe main aim of this study is to investigate if verbal, vocal, and facial information can be used to identify low-engaged second language learners in robot-led conversation practice. The experiments were performed on voice recordings and video data from 50 conversations, in which a robotic head talks with pairs of adult language learners using four different interaction strategies with varying robot-learner focus and initiative. It was found that these robot interaction strategies influenced learner activity and engagement. The verbal analysis indicated that learners with low activity rated the robot significantly lower on two out of four scales related to social competence. The acoustic vocal and video-based facial analysis, based on manual annotations or machine learning classification, both showed that learners with low engagement rated the robot’s social competencies consistently, and in several cases significantly, lower, and in addition rated the learning effectiveness lower. The agreement between manual and automatic identification of low-engaged learners based on voice recordings or face videos was further found to be adequate for future use. These experiments constitute a first step towards enabling adaption to learners’ activity and engagement through within- and between-strategy changes of the robot’s interaction with learners. Olov Engwall, Ronald Cumbal, José Lopes 0001, Mikael Ljung, Linnea Månsson |
ACM Trans. Hum. Robot Interact. | 1 |
| 2021 | Robot Gaze Can Mediate Participation Imbalance in Groups with Different Skill LevelsabstractMany small group activities, like working teams or study groups, have a high dependency on the skill of each group member. Differences in skill level among participants can affect not only the performance of a team but also influence the social interaction of its members. In these circumstances, an active member could balance individual participation without exerting direct pressure on specific members by using indirect means of communication, such as gaze behaviors. Similarly, in this study, we evaluate whether a social robot can balance the level of participation in a language skill-dependent game, played by a native speaker and a second language learner. In a between-subjects study (N = 72), we compared an adaptive robot gaze behavior, that was targeted to increase the level of contribution of the least active player, with a non-adaptive gaze behavior. Our results imply that, while overall levels of speech participation were influenced predominantly by personal traits of the participants, the robot's adaptive gaze behavior could shape the interaction among participants which lead to more even participation during the game. Sarah Gillet, Ronald Cumbal, André Pereira 0001, José Lopes 0001, Olov Engwall, Iolanda Leite |
HRI | 5 |
| 2021 | "You don't understand me!": Comparing ASR Results for L1 and L2 Speakers of SwedishabstractThe performance of Automatic Speech Recognition (ASR)systems has constantly increased in state-of-the-art develop-ment. However, performance tends to decrease considerably inmore challenging conditions (e.g., background noise, multiplespeaker social conversations) and with more atypical speakers(e.g., children, non-native speakers or people with speech dis-orders), which signifies that general improvements do not nec-essarily transfer to applications that rely on ASR, e.g., educa-tional software for younger students or language learners. Inthis study, we focus on the gap in performance between recog-nition results for native and non-native, read and spontaneous,Swedish utterances transcribed by different ASR services. Wecompare the recognition results using Word Error Rate and an-alyze the linguistic factors that may generate the observed tran-scription errors. Ronald Cumbal, Birger Moëll, José Lopes 0001, Olov Engwall |
Interspeech | 4 |
| 2020 | Detection of Listener Uncertainty in Robot-Led Second Language Conversation PracticeabstractUncertainty is a frequently occurring affective state that learners experience during the acquisition of a second language. This state can constitute both a learning opportunity and a source of learner frustration. An appropriate detection could therefore benefit the learning process by reducing cognitive instability. In this study, we use a dyadic practice conversation between an adult second-language learner and a social robot to elicit events of uncertainty through the manipulation of the robot's spoken utterances (increased lexical complexity or prosody modifications). The characteristics of these events are then used to analyze multi-party practice conversations between a robot and two learners. Classification models are trained with multimodal features from annotated events of listener (un)certainty. We report the performance of our models on different settings, (sub)turn segments and multimodal inputs. Ronald Cumbal, José Lopes 0001, Olov Engwall |
ICMI | 3 |
| 2019 | MRI-Based Vocal Tract Representations for the Three-Dimensional Finite Element Synthesis of DiphthongsabstractThe synthesis of diphthongs in three-dimensions (3D) involves the simulation of acoustic waves propagating through a complex 3D vocal tract geometry that deforms over time. Accurate 3D vocal tract geometries can be extracted from Magnetic Resonance Imaging (MRI), but due to long acquisition times, only static sounds can be currently studied with an adequate spatial resolution. In this work, 3D dynamic vocal tract representations are built to generate diphthongs, based on a set of cross-sections extracted from MRI-based vocal tract geometries of static vowel sounds. A diphthong can then be easily generated by interpolating the location, orientation and shape of these cross-sections, thus avoiding the interpolation of full 3D geometries. Two options are explored to extract the cross-sections. The first one is based on an adaptive grid (AG), which extracts the cross-sections perpendicular to the vocal tract midline, whereas the second one resorts to a semi-polar grid (SPG) strategy, which fixes the cross-section orientations. The finite element method (FEM) has been used to solve the mixed wave equation and synthesize diphthongs [Ai] and [Au] in the dynamic 3D vocal tracts. The outputs from a1D acoustic model based on the Transfer Matrix Method have also been included for comparison. The results show that the SPG and AG provide very close solutions in 3D, whereas significant differences are observed when using them in 1D. The SPG dynamic vocal tract representation is recommended for 3D simulations because it helps to prevent the collision of adjacent cross-sections. Marc Arnela, Saeed Dabbaghchian, Oriol Guasch, Olov Engwall |
IEEE ACM Trans. Audio Speech Lang. Process. | 4 |
| 2017 | A Semi-Polar Grid Strategy for the Three-Dimensional Finite Element Simulation of Vowel-Vowel SequencesabstractThree-dimensional computational acoustic models need very detailed 3D vocal tract geometries to generate high quality sounds. Static geometries can be obtained from Magnetic Resonance Imaging (MRI), but it is not currently possible to capture dynamic MRI-based geometries with sufficient spatial and time resolution. One possible solution consists in interpolating between static geometries, but this is a complex task. We instead propose herein to use a semi-polar grid to extract 2D cross-sections from the static 3D geometries, and then interpolate them to obtain the vocal tract dynamics. Other approaches such as the adaptive grid have also been explored. In this method, cross-sections are defined perpendicular to the vocal tract midline, as typically done in 1D to obtain the vocal tract area functions. However, intersections between adjacent cross-sections may occur during the interpolation process, especially when the vocal tract midline quickly changes its orientation. In contrast, the semi-polar grid prevents these intersections because the plane orientations are fixed over time. Finite element simulations of static vowels are first conducted, showing that 3D acoustic wave propagation is not significantly altered when the semi-polar grid is used instead of the adaptive grid. The vowel-vowel sequence [ɑi] is finally simulated to demonstrate the method. Marc Arnela, Saeed Dabbaghchian, Oriol Guasch, Olov Engwall |
INTERSPEECH | 4 |
| 2017 | Synthesis of VV Utterances from Muscle Activation to Sound with a 3D ModelabstractWe propose a method to automatically generate deformable 3D vocal tract geometries from the surrounding structures in a biomechanical model. This allows us to couple 3D biomechanics and acoustics simulations. The basis of the simulations is muscle activation trajectories in the biomechanical model, which move the articulators to the desired articulatory positions. The muscle activation trajectories for a vowel-vowel utterance are here defined through interpolation between the determined activations of the start and end vowel. The resulting articulatory trajectories of flesh points on the tongue surface and jaw are similar to corresponding trajectories measured using Electromagnetic Articulography, hence corroborating the validity of interpolating muscle activation. At each time step in the articulatory transition, a 3D vocal tract tube is created through a cavity extraction method based on first slicing the geometry of the articulators with a semi-polar grid to extract the vocal tract contour in each plane and then reconstructing the vocal tract through a smoothed 3D mesh-generation using the extracted contours. A finite element method applied to these changing 3D geometries simulates the acoustic wave propagation. We present the resulting acoustic pressure changes on the vocal tract boundary and the formant transitions for the utterance [Ai]. Saeed Dabbaghchian, Marc Arnela, Olov Engwall, Oriol Guasch |
INTERSPEECH | 3 |
| 2016 | Using a Biomechanical Model and Articulatory Data for the Numerical Production of VowelsabstractInternational audience Saeed Dabbaghchian, Marc Arnela, Olov Engwall, Oriol Guasch, Ian Stavness, Pierre Badin |
INTERSPEECH | 3 |
| 2013 | On mispronunciation analysis of individual foreign speakers using auditory periphery models
Christos Koniaris, Giampiero Salvi, Olov Engwall |
Speech Commun. | 3 |
| 2012 | Auditory and Dynamic Modeling Paradigms to Detect L2 MispronunciationsabstractThis paper expands our previous work on automatic pronunciation error detection that exploits knowledge from psychoacoustic auditory models. The new system has two additional important features, i.e., auditory and acoustic processing of the temporal cues of the speech signal, and classification feedback from a trained linear dynamic model. We also perform a pronunciation analysis by considering the task as a classification problem. Finally, we evaluate the proposed methods conducting a listening test on the same speech material and compare the judgment of the listeners and the methods. The automatic analysis based on spectro-temporal cues is shown to have the best agreement with the human evaluation, particularly with that of language teachers, and with previous plenary linguistic studies. Index Terms: L2 pronunciation error, auditory model, linear dynamic model, distortion measure, phoneme. Christos Koniaris, Olov Engwall, Giampiero Salvi |
INTERSPEECH | 2 |
| 2012 | Exploring the Predictability of Non-Unique Acoustic-to-Articulatory MappingsabstractThis paper explores statistical tools that help analyze the predictability in the acoustic-to-articulatory inversion of speech, using an Electromagnetic Articulography database of simultaneously recorded acoustic and articulatory data. Since it has been shown that speech acoustics can be mapped to non-unique articulatory modes, the variance of the articulatory parameters is not sufficient to understand the predictability of the inverse mapping. We, therefore, estimate an upper bound to the conditional entropy of the articulatory distribution. This provides a probabilistic estimate of the range of articulatory values (either over a continuum or over discrete non-unique regions) for a given acoustic vector in the database. The analysis is performed for different British/Scottish English consonants with respect to which articulators (lips, jaws or the tongue) are important for producing the phoneme. The paper shows that acoustic-articulatory mappings for the important articulators have a low upper bound on the entropy, but can still have discrete non-unique configurations. Gopal Ananthakrishnan, Olov Engwall, Daniel Neiberg |
IEEE Trans. Speech Audio Process. | 2 |
| 2011 | Resolving non-uniqueness in the acoustic-to-articulatory mappingabstractThis paper studies the role of non-uniqueness in the Acoustic-to-Articulatory Inversion. It is generally believed that applying continuity constraints to the estimates of the articulatory parameters can resolve the problem of non-uniqueness. This paper tries to find out whether all instances of non-uniqueness can be resolved using continuity constraints. The investigation reveals that applying continuity constraints provides the best estimate in roughly around 50 to 53% of the non-unique mappings. Roughly around 8 to 13% of the non-unique mappings are best estimated by choosing discontinuous paths along the hypothetical high probability estimates of articulatory trajectories. Gopal Ananthakrishnan, Olov Engwall |
ICASSP | 2 |
| 2011 | Perceptual differentiation modeling explains phoneme mispronunciation by non-native speakersabstractOne of the difficulties in second language (L2) learning is the weakness in discriminating between acoustic diversity within an L2 phoneme category and between different categories. In this paper, we describe a general method to quantitatively measure the perceptual difference between a group of native and individual nonnative speakers. Normally, this task includes subjective listening tests and/or a thorough linguistic study. We instead use a totally automated method based on a psycho-acoustic auditory model. For a certain phoneme class, we measure the similarity of the Euclidean space spanned by the power spectrum of a native speech signal and the Euclidean space spanned by the auditory model output. We do the same for a non-native speech signal. Comparing the two similarity measurements, we find problematic phonemes for a given speaker. To validate our method, we apply it to different groups of non-native speakers of various first language (L1) backgrounds. Our results are verified by the theoretical findings in literature obtained from linguistic studies. Christos Koniaris, Olov Engwall |
ICASSP | 2 |
| 2011 | Phoneme Level Non-Native Pronunciation Analysis by an Auditory Model-Based Native Assessment SchemeabstractWe introduce a general method for automatic diagnostic evaluation of the pronunciation of individual non-native speakers based on a model of the human auditory system trained with native data stimuli. For each phoneme class, the Euclidean geometry similarity between the native perceptual domain and the non-native speech power spectrum domain is measured. The problematic phonemes for a given second language speaker are found by comparing this measure to the Euclidean geometry similarity for the same phonemes produced by native speakers only. The method is applied to different groups of non-native speakers of various language backgrounds and the experimental results are in agreement with theoretical findings of linguistic studies. Christos Koniaris, Olov Engwall |
INTERSPEECH | 2 |
| 2011 | Mapping between acoustic and articulatory gestures
Gopal Ananthakrishnan, Olov Engwall |
Speech Commun. | 2 |
| 2010 | Predicting unseen articulations from multi-speaker articulatory modelsabstractIn order to study inter-speaker variability, this work aims to assess the generalization capabilities of data-based multi-speaker articulatory models. We use various three-mode factor analysis techniques to model the variations of midsagittal vocal tract contours obtained from MRI images for three French speakers articulating 73 vowels and consonants. Articulations of a given speaker for phonemes not present in the training set are then predicted by inversion of the models from measurements of these phonemes articulated by the other subjects. On the average, the prediction RMSE was 5.25 mm for tongue contours, and 3.3 mm for 2D midsagittal vocal tract distances. Besides, this study has established a methodology to determine the optimal number of factors for such models. Gopal Ananthakrishnan, Pierre Badin, Julián Andrés Valdés Vargas, Olov Engwall |
INTERSPEECH | 4 |
| 2009 | In search of non-uniqueness in the acoustic-to-articulatory mappingabstractThis paper explores the possibility and extent of non-uniqueness in the acoustic-to-articulatory inversion of speech, from a statistical point of view. It proposes a technique to estimate the non-uniqueness, based on finding peaks in the conditional probability function of the articulatory space. The paper corroborates the existence of non-uniqueness in a statistical sense, especially in stop consonants, nasals and fricatives. The relationship between the importance of the articulator position and nonuniqueness at each instance is also explored. Index Terms: acoustic-to-articulatory inversion, nonuniqueness, gaussian mixture modeling. Gopal Ananthakrishnan, Daniel Neiberg, Olov Engwall |
INTERSPEECH | 3 |
| 2009 | Are real tongue movements easier to speech read than synthesized?abstractSpeech perception studies with augmented reality displays in talking heads have shown that tongue reading abilities are weak initially, but that subjects become able to extract some information from intra-oral visualizations after a short training session. In this study, we investigate how the nature of the tongue movements influences the results, by comparing synthetic rulebased and actual, measured movements. The subjects were significantly better at perceiving sentences accompanied by real movements, indicating that the current coarticulation model developed for facial movements is not optimal for the tongue. Index Terms: multimodal speech perception, augmented reality, visual speech synthesis 1. Olov Engwall, Preben Wik |
INTERSPEECH | 1 |
| 2009 | Audiovisual-to-articulatory inversion
Hedvig Kjellström, Olov Engwall |
Speech Commun. | 2 |
| 2008 | Can audio-visual instructions help learners improve their articulation? - an ultrasound study of short term changesabstractThis paper describes how seven French subjects change their pronunciation and articulation when practising Swedish words with a computer-animated virtual teacher. The teacher gives feedback on the ... Olov Engwall |
INTERSPEECH | 1 |
| 2008 | The acoustic to articulation mapping: non-linear or non-unique?abstractThis paper studies the hypothesis that the acoustic-toarticulatory mapping is non-unique, statistically. The distributions of the acoustic and articulatory spaces are obtained by fitting the data into a Gaussian Mixture Model. The kurtosis is used to measure the non-Gaussianity of the distributions and the Bhattacharya distance is used to find the difference between distributions of the acoustic vectors producing non-unique articulator configurations. It is found that stop consonants and alveolar fricatives are generally not only non-linear but also nonunique, while dental fricatives are found to be highly non-linear but fairly unique. Two more investigations are also discussed: the first is on how well the best possible piecewise linear regression is likely to perform, the second is on whether the dynamic constraints improve the ability to predict different articulatory regions corresponding to the same region in the acoustic space. Daniel Neiberg, Gopal Ananthakrishnan, Olov Engwall |
INTERSPEECH | 3 |
| 2008 | Can visualization of internal articulators support speech perception?abstractThis paper describes the contribution to speech perception given by animations of intra-oral articulations. 18 subjects were asked to identify the words in acoustically degraded sentences in three different presentation modes: acoustic signal only, audiovisual with a front view of a synthetic face and an audiovisual with both front face view and a side view, where tongue movements were visible by making parts of the cheek transparent. The augmented reality side-view did not help subjects perform better overall than with the front view only, but it seems to have been beneficial for the perception of palatal plosives, liquids and rhotics, especially in clusters. The results indicate that it cannot be expected that intra-oral animations support speech perception in general, but that information on some articulatory features can be extracted. Animations of tongue movements have hence more potential for use in computer-assisted pronunciation and perception training than as a communication aid for the hearingimpaired. Index Terms: talking head, speech perception, speech visualization, audiovisual speech, internal articulation Preben Wik, Olov Engwall |
INTERSPEECH | 2 |
| 2007 | Audio-visual phoneme classification for pronunciation training applicationsabstractWe present a method for audio-visual classification of Swedish phonemes, to be used in computer-assisted pronunciation training. The probabilistic kernel-based method is applied to the audio signal and/or either a principal or an independent component (PCA or ICA) representation of the mouth region in video images. We investigate which representation (PCA or ICA) that may be most suitable and the number of components required in the base, in order to be able to automatically detect pronunciation errors in Swedish from audio-visual input. Experiments performed on one speaker show that the visual information help avoiding classification errors that would lead to gravely erroneous feedback to the user; that it is better to perform phoneme classification on audio and video separately and then fuse the results, rather than combining them before classification; and that PCA outperforms ICA for fewer than 50 components. Index Terms: audiovisual phoneme classification, pronunciation error detection, PCA, ICA Hedvig Kjellström, Olov Engwall, Sherif M. Abdou, Olle Bälter |
INTERSPEECH | 2 |
| 2006 | Reconstructing tongue movements from audio and videoabstractThis paper presents an approach to articulatory inversion using audio and video of the user’s face, requiring no special markers. The video is stabilized with respect to the face, and the mouth region cropped out. The mouth image is projected into a learned independent component subspace to obtain a low-dimensional representation of the mouth appearance. The inversion problem is treated as one of regression; a non-linear regressor using relevance vector machines is trained with a dataset of simultaneous images of a subject’s face, acoustic features and positions of magnetic coils glued to the subjects’s tongue. The results show the benefit of using both cues for inversion. We envisage the inversion method to be part of a pronunciation training system with articulatory feedback. Index Terms: audio-visual to articulatory inversion. 1. Hedvig Kjellström, Olov Engwall, Olle Bälter |
INTERSPEECH | 2 |
| 2006 | Designing the user interface of the computer-based speech training system ARTUR based on early user testsabstractThis study has been performed in order to evaluate a prototype for the human – computer interface of a computer-based speech training aid named ARTUR. The main feature of the aid is that it can give suggestions on how to improve articulations. Two user groups were involved: three children aged 9 – 14 with extensive experience of speech training with therapists and computers, and three children aged 6, with little or no prior experience of computer-based speech training. All children had general language disorders. The study indicates that the present interface is usable without prior training or instructions, even for the younger children, but that more motivational factors should be introduced. The granularity of the mesh that classifies mispronunciations was satisfactory, but the flexibility and level of detail of the feedback should be developed further. Olov Engwall, Olle Bälter, Anne-Marie Öster, Hedvig Kjellström |
Behav. Inf. Technol. | 1 |
| 2005 | Wizard-of-Oz test of ARTUR: a computer-based speech training system with articulation correctionabstractThis study has been performed in order to test the human-machine interface of a computer-based speech training aid named ARTUR with the main feature that it can give suggestions on how to improve articulation. Two user groups were involved: three children aged 9-14 with extensive experience of speech training, and three children aged 6. All children had general language disorders.The study indicates that the present interface is usable without prior training or instructions, even for the younger children, although it needs some improvement to fit illiterate children. The granularity of the mesh that classifies mispronunciations was satisfactory, but can be developed further. Olle Bälter, Olov Engwall, Anne-Marie Öster, Hedvig Kjellström |
ASSETS | 2 |
| 2005 | Articulatory synthesis using corpus-based estimation of line spectrum pairsabstractAn attempt to define a new articulatory synthesis method, in which the speech signal is generated through a statistical estimation of its relation with articulatory parameters, is presented. A corpus containing acoustic material and simultaneous recordings of the tongue and facial movements was used to train and test the articulatory synthesis of VCV words and short sentences. Tongue and facial motion data, captured with electromagnetic articulography and three-dimensional optical motion tracking, respectively, define articulatory parameters of a talking head. These articulatory parameters are then used as estimators of the speech signal, represented by line spectrum pairs. The statistical link between the articulatory parameters and the speech signal was established using either linear estimation or artificial neural networks. The results show that the linear estimation was only enough to synthesize identifiable vowels, but not consonants, whereas the neural networks gave a perceptually better synthesis. 1. Olov Engwall |
INTERSPEECH | 1 |
| 2005 | Introducing visual cues in acoustic-to-articulatory inversionabstractThe contribution of facial measures in a statistical acoustic-toarticulatory inversion has been investigated. The tongue contour was estimated using a linear estimation from either acoustics or acoustics and facial measures. Measures of the lateral movement of lip corners and the vertical movement of the upper and lower lip and the jaw gave a substantial improvement over the audio-only case. It was further found that adding the corresponding articulatory measures that could be extracted from a profile view of the face; i.e. the protrusion of the lips, lip corners and the jaw, did not give any additional improvement of the inversion result. The present study hence suggests that audiovisual-to-articulatory inversion can as well be performed using front view monovision of the face, rather than stereovision of both the front and profile view. Olov Engwall |
INTERSPEECH | 1 |
| 2004 | Design strategies for a virtual language tutorabstractIn this paper we discuss work in progress on an interactive talking agent as a virtual language tutor in CALL applications. The ambition is to create a tutor that can be engaged in many aspects of language learning from detailed pronunciation to conversational training. Some of the crucial components of such a system is described. An initial implementation of a stress/quantity training scheme will be presented. Jonas Beskow, Olov Engwall, Björn Granström, Preben Wik |
INTERSPEECH | 2 |
| 2004 | Speaker adaptation of a three-dimensional tongue modelabstractMagnetic Resonance Images of nine subjects have been collected to determine scaling factors that can adapt a 3D tongue model to new subjects. The aim is to define few and simple measures that will allow for an automatic, but accurate, scaling of the model. The scaling should be automatic in order to be useful in an application for articulation training, in which the model must replicate the user’s articulators without involving the user in a complicated speaker adaptation. It should further be accurate enough to allow for correct acoustic-to-articulatory inversion. The evaluation shows that the defined scaling technique is able to estimate a tongue shape that was not included in the training with an accuracy of 1.5 mm in the midsagittal plane and 1.7 mm for the whole 3D tongue, based on four articulatory measures. Olov Engwall |
INTERSPEECH | 1 |
| 2004 | From real-time MRI to 3d tongue movementsabstractReal-time Magnetic Resonance Imaging (MRI) at 9 images/s of the midsagittal plane is used as input to a threedimensional tongue model, previously generated based on sustained articulations imaged with static MRI. The aim is two-fold, firstly to use articulatory inversion to extrapolate the midsagittal tongue movements to three-dimensional movements, secondly to determine the accuracy of the tongue model in replicating the real-time midsagittal tongue shapes. The evaluation of the inversion shows that the realtime midsagittal contour is reproduced with acceptable accuracy. This means that the 3D model can be used to represent real-time articulations, eventhough the artificially sustained articulations on which it was based were hyperarticulated and had a backward displacement of the tongue. 1. Olov Engwall |
INTERSPEECH | 1 |
| 2003 | Resynthesis of 3d tongue movements from facial dataabstractSimultaneous measurements of tongue and facial motion, using a combination of electromagnetic articulography (EMA) and optical motion tracking, are analysed to investigate the possibility to resynthesize the subject’s tongue movements with a parametrically controlled 3D model using the facial data only. The recorded material consists of 63 VCV words spoken by one Swedish subject. The tongue movements are resynthesized using a combination of a linear estimation to predict the tongue data from the face and an inversion procedure to determine the articulatory parameters of the model. Olov Engwall, Jonas Beskow |
INTERSPEECH | 1 |
| 2003 | Combining MRI, EMA and EPG measurements in a three-dimensional tongue model
Olov Engwall |
Speech Commun. | 1 |
| 2002 | Evaluation of a system for concatenative articulatory visual speech synthesisabstractA method for concatenative articulatory visual speech synthesis has been evaluated. The method consists in using concatenated units of articulatory parameter transitions from the middle of one phoneme to the middle of the next as input to a 3D parametric tongue model. The units were created by segmentation of the Electromagnetic articulography (EMA) measures in a database of 460 phonetically balanced sentences collected at the University of Edinburgh. The evaluation was made against the EMA database on which the movements were based and against X-ray films of three other speakers. The results show that the model replicates the natural movements globally, but that the rare units in the concatenation database may cause large differences between the synthesized and the natural utterance and that the tongue root and tongue tip movements are too restricted in the model. Olov Engwall |
INTERSPEECH | 1 |
| 2001 | Making the tongue model talk: merging MRI & EMA measurements
Olov Engwall |
INTERSPEECH | 1 |
| 2001 | Using linguopalatal contact patterns to tune a 3d tongue modelabstractThe six articulatory parameters of a three-dimensional tongue model were adjusted to replicate linguopalatal contact patterns measured with Electropalatography (EPG). The tongue model is based on artificially sustained articulations measured with MRI and the EPG data provides one possibility to tune the parameters to dynamic speech. A 3D model was generated of the palate and the electrode distribution, allowing the synthetic contact patterns to be calculated. The tongue parameters were then adjusted to minimise the deviation from the natural contact patterns. Substantial reduction of the false and missing electrode contacts was made in the tuning and the synthetic linguopalatal contact pattern is shown to replicate the total characteristics of the natural patterns rather well. The remaining error is often due to lateral asymmetry or centralto -edge contact variations. Olov Engwall |
INTERSPEECH | 1 |
| 2000 | Are static MRI measurements representative of dynamic speech? results from a comparative study using MRI, EPG and EMA
Olov Engwall |
INTERSPEECH | 1 |
| 2000 | A 3d tongue model based on MRI dataabstractA new three-dimensional tongue model has been developed within the KTH 3D vocal tract project using manually extracted tongue contours from MR Images of a reference subject producing 43 artificially sustained Swedish articulations. The six linear parameters jaw height, tongue body, tongue dorsum, tongue tip, tongue advance and tongue width were determined using an ordered linear factor analysis controlled by articulatory measures. 88 % of the variation in the midsagittal plane and 78 % of the overall sagittal variation was explained by the first five factors of the analysis. The six parameter model is able to reconstruct the modeled articulations in 3D with an overall RMS reconstruction error of 0.13 cm sagittally and 0.12 cm laterally, and it specifically handles lateral differences and the observed asymmetries in tongue shape. 1. Olov Engwall |
INTERSPEECH | 1 |
| 1999 | Modeling of the vocal tract in three dimensionsabstractThis paper describes the development of a threedimensional articulatory vocal tract model at KTH. The model represents vocal and nasal tract walls, lips, teeth and tongue as parameterised polygon surfaces. This allows the geometry of the vocal tract to be set with a small number of articulatory parameters. As the crosssectional areas in addition are given directly from the vocal tract geometry, the model is suitable for articulatory synthesis in 3D. The second field of application is pronunciation training, where the model can provide visual feedback to heating-impaired children and adult second language learners. A 3D model can improve both articulatory and visual speech synthesis as it provides information lacking in the 2D models traditionally used. Correctness of the model will increase with the amount of articulatory data incorporated, as exemplified by this paper's description of the method to improve the tongue model. Olov Engwall |
EUROSPEECH | 1 |