EDBT 2026 Demo / reviewers in the wild / expert
Volker Dellwo
dblp:91/5042
· DBLP profile ↗
27ranked-venue papers
4as first author
8since 2021 · last 2026
0000-0002-8494-6025ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 24 · 4 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 24 · 4 first-author · 7 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Micro-Expression-Aware Avatar Fingerprinting via Inter-Frame Feature Differencing
Masoumeh Chapariniya, Jean-Marc Odobez, Volker Dellwo, Teodora Vukovic |
FG | 3 |
| 2025 | Voxplorer: Voice data exploration and projection in an interactive dashboard
Alessandro De Luca 0005, Srikanth R. Madikeri, Volker Dellwo |
INTERSPEECH | 3 |
| 2024 | Temporal Co-Registration of Simultaneous Electromagnetic Articulography and Electroencephalography for Precise Articulatory and Neural Data AlignmentabstractThis study presents a temporal co-registration method combining electromagnetic articulography (EMA) and electroencephalography (EEG) to capture the neural planning and execution phases of speech with high precision. Traditional EEG alignment based on acoustic vocal onset is often inaccurate due to the variable lag between articulatory and acoustic onsets. Our approach synchronizes EMA-derived speech kinematics with EEG data, addressing these challenges. We also examined the interaction between EMA and EEG systems, focusing on the integrity of EMA signals in the presence of EEG equipment and the electromagnetic influence of EMA on EEG signal quality. The method achieved a mean alignment delay of 2.7 ms (SD = 0.4 ms), enabling detailed analysis of pre-articulatory brain activities. Additionally, our evaluations confirmed the robustness of EMA signals and EEG event-related potentials, supporting the method's precision, feasibility, and reliability for speech planning research. Daniel Friedrichs, Monica Lancheros, Sam Kirkham, Lei He 0021, Clemens Lutz, Volker Dellwo, Steven Moran |
INTERSPEECH | 7 |
| 2024 | NumberLie: a game-based experiment to understand the acoustics of deception and truthfulnessabstractTo record clearly defined natural deceptive speech with precise knowledge of the ground truth and immediate consequences for the lying subject we present here the NumberLie game. The NumberLie design enables simultaneous and isolated audio recording of five players in our state-of-the-art laboratory, or adapted to any number of players in an online setting, playing against each other in a number-based game revolving around deception and trustworthiness. We describe the technical solutions employed to guarantee precise labelling of statements as truths or lies and immediate consequences to each interaction, backed by a performance-based financial reward to motivate participants. The design is easily manipulated to tailor to specific research questions, maintaining constant or eliminating completely additional sources of variability. Alessandro De Luca 0005, Volker Dellwo |
INTERSPEECH | 3 |
| 2024 | Deep neural networks for automatic speaker recognition do not learn supra-segmental temporal featuresabstractWhile deep neural networks have shown impressive results in automatic speaker recognition and related tasks, it is dissatisfactory how little is understood about what exactly is responsible for these results. Part of the success has been attributed in prior work to their capability to model supra-segmental temporal information (SST), i.e., learn rhythmic-prosodic characteristics of speech in addition to spectral features. In this paper, we (i) present and apply a novel test to quantify to what extent the performance of state-of-the-art neural networks for speaker recognition can be explained by modeling SST; and (ii) present several means to force respective nets to focus more on SST and evaluate their merits. We find that a variety of CNN- and RNN-based neural network architectures for speaker recognition do not model SST to any sufficient degree, even when forced. The results provide a highly relevant basis for impactful future research into better exploitation of the full speech signal and give insights into the inner workings of such networks, enhancing explainability of deep learning for speech technologies. Daniel Neururer, Volker Dellwo, Thilo Stadelmann |
Pattern Recognit. Lett. | 2 |
| 2024 | Forms, factors and functions of phonetic convergence: Editorial
Elisa Pellegrino, Volker Dellwo, Jennifer S. Pardo, Bernd Möbius |
Speech Commun. | 2 |
| 2022 | Fundamental Frequency Variability over Time in Telephone InteractionsabstractSpeech signals contain substantial fundamental frequency (f0) variability. Even within a single utterance, speakers modify f0 to create different intonational patterns. Previous studies have identified markers of increased f0 variability, such as the introduction of a new topic or greetings, but these are limited in the scope of their analyses. In the present study, we investigate f0 variability over the course of a telephone conversation, with a focus on the initial and medial utterances within the exchange. We examined f0 standard deviation of each utterance in over 2000 telephone conversations from 509 American English speakers from the Switchboard corpus. Findings showed that on average, speakers exhibit more f0 variability in the opening compared to mid-conversation utterances. Further, findings suggest that the inclusion of a greeting word in an initial turn, e.g., "hello” or "hi”, corresponds to an increase in f0 standard deviation. These results suggest that speakers employed more variable f0 in the initial few turns of a telephone conversation. The interpretation of this finding is multifaceted and may be linked to several communicative goals, including the placement of identity markers in conversation or the attraction of attention, or the role of openings as boundary markers. Leah Bradshaw, Eleanor Chodroff, Lena A. Jäger, Volker Dellwo |
INTERSPEECH | 4 |
| 2022 | Idiosyncratic lingual articulation of American English /æ/ and /ɑ/ using network analysisabstractFormant dynamics are believed to reflect the characteristic articulatory behavior of a speaker. The present study aims to explore individual articulatory behaviors when producing American English /æ/ and /ɑ/. The two vowels differ in the degree of inherent spectral change, a property believed to carry information about vowel-phoneme identity, which may be reflected in the articulatory movements. We measured first and second formants together with tongue blade and dorsum trajectories from 20 speakers producing 330 words in citation forms. Using the network analysis, the relationships between acoustic and kinematic variables were revealed. In particular, between-speaker articulatory behaviors were most dissimilar in /ɑ/ which requires less inherent spectral change. Moreover, when networks of speakers with similar formant patterns were compared, it was revealed that their articulatory behaviors also shared similarities, although they seemed to be organized in characteristic ways. These findings contribute to our understanding of the complex interaction between articulatory variables and the acoustic outcome. Carolina Lins Machado, Volker Dellwo, Lei He 0021 |
INTERSPEECH | 2 |
| 2020 | Arabic Speech Rhythm Corpus: Read and Spontaneous Speaking StylesabstractDatabases for studying speech rhythm and tempo exist for numerous languages. The present corpus was built to allow comparisons between Arabic speech rhythm and other languages. 10 Egyptian speakers (gender-balanced) produced speech in two different speaking styles (read and spontaneous). The design of the reading task replicates the methodology used in the creation of BonnTempo corpus (BTC). During the spontaneous task, speakers talked freely for more than one minute about their daily life and/or their studies, then they described the directions to come to the university from a famous near location using a map as a visual stimulus. For corpus annotation, the database has been manually and automatically time-labeled, which makes it feasible to perform a quantitative analysis of the rhythm of Arabic in both Modern Standard Arabic (MSA) and Egyptian dialect variety. The database serves as a phonetic resource, which allows researchers to examine various aspects of Arabic supra-segmental features and it can be used for forensic phonetic research, for comparison of different speakers, analyzing variability in different speaking styles, and automatic speech and speaker recognition. Omnia Ibrahim, Homa Asadi, Eman Kassem, Volker Dellwo |
LREC | 4 |
| 2019 | Fundamental Frequency Accommodation in Multi-Party Human-Robot Game Interactions: The Effect of Winning or LosingabstractIn human-human interactions, the situational context plays a large role in the degree of speakers’ accommodation. In this paper, we investigate whether the degree of accommodation in a human-robot computer game is affected by (a) the duration of the interaction and (b) the success of the players in the game. 30 teams of two players played two card games with a conversational robot in which they had to find a correct order of five cards. After game 1, the players received the result of the game on a success scale from 1 (lowest success) to 5 (highest). Speakers’ fo accommodation was measured as the Euclidean distance between the human speakers and each human and the robot. Results revealed that (a) the duration of the game had no influence on the degree of fo accommodation and (b) the result of Game 1 correlated with the degree of fo accommodation in Game 2 (higher success equals lower Euclidean distance). We argue that game success is most likely considered as a sign of the success of players’ cooperation during the discussion, which leads to a higher accommodation behavior in speech. Omnia Ibrahim, Gabriel Skantze, Sabine Stoll, Volker Dellwo |
INTERSPEECH | 4 |
| 2019 | Formant Pattern and Spectral Shape Ambiguity of Vowel Sounds, and Related Phenomena of Vowel Acoustics - Exemplary Evidence
Dieter Maurer, Heidy Suter, Christian d'Heureuse, Volker Dellwo |
INTERSPEECH | 4 |
| 2019 | Evaluation of VOCALISE under conditions reflecting those of a real forensic voice comparison case (forensic_eval_01)
Finnian Kelly, Andrea Fröhlich, Volker Dellwo, Oscar Forth, Samuel Kent, Anil Alexander |
Speech Commun. | 3 |
| 2018 | Influences of Fundamental Oscillation on Speaker Identification in Vocalic Utterances by Humans and ComputersabstractWe tested the influence of fundamental oscillation (fo) on human and machine speaker recognition performance in vocalic test utterances. In experiment I, we trained a Gaussian-Mixture model on 15 speakers (80 multi-word utterances each) and tested it with sustained vowel utterances (/a:/, /i:/ and /u:/) under six fo conditions, three changing (fall, rise, fall-rise) and three steady-state (high, mid, low). Results revealed better performance for the steady-state compared to the changing conditions and within the steady-state condition, performance was poorest for high fo. In experiment II, we tested 9 human listeners on a subset of 4 speakers from experiment I. They went through two training tasks (training 1: multi-word utterances; training 2: words). In the test, they recognized speakers based on the same vocalic utterances as in experiment I (for these 4 speakers). Results showed that performance was about equally high for the changing and steady-state vowels, however, in the steady-state condition performance was best for high fo vowels. The experiments suggest that (a) fo has an influence on the strength of speaker specific characteristics in vowels and (b) humans - compared to machines - pay attention to different acoustic information in vocalic utterances for speaker recognition. Volker Dellwo, Thayabaran Kathiresan, Elisa Pellegrino, Lei He 0021, Sandra Schwab, Dieter Maurer |
INTERSPEECH | 1 |
| 2018 | The Zurich Corpus of Vowel and Voice Quality, Version 1.0abstractExisting databases of isolated vowel sounds or vowel sounds embedded in consonantal context generally document only limited variation of basic production parameters. Thus, concerning the possible variation range of vowel and voice quality-related sound characteristics, there is a lack of broad phenomenological and descriptive references that allow for a comprehensive understanding of vowel acoustics and for an evaluation of the extent to which corresponding existing approaches and models can be generalised. In order to contribute to the building up of such references, a novel database of vowel sounds that exceeds any existing collection by size and diversity of vocalic characteristicsis presented here, comprised of c. 34600 utterances of 70 speakers (46 nonprofessional speakers, children, women and men, and 24 professional actors/actresses and singers of straight theatre, contemporary singing, and European classical singing). The database focuses on sounds of the long Standard German vowels /i–y–e–ø–ɛ–a–o–u/ produced with varying basic production parameters such as phonation type, vocal effort, fundamental frequency, vowel context and speaking or singing style. In addition, a read text and, for professionals, songs are also included. The database is accessible for scientific use, and further extensions are in progress. Dieter Maurer, Christian d'Heureuse, Heidy Suter, Volker Dellwo, Daniel Friedrichs, Thayabaran Kathiresan |
INTERSPEECH | 4 |
| 2017 | Listeners use temporal information to identify French- and English-accented speech
Marie-José Kolly, Philippe Boula de Mareüil, Adrian Leemann, Volker Dellwo |
Speech Commun. | 4 |
| 2016 | A Praat-Based Algorithm to Extract the Amplitude Envelope and Temporal Fine Structure Using the Hilbert TransformabstractA speech signal can be viewed as a high frequency carrier signal containing the temporal fine structure (TFS) that is modulated by a low frequency envelope (ENV). A widely used method to decompose a speech signal into the TFS and ENV is the Hilbert transform. Although this method has been available for about one century and is widely applied in various kinds of speech processing tasks (e.g. speech chimeras), there are only very few speech processing packages that contain readily available functions for the Hilbert transform, and there is very little textbook type literature tailored for speech scientists to explain the processes behind the transform. With this paper we provide the code for carrying out the Hilbert operation to obtain the TFS and ENV in the widely used speech processing software Praat, and explain the basics of the procedure. To verify our code, we compare the Hilbert transform in Praat with a widely applied function for the same purpose in MATLAB (“hilbert(...)”). We can confirm that both methods arrive at identical outputs. Lei He 0021, Volker Dellwo |
INTERSPEECH | 2 |
| 2015 | Stable and unstable intervals as a basic segmentation procedure of the speech signalabstractThe concept of acoustically stable and unstable intervals to structure continuous speech is introduced. We present a method to compute stable intervals efficiently and reliably as a bottom-up approach at an early processing stage. We argue that such intervals stand in close relation to the rhythm of speech as they contribute to the overall temporal organization of the speech production process and the acoustic signal (stable intervals = intervals of reduced movement of certain articulators; unstable intervals = intervals of enhanced movement of certain articulators). To test the relationship of stability intervals with speech rhythm we investigated the between-speaker variability of stable and unstable intervals in the TEVOID corpus. Results revealed that significant between-speaker variability exists. We hypothesize from our findings that the basic segmentation of speech into stable and unstable intervals is a process that might play a role in human perception and processing of speech. Ulrike Glavitsch, Lei He 0021, Volker Dellwo |
INTERSPEECH | 3 |
| 2015 | Voice Äpp: a mobile app for crowdsourcing Swiss German dialect data
Adrian Leemann, Marie-José Kolly, Jean-Philippe Goldman, Volker Dellwo, Ingrid Hove, Ibrahim Almajai, Sarah Grimm, Sylvain Robert, Daniel Wanitsch |
INTERSPEECH | 4 |
| 2015 | Swiss graphogame: concept and design presentation of a computerised reading intervention for children with high risk for poor reading outcomes
Martina Röthlisberger, Iliana I. Karipidis, Georgette Pleisch, Volker Dellwo, Ulla Richardson, Silvia Brem |
INTERSPEECH | 4 |
| 2014 | Rhythmic variability between some asian languages: results from an automatic analysis of temporal characteristicsabstractThe rhythmic organization of speech can vary between languages.In the present research we studied rhythmic variability between Mandarin, Cantonese and Thai using automatically retrieved prosodic temporal characteristics from read speech.We measured the variability of intervals between amplitude peaks in the amplitude envelope (<10 Hz) and the durational characteristics of intervals with and without glottal activity (voiced and unvoiced intervals) in speech.Results for between language comparisons revealed significant differences between languages in both amplitude peak interval variability and voiced-voiceless interval durational characteristics.Results are discussed in connection with language specific phonotactic/phonological properties and hypotheses about the perceptual significance of the acoustic measurements in terms of speech rhythm. Volker Dellwo, Peggy Mok, Mathias Jenny |
INTERSPEECH | 1 |
| 2014 | Speaker idiosyncratic variability of intensity across syllablesabstractThis study explored speaker idiosyncrasy by measuring the syllabic intensity variability in the speech signal.Sixteen speakers of the TEVOID corpus, each producing 256 read sentences, were analyzed.Characteristics of intensity variability (average or peak) between syllables were measured either holistically (standard deviation of intensity changes between syllables) or locally (pairwise variability indices of intensity changes between syllables).The results indicated significant effects of the speakers in all the metrics, suggesting a potential application of the methods for speaker recognition, and in particular for forensic speaker comparison. Lei He 0021, Volker Dellwo |
INTERSPEECH | 2 |
| 2014 | Foreign accent recognition based on temporal information contained in lowpass-filtered speechabstractCan the foreign accent of a speaker be recognized based on suprasegmental temporal information?For a perception experiment we created stimuli based on German sentences read by six French and six English speakers.These foreignaccented sentences were manipulated by (1) applying a lowpass filter with a cutoff frequency of 300 Hz and (2) applying the same lowpass filter and monotonizing F0.In a between-subject 2AFC perception experiment we tested the accent recognition ability of 15 Swiss German listeners per signal manipulation condition.The results showed that speakers' native language could be recognized above chance in both conditions.However, listeners obtained significantly lower recognition scores in the monotonized condition.Furthermore, higher recognition scores were obtained for French-accented speech in the monotonized condition, a result that is discussed in light of research on speech rhythm.We further report an effect for speaker within each accent group.The results suggest that suprasegmental temporal information allows for foreign accent recognition to some degree. Marie-José Kolly, Adrian Leemann, Volker Dellwo |
INTERSPEECH | 3 |
| 2014 | Intelligibility of high-pitched vowel sounds in the singing and speaking of a female Cantonese opera singer
Dieter Maurer, Peggy Mok, Daniel Friedrichs, Volker Dellwo |
INTERSPEECH | 4 |
| 2014 | A Crowdsourcing Smartphone Application for Swiss German: Putting Language Documentation in the Hands of the Users
Jean-Philippe Goldman, Adrian Leemann, Marie-José Kolly, Ingrid Hove, Ibrahim Almajai, Volker Dellwo, Steven Moran |
LREC | 6 |
| 2012 | Speaker idiosyncratic rhythmic features in the speech signalabstractSpeakers' voices are to a high degree individual.In the present paper we report about an ongoing research project in which we study how temporal characteristics of human speech (e.g.segmental or prosodic timing patterns, speech rhythmic characteristics and durational patterns of voicing) contribute to speaker individuality.We report about the creation of the TEVOID-Corpus (Temporal Voice Idiosyncrasy) that we are currently creating in our lab at Zurich University.8 speakers producing 16 spontaneous sentences each are currently in the database which is rapidly growing.The paper gives an overview of the general ideas for the data collection and first results showing that there are significant rhythmic differences (%V, %VO, VarcoPeak) in spontaneously produced sentences between speakers of Zurich German. Volker Dellwo, Adrian Leemann, Marie-José Kolly |
INTERSPEECH | 1 |
| 2010 | Expectations for discourse genre identification: a prosodic studyabstractSpeech can be divided into discourse genres based on the contextual environment it occurs in (e.g.political speech, sport commentary speech, etc.).The present study investigated whether listeners can distinguish between speech from different discourse genres on the basis of acoustic prosodic cues only 1 .In a perception experiment with delexicalized speech 70 listeners with varying experience in French (native speakers, nonnative speakers, and non-speakers) were asked to identify four different types of discourse genres (church service, political, journal, and sport commentary).Results revealed a fair identification ability with a significant increase in performance with increasing experience in French.Identification confusion was used to cluster discourse genres according to their perceptual similarity. Nicolas Obin, Volker Dellwo, Anne Lacheret, Xavier Rodet |
INTERSPEECH | 2 |
| 2004 | Bonntempo-corpus and bonntempo-tools: a database for the study of speech rhythm and rateabstractWork is currently being carried out on a speech database
constructed in order to study speech rhythm in connection with speech rate. The database, BonnTempo-Corpus, and the Praat based analysis tools, BonnTempo-Tools, are a powerful
instrument for examining various aspects of recently proposed rhythm measures (e.g. %V, C, nPVI, rPVI, etc.) in relation to speech rate among a wide range of languages and speakers.
First observations pose new problems on traditionally not well classifiable languages like Czech. Volker Dellwo, Bianca Aschenberner, Petra Wagner, Jana Dancovicova, Ingmar Steiner |
INTERSPEECH | 1 |