EDBT 2026 Demo / reviewers in the wild / expert
Barbara Schuppler
dblp:99/9229
· DBLP profile ↗
32ranked-venue papers
11as first author
15since 2021 · last 2025
0000-0003-4009-0832ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 27 · 9 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 26 · 9 first-author · 12 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Uncertainty prediction for prominence classification with chroma featuresabstractThis paper presents methods for prominence classification in conversational speech. Most existing tools rely on prosodic features extracted at syllable- or phone-level, performing well on read speech. This is not the case for conversational speech, where the quality of automatic segmentation is significantly worse. We introduce entropy-based chroma features, requiring only word-level segmentations. They perform equally well as a random forest classifier with prosodic features (requiring phone-level segmentation), with accuracies in the range of the human inter-rater agreement. We further use Bayesian deep learning to quantify the epistemic and aleatoric uncertainty of the prediction for prosodic and chroma features. Whereas the aleatoric uncertainty is, as expected, consistent with inter-rater agreement and similarly high for both feature sets, the epistemic uncertainty is lower for the classifier based on chroma features, indicating higher classification consistency across the corpus. Julian Linke, Sophie Steger, Philipp Steinwender, Gernot Kubin, Franz Pernkopf, Barbara Schuppler |
ICASSP | 6 |
| 2025 | Context is all you need? Low-resource conversational ASR profits from context, coming from the same or from the other speaker
Julian Linke, Jana Winkler, Barbara Schuppler |
INTERSPEECH | 3 |
| 2025 | Continuous prediction of backchannel timing for human-robot interaction
Michael Paierl, Martin Hagmüller, Barbara Schuppler |
INTERSPEECH | 3 |
| 2025 | What the Filler? Both ASR Systems and Humans Struggle More With Other Kinds of Disfluencies Than With Filler Particles
Saskia Wepner, Lucas Eckert, Gernot Kubin, Barbara Schuppler |
INTERSPEECH | 4 |
| 2025 | What's so complex about conversational speech? A comparison of HMM-based and transformer-based ASR architecturesabstractHighly performing speech recognition is important for more fluent human–machine interaction (e.g., dialogue systems). Modern ASR architectures achieve human-level recognition performance on read speech but still perform sub-par on conversational speech, which arguably is or, at least, will be instrumental for human–machine interaction. Understanding the factors behind this shortcoming of modern ASR systems may suggest directions for improving them. In this work, we compare the performances of HMM- vs. transformer-based ASR architectures on a corpus of Austrian German conversational speech. Specifically, we investigate how strongly utterance length, prosody, pronunciation, and utterance complexity as measured by perplexity affect different ASR architectures. Among other findings, we observe that single-word utterances – which are characteristic of conversational speech and constitute roughly 30% of the corpus – are recognized more accurately if their F0 contour is flat; for longer utterances, the effects of the F0 contour tend to be weaker. We further find that zero-shot systems require longer utterance lengths and are less robust to pronunciation variation, which indicates that pronunciation lexicons and fine-tuning on the respective corpus are essential ingredients for the successful recognition of conversational speech. Julian Linke, Bernhard C. Geiger, Gernot Kubin, Barbara Schuppler |
Comput. Speech Lang. | 4 |
| 2024 | On Disfluency and Non-lexical Sound Labeling for End-to-end Automatic Speech Recognition
Péter Mihajlik, Mate Kadar, Julian Linke, Barbara Schuppler, Katalin Mády |
INTERSPEECH | 5 |
| 2024 | An introduction to pluricentric languages in speech science and technologyabstractPluricentric languages are languages that are spoken in at least two countries where they have an official function and thus develop national varieties with specific linguistic and pragmatic features. Presently 43 languages have been identified as belonging to this category, for instance, English, Spanish, German, Bengali, Hindi and Urdu. This article forms an introduction to the special issue “Pluricentric Languages in Speech Science and Technology” by giving an overview of current challenges with respect to the development of speech and language resources, annotation and analysis tools, as well as speech technology services for pluricentric languages. The article discusses potential solutions that come from cross-fertilization: on the one hand, how phonetic and linguistic knowledge may contribute to advancements in speech technology, and on the other, how speech technology may facilitate phonetic and linguistic studies on pluricentric languages. In our discussion, we include the research methods and findings of the eight research articles of this special issue and point towards promising paths for future research in the field. Barbara Schuppler, Martine Adda-Decker, Catia Cucchiarini, Rudolf Muhr |
Speech Commun. | 1 |
| 2024 | The prosody of theme, rheme and focus in Egyptian Arabic: A quantitative investigation of tunes, configurations and speaker variabilityabstractThis paper investigates the prosody of sentences elicited in three Information Structure (IS) conditions: all-new, theme-rheme and rhematic focus-background. The sentences were produced by 18 speakers of Egyptian Arabic (EA). This is the first quantitative study to provide a comprehensive analysis of holistic f0 contours (by means of GAMM) and configurations of f0, duration and intensity (by means of FPCA) associated with the three IS conditions, both across and within speakers. A significant difference between focus-background and the other information structure conditions was found, but also strong inter-speaker variation in terms of strategies and the degree to which these strategies were applied. The results suggest that post-focus register lowering and the duration of the stressed syllables of the focused and the utterance-final word are more consistent cues to focus than a higher peak of the focus accent. In addition, some independence of duration and intensity from f0 could be identified. These results thus support the assumption that, when focus is marked prosodically in EA, it is marked by prominence. Nevertheless, the fact that a considerable number of EA speakers did not apply prosodic marking and the fact that prosodic focus marking was gradient rather than categorical suggest that EA does not have a fully conventionalized prosodic focus construction. Dina El Zarka, Anneliese Kelterer, Michele Gubian, Barbara Schuppler |
Speech Commun. | 4 |
| 2023 | Exploring Graph Theory Methods For the Analysis of Pronunciation Variation in Spontaneous Speech
Bernhard C. Geiger, Barbara Schuppler |
INTERSPEECH | 2 |
| 2023 | (Dis)agreement and Preference Structure are Reflected in Matching Along Distinct Acoustic-prosodic Features
Anneliese Kelterer, Margaret Zellers, Barbara Schuppler |
INTERSPEECH | 3 |
| 2023 | What do self-supervised speech representations encode? An analysis of languages, varieties, speaking styles and speakers
Julian Linke, Mate Kadar, Gergely Dosinszky, Péter Mihajlik, Gernot Kubin, Barbara Schuppler |
INTERSPEECH | 6 |
| 2022 | Homophone Disambiguation Profits from Durational Information
Barbara Schuppler, Emil Berger, Xenia Kogler, Franz Pernkopf |
INTERSPEECH | 1 |
| 2022 | Conversational Speech Recognition Needs Data? Experiments with Austrian GermanabstractConversational speech represents one of the most complex of automatic speech recognition (ASR) tasks owing to the high inter-speaker variation in both pronunciation and conversational dynamics. Such complexity is particularly sensitive to low-resourced (LR) scenarios. Recent developments in self-supervision have allowed such scenarios to take advantage of large amounts of otherwise unrelated data. In this study, we characterise an (LR) Austrian German conversational task. We begin with a non-pre-trained baseline and show that fine-tuning of a model pre-trained using self-supervision leads to improvements consistent with those in the literature; this extends to cases where a lexicon and language model are included. We also show that the advantage of pre-training indeed arises from the larger database rather than the self-supervision. Further, by use of a leave-one-conversation out technique, we demonstrate that robustness problems remain with respect to inter-speaker and inter-conversation variation. This serves to guide where future research might best be focused in light of the current state-of-the-art. Julian Linke, Philip N. Garner, Gernot Kubin, Barbara Schuppler |
LREC | 4 |
| 2022 | To laugh or not to laugh? The use of laughter to mark discourse structureabstractA number of cues, both linguistic and nonlinguistic, have been found to mark discourse structure in conversation.This paper investigates the role of laughter, one of the most encountered non-verbal vocalizations in human communication, in the signalling of turn boundaries.We employ a corpus of informal dyadic conversations to determine the likelihood of laughter at the end of speaker turns and to establish the potential role of laughter in discourse organization.Our results show that, on average, about 10% of the turns are marked by laughter, but also that the marking is subject to individual variation, as well as effects of other factors, such as the type of relationship between speakers.More importantly, we find that turn ends are twice more likely than transition relevance places to be marked by laughter, suggesting that, indeed, laughter plays a role in marking discourse structure. Bogdan Ludusan, Barbara Schuppler |
SIGDIAL | 2 |
| 2022 | An analysis of prosodic boundaries across speaking styles in two varieties of German
Bogdan Ludusan, Barbara Schuppler |
Speech Commun. | 2 |
| 2020 | An Analysis of Prosodic Prominence Cues to Information Structure in Egyptian Arabic
Dina El Zarka, Anneliese Kelterer, Barbara Schuppler |
INTERSPEECH | 3 |
| 2020 | Microprosodic Variability in Plosives in German and Austrian German
Margaret Zellers, Barbara Schuppler |
INTERSPEECH | 2 |
| 2020 | Towards Building an Automatic Transcription System for Language Documentation: Experiences from MuyuabstractSince at least half of the world’s 6000 plus languages will vanish during the 21st century, language documentation has become a rapidly growing field in linguistics. A fundamental challenge for language documentation is the ”transcription bottleneck”. Speech technology may deliver the decisive breakthrough for overcoming the transcription bottleneck. This paper presents first experiments from the development of ASR4LD, a new automatic speech recognition (ASR) based tool for language documentation (LD). The experiments are based on recordings from an ongoing documentation project for the endangered Muyu language in New Guinea. We compare phoneme recognition experiments with American English, Austrian German and Slovenian as source language and Muyu as target language. The Slovenian acoustic models achieve the by far best performance (43.71% PER) in comparison to 57.14% PER with American English, and 89.49% PER with Austrian German. Whereas part of the errors can be explained by phonetic variation, the recording mismatch poses a major problem. On the long term, ASR4LD will not only be an integral part of the ongoing documentation project of Muyu, but will be further developed in order to facilitate also the language documentation process of other language groups. Alexander Zahrer, Andrej Zgank, Barbara Schuppler |
LREC | 3 |
| 2019 | Acoustic Correlates of Phonation Type in Chichimec
Anneliese Kelterer, Barbara Schuppler |
INTERSPEECH | 2 |
| 2019 | Prosodic Effects on Plosive Duration in German and Austrian German
Barbara Schuppler, Margaret Zellers |
INTERSPEECH | 1 |
| 2019 | Acoustic Cues to Topic and Narrow Focus in Egyptian Arabic
Dina El Zarka, Barbara Schuppler, Francesco Cangemi |
INTERSPEECH | 2 |
| 2018 | On the use of acoustic features for automatic disambiguation of homophones in spontaneous German
Barbara Schuppler, Tobias Schrank |
Comput. Speech Lang. | 1 |
| 2017 | A corpus of read and conversational Austrian German
Barbara Schuppler, Martin Hagmüller, Alexander Zahrer |
Speech Commun. | 1 |
| 2015 | Automatic detection of uncertainty in spontaneous German dialogueabstractUncertainty is ubiquitous in natural human communication. Human listeners assess the speaker’s degree of uncertainty at any time in communication and use this information to shape dialogue. In contrast, currently available computer systems dealing with spoken language are usually not built to perform this task. The ability to detect uncertainty would likely lead to more natural human-computer dialogue. In order to detect un-certainty automatically, we extract linguistic, paralinguistic and dialogue-related features from the Kiel Corpus, a corpus of nat-uralistic task-oriented spoken German. We then use these fea-tures to train a random forests model. Our experimental results show that relatively high classification accuracy can be obtained while employing only 64 well-chosen features (73 % accuracy, 69 % F1). To our best knowledge, this is the first study of auto-matic uncertainty detection using German speech data as well as the first achieving good performance on everyday speech. Index Terms: uncertainty detection, emotion recognition, conversational speech, spontaneous dialogue, random forests, speech rate 1. Tobias Schrank, Barbara Schuppler |
INTERSPEECH | 2 |
| 2014 | Where /ar/ the /r/s in standard austrian German?
Anke Jackschina, Barbara Schuppler, Rudolf Muhr |
INTERSPEECH | 2 |
| 2014 | Pronunciation variation in read and conversational austrian GermanabstractInternational audience Barbara Schuppler, Martine Adda-Decker, Juan Andres Morales-Cordovilla |
INTERSPEECH | 1 |
| 2014 | GRASS: the Graz corpus of Read And Spontaneous Speech
Barbara Schuppler, Martin Hagmüller, Juan Andres Morales-Cordovilla, Hannes Pessentheiner |
LREC | 1 |
| 2010 | Morphological and predictability effects on schwa reduction: the case of dutch word-initial syllablesabstractContains fulltext : 86152.pdf (Publisher’s version ) (Open Access) Iris Hanique, Barbara Schuppler, Mirjam Ernestus |
INTERSPEECH | 2 |
| 2010 | Predicting human perception and ASR classification of word-final [t] by its acoustic sub-segmental propertiesabstractThis paper presents a study on the acoustic sub-segmental properties of word-final /t/ in conversational standard Dutch and how these properties contribute to whether humans and an ASR system classify the /t/ as acoustically present or absent. In general, humans and the ASR system use the same cues (presence of a constriction, a burst, and alveolar frication), but the ASR system is also less sensitive to fine cues (weak bursts, smoothly starting friction) than human listeners and misled by the presence of glottal vibration. These data inform the further development of models of human and automatic speech processing. Barbara Schuppler, Mirjam Ernestus, Wim A. van Dommelen, Jacques C. Koreman |
INTERSPEECH | 1 |
| 2009 | Using temporal information for improving articulatory-acoustic feature classificationabstractThis paper combines acoustic features with a high temporal and a high frequency resolution to reliably classify articulatory events of short duration, such as bursts in plosives. SVM classification experiments on TIMIT and SV Articulatory showed that articulatory-acoustic features (AFs) based on a combination of MFCCs derived from a long window of 25 ms and a short window of 5 ms that are both shifted with 2.5 ms steps (Both) outperform standard MFCCs derived with a window of 25 ms and a shift of 10 ms (Baseline). Finally, comparison of the TIMIT and SV Articulatory results showed that for classifiers trained on data that allows for asynchronously changing AFs (SV Articulatory) the improvement from Baseline to Both is larger than for classifiers trained on data where AFs change simultaneously with the phone boundaries (TIMIT). Barbara Schuppler, Joost van Doremalen, Odette Scharenborg, Bert Cranen, Lou Boves |
ASRU | 1 |
| 2009 | Word-final [t]-deletion: an analysis on the segmental and sub-segmental levelabstractThis paper presents a study on the reduction of word-final [t]s in conversational standard Dutch. Based on a large amount of tokens\nannotated on the segmental level, we show that the bigram\nfrequency and the segmental context are the main predictors for\nthe absence of [t]s. In a second study, we present an analysis of\nthe detailed acoustic properties of word-final [t]s and we show\nthat bigram frequency and context also play a role on the subsegmental\nlevel. This paper extends research on the realization\nof /t/ in spontaneous speech and shows the importance of incorporating\nsub-segmental properties in models of speech. Barbara Schuppler, Wim A. van Dommelen, Jacques C. Koreman, Mirjam Ernestus |
INTERSPEECH | 1 |
| 2008 | Preparing a corpus of dutch spontaneous dialogues for automatic phonetic analysisabstractThis paper presents the steps needed to make a corpus of Dutch spontaneous dialogues accessible for automatic phonetic research aimed at increasing our understanding of reduction phenomena and the role of fine phonetic detail. Since the corpus was not created with automatic processing in mind, it needed to be reshaped. The first part of this paper describes the actions needed for this reshaping in some detail. The second part reports the results of a preliminary analysis of the reduction phenomena in the corpus. For this purpose a phonemic transcription of the corpus was created by means of a forced alignment, first with a lexicon of canonical pronunciations and then with multiple pronunciation variants per word. In this study pronunciation variants were generated by applying a large set of phonetic processes that have been implicated in reduction to the canonical pronunciations of the words. This relatively straightforward procedure allows us to produce plausible pronunciation variants and to verify and extend the results of previous reduction studies reported in the literature. Barbara Schuppler, Mirjam Ernestus, Odette Scharenborg, Lou Boves |
INTERSPEECH | 1 |