Tim Schlippe

dblp:32/9230 · DBLP profile ↗
← Back
22ranked-venue papers
10as first author
3since 2021 · last 2023
0000-0002-9462-8610ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 18 · 10 first-author · 1 since 2021Artificial intelligence and machine learning · 9 · 5 first-authorHuman-computer interaction and ubiquitous computing · 4 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2023 Sentiment Analysis for Shona
abstract
No sentiment analysis system existed for Shona yet—even though it is a Bantu language spoken by nearly 17 million people. Consequently, we collected ShonaSenti—a new corpus of 16,000 tweets in Shona covering 8 different topics. In this paper, we describe our distant supervised labelling strategies to support the annotators to categorize the collected tweets into the 5 sentiment classes of very negative, negative, neutral, positive and very positive. Moreover, we leveraged the Shona sentiment analysis corpus to develop first mono-lingual and cross-lingual sentiment analysis systems for Shona. Our best sentiment analysis systems are cross-lingual and mono-lingual Transformer-based systems which achieve accuracies of 84% on 3 sentiment classes and 70% on 5 sentiment classes.
Barlette Makuwe, Koena Ronny Mabokela, Tim Schlippe
ACII3
2023 Exploring ChatGPT's Empathic Abilities
abstract
Empathy is often understood as the ability to share and understand another individual’s state of mind or emotion. With the increasing use of chatbots in various domains, e.g., children seeking help with homework, individuals looking for medical advice, and people using the chatbot as a daily source of everyday companionship, the importance of empathy in human-computer interaction has become more apparent. Therefore, our study investigates the extent to which ChatGPT based on GPT-3.5 can exhibit empathetic responses and emotional expressions. We analyzed the following three aspects: (1) understanding and expressing emotions, (2) parallel emotional response, and (3) empathic personality. Thus, we not only evaluate ChatGPT on various empathy aspects and compare it with human behavior but also show a possible way to analyze the empathy of chatbots in general. Our results show, that in 91.7% of the cases, ChatGPT was able to correctly identify emotions and produces appropriate answers. In conversations, ChatGPT reacted with a parallel emotion in 70.7% of cases. The empathic capabilities of ChatGPT were evaluated using a set of five questionnaires covering different aspects of empathy. Even though the results show, that the scores of ChatGPT are still worse than the average of healthy humans, it scores better than people who have been diagnosed with Asperger syndrome /high-functioning autism.
Kristina Schaaff, Caroline Reinig, Tim Schlippe
ACII3
2021 AI in Art: Simulating the Human Painting Process
Alexander Leiser, Tim Schlippe
ArtsIT2
2020 Visualizing Voice Characteristics with Type Design in Closed Captions for Arabic
abstract
Diversification of fonts in video captions based on the voice characteristics, namely loudness, speed and pauses, can affect the viewer receiving the content. This study evaluates a new method, WaveFont, which visualizes the voice characteristics for captions in an intuitive way. The study was specifically designed to test captions, which aims to add a new experience for Arabic viewers. The results indicate that our visualization is comprehensible and acceptable and provides significant added value-for hearing-impaired and non-hearing impaired participants: Significantly more participants stated that WaveFont improves their watching experience more than standard captions.
Tim Schlippe, Shaimaa Alessai, Ghanimeh El-Taweel, Matthias Wölfel, Wajdi Zaghouani
CW1
2016 Word segmentation and pronunciation extraction from phoneme sequences through cross-lingual word-to-phoneme alignment
Felix Stahlberg, Tim Schlippe, Stephan Vogel, Tanja Schultz
Comput. Speech Lang.2
2015 Cross-lingual lexical language discovery from audio data using multiple translations
abstract
Zero-resource Automatic Speech Recognition (ZR ASR) addresses target languages without given pronunciation dictionary, transcribed speech, and language model. Lexical discovery for ZR ASR aims to extract word-like chunks from speech. Lexical discovery benefits from the availability of written translations in another source language. In this paper, we improve lexical discovery even more by combining multiple source languages. We present a novel method for combining noisy word segmentations resulting in up to 11.2% relative F-score gain. When we extract word pronunciations from the combined segmentations to bootstrap an ASR system, we improve accuracy by 9.1% relative compared to the best system with only one translation, and by 50.1% compared to monolingual lexical discovery.
Felix Stahlberg, Tim Schlippe, Stephan Vogel, Tanja Schultz
ICASSP2
2014 Methods for efficient semi-automatic pronunciation dictionary bootstrapping
abstract
In this paper we propose efficient methods which contribute to a rapid and economic semi-automatic pronunciation dictionary development and evaluate them on English, German, Spanish, Vietnamese, Swahili, and Haitian Creole. First we determine optimal strategies for the word selection and the period for the grapheme-to-phoneme model retraining. In addition to the traditional concatenation of single phonemes most commonly associated with each grapheme, we show that web-derived pronunciations and cross-ligual grapheme-to-phoneme models can help to reduce the initial editing effort. Furthermore we show that our phoneme-level combination of the output of multiple grapheme-to-phoneme converters reduces the editing effort more than the best single converters. Totally, we report on average 15% relative editing effort reduction with our phonemelevel combination compared to conventional methods. An additional reduction of 6% relative was possible by including initial pronunciations from Wiktionary for English, German, and Spanish.
Tim Schlippe, Matthias Merz, Tanja Schultz
INTERSPEECH1
2014 BioKIT - real-time decoder for biosignal processing
abstract
We introduce BioKIT, a new Hidden Markov Model based toolkit to preprocess, model and interpret biosignals such as speech, motion, muscle and brain activities. The focus of this toolkit is to enable researchers from various communities to pursue their experiments and integrate real-time biosignal interpretation into their applications. BioKIT boosts a flexible two-layer structure with a modular C++ core that interfaces with a Python scripting layer, to facilitate development of new applications. BioKIT employs sequence-level parallelization and memory sharing across threads. Additionally, a fully integrated error blaming component facilitates in-depth analysis. A generic terminology keeps the barrier to entry for researchers from multiple fields to a minimum. We describe our onlinecapable dynamic decoder and report on initial experiments on three different tasks. The presented speech recognition experiments employ Kaldi [1] trained deep neural networks with the results set in relation to the real time factor needed to obtain them.
Dominic Telaar, Michael Wand 0002, Dirk Gehrig, Felix Putze, Christoph Amma, Dominic Heger, Ngoc Thang Vu, Mark Erhardt, Tim Schlippe, Matthias Janke, Christian Herff, Tanja Schultz
INTERSPEECH9
2014 GlobalPhone: Pronunciation Dictionaries in 20 Languages
Tanja Schultz, Tim Schlippe
LREC2
2014 Web-based tools and methods for rapid pronunciation dictionary creation
Tim Schlippe, Sebastian Ochs 0002, Tanja Schultz
Speech Commun.1
2013 Recurrent neural network language modeling for code switching conversational speech
abstract
Code-switching is a very common phenomenon in multilingual communities. In this paper, we investigate language modeling for conversational Mandarin-English code-switching (CS) speech recognition. First, we investigate the prediction of code switches based on textual features with focus on Part-of-Speech (POS) tags and trigger words. Second, we propose a structure of recurrent neural networks to predict code-switches. We extend the networks by adding POS information to the input layer and by factorizing the output layer into languages. The resulting models are applied to our task of code-switching language modeling. The final performance shows 10.8% relative improvement in perplexity on the SEAME development set which transforms into a 2% relative improvement in terms of Mixed Error Rate and a relative improvement of 16.9% in perplexity on the evaluation set which leads to a 2.7% relative improvement of MER.
Heike Adel, Ngoc Thang Vu, Franziska Kraus, Tim Schlippe, Haizhou Li 0001, Tanja Schultz
ICASSP4
2013 Rapid bootstrapping of a Ukrainian large vocabulary continuous speech recognition system
abstract
We report on our efforts toward an LVCSR system for the Slavic language Ukrainian. We describe the Ukrainian text and speech database recently collected as a part of our GlobalPhone corpus [1] with our Rapid Language Adaptation Toolkit [2]. The data was complemented by a large collection of text data crawled from various Ukrainian websites. For the production of the pronunciation dictionary, we investigate strategies using grapheme-to-phoneme (g2p) models derived from existing dictionaries of other languages, thereby reducing severely the necessary manual effort. Russian and Bulgarian g2p models even decrease the number of pronunciation rules to one fifth. We achieve significant improvement by applying state-of-the art techniques for acoustic modeling and our day-wise text collection and language model interpolation strategy [3]. Our best system achieves a word error rate of 11.21% on the test set on read newspaper speech.
Tim Schlippe, Mykola Volovyk, Kateryna Yurchenko, Tanja Schultz
ICASSP1
2013 Statistical machine translation based text normalization with crowdsourcing
abstract
In [1], we have proposed systems for text normalization based on statistical machine translation (SMT) methods which are constructed with the support of Internet users and evaluated those with French texts. Internet users normalize text displayed in a web interface in an annotation process, thereby providing a parallel corpus of normalized and non-normalized text. With this corpus, SMT models are generated to translate non-normalized into normalized text. In this paper, we analyze their efficiency for other languages. Additionally, we embedded the English annotation process for training data in Amazon Mechanical Turk and compare the quality of texts thoroughly annotated in our lab to those annotated by the Turkers. Finally, we investigate how to reduce the user effort by iteratively applying an SMT system to the next sentences to be edited, built from the sentences which have been annotated so far.
Tim Schlippe, Chenfei Zhu, Daniel Lemcke, Tanja Schultz
ICASSP1
2013 GlobalPhone: A multilingual text & speech database in 20 languages
abstract
This paper describes the advances in the multilingual text and speech database GlobalPhone, a multilingual database of high-quality read speech with corresponding transcriptions and pronunciation dictionaries in 20 languages. GlobalPhone was designed to be uniform across languages with respect to the amount of data, speech quality, the collection scenario, the transcription and phone set conventions. With more than 400 hours of transcribed audio data from more than 2000 native speakers GlobalPhone supplies an excellent basis for research in the areas of multilingual speech recognition, rapid deployment of speech processing systems to yet unsupported languages, language identification tasks, speaker recognition in multiple languages, multilingual speech synthesis, as well as monolingual speech recognition in a large variety of languages.
Tanja Schultz, Ngoc Thang Vu, Tim Schlippe
ICASSP3
2013 Unsupervised language model adaptation for automatic speech recognition of broadcast news using web 2.0
abstract
We improve the automatic speech recognition of broadcast news using paradigms from Web 2.0 to obtain timeand topicrelevant text data for language modeling. We elaborate an unsupervised text collection and decoding strategy that includes crawling appropriate texts from RSS Feeds, complementing it with texts from Twitter, language model and vocabulary adaptation, as well as a 2-pass decoding. The word error rates of the tested French broadcast news shows from Europe 1 are reduced by almost 32% relative with an underlying language model from the GlobalPhone project [1] and by almost 4% with an underlying language model from the Quaero project. The tools that we use for the text normalization, the collection of RSS Feeds together with the text on the related websites, a TF-IDF-based topic words extraction, as well as the opportunity for language model interpolation are available in our Rapid Language Adaptation Toolkit [2] [3].
Tim Schlippe, Lukasz Gren, Ngoc Thang Vu, Tanja Schultz
INTERSPEECH1
2012 Grapheme-to-phoneme model generation for Indo-European languages
abstract
In this paper, we evaluate grapheme-to-phoneme (g2p) models among languages and of different quality. We created g2p models for Indo-European languages with word-pronunciation pairs from the GlobalPhone project and from Wiktionary [1]. Then we checked their quality in terms of consistency and complexity as well as their impact on Czech, English, French, Spanish, Polish, and German ASR. While the GlobalPhone dictionaries were manually cross-checked and have been used successfully in LVCSR, Wiktionary pronunciations have been provided by the Internet community and can be used to rapidely and economically create pronunciation dictionaries for new languages and domains.
Tim Schlippe, Sebastian Ochs 0002, Tanja Schultz
ICASSP1
2012 A first speech recognition system for Mandarin-English code-switch conversational speech
abstract
This paper presents first steps toward a large vocabulary continuous speech recognition system (LVCSR) for conversational Mandarin-English code-switching (CS) speech. We applied state-of-the-art techniques such as speaker adaptive and discriminative training to build the first baseline system on the SEAME corpus [1] (South East Asia Mandarin-English). For acoustic modeling, we applied different phone merging approaches based on the International Phonetic Alphabet (IPA) and Bhattacharyya distance in combination with discriminative training to improve accuracy. On language model level, we investigated statistical machine translation (SMT) - based text generation approaches for building code-switching language models. Furthermore, we integrated the provided information from a language identification system (LID) into the decoding process by using a multi-stream approach. Our best 2-pass system achieves a Mixed Error Rate (MER) of 36.6% on the SEAME development set.
Ngoc Thang Vu, Dau-Cheng Lyu, Jochen Weiner, Dominic Telaar, Tim Schlippe, Fabian Blaicher, Chng Eng Siong, Tanja Schultz, Haizhou Li 0001
ICASSP5
2012 Automatic Error Recovery for Pronunciation Dictionaries
abstract
In this paper, we present our latest investigations on pronunciation modeling and its impact on ASR. We propose completely automatic methods to detect, remove, and substitute inconsistent or flawed entries in pronunciation dictionaries. The experiments were conducted on different tasks, namely (1) word-pronunciation pairs from the Czech, English, French, German, Polish, and Spanish Wiktionary [1], a multilingual wiki-based open content dictionary, (2) our GlobalPhone Hausa pronunciation dictionary [2], and (3) pronunciations to complement our Mandarin-English SEAME code-switch dictionary [3]. In the final results, we fairly observed on average an improvement of 2.0% relative in terms of word error rate and even 27.3% for the case of English Wiktionary word-pronunciation pairs.
Tim Schlippe, Sebastian Ochs 0002, Ngoc Thang Vu, Tanja Schultz
INTERSPEECH1
2012 Word segmentation through cross-lingual word-to-phoneme alignment
abstract
We present our new alignment model Model 3P for cross-lingual word-to-phoneme alignment, and show that unsupervised learning of word segmentation is more accurate when information of another language is used. Word segmentation with cross-lingual information is highly relevant to bootstrap pronunciation dictionaries from audio data for Automatic Speech Recognition, bypass the written form in Speech-to-Speech Translation or build the vocabulary of an unseen language, particularly in the context of under-resourced languages. Using Model 3P for the alignment between English words and Spanish phonemes outperforms a state-of-the-art monolingual word segmentation approach [1] on the BTEC corpus [2] by up to 42% absolute in F-Score on the phoneme level and a GIZA++ alignment based on IBM Model 3 by up to 17%.
Felix Stahlberg, Tim Schlippe, Stephan Vogel, Tanja Schultz
SLT2
2010 Wiktionary as a source for automatic pronunciation extraction
abstract
In this paper, we analyze whether dictionaries from the World Wide Web which contain phonetic notations, may support the rapid creation of pronunciation dictionaries within the speech recognition and speech synthesis system building process. As a representative dictionary, we selected Wiktionary [1] since it is at hand in multiple languages and, in addition to the definitions of the words, many phonetic notations in terms of the International Phonetic Alphabet (IPA) are available. Given word lists in four languages English, French, German, and Spanish, we calculated the percentage of words with phonetic notations in Wiktionary. Furthermore, two quality checks were performed: First, we compared pronunciations from Wiktionary to pronunciations from dictionaries based on the GlobalPhone project, which had been created in a rule-based fashion and were manually cross-checked [2]. Second, we analyzed the impact of Wiktionary pronunciations on automatic speech recognition (ASR) systems. French Wiktionary achieved the best pronunciation coverage, containing 92.58% phonetic notations for the French GlobalPhone word list as well as 76.12% and 30.16% for country and international city names. In our ASR systems evaluation, the Spanish system gained the most improvement from Wiktionary pronunciations with 7.22% relative word error rate reduction.
Tim Schlippe, Sebastian Ochs 0002, Tanja Schultz
INTERSPEECH1
2010 Text normalization based on statistical machine translation and internet user support
Tim Schlippe, Chenfei Zhu, Jan Gebhardt, Tanja Schultz
INTERSPEECH1
2010 Rapid bootstrapping of five eastern european languages using the rapid language adaptation toolkit
abstract
This paper presents our latest efforts toward LVCSR systems for five Eastern European languages such as Bulgarian, Croatian, Czech, Polish, and Russian using our Rapid Language Adaptation Toolkit (RLAT) [1]. We investigated the possibility of crawling large quantities of text material from the Internet, which is very cheap but also requires text post-processing steps due to the varying text quality. The goal of this study is to determine the best strategy for language model optimization on the given domain in a short time period with minimal human effort. Our results show that we can build an initial ASR system for these five languages in only twenty days using RLAT. On the multilingual GlobalPhone speech corpus [2], we achieved a word error rate (WER) of 16.9 % for Bulgarian, 32.8 % for
Ngoc Thang Vu, Tim Schlippe, Franziska Kraus, Tanja Schultz
INTERSPEECH2