VLDB 2026 Research / reviewers in the wild / expert
Thomas Schaaf
dblp:75/5600
· DBLP profile ↗
34ranked-venue papers
5as first author
8since 2021 · last 2025
0000-0002-9569-4759ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 21 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 18 · 3 first-author · 8 since 2021Computer networks · 5Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | LLM-Based Dictation Detection from Doctor-Patient ConversationsabstractClinical ambient documentation systems need to detect dictation from doctor-patient conversations to ensure critical statements are properly recorded. Existing methods struggle to detect dictation interleaved within conversational exchanges. We investigated LLM-based approaches: zero-shot, multi-shot, manual/auto prompt tuning, and supervised fine-tuning for medical dictation detection. A word alignment-based method is proposed to mitigate ASR errors and LLM hallucinations while performing utterance- and word-level evaluation. Supervised finetuning of a Qwen2-1.5B LLM achieves the best performance: $\mathbf{9 0. 4 \%}$ utterance-level and $\mathbf{8 6. 8 \%}$ word-level macro F1 on manual transcripts, with minimal loss on ASR. This surpasses Claude 3.5 Sonnet with auto-tuned prompts and multiple in-context learning examples, though the latter requires much less training data. We release model outputs on the ACI-Bench dataset, providing the first open resource for the evaluation of medical dictation detection. Mojtaba Kadkhodaie, Susanne Burger, Thomas Schaaf |
ASRU | 5 |
| 2024 | Annotate the Way You Think: An Incremental Note Generation Framework for the Summarization of Medical ConversationsabstractThe scarcity of public datasets for the summarization of medical conversations has been a limiting factor for advancing NLP research in the healthcare domain, and the structure of the existing data is largely limited to the simple format of conversation-summary pairs. We therefore propose a novel Incremental Note Generation (ING) annotation framework capable of greatly enriching summarization datasets in the healthcare domain and beyond. Our framework is designed to capture the human summarization process via an annotation task by instructing the annotators to first incrementally create a draft note as they accumulate information through a conversation transcript (Generation) and then polish the draft note into a reference note (Rewriting). The annotation results include both the reference note and a comprehensive editing history of the draft note in tabular format. Our pilot study on the task of SOAP note generation showed reasonable consistency between four expert annotators, established a solid baseline for quantitative targets of inter-rater agreement, and demonstrated the ING framework as an improvement over the traditional annotation process for future modeling of summarization. Longxiang Zhang, Caleb D. Hart, Susanne Burger, Thomas Schaaf |
LREC/COLING | 4 |
| 2024 | Reference-Free Estimation of the Quality of Clinical Notes Generated from Doctor-Patient Conversations
Mojtaba Kadkhodaie, John Glover, Thomas Schaaf |
INTERSPEECH | 3 |
| 2023 | Investigating the Utility of Synthetic Data for Doctor-Patient Conversation Summarization
Colin A. Grambow, Mojtaba Kadkhodaie, Federico Fancellu, Thomas Schaaf |
INTERSPEECH | 6 |
| 2022 | Extract and Abstract with BART for Clinical Notes from Doctor-Patient Conversations
Longxiang Zhang, Hamid Reza Hassanzadeh, Thomas Schaaf |
INTERSPEECH | 4 |
| 2022 | AdaFocal: Calibration-aware Adaptive Focal LossabstractMuch recent work has been devoted to the problem of ensuring that a neural network's confidence scores match the true probability of being correct, i.e. the calibration problem. Of note, it was found that training with focal loss leads to better calibration than cross-entropy while achieving similar level of accuracy \cite{mukhoti2020}. This success stems from focal loss regularizing the entropy of the model's prediction (controlled by the parameter $\gamma$), thereby reining in the model's overconfidence. Further improvement is expected if $\gamma$ is selected independently for each training sample (Sample-Dependent Focal Loss (FLSD-53) \cite{mukhoti2020}). However, FLSD-53 is based on heuristics and does not generalize well. In this paper, we propose a calibration-aware adaptive focal loss called AdaFocal that utilizes the calibration properties of focal (and inverse-focal) loss and adaptively modifies $\gamma_t$ for different groups of samples based on $\gamma_{t-1}$ from the previous step and the knowledge of model's under/over-confidence on the validation set. We evaluate AdaFocal on various image recognition and one NLP task, covering a wide variety of network architectures, to confirm the improvement in calibration while achieving similar levels of accuracy. Additionally, we show that models trained with AdaFocal achieve a significant boost in out-of-distribution detection. Thomas Schaaf, Matthew R. Gormley |
NeurIPS | 2 |
| 2021 | Are You Dictating to Me? Detecting Embedded Dictations in Doctor-Patient ConversationsabstractMedical scribes chart doctor-patient conversations in real time or by listening to an audio recording afterwards. Doctors sometimes dictate during a patient encounter, a highly informative part for a scribe. We introduced a light-weight annotation schema and ana-lyzed recordings of 105 randomly selected doctor-patient encounters from 21 physicians to quantify the frequency and automatically de-tect dictated regions. Dictation behavior of individual doctors was consistent but varied among them. A linguistic analysis is provided to describe differences of doctors speech when talking to a patient or dictating. A description of the data is given, highlighting challenges of segmenting audio into conversation and dictation regions. We in-vestigate different features and methods to segment conversations including keyword spotting, acoustic features and class-conditioned language models. Results are anchored to a majority class base-line. Using only acoustic features allows to predict dictated speech without the need of a speech recognition system performing com-parable to a rule-based approach using lexical features derived from a speech recognition system. Performance is assessed using leave-one-physician-out cross validation and an analysis using a random forest classifier indicates that language model derived features are most useful, and that a combination of acoustic and lexical features performed best. Thomas Schaaf, Longxiang Zhang, Alireza Bayestehtashk, Mark C. Fuhs, Shahid Durrani, Susanne Burger, Monika Woszczyna, Thomas Polzin |
ASRU | 1 |
| 2021 | Effective Convolutional Attention Network for Multi-label Clinical Document ClassificationabstractMulti-label document classification (MLDC) problems can be challenging, especially for long documents with a large label set and a long-tail distribution over labels.In this paper, we present an effective convolutional attention network for the MLDC problem with a focus on medical code prediction from clinical documents.Our innovations are three-fold: (1) we utilize a deep convolution-based encoder with the squeeze-and-excitation networks and residual networks to aggregate the information across the document and learn meaningful document representations that cover different ranges of texts; (2) we explore multilayer and sum-pooling attention to extract the most informative features from these multiscale representations; (3) we combine binary cross entropy loss and focal loss to improve performance for rare labels.We focus our evaluation study on MIMIC-III, a widely used dataset in the medical domain.Our models outperform prior work on medical coding and achieve new state-of-the-art results on multiple metrics.We also demonstrate the language independent nature of our approach by applying it to two non-English datasets.Our model outperforms prior best model and a multilingual Transformer model by a substantial margin. Russell Klopfer, Matthew R. Gormley, Thomas Schaaf |
EMNLP (1) | 5 |
| 2020 | Posterior Calibrated Training on Sentence Classification TasksabstractMost classification models work by first predicting a posterior probability distribution over all classes and then selecting that class with the largest estimated probability.In many settings however, the quality of posterior probability itself (e.g., 65% chance having diabetes), gives more reliable information than the final predicted class alone.When these methods are shown to be poorly calibrated, most fixes to date have relied on posterior calibration, which rescales the predicted probabilities but often has little impact on final classifications.Here we propose an end-to-end training procedure called posterior calibrated (PosCal) training that directly optimizes the objective while minimizing the difference between the predicted and empirical posterior probabilities.We show that PosCal not only helps reduce the calibration error but also improve task performance by penalizing drops in performance of both objectives.Our PosCal achieves about 2.5% of task performance gain and 16.1% of calibration error reduction on GLUE (Wang et al., 2018) compared to the baseline.We achieved the comparable task performance with 13.2% calibration error reduction on xSLUE (Kang and Hovy, 2019), but not outperforming the two-stage calibration baseline.PosCal training can be easily extendable to any types of classification tasks as a form of regularization term.Also, PosCal has the advantage that it incrementally tracks needed statistics for the calibration objective during the training process, making efficient use of large training sets 1 . Taehee Jung, Dongyeop Kang, Lucas K. Mentch, Thomas Schaaf |
ACL | 5 |
| 2019 | IT Service Management Frameworks Compared - Simplifying Service Portfolio Management
Michael Brenner 0002, Thomas Schaaf |
IM | 3 |
| 2014 | gSLM: The Initial Steps for the Specification of a Service Management Standard for Federated e-InfrastructuresabstractThis paper presents a methodology used to create a site independent assessment process of the capabilities of service management systems in federated e-infrastructures that can contribute to introduce or improve service management in these application domains. Based on ISO/IEC 20000 concepts it consist of an actors and relationships model, a set of management processes with their corresponding requirements and a capability model that all converge in an easy to use assessment tool. The methodology has been evaluated through its adoption in an existing federated e-infrastructure. Joan Serrat 0001, Tomasz Szepieniec, Adam Belloum, Javier Rubio-Loyola, Owen Appleton, Thomas Schaaf, Joanna Kocot |
iiWAS | 6 |
| 2014 | S3MS a simple service & security management systemabstractPlanning and implementing IT Service Management (ITSM) and Information Security Management according to the International Standards ISO/IEC 20000-1 and ISO/IEC 27001, following good practice approaches like ITIL, and considering IT governance controls as described in COBIT, is challenging in multiple ways. One of the most obvious difficulties in practice is to produce and maintain the required documentation in a way that it effectively supports the delivery of IT services and the implementation of security controls, by at the same time avoiding an amount of bureaucratic overhead that jeopardizes the efficiency of the management system in the end. The “S3MS” approach presented here is a practical approach of implementing ITSM and ISM in a consolidated and integrated way. It is driven by the methodology and requirements provided by the above mentioned standards and frameworks, but it complements them by offering a wide set of templates and samples that can be re-used, instantiated and/or refined to generate what is needed to deploy an effective documented service and security management system. Therefore, the S3MS framework is divided into a service module, a security module and a general management system module - all of which are fully aligned to each other. S3MS is the outcome from merging scientific/academic work with practical experiences and lessons learned in various ITSM- and ISM-related projects in industry and in the public service sector. Thomas Schaaf, Robert Kuhlig |
NOMS | 1 |
| 2012 | An approach to consolidate and optimize monitoring solutionsabstractLike most IT service providers, the Leibniz Supercomputing Centre (LRZ) is facing the challenge of managing complex monitoring solutions, consisting of various tools offering monitoring capabilities for the different systems and applications in support of the services provided. Monitoring encompasses a wide range of functional aspects, and therefore it is hardly possible to significantly reduce the portfolio and number of tools. High administration effort and expensive licensing are typical consequences. This short paper introduces a systematic and holistic approach that shall help an IT service provider to analyze their entire monitoring environment. Further, this method provides guidance on how to consolidate the tool landscape and optimize tool support for IT Service Management (ITSM) processes. Thomas Schaaf |
NOMS | 2 |
| 2011 | An information model for inter-organizational fault managementabstractIT service providers outsourcing (parts of) their IT services often have to face the side-effect of losing control over service delivery and support, and thus over the service quality. To respond to this problem, guidance on delivering IT services in inter-organizational environments is needed. One of the most important disciplines in this context is fault management. This paper presents the fundamentals of an information model for inter-organizational fault management and shows how this information model is an important component of a comprehensive management architecture for inter-organizational fault management. Patricia Marcu, Thomas Schaaf |
Integrated Network Management | 2 |
| 2011 | A maturity model for tool landscapes of IT service providersabstractIn the last years, several approaches and models for IT service management (ITSM) have emerged from research under the business-driven IT management (BDIM) paradigm as well as from industry initiatives. Effective and comprehensive tool support for ITSM is still one of the big remaining challenges. Nowadays, many IT service providers face the problem, that they are using dozens to hundreds of different tools to manage their infrastructure and services. Although such management tools are intended to increase the efficiency of ITSM processes, very complex and heterogeneous tool environments may have an inverse impact. This paper presents an approach to addresses this problem area. It aims at providing guidance in assessing and improving existing tool landscapes based on a global maturity model and specific capability models to be applied to the different topic areas in ITSM. Thomas Schaaf |
Integrated Network Management | 2 |
| 2011 | Analysis of Dialectal Influence in Pan-Arabic ASRabstractIn this paper, we analyze the impact of five Arabic dialects on the front-end and pronunciation dictionary component of an Automatic Speech Recognition (ASR) system.We use ASR"s phonetic decision tree as a diagnostic tool to compare the robustness of MFCC to MLP front-ends to dialectal variations in the speech data and found that MLP Bottle-Neck features are less robust to dialectal variation.We also perform a rulebased analysis of the pronunciation dictionary, which enables us to identify dialectal words in the vocabulary and automatically generate pronunciations for unseen words.We show that our technique produces pronunciations with an average phone error rate 9.2%. Udhyakumar Nallasamy, Michael Garbus, Florian Metze, Qin Jin, Thomas Schaaf, Tanja Schultz |
INTERSPEECH | 5 |
| 2011 | VTLN in the MFCC Domain: Band-Limited versus Local InterpolationabstractWe propose a new easy-to-implement method to compute a Lin-ear Transform (LT) to perform Vocal Tract Length Normalization (VTLN) on truncated Mel Frequency Cepstral Coefficients (MFCCs) normally used in distributed speech recognition. The method is based on a Local Interpolation which is independent of the Mel filter design. Local Interpolation (LILT) VTLN is theoretically and experimentally compared to a global scheme based on band-limited interpolation (BLI-VTLN) and the conventional frequency warp-ing scheme (FFT-VTLN). Investigating the interoperability of these methods shows that the performance of LILT-VTLN is on par with FFT-VTLN and BLI-VTLN. Models trained with LILT- and BLI-VTLN performance degrades if FFT-VTLN is used as a front-end. The degradation for LILT-VTLN is slightly less, indicating that it produces models that are a better match for FFT-VTLN. Index Terms — Automatic speech recognition, VTLN, fre-quency warping, linear transform Ehsan Variani, Thomas Schaaf |
INTERSPEECH | 2 |
| 2010 | Analysis of gender normalization using MLP and VTLN featuresabstractThis paper analyzes the capability of multilayer perceptron frontends to perform speaker normalization. We find the context decision tree to be a very useful tool to assess the speaker normalization power of different frontends. We introduce a gender question into the training of the phonetic context decision tree. After the context clustering the gender specific models are counted. We compare this for the following frontends: (1) Bottle-Neck (BN) with and without vocal tract length normalization (VTLN), (2) standard MFCC, (3) stacking of multiple MFCC frames with linear discriminant analysis (LDA). We find the BN-frontend to be even more effective in reducing the number of gender questions than VTLN. From this we conclude that a Bottle-Neck frontend is more effective for gender normalization. Combining VTLN and BN-features reduces the number of gender specific models further. Thomas Schaaf, Florian Metze |
INTERSPEECH | 1 |
| 2009 | Introducing process-oriented IT service management at an academic computing center: An interim reportabstractThe Leibniz Supercomputing Centre (Leibniz-Rechenzentrum, LRZ) is a service provider for a variety of academic institutions, mainly in the Munich (Germany) area. The services provided range from network services, server hosting, application services to specialized supercomputing services. Even in academia, computing services become ever more business critical : IT services for university spin-offs, virtual labs provided to other universities as an application service, and an increasing number of industry cooperation projects require highly available and reliable services. As scope, volume, complexity and required quality of services increase, financial and personal resources to provide these do not (at least not on the same scale). The only way to meet this challenge is to improve operational effectiveness and efficiency. Such improvements do not seem achievable just by purchasing or developing more management tools (the LRZ already uses an abundance of management software applications). Addressing the, in the past often somewhat neglected, organizational aspects of IT service management (ITSM), i.e. process-oriented ITSM, promises to yield much better gains in efficiency. A project to introduce process-oriented IT service management at the LRZ was started at the end of 2007. This presentation outlines the motivation and scope of this long-running (3-4 years), multi-faceted project, and presents and interim report on results and experiences in the introduction of process-oriented ITSM at a large academic computing center. Michael Brenner 0002, Heinz-Gerd Hegering, Helmut Reiser, Thomas Schaaf |
Integrated Network Management | 5 |
| 2009 | Towards an information model for ITIL and ISO/IEC 20000 processesabstractAs IT service providers are adopting more comprehensive approaches towards IT service management (ITSM), they increasingly need to rely on ITSM software solutions in their day-to-day operations. However, when wishing to integrate ITSM software from one vendor with that of another, the lack of underlying standards becomes woefully apparent. Without any standardized information model for ITSM processes, efficient and integrated ITSM will remain a vision. While in the telecommunications sector, a lot of work has been invested into developing the shared information/data model (SID), a companion model for the industry-specific process framework enhanced telecom operations map (eTOM), no equivalent for the more general process frameworks of ITIL and ISO/IEC 20000 is in sight. This paper introduces an approach towards an information model for ITSM processes. The presented method leverages work done for SID, by adapting and complementing SID concepts and content to produce an information model compliant to ISO/IEC 20000 requirements and ITIL recommendations. Michael Brenner 0002, Thomas Schaaf, Alexander Scherer |
Integrated Network Management | 2 |
| 2007 | An Adaptive Distance Measure for Similarity Based Playlist GenerationabstractNowadays, a large part of all music ever recorded is digitally available and due to MP3 already ten thousands of songs can be carried around on a mobile device. Intelligent automatic song selection is more and more required alternatively to random selection or manual playlist generation. We propose a system, that generates playlists including songs similar to accepted ones, discarding songs similar to rejected ones, where similar refers to timbre. Additional adaptivity is achieved with a user-adaptive distance function which in our case requires modeling features separately. After a seed-song (which is the first accepted song) is given by the user, the distance function is used by a song selection strategy to select songs. Minimal user feedback is collected with a skip button that is pressed to directly jump to the next song and explicitly reject the current one while acceptance is implicitly given by listening to a song. Daniel Gärtner, Florian Kraft, Thomas Schaaf |
ICASSP (1) | 3 |
| 2006 | A comparative study of Gaussian selection methods in large vocabulary continuous speech recognitionabstractGaussian mixture models are the most popular probability density used in automatic speech recognition. During decoding, often many Gaussians are evaluated. Only a small number of Gaussians contributes significantly to probability. Several promising methods to select relevant Gaussians are known. These methods have different properties in terms of required memory, overhead and quality of selected Gaussians. Projection search, bucket box intersection, and Gaussian clustering are investigated in a broadcast news system with focus on adaptation (MLLR). Index Terms: speech recognition, LVCSR, Gaussian selection, speaker adaptation, MLLR. Dirk Gehrig, Thomas Schaaf |
INTERSPEECH | 2 |
| 2005 | Temporal ICA for classification of acoustic events i a kitchen environmentabstractWe describe a feature extraction method for general audio modeling using a temporal extension of Independent Component Analysis (ICA) and demonstrate its utility in the context of a sound classification task in a kitchen environment. Our approach accounts for temporal dependencies over multiple analysis frames much like the standard audio modeling technique of adding first and second temporal derivatives to the feature set. Using a real-world dataset of kitchen sounds, we show that our approach outperforms a canonical version of this standard front end, the mel-frequency cepstral coefficients (MFCCs), which has found successful application in automatic speech recognition tasks. Florian Kraft, Robert G. Malkin, Thomas Schaaf, Alex Waibel |
INTERSPEECH | 3 |
| 2005 | Document driven machine translation enhanced ASRabstractIn human-mediated translation scenarios a human interpreter translates between a source and a target language using either a spoken or a written representation of the source language. In this paper we improve the recognition performance on the speech of the human translator spoken in the target language by taking advantage of the source language representations. We use machine translation techniques to translate between the source and target language resources and then bias the target language speech recognizer towards the gained knowledge, hence the name Machine Translation Enhanced Automatic Speech Recognition. We investigate several different techniques among which are restricting the search vocabulary, selecting hypotheses from n-best lists, applying cache and interpolation schemes to language modeling, and combining the most successful techniques into our final, iterative system. Overall we outperform the baseline system by a relative word error rate reduction of 37.6%. Matthias Paulik, Christian Fügen, Sebastian Stüker, Tanja Schultz, Thomas Schaaf, Alex Waibel |
INTERSPEECH | 5 |
| 2004 | Dictionary refinements based on phonetic consensus and non-uniform pronunciation reductionabstractIn this paper we present a procedure to refine the recognition dictionary based on a composite approach to prune the unneeded pronunciations. First, pruning is applied in a non-uniform manner according to the characteristics of each word. Even though this straightforward operation may produce high-quality dictionaries, it makes the refined dictionary heavily dependent on the data used in this process. For the words not observed in the data, we propose, in second place, to use multiple sequence alignment techniques in order to find phonetic consensus among the pronunciation variants and select the worthy pronunciations that will represent the unobserved words. Experimental results show that our dictionary refining method helps to improve the recognition performance in two relevant aspects: it increases the recognition accuracy by reducing the cross-word confusibility and it improves the recognition speed by reducing the complexity of the search space. 1. Gustavo Hernández Ábrego, Lex Olorenshaw, Raquel Tato, Thomas Schaaf |
INTERSPEECH | 4 |
| 2004 | Speaker adaptation with all-pass transforms
John W. McDonough, Thomas Schaaf, Alex Waibel |
Speech Commun. | 2 |
| 2002 | On maximum mutual information speaker-adapted trainingabstractIn this work, we combine maximum mutual information-based parameter estimation with speaker-adapted training (SAT). As will be shown, this can be achieved by performing unsupervised parameter estimation on the test data, a distinct advantage for many recognition tasks involving conversational speech. We also propose an approximation to the maximum likelihood and maximum mutual information SAT re-estimation formulae that greatly reduces the amount of disk space required to conduct training on corpora such as Broadcast News, which contains speech from thousands of speakers. We present the results of a set of speech recognition experiments on three test sets: the English Spontaneous Scheduling Task corpus, Broadcast News, and a new corpus of Meeting Room data collected at the Interactive Systems Laboratories of the Carnegie Mellon University. John W. McDonough, Thomas Schaaf, Alex Waibel |
ICASSP | 2 |
| 2002 | Lecture and Presentation Tracking in an Intelligent Meeting RoomabstractArchiving, indexing, and later browsing through stored presentations and lectures is increasingly being used. We have investigated the special problems and advantages of lectures and propose the design and adaptation of a speech recognizer to a lecture such that the recognition accuracy can be significantly improved by prior analysis of the presented documents using a special class-based language model. We define a tracking accuracy measure which measures how well a system can automatically align recognized words with parts of a presentation and show that by prior exploitation of the presented documents, the tracking accuracy can be improved. The system described in this paper is part of an intelligent meeting room developed in the European Union-sponsored project FAME (Facilitating Agent for Multicultural Exchange). Ivica Rogina, Thomas Schaaf |
ICMI | 2 |
| 2001 | The ISL evaluation system for Verbmobil-IIabstractDescribes the 2000 ISL large vocabulary speech recognition system for fast decoding of conversational speech which was used in the German Verbmobil-II project. The challenge of this task is to build robust acoustic models to handle different dialects, spontaneous effects, and crosstalk as occur in conversational speech. We present speaker incremental normalization and adaptation experiments close to real-time constraints. To reduce the number of consequential errors caused by out-of-vocabulary words, we conducted filler-model experiments to handle unknown proper names. The overall improvements from 1998 to 2000 resulted in a word error reduction from 40% to 17% on our development test set. Hagen Soltau, Thomas Schaaf, Florian Metze, Alex Waibel |
ICASSP | 2 |
| 2001 | Advances in automatic meeting record creation and accessabstractOral communication is transient, but many important decisions, social contracts and fact findings are first carried out in an oral setup, documented in written form and later retrieved. At Carnegie Mellon University's Interactive Systems Laboratories we have been experimenting with the documentation of meetings. The paper summarizes part of the progress that we have made in this test bed, specifically on the question of automatic transcription using large vocabulary continuous speech recognition, information access using non-keyword based methods, summarization and user interfaces. The system is capable of automatically constructing a searchable and browsable audio-visual database of meetings and provide access to these records. Alex Waibel, Michael Bett, Florian Metze, Klaus Ries 0001, Thomas Schaaf, Tanja Schultz, Hagen Soltau, Hua Yu 0008, Klaus Zechner |
ICASSP | 5 |
| 2001 | Detection of OOV words using generalized word models and a semantic class language modelabstractThis paper describes an approach to detect out-of-vocabulary words in spontaneous speech using a language model built on semantic categories and a new type of generalized word models consisting of a mixture of specific and general acoustic units. We demonstrate the construction of the generalized word models as replacements for surnames in a German spontaneous travel planning task GSST [1]. We show that the use of our generalized word models improves recognition accuracy in cases where out-of-vocabulary words appear and does not lead to a degradation of the overall recognition accuracy. In our experiments we measured recall and precision rates of OOV-detection which are close to their theoretic optimum. Furthermore, we compared the effect of using cross-word-triphones vs. using context-independent cross-word models. We show that when using generalized word models with cross-word-triphones, the expected number of consequential errors following an OOV word can be reduced significantly by 37%. Thomas Schaaf |
INTERSPEECH | 1 |
| 2000 | Confidence measure based language identificationabstractIn this paper we present a new application for confidence measures in spoken language processing. In today's computerized dialogue systems, language identification (LID) is typically achieved via dedicated modules. In our approach, LID is integrated into the speech recognizer, therefore profiting from high-level linguistic knowledge at very little extra cost. Our new approach is based on a word lattice based confidence measure (Kemp and Schaaf, 1997), which was originally devised for unsupervised training. In this work, we show that the confidence based language identification algorithm outperforms conventional score based methods. Also, this method is less dependent on the acoustic characteristics of the transmission channel than score based methods. By introducing additional parameters, unknown languages can be rejected. The proposed method is compared to a score based approach on the Verbmobil database, a three language task. Florian Metze, Thomas Kemp, Thomas Schaaf, Tanja Schultz, Hagen Soltau |
ICASSP | 3 |
| 1997 | Confidence measures for spontaneous speech recognitionabstractFor many practical applications of speech recognition systems, it is desirable to have an estimate of confidence for each hypothesized word, i.e. to have an estimate of which words of the output of the speech recognizer are likely to be correct and which are not reliable. We describe the development of the measure of the confidence tagger JANKA, which is able to provide confidence information for the words at the output of the speech recognizer JANUS-3-SR. On a spontaneous German human-to-human database, JANKA achieves a tagging accuracy of 90% at a baseline word accuracy of 82%. Thomas Schaaf, Thomas Kemp |
ICASSP | 1 |
| 1997 | Estimating confidence using word latticesabstractFor many practical applications of speech recognition systems, it is desirable to have an estimate of confidence for each hypothesized word, i.e. to have an estimate which words of the speech recognizer's output are likely to be correct and which are not reliable. Many of today's speech recognition systems use word lattices as a compact representation of a set of alternative hypothesis. We exploit the use of such word lattices as information sources for the measure-of-confidence tagger JANKA [1]. In experiments on spontaneous human -to-human speech data the use of word lattice related information significantly improves the tagging accuracy. 1. INTRODUCTION Current speech recognition systems are far from perfect. Unfortunately, number and location of the errors in their output is usually unknown. However, this information could be used in a number of applications. Examples are word selection for unsupervised adaptation schemes like MLLR [5], automatic weighting of additional, non-speec... Thomas Kemp, Thomas Schaaf |
EUROSPEECH | 2 |