VLDB 2026 Research / reviewers in the wild / expert
Klaus Zechner
dblp:69/5852
· DBLP profile ↗
33ranked-venue papers
9as first author
1since 2021 · last 2021
0000-0002-9465-6929ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 26 · 7 first-authorGraphics, computer vision, multimedia, augmented reality and games · 22 · 4 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
1 paper |
Information extraction and text analysis · 100% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% |
Topics — the 3 heaviest of 4, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Information extraction and text analysis › text classification
automated scoring |
0.1 | 1 | 2011 | Computing and Evaluating Syntactic Complexity Features for Automated Scoring of Spontaneous Non-Native Speech · ACL 2011 |
Information retrieval
text summarization |
0.0 | 1 | 2001 | Automatic Generation of Concise Summaries of Spoken Dialogues in Unrestricted Domains · SIGIR 2001 |
Information retrieval › text summarization
extractive summarization |
0.0 | 1 | 2001 | Automatic Generation of Concise Summaries of Spoken Dialogues in Unrestricted Domains · SIGIR 2001 |
Methods — techniques the papers use, named apart from their topics
syntactic parsing · 0.1feature extraction · 0.1maximum marginal relevance · 0.0TFIDF term weighting · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2021 | A Good Start is Half the Battle Won: Unsupervised Pre-training for Low Resource Children's Speech Recognition for an Interactive Reading Companion
Abhinav Misra, Anastassia Loukina, Beata Beigman Klebanov, Binod Gyawali, Klaus Zechner |
AIED (1) | 5 |
| 2020 | Do Face Masks Introduce Bias in Speech Technologies? The Case of Automated Scoring of Speaking ProficiencyabstractThe COVID-19 pandemic has led to a dramatic increase in the use of face masks worldwide. Face coverings can affect both acoustic properties of the signal as well as speech patterns and have unintended effects if the person wearing the mask attempts to use speech processing technologies. In this paper we explore the impact of wearing face masks on the automated assessment of English language proficiency. We use a dataset from a large-scale speaking test for which test-takers were required to wear face masks during the test administration, and we compare it to a matched control sample of test-takers who took the same test before the mask requirements were put in place. We find that the two samples differ across a range of acoustic measures and also show a small but significant difference in speech patterns. However, these differences do not lead to differences in human or automated scores of English language proficiency. Several measures of bias showed no differences in scores between the two groups. Anastassia Loukina, Keelan Evanini, Matthew Mulholland, Ian Blood, Klaus Zechner |
INTERSPEECH | 5 |
| 2020 | Targeted Content Feedback in Spoken Language Learning and Assessment
Klaus Zechner, Christopher Hamill |
INTERSPEECH | 2 |
| 2019 | Using Very Deep Convolutional Neural Networks to Automatically Detect Plagiarized Spoken ResponsesabstractThis study focuses on the automatic plagiarism detection in the context of high-stakes spoken language proficiency assessment, in which some test takers may attempt to game the test by memorizing prepared source materials before the test and then adapting them on-the-fly during the test to produce their spoken responses. When trying to identify such instances of plagiarism, experienced human raters attempt to find salient matching expressions that appear both in potential source materials and the test responses. This motivates an approach that visualizes a grid of lexical matches between a test response and a source and then applies state-of-the-art image recognition techniques to detect patterns of matching sequences. This study employs Inception networks-very deep convolutional neural networks-to build automatic detection models. The system achieves an F1-score of 79.6% on the class of plagiarized responses outperforming a baseline system based on word sequence matching (F1-score of 74.1%). Keelan Evanini, Yao Qian, Klaus Zechner |
ASRU | 4 |
| 2019 | Automated Estimation of Oral Reading Fluency During Summer Camp e-Book Reading with MyTurnToRead
Anastassia Loukina, Beata Beigman Klebanov, Patrick L. Lange, Yao Qian, Binod Gyawali, Nitin Madnani, Abhinav Misra, Klaus Zechner, John Sabatini 0001 |
INTERSPEECH | 8 |
| 2019 | Automatic Detection of Off-Topic Spoken Responses Using Very Deep Convolutional Neural Networks
Su-Youn Yoon, Keelan Evanini, Klaus Zechner, Yao Qian |
INTERSPEECH | 4 |
| 2019 | Development of Robust Automated Scoring Models Using Adversarial Input for Oral Proficiency Assessment
Su-Youn Yoon, Chong Min Lee, Klaus Zechner, Keelan Evanini |
INTERSPEECH | 3 |
| 2018 | Discourse Modeling of Non-Native Spontaneous Speech Using the Rhetorical Structure Theory FrameworkabstractThis study aims to model the discourse structure of spontaneous spoken responses within the context of an assessment of English speaking proficiency for non-native speakers. Rhetorical Structure Theory (RST) has been commonly used in the analysis of discourse organization of written texts; however, limited research has been conducted to date on RST annotation and parsing of spoken language, in particular, non-native spontaneous speech. Due to the fact that the measurement of discourse coherence is typically a key metric in human scoring rubrics for assessments of spoken language, we initiated a research to first obtain RST annotations on non-native spoken responses from a standardized assessment of academic English proficiency. Afterwards, based on the annotations obtained, automatic parsers were built to process non-native spontaneous speech. Finally, a set of effective features were extracted from both manually annotated and automatically generated RST trees to evaluate the discourse structure of non-native spontaneous speech, and then employed to further improve the validity of an automated speech scoring system. Binod Gyawali, James V. Bruno, Hillary Molloy, Keelan Evanini, Klaus Zechner |
SLT | 6 |
| 2017 | Combining human and automated scores for the improved assessment of non-native speech
Su-Youn Yoon, Klaus Zechner |
Speech Commun. | 2 |
| 2016 | Exploring deep learning architectures for automatically grading non-native spontaneous speechabstractWe investigate two deep learning architectures reported to have superior performance in ASR over the conventional GMM system, with respect to automatic speech scoring. We use an approximately 800-hour large-vocabulary non-native spontaneous English corpus to build three ASR systems. One system is in GMM, and two are in deep learning architectures - namely, DNN and Tandem with bottleneck features. The evaluation results show that the both deep learning systems significantly outperform the GMM ASR. These ASR systems are used as the front-end in building an automated speech scoring system. To examine the effectiveness of the deep learning ASR systems for automated scoring, another non-native spontaneous speech corpus is used to train and evaluate the scoring models. Using deep learning architectures, ASR accuracies drop significantly on the scoring corpus, whereas the performance of the scoring systems get closer to human raters, and consistently better than the GMM one. Compared to the DNN ASR, the Tandem performs slightly better on the scoring speech while it is a little less accurate on the ASR evaluation dataset. Furthermore, given the results of the improved scoring performance while using fewer scoring features, the Tandem system shows more robustness for scoring task than the DNN one. Jidong Tao, Shabnam Ghaffarzadegan, Lei Chen 0004, Klaus Zechner |
ICASSP | 4 |
| 2015 | Using bidirectional lstm recurrent neural networks to learn high-level abstractions of sequential features for automated scoring of non-native spontaneous speechabstractWe introduce a new method to grade non-native spoken language tests automatically. Traditional automated response grading approaches use manually engineered time-aggregated features (such as mean length of pauses). We propose to incorporate general time-sequence features (such as pitch) which preserve more information than time-aggregated features and do not require human effort to design. We use a type of recurrent neural network to jointly optimize the learning of high level abstractions from time-sequence features with the time-aggregated features. We first automatically learn high level abstractions from time-sequence features with a Bidirectional Long Short Term Memory (BLSTM) and then combine the high level abstractions with time-aggregated features in a Multilayer Perceptron (MLP)/Linear Regression (LR). We optimize the BLSTM and the MLP/LR jointly. We find such models reach the best performance in terms of correlation with human raters. We also find that when there are limited time-aggregated features available, our model that incorporates time-sequence features improves performance drastically. Zhou Yu 0005, Vikram Ramanarayanan, David Suendermann-Oeft, Klaus Zechner, Lei Chen 0004, Jidong Tao, Aliaksei Ivanou, Yao Qian |
ASRU | 5 |
| 2015 | Pronunciation accuracy and intelligibility of non-native speechabstractThis paper investigates the connection between intelligibility and pronunciation accuracy. We compare which words in non-native English speech are likely to be misrecognized and which words are likely to be marked as pronunciation errors. We found that only 16% of the variability in word-level intelligibility can be explained by the presence of obvious mispronunciations. In some cases, a word remained recognizable or could be identified from the context despite obvious pronunciation errors. In many other cases, the annotators were unable to identify the word when listening to the audio but did not perceive it as mispronounced when presented with its transcription. At the same time, we see high agreement when the results are aggregated across all words from the same speaker. Anastassia Loukina, Melissa Lopez, Keelan Evanini, David Suendermann-Oeft, Alexei V. Ivanov, Klaus Zechner |
INTERSPEECH | 6 |
| 2015 | Expert and crowdsourced annotation of pronunciation errors for automatic scoring systemsabstractThis paper evaluates and compares different approaches to collecting judgments about pronunciation accuracy of nonnative speech. We compare the common approach, which requires expert linguists to provide a detailed phonetic transcription of non-native English speech, with word-level judgments collected from multiple naive listeners using a crowdsourcing platform. In both cases we found low agreement between annotators on what words should be marked as errors. We compare the error detection task to a simple transcription task in which the annotators were asked to transcribe the same fragments using standard English spelling. We argue that the transcription task is a simpler and more practical way of collecting annotations which also leads to more valid data for training an automatic scoring system. Anastassia Loukina, Melissa Lopez, Keelan Evanini, David Suendermann-Oeft, Klaus Zechner |
INTERSPEECH | 5 |
| 2015 | Using F0 contours to assess nativeness in a sentence repeat taskabstractIn this paper, we conduct experiments using F0 contour features to assess the nativeness of responses provided by speakers from India and China to a Sentence Repeat task in an assessment of English speaking proficiency for non-native speakers. The results show that the coefficients from polynomial models of the pitch contours help distinguish between native and non-native speakers, especially among females. We find that the F0 contour can be represented adequately by using only basic statistical variables and the first three orders of polynomial coefficients. In addition, the most important features for classification are presented for each group of speakers. Finally, we discuss the differences among the gender-specific groups of the speakers. Keelan Evanini, Anastassia Loukina, Klaus Zechner |
INTERSPEECH | 5 |
| 2013 | Coherence Modeling for the Automated Assessment of Spontaneous Spoken Responses
Keelan Evanini, Klaus Zechner |
HLT-NAACL | 3 |
| 2012 | Exploring Content Features for Automated Speech Scoring
Shasha Xie, Keelan Evanini, Klaus Zechner |
HLT-NAACL | 3 |
| 2011 | Computing and Evaluating Syntactic Complexity Features for Automated Scoring of Spontaneous Non-Native Speech
Klaus Zechner |
ACL | 2 |
| 2011 | Evaluating prosodic features for automated scoring of non-native read speechabstractWe evaluate two types of prosodic features utilizing automatically generated stress and tone labels for non-native read speech in terms of their applicability for automated speech scoring. Both types of features have not been used in the context of automated scoring of non-native read speech to date. In our first experiment, we compute features based on a positional match between automatically identified stress and tone labels for 741 non-native read text passages with a human gold standard on the same texts read by a native speaker. Pearson correlations of up to r=0.54 between these features and human proficiency scores are observed. In our second experiment, we use stress and tone labels of the same non-native read speech corpus to compute derived features of rhythm and relative frequencies, which then again are correlated with human proficiency scores. Pearson correlations of up to r=-0.38 are observed. Klaus Zechner, Xiaoming Xi, Lei Chen 0004 |
ASRU | 1 |
| 2011 | Applying Rhythm Features to Automatically Assess Non-Native SpeechabstractSpeech rhythm measurements have been used in a limited num-ber of previous studies on automated speech assessment, an ap-proach using speech recognition technology to judge non-native speakers ’ proficiency levels. However, one of the most prob-lematic issues of these previous studies is a lack of a compar-ison of these rhythm features with other effective non-rhythm features found in decade-long previous research. In this paper, we extracted both non-rhythm and rhythm features and com-pared them with respect to their performances to predict profi-ciency scores rated by humans. We show that adding rhythm features significantly improves the performance of the scoring model based only on non-rhythm features. Index Terms: automated speech assessment, speech rhythm, prosody Lei Chen 0004, Klaus Zechner |
INTERSPEECH | 2 |
| 2011 | Using Crowdsourcing to Provide Prosodic Annotations for Non-Native SpeechabstractWe present the results of an experiment in which 2 expert and 11 naive annotators provided prosodic annotations for stress and boundary tones on a corpus of spontaneous speech produced by non-native speakers of English. The results show that agreement rates were higher for boundary tones than for stress. In addition, a crowdsourcing approach was implemented to combine the naive annotations to increase accuracy. The crowdsourcing approach was able to match expert agreement for stress (62.1%) with 3 naive annotators, and come within 7.2 % of expert agreement for boundary tones (82.4%) with 11 naive annotators. This experiment also demonstrates that noticeable improvements in naive annotations can be obtained with a small amount of additional training. Index Terms: crowdsourcing, prosodic annotation, stress, boundary tone Keelan Evanini, Klaus Zechner |
INTERSPEECH | 2 |
| 2011 | A three-stage approach to the automated scoring of spontaneous spoken responses
Derrick Higgins, Xiaoming Xi, Klaus Zechner, David M. Williamson |
Comput. Speech Lang. | 3 |
| 2010 | Predicting word accuracy for the automatic speech recognition of non-native speechabstractWe have developed an automated method that predicts the word accuracy of a speech recognition system for non-native speech, in the context of speaking proficiency scoring. A model was trained using features based on speech recognizer scores, func-tion word distributions, prosody, background noise, and speak-ing fluency. Since the method was implemented for non-native speech, fluency features, which have been used for non-native speak-ers ’ proficiency scoring, were implemented along with several feature groups used from past research. The fluency features showed promising performance by themselves, and improved the overall performance in tandem with other more traditional features. A model using stepwise regression achieved a correlation with word accuracy rates of 0.76, compared to a baseline of 0.63 using only confidence scores. A binary classifier for plac-ing utterances in high-or low-word accuracy bins achieved an accuracy of 84%, compared to a majority class baseline of 64%. Index Terms: speech recognition, word accuracy rate, non-native speakers ’ speech Su-Youn Yoon, Lei Chen 0004, Klaus Zechner |
INTERSPEECH | 3 |
| 2009 | Adapting the acoustic model of a speech recognizer for varied proficiency non-native spontaneous speech using read speech with language-specific pronunciation difficultyabstractThis paper presents a novel approach to acoustic model adaptation of a recognizer for non-native spontaneous speech in the context of recognizing candidates ’ responses in a test of spoken English. Instead of collecting and then transcribing spontaneous speech data, a read speech corpus is created where non-native speakers of English read English sentences of different degrees of pronunciation difficulty with respect to their native language. The motivation for this approach is (1) to save time and cost associated with transcribing spontaneous speech, and (2) to allow for a targeted training of the recognizer, focusing particularly on those phoneme environments which are difficult to pronounce correctly by non-native speakers and hence have a higher likelihood of being misrecognized. As a criterion for selecting the sentences to be read, we develop a novel score, the “phonetic challenge score”, consisting of a measure for native language-specific difficulties described in the second-language acquisition literature and also of a statistical measure based on the cross-entropy between phoneme sequences of the native language and English. We collected about 23,000 read sentences from 200 speakers in four language groups: Chinese, Japanese, Korean, and Spanish. We used this data for acoustic model adaptation of a spontaneous speech recognizer and compared recognition performance between the unadapted baseline and the system after adaptation on a held-out set from the English test responses data set. The results show that using this targeted read speech material for acoustic model adaptation does reduce the word error rate significantly for two of four language groups of the spontaneous speech test set, while changes of the two other language groups are not significant. Insdex Terms: acoustic model adaptation, non-native spontaneous speech, cross-lingual phonetic difficulty Klaus Zechner, Derrick Higgins, René Lawless, Yoko Futagi, Sarah Ohls, George Ivanov |
INTERSPEECH | 1 |
| 2009 | Improved pronunciation features for construct-driven assessment of non-native spontaneous speech
Lei Chen 0004, Klaus Zechner, Xiaoming Xi |
HLT-NAACL | 2 |
| 2009 | Automatic scoring of non-native spontaneous speech in tests of spoken English
Klaus Zechner, Derrick Higgins, Xiaoming Xi, David M. Williamson |
Speech Commun. | 1 |
| 2006 | Towards Automatic Scoring of Non-Native Spontaneous Speech
Klaus Zechner, Isaac Bejar |
HLT-NAACL | 1 |
| 2003 | Spoken language condensation in the 21st centuryabstractWhile the field of Information Retrieval originally had the search for the most relevant documents in mind, it has become increasingly clear that in many instances, what the user wants is a piece of coherent information, derived from a set of relevant documents and possibly other sources. Reducing relevant documents, passages, and sentences to their core is the task of text summarization or information condensation. Applying text-based technologies to speech is not always workable and often not enough to capture speech specific phenomena. In this paper, we will contrast speech summarization with text summarization, give an overview of the history of speech summarization, its current state, and, finally, sketch possible avenues as well as remaining challenges in future research. Klaus Zechner |
INTERSPEECH | 1 |
| 2002 | Automatic Summarization of Open-Domain Multiparty Dialogues in Diverse GenresabstractAutomatic summarization of open-domain spoken dialogues is a relatively new research area. This article introduces the task and the challenges involved and motivates and presents an approach for obtaining automatic-extract summaries for human transcripts of multiparty dialogues of four different genres, without any restriction on domain. We address the following issues, which are intrinsic to spoken-dialogue summarization and typically can be ignored when summarizing written text such as news wire data: (1) detection and removal of speech disfluencies; (2) detection and insertion of sentence boundaries; and (3) detection and linking of cross-speaker information units (question-answer pairs). A system evaluation is performed using a corpus of 23 dialogue excerpts with an average duration of about 10 minutes, comprising 80 topical segments and about 47,000 words total. The corpus was manually annotated for relevant text spans by six human annotators. The global evaluation shows that for the two more informal genres, our summarization system using dialogue-specific components significantly outperforms two baselines: (1) a maximum-marginal-relevance ranking algorithm using TF*IDF term weighting, and (2) a LEAD baseline that extracts the first n words from a text. Klaus Zechner |
Comput. Linguistics | 1 |
| 2001 | Advances in automatic meeting record creation and accessabstractOral communication is transient, but many important decisions, social contracts and fact findings are first carried out in an oral setup, documented in written form and later retrieved. At Carnegie Mellon University's Interactive Systems Laboratories we have been experimenting with the documentation of meetings. The paper summarizes part of the progress that we have made in this test bed, specifically on the question of automatic transcription using large vocabulary continuous speech recognition, information access using non-keyword based methods, summarization and user interfaces. The system is capable of automatically constructing a searchable and browsable audio-visual database of meetings and provide access to these records. Alex Waibel, Michael Bett, Florian Metze, Klaus Ries 0001, Thomas Schaaf, Tanja Schultz, Hagen Soltau, Hua Yu 0008, Klaus Zechner |
ICASSP | 9 |
| 2001 | Automatic Generation of Concise Summaries of Spoken Dialogues in Unrestricted DomainsabstractAutomatic summarization of open domain spoken dialogues is a new research area. This paper introduces the task, the challenges involved, and presents an approach to obtain automatic extract summaries for multi-party dialogues of four different genres, without any restriction on domain. We address the following issues which are intrinsic to spoken dialogue summarization and typically can be ignored when summarizing written text such as newswire data: (i) detection and removal of speech disfluencies; (ii) detection and insertion of sentence boundaries; (iii) detection and linking of cross-speaker information units (question-answer pairs). A global system evaluation using a corpus of 23 relevance annotated dialogues containing 80 topical segments shows that for the two more informal genres, our summarization system using dialogue specific components significantly outperforms a baseline using TFIDF term weighting with maximum marginal relevance ranking (MMR). Klaus Zechner |
SIGIR | 1 |
| 2000 | DIASUMM: Flexible Summarization of Spontaneous Dialogues in Unrestricted Domains
Klaus Zechner, Alex Waibel |
COLING | 1 |
| 1998 | A discourse coding scheme for conversational SpanishabstractThis paper describes a 3-level manual discourse coding scheme that we have devised for manual tagging of the CallHome Spanish (CHS) and CallFriend Spanish (CFS) databases used in the CLARITY project. The goal of CLARITY is to explore the use of discourse structure in understanding conversational speech. The project combines empirical methods for dialogue processing with state-of-the art LVCSR (using the JANUS recognizer). The three levels of the coding scheme are (1) a speech act level consisting of a tag set extended from DAMSL and Switchboard; (2) dialogue game level defined by initiative and intention; and (3) an activity level defined within topic units. The manually tagged dialogues are used to train automatic classifiers. We present preliminary results for statement categorization, and give an in-progress report of automatic speech act classification and topic boundary identification. 1. INTRODUCTION To appear in: International Conference on Spoken Language Processing (ICSLP'98... Lori S. Levin, Ann E. Thymé-Gobbel, Alon Lavie, Klaus Ries 0001, Klaus Zechner |
ICSLP | 5 |
| 1996 | Fast Generation of Abstracts from General Domain Text Corpora by Extracting Relevant Sentences
Klaus Zechner |
COLING | 1 |