EDBT 2026 Demo / reviewers in the wild / expert
Edward W. D. Whittaker
dblp:98/5870
· DBLP profile ↗
29ranked-venue papers
11as first author
0since 2021 · last 2014
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 8 first-authorGraphics, computer vision, multimedia, augmented reality and games · 19 · 7 first-authorDatabases, data management, data science and information retrieval · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Information extraction and text analysis · 54% Language models and text generation · 32% Question answering and dialogue systems · 14% | |
| Interdisciplinary, comprehensive, and emerging computing
3 papers |
Computing education · 66% Computational social science and digital humanities · 34% |
Topics — the 4 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation › text generation
grammatical error correction |
0.2 | 1 | 2014 | Correcting Preposition Errors in Learner English Using Error Case Frames and Feedback Messages · ACL (1) 2014 |
Computational social science and digital humanities
historical linguistics |
0.2 | 1 | 2013 | Reconstructing an Indo-European Family Tree from Non-native English Texts · ACL (1) 2013 |
Natural language and speech › Information extraction and text analysis › data annotation
corpus annotation |
0.1 | 1 | 2011 | Creating a manually error-tagged and shallow-parsed learner corpus · ACL 2011 |
Natural language and speech › Question answering and dialogue systems › reasoning-based question answering
why-question answering |
0.1 | 1 | 2008 | Collecting a Why-Question Corpus for Development and Evaluation of an Automatic QA-System · ACL 2008 |
Methods — techniques the papers use, named apart from their topics
learner-native corpus comparison · 0.4error case frames · 0.4phylogenetic reconstruction · 0.3shallow parsing · 0.2error tagging · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2014 | Correcting Preposition Errors in Learner English Using Error Case Frames and Feedback MessagesabstractThis paper presents a novel framework called error case frames for correcting preposition errors. They are case frames specially designed for describing and cor-recting preposition errors. Their most dis-tinct advantage is that they can correct er-rors with feedback messages explaining why the preposition is erroneous. This pa-per proposes a method for automatically generating them by comparing learner and native corpora. Experiments show (i) au-tomatically generated error case frames achieve a performance comparable to con-ventional methods; (ii) error case frames are intuitively interpretable and manually modifiable to improve them; (iii) feedback messages provided by error case frames are effective in language learning assis-tance. Considering these advantages and the fact that it has been difficult to provide feedback messages by automatically gen-erated rules, error case frames will likely be one of the major approaches for prepo-sition error correction. 1 Ryo Nagata, Mikko Vilenius, Edward W. D. Whittaker |
ACL (1) | 3 |
| 2013 | Reconstructing an Indo-European Family Tree from Non-native English Texts
Ryo Nagata, Edward W. D. Whittaker |
ACL (1) | 2 |
| 2012 | Question answering using statistical language modelling
Matthias H. Heie, Edward W. D. Whittaker, Sadaoki Furui |
Comput. Speech Lang. | 2 |
| 2011 | Creating a manually error-tagged and shallow-parsed learner corpus
Ryo Nagata, Edward W. D. Whittaker, Vera Sheinman |
ACL | 2 |
| 2008 | Collecting a Why-Question Corpus for Development and Evaluation of an Automatic QA-System
Joanna Mrozinski, Edward W. D. Whittaker, Sadaoki Furui |
ACL | 2 |
| 2007 | A language modeling approach to question answering on speech transcriptsabstractThis paper presents a language modeling approach to sentence retrieval for Question Answering (QA) that we used in Question Answering on speech transcripts (QAst), a pilot task at the Cross Language Evaluation Forum (CLEF) evaluations 2007. A language model (LM) is generated for each sentence and these models are combined with document LMs to take advantage of contextual information. A query expansion technique using class models is proposed and included in our framework. Finally, our method's impact on exact answer extraction is evaluated. We show that combining sentence LMs with document LMs significantly improves sentence retrieval performance, and that this sentence retrieval approach leads to better answer extraction performance. Matthias H. Heie, Edward W. D. Whittaker, Josef R. Novak, Sadaoki Furui |
ASRU | 2 |
| 2006 | Rapid Development of Web-Based Monolingual Question Answering Systems
Edward W. D. Whittaker, Julien Hamonic, Tor Klingberg, Sadaoki Furui |
ECIR | 1 |
| 2006 | Topic and Stylistic Adaptation for Speech SummarisationabstractContemporary approaches to automatic speech summarisation comprise several components, among them a linguistic model (LiM) component, which is unrelated to the language model used during the recognition process. This LiM component assigns a probability to word sequences from the source text according to their likelihood of appearing in the summarised text. In this paper we investigate LiM topic and stylistic adaptation using combinations of LiMs each trained on different adaptation data. Experiments are performed on 9 talks from the TED corpus of Eurospeech conference presentations, as well as 5 news stories from CNN broadcast news data, for all of which human (TRS) and speech recogniser (ASR) transcriptions along with human summaries were used. In all ASR cases, summarisation accuracy (SumACCY) of automatically generated summaries was significantly improved by automatic LiM adaptation, with relative improvements of at least 2.5% in all experiments Pierre Chatain, Edward W. D. Whittaker, Joanna Mrozinski, Sadaoki Furui |
ICASSP (1) | 2 |
| 2006 | Automatic Sentence Segmentation of Speech for Automatic SummarizationabstractThis paper presents an automatic sentence segmentation method for an automatic speech summarization system. The segmentation method is based on combining word- and class-based statistical language models to predict sentence and non-sentence boundaries. We study both the performance of the sentence segmentation system itself and the effect of the segmentation on the summarization accuracy. The sentence segmentation is done by modelling the probability of a sentence boundary given a certain word history with language models trained on transcriptions and texts from several sources. The resulting segmented data is used as the input to an existing automatic summarization system to determine the effect it has on the summarization process. We conduct all our experiments with two types of evaluation data: broadcast news and lecture transcriptions. The automatic summarizations are created with different sentence segmentations and different summarization ratios (30% and 40%) and evaluated by comparing them to human-made summaries. We show that a proper sentence segmentation is essential to achieve good performance with an automatic summarization system Joanna Mrozinski, Edward W. D. Whittaker, Pierre Chatain, Sadaoki Furui |
ICASSP (1) | 2 |
| 2006 | Perplexity based linguistic model adaptation for speech summarisation
Pierre Chatain, Edward W. D. Whittaker, Joanna Mrozinski, Sadaoki Furui |
INTERSPEECH | 2 |
| 2006 | Class Model Adaptation for Speech Summarisation
Pierre Chatain, Edward W. D. Whittaker, Joanna Mrozinski, Sadaoki Furui |
HLT-NAACL | 2 |
| 2006 | Factoid Question Answering with Web, Mobile and Speech Interfaces
Edward W. D. Whittaker, Joanna Mrozinski, Sadaoki Furui |
HLT-NAACL | 1 |
| 2005 | A Statistical Classification Approach to Question Answering using Web DataabstractIn this paper we treat question answering (QA) as a classification problem. Our motivation is to build systems for many languages without the need for highly tuned linguistic modules. Consequently, word tokens and Web data are used extensively but no explicit linguistic knowledge is incorporated. A mathematical model for answer retrieval, answer classification and answer length prediction is derived. The TREC 2002 QA task is used for system development where 33% of questions are answered correctly. Performance is then evaluated on the factoid questions of the TREC 2003 QA task where 23% of questions were answered correctly, which would rank the system in the top 10 of contemporary QA systems on the same task Edward W. D. Whittaker, Sadaoki Furui, Dietrich Klakow |
CW | 1 |
| 2005 | Cluster-based modeling for ubiquitous speech recognitionabstractIn order to realize speech recognition systems that can achieve high recognition accuracy for ubiquitous speech, it is crucial to make the systems flexible enough to cope with a large variability of spontaneous speech. This paper investigates two speech recognition methods that can adapt to speech variation using a large number of models trained based on clustering techniques; one automatically builds a model adapted to input speech using recognition hypotheses and clustered models, and the other directly uses clustered models in parallel. Both methods have been confirmed to be effective by evaluation experiments using presentation speech. Although the latter method needs a large amount of computation, it has an advantage in that it can be applied to online recognition, since it does not need recognition hypotheses. The former method can also be applied to online recognition, if the text of proceedings for the presentation can be used in place of recognition hypotheses. 1. Sadaoki Furui, Tomohisa Ichiba, Takahiro Shinozaki, Edward W. D. Whittaker, Koji Iwano |
INTERSPEECH | 4 |
| 2005 | Language model adaptation for resource deficient languages using translated dataabstractText corpus size is an important issue when building a language model (LM). This is a particularly important issue for languages where little data is available. This paper introduces a technique to improve a LM built using a small amount of task dependent text with the help of a machine-translated text corpus. Perplexity experiments were performed using data, machine translated (MT) from English to French on a sentence-by-sentence basis and using dictionary lookup on a word-by-word basis. Then perplexity and word error rate experiments using MT data from English to Icelandic were done on a word-by-word basis. For the latter, the baseline word error rate was 44.0%. LM interpolation reduced word error rate significantly to 39.2%. Arnar Thor Jensson, Edward W. D. Whittaker, Koji Iwano, Sadaoki Furui |
INTERSPEECH | 2 |
| 2004 | Unsupervised language model adaptation methods for spontaneous speech
Luc Lussier, Edward W. D. Whittaker, Sadaoki Furui |
INTERSPEECH | 2 |
| 2003 | Lossless compression of language model structure and word identifiersabstractVery large reductions in language model memory requirements have recently been reported for large vocabulary continuous speech recognition applications through the pruning and quantization of the floating-point components of the language model: the probabilities and back-off weights. In this paper that work is extended through the compression of the integer components: the word identifiers and storage structures. A novel algorithm is presented for converting ordered lists of monotonically increasing integer values (such as are commonly found in language models) into variable-bit width tree structures such that the most memory efficient configuration is obtained for each original list. By applying this new technique together with the techniques reported previously we obtain an 86% reduction in language model size to 10Mb for no increase in word error rate on the DARPA Hub4 1998 task and a 0.5% absolute increase on the Hub4 1997 task. Bhiksha Raj, Edward W. D. Whittaker |
ICASSP (1) | 2 |
| 2003 | Language modelling for Russian and English using words and classes
Edward W. D. Whittaker, Philip C. Woodland |
Comput. Speech Lang. | 1 |
| 2003 | Erratum: Language modelling for Russian and English using words and classes [Computer Speech and Language 17 (2003) 87-104]
Edward W. D. Whittaker, Philip C. Woodland |
Comput. Speech Lang. | 1 |
| 2002 | Efficient construction of long-range language models using log-linear interpolationabstractIn this paper we examine the construction of long-range language models using log-linear interpolation and how this can be achieved effectively. Particular attention is paid to the efficient computation of the normalisation in the models. Using the Penn Treebank for experiments we argue that the perplexity performance demonstrated recently in the literature using grammar-based approaches can actually be achieved with an appropriately smoothed 4-gram language model. Using such a model as the baseline, we demonstrate how further improvements can be obtained using loglinear interpolation to combine distance word and class models. We also examine the performance of similar model combinations for rescoring word lattices on a medium-sized vocabulary Wall Street Journal task. 1. Edward W. D. Whittaker, Dietrich Klakow |
INTERSPEECH | 1 |
| 2001 | Efficient class-based language modelling for very large vocabulariesabstractInvestigates the perplexity and word error rate performance of two different forms of class model and the respective data-driven algorithms for obtaining automatic word classifications. The computational complexity of the algorithm for the 'conventional' two-sided class model is found to be unsuitable for very large vocabularies (>100k) or large numbers of classes (>2000). A one-sided class model is therefore investigated and the complexity of its algorithm is found to be substantially less in such situations. Perplexity results are reported on both English and Russian data. For the latter both 65k and 430k vocabularies are used. Lattice rescoring experiments are also performed on an English language broadcast news task. These experimental results show that both models, when interpolated with a word model, perform similarly well. Moreover, classifications are obtained for the one-sided model in a fraction of the time required by the two-sided model, especially for very large vocabularies. Edward W. D. Whittaker, Philip C. Woodland |
ICASSP | 1 |
| 2001 | Quantization-based language model compressionabstractThis paper describes two techniques for reducing the size of statistical back-off gram language models in computer memory. Language model compression is achieved through a combination of quantizing language model probabilities and back-off weights and the pruning of parameters that are determined to be unnecessary after quantization. The recognition performance of the original and compressed language models is evaluated across three different language models and two different recognition tasks. The results show that the language models can be compressed by up to 60% of their original size with no significant loss in recognition performance. Moreover, the techniques that are described provide a principled method with which to compress language models further while minimising degradation in recognition performance. Edward W. D. Whittaker, Bhiksha Raj |
INTERSPEECH | 1 |
| 2001 | Comparison of width-wise and length-wise language model compressionabstractIn this paper we investigate the extent to which Katz backoff language models can be compressed through a combination of parameter quantization (width-wise compression) and parameter pruning (length-wise compression) methods while preserving performance. We compare the compression and performance that is achieved using entropy-based pruning against that achieved using only parameter quantization. We then compare combinations of both methods. It is shown that a broadcast news language model can be compressed by up to 83 % to only 12.6Mb with no loss in performance on a broadcast news task. Compressing the language model further by quantization to 10.3Mb resulted in only a 0.4 % degradation in word error rate which is better than can be achieved through entropy-based pruning alone. 1. Edward W. D. Whittaker, Bhiksha Raj |
INTERSPEECH | 1 |
| 2000 | An experimental study of an audio indexing system for the webabstractWe have developed a speech recognition based audio search engine for indexing spoken documents found on the World Wide Web. Our site (http://www.compaq.com/speechbot) indexes around 20 news and talk radio shows covering a wide range of topics, speaking styles and acoustic conditions from a selection of public Web sites with multimedia archives. In this paper, we describe our system and its performance, focusing on the speech recognition and retrieval aspects. We describe our training procedure in some detail and report our historical error rate since the site launch. We also investigate the impact of Out Of Vocabulary (OOV) words. Finally we report the results of retrieval experiments which demonstrate that our system can index effectively. Beth Logan, Pedro J. Moreno 0001, Jean-Manuel Van Thong, Edward W. D. Whittaker |
INTERSPEECH | 4 |
| 2000 | Particle-based language modellingabstractThis paper investigates the use of particle (sub-word) N-grams for language modelling. One linguistics-based and two datadriven algorithms are presented and evaluated in terms of perplexity for Russian and English. Interpolating word trigram and particle 6-gram models gives up to a 7.5% perplexity reduction over the baseline word trigram model for Russian. Lattice rescoring experiments are also performed on 1997 DARPA Hub4 evaluation lattices where the interpolated model gives a 0.4% absolute reduction in word error rate over the baseline word trigram model. 1. INTRODUCTION Most of the current approaches to language modelling for speech recognition tend to use words, or classes of words, as the modelling units. Words are a logical choice, since it is ultimately words that are to be output by a speech recognition system, but they are not necessarily the best units for capturing dependencies in a text. The optimal set of units will inevitably depend on the language, the sparsity of the ... Edward W. D. Whittaker, Philip C. Woodland |
INTERSPEECH | 1 |
| 1999 | The 1998 HTK system for transcription of conversational telephone speechabstractThis paper describes the 1998 HTK large vocabulary speech recognition system for conversational telephone speech as used in the NIST 1998 Hub5E evaluation. Front-end and language modelling experiments conducted using various training and test sets from both the Switchboard and Callhome English corpora are presented. Our complete system includes reduced bandwidth analysis, side-based cepstral feature normalisation, vocal tract length normalisation (VTLN), triphone and quinphone hidden Markov models (HMMs) built using speaker adaptive training (SAT), maximum likelihood linear regression (MLLR) speaker adaptation and a confidence score based system combination. A detailed description of the complete system together with experimental results for each stage of our multi-pass decoding scheme is presented. The word error rate obtained is almost 20% better than our 1997 system on the development set. Thomas Hain, Philip C. Woodland, Thomas Niesler, Edward W. D. Whittaker |
ICASSP | 4 |
| 1999 | Improvements in accuracy and speed in the HTK broadcast news transcription system
Philip C. Woodland, J. J. Odell, Thomas Hain, Gareth L. Moore, Thomas Niesler, Andreas Tuerk, Edward W. D. Whittaker |
EUROSPEECH | 7 |
| 1998 | Comparison of part-of-speech and automatically derived category-based language models for speech recognitionabstractThis paper compares various category-based language models when used in conjunction with a word-based trigram by means of linear interpolation. Categories corresponding to parts-of-speech as well as automatically clustered groupings are considered. The category-based model employs variable-length n-grams and permits each word to belong to multiple categories. Relative word error rate reductions of between 2 and 7% over the baseline are achieved in N-best rescoring experiments on the Wall Street Journal corpus. The largest improvement is obtained with a model using automatically determined categories. Perplexities continue to decrease as the number of different categories is increased, but improvements in the word error rate reach an optimum. Thomas Niesler, Edward W. D. Whittaker, Philip C. Woodland |
ICASSP | 2 |
| 1998 | Comparison of language modelling techniques for Russian and EnglishabstractIn this paper the main differences between language modelling of Russian and English are examined. A Russian corpus and a comparable English corpus are described. The effects of high inflectionality in Russian and the relationship between the outof -vocabulary rate and vocabulary size are investigated. Standard word and class N-gram language modelling techniques are applied to the two corpora and perplexity results are reported. A novel approach to the modelling of inflected languages is proposed and its efficacy compared with the other techniques. 1. INTRODUCTION Much work has been conducted in recent years on language modelling techniques for speech recognition of English. In contrast, less commercially attractive yet widely spoken languages like Russian have received comparatively little attention in the literature (the first reported large-vocabulary recogniser for Russian appeared only recently[3]). Moreover, there are important difficulties with modelling Russian which are also... Edward W. D. Whittaker, Philip C. Woodland |
ICSLP | 1 |