VLDB 2026 Research / reviewers in the wild / expert
William J. Byrne
dblp:b/WilliamJByrne · also Bill Byrne, William Byrne
· DBLP profile ↗
125ranked-venue papers
9as first author
21since 2021 · last 2026
0000-0003-1896-4492ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 97 · 6 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 48 · 4 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 4Human-computer interaction and ubiquitous computing · 3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Benchmarking Deflection and Hallucination in Large Vision-Language ModelsabstractNicholas Moratelli, Christopher Davis, Leonardo F. R. Ribeiro, Bill Byrne, Gonzalo Iglesias. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Nicholas Moratelli, Leonardo F. R. Ribeiro, William J. Byrne, Gonzalo Iglesias |
ACL (1) | 4 |
| 2026 | What Does LLM Refinement Actually Improve? A Systematic Study on Document-Level Literary TranslationabstractShaomu Tan, Dawei Zhu, Ke Tran, Michael Denkowski, Sony Trenous, Leonardo F. R. Ribeiro, Bill Byrne, Felix Hieber. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Shaomu Tan, Ke Tran, Michael J. Denkowski, Sony Trenous, Leonardo F. R. Ribeiro, William J. Byrne, Felix Hieber |
ACL (1) | 7 |
| 2026 | Retrieval-Augmented Defense: Adaptive and Controllable Jailbreak Prevention for Large Language ModelsabstractLarge Language Models (LLMs) remain vulnerable to jailbreak attacks, which attempt to elicit harmful responses from LLMs.The evolving nature and diversity of these attacks pose many challenges for defense systems, including (1) adaptation to counter emerging attack strategies without costly retraining, and (2) control of the trade-off between safety and utility.To address these challenges, we propose Retrieval-Augmented Defense (RAD), a novel framework for jailbreak detection that incorporates a database of known attack examples into Retrieval-Augmented Generation, which is used to infer the underlying, malicious user query and jailbreak strategy used to attack the system.RAD enables training-free updates for newly discovered jailbreak strategies and provides a mechanism to balance safety and utility.Experiments on StrongREJECT show that RAD substantially reduces the effectiveness of strong jailbreak attacks such as PAP and PAIR while maintaining low rejection rates for benign queries.We propose a novel evaluation scheme and show that RAD achieves a robust safety-utility trade-off across a range of operating points in a controllable manner. 1 This paper contains harmful jailbreak contents for demonstration purposes that can be offensive. Jinghong Chen, Jingbiao Mei, Weizhe Lin, William J. Byrne |
ACL (1) | 5 |
| 2025 | Learning to Reason Over Time: Timeline Self-Reflection for Improved Temporal Reasoning in Language ModelsabstractLarge Language Models (LLMs) have emerged as powerful tools for generating coherent text, understanding context, and performing reasoning tasks.However, they struggle with temporal reasoning, which requires processing time-related information such as event sequencing, durations, and inter-temporal relationships.These capabilities are critical for applications including question answering, scheduling, and historical analysis.In this paper, we introduce TISER, a novel framework that enhances the temporal reasoning abilities of LLMs through a multi-stage process that combines timeline construction with iterative self-reflection.Our approach leverages test-time scaling to extend the length of reasoning traces, enabling models to capture complex temporal dependencies more effectively.This strategy not only boosts reasoning accuracy but also improves the traceability of the inference process.Experimental results demonstrate state-of-the-art performance across multiple benchmarks, including out-of-distribution test sets, and reveal that TISER enables smaller open-source models to surpass larger closed-weight models on challenging temporal reasoning tasks. 1 * Work done during an internship at Amazon.Now at Microsoft. Adrian Bazaga, Rexhina Blloshmi, William J. Byrne, Adrià de Gispert |
ACL (1) | 3 |
| 2025 | RAGferee: Building Contextual Reward Models for Retrieval-Augmented GenerationabstractAndrei Catalin Coman, Ionut Teodor Sorodoc, Leonardo F. R. Ribeiro, Bill Byrne, James Henderson, Adrià de Gispert. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Andrei Catalin, Ionut Sorodoc, Leonardo F. R. Ribeiro, William J. Byrne, Adrià de Gispert |
EMNLP | 4 |
| 2025 | Robust Adaptation of Large Multimodal Models for Retrieval Augmented Hateful Meme DetectionabstractHateful memes have become a significant concern on the Internet, necessitating robust automated detection systems. While Large Multimodal Models (LMMs) have shown promise in hateful meme detection, they face notable challenges like sub-optimal performance and limited out-of-domain generalization capabilities. Recent studies further reveal the limitations of both supervised fine-tuning (SFT) and in-context learning when applied to LMMs in this setting. To address these issues, we propose a robust adaptation framework for hateful meme detection that enhances in-domain accuracy and cross-domain generalization while preserving the general vision-language capabilities of LMMs. Analysis reveals that our approach achieves improved robustness under adversarial attacks compared to SFT models. Experiments on six meme classification datasets show that our approach achieves state-of-the-art performance, outperforming larger agentic systems.Moreover, our method generates higher-quality rationales for explaining hateful content compared to standard SFT, enhancing model interpretability. Code available at https://github.com/JingbiaoMei/RGCL Jingbiao Mei, Jinghong Chen, Weizhe Lin, William J. Byrne |
EMNLP | 5 |
| 2025 | On Extending Direct Preference Optimization to Accommodate TiesabstractWe derive and investigate two DPO variants that explicitly model the possibility of declaring a tie in pair-wise comparisons. We replace the Bradley-Terry model in DPO with two well-known modeling extensions, by Rao and Kupper and by Davidson, that assign probability to ties as alternatives to clear preferences. Our experiments in neural machine translation and summarization show that explicitly labeled ties can be added to the datasets for these DPO variants without the degradation in task performance that is observed when the same tied pairs are presented to DPO. We find empirically that the inclusion of ties leads to stronger regularization with respect to the reference policy as measured by KL divergence, and we see this even for DPO in its original form. We provide a theoretical explanation for this regularization effect using ideal DPO policy theory. We further show performance improvements over DPO in translation and mathematical reasoning using our DPO variants. We find it can be beneficial to include ties in preference optimization rather than simply discard them, as is done in common practice. Jinghong Chen, Weizhe Lin, Jingbiao Mei, Chenxu Lyu, William J. Byrne |
NeurIPS | 6 |
| 2024 | PreFLMR: Scaling Up Fine-Grained Late-Interaction Multi-modal RetrieversabstractLarge Multimodal Models (LMMs) excel in natural language and visual understanding but are challenged by exacting tasks such as Knowledge-based Visual Question Answering (KB-VQA) which involve the retrieval of relevant information from document collections to use in shaping answers to questions.We present an extensive training and evaluation framework, M2KR, for KB-VQA.M2KR contains a collection of vision and language tasks which we have incorporated into a single suite of benchmark tasks for training and evaluating general-purpose multi-modal retrievers.We use M2KR to develop PreFLMR, a pretrained version of the recently developed Finegrained Late-interaction Multi-modal Retriever (FLMR) approach to KB-VQA, and we report new state-of-the-art results across a range of tasks.We also present investigations into the scaling behaviors of PreFLMR intended to be useful in future developments in generalpurpose multi-modal retrievers.The code, demo, dataset, and pre-trained checkpoints are available at https://preflmr.github.io/. Weizhe Lin, Jingbiao Mei, Jinghong Chen, William J. Byrne |
ACL (1) | 4 |
| 2024 | Improving Hateful Meme Detection through Retrieval-Guided Contrastive LearningabstractHateful memes have emerged as a significant concern on the Internet.Detecting hateful memes requires the system to jointly understand the visual and textual modalities.Our investigation reveals that the embedding space of existing CLIP-based systems lacks sensitivity to subtle differences in memes that are vital for correct hatefulness classification.We propose constructing a hatefulness-aware embedding space through retrieval-guided contrastive training.Our approach achieves state-of-theart performance on the HatefulMemes dataset with an AUROC of 87.0, outperforming much larger fine-tuned large multimodal models.We demonstrate a retrieval-based hateful memes detection system, which is capable of identifying hatefulness based on data unseen in training.This allows developers to update the hateful memes detection system by simply adding new examples without retraining -a desirable feature for real services in the constantly evolving landscape of hateful memes on the Internet.This paper contains content for demonstration purposes that may be disturbing for some readers. Jingbiao Mei, Jinghong Chen, Weizhe Lin, William J. Byrne, Marcus Tomalin |
ACL (1) | 4 |
| 2024 | The Fine-Tuning Paradox: Boosting Translation Quality Without Sacrificing LLM AbilitiesabstractFine-tuning large language models (LLMs) for machine translation has shown improvements in overall translation quality.However, it is unclear what is the impact of fine-tuning on desirable LLM behaviors that are not present in neural machine translation models, such as steerability, inherent document-level translation abilities, and the ability to produce less literal translations.We perform an extensive translation evaluation on the LLaMA and Falcon family of models with model size ranging from 7 billion up to 65 billion parameters.Our results show that while fine-tuning improves the general translation quality of LLMs, several abilities degrade.In particular, we observe a decline in the ability to perform formality steering, to produce technical translations through few-shot examples, and to perform documentlevel translation.On the other hand, we observe that the model produces less literal translations after fine-tuning on parallel data.We show that by including monolingual data as part of the fine-tuning data we can maintain the abilities while simultaneously enhancing overall translation quality.Our findings emphasize the need for fine-tuning strategies that preserve the benefits of LLMs for machine translation. David Stap, Eva Hasler, William J. Byrne, Christof Monz, Ke Tran |
ACL (1) | 3 |
| 2024 | A Preference-driven Paradigm for Enhanced Translation with Large Language ModelsabstractDawei Zhu, Sony Trenous, Xiaoyu Shen, Dietrich Klakow, Bill Byrne, Eva Hasler. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Sony Trenous, Xiaoyu Shen 0001, Dietrich Klakow, William J. Byrne, Eva Hasler |
NAACL-HLT | 5 |
| 2023 | An Inner Table Retriever for Robust Table Question AnsweringabstractRecent years have witnessed the thriving of pretrained Transformer-based language models for understanding semi-structured tables, with several applications, such as Table Question Answering (TableQA).These models are typically trained on joint tables and surrounding natural language text, by linearizing table content into sequences comprising special tokens and cell information.This yields very long sequences which increase system inefficiency, and moreover, simply truncating long sequences results in information loss for downstream tasks.We propose Inner Table Retriever (ITR), 1 a generalpurpose approach for handling long tables in TableQA that extracts sub-tables to preserve the most relevant information for a question.We show that ITR can be easily integrated into existing systems to improve their accuracy with up to 1.3-4.8%and achieve state-of-the-art results in two benchmarks, i.e., 63.4% in Wik-iTableQuestions and 92.1% in WikiSQL.Additionally, we show that ITR makes TableQA systems more robust to reduced model capacity and to different ordering of columns and rows. * Work done as an intern at Amazon Alexa AI. Weizhe Lin, Rexhina Blloshmi, William J. Byrne, Adrià de Gispert, Gonzalo Iglesias |
ACL (1) | 3 |
| 2023 | Uniform Training and Marginal Decoding for Multi-Reference Question-Answer GenerationabstractQuestion generation is an important task that helps to improve question answering performance and augment search interfaces with possible suggested questions. While multiple approaches have been proposed for this task, none addresses the goal of generating a diverse set of questions given the same input context. The main reason for this is the lack of multi-reference datasets for training such models. We propose to bridge this gap by seeding a baseline question generation model with named entities as candidate answers. This allows us to automatically synthesize an unlimited number of question-answer pairs. We then propose an approach designed to leverage such multi-reference annotations, and demonstrate its advantages over the standard training and decoding strategies used in question generation. An experimental evaluation on synthetic, as well as manually annotated data shows that our approach can be used in creating a single generative model that produces a diverse set of question-answer pairs per input sentence. Svitlana Vakulenko, William J. Byrne, Adrià de Gispert |
ECAI | 2 |
| 2023 | Fine-grained Late-interaction Multi-modal Retrieval for Retrieval Augmented Visual Question AnsweringabstractKnowledge-based Visual Question Answering (KB-VQA) requires VQA systems to utilize knowledge from external knowledge bases to answer visually-grounded questions. Retrieval-Augmented Visual Question Answering (RA-VQA), a strong framework to tackle KB-VQA, first retrieves related documents with Dense Passage Retrieval (DPR) and then uses them to answer questions. This paper proposes Fine-grained Late-interaction Multi-modal Retrieval (FLMR) which significantly improves knowledge retrieval in RA-VQA. FLMR addresses two major limitations in RA-VQA's retriever: (1) the image representations obtained via image-to-text transforms can be incomplete and inaccurate and (2) similarity scores between queries and documents are computed with one-dimensional embeddings, which can be insensitive to finer-grained similarities.
FLMR overcomes these limitations by obtaining image representations that complement those from the image-to-text transform using a vision model aligned with an existing text-based retriever through a simple alignment network. FLMR also encodes images and questions using multi-dimensional embeddings to capture finer-grained similarities between queries and documents.
FLMR significantly improves the original RA-VQA retriever's PRRecall@5 by approximately 8\%. Finally, we equipped RA-VQA with two state-of-the-art large multi-modal/language models to achieve $\sim62$% VQA score in the OK-VQA dataset. Weizhe Lin, Jinghong Chen, Jingbiao Mei, Alexandru Coca, William J. Byrne |
NeurIPS | 5 |
| 2023 | Grounding Description-Driven Dialogue State Trackers with Knowledge-Seeking TurnsabstractAlexandru Coca, Bo-Hsiang Tseng, Jinghong Chen, Weizhe Lin, Weixuan Zhang, Tisha Anders, Bill Byrne. Proceedings of the 24th Annual Meeting of the Special Interest Group on Discourse and Dialogue. 2023. Alexandru Coca, Bo-Hsiang Tseng, Jinghong Chen, Weizhe Lin, Weixuan Zhang, Tisha Anders, William J. Byrne |
SIGDIAL | 7 |
| 2022 | Retrieval Augmented Visual Question Answering with Outside KnowledgeabstractOutside-Knowledge Visual Question Answering (OK-VQA) is a challenging VQA task that requires retrieval of external knowledge to answer questions about images.Recent OK-VQA systems use Dense Passage Retrieval (DPR) to retrieve documents from external knowledge bases, such as Wikipedia, but with DPR trained separately from answer generation, introducing a potential limit on the overall system performance.Instead, we propose a joint training scheme which includes differentiable DPR integrated with answer generation so that the system can be trained in an end-to-end fashion.Our experiments show that our scheme outperforms recent OK-VQA systems with strong DPR for retrieval.We also introduce new diagnostic metrics to analyze how retrieval and generation interact.The strong retrieval ability of our model significantly reduces the number of retrieved documents needed in training, yielding significant benefits in answer quality and computation required for training. Weizhe Lin, William J. Byrne |
EMNLP | 2 |
| 2022 | The Devil is in the Details: On the Pitfalls of Vocabulary Selection in Neural Machine TranslationabstractTobias Domhan, Eva Hasler, Ke Tran, Sony Trenous, Bill Byrne, Felix Hieber. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Tobias Domhan, Eva Hasler, Ke Tran, Sony Trenous, William J. Byrne, Felix Hieber |
NAACL-HLT | 5 |
| 2021 | TicketTalk: Toward human-level performance with end-to-end, transaction-based dialog systemsabstractBill Byrne, Karthik Krishnamoorthi, Saravanan Ganesh, Mihir Kale. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. William J. Byrne, Karthik Krishnamoorthi, Saravanan Ganesh, Mihir Sanjay Kale |
ACL/IJCNLP (1) | 1 |
| 2021 | Transferable Dialogue Systems and User SimulatorsabstractBo-Hsiang Tseng, Yinpei Dai, Florian Kreyssig, Bill Byrne. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Bo-Hsiang Tseng, Yinpei Dai, Florian Kreyssig, William J. Byrne |
ACL/IJCNLP (1) | 4 |
| 2021 | Improving the Quality Trade-Off for Neural Machine Translation Multi-Domain AdaptationabstractBuilding neural machine translation systems to perform well on a specific target domain is a well-studied problem.Optimizing system performance for multiple, diverse target domains however remains a challenge.We study this problem in an adaptation setting where the goal is to preserve the existing system quality while incorporating data for domains that were not the focus of the original translation system.We find that we can improve over the performance trade-off offered by Elastic Weight Consolidation with a relatively simple data mixing strategy.At comparable performance on the new domains, catastrophic forgetting is mitigated significantly on strong WMT baselines.Combining both approaches improves the Pareto frontier on this task. Eva Hasler, Tobias Domhan, Jonay Trénous, Ke Tran, William J. Byrne, Felix Hieber |
EMNLP (1) | 5 |
| 2021 | Knowledge-Aware Graph-Enhanced GPT-2 for Dialogue State TrackingabstractDialogue State Tracking is central to multidomain task-oriented dialogue systems, responsible for extracting information from user utterances.We present a novel hybrid architecture that augments GPT-2 with representations derived from Graph Attention Networks in such a way to allow causal, sequential prediction of slot values.The model architecture captures inter-slot relationships and dependencies across domains that otherwise can be lost in sequential prediction.We report improvements in state tracking performance in Mul-tiWOZ 2.0 against a strong GPT-2 baseline and investigate a simplified sparse training scenario in which DST models are trained only on session-level annotations but evaluated at the turn level.We further report detailed analyses to demonstrate the effectiveness of graph models in DST by showing that the proposed graph modules capture inter-slot dependencies and improve the predictions of values that are common to multiple domains. Weizhe Lin, Bo-Hsiang Tseng, William J. Byrne |
EMNLP (1) | 3 |
| 2020 | Reducing Gender Bias in Neural Machine Translation as a Domain Adaptation ProblemabstractTraining data for NLP tasks often exhibits gender bias in that fewer sentences refer to women than to men.In Neural Machine Translation (NMT) gender bias has been shown to reduce translation quality, particularly when the target language has grammatical gender.The recent WinoMT challenge set allows us to measure this effect directly (Stanovsky et al., 2019).Ideally we would reduce system bias by simply debiasing all data prior to training, but achieving this effectively is itself a challenge.Rather than attempt to create a 'balanced' dataset, we use transfer learning on a small set of trusted, gender-balanced examples.This approach gives strong and consistent improvements in gender debiasing with much less computational cost than training from scratch.A known pitfall of transfer learning on new domains is 'catastrophic forgetting', which we address both in adaptation and in inference. During adaptation we show that Elastic Weight Consolidation allows a performance trade-off between general translation quality and bias reduction. During inference we propose a latticerescoring scheme which outperforms all systems evaluated in Stanovsky et al. (2019) onWinoMT with no degradation of general test set BLEU, and we show this scheme can be applied to remove gender bias in the output of 'black box' online commercial MT systems.We demonstrate our approach translating from English into three languages with varied linguistic properties and data availability. Danielle Saunders, William J. Byrne |
ACL | 2 |
| 2020 | Using Context in Neural Machine Translation Training ObjectivesabstractWe present Neural Machine Translation (NMT) training using document-level metrics with batch-level documents.Previous sequence-objective approaches to NMT training focus exclusively on sentence-level metrics like sentence BLEU which do not correspond to the desired evaluation metric, typically document BLEU.Meanwhile research into document-level NMT training focuses on data or model architecture rather than training procedure.We find that each of these lines of research has a clear space in it for the other, and propose merging them with a scheme that allows a document-level evaluation metric to be used in the NMT training objective.We first sample pseudo-documents from sentence samples.We then approximate the expected document BLEU gradient with Monte Carlo sampling for use as a cost function in Minimum Risk Training (MRT).This twolevel sampling procedure gives NMT performance gains over sequence MRT and maximum-likelihood training.We demonstrate that training is more robust for document-level metrics than with sequence metrics.We further demonstrate improvements on NMT with TER and Grammatical Error Correction (GEC) using GLEU, both metrics used at the document level for evaluations. Danielle Saunders, Felix Stahlberg, William J. Byrne |
ACL | 3 |
| 2019 | Domain Adaptive Inference for Neural Machine TranslationabstractWe investigate adaptive ensemble weighting for Neural Machine Translation, addressing the case of improving performance on a new and potentially unknown domain without sacrificing performance on the original domain.We adapt sequentially across two Spanish-English and three English-German tasks, comparing unregularized fine-tuning, L2 and Elastic Weight Consolidation.We then report a novel scheme for adaptive NMT ensemble decoding by extending Bayesian Interpolation with source information, and show strong improvements across test domains without access to the domain label. Danielle Saunders, Felix Stahlberg, Adrià de Gispert, William J. Byrne |
ACL (1) | 4 |
| 2019 | Taskmaster-1: Toward a Realistic and Diverse Dialog DatasetabstractBill Byrne, Karthik Krishnamoorthi, Chinnadhurai Sankar, Arvind Neelakantan, Ben Goodrich, Daniel Duckworth, Semih Yavuz, Amit Dubey, Kyu-Young Kim, Andy Cedilnik. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. William J. Byrne, Karthik Krishnamoorthi, Chinnadhurai Sankar, Arvind Neelakantan, Ben Goodrich, Daniel Duckworth, Semih Yavuz, Amit Dubey, Kyu-Young Kim, Andy Cedilnik |
EMNLP/IJCNLP (1) | 1 |
| 2019 | On NMT Search Errors and Model Errors: Cat Got Your Tongue?abstractFelix Stahlberg, Bill Byrne. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Felix Stahlberg, William J. Byrne |
EMNLP/IJCNLP (1) | 2 |
| 2019 | Semi-Supervised Bootstrapping of Dialogue State Trackers for Task-Oriented ModellingabstractBo-Hsiang Tseng, Marek Rei, Paweł Budzianowski, Richard Turner, Bill Byrne, Anna Korhonen. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Bo-Hsiang Tseng, Marek Rei, Pawel Budzianowski, Richard E. Turner, William J. Byrne, Anna Korhonen |
EMNLP/IJCNLP (1) | 5 |
| 2019 | Coached Conversational Preference Elicitation: A Case Study in Understanding Movie PreferencesabstractConversational recommendation has recently attracted significant attention.As systems must understand users' preferences, training them has called for conversational corpora, typically derived from task-oriented conversations.We observe that such corpora often do not reflect how people naturally describe preferences.We present a new approach to obtaining user preferences in dialogue: Coached Conversational Preference Elicitation.It allows collection of natural yet structured conversational preferences.Studying the dialogues in one domain, we present a brief quantitative analysis of how people describe movie preferences at scale.Demonstrating the methodology, we release the CCPE-M dataset to the community with over 500 movie preference dialogues expressing over 10,000 preferences.1 Filip Radlinski, Krisztian Balog, William J. Byrne, Karthik Krishnamoorthi |
SIGdial | 3 |
| 2017 | Unfolding and Shrinking Neural Machine Translation EnsemblesabstractEnsembling is a well-known technique in neural machine translation (NMT) to improve system performance.Instead of a single neural net, multiple neural nets with the same topology are trained separately, and the decoder generates predictions by averaging over the individual models.Ensembling often improves the quality of the generated translations drastically.However, it is not suitable for production systems because it is cumbersome and slow.This work aims to reduce the runtime to be on par with a single system without compromising the translation quality.First, we show that the ensemble can be unfolded into a single large neural network which imitates the output of the ensemble system.We show that unfolding can already improve the runtime in practice since more work can be done on the GPU.We proceed by describing a set of techniques to shrink the unfolded network by reducing the dimensionality of layers.On Japanese-English we report that the resulting network has the size and decoding speed of a single NMT network but performs on the level of a 3-ensemble system. Felix Stahlberg, William J. Byrne |
EMNLP | 2 |
| 2017 | Break it Down for Me: A Study in Automated Lyric AnnotationabstractComprehending lyrics, as found in songs and poems, can pose a challenge to human and machine readers alike. This motivates the need for systems that can understand the ambiguity and jargon found in such creative texts, and provide commentary to aid readers in reaching the correct interpretation. We introduce the task of automated lyric annotation (ALA). Like text simplification, a goal of ALA is to rephrase the original text in a more easily understandable manner. However, in ALA the system must often include additional information to clarify niche terminology and abstract concepts. To stimulate research on this task, we release a large collection of crowdsourced annotations for song lyrics. We analyze the performance of translation and retrieval models on this task, measuring performance with both automated and human evaluation. We find that each model captures a unique type of information important to the task. Lucas Sterckx, Jason Naradowsky, William J. Byrne, Thomas Demeester, Chris Develder |
EMNLP | 3 |
| 2017 | A Comparison of Neural Models for Word OrderingabstractWe compare several language models for the word-ordering task and propose a new bagto-sequence neural model based on attentionbased sequence-to-sequence models.We evaluate the model on a large German WMT data set where it significantly outperforms existing models.We also describe a novel search strategy for LM-based word ordering and report results on the English Penn Treebank.Our best model setup outperforms prior work both in terms of speed and quality. Eva Hasler, Felix Stahlberg, Marcus Tomalin, Adrià de Gispert, William J. Byrne |
INLG | 5 |
| 2017 | Source sentence simplification for statistical machine translationabstractLong sentences with complex syntax and long-distance dependencies pose difficulties for machine translation systems. Short sentences, on the other hand, are usually easier to translate. We study the potential of addressing this mismatch using text simplification: given a simplified version of the full input sentence, can we use it in addition to the full input to improve translation? We show that the spaces of original and simplified translations can be effectively combined using translation lattices and compare two decoding approaches to process both inputs at different levels of integration. We demonstrate on source-annotated portions of WMT test sets and on top of strong baseline systems combining hierarchical and neural translation for two language pairs that source simplification can help to improve translation quality. Eva Hasler, Adrià de Gispert, Felix Stahlberg, Aurelien Waite, William J. Byrne |
Comput. Speech Lang. | 5 |
| 2016 | Speed-Constrained Tuning for Statistical Machine Translation Using Bayesian OptimizationabstractDaniel Beck, Adrià de Gispert, Gonzalo Iglesias, Aurelien Waite, Bill Byrne. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016. Daniel Beck, Adrià de Gispert, Gonzalo Iglesias, Aurelien Waite, William J. Byrne |
HLT-NAACL | 5 |
| 2015 | Transducer Disambiguation with Sparse Topological FeaturesabstractWe describe a simple and efficient algorithm to disambiguate non-functional weighted finite state transducers (WFSTs), i.e. to generate a new WFST that contains a unique, best-scoring path for each hypothesis in the input labels along with the best output labels.The algorithm uses topological features combined with a tropical sparse tuple vector semiring.We empirically show that our algorithm is more efficient than previous work in a PoStagging disambiguation task.We use our method to rescore very large translation lattices with a bilingual neural network language model, obtaining gains in line with the literature. Gonzalo Iglesias, Adrià de Gispert, William J. Byrne |
EMNLP | 3 |
| 2015 | Fast and Accurate Preordering for SMT using Neural NetworksabstractWe propose the use of neural networks to model source-side preordering for faster and better statistical machine translation.The neural network trains a logistic regression model to predict whether two sibling nodes of the source-side parse tree should be swapped in order to obtain a more monotonic parallel corpus, based on samples extracted from the word-aligned parallel corpus.For multiple language pairs and domains, we show that this yields the best reordering performance against other state-of-the-art techniques, resulting in improved translation quality and very fast decoding. Adrià de Gispert, Gonzalo Iglesias, William J. Byrne |
HLT-NAACL | 3 |
| 2015 | The Geometry of Statistical Machine TranslationabstractMost modern statistical machine translation systems are based on linear statistical models.One extremely effective method for estimating the model parameters is minimum error rate training (MERT), which is an efficient form of line optimisation adapted to the highly nonlinear objective functions used in machine translation.We describe a polynomial-time generalisation of line optimisation that computes the error surface over a plane embedded in parameter space.The description of this algorithm relies on convex geometry, which is the mathematics of polytopes and their faces.Using this geometric representation of MERT we investigate whether the optimisation of linear models is tractable in general.Previous work on finding optimal solutions in MERT (Galley and Quirk, 2011) established a worstcase complexity that was exponential in the number of sentences, in contrast we show that exponential dependence in the worst-case complexity is mainly in the number of features.Although our work is framed with respect to MERT, the convex geometric description is also applicable to other error-based training methods for linear models.We believe our analysis has important ramifications because it suggests that the current trend in building statistical machine translation systems by introducing a very large number of sparse features is inherently not robust. Aurelien Waite, William J. Byrne |
HLT-NAACL | 2 |
| 2014 | Effective Incorporation of Source Syntax into Hierarchical Phrase-based Translation
Tong Xiao 0001, Adrià de Gispert, William J. Byrne |
COLING | 4 |
| 2014 | Word Ordering with Phrase-Based GrammarsabstractWe describe an approach to word ordering using modelling techniques from statistical machine translation. The system incorporates a phrase-based model of string generation that aims to take unordered bags of words and produce fluent, grammatical sentences. We describe the generation grammars and introduce parsing procedures that address the computational complexity of generation under permutation of phrases. Against the best previous results reported on this task, obtained using syntax driven models, we report huge quality improvements, with BLEU score gains of 20+ which we confirm with human fluency judgements. Our system incorporates dependency language models, large n-gram language models, and minimum Bayes risk decoding. © 2014 Association for Computational Linguistics. Adrià de Gispert, Marcus Tomalin, William J. Byrne |
EACL | 3 |
| 2014 | A Graph-Based Approach to String RegenerationabstractThe string regeneration problem is the problem of generating a fluent sentence from a bag of words. We explore the Ngram language model approach to string regeneration. The approach computes the highest probability permutation of the input bag of words under an N-gram language model. We describe a graph-based approach for finding the optimal permutation. The evaluation of the approach on a number of datasets yielded promising results, which were confirmed by conducting a manual evaluation study. Matic Horvat, William J. Byrne |
EACL | 2 |
| 2014 | Source-side Preordering for Translation using Logistic Regression and Depth-first Branch-and-Bound SearchabstractWe present a simple preordering approach for machine translation based on a featurerich logistic regression model to predict whether two children of the same node in the source-side parse tree should be swapped or not.Given the pair-wise children regression scores we conduct an efficient depth-first branch-and-bound search through the space of possible children permutations, avoiding using a cascade of classifiers or limiting the list of possible ordering outcomes.We report experiments in translating English to Japanese and Korean, demonstrating superior performance as (a) the number of crossing links drops by more than 10% absolute with respect to other state-of-the-art preordering approaches, (b) BLEU scores improve on 2.2 points over the baseline with lexicalised reordering model, and (c) decoding can be carried out 80 times faster. Laura Jehl, Adrià de Gispert, Mark Hopkins, William J. Byrne |
EACL | 4 |
| 2014 | Investigating automatic & human filled pause insertion for speech synthesisabstractFilled pauses are pervasive in conversational speech and have been shown to serve several psychological and structural purposes.Despite this, they are seldom modelled overtly by stateof-the-art speech synthesis systems.This paper seeks to motivate the incorporation of filled pauses into speech synthesis systems by exploring their use in conversational speech, and by comparing the performance of several automatic systems inserting filled pauses into fluent text.Two initial experiments are described which seek to determine whether people's predicted insertion points are consistent with actual practice and/or with each other.The experiments also investigate whether there are 'right' and 'wrong' places to insert filled pauses.The results show good consistency between people's predictions of usage and their actual practice, as well as a perceptual preference for the 'right' placement.The third experiment contrasts the performance of several automatic systems that insert filled pauses into fluent sentences.The best performance (determined by F-score) was achieved through the by-word interpolation of probabilities predicted by Recurrent Neural Network and 4gram Language Models.The results offer insights into the use and perception of filled pauses by humans, and how automatic systems can be used to predict their locations. Rasmus Dall, Marcus Tomalin, Mirjam Wester, William J. Byrne, Simon King 0001 |
INTERSPEECH | 4 |
| 2014 | Pushdown Automata in Statistical Machine TranslationabstractThis article describes the use of pushdown automata (PDA) in the context of statistical machine translation and alignment under a synchronous context-free grammar. We use PDAs to compactly represent the space of candidate translations generated by the grammar when applied to an input sentence. General-purpose PDA algorithms for replacement, composition, shortest path, and expansion are presented. We describe HiPDT, a hierarchical phrase-based decoder using the PDA representation and these algorithms. We contrast the complexity of this decoder with a decoder based on a finite state automata representation, showing that PDAs provide a more suitable framework to achieve exact decoding for larger synchronous context-free grammars and smaller language models. We assess this experimentally on a large-scale Chinese-to-English alignment and translation task. In translation, we propose a two-pass decoding strategy involving a weaker language model in the first-pass to address the results of PDA complexity analysis. We study in depth the experimental conditions and tradeoffs in which HiPDT can achieve state-of-the-art performance for large-scale SMT. Cyril Allauzen, William J. Byrne, Adrià de Gispert, Gonzalo Iglesias, Michael Riley 0001 |
Comput. Linguistics | 2 |
| 2013 | Visualising Multiple Data Sources in an Independent Open Learner Model
Susan Bull, Matthew D. Johnson 0002, Mohammad Alotaibi, William J. Byrne, Gabriele Cierniak |
AIED | 4 |
| 2013 | Learning, Learning Analytics, Activity Visualisation and Open Learner Model: Confusing?
Susan Bull, Michael D. Kickmeier-Rust, Ravikiran Vatrapu, Matthew D. Johnson 0002, Klaus Hammermueller, William J. Byrne, Luis Hernandez-Munoz, Fabrizio Giorgini, Gerhilde Meissl-Egghart |
EC-TEL | 6 |
| 2013 | Fast, low-artifact speech synthesis considering global varianceabstractSpeech parameter generation considering global variance (GV generation) is widely acknowledged to dramatically improve the quality of synthetic speech generated by HMM-based systems. However it is slower and has higher latency than the standard speech parameter generation algorithm. In addition it is known to produce artifacts, though existing approaches to prevent artifacts are effective. We present a simple new theoretical analysis of speech parameter generation considering global variance based on Lagrange multipliers. This analysis sheds light on one source of artifacts and suggests a way to reduce their occurrence. It also suggests an approximation to exact GV generation that allows fast, low latency synthesis. In a subjective evaluation our fast approximation shows no degradation in naturalness compared to conventional GV generation. Matt Shannon, William J. Byrne |
ICASSP | 2 |
| 2013 | Personalising speech-to-speech translation: Unsupervised cross-lingual speaker adaptation for HMM-based speech synthesis
John Dines, Lakshmi Babu Saheer, Matthew Gibson, William J. Byrne, Keiichiro Oura, Keiichi Tokuda, Junichi Yamagishi, Simon King 0001, Mirjam Wester, Teemu Hirsimäki, Reima Karhila, Mikko Kurimo |
Comput. Speech Lang. | 5 |
| 2013 | N-gram posterior probability confidence measures for statistical machine translation: an empirical studyabstractWe report an empirical study of n -gram posterior probability confidence measures for statistical machine translation (SMT). We first describe an efficient and practical algorithm for rapidly computing n -gram posterior probabilities from large translation word lattices. These probabilities are shown to be a good predictor of whether or not the n -gram is found in human reference translations, motivating their use as a confidence measure for SMT. Comprehensive n -gram precision and word coverage measurements are presented for a variety of different language pairs, domains and conditions. We analyze the effect on reference precision of using single or multiple references, and compare the precision of posteriors computed from k -best lists to those computed over the full evidence space of the lattice. We also demonstrate improved confidence by combining multiple lattices in a multi-source translation framework. Adrià de Gispert, Graeme W. Blackwood, Gonzalo Iglesias, William J. Byrne |
Mach. Transl. | 4 |
| 2013 | Autoregressive Models for Statistical Parametric Speech SynthesisabstractWe propose using the autoregressive hidden Markov model (HMM) for speech synthesis. The autoregressive HMM uses the same model for parameter estimation and synthesis in a consistent way, in contrast to the standard approach to statistical parametric speech synthesis. It supports easy and efficient parameter estimation using expectation maximization, in contrast to the trajectory HMM. At the same time its similarities to the standard approach allow use of established high quality synthesis algorithms such as speech parameter generation considering global variance. The autoregressive HMM also supports a speech parameter generation algorithm not available for the standard approach or the trajectory HMM and which has particular advantages in the domain of real-time, low latency synthesis. We show how to do efficient parameter estimation and synthesis with the autoregressive HMM and look at some of the similarities and differences between the standard approach, the trajectory HMM and the autoregressive HMM. We compare the three approaches in subjective and objective evaluations. We also systematically investigate which choices of parameters such as autoregressive order and number of states are optimal for the autoregressive HMM. Matt Shannon, Heiga Zen, William J. Byrne |
IEEE Trans. Speech Audio Process. | 3 |
| 2012 | Impacts of machine translation and speech synthesis on speech-to-speech translation
Kei Hashimoto, Junichi Yamagishi, William J. Byrne, Simon King 0001, Keiichi Tokuda |
Speech Commun. | 3 |
| 2011 | Hierarchical Phrase-based Translation Representations
Gonzalo Iglesias, Cyril Allauzen, William J. Byrne, Adrià de Gispert, Michael Riley 0001 |
EMNLP | 3 |
| 2011 | An analysis of machine translation and speech synthesis in speech-to-speech translation systemabstractThis paper provides an analysis of the impacts of machine translation and speech synthesis on speech-to-speech translation systems. The speech-to-speech translation system consists of three components: speech recognition, machine translation and speech synthesis. Many techniques for integration of speech recognition and machine translation have been proposed. However, speech synthesis has not yet been considered. Therefore, in this paper, we focus on machine translation and speech synthesis, and report a subjective evaluation to analyze the impact of each component. The results of these analyses show that the naturalness and intelligibility of synthesized speech are strongly affected by the fluency of the translated sentences. Kei Hashimoto, Junichi Yamagishi, William J. Byrne, Simon King 0001, Keiichi Tokuda |
ICASSP | 3 |
| 2011 | Choosing your Moment: Interruptions in Multimedia Annotation
Chris P. Bowers, William J. Byrne, Benjamin R. Cowan, Chris Creed, Robert J. Hendley, Russell Beale |
INTERACT (2) | 2 |
| 2011 | The Effect of Using Normalized Models in Statistical Speech SynthesisabstractThe standard approach to HMM-based speech synthesis is inconsistent in the enforcement of the deterministic constraints between static and dynamic features. The trajectory HMM and autoregressive HMM have been proposed as normalized models which rectify this inconsistency. This paper investigates the practical effects of using these normalized models, and examines the strengths and weaknesses of the different models as probabilistic models of speech. The most striking difference observed is that the standard approach greatly underestimates predictive variance. We argue that the normalized models have better predictive distributions than the standard approach, but that all the models we consider are still far from satisfactory probabilistic models of speech. We also present evidence that better intra-frame correlation modelling goes some way towards improving existing normalized models. Index terms: HMM-based speech synthesis, acoustic modelling, autoregressive HMM, trajectory HMM, normalization 1. Matt Shannon, Heiga Zen, William J. Byrne |
INTERSPEECH | 3 |
| 2011 | Unsupervised Intralingual and Cross-Lingual Speaker Adaptation for HMM-Based Speech Synthesis Using Two-Pass Decision Tree ConstructionabstractHidden Markov model (HMM)-based speech synthesis systems possess several advantages over concatenative synthesis systems. One such advantage is the relative ease with which HMM-based systems are adapted to speakers not present in the training dataset. Speaker adaptation methods used in the field of HMM-based automatic speech recognition (ASR) are adopted for this task. In the case of unsupervised speaker adaptation, previous work has used a supplementary set of acoustic models to estimate the transcription of the adaptation data. This paper first presents an approach to the unsupervised speaker adaptation task for HMM-based speech synthesis models which avoids the need for such supplementary acoustic models. This is achieved by defining a mapping between HMM-based synthesis models and ASR-style models, via a two-pass decision tree construction process. Second, it is shown that this mapping also enables unsupervised adaptation of HMM-based speech synthesis models without the need to perform linguistic analysis of the estimated transcription of the adaptation data. Third, this paper demonstrates how this technique lends itself to the task of unsupervised cross-lingual adaptation of HMM-based speech synthesis models, and explains the advantages of such an approach. Finally, listener evaluations reveal that the proposed unsupervised adaptation methods deliver performance approaching that of supervised adaptation. Matthew Gibson, William J. Byrne |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2010 | Fluency Constraints for Minimum Bayes-Risk Decoding of Statistical Machine Translation Lattices
Graeme W. Blackwood, Adrià de Gispert, William J. Byrne |
COLING | 3 |
| 2010 | Hierarchical Phrase-Based Translation Grammars Extracted from Alignment Posterior Probabilities
Adrià de Gispert, Juan Pino 0001, William J. Byrne |
EMNLP | 3 |
| 2010 | Unsupervised cross-lingual speaker adaptation for HMM-based speech synthesis using two-pass decision tree constructionabstractThis paper demonstrates how unsupervised cross-lingual adaptation of HMM-based speech synthesis models may be performed without explicit knowledge of the adaptation data language. A two-pass decision tree construction technique is deployed for this purpose. Using parallel translated datasets, cross-lingual and intralingual adaptation are compared in a controlled manner. Listener evaluations reveal that the proposed method delivers performance approaching that of unsupervised intralingual adaptation. Matthew Gibson, Teemu Hirsimäki, Reima Karhila, Mikko Kurimo, William J. Byrne |
ICASSP | 5 |
| 2010 | Autoregressive clustering for HMM speech synthesisabstractThe autoregressive HMM has been shown to provide efficient parameter estimation and high-quality synthesis, but in previous experiments decision trees derived from a non-autoregressive system were used. In this paper we investigate the use of autoregressive clustering for autoregressive HMM-based speech synthesis. We describe decision tree clustering for the autoregressive HMM and highlight differences to the standard clustering procedure. Subjective listening evaluation results suggest that autoregressive clustering improves the naturalness of the resulting speech. We find that the standard minimum description length (MDL) criterion for selecting model complexity is inappropriate for the autoregressive HMM. Investigating the effect of model complexity on naturalness, we find that a large degree of overfitting is tolerated without a substantial decrease in naturalness. Index terms: HMM-based speech synthesis, decision tree clustering, autoregressive HMM 1. Matt Shannon, William J. Byrne |
INTERSPEECH | 2 |
| 2010 | Hierarchical Phrase-Based Translation with Weighted Finite-State Transducers and Shallow-n GrammarsabstractIn this article we describe HiFST, a lattice-based decoder for hierarchical phrase-based translation and alignment. The decoder is implemented with standard Weighted Finite-State Transducer (WFST) operations as an alternative to the well-known cube pruning procedure. We find that the use of WFSTs rather than k-best lists requires less pruning in translation search, resulting in fewer search errors, better parameter optimization, and improved translation performance. The direct generation of translation lattices in the target language can improve subsequent rescoring procedures, yielding further gains when applying long-span language models and Minimum Bayes Risk decoding. We also provide insights as to how to control the size of the search space defined by hierarchical rules. We show that shallow-n grammars, low-level rule catenation, and other search constraints can help to match the power of the translation system to specific language pairs. Adrià de Gispert, Gonzalo Iglesias, Graeme W. Blackwood, Eduardo Rodríguez Banga, William J. Byrne |
Comput. Linguistics | 5 |
| 2009 | Rule Filtering by Pattern for Efficient Hierarchical Translation
Gonzalo Iglesias, Adrià de Gispert, Eduardo Rodríguez Banga, William J. Byrne |
EACL | 4 |
| 2009 | Autoregressive HMMs for speech synthesisabstractWe propose the autoregressive HMM for speech synthesis. We show that the autoregressive HMM supports efficient EM parameter estimation and that we can use established effective synthesis techniques such as synthesis considering global variance with minimal modification. The autoregressive HMM uses the same model for parameter estimation and synthesis in a consistent way, in contrast to the standard HMM synthesis framework, and supports easy and efficient parameter estimation, in contrast to the trajectory HMM. We find that the autoregressive HMM gives performance comparable to the standard HMM synthesis framework on a Blizzard Challenge-style naturalness evaluation. Matt Shannon, William J. Byrne |
INTERSPEECH | 2 |
| 2009 | Context-Dependent Alignment Models for Statistical Machine Translation
Jamie Brunning, Adrià de Gispert, William J. Byrne |
HLT-NAACL | 3 |
| 2009 | Hierarchical Phrase-Based Translation with Weighted Finite State Transducers
Gonzalo Iglesias, Adrià de Gispert, Eduardo Rodríguez Banga, William J. Byrne |
HLT-NAACL | 4 |
| 2008 | HMM Word and Phrase Alignment for Statistical Machine TranslationabstractEstimation and alignment procedures for word and phrase alignment hidden Markov models (HMMs) are developed for the alignment of parallel text. The development of these models is motivated by an analysis of the desirable features of IBM Model 4, one of the original and most effective models for word alignment. These models are formulated to capture the desirable aspects of Model 4 in an HMM alignment formalism. Alignment behavior is analyzed and compared to human-generated reference alignments, and the ability of these models to capture different types of alignment phenomena is evaluated. In analyzing alignment performance, Chinese-English word alignments are shown to be comparable to those of IBM Model 4 even when models are trained over large parallel texts. In translation performance, phrase-based statistical machine translation systems based on these HMM alignments can equal and exceed systems based on Model 4 alignments, and this is shown in Arabic-English and Chinese-English translation. These alignment models can also be used to generate posterior statistics over collections of parallel text, and this is used to refine and extend phrase translation tables with a resulting improvement in translation quality. Yonggang Deng, William J. Byrne |
IEEE Trans. Speech Audio Process. | 2 |
| 2007 | Discriminative language model adaptation for Mandarin broadcast speech transcription and translationabstractThis paper investigates unsupervised test-time adaptation of language models (LM) using discriminative methods for a Mandarin broadcast speech transcription and translation task. A standard approach to adapt interpolated language models to is to optimize the component weights byminimizing the perplexity on supervision data. This is a widely made approximation for language modeling in automatic speech recognition (ASR) systems. For speech translation tasks, it is unclear whether a strong correlation still exists between perplexity and various forms of error cost functions in recognition and translation stages. The proposed minimum Bayes risk (MBR) based approach provides a flexible framework for unsupervised LM adaptation. It generalizes to a variety of forms of recognition and translation error metrics. LM adaptation is performed at the audio document level using either the character error rate (CER), or translation edit rate (TER) as the cost function. An efficient parameter estimation scheme using the extended Baum-Welch (EBW) algorithm is proposed. Experimental results on a state-of-the-art speech recognition and translation system are presented. The MBR adapted language models gave the best recognition and translation performance and reduced the TER score by up to 0.54% absolute. Xunying Liu, William J. Byrne, Mark J. F. Gales, Adrià de Gispert, Marcus Tomalin, Philip C. Woodland, Kai Yu 0004 |
ASRU | 2 |
| 2007 | Consensus Network Decoding for Statistical Machine Translation System CombinationabstractThis paper presents a simple and robust consensus decoding approach for combining multiple machine translation (MT) system outputs. A consensus network is constructed from an N-best list by aligning the hypotheses against an alignment reference, where the alignment is based on minimising the translation edit rate (TER). The minimum Bayes risk (MBR) decoding technique is investigated for the selection of an appropriate alignment reference. Several alternative decoding strategies proposed to retain coherent phrases in the original translations. Experimental results are presented primarily based on three-way combination of Chinese-English translation outputs, and also presents results for six-way system combination. It is shown that worthwhile improvements in translation performance can be obtained using the methods discussed. Khe Chai Sim, William J. Byrne, Mark J. F. Gales, Hichem Sahbi, Philip C. Woodland |
ICASSP (4) | 2 |
| 2007 | Ginisupport vector machines for segmental minimum Bayes risk decoding of continuous speech
Veera Venkataramani, Shantanu Chakrabartty, William J. Byrne |
Comput. Speech Lang. | 3 |
| 2007 | Segmentation and alignment of parallel text for statistical machine translationabstractWe address the problem of extracting bilingual chunk pairs from parallel text to create training sets for statistical machine translation. We formulate the problem in terms of a stochastic generative process over text translation pairs, and derive two different alignment procedures based on the underlying alignment model. The first procedure is a now-standard dynamic programming alignment model which we use to generate an initial coarse alignment of the parallel text. The second procedure is a divisive clustering parallel text alignment procedure which we use to refine the first-pass alignments. This latter procedure is novel in that it permits the segmentation of the parallel text into sub-sentence units which are allowed to be reordered to improve the chunk alignment. The quality of chunk pairs are measured by the performance of machine translation systems trained from them. We show practical benefits of divisive clustering as well as how system performance can be improved by exploiting portions of the parallel text that otherwise would have to be discarded. We also show that chunk alignment as a first step in word alignment can significantly reduce word alignment error rate. Yonggang Deng, Shankar Kumar, William J. Byrne |
Nat. Lang. Eng. | 3 |
| 2006 | Statistical Phrase-Based Speech TranslationabstractA generative statistical model of speech-to-text translation is developed as an extension of existing models of phrase-based text translation. Speech is translated by mapping ASR word lattices to lattices of phrase sequences which are then translated using operations developed for text translation. Performance is reported on Chinese to English translation of Mandarin Broadcast News Lambert Mathias, William J. Byrne |
ICASSP (1) | 2 |
| 2006 | MTTK: An Alignment Toolkit for Statistical Machine Translation
Yonggang Deng, William J. Byrne |
HLT-NAACL | 2 |
| 2006 | A Dialectal Chinese Speech Recognition Framework
Thomas Zheng, William J. Byrne, Daniel Jurafsky |
J. Comput. Sci. Technol. | 3 |
| 2006 | A weighted finite state transducer translation template model for statistical machine translationabstractWe present a Weighted Finite State Transducer Translation Template Model for statistical machine translation. This is a source-channel model of translation inspired by the Alignment Template translation model. The model attempts to overcome the deficiencies of word-to-word translation models by considering phrases rather than words as units of translation. The approach we describe allows us to implement each constituent distribution of the model as a weighted finite state transducer or acceptor. We show that bitext word alignment and translation under the model can be performed with standard finite state machine operations involving these transducers. One of the benefits of using this framework is that it avoids the need to develop specialized search procedures, even for the generation of lattices or N-Best lists of bitext word alignments and translation hypotheses. We report and analyze bitext word alignment and translation performance on the Hansards French-English task and the FBIS Chinese-English task under the Alignment Error Rate, BLEU, NIST and Word Error-Rate metrics. These experiments identify the contribution of each of the model components to different aspects of alignment and translation performance. We finally discuss translation performance with large bitext training sets on the NIST 2004 Chinese-English and Arabic-English MT tasks. Shankar Kumar, Yonggang Deng, William J. Byrne |
Nat. Lang. Eng. | 3 |
| 2006 | Lattice segmentation and minimum Bayes risk discriminative training for large vocabulary continuous speech recognition
Vlasios Doumpiotis, William J. Byrne |
Speech Commun. | 2 |
| 2006 | Corrections to "Segmental minimum Bayes-risk decoding for automatic speech recognition"abstractThe purpose of this paper is to correct and expand upon the experimental results presented in our recently published paper [1]. In [1, Sec. III-B], we present a risk-based lattice cutting (RLC) procedure to segment ASR word lattices into sequences of smaller sublattices. The purpose of this procedure is to restructure the original lattice to improve the efficiency of minimum Bayes-risk (MBR) and other lattice rescoring procedures. Given that the segmented lattices are to be rescored, it is crucial that no paths from the original lattice be lost in the segmentation process. In the experiments reported in our original publication, some of the original paths were inadvertently discarded from the segmented lattices. This affected the performance of the MBR results presented. In this paper, we briefly review the segmentation algorithm and explain the flaw in our previous experiments. We find consistent minor improvements in word error rate (WER) under the corrected procedure. More importantly, we report experiments confirming that the lattice segmentation procedure does indeed preserve all the paths in the original lattice. Vaibhava Goel, Shankar Kumar, William J. Byrne |
IEEE Trans. Speech Audio Process. | 3 |
| 2005 | Acoustic Training from Heterogeneous Data Sources: Experiments in Mandarin Conversational Telephone Speech TranscriptionabstractIn this paper we investigate the use of heterogeneous data sources for acoustic training. We describe an acoustic normalization procedure for enlarging an ASR acoustic training set with out-of-domain acoustic data. A larger in-domain training set is created by effectively transforming the out-of-domain data before incorporation in training. Baseline experimental results in Mandarin conversational telephone speech transcription show that a simple attempt to add out-of-domain data degrades performance. Preliminary experiments assess the effectiveness of the proposed cross-corpus acoustic normalization. Furthermore, we investigate the behavior of speaker adaptive training in conjunction with the cross-corpus normalization procedure. Stavros Tsakalidis, William J. Byrne |
ICASSP (1) | 2 |
| 2005 | Lattice Segmentation and Support Vector Machines for Large Vocabulary Continuous Speech RecognitionabstractLattice segmentation procedures are used to spot possible recognition errors in first-pass recognition hypotheses produced by a large vocabulary continuous speech recognition system. This approach is analyzed in terms of its ability to reliably identify, and provide good alternatives for, incorrectly hypothesized words. A procedure is described to train and apply support vector machines to strengthen the first pass system where it was found to be weak, resulting in small but statistically significant recognition improvements on a large test set of conversational speech. Veera Venkataramani, William J. Byrne |
ICASSP (1) | 2 |
| 2005 | Automatic transcription of Czech, Russian, and Slovak spontaneous speech in the MALACH projectabstractThis paper describes the 3.5-years effort put into building LVCSR systems for recognition of spontaneous speech of Czech, Russian, and Slovak witnesses of the Holocaust in the MALACH project. For processing of colloquial, highly emotional and heavily accented speech of elderly people containing many non-speech events we have developed techniques that very effectively handle both non-speech events and colloquial and accented variants of uttered words. Manual transcripts as one of the main sources for language modeling were automatically „normalized ” using standardized lexicon, which brought about 2 to 3 % reduction of the word error rate (WER). The subsequent interpolation of such LMs with models built from an additional collection (consisting of topically selected sentences from general text corpora) resulted into an additional improvement of performance of up to 3 %. 1. Josef Psutka, Pavel Ircing, Josef V. Psutka, Jan Hajic 0001, William J. Byrne, Jirí Mírovský |
INTERSPEECH | 5 |
| 2005 | Convergence Theorems for Generalized Alternating Minimization ProceduresabstractThe EM algorithm is widely used to develop iterative parameter estimation procedures for statistical models. In cases where these procedures strictly follow the EM formulation, the convergence properties of the estimation procedures are well understood. In some instances there are practical reasons to develop procedures that do not strictly fall within the EM framework. We study EM variants in which the E-step is not performed exactly, either to obtain improved rates of convergence, or due to approximations needed to compute statistics under a model family over which E-steps cannot be realized. Since these variants are not EM procedures, the standard (G)EM convergence results do not apply to them. We present an information geometric framework for describing such algorithms and analyzing their convergence properties. We apply this framework to analyze the convergence properties of incremental EM and variational EM. For incremental EM, we discuss conditions under these algorithms converge in likelihood. For variational EM, we show how the E-step approximation prevents convergence to local maxima in likelihood. Asela Gunawardana, William J. Byrne |
J. Mach. Learn. Res. | 2 |
| 2005 | Editorial
Eric Fosler-Lussier, William J. Byrne, Daniel Jurafsky |
Speech Commun. | 2 |
| 2005 | Discriminative linear transforms for feature normalization and speaker adaptation in HMM estimationabstractLinear transforms have been used extensively for training and adaptation of HMM-based ASR systems. Recently procedures have been developed for the estimation of linear transforms under the Maximum Mutual Information (MMI) criterion. In this paper we introduce discriminative training procedures that employ linear transforms for feature normalization and for speaker adaptive training. We integrate these discriminative linear transforms into MMI estimation of HMM parameters for improvement of large vocabulary conversational speech recognition systems. Stavros Tsakalidis, Vlasios Doumpiotis, William J. Byrne |
IEEE Trans. Speech Audio Process. | 3 |
| 2004 | The development of ASR for Slavic languages in the MALACH projectabstractThe development of acoustic training material for Slavic languages within the MALACH project is described. Initial experience with the variety of speakers and the difficulties encountered in transcribing Czech, Slovak, and Russian language oral history are described along with automatic speech recognition results intended to investigate the effectiveness of different transcription conventions that address language specific phenomena within the task domain. Josef Psutka, Jan Hajic 0001, William J. Byrne |
ICASSP (3) | 3 |
| 2004 | Pinched lattice minimum Bayes risk discriminative training for large vocabulary continuous speech recognitionabstractIterative estimation procedures that minimize empirical risk based on general loss functions such as the Levenshtein distance have been derived as extensions of the Extended Baum Welch algorithm. While reducing expected loss on training data is a desirable training criterion, these algorithms can be difficult to apply. They are unlike MMI estimation in that they require an explicit listing of the hypotheses to be considered and in complex problems such lists tend to be prohibitively large. To overcome this difficulty, modeling techniques originally developed to improve search efficiency in Minimum Bayes Risk decoding can be used to transform these estimation algorithms so that exact update, risk minimization procedures can be used for complex recognition problems. Experimental results in two large vocabulary speech recognition tasks show improvements over conventionally trained MMIE models. Vlasios Doumpiotis, William J. Byrne |
INTERSPEECH | 2 |
| 2004 | Task-specific minimum Bayes-risk decoding using learned edit distanceabstractThis paper extends the minimum Bayes-risk framework to incorporate a loss function specific to the task and the ASR system. The errors are modeled as a noisy channel and the parameters are learned from the data. The resulting loss function is used in the risk criterion for decoding. Experiments on a large vocabulary conversational speech recognition system demonstrate significant gains of about 1% absolute over MAP hypothesis and about 0.6% absolute over untrained loss function. The approach is general enough to be applicable to other sequence recognition problems such as in Optical Character Recognition (OCR) and in analysis of biological sequences. Izhak Shafran, William J. Byrne |
INTERSPEECH | 2 |
| 2004 | Issues in Annotation of the Czech Spontaneous Speech Corpus in the MALACH project
Josef Psutka, Pavel Ircing, Jan Hajic 0001, Vlasta Radová, Josef V. Psutka, William J. Byrne, Samuel Gustman |
LREC | 6 |
| 2004 | Minimum Bayes-Risk Decoding for Statistical Machine Translation
Shankar Kumar, William J. Byrne |
HLT-NAACL | 2 |
| 2004 | Automatic recognition of spontaneous speech for access to multilingual oral history archivesabstractMuch is known about the design of automated systems to search broadcast news, but it has only recently become possible to apply similar techniques to large collections of spontaneous speech. This paper presents initial results from experiments with speech recognition, topic segmentation, topic categorization, and named entity detection using a large collection of recorded oral histories. The work leverages a massive manual annotation effort on 10 000 h of spontaneous speech to evaluate the degree to which automatic speech recognition (ASR)-based segmentation and categorization techniques can be adapted to approximate decisions made by human annotators. ASR word error rates near 40% were achieved for both English and Czech for heavily accented, emotional and elderly spontaneous speech based on 65-84 h of transcribed speech. Topical segmentation based on shifts in the recognized English vocabulary resulted in 80% agreement with manually annotated boundary positions at a 0.35 false alarm rate. Categorization was considerably more challenging, with a nearest-neighbor technique yielding F=0.3. This is less than half the value obtained by the same technique on a standard newswire categorization benchmark, but replication on human-transcribed interviews showed that ASR errors explain little of that difference. The paper concludes with a description of how these capabilities could be used together to search large collections of recorded oral histories. William J. Byrne, David S. Doermann, Martin Franz, Samuel Gustman, Jan Hajic 0001, Douglas W. Oard, Michael Picheny, Josef Psutka, Bhuvana Ramabhadran, Dagobert Soergel, Todd Ward, Wei-Jing Zhu |
IEEE Trans. Speech Audio Process. | 1 |
| 2004 | Segmental minimum Bayes-risk decoding for automatic speech recognitionabstractMinimum Bayes-risk (MBR) speech recognizers have been shown to yield improvements over the conventional maximum a-posteriori probability (MAP) decoders through N-best list rescoring and A/sup */ search over word lattices. We present a segmental minimum Bayes-risk decoding (SMBR) framework that simplifies the implementation of MBR recognizers through the segmentation of the N-best lists or lattices over which the recognition is to be performed. This paper presents lattice cutting procedures that underly SMBR decoding. Two of these procedures are based on a risk minimization criterion while a third one is guided by word-level confidence scores. In conjunction with SMBR decoding, these lattice segmentation procedures give consistent improvements in recognition word error rate (WER) on the Switchboard corpus. We also discuss an application of risk-based lattice cutting to multiple-system SMBR decoding and show that it is related to other system combination techniques such as ROVER. This strategy combines lattices produced from multiple ASR systems and is found to give WER improvements in a Switchboard evaluation system. Vaibhava Goel, Shankar Kumar, William J. Byrne |
IEEE Trans. Speech Audio Process. | 3 |
| 2003 | Discriminative training for segmental minimum Bayes risk decodingabstractA modeling approach is presented that incorporates discriminative training procedures within segmental minimum Bayes-risk decoding (SMBR). SMBR is used to segment lattices produced by a general automatic speech recognition (ASR) system into sequences of separate decision problems involving small sets of confusable words. Acoustic models specialized to discriminate between the competing words in these classes are then applied in subsequent SMBR rescoring passes. Refinement of the search space that allows the use of specialized discriminative models is shown to be an improvement over rescoring with conventionally trained discriminative models. Vlasios Doumpiotis, Stavros Tsakalidis, William J. Byrne |
ICASSP (1) | 3 |
| 2003 | Lattice segmentation and minimum Bayes risk discriminative trainingabstractLattice segmentation techniques developed for Minimum Bayes Risk decoding in large vocabulary speech recognition tasks are used to compute the statistics needed for discriminative training algorithms that estimate HMM parameters so as to reduce the overall risk over the training data. New estimation procedures are developed and evaluated for both small and large vocabulary recognition tasks, and additive performance improvements are shown relative to maximum mutual information estimation. These relative gains are explained through a detailed analysis of individual word recognition errors. Key words: Discriminative training, maximum mutual information (MMI) estimation, acoustic modeling, minimum Bayes risk decoding, risk minimization, large vocabulary speech recognition, lattice segmentation. 1 Vlasios Doumpiotis, Stavros Tsakalidis, William J. Byrne |
INTERSPEECH | 3 |
| 2003 | Large vocabulary ASR for spontaneous czech in the MALACH projectabstractThis paper describes LVCSR research into the automatic transcription of spontaneous Czech speech in the MALACH (Multilingual Access to Large Spoken Archives) project. This project attempts to provide improved access to the large multilingual spoken archives collected by the Survivors of the Shoah Visual History Foundation (VHF) (www.vhf.org) by advancing the state of the art in automated speech recognition. We describe a baseline ASR system and discuss the problems in language modeling that arise from the nature of Czech as a highly inflectional language that also exhibits diglossia between its written and spontaneous forms. The difficulties of this task are compounded by heavily accented, emotional and disfluent speech along with frequent switching between languages. To overcome the limited amount of relevant language model data we use statistical techniques for selecting an appropriate training corpus from a large unstructured text collection resulting in significant reductions in word error rate. 1. Josef Psutka, Pavel Ircing, Josef V. Psutka, Vlasta Radová, William J. Byrne, Jan Hajic 0001, Jirí Mírovský, Samuel Gustman |
INTERSPEECH | 5 |
| 2003 | A Generative Probabilistic OCR Model for NLP Applications
Okan Kolak, William J. Byrne, Philip Resnik |
HLT-NAACL | 2 |
| 2003 | A Weighted Finite State Transducer Implementation of the Alignment Template Model for Statistical Machine Translation
Shankar Kumar, William J. Byrne |
HLT-NAACL | 2 |
| 2003 | Desparately Seeking Cebuano
Douglas W. Oard, David S. Doermann, Bonnie J. Dorr, Daqing He, Philip Resnik, Amy Weinberg, William J. Byrne, Sanjeev Khudanpur, David Yarowsky, Anton Leuski, Philipp Koehn, Kevin Knight |
HLT-NAACL | 7 |
| 2002 | Minimum Bayes-Risk Word Alignments of Bilingual TextsabstractWe present Minimum Bayes-Risk word alignment for machine translation. This statistical, model-based approach attempts to minimize the expected risk of alignment errors under loss functions that measure alignment quality. We describe various loss functions, including some that incorporate linguistic analysis as can be obtained from parse trees, and show that these approaches can improve alignments of the English-French Hansards. Shankar Kumar, William J. Byrne |
EMNLP | 2 |
| 2002 | Structurally discriminative graphical models for automatic speech recognition - results from the 2001 Johns Hopkins Summer WorkshopabstractIn recent years there has been growing interest in discriminative parameter training techniques, resulting from notable improvements in speech recognition performance on tasks ranging in size from digit recognition to Switchboard. Typified by Maximum Mutual Information training, these methods assume a fixed statistical modeling structure, and then optimize only the associated numerical parameters (such as means, variances, and transition matrices). In this paper, we explore the significantly different methodology of discriminative structure learning. Here, the fundamental dependency relationships between random variables in a probabilistic model are learned in a discriminative fashion, and are learned separately from the numerical parameters. Tn order to apply the principles of structural discriminability, we adopt the framework of graphical models, which allows an arbitrary set of variables with arbitrary conditional independence relationships to be modeled at each time frame. We present results using a new graphical modeling toolkit (described in a companion paper) from the recent 2001 Johns Hopkins Summer Workshop. These results indicate that significant gains result from discriminative structural analysis of both conventional MFCC and novel AM-FM features on the Aurora continuous digits task. Geoffrey Zweig, Jeff A. Bilmes, Thomas Richardson 0001, Karim Filali, Karen Livescu, Kirk Jackson, Yigal Brandman, Eric D. Sandness, Eva Holtz, Jerry Torres, William J. Byrne |
ICASSP | 12 |
| 2002 | Risk based lattice cutting for segmental minimum Bayes-risk decodingabstractMinimum Bayes-Risk (MBR) speech recognizers have been shown to give improvements over the conventional maximum a-posteriori probability (MAP) decoders through N-best list rescoring and search over word lattices. Segmental MBR (SMBR) decoders simplify the implementation of MBR recognizers by segmenting the N-best lists or lattices over which the recognition is performed. We present a lattice cutting procedure that attempts to minimize the total Bayes-Risk of all word strings in the segmented lattice. We provide experimental results on the Switchboard conversational speech corpus showing that this segmentation procedure, in conjunction with SMBR decoding, gives modest but significant improvements over MAP decoders as well as MBR decoders on unsegmented lattices. Shankar Kumar, William J. Byrne |
INTERSPEECH | 2 |
| 2002 | Discriminative linear transforms for feature normalization and speaker adaptation in HMM estimationabstractLinear transforms have been used extensively for training and adaptation of HMM-based ASR systems. Recently procedures have been developed for the estimation of linear transforms under the Maximum Mutual Information (MMI) criterion. In this paper we introduce discriminative training procedures that employ linear transforms for feature normalization and for speaker adaptive training. We integrate these discriminative linear transforms into MMI estimation of HMM parameters for improvement of large vocabulary conversational speech recognition systems. Stavros Tsakalidis, Vlasios Doumpiotis, William J. Byrne |
INTERSPEECH | 3 |
| 2002 | Reducing pronunciation lexicon confusion and using more data without phonetic transcription for pronunciation modelingabstractThe multiple-pronunciation lexicon (MPL) is very important to model the pronunciation variations for spontaneous speech recognition. But the introduction of MPL brings out two problems. First, the MPL will increase the among-lexicon confusion and degrade the recognizer's performance. Second, the MPL needs more data with phonetic transcription so as to cover as many surface forms as possible. Accordingly, two solutions are proposed, they are the context-dependent weighting method and the iterative forced-alignment based transcription method. The use of them can compensate what the MPL causes and improve the overall performance. Experiments across a naturally spontaneous speech database show that the proposed methods are effective and better than other methods. Thomas Fang Zheng, Zhanjiang Song, Pascale Fung, William J. Byrne |
INTERSPEECH | 4 |
| 2002 | Mandarin Pronunciation Modeling Based on CASS Corpus
Thomas Fang Zheng, Zhanjiang Song, Pascale Fung, William J. Byrne |
J. Comput. Sci. Technol. | 4 |
| 2001 | Automatic generation of pronunciation lexicons for Mandarin spontaneous speechabstractPronunciation modeling for large vocabulary speech recognition attempts to improve recognition accuracy by identifying and modeling pronunciations that are not in the ASR systems pronunciation lexicon. Pronunciation variability in spontaneous Mandarin is studied using the newly created CASS corpus of phonetically annotated spontaneous speech. Pronunciation modeling techniques developed for English,are applied to this corpus to train pronunciation models which are then used for Mandarin broadcast news transcription. William J. Byrne, Veera Venkataramani, Terri Kamm, Thomas Fang Zheng, Zhanjiang Song, Pascale Fung, Yi Liu 0022, Umar Ruhi |
ICASSP | 1 |
| 2001 | Confidence based lattice segmentation and minimum Bayes-risk decodingabstractMinimum Bayes Risk (MBR) speech recognizers have been shown to yield improvements over the conventional maximum a-posteriori probability (MAP) decoders in the context of N-best list rescoring and search over recognition lattices. Segmental MBR (SMBR) procedures have been developed to simplify implementation of MBR recognizers, by segmenting the N-best list or lattice, to reduce the size of the search space over which MBR recognition is carried out. In this paper we describe lattice cutting as a method to segment recognition word lattices into regions of low confidence and high confidence. We present two SMBR decoding procedures that can be applied on low confidence segment sets. Results obtained on the Switchboard conversational telephone speech corpus show modest but significant improvements relative to MAP decoders. Vaibhava Goel, Shankar Kumar, William J. Byrne |
INTERSPEECH | 3 |
| 2001 | Discriminative speaker adaptation with conditional maximum likelihood linear regressionabstractWe present a simplified derivation of the extended Baum-Welch procedure, which shows that it can be used for Maximum Mutual Information (MMI) of a large class of continuous emission density hidden Markov models (HMMs). We use the extended Baum-Welch procedure for discriminative estimation of MLLR-type speaker adaptation transformations. The resulting adaptation procedure, termed Conditional Maximum Likelihood Linear Regression (CMLLR), is used successfully for supervised and unsupervised adaptation tasks on the Switchboard corpus, yielding an improvement over MLLR. The interaction of unsupervised CMLLR with segmental minimum Bayes risk lattice voting procedures is also explored, showing that the two procedures are complimentary. 1. Asela Gunawardana, William J. Byrne |
INTERSPEECH | 2 |
| 2001 | On large vocabulary continuous speech recognition of highly inflectional language - czechabstractThe thesis concerns the development of a large vocabulary continuous speech recognition (LVCSR) system for highly inflectional languages, with special emphasis on the language modeling. An idea and usage of the automatic speech recognition is introduced and the basic principles of the statistical approach to the speech recognition and the decomposition of the system into basic components are explained. An overview of the existing statistical language modeling techniques is given and methods of inferring reliable probability estimates from sparse data and measures of the language model quality are described. There are offered a theoretical background to the finite-state machinery and the application of the finite-state machine framework to LVCSR. The goals of the thesis were to build a LVCSR system for the Czech language using standard techniques that were used for English and to analyze the system performance and propose and implement techniques that would improve the recognition accuracy. The development of the baseline system is described. The Czech language properties, especially from the automatic speech recognition point of view, were analyzed. The outcomes of this theoretical analysis are exploited and language models that take into account the specific features of the Czech language are presented. There is given a description of the class-based language models that strengthen the language model robustness and therefore reduce the perplexity and consequently improve the recognition accuracy. And finally a model that uses subword parts (morphemes) as the basic language modeling units is introduced. Such model offers a better coverage of an unknown text in comparison with standard word-based models given the same vocabulary size. Pavel Ircing, Pavel Krbec, Jan Hajic 0001, Josef Psutka, Sanjeev Khudanpur, Frederick Jelinek, William J. Byrne |
INTERSPEECH | 7 |
| 2001 | Modeling pronunciation variation using context-dependent weighting and b/s refined acoustic modelingabstractThe pronunciation variability is an important issue that must be faced with when developing practical automatic spontaneous speech recognition systems. By studying the initial/final (IF) characteristics of Chinese language and developing the Bayesian equation, we propose the concepts of generalized initial/final (GIF) and generalized syllable (GS), the GIF modeling method and the IF-GIF modeling method, as well as the context-dependent pronunciation weighting method. By using these approaches, the IF-GIF modeling reduces the Chinese syllable error rate (SER) by 6.3 % and 4.2 % compared with the GIF modeling and IF modeling respectively when the language modeling, such as syllable or word N-gram, is not used. 1. Thomas Fang Zheng, Zhanjiang Song, Pascale Fung, William J. Byrne |
INTERSPEECH | 4 |
| 2001 | Discounted likelihood linear regression for rapid speaker adaptation
Asela Gunawardana, William J. Byrne |
Comput. Speech Lang. | 2 |
| 2000 | Towards language independent acoustic modelingabstractWe describe procedures and experimental results using speech from diverse source languages to build an ASR system for a single target language. This work is intended to improve ASR in languages for which large amounts of training data are not available. We have developed both knowledge-based and automatic methods to map phonetic units from the source languages to the target language. We employed HMM adaptation techniques and discriminative model combination to combine acoustic models from the individual source languages for recognition of speech in the target language. Experiments are described in which Czech Broadcast News is transcribed using acoustic models trained from small amounts of Czech read speech augmented by English, Spanish, Russian, and Mandarin acoustic models. William J. Byrne, Peter Beyerlein, Juan M. Huerta, Sanjeev Khudanpur, B. Marthi, John Morgan, Nino Peterek, Joseph Picone, Dimitra Vergyri |
ICASSP | 1 |
| 2000 | Robust estimation for rapid speaker adaptation using discounted likelihood techniquesabstractThe discounted likelihood procedure, which is a robust extension of the usual EM procedure, is presented, and two approximations which lead to two different variants of the usual maximum likelihood linear regression adaptation scheme are introduced. These schemes are shown to robustly estimate speaker adaptation transforms with very little data. The evaluation is carried out on the Switchboard corpus. Asela Gunawardana, William J. Byrne |
ICASSP | 2 |
| 2000 | On the incremental addition of regression classes for speaker adaptationabstractWe previously proposed the all-pass transform (APT) as the basis of a speaker adaptation scheme intended for use with a large vocabulary speech recognition system. It was shown that APT-based adaptation reduces to a linear transformation of cepstral means, much like the better known maximum likelihood linear regression (MLLR). Due to this linearity, APT-based adaptation can be used in conjunction with speaker-adapted training (SAT), an algorithm for performing maximum likelihood estimation of the parameters of a hidden Markov model when speaker adaptation is to be employed during both training and test. In other work, we proposed a refinement of SAT dubbed single-pass adapted training (SPAT) specifically-tailored for use with the APT. Here we introduce an incremental training procedure intended for use with the APT and multiple regression classes. In a set of speech recognition experiments conducted on the Switchboard Corpus, we obtained a word error rate of 37.9% using APT adaptation, a significant improvement over the 39.5% word error rate achieved with MLLR. John W. McDonough, Veera Venkataramani, William J. Byrne |
ICASSP | 3 |
| 2000 | Segmental minimum Bayes-risk ASR voting strategiesabstractROVER [1] and its successor voting procedures have been shown to be quite effective in reducing the recognition word error rate (WER). The success of these methods has been attributed to their minimum Bayes-risk (MBR) nature: they produce the hypothesis with the least expected word error. In this paper we develop a general procedure within the MBR framework, called segmental MBR recognition, that encompasses current voting techniques and allows further extensions that yield lower expected WER. It also allows incorporation of loss functions other than the WER. We present a derivation of voting procedure of N-best ROVER as an instance of segmental MBR recognition. We then present an extension, called e-ROVER, that alleviates some of the restrictions of N-best ROVER by better approximating the WER. e-ROVER is compared with N-best ROVER on multi-lingual acoustic modeling task and is shown to yield modest yet significant and easily obtained improvements. Vaibhava Goel, Shankar Kumar, William J. Byrne |
INTERSPEECH | 3 |
| 2000 | CASS: a phonetically transcribed corpus of mandarin spontaneous speech
Thomas Fang Zheng, William J. Byrne, Pascale Fung, Terri Kamm, Yi Liu 0022, Zhanjiang Song, Umar Ruhi, Veera Venkataramani |
INTERSPEECH | 3 |
| 2000 | Minimum risk acoustic clustering for multilingual acoustic model combinationabstractIn this paper we describe procedures for combining multiple acoustic models, obtained using training corpora from different languages, in order to improve ASR performance in languages for which large amounts of training data are not available. We treat these models as multiple sources of information whose scores are combined in a log-linear model to compute the hypothesis likelihood. The model combination can either be performed in a static way, with constant combination weights, or in a dynamic way, with parameters that can vary for different segments of a hypothesis. The aim is to optimize the parameters so as to achieve minimum word error rate. In order to achieve robust parameter estimation in the dynamic combination case, the parameters are defined to be piecewise constant on different phonetic classes that form a partition of the space of hypothesis segments. The partition is defined, using phonological knowledge, on segments that correspond to hypothesized phones. We examine different ways to define such a partition, including an automatic approach that gives a binary tree structured partition which tries to achieve the minimum WER with the minimum number of classes. 1. Dimitra Vergyri, Stavros Tsakalidis, William J. Byrne |
INTERSPEECH | 3 |
| 2000 | Minimum Bayes-risk automatic speech recognitionabstractIn this paper we address application of minimum Bayes-risk classifiers to tasks in automatic speech recognition (ASR). Minimum-risk classifiers are useful because they produce hypotheses in an attempt to be optimal under a specified task-dependent performance criterion. While the form of the optimal classifier is well known, its implementation is prohibitively expensive. We present efficient approximations that can be used to implement these procedures. In particular, anA * search over word lattices produced by a conventional ASR system is described. This algorithm is intended to extend the previously proposed N -best list rescoring approximation to minimum-risk classifiers. We provide experimental results showing that both the A * and N -best list rescoring implementations of minimum-risk classifiers yield better recognition accuracy than the commonly used maximum a posteriori probability (MAP) classifier in word transcription and identification of keywords. TheA * implementation is compared to the N -best list rescoring implementation and is found to obtain modest but significant improvements in accuracy at little additional computational cost. Another application of minimum-risk classifiers for the identification of named entities from speech is presented. Only the N -best list rescoring could be implemented for this task and was found to yield better named entity identification performance than the MAP classifier. Vaibhava Goel, William J. Byrne |
Comput. Speech Lang. | 2 |
| 2000 | Comments on "Efficient training algorithms for HMMs using incremental estimation"abstractThe paper entitled "Efficient training algorithms for HMMs using incremental estimation" by Gotoh et al. (IEEE Trans. Speech Audio Processing, vol.6, p.539-48, Nov. 1998) investigated expectation maximization (EM) procedures that increase training speed. The claim of Gotoh et al. that these procedures are generalized EM (Dempster et al. 1977) procedures is shown to be incorrect in the present paper. We discuss why this is so, provide an example of nonmonotonic convergence to a local maximum in likelihood, and outline conditions that guarantee such convergence. William J. Byrne, Asela Gunawardana |
IEEE Trans. Speech Audio Process. | 1 |
| 1999 | Rapid speech recognizer adaptation to new speakersabstractThis paper summarizes the work of the "Rapid Speech Recognizer Adaptation" team in the workshop held at Johns Hopkins University in the summer of 1998. The project addressed the modeling of dependencies between units of speech with the goal of making more effective use of small amounts of data for speaker adaptation. A variety of methods were investigated and their effectiveness in a rapid adaptation task defined on the Switchboard conversational speech corpus is reported. Vassilios Digalakis, Heather Collier, Sid Berkowitz, Adrian Corduneanu, Enrico Bocchieri, Ashvin Kannan, Constantinos Boulis, Sanjeev Khudanpur, William J. Byrne, Ananth Sankar |
ICASSP | 9 |
| 1999 | Speaker adaptation with all-pass transformsabstractIn previous work, a class of transforms were proposed which achieve a remapping of the frequency axis much like conventional vocal tract length normalization. These mappings, known collectively as all-pass transforms (APT), were shown to produce substantial improvements in the performance of a large vocabulary speech recognition system when used to normalize incoming speech prior to recognition. In this application, the most advantageous characteristic of the APT was its cepstral-domain linearity; this linearity makes speaker normalization simple to implement, and provides for the robust estimation of the parameters characterizing individual speakers. In the current work, we exploit the APT to develop a speaker adaptation scheme in which the cepstral means of a speech recognition model are transformed to better match the speech of a given speaker. In a set of speech recognition experiments conducted on the Switchboard corpus, we report reductions in word error rate of 3.7% absolute. John W. McDonough, William J. Byrne |
ICASSP | 2 |
| 1999 | Discounted likelihood linear regression for rapid adaptationabstractRapid adaptation schemes that employ the EM algorithm may suffer from overtraining problems when used with small amounts of adaptation data. An algorithm to alleviate this problem is derived within the information geometric framework of Csiszar and Tusnady, and is used to improve MLLR adaptation on NAB and Switchboard adaptation tasks. It is shown how this algorithm approximately optimizes a discounted likelihood criterion. 1. INTRODUCTION In speaker independent LVCSR systems, acoustic models are trained from a large amount of training data from many speakers. This is intended to yield robust models that work fairly well with a variety of channel conditions and speakers. As data becomes available for new speakers and channel conditions, model adaptation techniques can be used to adapt the models to the new conditions. We address here the difficult problem of rapid adaptation. A typical instance of this problem is a recognition task in which each speaker speaks only briefly so that onl... William J. Byrne, Asela Gunawardana |
EUROSPEECH | 1 |
| 1999 | Task dependent loss functions in speech recognition: a* search over recognition latticesabstractThis paper presents an approach to processing of ambiguous requests in spoken dialogs in information train timetable service systems. The main problem of human-machine interaction is in the fact, that speaking human does not express some essential information, because it results from the receiver’s knowledge and the dialog context. Therefore the real understanding system have to combine the incomplete semantic information extracted from the analyzed utterance with the previous data stored in the dialog history path. This is provided in the developed dialog system by the dialog module. Its is based on the representation and storage of facts extracted from the previous steps of dialog (dialog path), and on the creation of simple knowledge bases containing the reasoning rules linking the meaning detected in the analyzed user’s utterance with the facts of the dialog path. Vaibhava Goel, William J. Byrne |
EUROSPEECH | 2 |
| 1999 | Single-pass adapted training with all-pass transformsabstractIn recent work, the all-pass transform (APT) was proposed as the basis of a speaker adaptation scheme intended for use with a large vocabulary speech recognition system. It was shown that APT-based adaptation reduces to a linear transformation of cepstral means, much like the better known maximum likelihood linear regression (MLLR), but is specified by far fewer free parameters. Due to its linearity, APT-based adaptation can be used in conjunction with speaker-adapted training (SAT), an algorithm for performing maximum likelihood estimation of the parameters of an HMM when speaker adaptation is to be employed during both training and test. In this work, we propose a refinement of SAT called single-pass adapted training (SPAT) which achieves the same improvement in system performance as SAT but requires much less computation for HMM training. In a set of speech recognition experiments conducted on the Switchboard Corpus, we report a word error rate reduction of 5.3 % absolute using a single, global APT. 1. John W. McDonough, William J. Byrne |
EUROSPEECH | 2 |
| 1999 | Stochastic pronunciation modelling from hand-labelled phonetic corpora
Michael Riley 0001, William J. Byrne, Michael Finke, Sanjeev Khudanpur, Andrej Ljolje, John W. McDonough, Harriet J. Nock, Murat Saraclar, Charles Wooters, George Zavaliagkos |
Speech Commun. | 2 |
| 1998 | Pronunciation modelling using a hand-labelled corpus for conversational speech recognitionabstractAccurately modelling pronunciation variability in conversational speech is an important component of an automatic speech recognition system. We describe some of the projects undertaken in this direction during and after WS97, the Fifth LVCSR Summer Workshop, held at Johns Hopkins University, Baltimore, in July-August, 1997. We first illustrate a use of hand-labelled phonetic transcriptions of a portion of the Switchboard corpus, in conjunction with statistical techniques, to learn alternatives to canonical pronunciations of words. We then describe the use of these alternate pronunciations in an automatic speech recognition system. We demonstrate that the improvement in recognition performance from pronunciation modelling persists as the system is enhanced with better acoustic and language models. William J. Byrne, Michael Finke, Sanjeev Khudanpur, John W. McDonough, Harriet J. Nock, Michael Riley 0001, Murat Saraclar, Charles Wooters, George Zavaliagkos |
ICASSP | 1 |
| 1998 | LVCSR rescoring with modified loss functions: a decision theoretic perspectiveabstractThe problem of speech decoding is considered in a decision theoretic framework and a modified speech decoding procedure to minimize the expected risk under a general loss function is formulated. A specific word error rate loss function is considered and an implementation in an N-best list rescoring procedure is presented. Methods for estimation of the parameters of the resulting decision rules are provided for both supervised and unsupervised training. Preliminary experiments on an LVCSR task show small but statistically significant error rate improvements. Vaibhava Goel, William J. Byrne, Sanjeev Khudanpur |
ICASSP | 2 |
| 1998 | Speaker normalization with all-pass transformsabstractSpeaker normalization is a process in which the short-time features of speech from a given speaker are transformed so as to better match some speaker independent model. Vocal tract length normalization (VTLN) is a popular speaker normalization scheme wherein the frequency axis of the short-time spectrum associated with a particular speaker's speech is rescaled or warped prior to the extraction of cepstral features. In this work, we develop a novel speaker normalization scheme by exploiting the fact that frequency domain transformations similar to that inherent in VTLN can be accomplished entirely in the cepstral domain through the use of conformal maps. We propose a class of such maps, designated all-pass transforms for reasons given hereafter, and rigorously investigate their properties. Theoretical results are provided relating to the transformation of cepstral sequences under these maps. Additionally, all relations necessary to determine maximum likelihood estimates of the mapping parameters are derived for both speaker normalization and adaptation. 1 2 Speaker Normalization with All-Pass Transforms 1 John W. McDonough, William J. Byrne, Xiaoqiang Luo |
ICSLP | 2 |
| 1994 | Spontaneous speech recognition for the credit card corpus using the HTK toolkitabstractThis paper describes a speech recognition system for the credit card corpus which was provided as a baseline for the Summer Workshop on Robust Speech Processing held at the Rutgers CAIP Center in July/August 1993. The system was built using a portable HMM toolkit called HTK. This is described, along with details of the training and testing methods used. Benchmark performance results are presented for mixture Gaussian monophones and tied-state triphones using four different language models. The overall accuracy levels achieved were very low, indicating that this is a very difficult recognition task. The paper concludes with a discussion of some specific problems encountered with the Credit Card data and suggestions for future work.> Steve J. Young, Philip C. Woodland, William J. Byrne |
IEEE Trans. Speech Audio Process. | 3 |
| 1993 | Noise robustness in the auditory representation of speech signals
Kuansan Wang, Shihab A. Shamma, William J. Byrne |
ICASSP (2) | 3 |
| 1992 | Alternating minimization and Boltzmann machine learningabstractTraining a Boltzmann machine with hidden units is appropriately treated in information geometry using the information divergence and the technique of alternating minimization. The resulting algorithm is shown to be closely related to gradient descent Boltzmann machine learning rules, and the close relationship of both to the EM algorithm is described. An iterative proportional fitting procedure for training machines without hidden units is described and incorporated into the alternating minimization algorithm. William J. Byrne |
IEEE Trans. Neural Networks | 1 |