Andy Way

dblp:69/5430 · DBLP profile ↗
← Back
172ranked-venue papers
10as first author
17since 2021 · last 2024
0000-0001-5736-5930ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 170 · 10 first-author · 17 since 2021Databases, data management, data science and information retrieval · 5Graphics, computer vision, multimedia, augmented reality and games · 3
YearPublicationVenuePosition
2024 SignON - a Co-creative Machine Translation for Sign and Spoken Languages (end-of-project results, contributions and lessons learned)
abstract
SignON, a 3-year Horizon 20202 project addressing the lack of technology and services for MT between sign languages (SLs) and spoken languages (SpLs) ended in December 2023. SignON was unprecedented. Not only it addressed the wider complexity of the aforementioned problem – from research and development of recognition, translation and synthesis, through development of easy-to-use mobile applications and a cloud-based framework to do the “heavy lifting” as well as to establishing ethical, privacy and inclusivenesspolicies and operation guidelines – but also engaged with the deaf and hard of hearing communities in an effective co-creation approach where these main stakeholders drove the development in the right direction and had the final say.Currently we are witnessing advances in natural language processing for SLs, including MT. SignON was one of the largest projects that contributed to this surge with 17 partners and more than 60 consortium members, working in parallel with other international and European initiatives, such as project EASIER and others.
Dimitar Sht. Shterionov, Vincent Vandeghinste, Mirella De Sisto, Aoife Brady, Mathieu De Coster, Lorraine Leeson, Andy Way, Josep Blat, Frankie Picron, Davy Van Landuyt, Marcello Paolo Scipioni, Aditya Parikh, Louis ten Bosch, John J. O'Flaherty, Joni Dambre, Caro Brosens, Jorn Rijckaert, Víctor Ubieto Nogales, Bram Vanroy, Santiago Egea Gómez, Ineke Schuurman, Gorka Labaka, Adrián Núñez-Marcos, Irene Murtagh, Euan McGill, Horacio Saggion
EAMT (2)7
2023 Adaptive Machine Translation with Large Language Models
abstract
Consistency is a key requirement of high-quality translation. It is especially important to adhere to pre-approved terminology and adapt to corrected translations in domain-specific projects. Machine translation (MT) has achieved significant progress in the area of domain adaptation. However, real-time adaptation remains challenging. Large-scale language models (LLMs) have recently shown interesting capabilities of in-context learning, where they learn to replicate certain input-output text generation patterns, without further fine-tuning. By feeding an LLM at inference time with a prompt that consists of a list of translation pairs, it can then simulate the domain and style characteristics. This work aims to investigate how we can utilize in-context learning to improve real-time adaptive MT. Our extensive experiments show promising results at translation time. For example, GPT-3.5 can adapt to a set of in-domain sentence pairs and/or terminology while translating a new sentence. We observe that the translation quality with few-shot in-context learning can surpass that of strong encoder-decoder MT systems, especially for high-resource languages. Moreover, we investigate whether we can combine MT from strong encoder-decoder models with fuzzy matches, which can further improve translation quality, especially for less supported languages. We conduct our experiments across five diverse language pairs, namely English-to-Arabic (EN-AR), English-to-Chinese (EN-ZH), English-to-French (EN-FR), English-to-Kinyarwanda (EN-RW), and English-to-Spanish (EN-ES).
Yasmin Moslem, Rejwanul Haque, John D. Kelleher, Andy Way
EAMT4
2023 Instance-Based Domain Adaptation for Improving Terminology Translation
abstract
Terms are essential indicators of a domain, and domain term translation is dealt with priority in any translation workflow. Translation service providers who use machine translation (MT) expect term translation to be unambiguous and consistent with the context and domain in question. Although current state-of-the-art neural MT (NMT) models are able to produce high-quality translations for many languages, they are still not at the level required when it comes to translating domain-specific terms. This study presents a terminology-aware instance- based adaptation method for improving terminology translation in NMT. We conducted our experiments for French-to-English and found that our proposed approach achieves a statistically significant improvement over the baseline NMT system in translating domain-specific terms. Specifically, the translation of multi-word terms is improved by 6.7% compared to the strong baseline.
Prashanth Nayak, John D. Kelleher, Rejwanul Haque, Andy Way
MTSummit (1)4
2022 Overview of the ELE Project
abstract
This paper provides an overview of the ongoing European Language Equality(ELE) project, an 18-month action funded by the European Commission which involves 52 partners. The primary goal of ELE is to prepare the European Language Equality Programme, in the form of a strategic research, innovation and implementation agenda and a roadmap for achieving full digital language equality (DLE) in Europe by 2030.
Itziar Aldabe, Jane Dunne, Aritz Farwell, Owen Gallagher, Federico Gaspari, Maria Giagkou, Jan Hajic 0001, Jens Peter Kückens, Teresa Lynn, Georg Rehm, German Rigau, Katrin Marheinecke, Stelios Piperidis, Natália Resende, Tereza Vojtechová, Andy Way
EAMT16
2022 Achievements of the PRINCIPLE Project: Promoting MT for Croatian, Icelandic, Irish and Norwegian
abstract
This paper provides an overview of the main achievements of the completed PRINCIPLE project, a 2-year action funded by the European Commission under the Connecting Europe Facility (CEF) programme. PRINCIPLE focused on collecting high-quality language resources for Croatian, Icelandic, Irish and Norwegian, which are severely low-resource languages, especially for building effective machine translation (MT) systems. We report the achievements of the project, primarily, in terms of the large amounts of data collected for all four low-resource languages and of promoting the uptake of neural MT (NMT) for these languages.
Petra Bago, Sheila Castilho, Jane Dunne, Federico Gaspari, Andre Kåsen, Gauti Kristmannsson, Jon Arild Olsen, Natália Resende, Níels Rúnar Gíslason, Dana Davis Sheridan, Paraic Sheridan, John Tinsley, Andy Way
EAMT13
2022 Developing Machine Translation Engines for Multilingual Participatory Spaces
abstract
It is often a challenging task to build Machine Translation (MT) engines for a specific domain due to the lack of parallel data in that area. In this project, we develop a range of MT systems for 6 European languages (English, German, Italian, French, Polish and Irish) in all directions and in two domains (environment and economics).
Pintu Lohar, Guodong Xie, Andy Way
EAMT3
2022 gaHealth: An English-Irish Bilingual Corpus of Health Data
abstract
Machine Translation is a mature technology for many high-resource language pairs. However in the context of low-resource languages, there is a paucity of parallel data datasets available for developing translation models. Furthermore, the development of datasets for low-resource languages often focuses on simply creating the largest possible dataset for generic translation. The benefits and development of smaller in-domain datasets can easily be overlooked. To assess the merits of using in-domain data, a dataset for the specific domain of health was developed for the low-resource English to Irish language pair. Our study outlines the process used in developing the corpus and empirically demonstrates the benefits of using an in-domain dataset for the health domain. In the context of translating health-related data, models developed using the gaHealth corpus demonstrated a maximum BLEU score improvement of 22.2 points (40%) when compared with top performing models from the LoResMT2021 Shared Task. Furthermore, we define linguistic guidelines for developing gaHealth, the first bilingual corpus of health data for the Irish language, which we hope will be of use to other creators of low-resource data sets. gaHealth is now freely available online and is ready to be explored for further research.
Séamus Lankford, Haithem Afli, Orla Ni Loinsigh, Andy Way
LREC4
2022 Improved feature decay algorithms for statistical machine translation
abstract
Abstract In machine-learning applications, data selection is of crucial importance if good runtime performance is to be achieved. In a scenario where the test set is accessible when the model is being built, training instances can be selected so they are the most relevant for the test set. Feature Decay Algorithms (FDA) are a technique for data selection that has demonstrated excellent performance in a number of tasks. This method maximizes the diversity of the n-grams in the training set by devaluing those ones that have already been included. We focus on this method to undertake deeper research on how to select better training data instances. We give an overview of FDA and propose improvements in terms of speed and quality. Using German-to-English parallel data, first we create a novel approach that decreases the execution time of FDA when multiple computation units are available. In addition, we obtain improvements on translation quality by extending FDA using information from the parallel corpus that is generally ignored.
Alberto Poncelas, Gideon Maillette de Buy Wenniger, Andy Way
Nat. Lang. Eng.3
2021 Augmenting Training Data for Low-Resource Neural Machine Translation via Bilingual Word Embeddings and BERT Language Modelling
Akshai Ramesh, Haque Usuf Uhana, Venkatesh Balavadhani Parthasarathy, Rejwanul Haque, Andy Way
IJCNN5
2021 Transformers for Low-Resource Languages: Is Féidir Linn!
abstract
The Transformer model is the state-of-the-art in Machine Translation. However and in general and neural translation models often under perform on language pairs with insufficient training data. As a consequence and relatively few experiments have been carried out using this architecture on low-resource language pairs. In this study and hyperparameter optimization of Transformer models in translating the low-resource English-Irish language pair is evaluated. We demonstrate that choosing appropriate parameters leads to considerable performance improvements. Most importantly and the correct choice of subword model is shown to be the biggest driver of translation performance. SentencePiece models using both unigram and BPE approaches were appraised. Variations on model architectures included modifying the number of layers and testing various regularization techniques and evaluating the optimal number of heads for attention. A generic 55k DGT corpus and an in-domain 88k public admin corpus were used for evaluation. A Transformer optimized model demonstrated a BLEU score improvement of 7.8 points when compared with a baseline RNN model. Improvements were observed across a range of metrics and including TER and indicating a substantially reduced post editing effort for Transformer optimized models with 16k BPE subword models. Bench-marked against Google Translate and our translation engines demonstrated significant improvements. The question of whether or not Transformers can be used effectively in a low-resource setting of English-Irish translation has been addressed. Is féidir linn - yes we can.
Séamus Lankford, Haithem Alfi, Andy Way
MTSummit (1)3
2021 A review of the state-of-the-art in automatic post-editing
abstract
This article presents a review of the evolution of automatic post-editing, a term that describes methods to improve the output of machine translation systems, based on knowledge extracted from datasets that include post-edited content. The article describes the specificity of automatic post-editing in comparison with other tasks in machine translation, and it discusses how it may function as a complement to them. Particular detail is given in the article to the five-year period that covers the shared tasks presented in WMT conferences (2015-2019). In this period, discussion of automatic post-editing evolved from the definition of its main parameters to an announced demise, associated with the difficulties in improving output obtained by neural methods, which was then followed by renewed interest. The article debates the role and relevance of automatic post-editing, both as an academic endeavour and as a useful application in commercial workflows.
Félix do Carmo, Dimitar Sht. Shterionov, Joss Moorkens, Joachim Wagner 0001, Murhaf Hossari, Eric Paquin, Dag Schmidtke, Declan Groves, Andy Way
Mach. Transl.9
2021 Augmenting training data with syntactic phrasal-segments in low-resource neural machine translation
Kamal Kumar Gupta, Sukanta Sen, Rejwanul Haque, Asif Ekbal, Pushpak Bhattacharyya, Andy Way
Mach. Transl.6
2021 Recent advances of low-resource neural machine translation
Rejwanul Haque, Chao-Hong Liu, Andy Way
Mach. Transl.3
2021 Philipp Koehn: Neural Machine Translation
abstract
Abstract Neural machine translation (NMT) is an approach to machine translation (MT) that uses deep learning techniques, a broad area of machine learning based on deep artificial neural networks (NNs). The book Neural Machine Translation by Philipp Koehn targets a broad range of readers including researchers, scientists, academics, advanced undergraduate or postgraduate students, and users of MT, covering wider topics including fundamental and advanced neural network-based learning techniques and methodologies used to develop NMT systems. The book demonstrates different linguistic and computational aspects in terms of NMT with the latest practices and standards and investigates problems relating to NMT. Having read this book, the reader should be able to formulate, design, implement, critically assess and evaluate some of the fundamental and advanced deep learning techniques and methods used for MT. Koehn himself notes that he was somewhat overtaken by events, as originally this book was envisaged only as a chapter in a revised, extended version of his 2009 book Statistical Machine Translation . However, in the interim, NMT completely overtook this previously dominant paradigm, and this new book is likely to serve as the reference of note for the field for some time to come, despite the fact that new techniques are coming onstream all the time.
Wandri Jooste, Rejwanul Haque, Andy Way
Mach. Transl.3
2021 From MT to LREV: managing the transition
Andy Way
Mach. Transl.1
2021 Neural machine translation of low-resource languages using SMT phrase pair injection
abstract
Abstract Neural machine translation (NMT) has recently shown promising results on publicly available benchmark datasets and is being rapidly adopted in various production systems. However, it requires high-quality large-scale parallel corpus, and it is not always possible to have sufficiently large corpus as it requires time, money, and professionals. Hence, many existing large-scale parallel corpus are limited to the specific languages and domains. In this paper, we propose an effective approach to improve an NMT system in low-resource scenario without using any additional data. Our approach aims at augmenting the original training data by means of parallel phrases extracted from the original training data itself using a statistical machine translation (SMT) system. Our proposed approach is based on the gated recurrent unit (GRU) and transformer networks. We choose the Hindi–English, Hindi–Bengali datasets for Health, Tourism, and Judicial (only for Hindi–English) domains. We train our NMT models for 10 translation directions, each using only 5–23k parallel sentences. Experiments show the improvements in the range of 1.38–15.36 BiLingual Evaluation Understudy points over the baseline systems. Experiments show that transformer models perform better than GRU models in low-resource scenarios. In addition to that, we also find that our proposed method outperforms SMT—which is known to work better than the neural models in low-resource scenarios—for some translation directions. In order to further show the effectiveness of our proposed model, we also employ our approach to another interesting NMT task, for example, old-to-modern English translation, using a tiny parallel corpus of only 2.7K sentences. For this task, we use publicly available old-modern English text which is approximately 1000 years old. Evaluation for this task shows significant improvement over the baseline NMT.
Sukanta Sen, Mohammed Hasanuzzaman, Asif Ekbal, Pushpak Bhattacharyya, Andy Way
Nat. Lang. Eng.5
2021 Reinforced NMT for Sentiment and Content Preservation in Low-resource Scenario
abstract
The preservation of domain knowledge from source to the target is crucial in any translation workflows. Hence, translation service providers that use machine translation (MT) in production could reasonably expect that the translation process should transfer both the underlying pragmatics and the semantics of the source-side sentences into the target language. However, recent studies suggest that the MT systems often fail to preserve such crucial information (e.g., sentiment, emotion, gender traits) embedded in the source text in the target. In this context, the raw automatic translations are often directly fed to other natural language processing (NLP) applications (e.g., sentiment classifier) in a cross-lingual platform. Hence, the loss of such crucial information during the translation could negatively affect the performance of such downstream NLP tasks that heavily rely on the output of the MT systems. In our current research, we carefully balance both the sides (i.e., sentiment and semantics) during translation, by controlling a global-attention-based neural MT (NMT), to generate translations that encode the underlying sentiment of a source sentence while preserving its non-opinionated semantic content. Toward this, we use a state-of-the-art reinforcement learning method, namely, actor-critic , that includes a novel reward combination module, to fine-tune the NMT system so that it learns to generate translations that are best suited for a downstream task, viz. sentiment classification while ensuring the source-side semantics is intact in the process. Experimental results for Hindi–English language pair show that our proposed method significantly improves the performance of the sentiment classifier and alongside results in an improved NMT system.
Divya Kumari, Asif Ekbal, Rejwanul Haque, Pushpak Bhattacharyya, Andy Way
ACM Trans. Asian Low Resour. Lang. Inf. Process.5
2020 Selecting Backtranslated Data from Multiple Sources for Improved Neural Machine Translation
abstract
Machine translation (MT) has benefited from using synthetic training data originating from translating monolingual corpora, a technique known as backtranslation.Combining backtranslated data from different sources has led to better results than when using such data in isolation.In this work we analyse the impact that data translated with rule-based, phrasebased statistical and neural MT systems has on new MT systems.We use a real-world low-resource use-case (Basque-to-Spanish in the clinical domain) as well as a high-resource language pair (German-to-English) to test different scenarios with backtranslation and employ data selection to optimise the synthetic corpora.We exploit different data selection strategies in order to reduce the amount of data used, while at the same time maintaining highquality MT systems.We further tune the data selection method by taking into account the quality of the MT systems used for backtranslation and lexical diversity of the resulting corpora.Our experiments show that incorporating backtranslated data from different sources can be beneficial, and that availing of data selection can yield improved performance.
Xabier Soto, Dimitar Sht. Shterionov, Alberto Poncelas, Andy Way
ACL4
2020 A human evaluation of English-Irish statistical and neural machine translation
abstract
With official status in both Ireland and the EU, there is a need for high-quality English-Irish (EN-GA) machine translation (MT) systems which are suitable for use in a professional translation environment. While we have seen recent research on improving both statistical MT and neural MT for the EN-GA pair, the results of such systems have always been reported using automatic evaluation metrics. This paper provides the first human evaluation study of EN-GA MT using professional translators and in-domain (public administration) data for a more accurate depiction of the translation quality available via MT.
Meghan Dowling, Sheila Castilho, Joss Moorkens, Teresa Lynn, Andy Way
EAMT5
2020 Modelling Source- and Target- Language Syntactic Information as Conditional Context in Interactive Neural Machine Translation
abstract
In interactive machine translation (MT), human translators correct errors in automatic translations in collaboration with the MT systems, which is seen as an effective way to improve the productivity gain in translation. In this study, we model source-language syntactic constituency parse and target-language syntactic descriptions in the form of supertags as conditional context for interactive prediction in neural MT (NMT). We found that the supertags significantly improve productivity gain in translation in interactive-predictive NMT (INMT), while syntactic parsing somewhat found to be effective in reducing human effort in translation. Furthermore, when we model this source- and target-language syntactic information together as the conditional context, both types complement each other and our fully syntax-informed INMT model statistically significantly reduces human efforts in a French–to–English translation task, achieving 4.30 points absolute (corresponding to 9.18% relative) improvement in terms of word prediction accuracy (WPA) and 4.84 points absolute (corresponding to 9.01% relative) reduction in terms of word stroke ratio (WSR) over the baseline.
Kamal Kumar Gupta, Rejwanul Haque, Asif Ekbal, Pushpak Bhattacharyya, Andy Way
EAMT5
2020 MT syntactic priming effects on L2 English speakers
abstract
In this paper, we tested 20 Brazilian Portuguese speakers at intermediate and advanced English proficiency levels to investigate the influence of Google Translate’s MT system on the mental processing of English as a second language. To this end, we employed a syntactic priming experimental paradigm using a pretest-priming design which allowed us to compare participants’ linguistic behaviour before and after a translation task using Google Translate. Results show that, after performing a translation task with Google Translate, participants more frequently described images in English using the syntactic alternative previously seen in the output of Google Translate, compared to the translation task with no prior influence of the MT output. Results also show that this syntactic priming effect is modulated by English proficiency levels.
Natália Resende, Benjamin R. Cowan, Andy Way
EAMT3
2020 MTrill project: Machine Translation impact on language learning
abstract
Over the last decades, massive research investments have been made in the development of machine translation (MT) systems (Gupta and Dhawan, 2019). This has brought about a paradigm shift in the performance of these language tools, leading to widespread use of popular MT systems (Gaspari and Hutchins, 2007). Although the first MT engines were used for gisting purposes, in recent years, there has been an increasing interest in using MT tools, especially the freely available online MT tools, for language teaching and learning (Clifford et al., 2013). The literature on MT and Computer Assisted Language Learning (CALL) shows that, over the years, MT systems have been facilitating language teaching and also language learning (Nin ̃o, 2006). It has been shown that MT tools can increase awareness of grammatical linguistic features of a foreign language. Research also shows the positive role of MT systems in the development of writing skills in English as well as in improving communication skills in English(Garcia and Pena, 2011). However, to date, the cognitive impact of MT on language acquisition and on the syntactic aspects of language processing has not yet been investigated and deserves further scrutiny. The MTril project aims at filling this gap in the literature by examining whether MT is contributing to a central aspect of language acquisition: the so-called language binding, i.e., the ability to combine single words properly in a grammatical sentence (Heyselaar et al., 2017; Ferreira and Bock, 2006). The project focus on the initial stages (pre-intermediate and intermediate) of the acquisition of English syntax by Brazilian Portuguese native speakers using MT systems as a support for language learning.
Natália Resende, Andy Way
EAMT2
2020 Progress of the PRINCIPLE Project: Promoting MT for Croatian, Icelandic, Irish and Norwegian
abstract
This paper updates the progress made on the PRINCIPLE project, a 2-year action funded by the European Commission under the Connecting Europe Facility (CEF) programme. PRINCIPLE focuses on collecting high-quality language resources for Croatian, Icelandic, Irish and Norwegian, which have been identified as low-resource languages, especially for building effective machine translation (MT) systems. We report initial achievements of the project and ongoing activities aimed at promoting the uptake of neural MT for the low-resource languages of the project.
Andy Way, Petra Bago, Jane Dunne, Federico Gaspari, Andre Kåsen, Gauti Kristmannsson, Helen McHugh, Jon Arild Olsen, Dana Davis Sheridan, Paraic Sheridan, John Tinsley
EAMT1
2020 Syntax-Informed Interactive Neural Machine Translation
abstract
In interactive machine translation (MT), human translators correct errors in automatic translations in collaboration with the MT systems, and this is an effective way to improve productivity gain in translation. Phrase-based statistical MT (PB-SMT) has been the mainstream approach to MT for the past 30 years, both in academia and industry. Neural MT (NMT), an end-to-end learning approach to MT, represents the current state-of-the-art in MT research. The recent studies on interactive MT have indicated that NMT can significantly outperform PB-SMT. In this work, first we investigate the possibility of integrating lexical syntactic descriptions in the form of supertags into the state-of-the-art NMT model, Transformer. Then, we explore whether integration of supertags into Transformer could indeed reduce human efforts in translation in an interactive-predictive platform. From our investigation we found that our syntax-aware interactive NMT (INMT) framework significantly reduces simulated human efforts in the French-to-English and Hindi- to-English translation tasks, achieving a 2.65 point absolute corresponding to 5.65% relative improvement and a 6.55 point absolute corresponding to 19.1% relative improvement, respectively, in terms of word prediction accuracy (WPA) over the respective baselines.
Kamal Kumar Gupta, Rejwanul Haque, Asif Ekbal, Pushpak Bhattacharyya, Andy Way
IJCNN5
2020 On Context Span Needed for Machine Translation Evaluation
abstract
Despite increasing efforts to improve evaluation of machine translation (MT) by going beyond the sentence level to the document level, the definition of what exactly constitutes a “document level” is still not clear. This work deals with the context span necessary for a more reliable MT evaluation. We report results from a series of surveys involving three domains and 18 target languages designed to identify the necessary context span as well as issues related to it. Our findings indicate that, despite the fact that some issues and spans are strongly dependent on domain and on the target language, a number of common patterns can be observed so that general guidelines for context-aware MT evaluation can be drawn.
Sheila Castilho, Maja Popovic, Andy Way
LREC3
2020 The European Language Technology Landscape in 2020: Language-Centric and Human-Centric AI for Cross-Cultural Communication in Multilingual Europe
abstract
Multilingualism is a cultural cornerstone of Europe and firmly anchored in the European treaties including full language equality. However, language barriers impacting business, cross-lingual and cross-cultural communication are still omnipresent. Language Technologies (LTs) are a powerful means to break down these barriers. While the last decade has seen various initiatives that created a multitude of approaches and technologies tailored to Europe’s specific needs, there is still an immense level of fragmentation. At the same time, AI has become an increasingly important concept in the European Information and Communication Technology area. For a few years now, AI – including many opportunities, synergies but also misconceptions – has been overshadowing every other topic. We present an overview of the European LT landscape, describing funding programmes, activities, actions and challenges in the different countries with regard to LT, including the current state of play in industry and the LT market. We present a brief overview of the main LT-related activities on the EU level in the last ten years and develop strategic guidance with regard to four key dimensions.
Georg Rehm, Katrin Marheinecke, Stefanie Hegele, Stelios Piperidis, Kalina Bontcheva, Jan Hajic 0001, Khalid Choukri, Andrejs Vasiljevs, Gerhard Backfried, Christoph Prinz, José Manuél Gómez-Pérez, Luc Meertens, Paul Lukowicz, Josef van Genabith, Andrea Lösch, Philipp Slusallek, Morten Irgens, Patrick Gatellier, Joachim Köhler, Laure Le Bars, Dimitra Anastasiou, Albina Auksoriute, Núria Bel, António Branco, Gerhard Budin, Walter Daelemans, Koenraad De Smedt, Radovan Garabík, Maria Gavrilidou, Dagmar Gromann, Svetla Koeva, Simon Krek, Cvetana Krstev, Krister Lindén, Bernardo Magnini, Jan Odijk, Maciej Ogrodniczuk, Eiríkur Rögnvaldsson, Mike Rosner, Bolette S. Pedersen, Inguna Skadina, Marko Tadic, Dan Tufis, Tamás Váradi, Kadri Vider, Andy Way, François Yvon
LREC46
2020 Investigating Query Expansion and Coreference Resolution in Question Answering on BERT
Santanu Bhattacharjee, Rejwanul Haque, Gideon Maillette de Buy Wenniger, Andy Way
NLDB4
2020 Analysing terminology translation errors in statistical and neural machine translation
Rejwanul Haque, Mohammed Hasanuzzaman, Andy Way
Mach. Transl.3
2020 A roadmap to neural automatic post-editing: an empirical approach
abstract
In a translation workflow, machine translation (MT) is almost always followed by a human post-editing step, where the raw MT output is corrected to meet required quality standards. To reduce the number of errors human translators need to correct, automatic post-editing (APE) methods have been developed and deployed in such workflows. With the advances in deep learning, neural APE (NPE) systems have outranked more traditional, statistical, ones. However, the plethora of options, variables and settings, as well as the relation between NPE performance and train/test data makes it difficult to select the most suitable approach for a given use case. In this article, we systematically analyse these different parameters with respect to NPE performance. We build an NPE "roadmap" to trace the different decision points and train a set of systems selecting different options through the roadmap. We also propose a novel approach for APE with data augmentation. We then analyse the performance of 15 of these systems and identify the best ones. In fact, the best systems are the ones that follow the newly-proposed method. The work presented in this article follows from a collaborative project between Microsoft and the ADAPT centre. The data provided by Microsoft originates from phrase-based statistical MT (PBSMT) systems employed in production. All tested NPE systems significantly increase the translation quality, proving the effectiveness of neural post-editing in the context of a commercial translation workflow that leverages PBSMT.
Dimitar Sht. Shterionov, Félix do Carmo, Joss Moorkens, Murhaf Hossari, Joachim Wagner 0001, Eric Paquin, Dag Schmidtke, Declan Groves, Andy Way
Mach. Transl.9
2019 Evaluating Terminology Translation in MT
Rejwanul Haque, Mohammed Hasanuzzaman, Andy Way
CICLing (1)3
2019 Adaptation of Machine Translation Models with Back-Translated Data Using Transductive Data Selection Methods
Alberto Poncelas, Gideon Maillette de Buy Wenniger, Andy Way
CICLing (1)3
2019 Take Help from Elder Brother: Old to Modern English NMT with Phrase Pair Feedback
Sukanta Sen, Mohammed Hasanuzzaman, Asif Ekbal, Pushpak Bhattacharyya, Andy Way
CICLing (1)5
2019 No Padding Please: Efficient Neural Handwriting Recognition
abstract
Neural handwriting recognition (NHR) is the recognition of handwritten text with deep learning models, such as multi-dimensional long short-term memory (MDLSTM) recurrent neural networks. Models with MDLSTM layers have achieved state-of-the art results on handwritten text recognition tasks. While multi-directional MDLSTM-layers have an unbeaten ability to capture the complete context in all directions, this strength limits the possibilities for parallelization, and therefore comes at a high computational cost. In this work we develop methods to create efficient MDLSTM-based models for NHR, particularly a method aimed at eliminating computation waste that results from padding. This proposed method, called example packing, replaces wasteful stacking of padded examples with efficient tiling in a 2-dimensional grid. For word-based NHR this yields a speed improvement of factor 6.6 over an already efficient baseline of minimal padding for each batch separately. For line-based NHR the savings are more modest, but still significant. In addition to example packing, we propose: 1) a technique to optimize parallelization for dynamic graph definition frameworks including PyTorch, using convolutions with grouping, 2) a method for parallelization across GPUs for variable-length example batches. All our techniques are thoroughly tested on our own PyTorch re-implementation of MDLSTM-based NHR models. A thorough evaluation on the IAM dataset shows that our models are performing similar to earlier implementations of state-of-the art models. Our efficient NHR model and some of the reusable techniques discussed with it offer ways to realize relatively efficient models for the omnipresent scenario of variable-length inputs in deep learning.
Gideon Maillette de Buy Wenniger, Lambert Schomaker, Andy Way
ICDAR3
2019 Selecting Artificially-Generated Sentences for Fine-Tuning Neural Machine Translation
abstract
Neural Machine Translation (NMT) models tend to achieve best performance when larger sets of parallel sentences are provided for training.For this reason, augmenting the training set with artificially-generated sentence pairs can boost performance.Nonetheless, the performance can also be improved with a small number of sentences if they are in the same domain as the test set.Accordingly, we want to explore the use of artificially-generated sentences along with data-selection algorithms to improve Germanto-English NMT models trained solely with authentic data.In this work, we show how artificiallygenerated sentences can be more beneficial than authentic pairs, and demonstrate their advantages when used in combination with dataselection algorithms.
Alberto Poncelas, Andy Way
INLG2
2019 Large-scale Machine Translation Evaluation of the iADAATPA Project
Sheila Castilho, Natália Resende, Federico Gaspari, Andy Way, Tony O'Dowd, Marek Mazur, Manuel Herranz, Alexandre Helle, Gema Ramírez-Sánchez, Víctor M. Sánchez-Cartagena, Marcis Pinnis, Valters Sics
MTSummit (2)4
2019 Pivot Machine Translation in INTERACT Project
Chao-Hong Liu, Andy Way, Catarina Cruz Silva
MTSummit (2)2
2019 When less is more in Neural Quality Estimation of Machine Translation. An industry case study
Dimitar Sht. Shterionov, Félix do Carmo, Joss Moorkens, Eric Paquin, Dag Schmidtke, Declan Groves, Andy Way
MTSummit (2)7
2019 Lost in Translation: Loss and Decay of Linguistic Richness in Machine Translation
Eva Vanmassenhove, Dimitar Sht. Shterionov, Andy Way
MTSummit (1)3
2019 PRINCIPLE: Providing Resources in Irish, Norwegian, Croatian and Icelandic for the Purposes of Language Engineering
Andy Way, Federico Gaspari
MTSummit (2)1
2019 Ruslan Mitkov, Johanna Monti, Gloria Corpas Pastor, and Violeta Seretan (eds): Multiword units in machine translation and translation technology - Current Issues in Linguistic Theory, Volume 341, John Benjamin Publishing Company, Amsterdam & Philadelphia, 2018, ix+259 pp, ISBN 978-90-272-0060-0 (HB), ISBN 978-90-272-6420-6 (e-book)
Rejwanul Haque, Mohammed Hasanuzzaman, Andy Way
Mach. Transl.3
2019 Post-editing neural machine translation versus translation memory segments
Pilar Sánchez-Gijón, Joss Moorkens, Andy Way
Mach. Transl.3
2018 A Decision-Level Approach to Multimodal Sentiment Analysis
Haithem Afli, Jason Burns, Andy Way
CICLing (2)3
2018 Incorporating Deep Visual Features into Multiobjective based Multi-view Search Results Clustering
abstract
Current paper explores the use of multi-view learning for search result clustering. A web-snippet can be represented using multiple views. Apart from textual view cued by both the semantic and syntactic information, a complimentary view extracted from images contained in the web-snippets is also utilized in the current framework. A single consensus partitioning is finally obtained after consulting these two individual views by the deployment of a multiobjective based clustering technique. Several objective functions including the values of a cluster quality measure measuring the goodness of partitionings obtained using different views and an agreement-disagreement index, quantifying the amount of oneness among multiple views in generating partitionings are optimized simultaneously using AMOSA. In order to detect the number of clusters automatically, concepts of variable length solutions and a vast range of permutation operators are introduced in the clustering process. Finally, a set of alternative partitioning are obtained on the final Pareto front by the proposed multi-view based multiobjective technique. Experimental results by the proposed approach on several benchmark test datasets of SRC with respect to different performance metrics evidently establish the power of visual and text-based views in achieving better search result clustering.
Sayantan Mitra, Mohammed Hasanuzzaman, Sriparna Saha 0001, Andy Way
COLING4
2018 Tailoring Neural Architectures for Translating from Morphologically Rich Languages
abstract
A morphologically complex word (MCW) is a hierarchical constituent with meaning-preserving subunits, so word-based models which rely on surface forms might not be powerful enough to translate such structures. When translating from morphologically rich languages (MRLs), a source word could be mapped to several words or even a full sentence on the target side, which means an MCW should not be treated as an atomic unit. In order to provide better translations for MRLs, we boost the existing neural machine translation (NMT) architecture with a double- channel encoder and a double-attentive decoder. The main goal targeted in this research is to provide richer information on the encoder side and redesign the decoder accordingly to benefit from such information. Our experimental results demonstrate that we could achieve our goal as the proposed model outperforms existing subword- and character-based architectures and showed significant improvements on translating from German, Russian, and Turkish into English.
Peyman Passban, Andy Way, Qun Liu 0001
COLING2
2018 ELRI - European Language Resources Infrastructure
abstract
We describe the European Language Resources Infrastructure project, whose main aim is the provision of an infrastructure to help collect, prepare and share language resources that can in turn improve translation services in Europe.
Thierry Etchegoyhen, Borja Anza Porras, Andoni Azpeitia, Eva Martínez Garcia, Paulo Vale, José Luis Fonseca, Teresa Lynn, Jane Dunne, Federico Gaspari, Andy Way, Victoria Arranz, Khalid Choukri, Vladimir Popescu, Pedro Neiva, Rui Neto, Maite Melero, David Pérez-Fernández, António Branco, Ruben Branco, Luís Gomes 0002
EAMT10
2018 Investigating Backtranslation in Neural Machine Translation
abstract
A prerequisite for training corpus-based machine translation (MT) systems – either Statistical MT (SMT) or Neural MT (NMT) – is the availability of high-quality parallel data. This is arguably more important today than ever before, as NMT has been shown in many studies to outperform SMT, but mostly when large parallel corpora are available; in cases where data is limited, SMT can still outperform NMT. Recently researchers have shown that back-translating monolingual data can be used to create synthetic parallel corpora, which in turn can be used in combination with authentic parallel data to train a highquality NMT system. Given that large collections of new parallel text become available only quite rarely, backtranslation has become the norm when building state-of-the-art NMT systems, especially in resource-poor scenarios. However, we assert that there are many unknown factors regarding the actual effects of back-translated data on the translation capabilities of an NMT model. Accordingly, in this work we investigate how using back-translated data as a training corpus – both as a separate standalone dataset as well as combined with human-generated parallel data – affects the performance of an NMT model. We use incrementally larger amounts of back-translated data to train a range of NMT systems for German-to-English, and analyse the resulting translation performance.
Alberto Poncelas, Dimitar Sht. Shterionov, Andy Way, Gideon Maillette de Buy Wenniger, Peyman Passban
EAMT3
2018 Feature Decay Algorithms for Neural Machine Translation
abstract
Neural Machine Translation (NMT) systems require a lot of data to be competitive. For this reason, data selection techniques are used only for finetuning systems that have been trained with larger amounts of data. In this work we aim to use Feature Decay Algorithms (FDA) data selection techniques not only to fine-tune a system but also to build a complete system with less data. Our findings reveal that it is possible to find a subset of sentence pairs, that outperforms by 1.11 BLEU points the full training corpus, when used for training a German-English NMT system .
Alberto Poncelas, Gideon Maillette de Buy Wenniger, Andy Way
EAMT3
2018 Perception vs. Acceptability of TM and SMT Output: What do translators prefer?
abstract
This paper reports the results of two studies carried out with two different group of professional translators to find out how professionals perceive and accept SMT in comparison with TM. The first group translated and post-edited segments from English into German, and the second group from English into Spanish. Both studies had equivalent settings in order to guarantee the comparability of the results. It will also help to shed light upon the real benefit of SMT from which translators may take advantage.
Pilar Sánchez-Gijón, Joss Moorkens, Andy Way
EAMT3
2018 Project PiPeNovel: Pilot on Post-editing Novels
abstract
Given (i) the rise of a new paradigm to machine translation based on neural networks that results in more fluent and less literal output than previous models and (ii) the maturity of machine-assisted translation via post-editing in industry, project PiPeNovel studies the feasibility of the post-editing workflow for literary text conducting experiments with professional literary translators.
Antonio Toral, Martijn Wieling 0001, Sheila Castilho, Joss Moorkens, Andy Way
EAMT5
2018 Multi-Level Structured Self-Attentions for Distantly Supervised Relation Extraction
abstract
Attention mechanisms are often used in deep neural networks for distantly supervised relation extraction (DS-RE) to distinguish valid from noisy instances.However, traditional 1-D vector attention models are insufficient for the learning of different contexts in the selection of valid instances to predict the relationship for an entity pair.To alleviate this issue, we propose a novel multi-level structured (2-D matrix) self-attention mechanism for DS-RE in a multi-instance learning (MIL) framework using bidirectional recurrent neural networks.In the proposed method, a structured word-level self-attention mechanism learns a 2-D matrix where each row vector represents a weight distribution for different aspects of an instance regarding two entities.Targeting the MIL issue, the structured sentence-level attention learns a 2-D matrix where each row vector represents a weight distribution on selection of different valid instances.Experiments conducted on two publicly available DS-RE datasets show that the proposed framework with a multi-level structured self-attention mechanism significantly outperform state-of-the-art baselines in terms of PR curves, P@N and F1 measures.
Jinhua Du, Jingguang Han, Andy Way, Dadong Wan
EMNLP3
2018 Getting Gender Right in Neural MT
abstract
Speakers of different languages must attend to and encode strikingly different aspects of the world in order to use their language correctly (Sapir, 1921;Slobin, 1996).One such difference is related to the way gender is expressed in a language.Saying "I am happy" in English, does not encode any additional knowledge of the speaker that uttered the sentence.However, many other languages do have grammatical gender systems and so such knowledge would be encoded.In order to correctly translate such a sentence into, say, French, the inherent gender information needs to be retained/recovered.The same sentence would become either "Je suis heureux", for a male speaker or "Je suis heureuse" for a female one.Apart from morphological agreement, demographic factors (gender, age, etc.) also influence our use of language in terms of word choices or even on the level of syntactic constructions (Tannen, 1991;Pennebaker et al., 2003).We integrate gender information into NMT systems.Our contribution is twofold: (1) the compilation of large datasets with speaker information for 20 language pairs, and (2) a simple set of experiments that incorporate gender information into NMT for multiple language pairs.Our experiments show that adding a gender feature to an NMT system significantly improves the translation quality for some language pairs.
Eva Vanmassenhove, Christian Hardmeier, Andy Way
EMNLP3
2018 Learning to Jointly Translate and Predict Dropped Pronouns with a Shared Reconstruction Mechanism
abstract
Pronouns are frequently omitted in pro-drop languages, such as Chinese, generally leading to significant challenges with respect to the production of complete translations.Recently, Wang et al. (2018) proposed a novel reconstruction-based approach to alleviating dropped pronoun (DP) translation problems for neural machine translation models.In this work, we improve the original model from two perspectives.First, we employ a shared reconstructor to better exploit encoder and decoder representations.Second, we jointly learn to translate and predict DPs in an end-to-end manner, to avoid the errors propagated from an external DP prediction model.Experimental results show that our approach significantly improves both translation performance and DP prediction accuracy.
Longyue Wang, Zhaopeng Tu, Andy Way, Qun Liu 0001
EMNLP3
2018 FooTweets: A Bilingual Parallel Corpus of World Cup Tweets
Henny Sluyter-Gäthje, Pintu Lohar, Haithem Afli, Andy Way
LREC4
2018 Fine-Grained Temporal Orientation and its Relationship with Psycho-Demographic Correlates
abstract
Sabyasachi Kamila, Mohammed Hasanuzzaman, Asif Ekbal, Pushpak Bhattacharyya, Andy Way. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Sabyasachi Kamila, Mohammed Hasanuzzaman, Asif Ekbal, Pushpak Bhattacharyya, Andy Way
NAACL-HLT5
2018 Improving Character-Based Decoding Using Target-Side Morphological Information for Neural Machine Translation
abstract
Peyman Passban, Qun Liu, Andy Way. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Peyman Passban, Qun Liu 0001, Andy Way
NAACL-HLT3
2018 IDEA: An Interactive Dialogue Translation Demo System Using Furhat Robots
Jinhua Du, Darragh Blake, Longyue Wang, Clare Conran, Declan McKibben, Andy Way
ECML/PKDD (3)6
2018 Evaluating MT for massive open online courses - A multifaceted comparison between PBSMT and NMT systems
Sheila Castilho, Joss Moorkens, Federico Gaspari, Rico Sennrich, Andy Way, Panayota Georgakopoulou
Mach. Transl.5
2018 Human versus automatic quality evaluation of NMT and PBSMT
Dimitar Sht. Shterionov, Riccardo Superbo, Pat Nagle, Laura Casanellas, Tony O'Dowd, Andy Way
Mach. Transl.6
2018 Editors' foreword to the invited issue on SMT and NMT
Andy Way, Mikel L. Forcada
Mach. Transl.1
2017 Exploiting Cross-Sentence Context for Neural Machine Translation
abstract
In translation, considering the document as a whole can help to resolve ambiguities and inconsistencies.In this paper, we propose a cross-sentence context-aware approach and investigate the influence of historical contextual information on the performance of neural machine translation (NMT).First, this history is summarized in a hierarchical way.We then integrate the historical representation into NMT in two strategies: 1) a warm-start of encoder and decoder states, and 2) an auxiliary context source for updating decoder states.Experimental results on a large Chinese-English translation task show that our approach significantly improves upon a strong attention-based NMT system by up to +2.1 BLEU points.
Longyue Wang, Zhaopeng Tu, Andy Way, Qun Liu 0001
EMNLP3
2017 Demographic Word Embeddings for Racism Detection on Twitter
abstract
Most social media platforms grant users freedom of speech by allowing them to freely express their thoughts, beliefs, and opinions. Although this represents incredible and unique communication opportunities, it also presents important challenges. Online racism is such an example. In this study, we present a supervised learning strategy to detect racist language on Twitter based on word embedding that incorporate demographic (Age, Gender, and Location) information. Our methodology achieves reasonable classification accuracy over a gold standard dataset (F1=76.3%) and significantly improves over the classification performance of demographic-agnostic models.
Mohammed Hasanuzzaman, Gaël Dias, Andy Way
IJCNLP(1)3
2017 Local Event Discovery from Tweets Metadata
abstract
We present a two-step strategy that addresses fundamental deficiencies in social media-based event detection and achieves effective local event by taking advantage of geo-located data from Twitter. While previous work has mainly relied on an analysis of tweet text to identify local events, we show how to reliably detect events using meta-data analysis of geo-tagged tweets. The first step of the method identifies several spatio-temporal clusters within the dataset across both space and time using metadata to form potential candidate events. In the second step, it ranks all the candidates by the amount of hashtag/entity inequality. We used crowdsourcing to evaluate the proposed approach on a data set that contains millions of geo-tagged tweets. The results show that our framework performs reasonably well in terms of precision and discovers local events faster.
Mohammed Hasanuzzaman, Andy Way
K-CAP2
2017 A Comparative Quality Evaluation of PBSMT and NMT using Professional Translators
Sheila Castilho, Joss Moorkens, Federico Gaspari, Rico Sennrich, Vilelmini Sosoni, Panayota Georgakopoulou, Pintu Lohar, Andy Way, Antonio Valerio Miceli Barone, Maria Gialama
MTSummit (1)8
2017 Neural Pre-Translation for Hybrid Machine Translation
Jinhua Du, Andy Way
MTSummit (1)2
2017 Temporality as Seen through Translation: A Case Study on Hindi Texts
Sabyasachi Kamila, Sukanta Sen, Mohammed Hasanuzzaman, Asif Ekbal, Andy Way, Pushpak Bhattacharyya
MTSummit (1)5
2017 The INTERACT Project and Crisis MT
Sharon O'Brien, Chao-Hong Liu, Andy Way, João Graça, Helena Moniz, Ellie Kemp, Rebecca Petras
MTSummit (2)3
2017 Elastic-substitution decoding for Hierarchical SMT: efficiency, richer search and double labels
Gideon Maillette de Buy Wenniger, Khalil Sima'an, Andy Way
MTSummit (1)3
2017 Syntax- and semantic-based reordering in hierarchical phrase-based statistical machine translation
abstract
We present a syntax-based reordering model (RM) for hierarchical phrase-based statistical machine translation (HPB-SMT) enriched with semantic features. Our model brings a number of novel contributions: (i) while the previous dependency-based RM is limited to the reordering of head and dependant constituent pairs, we also model the reordering of pairs of dependants; (ii) Our model is enriched with semantic features (Wordnet synsets) in order to allow the reordering model to generalize to pairs not seen in training but with equivalent meaning. (iii) We evaluate our model on two language directions: English-to-Farsi and English-to-Turkish. These language pairs are particularly challenging due to the free word order, rich morphology and lack of resources of the target languages. We evaluate our RM both intrinsically (accuracy of the RM classifier) and extrinsically (MT). Our best configuration outperforms the baseline classifier by 5–29% on pairs of dependants and by 12–30% on head and dependant pairs while the improvement on MT ranges between 1.6% and 5.5% relative in terms of BLEU depending on language pair and domain. We also analyze the value of the feature weights to obtain further insights on the impact of the reordering-related features in the HPB-SMT model. We observe that the features of our RM are assigned significant weights and that our features are complementary to the reordering feature included by default in the HPB-SMT model.
Arefeh Kazemi, Antonio Toral, Andy Way, S. Amirhassan Monadjemi, Mohammad Ali Nematbakhsh
Expert Syst. Appl.3
2017 A novel and robust approach for pro-drop language translation
abstract
A significant challenge for machine translation (MT) is the phenomena of dropped pronouns (DPs), where certain classes of pronouns are frequently dropped in the source language but should be retained in the target language. In response to this common problem, we propose a semi-supervised approach with a universal framework to recall missing pronouns in translation. Firstly, we build training data for DP generation in which the DPs are automatically labelled according to the alignment information from a parallel corpus. Secondly, we build a deep learning-based DP generator for input sentences in decoding when no corresponding references exist. More specifically, the generation has two phases: (1) DP position detection, which is modeled as a sequential labelling task with recurrent neural networks; and (2) DP prediction, which employs a multilayer perceptron with rich features. Finally, we integrate the above outputs into our statistical MT (SMT) system to recall missing pronouns by both extracting rules from the DP-labelled training data and translating the DP-generated input sentences. To validate the robustness of our approach, we investigate our approach on both Chinese–English and Japanese–English corpora extracted from movie subtitles. Compared with an SMT baseline system, experimental results show that our approach achieves a significant improvement of $$+$$ 1.58 BLEU points in translation performance with 66% F-score for DP generation accuracy for Chinese–English, and nearly $$+$$ 1 BLEU point with 58% F-score for Japanese–English. We believe that this work could help both MT researchers and industries to boost the performance of MT systems between pro-drop and non-pro-drop languages.
Longyue Wang, Zhaopeng Tu, Siyou Liu, Hang Li 0001, Andy Way, Qun Liu 0001
Mach. Transl.6
2017 Translating Low-Resource Languages by Vocabulary Adaptation from Close Counterparts
abstract
Some natural languages belong to the same family or share similar syntactic and/or semantic regularities. This property persuades researchers to share computational models across languages and benefit from high-quality models to boost existing low-performance counterparts. In this article, we follow a similar idea, whereby we develop statistical and neural machine translation (MT) engines that are trained on one language pair but are used to translate another language. First we train a reliable model for a high-resource language, and then we exploit cross-lingual similarities and adapt the model to work for a close language with almost zero resources. We chose Turkish (Tr) and Azeri or Azerbaijani (Az) as the proposed pair in our experiments. Azeri suffers from lack of resources as there is almost no bilingual corpus for this language. Via our techniques, we are able to train an engine for the Az → English (En) direction, which is able to outperform all other existing models.
Peyman Passban, Qun Liu 0001, Andy Way
ACM Trans. Asian Low Resour. Lang. Inf. Process.3
2016 Graph-Based Translation Via Graph Segmentation
abstract
One major drawback of phrase-based translation is that it segments an input sentence into continuous phrases.To support linguistically informed source discontinuity, in this paper we construct graphs which combine bigram and dependency relations and propose a graph-based translation model.The model segments an input graph into connected subgraphs, each of which may cover a discontinuous phrase.We use beam search to combine translations of each subgraph left-to-right to produce a complete translation.Experiments on Chinese-English and German-English tasks show that our system is significantly better than the phrase-based model by up to +1.5/+0.5 BLEU scores.By explicitly modeling the graph segmentation, our system obtains further improvement, especially on German-English.
Liangyou Li, Andy Way, Qun Liu 0001
ACL (1)2
2016 Enriching Phrase Tables for Statistical Machine Translation Using Mixed Embeddings
abstract
The phrase table is considered to be the main bilingual resource for the phrase-based statistical machine translation (PBSMT) model. During translation, a source sentence is decomposed into several phrases. The best match of each source phrase is selected among several target-side counterparts within the phrase table, and processed by the decoder to generate a sentence-level translation. The best match is chosen according to several factors, including a set of bilingual features. PBSMT engines by default provide four probability scores in phrase tables which are considered as the main set of bilingual features. Our goal is to enrich that set of features, as a better feature set should yield better translations. We propose new scores generated by a Convolutional Neural Network (CNN) which indicate the semantic relatedness of phrase pairs. We evaluate our model in different experimental settings with different language pairs. We observe significant improvements when the proposed features are incorporated into the PBSMT pipeline.
Peyman Passban, Qun Liu 0001, Andy Way
COLING3
2016 Topic-Informed Neural Machine Translation
abstract
In recent years, neural machine translation (NMT) has demonstrated state-of-the-art machine translation (MT) performance. It is a new approach to MT, which tries to learn a set of parameters to maximize the conditional probability of target sentences given source sentences. In this paper, we present a novel approach to improve the translation performance in NMT by conveying topic knowledge during translation. The proposed topic-informed NMT can increase the likelihood of selecting words from the same topic and domain for translation. Experimentally, we demonstrate that topic-informed NMT can achieve a 1.15 (3.3% relative) and 1.67 (5.4% relative) absolute improvement in BLEU score on the Chinese-to-English language pair using NIST 2004 and 2005 test sets, respectively, compared to NMT without topic information.
Jian Zhang 0003, Liangyou Li, Andy Way, Qun Liu 0001
COLING3
2016 Fast Gated Neural Domain Adaptation: Language Model as a Case Study
abstract
Neural network training has been shown to be advantageous in many natural language processing applications, such as language modelling or machine translation. In this paper, we describe in detail a novel domain adaptation mechanism in neural network training. Instead of learning and adapting the neural network on millions of training sentences – which can be very time-consuming or even infeasible in some cases – we design a domain adaptation gating mechanism which can be used in recurrent neural networks and quickly learn the out-of-domain knowledge directly from the word vector representations with little speed overhead. In our experiments, we use the recurrent neural network language model (LM) as a case study. We show that the neural LM perplexity can be reduced by 7.395 and 12.011 using the proposed domain adaptation mechanism on the Penn Treebank and News data, respectively. Furthermore, we show that using the domain-adapted neural LM to re-rank the statistical machine translation n-best list on the French-to-English language pair can significantly improve translation quality.
Jian Zhang 0003, Andy Way, Qun Liu 0001
COLING3
2016 Identifying Temporal Orientation of Word Senses
Mohammed Hasanuzzaman, Gaël Dias, Stéphane Ferrari, Yann Mathet, Andy Way
CoNLL5
2016 Comparing Translator Acceptability of TM and SMT Outputs
Joss Moorkens, Andy Way
EAMT2
2016 Improving Phrase-Based SMT Using Cross-Granularity Embedding Similarity
Peyman Passban, Chris Hokamp, Andy Way, Qun Liu 0001
EAMT3
2016 Using SMT for OCR Error Correction of Historical Texts
Haithem Afli, Zhengwei Qiu, Andy Way, Paraic Sheridan
LREC3
2016 Using BabelNet to Improve OOV Coverage in SMT
Jinhua Du, Andy Way, Andrzej Zydron
LREC2
2016 Enhancing Access to Online Education: Quality Machine Translation of MOOC Content
Valia Kordoni, Antal van den Bosch, Katia Kermanidis, Vilelmini Sosoni, Kostadin Cholakov, Iris Hendrickx, Matthias Huck, Andy Way
LREC8
2016 Automatic Construction of Discourse Corpora for Dialogue Translation
Longyue Wang, Zhaopeng Tu, Andy Way, Qun Liu 0001
LREC4
2016 ProphetMT: A Tree-based SMT-driven Controlled Language Authoring/Post-Editing Tool
Jinhua Du, Qun Liu 0001, Andy Way
LREC4
2016 A Novel Approach to Dropped Pronoun Translation
abstract
Longyue Wang, Zhaopeng Tu, Xiaojun Zhang, Hang Li, Andy Way, Qun Liu. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016.
Longyue Wang, Zhaopeng Tu, Hang Li 0001, Andy Way, Qun Liu 0001
HLT-NAACL5
2016 Using Wordnet to Improve Reordering in Hierarchical Phrase-Based Statistical Machine Translation
abstract
We propose the use of WordNet synsets in a syntax-based reordering model for hierarchical statistical machine translation (HPB-SMT) to enable the model to generalize to phrases not seen in the training data but that have equivalent meaning.We detail our methodology to incorporate synsets' knowledge in the reordering model and evaluate the resulting WordNetenhanced SMT systems on the English-to-Farsi language direction.The inclusion of synsets leads to the best BLEU score, outperforming the baseline (standard HPB-SMT) by 0.6 points absolute.
Arefeh Kazemi, Antonio Toral, Andy Way
GWC3
2016 Combining translation memories and statistical machine translation using sparse features
Liangyou Li, Carla Parra Escartín, Andy Way, Qun Liu 0001
Mach. Transl.3
2016 Boosting Neural POS Tagger for Farsi Using Morphological Information
abstract
Farsi (Persian) is a low-resource language that suffers from the data sparsity problem and a lack of efficient processing tools. Due to their broad application in natural language processing tasks, part-of-speech (POS) taggers are one of those important tools that should be considered in this respect. Despite recent work on Farsi tagging, there is still room for improvement. The best reported accuracy so far is 96%, which in special cases can rise to 96.9%. The main problem with existing taggers is their inefficiency in coping with out-of-vocabulary (OOV) words. Addressing both problems of accuracy and OOV words, we developed a neural network-based POS tagger (NPT) that performs efficiently on Farsi. Despite using less data, NPT provides better results in comparison to state-of-the-art systems. Our proposed tagger performs with an accuracy of 97.4%, with performance highly influenced by morphological features. We carry out a shallow morphological analysis and show considerable improvement over the baseline configuration.
Peyman Passban, Qun Liu 0001, Andy Way
ACM Trans. Asian Low Resour. Lang. Inf. Process.3
2015 Dependency-based Reordering Model for Constituent Pairs in Hierarchical SMT
Arefeh Kazemi, Antonio Toral, Andy Way, S. Amirhassan Monadjemi, Mohammad Ali Nematbakhsh
EAMT3
2015 TraMOOC: Translation for Massive Open Online Courses
Valia Kordoni, Kostadin Cholakov, Markus Egg, Andy Way, Lexi Birch, Katia Kermanidis, Vilelmini Sosoni, Dimitrios Tsoumakos, Antal van den Bosch, Iris Hendrickx, Michael Papadopoulos, Panayota Georgakopoulou, Maria Gialama, Menno van Zaanen, Ioana Buliga, Mitja Jermol, Davor Orlic
EAMT4
2015 Benchmarking SMT Performance for Farsi Using the TEP++ Corpus
Peyman Passban, Andy Way, Qun Liu 0001
EAMT2
2015 Abu-MaTran: Automatic building of Machine Translation
Antonio Toral, Flammie A. Pirinen, Andy Way, Gema Ramírez-Sánchez, Sergio Ortiz-Rojas, Raphaël Rubino, Miquel Esplà-Gomis, Mikel L. Forcada, Vassilis Papavassiliou, Prokopis Prokopidis, Nikola Ljubesic
EAMT3
2015 Dependency Graph-to-String Translation
abstract
Compared to tree grammars, graph grammars have stronger generative capacity over structures.Based on an edge replacement grammar, in this paper we propose to use a synchronous graph-to-string grammar for statistical machine translation.The graph we use is directly converted from a dependency tree by labelling edges.We build our translation model in the log-linear framework with standard features.Large-scale experiments on Chinese-English and German-English tasks show that our model is significantly better than the state-of-the-art hierarchical phrase-based (HPB) model and a recently improved dependency tree-to-string model on BLEU, METEOR and TER scores.Experiments also suggest that our model has better capability to perform long-distance reordering and is more suitable for translating long sentences.
Liangyou Li, Andy Way, Qun Liu 0001
EMNLP2
2015 An empirical study of segment prioritization for incrementally retrained post-editing-based SMT
Jinhua Du, Ankit K. Srivastava, Andy Way, Alfredo Maldonado-Guerra
MTSummit3
2014 Standard language variety conversion for content localisation via SMT
Federico Fancellu, Andy Way, Morgan O'Brien
EAMT2
2014 Extrinsic evaluation of web-crawlers in machine translation: a study on Croatian-English for the tourism domain
Antonio Toral, Raphaël Rubino, Miquel Esplà-Gomis, Flammie A. Pirinen, Andy Way, Gema Ramírez-Sánchez
EAMT5
2013 Manual labour: tackling machine translation for sign languages
Sara Morrissey, Andy Way
Mach. Transl.2
2012 Translation Quality-Based Supplementary Data Selection by Incremental Update of Translation Models
Pratyush Banerjee, Sudip Kumar Naskar, Johann Roturier, Andy Way, Josef van Genabith
COLING4
2012 Extending CCG-based Syntactic Constraints in Hierarchical Phrase-Based SMT
Hala Almaghout, Jie Jiang 0002, Andy Way
EAMT3
2012 Domain Adaptation in SMT of User-Generated Forum Content Guided by OOV Word Reduction: Normalization and/or Supplementary Data
Pratyush Banerjee, Sudip Kumar Naskar, Johann Roturier, Andy Way, Josef van Genabith
EAMT4
2012 From Subtitles to Parallel Corpora
Mark Fishel, Panayota Georgakopoulou, Sergio Penkale, Volha Petukhova, Matej Rojc, Martin Volk 0001, Andy Way
EAMT7
2012 SUMAT: Data Collection and Parallel Corpus Compilation for Machine Translation of Subtitles
Volha Petukhova, Rodrigo Agerri, Mark Fishel, Sergio Penkale, Arantza del Pozo, Mirjam Sepesy Maucec, Andy Way, Panayota Georgakopoulou, Martin Volk 0001
LREC7
2012 Efficient accurate syntactic direct translation models: one tree at a time
Hany Hassan, Khalil Sima'an, Andy Way
Mach. Transl.3
2012 What types of word alignment improve statistical machine translation?
Patrik Lambert, Simon Petit-Renaud, Yanjun Ma, Andy Way
Mach. Transl.4
2012 David Bellos (ed): Is that a fish in your ear: translation and the meaning of everything - Particular Books, Penguin Group, London, 2011, ix + 390 pp, ISBN 978-1-846-14464-6
Andy Way
Mach. Transl.1
2011 Consistent Translation using Discriminative Learning - A Translation Memory-inspired Approach
Yanjun Ma, Yifan He 0007, Andy Way, Josef van Genabith
ACL3
2011 CCG Contextual labels in Hierarchical Phrase-Based SMT
Hala Almaghout, Jie Jiang 0002, Andy Way
EAMT3
2011 Experiments on Domain Adaptation for Patent Machine Translation in the PLuTO project
Alexandru Ceausu, John Tinsley, Jian Zhang 0003, Andy Way
EAMT4
2011 Using Example-Based MT to Support Statistical MT when Translating Homogeneous Data in a Resource-Poor Setting
Sandipan Dandapat, Sara Morrissey, Andy Way, Mikel L. Forcada
EAMT3
2011 Combining Semantic and Syntactic Generalization in Example-Based Machine Translation
Sarah Ebling, Andy Way, Martin Volk 0001, Sudip Kumar Naskar
EAMT2
2011 Towards Using Web-Crawled Data for Domain Adaptation in Statistical Machine Translation
Pavel Pecina, Antonio Toral, Andy Way, Vassilis Papavassiliou, Prokopis Prokopidis, Maria Giagkou
EAMT3
2011 Preliminary Experiments on Using Users' Post-Editions to Enhance a SMT System Oracle-based Training for Phrase-based Statistical Machine Translation
Ankit K. Srivastava, Yanjun Ma, Andy Way
EAMT3
2011 A Comparative Evaluation of Research vs. Online MT Systems
Antonio Toral, Federico Gaspari, Sudip Kumar Naskar, Andy Way
EAMT4
2011 Towards a User-Friendly Webservice Architecture for Statistical Machine Translation in the PANACEA project
Antonio Toral, Pavel Pecina, Marc Poch, Andy Way
EAMT4
2011 Domain Adaptation in Statistical Machine Translation of User-Forum Data using Component Level Mixture Modelling
Pratyush Banerjee, Sudip Kumar Naskar, Johann Roturier, Andy Way, Josef van Genabith
MTSummit4
2011 Rich Linguistic Features for Translation Memory-Inspired Consistent Translation
Yifan He 0007, Yanjun Ma, Andy Way, Josef van Genabith
MTSummit3
2011 Phonetic Representation-Based Speech Translation
Jie Jiang 0002, Julie Carson-Berndsen, Peter Cahill, Andy Way
MTSummit5
2011 A Framework for Diagnostic Evaluation of MT Based on Linguistic Checkpoints
Sudip Kumar Naskar, Antonio Toral, Federico Gaspari, Andy Way
MTSummit4
2011 Integrating source-language context into phrase-based statistical machine translation
Rejwanul Haque, Sudip Kumar Naskar, Antal van den Bosch, Andy Way
Mach. Transl.4
2011 Improved Chinese-English SMT with Chinese "DE" Construction Classification and Reordering
abstract
Syntactic reordering on the source side has been demonstrated to be helpful and effective for handling different word orders between source and target languages in SMT. In this article, we focus on the Chinese (DE) construction which is flexible and ubiquitous in Chinese and has many different ways to be translated into English so that it is a major source of word order differences in terms of translation quality. This article carries out the Chinese “DE” construction study for Chinese--English SMT in which we propose a new classifier model---discriminative latent variable model (DPLVM)---with new features to improve the classification accuracy and indirectly improve the translation quality compared to a log-linear classifier. The DE classifier is used to recognize DE structures in both training and test sentences of Chinese, and then perform word reordering to make the Chinese sentences better match the word order of English. In order to investigate the impact of the DE classification and reordering in the source side on different types of SMT systems (namely PB-SMT, hierarchical PB-SMT (HPB-SMT) as well as the syntax-based SMT (SAMT)), we conduct a series of experiments on NIST 2005 and 2008 test sets to verify the effectiveness of our proposed model. The experimental results show that the MT systems using the data reordered by our proposed model outperform the baseline systems by 3.01% and 4.03% relative points on the NIST 2005 test set, 4.64% and 4.62% relative points on the NIST 2008 test set in terms of BLEU score for PB-SMT and HPB-SMT respectively. However, the DE classification method does not perform significantly well for SAMT. Additionally, we also conducted some experiments to evaluate our DE classification and reordering approach on the word alignment and phrase table in terms of these three types of SMT systems.
Jinhua Du, Andy Way
ACM Trans. Asian Lang. Inf. Process.2
2010 Bridging SMT and TM with Translation Recommendation
Yifan He 0007, Yanjun Ma, Josef van Genabith, Andy Way
ACL4
2010 A Discriminative Latent Variable-Based "DE" Classifier for Chinese-English SMT
Jinhua Du, Andy Way
COLING2
2010 TMX Markup: A Challenge When Adapting SMT to the Localisation Environment
Jinhua Du, Johann Roturier, Andy Way
EAMT3
2010 The Impact of Source-Side Syntactic Reordering on Hierarchical Phrase-based SMT
Jinhua Du, Andy Way
EAMT2
2010 Lattice Score Based Data Cleaning for Phrase-Based Statistical Machine Translation
Jie Jiang 0002, Julie Carson-Berndsen, Andy Way
EAMT3
2010 Statistical Analysis of Alignment Characteristics for Phrase-based Machine Translation
Patrik Lambert, Simon Petit-Renaud, Yanjun Ma, Andy Way
EAMT4
2010 Facilitating Translation Using Source Language Paraphrase Lattices
Jinhua Du, Jie Jiang 0002, Andy Way
EMNLP3
2010 Metric and reference factors in minimum error rate training
Yifan He 0007, Andy Way
Mach. Transl.2
2010 Panning for EBMT gold, or "Remembering not to forget"
Andy Way
Mach. Transl.1
2009 Exploiting Parallel Treebanks to Improve Phrase-Based Statistical Machine Translation
John Tinsley, Mary Hearne, Andy Way
CICLing3
2009 Bilingually Motivated Domain-Adapted Word Segmentation for Statistical Machine Translation
Yanjun Ma, Andy Way
EACL2
2009 Using Supertags as Source Language Context in SMT
Rejwanul Haque, Sudip Kumar Naskar, Yanjun Ma, Andy Way
EAMT4
2009 Learning Labelled Dependencies in Machine Translation Evaluation
Yifan He 0007, Andy Way
EAMT2
2009 Tuning Syntactically Enhanced Word Alignment for Statistical Machine Translation
Yanjun Ma, Patrik Lambert, Andy Way
EAMT3
2009 Optimal Bilingual Data for French-English PB-SMT
Sylwia Ozdowska, Andy Way
EAMT2
2009 Marker-Based Filtering of Bilingual Phrase Pairs for SMT
Felipe Sánchez-Martínez, Andy Way
EAMT2
2009 Accuracy-Based Scoring for DOT: Towards Direct Error Minimization for Data-Oriented Translation
Daniel Galron, Sergio Penkale, Andy Way, I. Dan Melamed
EMNLP3
2009 A Syntactified Direct Translation Model with Linear-time Decoding
Hany Hassan, Khalil Sima'an, Andy Way
EMNLP3
2009 Using same-language machine translation to create alternative target sequences for text-to-speech synthesis
abstract
Modern speech synthesis systems attempt to produce\nspeech utterances from an open domain of words. In some situations, the synthesiser will not have the appropriate units to pronounce some words or phrases accurately but it still must attempt to pronounce them. This paper presents a hybrid machine translation and unit selection speech synthesis system. The machine translation system was trained with English as the source and target language. Rather than the synthesiser only saying the input text as would happen in conventional synthesis systems, the synthesiser may say an alternative utterance with the same\nmeaning. This method allows the synthesiser to overcome the\nproblem of insufficient units in runtime.
Peter Cahill, Jinhua Du, Andy Way, Julie Carson-Berndsen
INTERSPEECH3
2009 Capturing Lexical Variation in MT Evaluation Using Automatically Built Sense-Cluster Inventories
Marianna Apidianaki, Yifan He 0007, Andy Way
PACLIC3
2009 Dependency Relations as Source Context in Phrase-Based SMT
Rejwanul Haque, Sudip Kumar Naskar, Antal van den Bosch, Andy Way
PACLIC4
2009 Experiments on Domain Adaptation for English--Hindi SMT
Rejwanul Haque, Sudip Kumar Naskar, Josef van Genabith, Andy Way
PACLIC4
2009 Automatically generated parallel treebanks and their exploitability in machine translation
John Tinsley, Andy Way
Mach. Transl.2
2009 Bilingually Motivated Word Segmentation for Statistical Machine Translation
abstract
We introduce a bilingually motivated word segmentation approach to languages where word boundaries are not orthographically marked, with application to Phrase-Based Statistical Machine Translation (PB-SMT). Our approach is motivated from the insight that PB-SMT systems can be improved by optimizing the input representation to reduce the predictive power of translation models. We firstly present an approach to optimize the existing segmentation of both source and target languages for PB-SMT and demonstrate the effectiveness of this approach using a Chinese--English MT task, that is, to measure the influence of the segmentation on the performance of PB-SMT systems. We report a 5.44% relative increase in Bleu score and a consistent increase according to other metrics. We then generalize this method for Chinese word segmentation without relying on any segmenters and show that using our segmentation PB-SMT can achieve more consistent state-of-the-art performance across two domains. There are two main advantages of our approach. First of all, it is adapted to the specific translation task at hand by taking the corresponding source (target) language into account. Second, this approach does not rely on manually segmented training data so that it can be automatically adapted for different domains.
Yanjun Ma, Andy Way
ACM Trans. Asian Lang. Inf. Process.2
2008 Automatic Generation of Parallel Treebanks
Ventsislav Zhechev, Andy Way
COLING2
2008 The ATIS Sign Language Corpus
Jan Bungeroth, Daniel Stein, Philippe Dreuw, Hermann Ney, Sara Morrissey, Andy Way, Lynette van Zijl
LREC6
2008 A syntactic language model based on incremental CCG parsing
abstract
Syntactically-enriched language models (parsers) constitute a promising component in applications such as machine translation and speech-recognition. To maintain a useful level of accuracy, existing parsers are non-incremental and must span a combinatorially growing space of possible structures as every input word is processed. This prohibits their incorporation into standard linear-time decoders. In this paper, we present an incremental, linear-time dependency parser based on Combinatory Categorial Grammar (CCG) and classification techniques. We devise a deterministic transform of CCG-bank canonical derivations into incremental ones, and train our parser on this data. We discover that a cascaded, incremental version provides an appealing balance between efficiency and accuracy.
Hany Hassan, Khalil Sima'an, Andy Way
SLT3
2008 Wide-Coverage Deep Statistical Parsing Using Automatic Dependency Structure Annotation
abstract
A number of researchers have recently conducted experiments comparing “deep” hand-crafted wide-coverage with “shallow” treebank- and machine-learning-based parsers at the level of dependencies, using simple and automatic methods to convert tree output generated by the shallow parsers into dependencies. In this article, we revisit such experiments, this time using sophisticated automatic LFG f-structure annotation methodologies with surprising results. We compare various PCFG and history-based parsers to find a baseline parsing system that fits best into our automatic dependency structure annotation technique. This combined system of syntactic parser and dependency structure annotation is compared to two hand-crafted, deep constraint-based parsers, RASP and XLE. We evaluate using dependency-based gold standards and use the Approximate Randomization Test to test the statistical significance of the results. Our experiments show that machine-learning-based shallow grammars augmented with sophisticated automatic dependency annotation technology outperform hand-crafted, deep, wide-coverage constraint grammars. Currently our best system achieves an f-score of 82.73% against the PARC 700 Dependency Bank, a statistically significant improvement of 2.18% over the most recent results of 80.55% for the hand-crafted LFG grammar and XLE parsing system and an f-score of 80.23% against the CBS 500 Dependency Bank, a statistically significant 3.66% improvement over the 76.57% achieved by the hand-crafted RASP grammar and parsing system.
Aoife Cahill, Michael Burke, Ruth O'Donovan, Stefan Riezler, Josef van Genabith, Andy Way
Comput. Linguistics6
2008 Syntactically Lexicalized Phrase-Based SMT
abstract
Until quite recently, extending phrase-based statistical machine translation (PBSMT) with syntactic knowledge caused system performance to deteriorate. The most recent successful enrichments of PBSMT with hierarchical structure either employ nonlinguistically motivated syntax for capturing hierarchical reordering phenomena, or extend the phrase translation table with redundantly ambiguous syntactic structures over phrase pairs. In this paper, we present an extended, harmonized account of our previous work which showed that incorporating linguistically motivated lexical syntactic descriptions, calledsupertags, can yield significantly better PBSMT systems at insignificant extra computational cost. We describe a novel PBSMT model that integrates supertags into the target language model and the target side of the translation model. Two kinds of supertags are employed: those from lexicalized tree-adjoining grammar and combinatory categorial grammar. Despite the differences between the two sets of supertags, they give similar improvements. In addition to integrating the Markov supertagging approach in PBSMT, we explore the utility of a new surface grammaticality measure based on combinatory operators. We perform various experiments on the Arabic-to-English NIST 2005 test set addressing the issues of sparseness, scalability, and the utility of system subcomponents. We show that even when the parallel training data grows very large, the supertagged system retains a relatively stable absolute performance advantage over the unadorned PBSMT system. Arguably, this hints at a performance gap that cannot be bridged by acquiring more phrase pairs. Our best result shows a relative improvement of 6.1% over a state-of-the-art PBSMT model, which compares favorably with the leading systems on the NIST 2005 task. We also demonstrate that the advantages of a supertag-based system carry over to German-English, where improvements of up to 8.9% relative to the baseline system are observed.
Hany Hassan, Khalil Sima'an, Andy Way
IEEE Trans. Speech Audio Process.3
2007 Supertagged Phrase-Based Statistical Machine Translation
Hany Hassan, Khalil Sima'an, Andy Way
ACL3
2007 Bootstrapping Word Alignment via Word Packing
Yanjun Ma, Nicolas Stroppa, Andy Way
ACL3
2007 Comparing rule-based and data-driven approaches to Spanish-to-Basque machine translation
Gorka Labaka, Nicolas Stroppa, Andy Way, Kepa Sarasola
MTSummit3
2007 Combining data-driven MT systems for improved sign language translation
Sara Morrissey, Andy Way, Daniel Stein, Jan Bungeroth, Hermann Ney
MTSummit2
2007 Robust language pair-independent sub-tree alignment
John Tinsley, Ventsislav Zhechev, Mary Hearne, Andy Way
MTSummit4
2007 Evaluating machine translation with LFG dependencies
Karolina Owczarzak, Josef van Genabith, Andy Way
Mach. Transl.3
2006 Hybridity in MT. Experiments on the Europarl Corpus
Declan Groves, Andy Way
EAMT2
2006 Disambiguation Strategies for Data-Oriented Translation
Mary Hearne, Andy Way
EAMT2
2006 A Syntactic Skeleton for Statistical Machine Translation
Bart Mellebeek, Karolina Owczarzak, Declan Groves, Josef van Genabith, Andy Way
EAMT5
2006 Syntactic Phrase-Based Statistical Machine Translation
abstract
Phrase-based statistical machine translation (PBSMT) systems represent the dominant approach in MT today. However, unlike systems in other paradigms, it has proven difficult to date to incorporate syntactic knowledge in order to improve translation quality. This paper improves on recent research which uses 'syntactified' target language phrases, by incorporating supertags as constraints to better resolve parse tree fragments. In addition, we do not impose any sentence-length limit, and using a log-linear decoder, we outperform a state-of-the-art PBSMT system by over 1.3 BLEU points (or 3.51% relative) on the NIST 2003 Arabic-English test corpus.
Hany Hassan, Mary Hearne, Andy Way, Khalil Sima'an
SLT3
2005 TransBooster: boosting the performance of wide-coverage machine translation systems
Bart Mellebeek, Anna Khasin, Josef van Genabith, Andy Way
EAMT4
2005 Improving Online Machine Translation Systems
abstract
In (Mellebeek et al., 2005), we proposed the design, implementation and evaluation of a novel and modular approach to boost the translation performance of existing, wide-coverage, freely available machine translation systems, based on reliable and fast automatic decomposition of the translation input and corresponding composition of translation output. Despite showing some initial promise, our method did not improve on the baseline Logomedia1 and Systran2 MT systems. In this paper, we improve on the algorithm presented in (Mellebeek et al., 2005), and on the same test data, show increased scores for a range of automatic evaluation metrics. Our algorithm now outperforms Logomedia, obtains similar results to SDL3 and falls tantalisingly short of the performance achieved by Systran.
Bart Mellebeek, Anna Khasin, Karolina Owczarzak, Josef van Genabith, Andy Way
MTSummit5
2005 Large-Scale Induction and Evaluation of Lexical Resources from the Penn-II and Penn-III Treebanks
abstract
We present a methodology for extracting subcategorization frames based on an automatic lexical-functional grammar (LFG) f-structure annotation algorithm for the Penn-II and Penn-III Treebanks. We extract syntactic-function-based subcategorization frames (LFG semantic forms) and traditional CFG category-based subcategorization frames as well as mixed function/category-based frames, with or without preposition information for obliques and particle information for particle verbs. Our approach associates probabilities with frames conditional on the lemma, distinguishes between active and passive frames, and fully reflects the effects of long-distance dependencies in the source data structures. In contrast to many other approaches, ours does not predefine the subcategorization frame types extracted, learning them instead from the source data. Including particles and prepositions, we extract 21,005 lemma frame types for 4,362 verb lemmas, with a total of 577 frame types and an average of 4.8 frame types per verb. We present a large-scale evaluation of the complete set of forms extracted against the full COMLEX resource. To our knowledge, this is the largest and most complete evaluation of subcategorization frames acquired automatically for English.
Ruth O'Donovan, Michael Burke, Aoife Cahill, Josef van Genabith, Andy Way
Comput. Linguistics5
2005 Introduction to special issue on example-based machine translation
Michael Carl, Andy Way
Mach. Transl.2
2005 Hybrid data-driven models of machine translation
Declan Groves, Andy Way
Mach. Transl.2
2005 Controlled Translation in an Example-based Environment: What do Automatic Evaluation Metrics Tell Us?
Andy Way, Nano Gough
Mach. Transl.1
2005 Comparing example-based and statistical machine translation
abstract
In previous work (Gough and Way 2004), we showed that our Example-Based Machine Translation (EBMT) system improved with respect to both coverage and quality when seeded with increasing amounts of training data, so that it significantly outperformed the on-line MT system Logomedia according to a wide variety of automatic evaluation metrics. While it is perhaps unsurprising that system performance is correlated with the amount of training data, we address in this paper the question of whether a large-scale, robust EBMT system such as ours can outperform a Statistical Machine Translation (SMT) system. We obtained a large English-French translation memory from Sun Microsystems from which we randomly extracted a near 4K test set. The remaining data was split into three training sets, of roughly 50K, 100K and 200K sentence-pairs in order to measure the effect of increasing the size of the training data on the performance of the two systems. Our main observation is that contrary to perceived wisdom in the field, there appears to be little substance to the claim that SMT systems are guaranteed to outperform EBMT systems when confronted with ‘enough’ training data. Our tests on a 4.8 million word bitext indicate that while SMT appears to outperform our system for French-English on a number of metrics, for English-French, on all but one automatic evaluation metric, the performance of our EBMT system is superior to the baseline SMT model.
Andy Way, Nano Gough
Nat. Lang. Eng.1
2004 Long-Distance Dependency Resolution in Automatically Acquired Wide-Coverage PCFG-Based LFG Approximations
abstract
This paper shows how finite approximations of long distance dependency (LDD) resolution can be obtained automatically for wide-coverage, robust, probabilistic Lexical-Functional Grammar (LFG) resources acquired from treebanks. We extract LFG subcategorisation frames and paths linking LDD reentrancies from f-structures generated automatically for the Penn-II treebank trees and use them in an LDD resolution algorithm to parse new text. Unlike (Collins, 1999; Johnson, 2000), in our approach resolution of LDDs is done at f-structure (attribute-value structure representations of basic predicate-argument or dependency structure) without empty productions, traces and coindexation in CFG parse trees. Currently our best automatically induced grammars achieve 80.97% f-score for f-structures parsing section 23 of the WSJ part of the Penn-II treebank and evaluating against the DCU 1051 and 80.24% against the PARC 700 Dependency Bank (King et al., 2003), performing at the same or a slightly better level than state-of-the-art hand-crafted grammars (Kaplan et al., 2004).
Aoife Cahill, Michael Burke, Ruth O'Donovan, Josef van Genabith, Andy Way
ACL5
2004 Large-Scale Induction and Evaluation of Lexical Resources from the Penn-II Treebank
abstract
In this paper we present a methodology for extracting subcategorisation frames based on an automatic LFG f-structure annotation algorithm for the Penn-II Treebank. We extract abstract syntactic function-based subcategorisation frames (LFG semantic forms), traditional CFG category-based subcategorisation frames as well as mixed function/category-based frames, with or without preposition information for obliques and particle information for particle verbs. Our approach does not predefine frames, associates probabilities with frames conditional on the lemma, distinguishes between active and passive frames, and fully reflects the effects of long-distance dependencies in the source data structures. We extract 3586 verb lemmas, 14348 semantic form types (an average of 4 per lemma) with 577 frame types. We present a large-scale evaluation of the complete set of forms extracted against the full COMLEX resource.
Ruth O'Donovan, Michael Burke, Aoife Cahill, Josef van Genabith, Andy Way
ACL5
2004 Robust Sub-Sentential Alignment of Phrase-Structure Trees
Declan Groves, Mary Hearne, Andy Way
COLING3
2004 Treebank-Based Acquisition of a Chinese Lexical-Functional Grammar
Michael Burke, Olivia S.-C. Lam, Aoife Cahill, Rowena Chan, Ruth O'Donovan, Adams Bodomo, Josef van Genabith, Andy Way
PACLIC8
2003 Controlled generation in example-based machine translation
abstract
The theme of controlled translation is currently in vogue in the area of MT. Recent research (Scha ̈ler et al., 2003; Carl, 2003) hypothesises that EBMT systems are perhaps best suited to this challenging task. In this paper, we present an EBMT system where the generation of the target string is filtered by data written according to controlled language specifications. As far as we are aware, this is the only research available on this topic. In the field of controlled language applications, it is more usual to constrain the source language in this way rather than the target. We translate a small corpus of controlled English into French using the on-line MT system Logomedia, and seed the memories of our EBMT system with a set of automatically induced lexical resources using the Marker Hypothesis as a segmentation tool. We test our system on a large set of sentences extracted from a Sun Translation Memory, and provide both an automatic and a human evaluation. For comparative purposes, we also provide results for Logomedia itself.
Nano Gough, Andy Way
MTSummit2
2003 Seeing the wood for the trees: data-oriented translation
abstract
Data-Oriented Translation (DOT), which is based on Data-Oriented Parsing (DOP), comprises an experience-based approach to translation, where new translations are derived with reference to grammatical analyses of previous translations. Previous DOT experiments [Poutsma, 1998, Poutsma, 2000a, Poutsma, 2000b] were small in scale because important advances in DOP technology were not incorporated into the translation model. Despite this, related work [Way, 1999, Way, 2003a, Way, 2003b] reports that DOT models are viable in that solutions to ‘hard’ translation cases are readily available. However, it has not been shown to date that DOT models scale to larger datasets. In this work, we describe a novel DOT system, inspired by recent advances in DOP parsing technology. We test our system on larger, more complex corpora than have been used heretofore, and present both automatic and human evaluations which show that high quality translations can be achieved at reasonable speeds.
Mary Hearne, Andy Way
MTSummit2
2003 wEBMT: Developing and Validating an Example-Based Machine Translation System using the World Wide Web
abstract
We have developed an example-based machine translation (EBMT) system that uses the World Wide Web for two different purposes: First, we populate the system's memory with translations gathered from rule-based MT systems located on the Web. The source strings input to these systems were extracted automatically from an extremely small subset of the rule types in the Penn-II Treebank. In subsequent stages, the source, target translation pairs obtained are automatically transformed into a series of resources that render the translation process more successful. Despite the fact that the output from on-line MT systems is often faulty, we demonstrate in a number of experiments that when used to seed the memories of an EBMT system, they can in fact prove useful in generating translations of high quality in a robust fashion. In addition, we demonstrate the relative gain of EBMT in comparison to on-line systems. Second, despite the perception that the documents available on the Web are of questionable quality, we demonstrate in contrast that such resources are extremely useful in automatically postediting translation candidates proposed by our system.
Andy Way, Nano Gough
Comput. Linguistics1
1999 A hybrid architecture for robust MT using LFG-DOP
Andy Way
J. Exp. Theor. Artif. Intell.1