EDBT 2026 Demo / reviewers in the wild / expert
Andy Way
dblp:69/5430
· DBLP profile ↗
172ranked-venue papers
10as first author
17since 2021 · last 2024
0000-0001-5736-5930ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 170 · 10 first-author · 17 since 2021Databases, data management, data science and information retrieval · 5Graphics, computer vision, multimedia, augmented reality and games · 3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | SignON - a Co-creative Machine Translation for Sign and Spoken Languages (end-of-project results, contributions and lessons learned)abstractSignON, a 3-year Horizon 20202 project addressing the lack of technology and services for MT between sign languages (SLs) and spoken languages (SpLs) ended in December 2023. SignON was unprecedented. Not only it addressed the wider complexity of the aforementioned problem – from research and development of recognition, translation and synthesis, through development of easy-to-use mobile applications and a cloud-based framework to do the “heavy lifting” as well as to establishing ethical, privacy and inclusivenesspolicies and operation guidelines – but also engaged with the deaf and hard of hearing communities in an effective co-creation approach where these main stakeholders drove the development in the right direction and had the final say.Currently we are witnessing advances in natural language processing for SLs, including MT. SignON was one of the largest projects that contributed to this surge with 17 partners and more than 60 consortium members, working in parallel with other international and European initiatives, such as project EASIER and others. Dimitar Sht. Shterionov, Vincent Vandeghinste, Mirella De Sisto, Aoife Brady, Mathieu De Coster, Lorraine Leeson, Andy Way, Josep Blat, Frankie Picron, Davy Van Landuyt, Marcello Paolo Scipioni, Aditya Parikh, Louis ten Bosch, John J. O'Flaherty, Joni Dambre, Caro Brosens, Jorn Rijckaert, Víctor Ubieto Nogales, Bram Vanroy, Santiago Egea Gómez, Ineke Schuurman, Gorka Labaka, Adrián Núñez-Marcos, Irene Murtagh, Euan McGill, Horacio Saggion |
EAMT (2) | 7 |
| 2023 | Adaptive Machine Translation with Large Language ModelsabstractConsistency is a key requirement of high-quality translation. It is especially important to adhere to pre-approved terminology and adapt to corrected translations in domain-specific projects. Machine translation (MT) has achieved significant progress in the area of domain adaptation. However, real-time adaptation remains challenging. Large-scale language models (LLMs) have recently shown interesting capabilities of in-context learning, where they learn to replicate certain input-output text generation patterns, without further fine-tuning. By feeding an LLM at inference time with a prompt that consists of a list of translation pairs, it can then simulate the domain and style characteristics. This work aims to investigate how we can utilize in-context learning to improve real-time adaptive MT. Our extensive experiments show promising results at translation time. For example, GPT-3.5 can adapt to a set of in-domain sentence pairs and/or terminology while translating a new sentence. We observe that the translation quality with few-shot in-context learning can surpass that of strong encoder-decoder MT systems, especially for high-resource languages. Moreover, we investigate whether we can combine MT from strong encoder-decoder models with fuzzy matches, which can further improve translation quality, especially for less supported languages. We conduct our experiments across five diverse language pairs, namely English-to-Arabic (EN-AR), English-to-Chinese (EN-ZH), English-to-French (EN-FR), English-to-Kinyarwanda (EN-RW), and English-to-Spanish (EN-ES). Yasmin Moslem, Rejwanul Haque, John D. Kelleher, Andy Way |
EAMT | 4 |
| 2023 | Instance-Based Domain Adaptation for Improving Terminology TranslationabstractTerms are essential indicators of a domain, and domain term translation is dealt with priority in any translation workflow. Translation service providers who use machine translation (MT) expect term translation to be unambiguous and consistent with the context and domain in question. Although current state-of-the-art neural MT (NMT) models are able to produce high-quality translations for many languages, they are still not at the level required when it comes to translating domain-specific terms. This study presents a terminology-aware instance- based adaptation method for improving terminology translation in NMT. We conducted our experiments for French-to-English and found that our proposed approach achieves a statistically significant improvement over the baseline NMT system in translating domain-specific terms. Specifically, the translation of multi-word terms is improved by 6.7% compared to the strong baseline. Prashanth Nayak, John D. Kelleher, Rejwanul Haque, Andy Way |
MTSummit (1) | 4 |
| 2022 | Overview of the ELE ProjectabstractThis paper provides an overview of the ongoing European Language Equality(ELE) project, an 18-month action funded by the European Commission which involves 52 partners. The primary goal of ELE is to prepare the European Language Equality Programme, in the form of a strategic research, innovation and implementation agenda and a roadmap for achieving full digital language equality (DLE) in Europe by 2030. Itziar Aldabe, Jane Dunne, Aritz Farwell, Owen Gallagher, Federico Gaspari, Maria Giagkou, Jan Hajic 0001, Jens Peter Kückens, Teresa Lynn, Georg Rehm, German Rigau, Katrin Marheinecke, Stelios Piperidis, Natália Resende, Tereza Vojtechová, Andy Way |
EAMT | 16 |
| 2022 | Achievements of the PRINCIPLE Project: Promoting MT for Croatian, Icelandic, Irish and NorwegianabstractThis paper provides an overview of the main achievements of the completed PRINCIPLE project, a 2-year action funded by the European Commission under the Connecting Europe Facility (CEF) programme. PRINCIPLE focused on collecting high-quality language resources for Croatian, Icelandic, Irish and Norwegian, which are severely low-resource languages, especially for building effective machine translation (MT) systems. We report the achievements of the project, primarily, in terms of the large amounts of data collected for all four low-resource languages and of promoting the uptake of neural MT (NMT) for these languages. Petra Bago, Sheila Castilho, Jane Dunne, Federico Gaspari, Andre Kåsen, Gauti Kristmannsson, Jon Arild Olsen, Natália Resende, Níels Rúnar Gíslason, Dana Davis Sheridan, Paraic Sheridan, John Tinsley, Andy Way |
EAMT | 13 |
| 2022 | Developing Machine Translation Engines for Multilingual Participatory SpacesabstractIt is often a challenging task to build Machine Translation (MT) engines for a specific domain due to the lack of parallel data in that area. In this project, we develop a range of MT systems for 6 European languages (English, German, Italian, French, Polish and Irish) in all directions and in two domains (environment and economics). Pintu Lohar, Guodong Xie, Andy Way |
EAMT | 3 |
| 2022 | gaHealth: An English-Irish Bilingual Corpus of Health DataabstractMachine Translation is a mature technology for many high-resource language pairs. However in the context of low-resource languages, there is a paucity of parallel data datasets available for developing translation models. Furthermore, the development of datasets for low-resource languages often focuses on simply creating the largest possible dataset for generic translation. The benefits and development of smaller in-domain datasets can easily be overlooked. To assess the merits of using in-domain data, a dataset for the specific domain of health was developed for the low-resource English to Irish language pair. Our study outlines the process used in developing the corpus and empirically demonstrates the benefits of using an in-domain dataset for the health domain. In the context of translating health-related data, models developed using the gaHealth corpus demonstrated a maximum BLEU score improvement of 22.2 points (40%) when compared with top performing models from the LoResMT2021 Shared Task. Furthermore, we define linguistic guidelines for developing gaHealth, the first bilingual corpus of health data for the Irish language, which we hope will be of use to other creators of low-resource data sets. gaHealth is now freely available online and is ready to be explored for further research. Séamus Lankford, Haithem Afli, Orla Ni Loinsigh, Andy Way |
LREC | 4 |
| 2022 | Improved feature decay algorithms for statistical machine translationabstractAbstract In machine-learning applications, data selection is of crucial importance if good runtime performance is to be achieved. In a scenario where the test set is accessible when the model is being built, training instances can be selected so they are the most relevant for the test set. Feature Decay Algorithms (FDA) are a technique for data selection that has demonstrated excellent performance in a number of tasks. This method maximizes the diversity of the n-grams in the training set by devaluing those ones that have already been included. We focus on this method to undertake deeper research on how to select better training data instances. We give an overview of FDA and propose improvements in terms of speed and quality. Using German-to-English parallel data, first we create a novel approach that decreases the execution time of FDA when multiple computation units are available. In addition, we obtain improvements on translation quality by extending FDA using information from the parallel corpus that is generally ignored. Alberto Poncelas, Gideon Maillette de Buy Wenniger, Andy Way |
Nat. Lang. Eng. | 3 |
| 2021 | Augmenting Training Data for Low-Resource Neural Machine Translation via Bilingual Word Embeddings and BERT Language Modelling
Akshai Ramesh, Haque Usuf Uhana, Venkatesh Balavadhani Parthasarathy, Rejwanul Haque, Andy Way |
IJCNN | 5 |
| 2021 | Transformers for Low-Resource Languages: Is Féidir Linn!abstractThe Transformer model is the state-of-the-art in Machine Translation. However and in general and neural translation models often under perform on language pairs with insufficient training data. As a consequence and relatively few experiments have been carried out using this architecture on low-resource language pairs. In this study and hyperparameter optimization of Transformer models in translating the low-resource English-Irish language pair is evaluated. We demonstrate that choosing appropriate parameters leads to considerable performance improvements. Most importantly and the correct choice of subword model is shown to be the biggest driver of translation performance. SentencePiece models using both unigram and BPE approaches were appraised. Variations on model architectures included modifying the number of layers and testing various regularization techniques and evaluating the optimal number of heads for attention. A generic 55k DGT corpus and an in-domain 88k public admin corpus were used for evaluation. A Transformer optimized model demonstrated a BLEU score improvement of 7.8 points when compared with a baseline RNN model. Improvements were observed across a range of metrics and including TER and indicating a substantially reduced post editing effort for Transformer optimized models with 16k BPE subword models. Bench-marked against Google Translate and our translation engines demonstrated significant improvements. The question of whether or not Transformers can be used effectively in a low-resource setting of English-Irish translation has been addressed. Is féidir linn - yes we can. Séamus Lankford, Haithem Alfi, Andy Way |
MTSummit (1) | 3 |
| 2021 | A review of the state-of-the-art in automatic post-editingabstractThis article presents a review of the evolution of automatic post-editing, a term that describes methods to improve the output of machine translation systems, based on knowledge extracted from datasets that include post-edited content. The article describes the specificity of automatic post-editing in comparison with other tasks in machine translation, and it discusses how it may function as a complement to them. Particular detail is given in the article to the five-year period that covers the shared tasks presented in WMT conferences (2015-2019). In this period, discussion of automatic post-editing evolved from the definition of its main parameters to an announced demise, associated with the difficulties in improving output obtained by neural methods, which was then followed by renewed interest. The article debates the role and relevance of automatic post-editing, both as an academic endeavour and as a useful application in commercial workflows. Félix do Carmo, Dimitar Sht. Shterionov, Joss Moorkens, Joachim Wagner 0001, Murhaf Hossari, Eric Paquin, Dag Schmidtke, Declan Groves, Andy Way |
Mach. Transl. | 9 |
| 2021 | Augmenting training data with syntactic phrasal-segments in low-resource neural machine translation
Kamal Kumar Gupta, Sukanta Sen, Rejwanul Haque, Asif Ekbal, Pushpak Bhattacharyya, Andy Way |
Mach. Transl. | 6 |
| 2021 | Recent advances of low-resource neural machine translation
Rejwanul Haque, Chao-Hong Liu, Andy Way |
Mach. Transl. | 3 |
| 2021 | Philipp Koehn: Neural Machine TranslationabstractAbstract Neural machine translation (NMT) is an approach to machine translation (MT) that uses deep learning techniques, a broad area of machine learning based on deep artificial neural networks (NNs). The book Neural Machine Translation by Philipp Koehn targets a broad range of readers including researchers, scientists, academics, advanced undergraduate or postgraduate students, and users of MT, covering wider topics including fundamental and advanced neural network-based learning techniques and methodologies used to develop NMT systems. The book demonstrates different linguistic and computational aspects in terms of NMT with the latest practices and standards and investigates problems relating to NMT. Having read this book, the reader should be able to formulate, design, implement, critically assess and evaluate some of the fundamental and advanced deep learning techniques and methods used for MT. Koehn himself notes that he was somewhat overtaken by events, as originally this book was envisaged only as a chapter in a revised, extended version of his 2009 book Statistical Machine Translation . However, in the interim, NMT completely overtook this previously dominant paradigm, and this new book is likely to serve as the reference of note for the field for some time to come, despite the fact that new techniques are coming onstream all the time. Wandri Jooste, Rejwanul Haque, Andy Way |
Mach. Transl. | 3 |
| 2021 | From MT to LREV: managing the transition
Andy Way |
Mach. Transl. | 1 |
| 2021 | Neural machine translation of low-resource languages using SMT phrase pair injectionabstractAbstract Neural machine translation (NMT) has recently shown promising results on publicly available benchmark datasets and is being rapidly adopted in various production systems. However, it requires high-quality large-scale parallel corpus, and it is not always possible to have sufficiently large corpus as it requires time, money, and professionals. Hence, many existing large-scale parallel corpus are limited to the specific languages and domains. In this paper, we propose an effective approach to improve an NMT system in low-resource scenario without using any additional data. Our approach aims at augmenting the original training data by means of parallel phrases extracted from the original training data itself using a statistical machine translation (SMT) system. Our proposed approach is based on the gated recurrent unit (GRU) and transformer networks. We choose the Hindi–English, Hindi–Bengali datasets for Health, Tourism, and Judicial (only for Hindi–English) domains. We train our NMT models for 10 translation directions, each using only 5–23k parallel sentences. Experiments show the improvements in the range of 1.38–15.36 BiLingual Evaluation Understudy points over the baseline systems. Experiments show that transformer models perform better than GRU models in low-resource scenarios. In addition to that, we also find that our proposed method outperforms SMT—which is known to work better than the neural models in low-resource scenarios—for some translation directions. In order to further show the effectiveness of our proposed model, we also employ our approach to another interesting NMT task, for example, old-to-modern English translation, using a tiny parallel corpus of only 2.7K sentences. For this task, we use publicly available old-modern English text which is approximately 1000 years old. Evaluation for this task shows significant improvement over the baseline NMT. Sukanta Sen, Mohammed Hasanuzzaman, Asif Ekbal, Pushpak Bhattacharyya, Andy Way |
Nat. Lang. Eng. | 5 |
| 2021 | Reinforced NMT for Sentiment and Content Preservation in Low-resource ScenarioabstractThe preservation of domain knowledge from source to the target is crucial in any translation workflows. Hence, translation service providers that use machine translation (MT) in production could reasonably expect that the translation process should transfer both the underlying pragmatics and the semantics of the source-side sentences into the target language. However, recent studies suggest that the MT systems often fail to preserve such crucial information (e.g., sentiment, emotion, gender traits) embedded in the source text in the target. In this context, the raw automatic translations are often directly fed to other natural language processing (NLP) applications (e.g., sentiment classifier) in a cross-lingual platform. Hence, the loss of such crucial information during the translation could negatively affect the performance of such downstream NLP tasks that heavily rely on the output of the MT systems. In our current research, we carefully balance both the sides (i.e., sentiment and semantics) during translation, by controlling a global-attention-based neural MT (NMT), to generate translations that encode the underlying sentiment of a source sentence while preserving its non-opinionated semantic content. Toward this, we use a state-of-the-art reinforcement learning method, namely, actor-critic , that includes a novel reward combination module, to fine-tune the NMT system so that it learns to generate translations that are best suited for a downstream task, viz. sentiment classification while ensuring the source-side semantics is intact in the process. Experimental results for Hindi–English language pair show that our proposed method significantly improves the performance of the sentiment classifier and alongside results in an improved NMT system. Divya Kumari, Asif Ekbal, Rejwanul Haque, Pushpak Bhattacharyya, Andy Way |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 5 |
| 2020 | Selecting Backtranslated Data from Multiple Sources for Improved Neural Machine TranslationabstractMachine translation (MT) has benefited from using synthetic training data originating from translating monolingual corpora, a technique known as backtranslation.Combining backtranslated data from different sources has led to better results than when using such data in isolation.In this work we analyse the impact that data translated with rule-based, phrasebased statistical and neural MT systems has on new MT systems.We use a real-world low-resource use-case (Basque-to-Spanish in the clinical domain) as well as a high-resource language pair (German-to-English) to test different scenarios with backtranslation and employ data selection to optimise the synthetic corpora.We exploit different data selection strategies in order to reduce the amount of data used, while at the same time maintaining highquality MT systems.We further tune the data selection method by taking into account the quality of the MT systems used for backtranslation and lexical diversity of the resulting corpora.Our experiments show that incorporating backtranslated data from different sources can be beneficial, and that availing of data selection can yield improved performance. Xabier Soto, Dimitar Sht. Shterionov, Alberto Poncelas, Andy Way |
ACL | 4 |
| 2020 | A human evaluation of English-Irish statistical and neural machine translationabstractWith official status in both Ireland and the EU, there is a need for high-quality English-Irish (EN-GA) machine translation (MT) systems which are suitable for use in a professional translation environment. While we have seen recent research on improving both statistical MT and neural MT for the EN-GA pair, the results of such systems have always been reported using automatic evaluation metrics. This paper provides the first human evaluation study of EN-GA MT using professional translators and in-domain (public administration) data for a more accurate depiction of the translation quality available via MT. Meghan Dowling, Sheila Castilho, Joss Moorkens, Teresa Lynn, Andy Way |
EAMT | 5 |
| 2020 | Modelling Source- and Target- Language Syntactic Information as Conditional Context in Interactive Neural Machine TranslationabstractIn interactive machine translation (MT), human translators correct errors in automatic translations in collaboration with the MT systems, which is seen as an effective way to improve the productivity gain in translation. In this study, we model source-language syntactic constituency parse and target-language syntactic descriptions in the form of supertags as conditional context for interactive prediction in neural MT (NMT). We found that the supertags significantly improve productivity gain in translation in interactive-predictive NMT (INMT), while syntactic parsing somewhat found to be effective in reducing human effort in translation. Furthermore, when we model this source- and target-language syntactic information together as the conditional context, both types complement each other and our fully syntax-informed INMT model statistically significantly reduces human efforts in a French–to–English translation task, achieving 4.30 points absolute (corresponding to 9.18% relative) improvement in terms of word prediction accuracy (WPA) and 4.84 points absolute (corresponding to 9.01% relative) reduction in terms of word stroke ratio (WSR) over the baseline. Kamal Kumar Gupta, Rejwanul Haque, Asif Ekbal, Pushpak Bhattacharyya, Andy Way |
EAMT | 5 |
| 2020 | MT syntactic priming effects on L2 English speakersabstractIn this paper, we tested 20 Brazilian Portuguese speakers at intermediate and advanced English proficiency levels to investigate the influence of Google Translate’s MT system on the mental processing of English as a second language. To this end, we employed a syntactic priming experimental paradigm using a pretest-priming design which allowed us to compare participants’ linguistic behaviour before and after a translation task using Google Translate. Results show that, after performing a translation task with Google Translate, participants more frequently described images in English using the syntactic alternative previously seen in the output of Google Translate, compared to the translation task with no prior influence of the MT output. Results also show that this syntactic priming effect is modulated by English proficiency levels. Natália Resende, Benjamin R. Cowan, Andy Way |
EAMT | 3 |
| 2020 | MTrill project: Machine Translation impact on language learningabstractOver the last decades, massive research investments have been made in the development of machine translation (MT) systems (Gupta and Dhawan, 2019). This has brought about a paradigm shift in the performance of these language tools, leading to widespread use of popular MT systems (Gaspari and Hutchins, 2007). Although the first MT engines were used for gisting purposes, in recent years, there has been an increasing interest in using MT tools, especially the freely available online MT tools, for language teaching and learning (Clifford et al., 2013). The literature on MT and Computer Assisted Language Learning (CALL) shows that, over the years, MT systems have been facilitating language teaching and also language learning (Nin ̃o, 2006). It has been shown that MT tools can increase awareness of grammatical linguistic features of a foreign language. Research also shows the positive role of MT systems in the development of writing skills in English as well as in improving communication skills in English(Garcia and Pena, 2011). However, to date, the cognitive impact of MT on language acquisition and on the syntactic aspects of language processing has not yet been investigated and deserves further scrutiny. The MTril project aims at filling this gap in the literature by examining whether MT is contributing to a central aspect of language acquisition: the so-called language binding, i.e., the ability to combine single words properly in a grammatical sentence (Heyselaar et al., 2017; Ferreira and Bock, 2006). The project focus on the initial stages (pre-intermediate and intermediate) of the acquisition of English syntax by Brazilian Portuguese native speakers using MT systems as a support for language learning. Natália Resende, Andy Way |
EAMT | 2 |
| 2020 | Progress of the PRINCIPLE Project: Promoting MT for Croatian, Icelandic, Irish and NorwegianabstractThis paper updates the progress made on the PRINCIPLE project, a 2-year action funded by the European Commission under the Connecting Europe Facility (CEF) programme. PRINCIPLE focuses on collecting high-quality language resources for Croatian, Icelandic, Irish and Norwegian, which have been identified as low-resource languages, especially for building effective machine translation (MT) systems. We report initial achievements of the project and ongoing activities aimed at promoting the uptake of neural MT for the low-resource languages of the project. Andy Way, Petra Bago, Jane Dunne, Federico Gaspari, Andre Kåsen, Gauti Kristmannsson, Helen McHugh, Jon Arild Olsen, Dana Davis Sheridan, Paraic Sheridan, John Tinsley |
EAMT | 1 |
| 2020 | Syntax-Informed Interactive Neural Machine TranslationabstractIn interactive machine translation (MT), human translators correct errors in automatic translations in collaboration with the MT systems, and this is an effective way to improve productivity gain in translation. Phrase-based statistical MT (PB-SMT) has been the mainstream approach to MT for the past 30 years, both in academia and industry. Neural MT (NMT), an end-to-end learning approach to MT, represents the current state-of-the-art in MT research. The recent studies on interactive MT have indicated that NMT can significantly outperform PB-SMT. In this work, first we investigate the possibility of integrating lexical syntactic descriptions in the form of supertags into the state-of-the-art NMT model, Transformer. Then, we explore whether integration of supertags into Transformer could indeed reduce human efforts in translation in an interactive-predictive platform. From our investigation we found that our syntax-aware interactive NMT (INMT) framework significantly reduces simulated human efforts in the French-to-English and Hindi- to-English translation tasks, achieving a 2.65 point absolute corresponding to 5.65% relative improvement and a 6.55 point absolute corresponding to 19.1% relative improvement, respectively, in terms of word prediction accuracy (WPA) over the respective baselines. Kamal Kumar Gupta, Rejwanul Haque, Asif Ekbal, Pushpak Bhattacharyya, Andy Way |
IJCNN | 5 |
| 2020 | On Context Span Needed for Machine Translation EvaluationabstractDespite increasing efforts to improve evaluation of machine translation (MT) by going beyond the sentence level to the document level, the definition of what exactly constitutes a “document level” is still not clear. This work deals with the context span necessary for a more reliable MT evaluation. We report results from a series of surveys involving three domains and 18 target languages designed to identify the necessary context span as well as issues related to it. Our findings indicate that, despite the fact that some issues and spans are strongly dependent on domain and on the target language, a number of common patterns can be observed so that general guidelines for context-aware MT evaluation can be drawn. Sheila Castilho, Maja Popovic, Andy Way |
LREC | 3 |
| 2020 | The European Language Technology Landscape in 2020: Language-Centric and Human-Centric AI for Cross-Cultural Communication in Multilingual EuropeabstractMultilingualism is a cultural cornerstone of Europe and firmly anchored in the European treaties including full language equality. However, language barriers impacting business, cross-lingual and cross-cultural communication are still omnipresent. Language Technologies (LTs) are a powerful means to break down these barriers. While the last decade has seen various initiatives that created a multitude of approaches and technologies tailored to Europe’s specific needs, there is still an immense level of fragmentation. At the same time, AI has become an increasingly important concept in the European Information and Communication Technology area. For a few years now, AI – including many opportunities, synergies but also misconceptions – has been overshadowing every other topic. We present an overview of the European LT landscape, describing funding programmes, activities, actions and challenges in the different countries with regard to LT, including the current state of play in industry and the LT market. We present a brief overview of the main LT-related activities on the EU level in the last ten years and develop strategic guidance with regard to four key dimensions. Georg Rehm, Katrin Marheinecke, Stefanie Hegele, Stelios Piperidis, Kalina Bontcheva, Jan Hajic 0001, Khalid Choukri, Andrejs Vasiljevs, Gerhard Backfried, Christoph Prinz, José Manuél Gómez-Pérez, Luc Meertens, Paul Lukowicz, Josef van Genabith, Andrea Lösch, Philipp Slusallek, Morten Irgens, Patrick Gatellier, Joachim Köhler, Laure Le Bars, Dimitra Anastasiou, Albina Auksoriute, Núria Bel, António Branco, Gerhard Budin, Walter Daelemans, Koenraad De Smedt, Radovan Garabík, Maria Gavrilidou, Dagmar Gromann, Svetla Koeva, Simon Krek, Cvetana Krstev, Krister Lindén, Bernardo Magnini, Jan Odijk, Maciej Ogrodniczuk, Eiríkur Rögnvaldsson, Mike Rosner, Bolette S. Pedersen, Inguna Skadina, Marko Tadic, Dan Tufis, Tamás Váradi, Kadri Vider, Andy Way, François Yvon |
LREC | 46 |
| 2020 | Investigating Query Expansion and Coreference Resolution in Question Answering on BERT
Santanu Bhattacharjee, Rejwanul Haque, Gideon Maillette de Buy Wenniger, Andy Way |
NLDB | 4 |
| 2020 | Analysing terminology translation errors in statistical and neural machine translation
Rejwanul Haque, Mohammed Hasanuzzaman, Andy Way |
Mach. Transl. | 3 |
| 2020 | A roadmap to neural automatic post-editing: an empirical approachabstractIn a translation workflow, machine translation (MT) is almost always followed by a human post-editing step, where the raw MT output is corrected to meet required quality standards. To reduce the number of errors human translators need to correct, automatic post-editing (APE) methods have been developed and deployed in such workflows. With the advances in deep learning, neural APE (NPE) systems have outranked more traditional, statistical, ones. However, the plethora of options, variables and settings, as well as the relation between NPE performance and train/test data makes it difficult to select the most suitable approach for a given use case. In this article, we systematically analyse these different parameters with respect to NPE performance. We build an NPE "roadmap" to trace the different decision points and train a set of systems selecting different options through the roadmap. We also propose a novel approach for APE with data augmentation. We then analyse the performance of 15 of these systems and identify the best ones. In fact, the best systems are the ones that follow the newly-proposed method. The work presented in this article follows from a collaborative project between Microsoft and the ADAPT centre. The data provided by Microsoft originates from phrase-based statistical MT (PBSMT) systems employed in production. All tested NPE systems significantly increase the translation quality, proving the effectiveness of neural post-editing in the context of a commercial translation workflow that leverages PBSMT. Dimitar Sht. Shterionov, Félix do Carmo, Joss Moorkens, Murhaf Hossari, Joachim Wagner 0001, Eric Paquin, Dag Schmidtke, Declan Groves, Andy Way |
Mach. Transl. | 9 |
| 2019 | Evaluating Terminology Translation in MT
Rejwanul Haque, Mohammed Hasanuzzaman, Andy Way |
CICLing (1) | 3 |
| 2019 | Adaptation of Machine Translation Models with Back-Translated Data Using Transductive Data Selection Methods
Alberto Poncelas, Gideon Maillette de Buy Wenniger, Andy Way |
CICLing (1) | 3 |
| 2019 | Take Help from Elder Brother: Old to Modern English NMT with Phrase Pair Feedback
Sukanta Sen, Mohammed Hasanuzzaman, Asif Ekbal, Pushpak Bhattacharyya, Andy Way |
CICLing (1) | 5 |
| 2019 | No Padding Please: Efficient Neural Handwriting RecognitionabstractNeural handwriting recognition (NHR) is the recognition of handwritten text with deep learning models, such as multi-dimensional long short-term memory (MDLSTM) recurrent neural networks. Models with MDLSTM layers have achieved state-of-the art results on handwritten text recognition tasks. While multi-directional MDLSTM-layers have an unbeaten ability to capture the complete context in all directions, this strength limits the possibilities for parallelization, and therefore comes at a high computational cost. In this work we develop methods to create efficient MDLSTM-based models for NHR, particularly a method aimed at eliminating computation waste that results from padding. This proposed method, called example packing, replaces wasteful stacking of padded examples with efficient tiling in a 2-dimensional grid. For word-based NHR this yields a speed improvement of factor 6.6 over an already efficient baseline of minimal padding for each batch separately. For line-based NHR the savings are more modest, but still significant. In addition to example packing, we propose: 1) a technique to optimize parallelization for dynamic graph definition frameworks including PyTorch, using convolutions with grouping, 2) a method for parallelization across GPUs for variable-length example batches. All our techniques are thoroughly tested on our own PyTorch re-implementation of MDLSTM-based NHR models. A thorough evaluation on the IAM dataset shows that our models are performing similar to earlier implementations of state-of-the art models. Our efficient NHR model and some of the reusable techniques discussed with it offer ways to realize relatively efficient models for the omnipresent scenario of variable-length inputs in deep learning. Gideon Maillette de Buy Wenniger, Lambert Schomaker, Andy Way |
ICDAR | 3 |
| 2019 | Selecting Artificially-Generated Sentences for Fine-Tuning Neural Machine TranslationabstractNeural Machine Translation (NMT) models tend to achieve best performance when larger sets of parallel sentences are provided for training.For this reason, augmenting the training set with artificially-generated sentence pairs can boost performance.Nonetheless, the performance can also be improved with a small number of sentences if they are in the same domain as the test set.Accordingly, we want to explore the use of artificially-generated sentences along with data-selection algorithms to improve Germanto-English NMT models trained solely with authentic data.In this work, we show how artificiallygenerated sentences can be more beneficial than authentic pairs, and demonstrate their advantages when used in combination with dataselection algorithms. Alberto Poncelas, Andy Way |
INLG | 2 |
| 2019 | Large-scale Machine Translation Evaluation of the iADAATPA Project
Sheila Castilho, Natália Resende, Federico Gaspari, Andy Way, Tony O'Dowd, Marek Mazur, Manuel Herranz, Alexandre Helle, Gema Ramírez-Sánchez, Víctor M. Sánchez-Cartagena, Marcis Pinnis, Valters Sics |
MTSummit (2) | 4 |
| 2019 | Pivot Machine Translation in INTERACT Project
Chao-Hong Liu, Andy Way, Catarina Cruz Silva |
MTSummit (2) | 2 |
| 2019 | When less is more in Neural Quality Estimation of Machine Translation. An industry case study
Dimitar Sht. Shterionov, Félix do Carmo, Joss Moorkens, Eric Paquin, Dag Schmidtke, Declan Groves, Andy Way |
MTSummit (2) | 7 |
| 2019 | Lost in Translation: Loss and Decay of Linguistic Richness in Machine Translation
Eva Vanmassenhove, Dimitar Sht. Shterionov, Andy Way |
MTSummit (1) | 3 |
| 2019 | PRINCIPLE: Providing Resources in Irish, Norwegian, Croatian and Icelandic for the Purposes of Language Engineering
Andy Way, Federico Gaspari |
MTSummit (2) | 1 |
| 2019 | Ruslan Mitkov, Johanna Monti, Gloria Corpas Pastor, and Violeta Seretan (eds): Multiword units in machine translation and translation technology - Current Issues in Linguistic Theory, Volume 341, John Benjamin Publishing Company, Amsterdam & Philadelphia, 2018, ix+259 pp, ISBN 978-90-272-0060-0 (HB), ISBN 978-90-272-6420-6 (e-book)
Rejwanul Haque, Mohammed Hasanuzzaman, Andy Way |
Mach. Transl. | 3 |
| 2019 | Post-editing neural machine translation versus translation memory segments
Pilar Sánchez-Gijón, Joss Moorkens, Andy Way |
Mach. Transl. | 3 |
| 2018 | A Decision-Level Approach to Multimodal Sentiment Analysis
Haithem Afli, Jason Burns, Andy Way |
CICLing (2) | 3 |
| 2018 | Incorporating Deep Visual Features into Multiobjective based Multi-view Search Results ClusteringabstractCurrent paper explores the use of multi-view learning for search result clustering. A web-snippet can be represented using multiple views. Apart from textual view cued by both the semantic and syntactic information, a complimentary view extracted from images contained in the web-snippets is also utilized in the current framework. A single consensus partitioning is finally obtained after consulting these two individual views by the deployment of a multiobjective based clustering technique. Several objective functions including the values of a cluster quality measure measuring the goodness of partitionings obtained using different views and an agreement-disagreement index, quantifying the amount of oneness among multiple views in generating partitionings are optimized simultaneously using AMOSA. In order to detect the number of clusters automatically, concepts of variable length solutions and a vast range of permutation operators are introduced in the clustering process. Finally, a set of alternative partitioning are obtained on the final Pareto front by the proposed multi-view based multiobjective technique. Experimental results by the proposed approach on several benchmark test datasets of SRC with respect to different performance metrics evidently establish the power of visual and text-based views in achieving better search result clustering. Sayantan Mitra, Mohammed Hasanuzzaman, Sriparna Saha 0001, Andy Way |
COLING | 4 |
| 2018 | Tailoring Neural Architectures for Translating from Morphologically Rich LanguagesabstractA morphologically complex word (MCW) is a hierarchical constituent with meaning-preserving subunits, so word-based models which rely on surface forms might not be powerful enough to translate such structures. When translating from morphologically rich languages (MRLs), a source word could be mapped to several words or even a full sentence on the target side, which means an MCW should not be treated as an atomic unit. In order to provide better translations for MRLs, we boost the existing neural machine translation (NMT) architecture with a double- channel encoder and a double-attentive decoder. The main goal targeted in this research is to provide richer information on the encoder side and redesign the decoder accordingly to benefit from such information. Our experimental results demonstrate that we could achieve our goal as the proposed model outperforms existing subword- and character-based architectures and showed significant improvements on translating from German, Russian, and Turkish into English. Peyman Passban, Andy Way, Qun Liu 0001 |
COLING | 2 |
| 2018 | ELRI - European Language Resources InfrastructureabstractWe describe the European Language Resources Infrastructure project, whose main aim is the provision of an infrastructure to help collect, prepare and share language resources that can in turn improve translation services in Europe. Thierry Etchegoyhen, Borja Anza Porras, Andoni Azpeitia, Eva Martínez Garcia, Paulo Vale, José Luis Fonseca, Teresa Lynn, Jane Dunne, Federico Gaspari, Andy Way, Victoria Arranz, Khalid Choukri, Vladimir Popescu, Pedro Neiva, Rui Neto, Maite Melero, David Pérez-Fernández, António Branco, Ruben Branco, Luís Gomes 0002 |
EAMT | 10 |
| 2018 | Investigating Backtranslation in Neural Machine TranslationabstractA prerequisite for training corpus-based machine translation (MT) systems – either Statistical MT (SMT) or Neural MT (NMT) – is the availability of high-quality parallel data. This is arguably more important today than ever before, as NMT has been shown in many studies to outperform SMT, but mostly when large parallel corpora are available; in cases where data is limited, SMT can still outperform NMT. Recently researchers have shown that back-translating monolingual data can be used to create synthetic parallel corpora, which in turn can be used in combination with authentic parallel data to train a highquality NMT system. Given that large collections of new parallel text become available only quite rarely, backtranslation has become the norm when building state-of-the-art NMT systems, especially in resource-poor scenarios. However, we assert that there are many unknown factors regarding the actual effects of back-translated data on the translation capabilities of an NMT model. Accordingly, in this work we investigate how using back-translated data as a training corpus – both as a separate standalone dataset as well as combined with human-generated parallel data – affects the performance of an NMT model. We use incrementally larger amounts of back-translated data to train a range of NMT systems for German-to-English, and analyse the resulting translation performance. Alberto Poncelas, Dimitar Sht. Shterionov, Andy Way, Gideon Maillette de Buy Wenniger, Peyman Passban |
EAMT | 3 |
| 2018 | Feature Decay Algorithms for Neural Machine TranslationabstractNeural Machine Translation (NMT) systems require a lot of data to be competitive. For this reason, data selection techniques are used only for finetuning systems that have been trained with larger amounts of data. In this work we aim to use Feature Decay Algorithms (FDA) data selection techniques not only to fine-tune a system but also to build a complete system with less data. Our findings reveal that it is possible to find a subset of sentence pairs, that outperforms by 1.11 BLEU points the full training corpus, when used for training a German-English NMT system . Alberto Poncelas, Gideon Maillette de Buy Wenniger, Andy Way |
EAMT | 3 |
| 2018 | Perception vs. Acceptability of TM and SMT Output: What do translators prefer?abstractThis paper reports the results of two studies carried out with two different group of professional translators to find out how professionals perceive and accept SMT in comparison with TM. The first group translated and post-edited segments from English into German, and the second group from English into Spanish. Both studies had equivalent settings in order to guarantee the comparability of the results. It will also help to shed light upon the real benefit of SMT from which translators may take advantage. Pilar Sánchez-Gijón, Joss Moorkens, Andy Way |
EAMT | 3 |
| 2018 | Project PiPeNovel: Pilot on Post-editing NovelsabstractGiven (i) the rise of a new paradigm to machine translation based on neural networks that results in more fluent and less literal output than previous models and (ii) the maturity of machine-assisted translation via post-editing in industry, project PiPeNovel studies the feasibility of the post-editing workflow for literary text conducting experiments with professional literary translators. Antonio Toral, Martijn Wieling 0001, Sheila Castilho, Joss Moorkens, Andy Way |
EAMT | 5 |
| 2018 | Multi-Level Structured Self-Attentions for Distantly Supervised Relation ExtractionabstractAttention mechanisms are often used in deep neural networks for distantly supervised relation extraction (DS-RE) to distinguish valid from noisy instances.However, traditional 1-D vector attention models are insufficient for the learning of different contexts in the selection of valid instances to predict the relationship for an entity pair.To alleviate this issue, we propose a novel multi-level structured (2-D matrix) self-attention mechanism for DS-RE in a multi-instance learning (MIL) framework using bidirectional recurrent neural networks.In the proposed method, a structured word-level self-attention mechanism learns a 2-D matrix where each row vector represents a weight distribution for different aspects of an instance regarding two entities.Targeting the MIL issue, the structured sentence-level attention learns a 2-D matrix where each row vector represents a weight distribution on selection of different valid instances.Experiments conducted on two publicly available DS-RE datasets show that the proposed framework with a multi-level structured self-attention mechanism significantly outperform state-of-the-art baselines in terms of PR curves, P@N and F1 measures. Jinhua Du, Jingguang Han, Andy Way, Dadong Wan |
EMNLP | 3 |
| 2018 | Getting Gender Right in Neural MTabstractSpeakers of different languages must attend to and encode strikingly different aspects of the world in order to use their language correctly (Sapir, 1921;Slobin, 1996).One such difference is related to the way gender is expressed in a language.Saying "I am happy" in English, does not encode any additional knowledge of the speaker that uttered the sentence.However, many other languages do have grammatical gender systems and so such knowledge would be encoded.In order to correctly translate such a sentence into, say, French, the inherent gender information needs to be retained/recovered.The same sentence would become either "Je suis heureux", for a male speaker or "Je suis heureuse" for a female one.Apart from morphological agreement, demographic factors (gender, age, etc.) also influence our use of language in terms of word choices or even on the level of syntactic constructions (Tannen, 1991;Pennebaker et al., 2003).We integrate gender information into NMT systems.Our contribution is twofold: (1) the compilation of large datasets with speaker information for 20 language pairs, and (2) a simple set of experiments that incorporate gender information into NMT for multiple language pairs.Our experiments show that adding a gender feature to an NMT system significantly improves the translation quality for some language pairs. Eva Vanmassenhove, Christian Hardmeier, Andy Way |
EMNLP | 3 |
| 2018 | Learning to Jointly Translate and Predict Dropped Pronouns with a Shared Reconstruction MechanismabstractPronouns are frequently omitted in pro-drop languages, such as Chinese, generally leading to significant challenges with respect to the production of complete translations.Recently, Wang et al. (2018) proposed a novel reconstruction-based approach to alleviating dropped pronoun (DP) translation problems for neural machine translation models.In this work, we improve the original model from two perspectives.First, we employ a shared reconstructor to better exploit encoder and decoder representations.Second, we jointly learn to translate and predict DPs in an end-to-end manner, to avoid the errors propagated from an external DP prediction model.Experimental results show that our approach significantly improves both translation performance and DP prediction accuracy. Longyue Wang, Zhaopeng Tu, Andy Way, Qun Liu 0001 |
EMNLP | 3 |
| 2018 | FooTweets: A Bilingual Parallel Corpus of World Cup Tweets
Henny Sluyter-Gäthje, Pintu Lohar, Haithem Afli, Andy Way |
LREC | 4 |
| 2018 | Fine-Grained Temporal Orientation and its Relationship with Psycho-Demographic CorrelatesabstractSabyasachi Kamila, Mohammed Hasanuzzaman, Asif Ekbal, Pushpak Bhattacharyya, Andy Way. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Sabyasachi Kamila, Mohammed Hasanuzzaman, Asif Ekbal, Pushpak Bhattacharyya, Andy Way |
NAACL-HLT | 5 |
| 2018 | Improving Character-Based Decoding Using Target-Side Morphological Information for Neural Machine TranslationabstractPeyman Passban, Qun Liu, Andy Way. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Peyman Passban, Qun Liu 0001, Andy Way |
NAACL-HLT | 3 |
| 2018 | IDEA: An Interactive Dialogue Translation Demo System Using Furhat Robots
Jinhua Du, Darragh Blake, Longyue Wang, Clare Conran, Declan McKibben, Andy Way |
ECML/PKDD (3) | 6 |
| 2018 | Evaluating MT for massive open online courses - A multifaceted comparison between PBSMT and NMT systems
Sheila Castilho, Joss Moorkens, Federico Gaspari, Rico Sennrich, Andy Way, Panayota Georgakopoulou |
Mach. Transl. | 5 |
| 2018 | Human versus automatic quality evaluation of NMT and PBSMT
Dimitar Sht. Shterionov, Riccardo Superbo, Pat Nagle, Laura Casanellas, Tony O'Dowd, Andy Way |
Mach. Transl. | 6 |
| 2018 | Editors' foreword to the invited issue on SMT and NMT
Andy Way, Mikel L. Forcada |
Mach. Transl. | 1 |
| 2017 | Exploiting Cross-Sentence Context for Neural Machine TranslationabstractIn translation, considering the document as a whole can help to resolve ambiguities and inconsistencies.In this paper, we propose a cross-sentence context-aware approach and investigate the influence of historical contextual information on the performance of neural machine translation (NMT).First, this history is summarized in a hierarchical way.We then integrate the historical representation into NMT in two strategies: 1) a warm-start of encoder and decoder states, and 2) an auxiliary context source for updating decoder states.Experimental results on a large Chinese-English translation task show that our approach significantly improves upon a strong attention-based NMT system by up to +2.1 BLEU points. Longyue Wang, Zhaopeng Tu, Andy Way, Qun Liu 0001 |
EMNLP | 3 |
| 2017 | Demographic Word Embeddings for Racism Detection on TwitterabstractMost social media platforms grant users freedom of speech by allowing them to freely express their thoughts, beliefs, and opinions. Although this represents incredible and unique communication opportunities, it also presents important challenges. Online racism is such an example. In this study, we present a supervised learning strategy to detect racist language on Twitter based on word embedding that incorporate demographic (Age, Gender, and Location) information. Our methodology achieves reasonable classification accuracy over a gold standard dataset (F1=76.3%) and significantly improves over the classification performance of demographic-agnostic models. Mohammed Hasanuzzaman, Gaël Dias, Andy Way |
IJCNLP(1) | 3 |
| 2017 | Local Event Discovery from Tweets MetadataabstractWe present a two-step strategy that addresses fundamental deficiencies in social media-based event detection and achieves effective local event by taking advantage of geo-located data from Twitter. While previous work has mainly relied on an analysis of tweet text to identify local events, we show how to reliably detect events using meta-data analysis of geo-tagged tweets. The first step of the method identifies several spatio-temporal clusters within the dataset across both space and time using metadata to form potential candidate events. In the second step, it ranks all the candidates by the amount of hashtag/entity inequality. We used crowdsourcing to evaluate the proposed approach on a data set that contains millions of geo-tagged tweets. The results show that our framework performs reasonably well in terms of precision and discovers local events faster. Mohammed Hasanuzzaman, Andy Way |
K-CAP | 2 |
| 2017 | A Comparative Quality Evaluation of PBSMT and NMT using Professional Translators
Sheila Castilho, Joss Moorkens, Federico Gaspari, Rico Sennrich, Vilelmini Sosoni, Panayota Georgakopoulou, Pintu Lohar, Andy Way, Antonio Valerio Miceli Barone, Maria Gialama |
MTSummit (1) | 8 |
| 2017 | Neural Pre-Translation for Hybrid Machine Translation
Jinhua Du, Andy Way |
MTSummit (1) | 2 |
| 2017 | Temporality as Seen through Translation: A Case Study on Hindi Texts
Sabyasachi Kamila, Sukanta Sen, Mohammed Hasanuzzaman, Asif Ekbal, Andy Way, Pushpak Bhattacharyya |
MTSummit (1) | 5 |
| 2017 | The INTERACT Project and Crisis MT
Sharon O'Brien, Chao-Hong Liu, Andy Way, João Graça, Helena Moniz, Ellie Kemp, Rebecca Petras |
MTSummit (2) | 3 |
| 2017 | Elastic-substitution decoding for Hierarchical SMT: efficiency, richer search and double labels
Gideon Maillette de Buy Wenniger, Khalil Sima'an, Andy Way |
MTSummit (1) | 3 |
| 2017 | Syntax- and semantic-based reordering in hierarchical phrase-based statistical machine translationabstractWe present a syntax-based reordering model (RM) for hierarchical phrase-based statistical machine translation (HPB-SMT) enriched with semantic features. Our model brings a number of novel contributions: (i) while the previous dependency-based RM is limited to the reordering of head and dependant constituent pairs, we also model the reordering of pairs of dependants; (ii) Our model is enriched with semantic features (Wordnet synsets) in order to allow the reordering model to generalize to pairs not seen in training but with equivalent meaning. (iii) We evaluate our model on two language directions: English-to-Farsi and English-to-Turkish. These language pairs are particularly challenging due to the free word order, rich morphology and lack of resources of the target languages. We evaluate our RM both intrinsically (accuracy of the RM classifier) and extrinsically (MT). Our best configuration outperforms the baseline classifier by 5–29% on pairs of dependants and by 12–30% on head and dependant pairs while the improvement on MT ranges between 1.6% and 5.5% relative in terms of BLEU depending on language pair and domain. We also analyze the value of the feature weights to obtain further insights on the impact of the reordering-related features in the HPB-SMT model. We observe that the features of our RM are assigned significant weights and that our features are complementary to the reordering feature included by default in the HPB-SMT model. Arefeh Kazemi, Antonio Toral, Andy Way, S. Amirhassan Monadjemi, Mohammad Ali Nematbakhsh |
Expert Syst. Appl. | 3 |
| 2017 | A novel and robust approach for pro-drop language translationabstractA significant challenge for machine translation (MT) is the phenomena of dropped pronouns (DPs), where certain classes of pronouns are frequently dropped in the source language but should be retained in the target language. In response to this common problem, we propose a semi-supervised approach with a universal framework to recall missing pronouns in translation. Firstly, we build training data for DP generation in which the DPs are automatically labelled according to the alignment information from a parallel corpus. Secondly, we build a deep learning-based DP generator for input sentences in decoding when no corresponding references exist. More specifically, the generation has two phases: (1) DP position detection, which is modeled as a sequential labelling task with recurrent neural networks; and (2) DP prediction, which employs a multilayer perceptron with rich features. Finally, we integrate the above outputs into our statistical MT (SMT) system to recall missing pronouns by both extracting rules from the DP-labelled training data and translating the DP-generated input sentences. To validate the robustness of our approach, we investigate our approach on both Chinese–English and Japanese–English corpora extracted from movie subtitles. Compared with an SMT baseline system, experimental results show that our approach achieves a significant improvement of $$+$$ 1.58 BLEU points in translation performance with 66% F-score for DP generation accuracy for Chinese–English, and nearly $$+$$ 1 BLEU point with 58% F-score for Japanese–English. We believe that this work could help both MT researchers and industries to boost the performance of MT systems between pro-drop and non-pro-drop languages. Longyue Wang, Zhaopeng Tu, Siyou Liu, Hang Li 0001, Andy Way, Qun Liu 0001 |
Mach. Transl. | 6 |
| 2017 | Translating Low-Resource Languages by Vocabulary Adaptation from Close CounterpartsabstractSome natural languages belong to the same family or share similar syntactic and/or semantic regularities. This property persuades researchers to share computational models across languages and benefit from high-quality models to boost existing low-performance counterparts. In this article, we follow a similar idea, whereby we develop statistical and neural machine translation (MT) engines that are trained on one language pair but are used to translate another language. First we train a reliable model for a high-resource language, and then we exploit cross-lingual similarities and adapt the model to work for a close language with almost zero resources. We chose Turkish (Tr) and Azeri or Azerbaijani (Az) as the proposed pair in our experiments. Azeri suffers from lack of resources as there is almost no bilingual corpus for this language. Via our techniques, we are able to train an engine for the Az → English (En) direction, which is able to outperform all other existing models. Peyman Passban, Qun Liu 0001, Andy Way |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2016 | Graph-Based Translation Via Graph SegmentationabstractOne major drawback of phrase-based translation is that it segments an input sentence into continuous phrases.To support linguistically informed source discontinuity, in this paper we construct graphs which combine bigram and dependency relations and propose a graph-based translation model.The model segments an input graph into connected subgraphs, each of which may cover a discontinuous phrase.We use beam search to combine translations of each subgraph left-to-right to produce a complete translation.Experiments on Chinese-English and German-English tasks show that our system is significantly better than the phrase-based model by up to +1.5/+0.5 BLEU scores.By explicitly modeling the graph segmentation, our system obtains further improvement, especially on German-English. Liangyou Li, Andy Way, Qun Liu 0001 |
ACL (1) | 2 |
| 2016 | Enriching Phrase Tables for Statistical Machine Translation Using Mixed EmbeddingsabstractThe phrase table is considered to be the main bilingual resource for the phrase-based statistical machine translation (PBSMT) model. During translation, a source sentence is decomposed into several phrases. The best match of each source phrase is selected among several target-side counterparts within the phrase table, and processed by the decoder to generate a sentence-level translation. The best match is chosen according to several factors, including a set of bilingual features. PBSMT engines by default provide four probability scores in phrase tables which are considered as the main set of bilingual features. Our goal is to enrich that set of features, as a better feature set should yield better translations. We propose new scores generated by a Convolutional Neural Network (CNN) which indicate the semantic relatedness of phrase pairs. We evaluate our model in different experimental settings with different language pairs. We observe significant improvements when the proposed features are incorporated into the PBSMT pipeline. Peyman Passban, Qun Liu 0001, Andy Way |
COLING | 3 |
| 2016 | Topic-Informed Neural Machine TranslationabstractIn recent years, neural machine translation (NMT) has demonstrated state-of-the-art machine translation (MT) performance. It is a new approach to MT, which tries to learn a set of parameters to maximize the conditional probability of target sentences given source sentences. In this paper, we present a novel approach to improve the translation performance in NMT by conveying topic knowledge during translation. The proposed topic-informed NMT can increase the likelihood of selecting words from the same topic and domain for translation. Experimentally, we demonstrate that topic-informed NMT can achieve a 1.15 (3.3% relative) and 1.67 (5.4% relative) absolute improvement in BLEU score on the Chinese-to-English language pair using NIST 2004 and 2005 test sets, respectively, compared to NMT without topic information. Jian Zhang 0003, Liangyou Li, Andy Way, Qun Liu 0001 |
COLING | 3 |
| 2016 | Fast Gated Neural Domain Adaptation: Language Model as a Case StudyabstractNeural network training has been shown to be advantageous in many natural language processing applications, such as language modelling or machine translation. In this paper, we describe in detail a novel domain adaptation mechanism in neural network training. Instead of learning and adapting the neural network on millions of training sentences – which can be very time-consuming or even infeasible in some cases – we design a domain adaptation gating mechanism which can be used in recurrent neural networks and quickly learn the out-of-domain knowledge directly from the word vector representations with little speed overhead. In our experiments, we use the recurrent neural network language model (LM) as a case study. We show that the neural LM perplexity can be reduced by 7.395 and 12.011 using the proposed domain adaptation mechanism on the Penn Treebank and News data, respectively. Furthermore, we show that using the domain-adapted neural LM to re-rank the statistical machine translation n-best list on the French-to-English language pair can significantly improve translation quality. Jian Zhang 0003, Andy Way, Qun Liu 0001 |
COLING | 3 |
| 2016 | Identifying Temporal Orientation of Word Senses
Mohammed Hasanuzzaman, Gaël Dias, Stéphane Ferrari, Yann Mathet, Andy Way |
CoNLL | 5 |
| 2016 | Comparing Translator Acceptability of TM and SMT Outputs
Joss Moorkens, Andy Way |
EAMT | 2 |
| 2016 | Improving Phrase-Based SMT Using Cross-Granularity Embedding Similarity
Peyman Passban, Chris Hokamp, Andy Way, Qun Liu 0001 |
EAMT | 3 |
| 2016 | Using SMT for OCR Error Correction of Historical Texts
Haithem Afli, Zhengwei Qiu, Andy Way, Paraic Sheridan |
LREC | 3 |
| 2016 | Using BabelNet to Improve OOV Coverage in SMT
Jinhua Du, Andy Way, Andrzej Zydron |
LREC | 2 |
| 2016 | Enhancing Access to Online Education: Quality Machine Translation of MOOC Content
Valia Kordoni, Antal van den Bosch, Katia Kermanidis, Vilelmini Sosoni, Kostadin Cholakov, Iris Hendrickx, Matthias Huck, Andy Way |
LREC | 8 |
| 2016 | Automatic Construction of Discourse Corpora for Dialogue Translation
Longyue Wang, Zhaopeng Tu, Andy Way, Qun Liu 0001 |
LREC | 4 |
| 2016 | ProphetMT: A Tree-based SMT-driven Controlled Language Authoring/Post-Editing Tool
Jinhua Du, Qun Liu 0001, Andy Way |
LREC | 4 |
| 2016 | A Novel Approach to Dropped Pronoun TranslationabstractLongyue Wang, Zhaopeng Tu, Xiaojun Zhang, Hang Li, Andy Way, Qun Liu. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016. Longyue Wang, Zhaopeng Tu, Hang Li 0001, Andy Way, Qun Liu 0001 |
HLT-NAACL | 5 |
| 2016 | Using Wordnet to Improve Reordering in Hierarchical Phrase-Based Statistical Machine TranslationabstractWe propose the use of WordNet synsets in a syntax-based reordering model for hierarchical statistical machine translation (HPB-SMT) to enable the model to generalize to phrases not seen in the training data but that have equivalent meaning.We detail our methodology to incorporate synsets' knowledge in the reordering model and evaluate the resulting WordNetenhanced SMT systems on the English-to-Farsi language direction.The inclusion of synsets leads to the best BLEU score, outperforming the baseline (standard HPB-SMT) by 0.6 points absolute. Arefeh Kazemi, Antonio Toral, Andy Way |
GWC | 3 |
| 2016 | Combining translation memories and statistical machine translation using sparse features
Liangyou Li, Carla Parra Escartín, Andy Way, Qun Liu 0001 |
Mach. Transl. | 3 |
| 2016 | Boosting Neural POS Tagger for Farsi Using Morphological InformationabstractFarsi (Persian) is a low-resource language that suffers from the data sparsity problem and a lack of efficient processing tools. Due to their broad application in natural language processing tasks, part-of-speech (POS) taggers are one of those important tools that should be considered in this respect. Despite recent work on Farsi tagging, there is still room for improvement. The best reported accuracy so far is 96%, which in special cases can rise to 96.9%. The main problem with existing taggers is their inefficiency in coping with out-of-vocabulary (OOV) words. Addressing both problems of accuracy and OOV words, we developed a neural network-based POS tagger (NPT) that performs efficiently on Farsi. Despite using less data, NPT provides better results in comparison to state-of-the-art systems. Our proposed tagger performs with an accuracy of 97.4%, with performance highly influenced by morphological features. We carry out a shallow morphological analysis and show considerable improvement over the baseline configuration. Peyman Passban, Qun Liu 0001, Andy Way |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 3 |
| 2015 | Dependency-based Reordering Model for Constituent Pairs in Hierarchical SMT
Arefeh Kazemi, Antonio Toral, Andy Way, S. Amirhassan Monadjemi, Mohammad Ali Nematbakhsh |
EAMT | 3 |
| 2015 | TraMOOC: Translation for Massive Open Online Courses
Valia Kordoni, Kostadin Cholakov, Markus Egg, Andy Way, Lexi Birch, Katia Kermanidis, Vilelmini Sosoni, Dimitrios Tsoumakos, Antal van den Bosch, Iris Hendrickx, Michael Papadopoulos, Panayota Georgakopoulou, Maria Gialama, Menno van Zaanen, Ioana Buliga, Mitja Jermol, Davor Orlic |
EAMT | 4 |
| 2015 | Benchmarking SMT Performance for Farsi Using the TEP++ Corpus
Peyman Passban, Andy Way, Qun Liu 0001 |
EAMT | 2 |
| 2015 | Abu-MaTran: Automatic building of Machine Translation
Antonio Toral, Flammie A. Pirinen, Andy Way, Gema Ramírez-Sánchez, Sergio Ortiz-Rojas, Raphaël Rubino, Miquel Esplà-Gomis, Mikel L. Forcada, Vassilis Papavassiliou, Prokopis Prokopidis, Nikola Ljubesic |
EAMT | 3 |
| 2015 | Dependency Graph-to-String TranslationabstractCompared to tree grammars, graph grammars have stronger generative capacity over structures.Based on an edge replacement grammar, in this paper we propose to use a synchronous graph-to-string grammar for statistical machine translation.The graph we use is directly converted from a dependency tree by labelling edges.We build our translation model in the log-linear framework with standard features.Large-scale experiments on Chinese-English and German-English tasks show that our model is significantly better than the state-of-the-art hierarchical phrase-based (HPB) model and a recently improved dependency tree-to-string model on BLEU, METEOR and TER scores.Experiments also suggest that our model has better capability to perform long-distance reordering and is more suitable for translating long sentences. Liangyou Li, Andy Way, Qun Liu 0001 |
EMNLP | 2 |
| 2015 | An empirical study of segment prioritization for incrementally retrained post-editing-based SMT
Jinhua Du, Ankit K. Srivastava, Andy Way, Alfredo Maldonado-Guerra |
MTSummit | 3 |
| 2014 | Standard language variety conversion for content localisation via SMT
Federico Fancellu, Andy Way, Morgan O'Brien |
EAMT | 2 |
| 2014 | Extrinsic evaluation of web-crawlers in machine translation: a study on Croatian-English for the tourism domain
Antonio Toral, Raphaël Rubino, Miquel Esplà-Gomis, Flammie A. Pirinen, Andy Way, Gema Ramírez-Sánchez |
EAMT | 5 |
| 2013 | Manual labour: tackling machine translation for sign languages
Sara Morrissey, Andy Way |
Mach. Transl. | 2 |
| 2012 | Translation Quality-Based Supplementary Data Selection by Incremental Update of Translation Models
Pratyush Banerjee, Sudip Kumar Naskar, Johann Roturier, Andy Way, Josef van Genabith |
COLING | 4 |
| 2012 | Extending CCG-based Syntactic Constraints in Hierarchical Phrase-Based SMT
Hala Almaghout, Jie Jiang 0002, Andy Way |
EAMT | 3 |
| 2012 | Domain Adaptation in SMT of User-Generated Forum Content Guided by OOV Word Reduction: Normalization and/or Supplementary Data
Pratyush Banerjee, Sudip Kumar Naskar, Johann Roturier, Andy Way, Josef van Genabith |
EAMT | 4 |
| 2012 | From Subtitles to Parallel Corpora
Mark Fishel, Panayota Georgakopoulou, Sergio Penkale, Volha Petukhova, Matej Rojc, Martin Volk 0001, Andy Way |
EAMT | 7 |
| 2012 | SUMAT: Data Collection and Parallel Corpus Compilation for Machine Translation of Subtitles
Volha Petukhova, Rodrigo Agerri, Mark Fishel, Sergio Penkale, Arantza del Pozo, Mirjam Sepesy Maucec, Andy Way, Panayota Georgakopoulou, Martin Volk 0001 |
LREC | 7 |
| 2012 | Efficient accurate syntactic direct translation models: one tree at a time
Hany Hassan, Khalil Sima'an, Andy Way |
Mach. Transl. | 3 |
| 2012 | What types of word alignment improve statistical machine translation?
Patrik Lambert, Simon Petit-Renaud, Yanjun Ma, Andy Way |
Mach. Transl. | 4 |
| 2012 | David Bellos (ed): Is that a fish in your ear: translation and the meaning of everything - Particular Books, Penguin Group, London, 2011, ix + 390 pp, ISBN 978-1-846-14464-6
Andy Way |
Mach. Transl. | 1 |
| 2011 | Consistent Translation using Discriminative Learning - A Translation Memory-inspired Approach
Yanjun Ma, Yifan He 0007, Andy Way, Josef van Genabith |
ACL | 3 |
| 2011 | CCG Contextual labels in Hierarchical Phrase-Based SMT
Hala Almaghout, Jie Jiang 0002, Andy Way |
EAMT | 3 |
| 2011 | Experiments on Domain Adaptation for Patent Machine Translation in the PLuTO project
Alexandru Ceausu, John Tinsley, Jian Zhang 0003, Andy Way |
EAMT | 4 |
| 2011 | Using Example-Based MT to Support Statistical MT when Translating Homogeneous Data in a Resource-Poor Setting
Sandipan Dandapat, Sara Morrissey, Andy Way, Mikel L. Forcada |
EAMT | 3 |
| 2011 | Combining Semantic and Syntactic Generalization in Example-Based Machine Translation
Sarah Ebling, Andy Way, Martin Volk 0001, Sudip Kumar Naskar |
EAMT | 2 |
| 2011 | Towards Using Web-Crawled Data for Domain Adaptation in Statistical Machine Translation
Pavel Pecina, Antonio Toral, Andy Way, Vassilis Papavassiliou, Prokopis Prokopidis, Maria Giagkou |
EAMT | 3 |
| 2011 | Preliminary Experiments on Using Users' Post-Editions to Enhance a SMT System Oracle-based Training for Phrase-based Statistical Machine Translation
Ankit K. Srivastava, Yanjun Ma, Andy Way |
EAMT | 3 |
| 2011 | A Comparative Evaluation of Research vs. Online MT Systems
Antonio Toral, Federico Gaspari, Sudip Kumar Naskar, Andy Way |
EAMT | 4 |
| 2011 | Towards a User-Friendly Webservice Architecture for Statistical Machine Translation in the PANACEA project
Antonio Toral, Pavel Pecina, Marc Poch, Andy Way |
EAMT | 4 |
| 2011 | Domain Adaptation in Statistical Machine Translation of User-Forum Data using Component Level Mixture Modelling
Pratyush Banerjee, Sudip Kumar Naskar, Johann Roturier, Andy Way, Josef van Genabith |
MTSummit | 4 |
| 2011 | Rich Linguistic Features for Translation Memory-Inspired Consistent Translation
Yifan He 0007, Yanjun Ma, Andy Way, Josef van Genabith |
MTSummit | 3 |
| 2011 | Phonetic Representation-Based Speech Translation
Jie Jiang 0002, Julie Carson-Berndsen, Peter Cahill, Andy Way |
MTSummit | 5 |
| 2011 | A Framework for Diagnostic Evaluation of MT Based on Linguistic Checkpoints
Sudip Kumar Naskar, Antonio Toral, Federico Gaspari, Andy Way |
MTSummit | 4 |
| 2011 | Integrating source-language context into phrase-based statistical machine translation
Rejwanul Haque, Sudip Kumar Naskar, Antal van den Bosch, Andy Way |
Mach. Transl. | 4 |
| 2011 | Improved Chinese-English SMT with Chinese "DE" Construction Classification and ReorderingabstractSyntactic reordering on the source side has been demonstrated to be helpful and effective for handling different word orders between source and target languages in SMT. In this article, we focus on the Chinese (DE) construction which is flexible and ubiquitous in Chinese and has many different ways to be translated into English so that it is a major source of word order differences in terms of translation quality. This article carries out the Chinese “DE” construction study for Chinese--English SMT in which we propose a new classifier model---discriminative latent variable model (DPLVM)---with new features to improve the classification accuracy and indirectly improve the translation quality compared to a log-linear classifier. The DE classifier is used to recognize DE structures in both training and test sentences of Chinese, and then perform word reordering to make the Chinese sentences better match the word order of English. In order to investigate the impact of the DE classification and reordering in the source side on different types of SMT systems (namely PB-SMT, hierarchical PB-SMT (HPB-SMT) as well as the syntax-based SMT (SAMT)), we conduct a series of experiments on NIST 2005 and 2008 test sets to verify the effectiveness of our proposed model. The experimental results show that the MT systems using the data reordered by our proposed model outperform the baseline systems by 3.01% and 4.03% relative points on the NIST 2005 test set, 4.64% and 4.62% relative points on the NIST 2008 test set in terms of BLEU score for PB-SMT and HPB-SMT respectively. However, the DE classification method does not perform significantly well for SAMT. Additionally, we also conducted some experiments to evaluate our DE classification and reordering approach on the word alignment and phrase table in terms of these three types of SMT systems. Jinhua Du, Andy Way |
ACM Trans. Asian Lang. Inf. Process. | 2 |
| 2010 | Bridging SMT and TM with Translation Recommendation
Yifan He 0007, Yanjun Ma, Josef van Genabith, Andy Way |
ACL | 4 |
| 2010 | A Discriminative Latent Variable-Based "DE" Classifier for Chinese-English SMT
Jinhua Du, Andy Way |
COLING | 2 |
| 2010 | TMX Markup: A Challenge When Adapting SMT to the Localisation Environment
Jinhua Du, Johann Roturier, Andy Way |
EAMT | 3 |
| 2010 | The Impact of Source-Side Syntactic Reordering on Hierarchical Phrase-based SMT
Jinhua Du, Andy Way |
EAMT | 2 |
| 2010 | Lattice Score Based Data Cleaning for Phrase-Based Statistical Machine Translation
Jie Jiang 0002, Julie Carson-Berndsen, Andy Way |
EAMT | 3 |
| 2010 | Statistical Analysis of Alignment Characteristics for Phrase-based Machine Translation
Patrik Lambert, Simon Petit-Renaud, Yanjun Ma, Andy Way |
EAMT | 4 |
| 2010 | Facilitating Translation Using Source Language Paraphrase Lattices
Jinhua Du, Jie Jiang 0002, Andy Way |
EMNLP | 3 |
| 2010 | Metric and reference factors in minimum error rate training
Yifan He 0007, Andy Way |
Mach. Transl. | 2 |
| 2010 | Panning for EBMT gold, or "Remembering not to forget"
Andy Way |
Mach. Transl. | 1 |
| 2009 | Exploiting Parallel Treebanks to Improve Phrase-Based Statistical Machine Translation
John Tinsley, Mary Hearne, Andy Way |
CICLing | 3 |
| 2009 | Bilingually Motivated Domain-Adapted Word Segmentation for Statistical Machine Translation
Yanjun Ma, Andy Way |
EACL | 2 |
| 2009 | Using Supertags as Source Language Context in SMT
Rejwanul Haque, Sudip Kumar Naskar, Yanjun Ma, Andy Way |
EAMT | 4 |
| 2009 | Learning Labelled Dependencies in Machine Translation Evaluation
Yifan He 0007, Andy Way |
EAMT | 2 |
| 2009 | Tuning Syntactically Enhanced Word Alignment for Statistical Machine Translation
Yanjun Ma, Patrik Lambert, Andy Way |
EAMT | 3 |
| 2009 | Optimal Bilingual Data for French-English PB-SMT
Sylwia Ozdowska, Andy Way |
EAMT | 2 |
| 2009 | Marker-Based Filtering of Bilingual Phrase Pairs for SMT
Felipe Sánchez-Martínez, Andy Way |
EAMT | 2 |
| 2009 | Accuracy-Based Scoring for DOT: Towards Direct Error Minimization for Data-Oriented Translation
Daniel Galron, Sergio Penkale, Andy Way, I. Dan Melamed |
EMNLP | 3 |
| 2009 | A Syntactified Direct Translation Model with Linear-time Decoding
Hany Hassan, Khalil Sima'an, Andy Way |
EMNLP | 3 |
| 2009 | Using same-language machine translation to create alternative target sequences for text-to-speech synthesisabstractModern speech synthesis systems attempt to produce\nspeech utterances from an open domain of words. In some situations, the synthesiser will not have the appropriate units to pronounce some words or phrases accurately but it still must attempt to pronounce them. This paper presents a hybrid machine translation and unit selection speech synthesis system. The machine translation system was trained with English as the source and target language. Rather than the synthesiser only saying the input text as would happen in conventional synthesis systems, the synthesiser may say an alternative utterance with the same\nmeaning. This method allows the synthesiser to overcome the\nproblem of insufficient units in runtime. Peter Cahill, Jinhua Du, Andy Way, Julie Carson-Berndsen |
INTERSPEECH | 3 |
| 2009 | Capturing Lexical Variation in MT Evaluation Using Automatically Built Sense-Cluster Inventories
Marianna Apidianaki, Yifan He 0007, Andy Way |
PACLIC | 3 |
| 2009 | Dependency Relations as Source Context in Phrase-Based SMT
Rejwanul Haque, Sudip Kumar Naskar, Antal van den Bosch, Andy Way |
PACLIC | 4 |
| 2009 | Experiments on Domain Adaptation for English--Hindi SMT
Rejwanul Haque, Sudip Kumar Naskar, Josef van Genabith, Andy Way |
PACLIC | 4 |
| 2009 | Automatically generated parallel treebanks and their exploitability in machine translation
John Tinsley, Andy Way |
Mach. Transl. | 2 |
| 2009 | Bilingually Motivated Word Segmentation for Statistical Machine TranslationabstractWe introduce a bilingually motivated word segmentation approach to languages where word boundaries are not orthographically marked, with application to Phrase-Based Statistical Machine Translation (PB-SMT). Our approach is motivated from the insight that PB-SMT systems can be improved by optimizing the input representation to reduce the predictive power of translation models. We firstly present an approach to optimize the existing segmentation of both source and target languages for PB-SMT and demonstrate the effectiveness of this approach using a Chinese--English MT task, that is, to measure the influence of the segmentation on the performance of PB-SMT systems. We report a 5.44% relative increase in Bleu score and a consistent increase according to other metrics. We then generalize this method for Chinese word segmentation without relying on any segmenters and show that using our segmentation PB-SMT can achieve more consistent state-of-the-art performance across two domains. There are two main advantages of our approach. First of all, it is adapted to the specific translation task at hand by taking the corresponding source (target) language into account. Second, this approach does not rely on manually segmented training data so that it can be automatically adapted for different domains. Yanjun Ma, Andy Way |
ACM Trans. Asian Lang. Inf. Process. | 2 |
| 2008 | Automatic Generation of Parallel Treebanks
Ventsislav Zhechev, Andy Way |
COLING | 2 |
| 2008 | The ATIS Sign Language Corpus
Jan Bungeroth, Daniel Stein, Philippe Dreuw, Hermann Ney, Sara Morrissey, Andy Way, Lynette van Zijl |
LREC | 6 |
| 2008 | A syntactic language model based on incremental CCG parsingabstractSyntactically-enriched language models (parsers) constitute a promising component in applications such as machine translation and speech-recognition. To maintain a useful level of accuracy, existing parsers are non-incremental and must span a combinatorially growing space of possible structures as every input word is processed. This prohibits their incorporation into standard linear-time decoders. In this paper, we present an incremental, linear-time dependency parser based on Combinatory Categorial Grammar (CCG) and classification techniques. We devise a deterministic transform of CCG-bank canonical derivations into incremental ones, and train our parser on this data. We discover that a cascaded, incremental version provides an appealing balance between efficiency and accuracy. Hany Hassan, Khalil Sima'an, Andy Way |
SLT | 3 |
| 2008 | Wide-Coverage Deep Statistical Parsing Using Automatic Dependency Structure AnnotationabstractA number of researchers have recently conducted experiments comparing “deep” hand-crafted wide-coverage with “shallow” treebank- and machine-learning-based parsers at the level of dependencies, using simple and automatic methods to convert tree output generated by the shallow parsers into dependencies. In this article, we revisit such experiments, this time using sophisticated automatic LFG f-structure annotation methodologies with surprising results. We compare various PCFG and history-based parsers to find a baseline parsing system that fits best into our automatic dependency structure annotation technique. This combined system of syntactic parser and dependency structure annotation is compared to two hand-crafted, deep constraint-based parsers, RASP and XLE. We evaluate using dependency-based gold standards and use the Approximate Randomization Test to test the statistical significance of the results. Our experiments show that machine-learning-based shallow grammars augmented with sophisticated automatic dependency annotation technology outperform hand-crafted, deep, wide-coverage constraint grammars. Currently our best system achieves an f-score of 82.73% against the PARC 700 Dependency Bank, a statistically significant improvement of 2.18% over the most recent results of 80.55% for the hand-crafted LFG grammar and XLE parsing system and an f-score of 80.23% against the CBS 500 Dependency Bank, a statistically significant 3.66% improvement over the 76.57% achieved by the hand-crafted RASP grammar and parsing system. Aoife Cahill, Michael Burke, Ruth O'Donovan, Stefan Riezler, Josef van Genabith, Andy Way |
Comput. Linguistics | 6 |
| 2008 | Syntactically Lexicalized Phrase-Based SMTabstractUntil quite recently, extending phrase-based statistical machine translation (PBSMT) with syntactic knowledge caused system performance to deteriorate. The most recent successful enrichments of PBSMT with hierarchical structure either employ nonlinguistically motivated syntax for capturing hierarchical reordering phenomena, or extend the phrase translation table with redundantly ambiguous syntactic structures over phrase pairs. In this paper, we present an extended, harmonized account of our previous work which showed that incorporating linguistically motivated lexical syntactic descriptions, calledsupertags, can yield significantly better PBSMT systems at insignificant extra computational cost. We describe a novel PBSMT model that integrates supertags into the target language model and the target side of the translation model. Two kinds of supertags are employed: those from lexicalized tree-adjoining grammar and combinatory categorial grammar. Despite the differences between the two sets of supertags, they give similar improvements. In addition to integrating the Markov supertagging approach in PBSMT, we explore the utility of a new surface grammaticality measure based on combinatory operators. We perform various experiments on the Arabic-to-English NIST 2005 test set addressing the issues of sparseness, scalability, and the utility of system subcomponents. We show that even when the parallel training data grows very large, the supertagged system retains a relatively stable absolute performance advantage over the unadorned PBSMT system. Arguably, this hints at a performance gap that cannot be bridged by acquiring more phrase pairs. Our best result shows a relative improvement of 6.1% over a state-of-the-art PBSMT model, which compares favorably with the leading systems on the NIST 2005 task. We also demonstrate that the advantages of a supertag-based system carry over to German-English, where improvements of up to 8.9% relative to the baseline system are observed. Hany Hassan, Khalil Sima'an, Andy Way |
IEEE Trans. Speech Audio Process. | 3 |
| 2007 | Supertagged Phrase-Based Statistical Machine Translation
Hany Hassan, Khalil Sima'an, Andy Way |
ACL | 3 |
| 2007 | Bootstrapping Word Alignment via Word Packing
Yanjun Ma, Nicolas Stroppa, Andy Way |
ACL | 3 |
| 2007 | Comparing rule-based and data-driven approaches to Spanish-to-Basque machine translation
Gorka Labaka, Nicolas Stroppa, Andy Way, Kepa Sarasola |
MTSummit | 3 |
| 2007 | Combining data-driven MT systems for improved sign language translation
Sara Morrissey, Andy Way, Daniel Stein, Jan Bungeroth, Hermann Ney |
MTSummit | 2 |
| 2007 | Robust language pair-independent sub-tree alignment
John Tinsley, Ventsislav Zhechev, Mary Hearne, Andy Way |
MTSummit | 4 |
| 2007 | Evaluating machine translation with LFG dependencies
Karolina Owczarzak, Josef van Genabith, Andy Way |
Mach. Transl. | 3 |
| 2006 | Hybridity in MT. Experiments on the Europarl Corpus
Declan Groves, Andy Way |
EAMT | 2 |
| 2006 | Disambiguation Strategies for Data-Oriented Translation
Mary Hearne, Andy Way |
EAMT | 2 |
| 2006 | A Syntactic Skeleton for Statistical Machine Translation
Bart Mellebeek, Karolina Owczarzak, Declan Groves, Josef van Genabith, Andy Way |
EAMT | 5 |
| 2006 | Syntactic Phrase-Based Statistical Machine TranslationabstractPhrase-based statistical machine translation (PBSMT) systems represent the dominant approach in MT today. However, unlike systems in other paradigms, it has proven difficult to date to incorporate syntactic knowledge in order to improve translation quality. This paper improves on recent research which uses 'syntactified' target language phrases, by incorporating supertags as constraints to better resolve parse tree fragments. In addition, we do not impose any sentence-length limit, and using a log-linear decoder, we outperform a state-of-the-art PBSMT system by over 1.3 BLEU points (or 3.51% relative) on the NIST 2003 Arabic-English test corpus. Hany Hassan, Mary Hearne, Andy Way, Khalil Sima'an |
SLT | 3 |
| 2005 | TransBooster: boosting the performance of wide-coverage machine translation systems
Bart Mellebeek, Anna Khasin, Josef van Genabith, Andy Way |
EAMT | 4 |
| 2005 | Improving Online Machine Translation SystemsabstractIn (Mellebeek et al., 2005), we proposed the design, implementation and evaluation of a novel and modular approach to boost the translation performance of existing, wide-coverage, freely available machine translation systems, based on reliable and fast automatic decomposition of the translation input and corresponding composition of translation output. Despite showing some initial promise, our method did not improve on the baseline Logomedia1 and Systran2 MT systems. In this paper, we improve on the algorithm presented in (Mellebeek et al., 2005), and on the same test data, show increased scores for a range of automatic evaluation metrics. Our algorithm now outperforms Logomedia, obtains similar results to SDL3 and falls tantalisingly short of the performance achieved by Systran. Bart Mellebeek, Anna Khasin, Karolina Owczarzak, Josef van Genabith, Andy Way |
MTSummit | 5 |
| 2005 | Large-Scale Induction and Evaluation of Lexical Resources from the Penn-II and Penn-III TreebanksabstractWe present a methodology for extracting subcategorization frames based on an automatic lexical-functional grammar (LFG) f-structure annotation algorithm for the Penn-II and Penn-III Treebanks. We extract syntactic-function-based subcategorization frames (LFG semantic forms) and traditional CFG category-based subcategorization frames as well as mixed function/category-based frames, with or without preposition information for obliques and particle information for particle verbs. Our approach associates probabilities with frames conditional on the lemma, distinguishes between active and passive frames, and fully reflects the effects of long-distance dependencies in the source data structures. In contrast to many other approaches, ours does not predefine the subcategorization frame types extracted, learning them instead from the source data. Including particles and prepositions, we extract 21,005 lemma frame types for 4,362 verb lemmas, with a total of 577 frame types and an average of 4.8 frame types per verb. We present a large-scale evaluation of the complete set of forms extracted against the full COMLEX resource. To our knowledge, this is the largest and most complete evaluation of subcategorization frames acquired automatically for English. Ruth O'Donovan, Michael Burke, Aoife Cahill, Josef van Genabith, Andy Way |
Comput. Linguistics | 5 |
| 2005 | Introduction to special issue on example-based machine translation
Michael Carl, Andy Way |
Mach. Transl. | 2 |
| 2005 | Hybrid data-driven models of machine translation
Declan Groves, Andy Way |
Mach. Transl. | 2 |
| 2005 | Controlled Translation in an Example-based Environment: What do Automatic Evaluation Metrics Tell Us?
Andy Way, Nano Gough |
Mach. Transl. | 1 |
| 2005 | Comparing example-based and statistical machine translationabstractIn previous work (Gough and Way 2004), we showed that our Example-Based Machine Translation (EBMT) system improved with respect to both coverage and quality when seeded with increasing amounts of training data, so that it significantly outperformed the on-line MT system Logomedia according to a wide variety of automatic evaluation metrics. While it is perhaps unsurprising that system performance is correlated with the amount of training data, we address in this paper the question of whether a large-scale, robust EBMT system such as ours can outperform a Statistical Machine Translation (SMT) system. We obtained a large English-French translation memory from Sun Microsystems from which we randomly extracted a near 4K test set. The remaining data was split into three training sets, of roughly 50K, 100K and 200K sentence-pairs in order to measure the effect of increasing the size of the training data on the performance of the two systems. Our main observation is that contrary to perceived wisdom in the field, there appears to be little substance to the claim that SMT systems are guaranteed to outperform EBMT systems when confronted with ‘enough’ training data. Our tests on a 4.8 million word bitext indicate that while SMT appears to outperform our system for French-English on a number of metrics, for English-French, on all but one automatic evaluation metric, the performance of our EBMT system is superior to the baseline SMT model. Andy Way, Nano Gough |
Nat. Lang. Eng. | 1 |
| 2004 | Long-Distance Dependency Resolution in Automatically Acquired Wide-Coverage PCFG-Based LFG ApproximationsabstractThis paper shows how finite approximations of long distance dependency (LDD) resolution can be obtained automatically for wide-coverage, robust, probabilistic Lexical-Functional Grammar (LFG) resources acquired from treebanks. We extract LFG subcategorisation frames and paths linking LDD reentrancies from f-structures generated automatically for the Penn-II treebank trees and use them in an LDD resolution algorithm to parse new text. Unlike (Collins, 1999; Johnson, 2000), in our approach resolution of LDDs is done at f-structure (attribute-value structure representations of basic predicate-argument or dependency structure) without empty productions, traces and coindexation in CFG parse trees. Currently our best automatically induced grammars achieve 80.97% f-score for f-structures parsing section 23 of the WSJ part of the Penn-II treebank and evaluating against the DCU 1051 and 80.24% against the PARC 700 Dependency Bank (King et al., 2003), performing at the same or a slightly better level than state-of-the-art hand-crafted grammars (Kaplan et al., 2004). Aoife Cahill, Michael Burke, Ruth O'Donovan, Josef van Genabith, Andy Way |
ACL | 5 |
| 2004 | Large-Scale Induction and Evaluation of Lexical Resources from the Penn-II TreebankabstractIn this paper we present a methodology for extracting subcategorisation frames based on an automatic LFG f-structure annotation algorithm for the Penn-II Treebank. We extract abstract syntactic function-based subcategorisation frames (LFG semantic forms), traditional CFG category-based subcategorisation frames as well as mixed function/category-based frames, with or without preposition information for obliques and particle information for particle verbs. Our approach does not predefine frames, associates probabilities with frames conditional on the lemma, distinguishes between active and passive frames, and fully reflects the effects of long-distance dependencies in the source data structures. We extract 3586 verb lemmas, 14348 semantic form types (an average of 4 per lemma) with 577 frame types. We present a large-scale evaluation of the complete set of forms extracted against the full COMLEX resource. Ruth O'Donovan, Michael Burke, Aoife Cahill, Josef van Genabith, Andy Way |
ACL | 5 |
| 2004 | Robust Sub-Sentential Alignment of Phrase-Structure Trees
Declan Groves, Mary Hearne, Andy Way |
COLING | 3 |
| 2004 | Treebank-Based Acquisition of a Chinese Lexical-Functional Grammar
Michael Burke, Olivia S.-C. Lam, Aoife Cahill, Rowena Chan, Ruth O'Donovan, Adams Bodomo, Josef van Genabith, Andy Way |
PACLIC | 8 |
| 2003 | Controlled generation in example-based machine translationabstractThe theme of controlled translation is currently in vogue in the area of MT. Recent research (Scha ̈ler et al., 2003; Carl, 2003) hypothesises that EBMT systems are perhaps best suited to this challenging task. In this paper, we present an EBMT system where the generation of the target string is filtered by data written according to controlled language specifications. As far as we are aware, this is the only research available on this topic. In the field of controlled language applications, it is more usual to constrain the source language in this way rather than the target. We translate a small corpus of controlled English into French using the on-line MT system Logomedia, and seed the memories of our EBMT system with a set of automatically induced lexical resources using the Marker Hypothesis as a segmentation tool. We test our system on a large set of sentences extracted from a Sun Translation Memory, and provide both an automatic and a human evaluation. For comparative purposes, we also provide results for Logomedia itself. Nano Gough, Andy Way |
MTSummit | 2 |
| 2003 | Seeing the wood for the trees: data-oriented translationabstractData-Oriented Translation (DOT), which is based on Data-Oriented Parsing (DOP), comprises an experience-based approach to translation, where new translations are derived with reference to grammatical analyses of previous translations. Previous DOT experiments [Poutsma, 1998, Poutsma, 2000a, Poutsma, 2000b] were small in scale because important advances in DOP technology were not incorporated into the translation model. Despite this, related work [Way, 1999, Way, 2003a, Way, 2003b] reports that DOT models are viable in that solutions to ‘hard’ translation cases are readily available. However, it has not been shown to date that DOT models scale to larger datasets. In this work, we describe a novel DOT system, inspired by recent advances in DOP parsing technology. We test our system on larger, more complex corpora than have been used heretofore, and present both automatic and human evaluations which show that high quality translations can be achieved at reasonable speeds. Mary Hearne, Andy Way |
MTSummit | 2 |
| 2003 | wEBMT: Developing and Validating an Example-Based Machine Translation System using the World Wide WebabstractWe have developed an example-based machine translation (EBMT) system that uses the World Wide Web for two different purposes: First, we populate the system's memory with translations gathered from rule-based MT systems located on the Web. The source strings input to these systems were extracted automatically from an extremely small subset of the rule types in the Penn-II Treebank. In subsequent stages, the source, target translation pairs obtained are automatically transformed into a series of resources that render the translation process more successful. Despite the fact that the output from on-line MT systems is often faulty, we demonstrate in a number of experiments that when used to seed the memories of an EBMT system, they can in fact prove useful in generating translations of high quality in a robust fashion. In addition, we demonstrate the relative gain of EBMT in comparison to on-line systems. Second, despite the perception that the documents available on the Web are of questionable quality, we demonstrate in contrast that such resources are extremely useful in automatically postediting translation candidates proposed by our system. Andy Way, Nano Gough |
Comput. Linguistics | 1 |
| 1999 | A hybrid architecture for robust MT using LFG-DOP
Andy Way |
J. Exp. Theor. Artif. Intell. | 1 |