Josef van Genabith

dblp:82/3447 · DBLP profile ↗
← Back
122ranked-venue papers
2as first author
24since 2021 · last 2026
0000-0003-1322-7944ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 120 · 2 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1Theory of computation · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 Why Does Reinforcement Learning Generalize? A Feature-Level Mechanistic Study of Post-Training in Large Language Models
abstract
Dan Shi, Zhuowen Han, Simon Ostermann, Renren Jin, Josef Van Genabith, Deyi Xiong. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Dan Shi 0001, Zhuowen Han, Simon Ostermann 0002, Renren Jin, Josef van Genabith, Deyi Xiong
ACL (1)5
2026 PETra: A Multilingual Corpus of Pragmatic Explicitation in Translation
abstract
Translators often enrich texts with background details that make implicit cultural meanings explicit for new audiences. This phenomenon, known as pragmatic explicitation, has been widely discussed in translation theory but rarely modeled computationally. We introduce PragExTra, the first multilingual corpus and detection framework for pragmatic explicitation. The corpus covers eight language pairs from TED-Multi and Europarl and includes additions such as entity descriptions, measurement conversions, and translator remarks. We identify candidate explicitation cases through null alignments and refined using active learning with human annotation. Our results show that entity and system-level explicitations are most frequent, and that active learning improves classifier accuracy by 7-8 percentage points, achieving up to 0.88 accuracy and 0.82 F1 across languages. PragExTra establishes pragmatic explicitation as a measurable, cross-linguistic phenomenon and takes a step towards building culturally aware machine translation. Keywords: translation, multilingualism, explicitation
Doreen Osmelak, Koel Dutta Chowdhury, Uliana Sentsova, Cristina España-Bonet, Josef van Genabith
LREC5
2026 Dialectal Filtering: Synthesizing Kurdish Corpora for Low-Resource Varieties by Utilizing "Noise" in Large Textual Data
Christian Schuler, Raman Ahmad, Anrán Wáng, Daniil Gurgurov, Timo Baumann, Simon Ostermann 0002, Josef van Genabith
LREC7
2026 A Critical Study of Automatic Evaluation in Sign Language Translation
Shakib Yazdani, Yasser Hamidullah, Cristina España-Bonet, Eleftherios Avramidis, Josef van Genabith
LREC5
2025 Continual Learning in Multilingual Sign Language Translation
abstract
Shakib Yazdani, Josef Van Genabith, Cristina España-Bonet. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Shakib Yazdani, Josef van Genabith, Cristina España-Bonet
NAACL (Long Papers)2
2025 Very-Long-Distance Dependency Capturing Evaluation via Language Modeling Based on Gender Consistency
Hongfei Xu, Zhuofei Liang, Josef van Genabith, Deyi Xiong, Hongying Zan, Qiuhui Liu, Tengxun Zhang
NLPCC (4)3
2024 When Your Cousin Has the Right Connections: Unsupervised Bilingual Lexicon Induction for Related Data-Imbalanced Languages
abstract
Most existing approaches for unsupervised bilingual lexicon induction (BLI) depend on good quality static or contextual embeddings requiring large monolingual corpora for both languages. However, unsupervised BLI is most likely to be useful for low-resource languages (LRLs), where large datasets are not available. Often we are interested in building bilingual resources for LRLs against related high-resource languages (HRLs), resulting in severely imbalanced data settings for BLI. We first show that state-of-the-art BLI methods in the literature exhibit near-zero performance for severely data-imbalanced language pairs, indicating that these settings require more robust techniques. We then present a new method for unsupervised BLI between a related LRL and HRL that only requires inference on a masked language model of the HRL, and demonstrate its effectiveness on truly low-resource languages Bhojpuri and Magahi (with <5M monolingual tokens each), against Hindi. We further present experiments on (mid-resource) Marathi and Nepali to compare approach performances by resource range, and release our resulting lexicons for five low-resource Indic languages: Bhojpuri, Magahi, Awadhi, Braj, and Maithili, against Hindi.
Niyati Bafna, Cristina España-Bonet, Josef van Genabith, Benoît Sagot, Rachel Bawden
LREC/COLING3
2024 Rewiring the Transformer with Depth-Wise LSTMs
abstract
Stacking non-linear layers allows deep neural networks to model complicated functions, and including residual connections in Transformer layers is beneficial for convergence and performance. However, residual connections may make the model “forget” distant layers and fail to fuse information from previous layers effectively. Selectively managing the representation aggregation of Transformer layers may lead to better performance. In this paper, we present a Transformer with depth-wise LSTMs connecting cascading Transformer layers and sub-layers. We show that layer normalization and feed-forward computation within a Transformer layer can be absorbed into depth-wise LSTMs connecting pure Transformer attention layers. Our experiments with the 6-layer Transformer show significant BLEU improvements in both WMT 14 English-German / French tasks and the OPUS-100 many-to-many multilingual NMT task, and our deep Transformer experiments demonstrate the effectiveness of depth-wise LSTM on the convergence and performance of deep Transformers.
Hongfei Xu, Qiuhui Liu, Josef van Genabith, Deyi Xiong
LREC/COLING4
2024 Mitigating Translationese with GPT-4: Strategies and Performance
abstract
Translations differ in systematic ways from texts originally authored in the same language.These differences, collectively known as translationese, can pose challenges in cross-lingual natural language processing: models trained or tested on translated input might struggle when presented with non-translated language. Translationese mitigation can alleviate this problem. This study investigates the generative capacities of GPT-4 to reduce translationese in human-translated texts. The task is framed as a rewriting process aimed at modified translations indistinguishable from the original text in the target language. Our focus is on prompt engineering that tests the utility of linguistic knowledge as part of the instruction for GPT-4. Through a series of prompt design experiments, we show that GPT4-generated revisions are more similar to originals in the target language when the prompts incorporate specific linguistic instructions instead of relying solely on the model’s internal knowledge. Furthermore, we release the segment-aligned bidirectional German-English data built from the Europarl corpus that underpins this study.
Maria Kunilovskaya, Koel Dutta Chowdhury, Heike Przybyl, Cristina España-Bonet, Josef van Genabith
EAMT (1)5
2023 Rescuespeech: A German Corpus for Speech Recognition in Search and Rescue Domain
abstract
Despite the recent advancements in speech recognition, there are still difficulties in accurately transcribing conversational and emotional speech in noisy and reverberant acoustic environments. This poses a particular challenge in the search and rescue (SAR) domain, where transcribing conversations among rescue team members is crucial to support real-time decision-making. The scarcity of speech data and associated background noise in SAR scenarios make it difficult to deploy robust speech recognition systems.To address this issue, we have created and made publicly available a German speech dataset called RescueSpeech. This dataset includes real speech recordings from simulated rescue exercises. Additionally, we have released competitive training recipes and pre-trained models. Our study highlights that the performance attained by state-of-the-art methods in this challenging scenario is still far from reaching an acceptable level.
Sangeet Sagar, Mirco Ravanelli, Bernd Kiefer, Ivana Kruijff-Korbayová, Josef van Genabith
ASRU5
2023 Exploring Paracrawl for Document-level Neural Machine Translation
abstract
Document-level neural machine translation (NMT) has outperformed sentence-level NMT on a number of datasets.However, documentlevel NMT is still not widely adopted in realworld translation systems mainly due to the lack of large-scale general-domain training data for document-level NMT.We examine the effectiveness of using Paracrawl for learning document-level translation.Paracrawl is a large-scale parallel corpus crawled from the Internet and contains data from various domains.The official Paracrawl corpus was released as parallel sentences (extracted from parallel webpages) and therefore previous works only used Paracrawl for learning sentence-level translation.In this work, we extract parallel paragraphs from Paracrawl parallel webpages using automatic sentence alignments and we use the extracted parallel paragraphs as parallel documents for training document-level translation models.We show that document-level NMT models trained with only parallel paragraphs from Paracrawl can be used to translate real documents from TED, News and Europarl, outperforming sentence-level NMT models.We also perform a targeted pronoun evaluation and show that document-level models trained with Paracrawl data can help context-aware pronoun translation.We release our data and code here 1 .
Yusser Al Ghussin, Jingyi Zhang 0002, Josef van Genabith
EACL3
2023 Translating away Translationese without Parallel Data
abstract
Translated texts exhibit systematic linguistic differences compared to original texts in the same language, and these differences are referred to as translationese.Translationese has effects on various cross-lingual natural language processing tasks, potentially leading to biased results.In this paper, we explore a novel approach to reduce translationese in translated texts: translation-based style transfer.As there are no parallel human-translated and original data in the same language, we use a selfsupervised approach that can learn from comparable (rather than parallel) mono-lingual original and translated data.However, even this self-supervised approach requires some parallel data for validation.We show how we can eliminate the need for parallel validation data by combining the self-supervised loss with an unsupervised loss.This unsupervised loss leverages the original language model loss over the style-transferred output and a semantic similarity loss between the input and style-transferred output.We evaluate our approach in terms of original vs. translationese binary classification in addition to measuring content preservation and target-style fluency.The results show that our approach is able to reduce translationese classifier accuracy to a level of a random classifier after style transfer while adequately preserving the content and fluency in the target original style.
Rricha Jalota, Koel Dutta Chowdhury, Cristina España-Bonet, Josef van Genabith
EMNLP4
2023 Find-2-Find: Multitask Learning for Anaphora Resolution and Object Localization
abstract
In multimodal understanding tasks, visual and linguistic ambiguities can arise.Visual ambiguity can occur when visual objects require a model to ground a referring expression in a video without strong supervision, while linguistic ambiguity can occur from changes in entities in action flows.As an example from the cooking domain, "oil" mixed with "salt" and "pepper" could later be referred to as a "mixture".Without a clear visual-linguistic alignment, we cannot know which among several objects shown is referred to by the language expression "mixture", and without resolved antecedents, we cannot pinpoint what the mixture is.We define this chicken-and-egg problem as visual-linguistic ambiguity.In this paper, we present Find2Find, a joint anaphora resolution and object localization dataset targeting the problem of visual-linguistic ambiguity, consisting of 500 anaphora-annotated recipes with corresponding videos.We present experimental results of a novel end-to-end joint multitask learning framework for Find2Find that fuses visual and textual information and shows improvements both for anaphora resolution and object localization as compared to a strong single-task baseline.
Cennet Oguz, Pascal Denis, Emmanuel Vincent 0001, Simon Ostermann 0002, Josef van Genabith
EMNLP5
2023 NAPG: Non-Autoregressive Program Generation for Hybrid Tabular-Textual Question Answering
Tengxun Zhang, Hongfei Xu, Josef van Genabith, Deyi Xiong, Hongying Zan
NLPCC (1)3
2023 JoinER-BART: Joint Entity and Relation Extraction With Constrained Decoding, Representation Reuse and Fusion
abstract
Joint Entity and Relation Extraction (JERE) is an important research direction in Information Extraction (IE). Given the surprising performance with fine-tuning of pre-trained BERT in a wide range of NLP tasks, nowadays most studies for JERE are based on the BERT model. Rather than predicting a simple tag for each word, these approaches are usually forced to design complex tagging schemes, as they may have to extract entity-relation pairs which may overlap with others from the same sequence of word representations in a sentence. Recently, sequence-to-sequence (seq2seq) pre-trained BART models show better performance than BERT models in many NLP tasks. Importantly, a seq2seq BART model can simply generate sequences of (many) entity-relation triplets with its decoder, rather than just tag input words. In this paper, we present a new generative JERE framework based on pre-trained BART. Different from the basic seq2seq BART architecture: 1) our framework employs a constrained classifier which only predicts either a token of the input sentence or a relation in each decoding step, and 2) we reuse representations from the pre-trained BART encoder in the classifier instead of a newly trained weight matrix, as this better utilizes the knowledge of the pre-trained model and context-aware representations for classification, and empirically leads to better performance. In our experiments on the widely studied NYT and WebNLG datasets, we show that our approach outperforms previous studies and establishes a new state-of-the-art (92.91 and 91.37 F1 respectively in exact match evaluation).
Hongyang Chang, Hongfei Xu, Josef van Genabith, Deyi Xiong, Hongying Zan
IEEE ACM Trans. Audio Speech Lang. Process.3
2022 Combining Noisy Semantic Signals with Orthographic Cues: Cognate Induction for the Indic Dialect Continuum
abstract
We present a novel method for unsupervised cognate/borrowing identification from monolingual corpora designed for low and extremely low resource scenarios, based on combining noisy semantic signals from joint bilingual spaces with orthographic cues modelling sound change.We apply our method to the North Indian dialect continuum, containing several dozens of dialects and languages spoken by more than 100 million people.Many of these languages are zero-resource and therefore natural language processing for them is nonexistent.We first collect monolingual data for 26 Indic languages, 16 of which were previously zero-resource, and perform exploratory character, lexical and subword cross-lingual alignment experiments for the first time at this scale on this dialect continuum.We create bilingual evaluation lexicons against Hindi for 20 of the languages.We then apply our cognate identification method on the data, and show that our method outperforms both traditional orthography baselines as well as EM-style learnt edit distance matrices.To the best of our knowledge, this is the first work to combine traditional orthographic cues with noisy bilingual embeddings to tackle unsupervised cognate detection in a (truly) low-resource setup, showing that even noisy bilingual embeddings can act as good guides for this task.We release our multilingual dialect corpus, called HinDialect, as well as our scripts for evaluation data collection and cognate induction.2
Niyati Bafna, Josef van Genabith, Cristina España-Bonet, Zdenek Zabokrtský
CoNLL2
2022 Towards Debiasing Translation Artifacts
abstract
Koel Dutta Chowdhury, Rricha Jalota, Cristina España-Bonet, Josef Genabith. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Koel Dutta Chowdhury, Rricha Jalota, Cristina España-Bonet, Josef van Genabith
NAACL-HLT4
2021 Mid-Air Hand Gestures for Post-Editing of Machine Translation
abstract
Rashad Albo Jamara, Nico Herbig, Antonio Krüger, Josef van Genabith. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Rashad Albo Jamara, Nico Herbig 0001, Antonio Krüger, Josef van Genabith
ACL/IJCNLP (1)4
2021 Multi-Head Highly Parallelized LSTM Decoder for Neural Machine Translation
abstract
Hongfei Xu, Qiuhui Liu, Josef van Genabith, Deyi Xiong, Meng Zhang. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Hongfei Xu, Qiuhui Liu, Josef van Genabith, Deyi Xiong, Meng Zhang 0019
ACL/IJCNLP (1)3
2021 A Bidirectional Transformer Based Alignment Model for Unsupervised Word Alignment
abstract
Jingyi Zhang, Josef van Genabith. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Jingyi Zhang 0002, Josef van Genabith
ACL/IJCNLP (1)2
2021 Comparing Feature-Engineering and Feature-Learning Approaches for Multilingual Translationese Classification
abstract
Daria Pylypenko, Kwabena Amponsah-Kaakyire, Koel Dutta Chowdhury, Josef van Genabith, Cristina España-Bonet. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021.
Daria Pylypenko, Kwabena Amponsah-Kaakyire, Koel Dutta Chowdhury, Josef van Genabith, Cristina España-Bonet
EMNLP (1)4
2021 Investigating the Helpfulness of Word-Level Quality Estimation for Post-Editing Machine Translation Output
abstract
Compared to fully manual translation, postediting (PE) machine translation (MT) output can save time and reduce errors.Automatic word-level quality estimation (QE) aims to predict the correctness of words in MT output and holds great promise to aid PE by flagging problematic output.Quality of QE is crucial, as incorrect QE might lead to translators missing errors or wasting time on already correct MT output.Achieving accurate automatic word-level QE is very hard, and it is currently not known (i) at what quality threshold QE is actually beginning to be useful for human PE, and (ii), how to best present wordlevel QE information to translators.In particular, should word-level QE visualization indicate uncertainty of the QE model or not?In this paper, we address both research questions with real and simulated word-level QE, visualizations, and user studies, where time, subjective ratings, and quality of the final translations are assessed.Results show that current wordlevel QE models are not yet good enough to support PE. Instead, quality levels of ≥ 80% F1 are required.For helpful quality levels, a visualization reflecting the uncertainty of the QE model is preferred.Our analysis further shows that speed gains achieved through QE are not merely a result of blindly trusting the QE system, but that the quality of the final translations also improves.The threshold results from the paper establish a quality goal for future wordlevel QE research.
Raksha Shenoy, Nico Herbig 0001, Antonio Krüger, Josef van Genabith
EMNLP (1)4
2021 Integrating Unsupervised Data Generation into Self-Supervised Neural Machine Translation for Low-Resource Languages
abstract
For most language combinations and parallel data is either scarce or simply unavailable. To address this and unsupervised machine translation (UMT) exploits large amounts of monolingual data by using synthetic data generation techniques such as back-translation and noising and while self-supervised NMT (SSNMT) identifies parallel sentences in smaller comparable data and trains on them. To this date and the inclusion of UMT data generation techniques in SSNMT has not been investigated. We show that including UMT techniques into SSNMT significantly outperforms SSNMT (up to +4.3 BLEU and af2en) as well as statistical (+50.8 BLEU) and hybrid UMT (+51.5 BLEU) baselines on related and distantly-related and unrelated language pairs.
Dana Ruiter, Dietrich Klakow, Josef van Genabith, Cristina España-Bonet
MTSummit (1)3
2021 Probing Word Translations in the Transformer and Trading Decoder for Encoder Layers
abstract
Hongfei Xu, Josef van Genabith, Qiuhui Liu, Deyi Xiong. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Hongfei Xu, Josef van Genabith, Qiuhui Liu, Deyi Xiong
NAACL-HLT2
2020 MMPE: A Multi-Modal Interface for Post-Editing Machine Translation
abstract
Nico Herbig, Tim Düwel, Santanu Pal, Kalliopi Meladaki, Mahsa Monshizadeh, Antonio Krüger, Josef van Genabith. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020.
Nico Herbig 0001, Tim Düwel, Santanu Pal, Kalliopi Meladaki, Mahsa Monshizadeh, Antonio Krüger, Josef van Genabith
ACL7
2020 Dynamically Adjusting Transformer Batch Size by Monitoring Gradient Direction Change
abstract
The choice of hyper-parameters affects the performance of neural models.While much previous research (Sutskever et al., 2013;Duchi et al., 2011;Kingma and Ba, 2015) focuses on accelerating convergence and reducing the effects of the learning rate, comparatively few papers concentrate on the effect of batch size.In this paper, we analyze how increasing batch size affects gradient direction, and propose to evaluate the stability of gradients with their angle change.Based on our observations, the angle change of gradient direction first tends to stabilize (i.e.gradually decrease) while accumulating mini-batches, and then starts to fluctuate.We propose to automatically and dynamically determine batch sizes by accumulating gradients of mini-batches and performing an optimization step at just the time when the direction of gradients starts to fluctuate.To improve the efficiency of our approach for large models, we propose a sampling approach to select gradients of parameters sensitive to the batch size.Our approach dynamically determines proper and efficient batch sizes during training.In our experiments on the WMT 14 English to German and English to French tasks, our approach improves the Transformer with a fixed 25k batch size by +0.73 and +0.82 BLEU respectively.
Hongfei Xu, Josef van Genabith, Deyi Xiong, Qiuhui Liu
ACL2
2020 Learning Source Phrase Representations for Neural Machine Translation
abstract
The Transformer translation model (Vaswani et al., 2017) based on a multi-head attention mechanism can be computed effectively in parallel and has significantly pushed forward the performance of Neural Machine Translation (NMT).Though intuitively the attentional network can connect distant words via shorter network paths than RNNs, empirical analysis demonstrates that it still has difficulty in fully capturing long-distance dependencies (Tang et al., 2018).Considering that modeling phrases instead of words has significantly improved the Statistical Machine Translation (SMT) approach through the use of larger translation blocks ("phrases") and its reordering ability, modeling NMT at phrase level is an intuitive proposal to help the model capture long-distance relationships.In this paper, we first propose an attentive phrase representation generation mechanism which is able to generate phrase representations from corresponding token representations.In addition, we incorporate the generated phrase representations into the Transformer translation model to enhance its ability to capture long-distance relationships.In our experiments, we obtain significant improvements on the WMT 14 English-German and English-French tasks on top of the strong Transformer baseline, which shows the effectiveness of our approach.Our approach helps Transformer Base models perform at the level of Transformer Big models, and even significantly better for long sentences, but with substantially fewer parameters and training steps.The fact that phrase representations help even in the big setting further supports our conjecture that they make a valuable contribution to long-distance relations.
Hongfei Xu, Josef van Genabith, Deyi Xiong, Qiuhui Liu, Jingyi Zhang 0002
ACL2
2020 Lipschitz Constrained Parameter Initialization for Deep Transformers
abstract
The Transformer translation model employs residual connection and layer normalization to ease the optimization difficulties caused by its multi-layer encoder/decoder structure. Previous research shows that even with residual connection and layer normalization, deep Transformers still have difficulty in training, and particularly Transformer models with more than 12 encoder/decoder layers fail to converge. In this paper, we first empirically demonstrate that a simple modification made in the official implementation, which changes the computation order of residual connection and layer normalization, can significantly ease the optimization of deep Transformers. We then compare the subtle differences in computation order in considerable detail, and present a parameter initialization method that leverages the Lipschitz constraint on the initialization of Transformer parameters that effectively ensures training convergence. In contrast to findings in previous research we further demonstrate that with Lipschitz parameter initialization, deep Transformers with the original computation order can converge, and obtain significant BLEU improvements with up to 24 layers. In contrast to previous research which focuses on deep encoders, our approach additionally enables Transformers to also benefit from deep decoders.
Hongfei Xu, Qiuhui Liu, Josef van Genabith, Deyi Xiong, Jingyi Zhang 0002
ACL3
2020 Understanding Translationese in Multi-view Embedding Spaces
abstract
Recent studies use a combination of lexical and syntactic features to show that footprints of the source language remain visible in translations, to the extent that it is possible to predict the original source language from the translation.In this paper, we focus on embedding-based semantic spaces, exploiting departures from isomorphism between spaces built from original target language and translations into this target language to predict relations between languages in an unsupervised way.We use different views of the data -words, parts of speech, semantic tags and synsets -to track translationese.Our analysis shows that (i) semantic distances between original target language and translations into this target language can be detected using the notion of isomorphism, (ii) language family ties with characteristics similar to linguistically motivated phylogenetic trees can be inferred from the distances and (iii) with delexicalised embeddings exhibiting source-language interference most significantly, other levels of abstraction display the same tendency, indicating the lexicalised results to be not "just" due to possible topic differences between original and translated texts.To the best of our knowledge, this is the first time departures from isomorphism between embedding spaces are used to track translationese.
Koel Dutta Chowdhury, Cristina España-Bonet, Josef van Genabith
COLING3
2020 The Transference Architecture for Automatic Post-Editing
abstract
In automatic post-editing (APE) it makes sense to condition post-editing (pe) decisions on both the source (src) and the machine translated text (mt) as input.This has led to multi-encoder based neural APE approaches.A research challenge now is the search for architectures that best support the capture, preparation and provision of src and mt information and its integration with pe decisions.In this paper we present an efficient multi-encoder based APE model, called transference.Unlike previous approaches, it (i) uses a transformer encoder block for src, (ii) followed by a decoder block, but without masking for self-attention on mt, which effectively acts as second encoder combining src → mt, and (iii) feeds this representation into a final decoder block generating pe.Our model outperforms the best performing systems by 1 BLEU point on the WMT 2016, 2017, and 2018 English-German APE shared tasks (PBSMT and NMT).Furthermore, the results of our model on the WMT 2019 APE task using NMT data shows performance at the level of the state-of-the-art.The inference time of our model is similar to the vanilla transformer-based NMT system although our model deals with two separate encoders.We further investigate the importance of our newly introduced second encoder and find that decreasing the number of layers hurts performance, while reducing the number of layers of the decoder does not matter much.
Santanu Pal, Hongfei Xu, Nico Herbig 0001, Sudip Kumar Naskar, Antonio Krüger, Josef van Genabith
COLING6
2020 Self-Induced Curriculum Learning in Self-Supervised Neural Machine Translation
abstract
Self-supervised neural machine translation (SSNMT) jointly learns to identify and select suitable training data from comparable (rather than parallel) corpora and to translate, in a way that the two tasks support each other in a virtuous circle.In this study, we provide an in-depth analysis of the sampling choices the SSNMT model makes during training.We show how, without it having been told to do so, the model self-selects samples of increasing (i) complexity and (ii) task-relevance in combination with (iii) performing a denoising curriculum.We observe that the dynamics of the mutual-supervision signals of both system internal representation types are vital for the extraction and translation performance.We show that in terms of the Gunning-Fog Readability index, SSNMT starts extracting and learning from Wikipedia data suitable for high school students and quickly moves towards content suitable for first year undergraduate students.
Dana Ruiter, Josef van Genabith, Cristina España-Bonet
EMNLP (1)2
2020 Translation Quality Estimation by Jointly Learning to Score and Rank
abstract
The translation quality estimation (QE) task, particularly the QE as a Metric task, aims to evaluate the general quality of a translation based on the translation and the source sentence without using reference translations.Supervised learning of this QE task requires human evaluation of translation quality as training data.Human evaluation of translation quality can be performed in different ways, including assigning an absolute score to a translation or ranking different translations.In order to make use of different types of human evaluation data for supervised learning, we present a multi-task learning QE model that jointly learns two tasks: score a translation and rank two translations.Our QE model exploits crosslingual sentence embeddings from pretrained multilingual language models.We obtain new state-of-the-art results on the WMT 2019 QE as a Metric task and outperform sentBLEU on the WMT 2019 Metrics task.
Jingyi Zhang 0002, Josef van Genabith
EMNLP (1)2
2020 Efficient Context-Aware Neural Machine Translation with Layer-Wise Weighting and Input-Aware Gating
abstract
Existing Neural Machine Translation (NMT) systems are generally trained on a large amount of sentence-level parallel data, and during prediction sentences are independently translated, ignoring cross-sentence contextual information. This leads to inconsistency between translated sentences. In order to address this issue, context-aware models have been proposed. However, document-level parallel data constitutes only a small part of the parallel data available, and many approaches build context-aware models based on a pre-trained frozen sentence-level translation model in a two-step training manner. The computational cost of these approaches is usually high. In this paper, we propose to make the most of layers pre-trained on sentence-level data in contextual representation learning, reusing representations from the sentence-level Transformer and significantly reducing the cost of incorporating contexts in translation. We find that representations from shallow layers of a pre-trained sentence-level encoder play a vital role in source context encoding, and propose to perform source context encoding upon weighted combinations of pre-trained encoder layers' outputs. Instead of separately performing source context and input encoding, we propose to iteratively and jointly encode the source input and its contexts and to generate input-aware context representations with a cross-attention layer and a gating mechanism, which resets irrelevant information in context encoding. Our context-aware Transformer model outperforms the recent CADec [Voita et al., 2019c] on the English-Russian subtitle data and is about twice as fast in training and decoding.
Hongfei Xu, Deyi Xiong, Josef van Genabith, Qiuhui Liu
IJCAI3
2020 The European Language Technology Landscape in 2020: Language-Centric and Human-Centric AI for Cross-Cultural Communication in Multilingual Europe
abstract
Multilingualism is a cultural cornerstone of Europe and firmly anchored in the European treaties including full language equality. However, language barriers impacting business, cross-lingual and cross-cultural communication are still omnipresent. Language Technologies (LTs) are a powerful means to break down these barriers. While the last decade has seen various initiatives that created a multitude of approaches and technologies tailored to Europe’s specific needs, there is still an immense level of fragmentation. At the same time, AI has become an increasingly important concept in the European Information and Communication Technology area. For a few years now, AI – including many opportunities, synergies but also misconceptions – has been overshadowing every other topic. We present an overview of the European LT landscape, describing funding programmes, activities, actions and challenges in the different countries with regard to LT, including the current state of play in industry and the LT market. We present a brief overview of the main LT-related activities on the EU level in the last ten years and develop strategic guidance with regard to four key dimensions.
Georg Rehm, Katrin Marheinecke, Stefanie Hegele, Stelios Piperidis, Kalina Bontcheva, Jan Hajic 0001, Khalid Choukri, Andrejs Vasiljevs, Gerhard Backfried, Christoph Prinz, José Manuél Gómez-Pérez, Luc Meertens, Paul Lukowicz, Josef van Genabith, Andrea Lösch, Philipp Slusallek, Morten Irgens, Patrick Gatellier, Joachim Köhler, Laure Le Bars, Dimitra Anastasiou, Albina Auksoriute, Núria Bel, António Branco, Gerhard Budin, Walter Daelemans, Koenraad De Smedt, Radovan Garabík, Maria Gavrilidou, Dagmar Gromann, Svetla Koeva, Simon Krek, Cvetana Krstev, Krister Lindén, Bernardo Magnini, Jan Odijk, Maciej Ogrodniczuk, Eiríkur Rögnvaldsson, Mike Rosner, Bolette S. Pedersen, Inguna Skadina, Marko Tadic, Dan Tufis, Tamás Váradi, Kadri Vider, Andy Way, François Yvon
LREC14
2020 Language Data Sharing in European Public Services - Overcoming Obstacles and Creating Sustainable Data Sharing Infrastructures
abstract
Data is key in training modern language technologies. In this paper, we summarise the findings of the first pan-European study on obstacles to sharing language data across 29 EU Member States and CEF-affiliated countries carried out under the ELRC White Paper action on Sustainable Language Data Sharing to Support Language Equality in Multilingual Europe. Why Language Data Matters. We present the methodology of the study, the obstacles identified and report on recommendations on how to overcome those. The obstacles are classified into (1) lack of appreciation of the value of language data, (2) structural challenges, (3) disposition towards CAT tools and lack of digital skills, (4) inadequate language data management practices, (5) limited access to outsourced translations, and (6) legal concerns. Recommendations are grouped into addressing the European/national policy level, and the organisational/institutional level.
Lilli Smal, Andrea Lösch, Josef van Genabith, Maria Giagkou, Thierry Declerck, Stephan Busemann
LREC3
2019 Self-Supervised Neural Machine Translation
abstract
We present a simple new method where an emergent NMT system is used for simultaneously selecting training data and learning internal NMT representations.This is done in a self-supervised way without parallel data, in such a way that both tasks enhance each other during training.The method is language independent, introduces no additional hyper-parameters, and achieves BLEU scores of 29.21 (en2f r) and 27.36 (f r2en) on new-stest2014 using English and French Wikipedia data for training.
Dana Ruiter, Cristina España-Bonet, Josef van Genabith
ACL (1)3
2019 Multi-Modal Approaches for Post-Editing Machine Translation
abstract
Current advances in machine translation increase the need for translators to switch from traditional translation to post-editing (PE) of machine-translated text, a process that saves time and improves quality. This affects the design of translation interfaces, as the task changes from mainly generating text to correcting errors within otherwise helpful translation proposals. Our results of an elicitation study with professional translators indicate that a combination of pen, touch, and speech could well support common PE tasks, and received high subjective ratings by our participants. Therefore, we argue that future translation environment research should focus more strongly on these modalities in addition to mouse- and keyboard-based approaches. On the other hand, eye tracking and gesture modalities seem less important. An additional interview regarding interface design revealed that most translators would also see value in automatically receiving additional resources when a high cognitive load is detected during PE.
Nico Herbig 0001, Santanu Pal, Josef van Genabith, Antonio Krüger
CHI3
2019 Improving CAT Tools in the Translation Workflow: New Approaches and Evaluation
Mihaela Vela, Santanu Pal, Marcos Zampieri, Sudip Kumar Naskar, Josef van Genabith
MTSummit (2)5
2019 Multi-modal indicators for estimating perceived cognitive load in post-editing of machine translation
Nico Herbig 0001, Santanu Pal, Mihaela Vela, Antonio Krüger, Josef van Genabith
Mach. Transl.5
2018 European Language Resource Coordination: Collecting Language Resources for Public Sector Multilingual Information Management
Andrea Lösch, Valérie Mapelli, Stelios Piperidis, Andrejs Vasiljevs, Lilli Smal, Thierry Declerck, Eileen Schnur, Khalid Choukri, Josef van Genabith
LREC9
2018 Neural machine translation for low-resource languages without parallel corpora
abstract
The problem of a total absence of parallel data is present for a large number of language pairs and can severely detriment the quality of machine translation. We describe a language-independent method to enable machine translation between a low-resource language (LRL) and a third language, e.g. English. We deal with cases of LRLs for which there is no readily available parallel data between the low-resource language and any other language, but there is ample training data between a closely-related high-resource language (HRL) and the third language. We take advantage of the similarities between the HRL and the LRL in order to transform the HRL data into data similar to the LRL using transliteration. The transliteration models are trained on transliteration pairs extracted from Wikipedia article titles. Then, we automatically back-translate monolingual LRL data with the models trained on the transliterated HRL data and use the resulting parallel corpus to train our final models. Our method achieves significant improvements in translation quality, close to the results that can be achieved by a general purpose neural machine translation system trained on a significant amount of parallel data. Moreover, the method does not rely on the existence of any parallel data for training, but attempts to bootstrap already existing resources in a related language.
Alina Karakanta, Jon Dehdari, Josef van Genabith
Mach. Transl.3
2017 An Extensive Empirical Evaluation of Character-Based Morphological Tagging for 14 Languages
abstract
This paper investigates neural characterbased morphological tagging for languages with complex morphology and large tag sets.Character-based approaches are attractive as they can handle rarelyand unseen words gracefully.We evaluate on 14 languages and observe consistent gains over a state-of-the-art morphological tagger across all languages except for English and French, where we match the state-of-the-art.We compare two architectures for computing characterbased word vectors using recurrent (RNN) and convolutional (CNN) nets.We show that the CNN based approach performs slightly worse and less consistently than the RNN based approach.Small but systematic gains are observed when combining the two architectures by ensembling.
Georg Heigold, Günter Neumann, Josef van Genabith
EACL (1)3
2016 Scaling character-based morphological tagging to fourteen languages
abstract
This paper investigates neural character-based morphological tagging for languages with complex morphology and large tag sets. Character-based approaches are attractive as they can handle rarely- and unseen words gracefully. More specifically, beside a rich morphology, non-canonical language, change of language or other linguistic variability can heavily degrade the accuracy of natural language processing of web and CMC data. We evaluate on 14 languages and observe consistent gains over a state-of-the-art morphological tagger across all languages except for English and French, where we match the state-of-the-art. The gains are clearly correlated with the amount of training data. We present supplementary experiments to explore whether and to what extent unsupervised data through pre-trained word vectors can compensate for limited amounts of supervised data. Moreover, we show preliminary results to study the effect of noisy input data by flipping characters at random.
Georg Heigold, Josef van Genabith, Günter Neumann
IEEE BigData2
2016 Forest to String Based Statistical Machine Translation with Hybrid Word Alignments
Santanu Pal, Sudip Kumar Naskar, Josef van Genabith
CICLing (2)3
2016 Multi-Engine and Multi-Alignment Based Automatic Post-Editing and its Impact on Translation Productivity
abstract
In this paper we combine two strands of machine translation (MT) research: automatic post-editing (APE) and multi-engine (system combination) MT. APE systems learn a target-language-side second stage MT system from the data produced by human corrected output of a first stage MT system, to improve the output of the first stage MT in what is essentially a sequential MT system combination architecture. At the same time, there is a rich research literature on parallel MT system combination where the same input is fed to multiple engines and the best output is selected or smaller sections of the outputs are combined to obtain improved translation output. In the paper we show that parallel system combination in the APE stage of a sequential MT-APE combination yields substantial translation improvements both measured in terms of automatic evaluation metrics as well as in terms of productivity improvements measured in a post-editing experiment. We also show that system combination on the level of APE alignments yields further improvements. Overall our APE system yields statistically significant improvement of 5.9% relative BLEU over a strong baseline (English–Italian Google MT) and 21.76% productivity increase in a human post-editing experiment with professional translators.
Santanu Pal, Sudip Kumar Naskar, Josef van Genabith
COLING3
2016 Modeling Diachronic Change in Scientific Writing with Information Density
abstract
Previous linguistic research on scientific writing has shown that language use in the scientific domain varies considerably in register and style over time. In this paper we investigate the introduction of information theory inspired features to study long term diachronic change on three levels: lexis, part-of-speech and syntax. Our approach is based on distinguishing between sentences from 19th and 20th century scientific abstracts using supervised classification models. To the best of our knowledge, the introduction of information theoretic features to this task is novel. We show that these features outperform more traditional features, such as token or character n-grams, while leading to more compact models. We present a detailed analysis of feature informativeness in order to gain a better understanding of diachronic change on different linguistic levels.
Raphaël Rubino, Stefania Degaetano-Ortlieb, Elke Teich, Josef van Genabith
COLING4
2016 CATaLog Online: Porting a Post-editing Tool to the Web
Santanu Pal, Marcos Zampieri, Sudip Kumar Naskar, Tapas Nayak, Mihaela Vela, Josef van Genabith
LREC6
2016 Fostering the Next Generation of European Language Technology: Recent Developments ― Emerging Initiatives ― Challenges and Opportunities
Georg Rehm, Jan Hajic 0001, Josef van Genabith, Andrejs Vasiljevs
LREC3
2016 BIRA: Improved Predictive Exchange Word Clustering
abstract
Word clusters are useful for many NLP tasks including training neural network language models, but current increases in datasets are outpacing the ability of word clusterers to handle them.Little attention has been paid thus far on inducing high-quality word clusters at a large scale.The predictive exchange algorithm is quite scalable, but sometimes does not provide as good perplexity as other slower clustering algorithms.We introduce the bidirectional, interpolated, refining, and alternating (BIRA) predictive exchange algorithm.It improves upon the predictive exchange algorithm's perplexity by up to 18%, giving it perplexities comparable to the slower two-sided exchange algorithm, and better perplexities than the slower Brown clustering algorithm.Our BIRA implementation is fast, clustering a 2.5 billion token English News Crawl corpus in 3 hours.It also reduces machine translation training time while preserving translation quality.Our implementation is portable and freely available.
Jon Dehdari, Liling Tan, Josef van Genabith
HLT-NAACL3
2016 Information Density and Quality Estimation Features as Translationese Indicators for Human Translation Classification
abstract
Raphael Rubino, Ekaterina Lapshinova-Koltunski, Josef van Genabith. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016.
Raphaël Rubino, Ekaterina Lapshinova-Koltunski, Josef van Genabith
HLT-NAACL3
2016 Improving translation memory matching and retrieval using paraphrases
Rohit Gupta 0007, Constantin Orasan, Marcos Zampieri, Mihaela Vela, Josef van Genabith, Ruslan Mitkov
Mach. Transl.5
2016 Arabic spelling error detection and correction
abstract
Abstract A spelling error detection and correction application is typically based on three main components: a dictionary (or reference word list), an error model and a language model. While most of the attention in the literature has been directed to the language model, we show how improvements in any of the three components can lead to significant cumulative improvements in the overall performance of the system. We develop our dictionary of 9.2 million fully-inflected Arabic words (types) from a morphological transducer and a large corpus, validated and manually revised. We improve the error model by analyzing error types and creating an edit distance re-ranker. We also improve the language model by analyzing the level of noise in different data sources and selecting an optimal subset to train the system on. Testing and evaluation experiments show that our system significantly outperforms Microsoft Word 2013, OpenOffice Ayaspell 3.4 and Google Docs.
Pavel Pecina, Younes Samih, Khaled Shaalan, Josef van Genabith
Nat. Lang. Eng.5
2015 Mining Parallel Resources for Machine Translation from Comparable Corpora
Santanu Pal, Partha Pakray, Alexander F. Gelbukh, Josef van Genabith
CICLing (1)4
2015 Can Translation Memories afford not to use paraphrasing?
Rohit Gupta 0007, Constantin Orasan, Marcos Zampieri, Mihaela Vela, Josef van Genabith
EAMT5
2015 Searching for Context: a Study on Document-Level Labels for Translation Quality Estimation
Carolina Scarton, Marcos Zampieri, Mihaela Vela, Josef van Genabith, Lucia Specia
EAMT4
2015 Re-assessing the WMT2013 Human Evaluation with Professional Translators Trainees
Mihaela Vela, Josef van Genabith
EAMT2
2015 ReVal: A Simple and Effective Machine Translation Evaluation Metric Based on Recurrent Neural Networks
abstract
Many state-of-the-art Machine Translation (MT) evaluation metrics are complex, involve extensive external resources (e.g. for paraphrasing) and require tuning to achieve best results.We present a simple alternative approach based on dense vector spaces and recurrent neural networks (RNNs), in particular Long Short Term Memory (LSTM) networks.For WMT-14, our new metric scores best for two out of five language pairs, and overall best and second best on all language pairs, using Spearman and Pearson correlation, respectively.We also show how training data is computed automatically from WMT ranks data.
Rohit Gupta 0007, Constantin Orasan, Josef van Genabith
EMNLP3
2015 Linguistically-augmented perplexity-based data selection for language models
abstract
This paper explores the use of linguistic information for the selection of data to train language models. We depart from the state-of-the-art method in perplexity-based data selection and extend it in order to use word-level linguistic units (i.e. lemmas, named entity categories and part-of-speech tags) instead of surface forms. We then present two methods that combine the different types of linguistic knowledge as well as the surface forms (1, naïve selection of the top ranked sentences selected by each method; 2, linear interpolation of the datasets selected by the different methods). The paper presents detailed results and analysis for four languages with different levels of morphologic complexity (English, Spanish, Czech and Chinese). The interpolation-based combination outperforms the purely statistical baseline in all the scenarios, resulting in language models with lower perplexity. In relative terms the improvements are similar regardless of the language, with perplexity reductions achieved in the range 7.72–13.02%. In absolute terms the reduction is higher for languages with high type-token ratio (Chinese, 202.16) or rich morphology (Czech, 81.53) and lower for the remaining languages, Spanish (55.2) and English (34.43 on the English side of the same parallel dataset as for Czech and 61.90 on the same parallel dataset as for Spanish).
Antonio Toral, Pavel Pecina, Longyue Wang, Josef van Genabith
Comput. Speech Lang.4
2015 Quality estimation-guided supplementary data selection for domain adaptation of statistical machine translation
Pratyush Banerjee, Raphaël Rubino, Johann Roturier, Josef van Genabith
Mach. Transl.4
2014 Active Learning for Post-Editing Based Incrementally Retrained MT
abstract
Aswarth Abhilash Dara, Josef van Genabith, Qun Liu, John Judge, Antonio Toral. Proceedings of the 14th Conference of the European Chapter of the Association for Computational Linguistics, volume 2: Short Papers. 2014.
Aswarth Abhilash Dara, Josef van Genabith, Qun Liu 0001, John Judge, Antonio Toral
EACL2
2014 The Strategic Impact of META-NET on the Regional, National and International Level
Georg Rehm, Hans Uszkoreit, Sophia Ananiadou, Núria Bel, Audroné Bieleviciené, Lars Borin, António Branco, Gerhard Budin, Nicoletta Calzolari, Walter Daelemans, Radovan Garabík, Marko Grobelnik, Carmen García-Mateo, Josef van Genabith, Jan Hajic 0001, Inma Hernáez Rioja, John Judge, Svetla Koeva, Simon Krek, Cvetana Krstev, Krister Lindén, Bernardo Magnini, Joseph Mariani, John McNaught, Maite Melero, Monica Monachini, Asunción Moreno, Jan Odijk, Maciej Ogrodniczuk, Piotr Pezik, Stelios Piperidis, Adam Przepiórkowski, Eiríkur Rögnvaldsson, Mike Rosner, Bolette S. Pedersen, Inguna Skadina, Koenraad De Smedt, Marko Tadic, Paul Thompson 0002, Dan Tufis, Tamás Váradi, Andrejs Vasiljevs, Kadri Vider, Jolanta Zabarskaite
LREC14
2014 A corpus-based finite-state morphological toolkit for contemporary arabic
abstract
We develop an open-source large-scale finite-state morphological processing toolkit (AraComLex) for Modern Standard Arabic (MSA) distributed under the GPLv3 license (http://aracomlex.sourceforge.net). The morphological transducer is based on a lexical database specifically constructed for this purpose. In contrast to previous resources, the database is tuned to MSA, eliminating lexical entries no longer attested in contemporary use. The database is built using a corpus of 1,089,111,204 word tokens, a pre-annotation tool, machine learning techniques and knowledge-based pattern matching to automatically acquire lexical knowledge. Our morphological transducer is evaluated and compared to LDC's SAMA(Standard Arabic Morphological Analyser). We also develop a finite-state morphological guesser as part of a methodology for extracting unknown word forms, lemmatizing them, and giving them a priority weight for inclusion in the lexicon.
Pavel Pecina, Antonio Toral, Josef van Genabith
J. Log. Comput.4
2013 Quality Estimation-guided Data Selection for Domain Adaptation of SMT
Pratyush Banerjee, Raphaël Rubino, Johann Roturier, Josef van Genabith
MTSummit4
2013 TMTprime: A Recommender System for MT and TM Integration
Aswarth Abhilash Dara, Sandipan Dandapat, Declan Groves, Josef van Genabith
HLT-NAACL4
2013 Predicting sentence translation quality using extrinsic and language independent features
Ergun Biçici, Declan Groves, Josef van Genabith
Mach. Transl.3
2012 Minimum Bayes Risk Decoding with Enlarged Hypothesis Space in System Combination
Tsuyoshi Okita 0002, Josef van Genabith
CICLing (2)2
2012 The Floating Arabic Dictionary: An Automatic Method for Updating a Lexical Database through the Detection and Lemmatization of Unknown Words
Younes Samih, Khaled Shaalan, Josef van Genabith
COLING4
2012 Translation Quality-Based Supplementary Data Selection by Incremental Update of Translation Models
Pratyush Banerjee, Sudip Kumar Naskar, Johann Roturier, Andy Way, Josef van Genabith
COLING5
2012 An Evaluation of Statistical Post-Editing Systems Applied to RBMT and SMT Systems
Hannah Béchara, Raphaël Rubino, Yifan He 0007, Yanjun Ma, Josef van Genabith
COLING5
2012 Simple and Effective Parameter Tuning for Domain Adaptation of Statistical Machine Translation
Pavel Pecina, Antonio Toral, Josef van Genabith
COLING3
2012 Domain Adaptation in SMT of User-Generated Forum Content Guided by OOV Word Reduction: Normalization and/or Supplementary Data
Pratyush Banerjee, Sudip Kumar Naskar, Johann Roturier, Andy Way, Josef van Genabith
EAMT5
2012 Domain Adaptation of Statistical Machine Translation using Web-Crawled Resources: A Case Study
Pavel Pecina, Antonio Toral, Vassilis Papavassiliou, Prokopis Prokopidis, Josef van Genabith
EAMT5
2012 Automatic Extraction and Evaluation of Arabic LFG Resources
Khaled Shaalan, Lamia Tounsi, Josef van Genabith
LREC4
2012 A Richly Annotated, Multilingual Parallel Corpus for Hybrid Machine Translation
Eleftherios Avramidis, Marta R. Costa-jussà, Christian Federmann, Josef van Genabith, Maite Melero, Pavel Pecina
LREC4
2012 The ML4HMT Workshop on Optimising the Division of Labour in Hybrid Machine Translation
Christian Federmann, Eleftherios Avramidis, Marta R. Costa-jussà, Josef van Genabith, Maite Melero, Pavel Pecina
LREC4
2012 Irish Treebanking and Parsing: A Preliminary Evaluation
Teresa Lynn, Özlem Çetinoglu, Jennifer Foster, Elaine Uí Dhonnchadha, Mark Dras, Josef van Genabith
LREC6
2012 Arabic Word Generation and Modelling for Spell Checking
Khaled Shaalan, Pavel Pecina, Younes Samih, Josef van Genabith
LREC5
2011 Consistent Translation using Discriminative Learning - A Translation Memory-inspired Approach
Yanjun Ma, Yifan He 0007, Andy Way, Josef van Genabith
ACL4
2011 From News to Comment: Resources and Benchmarks for Parsing the Language of Web 2.0
Jennifer Foster, Özlem Çetinoglu, Joachim Wagner 0001, Joseph Le Roux, Joakim Nivre, Deirdre Hogan, Josef van Genabith
IJCNLP7
2011 Domain Adaptation in Statistical Machine Translation of User-Forum Data using Component Level Mixture Modelling
Pratyush Banerjee, Sudip Kumar Naskar, Johann Roturier, Andy Way, Josef van Genabith
MTSummit5
2011 Statistical Post-Editing for a Statistical MT System
Hannah Béchara, Yanjun Ma, Josef van Genabith
MTSummit3
2011 Rich Linguistic Features for Translation Memory-Inspired Consistent Translation
Yifan He 0007, Yanjun Ma, Andy Way, Josef van Genabith
MTSummit4
2011 Dependency-based n-gram models for general purpose sentence realisation
abstract
Abstract This paper presents a general-purpose, wide-coverage, probabilistic sentence generator based on dependency n-gram models. This is particularly interesting as many semantic or abstract syntactic input specifications for sentence realisation can be represented as labelled bi-lexical dependencies or typed predicate-argument structures. Our generation method captures the mapping between semantic representations and surface forms by linearising a set of dependencies directly, rather than via the application of grammar rules as in more traditional chart-style or unification-based generators. In contrast to conventional n-gram language models over surface word forms, we exploit structural information and various linguistic features inherent in the dependency representations to constrain the generation space and improve the generation quality. A series of experiments shows that dependency-based n-gram models generalise well to different languages (English and Chinese) and representations (LFG and CoNLL). Compared with state-of-the-art generation systems, our general-purpose sentence realiser is highly competitive with the added advantages of being simple, fast, robust and accurate.
Yuqing Guo 0004, Haifeng Wang 0001, Josef van Genabith
Nat. Lang. Eng.3
2010 Bridging SMT and TM with Translation Recommendation
Yifan He 0007, Yanjun Ma, Josef van Genabith, Andy Way
ACL3
2010 Hard Constraints for Grammatical Function Labelling
Wolfgang Seeker, Ines Rehbein, Jonas Kuhn, Josef van Genabith
ACL4
2010 Finding Common Ground: Towards a Surface Realisation Shared Task
Anya Belz, Mike White, Josef van Genabith, Deirdre Hogan, Amanda Stent
INLG3
2010 An Automatically Built Named Entity Lexicon for Arabic
Antonio Toral, Lamia Tounsi, Monica Monachini, Josef van Genabith
LREC5
2010 Partial Dependency Parsing for Irish
Elaine Uí Dhonnchadha, Josef van Genabith
LREC2
2010 Arabic Parsing Using Grammar Transforms
Lamia Tounsi, Josef van Genabith
LREC2
2010 A Linguistically Inspired Statistical Model for Chinese Punctuation Generation
abstract
This article investigates a relatively underdeveloped subject in natural language processing---the generation of punctuation marks. From a theoretical perspective, we study 16 Chinese punctuation marks as defined in the Chinese national standard of punctuation usage, and categorize these punctuation marks into three different types according to their syntactic properties. We implement a three-tier maximum entropy model incorporating linguistically-motivated features for generating the commonly used Chinese punctuation marks in unpunctuated sentences output by a surface realizer. Furthermore, we present a method to automatically extract cue words indicating sentence-final punctuation marks as a specialized feature to construct a more precise model. Evaluating on the Penn Chinese Treebank data, the MaxEnt model achieves an f -score of 79.83% for punctuation insertion and 74.61% for punctuation restoration using gold data input, 79.50% for insertion and 73.32% for restoration using parser-based imperfect input. The experiments show that the MaxEnt model significantly outperforms a baseline 5-gram language model that scores 54.99% for punctuation insertion and 52.01% for restoration. We show that our results are not far from human performance on the same task with human insertion f -scores in the range of 81-87% and human restoration in the range of 71-82%. Finally, a manual error analysis of the generation output shows that close to 40% of the mismatched punctuation marks do in fact result in acceptable choices, a fact obscured in the automatic string-matching based evaluation scores.
Yuqing Guo 0004, Haifeng Wang 0001, Josef van Genabith
ACM Trans. Asian Lang. Inf. Process.3
2009 Experiments on Domain Adaptation for English--Hindi SMT
Rejwanul Haque, Sudip Kumar Naskar, Josef van Genabith, Andy Way
PACLIC3
2008 Dependency-Based N-Gram Models for General Purpose Sentence Realisation
Yuqing Guo 0004, Josef van Genabith, Haifeng Wang 0001
COLING2
2008 Packed rules for automatic transfer-rule induction
Yvette Graham, Josef van Genabith
EAMT2
2008 Accurate and Robust LFG-Based Generation for Chinese
Yuqing Guo 0004, Haifeng Wang 0001, Josef van Genabith
INLG3
2008 Parser-Based Retraining for Domain Adaptation of Probabilistic Generators
Deirdre Hogan, Jennifer Foster, Joachim Wagner 0001, Josef van Genabith
INLG4
2008 Learning Morphology with Morfette
Grzegorz Chrupala, Georgiana Dinu, Josef van Genabith
LREC3
2008 Parser Evaluation and the BNC: Evaluating 4 constituency parsers with 3 metrics
Jennifer Foster, Josef van Genabith
LREC2
2008 Treebank-Based Acquisition of LFG Parsing Resources for French
Natalie Schluter, Josef van Genabith
LREC2
2008 Wide-Coverage Deep Statistical Parsing Using Automatic Dependency Structure Annotation
abstract
A number of researchers have recently conducted experiments comparing “deep” hand-crafted wide-coverage with “shallow” treebank- and machine-learning-based parsers at the level of dependencies, using simple and automatic methods to convert tree output generated by the shallow parsers into dependencies. In this article, we revisit such experiments, this time using sophisticated automatic LFG f-structure annotation methodologies with surprising results. We compare various PCFG and history-based parsers to find a baseline parsing system that fits best into our automatic dependency structure annotation technique. This combined system of syntactic parser and dependency structure annotation is compared to two hand-crafted, deep constraint-based parsers, RASP and XLE. We evaluate using dependency-based gold standards and use the Approximate Randomization Test to test the statistical significance of the results. Our experiments show that machine-learning-based shallow grammars augmented with sophisticated automatic dependency annotation technology outperform hand-crafted, deep, wide-coverage constraint grammars. Currently our best system achieves an f-score of 82.73% against the PARC 700 Dependency Bank, a statistically significant improvement of 2.18% over the most recent results of 80.55% for the hand-crafted LFG grammar and XLE parsing system and an f-score of 80.23% against the CBS 500 Dependency Bank, a statistically significant 3.66% improvement over the 76.57% achieved by the hand-crafted RASP grammar and parsing system.
Aoife Cahill, Michael Burke, Ruth O'Donovan, Stefan Riezler, Josef van Genabith, Andy Way
Comput. Linguistics5
2007 Recovering Non-Local Dependencies for Chinese
Yuqing Guo 0004, Haifeng Wang 0001, Josef van Genabith
EMNLP-CoNLL3
2007 Exploiting Multi-Word Units in History-Based Probabilistic Generation
Deirdre Hogan, Conor Cafferkey, Aoife Cahill, Josef van Genabith
EMNLP-CoNLL4
2007 Treebank Annotation Schemes and Parser Evaluation for German
Ines Rehbein, Josef van Genabith
EMNLP-CoNLL2
2007 A Comparative Evaluation of Deep and Shallow Approaches to the Automatic Detection of Common Grammatical Errors
Joachim Wagner 0001, Jennifer Foster, Josef van Genabith
EMNLP-CoNLL3
2007 Automatic Acquisition of Lexical-Functional Grammar Resources from a Japanese Dependency Corpus
Masanori Oya, Josef van Genabith
PACLIC2
2007 Evaluating machine translation with LFG dependencies
Karolina Owczarzak, Josef van Genabith, Andy Way
Mach. Transl.2
2006 Robust PCFG-Based Generation Using Automatically Acquired LFG Approximations
abstract
We present a novel PCFG-based architecture for robust probabilistic generation based on wide-coverage LFG approximations (Cahill et al., 2004) automatically extracted from treebanks, maximising the probability of a tree given an f-structure. We evaluate our approach using string-based evaluation. We currently achieve coverage of 95.26%, a BLEU score of 0.7227 and string accuracy of 0.7476 on the Penn-II WSJ Section 23 sentences of length ≤20.
Aoife Cahill, Josef van Genabith
ACL2
2006 Using Machine-Learning to Assign Function Labels to Parser Output for Spanish
Grzegorz Chrupala, Josef van Genabith
ACL2
2006 QuestionBank: Creating a Corpus of Parse-Annotated Questions
abstract
This paper describes the development of QuestionBank, a corpus of 4000 parse-annotated questions for (i) use in training parsers employed in QA, and (ii) evaluation of question parsing. We present a series of experiments to investigate the effectiveness of QuestionBank as both an exclusive and supplementary training resource for a state-of-the-art parser in parsing both question and non-question test sets. We introduce a new method for recovering empty nodes and their antecedents (capturing long distance dependencies) from parser output in CFG trees using LFG f-structure reentrancies. Our main findings are (i) using QuestionBank training data improves parser performance to 89.75% labelled bracketing f-score, an increase of almost 11% over the baseline; (ii) back-testing experiments on non-question data (Penn-II WSJ Section 23) shows that the retrained parser does not suffer a performance drop on non-question material; (iii) ablation experiments show that the size of training material provided by QuestionBank is sufficient to achieve optimal results; (iv) our method for recovering empty nodes captures long distance dependencies in questions from the ATIS corpus with high precision (96.82%) and low recall (39.38%). In summary, QuestionBank provides a useful new resource in parser-based QA research.
John Judge, Aoife Cahill, Josef van Genabith
ACL3
2006 A Syntactic Skeleton for Statistical Machine Translation
Bart Mellebeek, Karolina Owczarzak, Declan Groves, Josef van Genabith, Andy Way
EAMT4
2006 A Part-of-speech tagger for Irish using Finite-State Morphology and Constraint Grammar Disambiguation
Elaine Uí Dhonnchadha, Josef van Genabith
LREC2
2005 TransBooster: boosting the performance of wide-coverage machine translation systems
Bart Mellebeek, Anna Khasin, Josef van Genabith, Andy Way
EAMT3
2005 Improving Online Machine Translation Systems
abstract
In (Mellebeek et al., 2005), we proposed the design, implementation and evaluation of a novel and modular approach to boost the translation performance of existing, wide-coverage, freely available machine translation systems, based on reliable and fast automatic decomposition of the translation input and corresponding composition of translation output. Despite showing some initial promise, our method did not improve on the baseline Logomedia1 and Systran2 MT systems. In this paper, we improve on the algorithm presented in (Mellebeek et al., 2005), and on the same test data, show increased scores for a range of automatic evaluation metrics. Our algorithm now outperforms Logomedia, obtains similar results to SDL3 and falls tantalisingly short of the performance achieved by Systran.
Bart Mellebeek, Anna Khasin, Karolina Owczarzak, Josef van Genabith, Andy Way
MTSummit4
2005 Dynamically structuring, updating and interrelating representations of visual and linguistic discourse context
John D. Kelleher, Fintan J. Costello, Josef van Genabith
Artif. Intell.3
2005 Large-Scale Induction and Evaluation of Lexical Resources from the Penn-II and Penn-III Treebanks
abstract
We present a methodology for extracting subcategorization frames based on an automatic lexical-functional grammar (LFG) f-structure annotation algorithm for the Penn-II and Penn-III Treebanks. We extract syntactic-function-based subcategorization frames (LFG semantic forms) and traditional CFG category-based subcategorization frames as well as mixed function/category-based frames, with or without preposition information for obliques and particle information for particle verbs. Our approach associates probabilities with frames conditional on the lemma, distinguishes between active and passive frames, and fully reflects the effects of long-distance dependencies in the source data structures. In contrast to many other approaches, ours does not predefine the subcategorization frame types extracted, learning them instead from the source data. Including particles and prepositions, we extract 21,005 lemma frame types for 4,362 verb lemmas, with a total of 577 frame types and an average of 4.8 frame types per verb. We present a large-scale evaluation of the complete set of forms extracted against the full COMLEX resource. To our knowledge, this is the largest and most complete evaluation of subcategorization frames acquired automatically for English.
Ruth O'Donovan, Michael Burke, Aoife Cahill, Josef van Genabith, Andy Way
Comput. Linguistics4
2004 Long-Distance Dependency Resolution in Automatically Acquired Wide-Coverage PCFG-Based LFG Approximations
abstract
This paper shows how finite approximations of long distance dependency (LDD) resolution can be obtained automatically for wide-coverage, robust, probabilistic Lexical-Functional Grammar (LFG) resources acquired from treebanks. We extract LFG subcategorisation frames and paths linking LDD reentrancies from f-structures generated automatically for the Penn-II treebank trees and use them in an LDD resolution algorithm to parse new text. Unlike (Collins, 1999; Johnson, 2000), in our approach resolution of LDDs is done at f-structure (attribute-value structure representations of basic predicate-argument or dependency structure) without empty productions, traces and coindexation in CFG parse trees. Currently our best automatically induced grammars achieve 80.97% f-score for f-structures parsing section 23 of the WSJ part of the Penn-II treebank and evaluating against the DCU 1051 and 80.24% against the PARC 700 Dependency Bank (King et al., 2003), performing at the same or a slightly better level than state-of-the-art hand-crafted grammars (Kaplan et al., 2004).
Aoife Cahill, Michael Burke, Ruth O'Donovan, Josef van Genabith, Andy Way
ACL4
2004 Large-Scale Induction and Evaluation of Lexical Resources from the Penn-II Treebank
abstract
In this paper we present a methodology for extracting subcategorisation frames based on an automatic LFG f-structure annotation algorithm for the Penn-II Treebank. We extract abstract syntactic function-based subcategorisation frames (LFG semantic forms), traditional CFG category-based subcategorisation frames as well as mixed function/category-based frames, with or without preposition information for obliques and particle information for particle verbs. Our approach does not predefine frames, associates probabilities with frames conditional on the lemma, distinguishes between active and passive frames, and fully reflects the effects of long-distance dependencies in the source data structures. We extract 3586 verb lemmas, 14348 semantic form types (an average of 4 per lemma) with 577 frame types. We present a large-scale evaluation of the complete set of forms extracted against the full COMLEX resource.
Ruth O'Donovan, Michael Burke, Aoife Cahill, Josef van Genabith, Andy Way
ACL4
2004 Treebank-Based Acquisition of a Chinese Lexical-Functional Grammar
Michael Burke, Olivia S.-C. Lam, Aoife Cahill, Rowena Chan, Ruth O'Donovan, Adams Bodomo, Josef van Genabith, Andy Way
PACLIC7
2003 Design, Implementation and Evaluation of an Inflectional Morphology Finite State Transducer for Irish
Elaine Uí Dhonnchadha, Caoilfhionn Nic Pháidín, Josef van Genabith
Mach. Transl.3
2002 TTS - A Treebank Tool Suite
Aoife Cahill, Josef van Genabith
LREC2
1997 On Interpreting F-Structures as UDRSs
abstract
We describe a method for interpreting abstract flat syntactic representations, LFG f-structures, as underspecified semantic representations, here Underspecified Discourse Representation Structures (UDRSs). The method establishes a one-to-one correspondence between subsets of the LFG and UDRS formalisms. It provides a model theoretic interpretation and an inferential component which operates directly on underspecified representations for f-structures through the translation images of f-structures as UDRSs.
Josef van Genabith, Richard S. Crouch
ACL1
1996 Direct and Underspecified Interpretations of LFG f-structures
Josef van Genabith, Richard S. Crouch
COLING1
1993 Experiments in Reusability of Grammatical Resources
Doug Arnold, Toni Badia, Josef van Genabith, Stella Markantonatou, Stefan Momma, Louisa Sadler, Paul Schmidt
EACL3