EDBT 2026 Demo / reviewers in the wild / expert
Christof Monz
dblp:m/ChristofMonz
· DBLP profile ↗
65ranked-venue papers
9as first author
20since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 53 · 3 first-author · 19 since 2021Databases, data management, data science and information retrieval · 10 · 4 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-authorTheory of computation · 2 · 2 first-authorApplied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ReMedy: Learning Machine Translation Evaluation from Human Preferences with Reward ModelingabstractA key challenge in MT evaluation is the inherent noise and inconsistency of human ratings.Regression-based neural metrics struggle with this noise, while prompting LLMs shows promise at system-level evaluation but performs poorly at segment level.In this work, we propose ReMedy, a novel MT metric framework that reformulates translation evaluation as a reward modeling task.Instead of regressing on imperfect human ratings directly, ReMedy learns relative translation quality using pairwise preference data, resulting in a more reliable evaluation.In extensive experiments across WMT22-24 shared tasks (39 language pairs, 111 MT systems), ReMedy achieves stateof-the-art performance at both segment-and system-level evaluation.Specifically, ReMedy-9B surpasses larger WMT winners and massive closed LLMs such as MetricX-13B, XCOMET-Ensemble, GEMBA-GPT-4, PaLM-540B, and finetuned PaLM2.Further analyses demonstrate that ReMedy delivers superior capability in detecting translation errors and evaluating low-quality translations.1 Shaomu Tan, Christof Monz |
EMNLP | 2 |
| 2025 | Please Translate Again: Two Simple Experiments on Whether Human-Like Reasoning Helps TranslationabstractLarge Language Models (LLMs) demonstrate strong reasoning capabilities for many tasks, often by explicitly decomposing the task via Chain-of-Thought (CoT) reasoning.Recent work on LLM-based translation designs handcrafted prompts to decompose translation, or trains models to incorporate intermediate steps.Translating Step-by-step (Briakou et al., 2024), for instance, introduces a multi-step prompt with decomposition and refinement of translation with LLMs, which achieved stateof-the-art results on WMT24 test data.In this work, we scrutinise this strategy's effectiveness.Empirically, we find no clear evidence that performance gains stem from explicitly decomposing the translation process via CoT, at least for the models on test; and we show prompting LLMs to "translate again" and self-refine yields even better results than human-like stepby-step prompting.While the decomposition influences translation behaviour, faithfulness to the decomposition has both positive and negative effects on translation.Our analysis therefore suggests a divergence between the optimal translation strategies for humans and LLMs. Seth Aycock, Christof Monz |
EMNLP | 3 |
| 2025 | Can LLMs Really Learn to Translate a Low-Resource Language from One Grammar Book?abstractExtremely low-resource (XLR) languages lack substantial corpora for training NLP models, motivating the use of all available resources such as dictionaries and grammar books. Machine Translation from One Book (Tanzer et al., 2024) suggests that prompting long-context LLMs with one grammar book enables English–Kalamang translation, an XLR language unseen by LLMs—a noteworthy case of linguistics helping an NLP task. We investigate the source of this translation ability, finding almost all improvements stem from the book’s parallel examples rather than its grammatical explanations. We find similar results for Nepali and Guarani, seen low-resource languages, and we achieve performance comparable to an LLM with a grammar book by simply fine-tuning an encoder-decoder translation model. We then investigate where grammar books help by testing two linguistic tasks, grammaticality judgment and gloss prediction, and we explore what kind of grammatical knowledge helps by introducing a typological feature prompt that achieves leading results on these more relevant tasks. We thus emphasise the importance of task-appropriate data for XLR languages: parallel examples for translation, and grammatical data for linguistic tasks. As we find no evidence that long-context LLMs can make effective use of grammatical explanations for XLR translation, we conclude data collection for multilingual XLR tasks such as translation is best focused on parallel data over linguistic description. Seth Aycock, David Stap, Christof Monz, Khalil Sima'an |
ICLR | 4 |
| 2025 | Reward-Guided Speculative Decoding for Efficient LLM ReasoningabstractWe introduce Reward-Guided Speculative Decoding (RSD), a novel framework aimed at improving the efficiency of inference in large language models (LLMs). RSD synergistically combines a lightweight draft model with a more powerful target model, incorporating a controlled bias to prioritize high-reward outputs, in contrast to existing speculative decoding methods that enforce strict unbiasedness. RSD employs a process reward model to evaluate intermediate decoding steps and dynamically decide whether to invoke the target model, optimizing the trade-off between computational cost and output quality. We theoretically demonstrate that a threshold-based mixture strategy achieves an optimal balance between resource utilization and performance. Extensive evaluations on challenging reasoning benchmarks, including Olympiad-level tasks, show that RSD delivers significant efficiency gains against decoding with the target model only (up to 4.4X fewer FLOPs), while achieving significant better accuracy than parallel decoding method on average (up to +3.5). These results highlight RSD as a robust and cost-effective approach for deploying LLMs in resource-intensive scenarios. Baohao Liao, Hanze Dong, Junnan Li 0001, Christof Monz, Silvio Savarese, Doyen Sahoo, Caiming Xiong |
ICML | 5 |
| 2025 | Calibrating Translation Decoding with Quality Estimation on LLMsabstractNeural machine translation (NMT) systems typically employ maximum *a posteriori* (MAP) decoding to select the highest-scoring translation from the distribution. However, recent evidence highlights the inadequacy of MAP decoding, often resulting in low-quality or even pathological hypotheses as the decoding objective is only weakly aligned with real-world translation quality. This paper proposes to directly calibrate hypothesis likelihood with translation quality from a distributional view by directly optimizing their Pearson correlation, thereby enhancing decoding effectiveness. With our method, translation with large language models (LLMs) improves substantially after limited training (2K instances per direction). This improvement is orthogonal to those achieved through supervised fine-tuning, leading to substantial gains across a broad range of metrics and human evaluations. This holds even when applied to top-performing translation-specialized LLMs fine-tuned on high-quality translation data, such as Tower, or when compared to recent preference optimization methods, like CPO. Moreover, the calibrated translation likelihood can directly serve as a strong proxy for translation quality, closely approximating or even surpassing some state-of-the-art translation quality estimation models, like CometKiwi.
Lastly, our in-depth analysis demonstrates that calibration enhances the effectiveness of MAP decoding, thereby enabling greater efficiency in real-world deployment. The resulting state-of-the-art translation model, which covers 10 languages, along with the accompanying code and human evaluation data, has been released: https://github.com/moore3930/calibrating-llm-mt. Yibin Lei, Christof Monz |
NeurIPS | 3 |
| 2024 | The Fine-Tuning Paradox: Boosting Translation Quality Without Sacrificing LLM AbilitiesabstractFine-tuning large language models (LLMs) for machine translation has shown improvements in overall translation quality.However, it is unclear what is the impact of fine-tuning on desirable LLM behaviors that are not present in neural machine translation models, such as steerability, inherent document-level translation abilities, and the ability to produce less literal translations.We perform an extensive translation evaluation on the LLaMA and Falcon family of models with model size ranging from 7 billion up to 65 billion parameters.Our results show that while fine-tuning improves the general translation quality of LLMs, several abilities degrade.In particular, we observe a decline in the ability to perform formality steering, to produce technical translations through few-shot examples, and to perform documentlevel translation.On the other hand, we observe that the model produces less literal translations after fine-tuning on parallel data.We show that by including monolingual data as part of the fine-tuning data we can maintain the abilities while simultaneously enhancing overall translation quality.Our findings emphasize the need for fine-tuning strategies that preserve the benefits of LLMs for machine translation. David Stap, Eva Hasler, William J. Byrne, Christof Monz, Ke Tran |
ACL (1) | 4 |
| 2024 | Disentangling the Roles of Target-side Transfer and Regularization in Multilingual Machine TranslationabstractMultilingual Machine Translation (MMT) benefits from knowledge transfer across different language pairs.However, improvements in oneto-many translation compared to many-to-one translation are only marginal and sometimes even negligible.This performance discrepancy raises the question of to what extent positive transfer plays a role on the target-side for oneto-many MT.In this paper, we conduct a largescale study that varies the auxiliary target-side languages along two dimensions, i.e., linguistic similarity and corpus size, to show the dynamic impact of knowledge transfer on the main language pairs.We show that linguistically similar auxiliary target languages exhibit strong ability to transfer positive knowledge.With an increasing size of similar target languages, the positive transfer is further enhanced to benefit the main language pairs.Meanwhile, we find distant auxiliary target languages can also unexpectedly benefit main language pairs, even with minimal positive transfer ability.Apart from transfer, we show distant auxiliary target languages can act as a regularizer to benefit translation performance by enhancing the generalization and model inference calibration. Christof Monz |
EACL (1) | 2 |
| 2024 | Analyzing the Evaluation of Cross-Lingual Knowledge Transfer in Multilingual Language ModelsabstractRecent advances in training multilingual language models on large datasets seem to have shown promising results in knowledge transfer across languages and achieve high performance on downstream tasks.However, we question to what extent the current evaluation benchmarks and setups accurately measure zero-shot crosslingual knowledge transfer.In this work, we challenge the assumption that high zero-shot performance on target tasks reflects high crosslingual ability by introducing more challenging setups involving instances with multiple languages.Through extensive experiments and analysis, we show that the observed high performance of multilingual models can be largely attributed to factors not requiring the transfer of actual linguistic knowledge, such as task-and surface-level knowledge.More specifically, we observe what has been transferred across languages is mostly data artifacts and biases, especially for low-resource languages.Our findings highlight the overlooked drawbacks of existing cross-lingual test data and evaluation setups, calling for a more nuanced understanding of the cross-lingual capabilities of multilingual models. 1 Sara Rajaee, Christof Monz |
EACL (1) | 2 |
| 2024 | ApiQ: Finetuning of 2-Bit Quantized Large Language ModelabstractMemory-efficient finetuning of large language models (LLMs) has recently attracted huge attention with the increasing size of LLMs, primarily due to the constraints posed by GPU memory limitations and the effectiveness of these methods compared to full finetuning.Despite the advancements, current strategies for memory-efficient finetuning, such as QLoRA, exhibit inconsistent performance across diverse bit-width quantizations and multifaceted tasks.This inconsistency largely stems from the detrimental impact of the quantization process on preserved knowledge, leading to catastrophic forgetting and undermining the utilization of pretrained models for finetuning purposes.In this work, we introduce a novel quantization framework named ApiQ, designed to restore the lost information from quantization by concurrently initializing the LoRA components and quantizing the weights of LLMs.This approach ensures the maintenance of the original LLM's activation precision while mitigating the error propagation from shallower into deeper layers.Through comprehensive evaluations conducted on a spectrum of language tasks with various LLMs, ApiQ demonstrably minimizes activation error during quantization.Consequently, it consistently achieves superior finetuning results across various bit-widths.Notably, one can even finetune a 2-bit Llama-2-70b with ApiQ on a single NVIDIA A100-80GB GPU without any memory-saving techniques, and achieve promising results. Baohao Liao, Christian Herold, Shahram Khadivi, Christof Monz |
EMNLP | 4 |
| 2024 | Communicating with Speakers and Listeners of Different Pragmatic LevelsabstractThis paper explores the impact of variable pragmatic competence on communicative success through simulating language learning and conversing between speakers and listeners with different levels of reasoning abilities.Through studying this interaction, we hypothesize that matching levels of reasoning between communication partners would create a more beneficial environment for communicative success and language learning.Our research findings indicate that learning from more explicit, literal language is advantageous, irrespective of the learner's level of pragmatic competence.Furthermore, we find that integrating pragmatic reasoning during language learning, not just during evaluation, significantly enhances overall communication performance.This paper provides key insights into the importance of aligning reasoning levels and incorporating pragmatic reasoning in optimizing communicative interactions. Kata Naszádi, Frans A. Oliehoek, Christof Monz |
EMNLP | 3 |
| 2024 | Neuron Specialization: Leveraging Intrinsic Task Modularity for Multilingual Machine TranslationabstractTraining a unified multilingual model promotes knowledge transfer but inevitably introduces negative interference.Language-specific modeling methods show promise in reducing interference.However, they often rely on heuristics to distribute capacity and struggle to foster cross-lingual transfer via isolated modules.In this paper, we explore intrinsic task modularity within multilingual networks and leverage these observations to circumvent interference under multilingual translation.We show that neurons in the feed-forward layers tend to be activated in a language-specific manner.Meanwhile, these specialized neurons exhibit structural overlaps that reflect language proximity, which progress across layers.Based on these findings, we propose Neuron Specialization, an approach that identifies specialized neurons to modularize feed-forward layers and then continuously updates them through sparse networks.Extensive experiments show that our approach achieves consistent performance gains over strong baselines with additional analyses demonstrating reduced interference and increased knowledge transfer.1 Shaomu Tan, Christof Monz |
EMNLP | 3 |
| 2024 | 3-in-1: 2D Rotary Adaptation for Efficient Finetuning, Efficient Batching and ComposabilityabstractParameter-efficient finetuning (PEFT) methods effectively adapt large language models (LLMs) to diverse downstream tasks, reducing storage and GPU memory demands. Despite these advantages, several applications pose new challenges to PEFT beyond mere parameter efficiency. One notable challenge involves the efficient deployment of LLMs equipped with multiple task- or user-specific adapters, particularly when different adapters are needed for distinct requests within the same batch. Another challenge is the interpretability of LLMs, which is crucial for understanding how LLMs function. Previous studies introduced various approaches to address different challenges. In this paper, we introduce a novel method, RoAd, which employs a straightforward 2D rotation to adapt LLMs and addresses all the above challenges: (1) RoAd is remarkably parameter-efficient, delivering optimal performance on GLUE, eight commonsense reasoning tasks and four arithmetic reasoning tasks with <0.1% trainable parameters; (2) RoAd facilitates the efficient serving of requests requiring different adapters within a batch, with an overhead comparable to element-wise multiplication instead of batch matrix multiplication; (3) RoAd enhances LLM's interpretability through integration within a framework of distributed interchange intervention, demonstrated via composition experiments. Baohao Liao, Christof Monz |
NeurIPS | 2 |
| 2023 | Parameter-Efficient Fine-Tuning without Introducing New LatencyabstractParameter-efficient fine-tuning (PEFT) of pretrained language models has recently demonstrated remarkable achievements, effectively matching the performance of full fine-tuning while utilizing significantly fewer trainable parameters, and consequently addressing the storage and communication constraints.Nonetheless, various PEFT methods are limited by their inherent characteristics.In the case of sparse fine-tuning, which involves modifying only a small subset of the existing parameters, the selection of fine-tuned parameters is task-and domain-specific, making it unsuitable for federated learning.On the other hand, PEFT methods with adding new parameters typically introduce additional inference latency.In this paper, we demonstrate the feasibility of generating a sparse mask in a task-agnostic manner, wherein all downstream tasks share a common mask.Our approach, which relies solely on the magnitude information of pre-trained parameters, surpasses existing methodologies by a significant margin when evaluated on the GLUE benchmark.Additionally, we introduce a novel adapter technique that directly applies the adapter to pre-trained parameters instead of the hidden representation, thereby achieving identical inference speed to that of full finetuning.Through extensive experiments, our proposed method attains a new state-of-the-art outcome in terms of both performance and storage efficiency, storing only 0.03% parameters of full fine-tuning.1 Baohao Liao, Christof Monz |
ACL (1) | 3 |
| 2023 | Multilingual k-Nearest-Neighbor Machine Translationabstractk-nearest-neighbor machine translation has demonstrated remarkable improvements in machine translation quality by creating a datastore of cached examples.However, these improvements have been limited to high-resource language pairs, with large datastores, and remain a challenge for low-resource languages.In this paper, we address this issue by combining representations from multiple languages into a single datastore.Our results consistently demonstrate substantial improvements not only in low-resource translation quality (up to +3.6 BLEU), but also for high-resource translation quality (up to +0.5 BLEU).Our experiments show that it is possible to create multilingual datastores that are a quarter of the size, achieving a 5.3x speed improvement, by using linguistic similarities for datastore creation.1 David Stap, Christof Monz |
EMNLP | 2 |
| 2023 | Towards a Better Understanding of Variations in Zero-Shot Neural Machine Translation PerformanceabstractMultilingual Neural Machine Translation (MNMT) facilitates knowledge sharing but often suffers from poor zero-shot (ZS) translation qualities.While prior work has explored the causes of overall low zero-shot translation qualities, our work introduces a fresh perspective: the presence of significant variations in zeroshot performance.This suggests that MNMT does not uniformly exhibit poor zero-shot capability; instead, certain translation directions yield reasonable results.Through systematic experimentation, spanning 1,560 language directions across 40 languages, we identify three key factors contributing to high variations in ZS NMT performance: 1) target-side translation quality, 2) vocabulary overlap, and 3) linguistic properties.Our findings highlight that the target side translation quality is the most influential factor, with vocabulary overlap consistently impacting zero-shot capabilities.Additionally, linguistic properties, such as language family and writing system, play a role, particularly with smaller models.Furthermore, we suggest that the off-target issue is a symptom of inadequate performance, emphasizing that zero-shot translation challenges extend beyond addressing the off-target problem.To support future research, we release the data and models as a benchmark for the study of ZS NMT. 1 Shaomu Tan, Christof Monz |
EMNLP | 2 |
| 2023 | Beyond Shared Vocabulary: Increasing Representational Word Similarities across Languages for Multilingual Machine TranslationabstractUsing a vocabulary that is shared across languages is common practice in Multilingual Neural Machine Translation (MNMT).In addition to its simple design, shared tokens play an important role in positive knowledge transfer, assuming that shared tokens refer to similar meanings across languages.However, when word overlap is small, especially due to different writing systems, transfer is inhibited.In this paper, we define word-level information transfer pathways via word equivalence classes and rely on graph networks to fuse word embeddings across languages.Our experiments demonstrate the advantages of our approach: 1) embeddings of words with similar meanings are better aligned across languages, 2) our method achieves consistent BLEU improvements of up to 2.3 points for high-and low-resource MNMT, and 3) less than 1.0% additional trainable parameters are required with a limited increase in computational costs, while inference time remains identical to the baseline.We release the codebase to the community.1 Christof Monz |
EMNLP | 2 |
| 2023 | Joint Dropout: Improving Generalizability in Low-Resource Neural Machine Translation through Phrase Pair VariablesabstractDespite the tremendous success of Neural Machine Translation (NMT), its performance on low- resource language pairs still remains subpar, partly due to the limited ability to handle previously unseen inputs, i.e., generalization. In this paper, we propose a method called Joint Dropout, that addresses the challenge of low-resource neural machine translation by substituting phrases with variables, resulting in significant enhancement of compositionality, which is a key aspect of generalization. We observe a substantial improvement in translation quality for language pairs with minimal resources, as seen in BLEU and Direct Assessment scores. Furthermore, we conduct an error analysis, and find Joint Dropout to also enhance generalizability of low-resource NMT in terms of robustness and adaptability across different domains. Ali Araabi, Vlad Niculae, Christof Monz |
MTSummit (1) | 3 |
| 2023 | Make Pre-trained Model Reversible: From Parameter to Memory Efficient Fine-TuningabstractParameter-efficient fine-tuning (PEFT) of pre-trained language models (PLMs) has emerged as a highly successful approach, with training only a small number of parameters without sacrificing performance and becoming the de-facto learning paradigm with the increasing size of PLMs. However, existing PEFT methods are not memory-efficient, because they still require caching most of the intermediate activations for the gradient calculation, akin to fine-tuning. One effective way to reduce the activation memory is to apply a reversible model, so the intermediate activations are not necessary to be cached and can be recomputed. Nevertheless, modifying a PLM to its reversible variant is not straightforward, since the reversible model has a distinct architecture from the currently released PLMs. In this paper, we first investigate what is a key factor for the success of existing PEFT methods, and realize that it's essential to preserve the PLM's starting point when initializing a PEFT method. With this finding, we propose memory-efficient fine-tuning (MEFT) that inserts adapters into a PLM, preserving the PLM's starting point and making it reversible without additional pre-training. We evaluate MEFT on the GLUE benchmark and five question-answering tasks with various backbones, BERT, RoBERTa, BART and OPT. MEFT significantly reduces the activation memory up to 84% of full fine-tuning with a negligible amount of trainable parameters. Moreover, MEFT achieves the same score on GLUE and a comparable score on the question-answering tasks as full fine-tuning. A similar finding is also observed for the image classification task. Baohao Liao, Shaomu Tan, Christof Monz |
NeurIPS | 3 |
| 2021 | NLQuAD: A Non-Factoid Long Question Answering Data SetabstractWe introduce NLQuAD, the first data set with baseline methods for non-factoid long question answering, a task requiring documentlevel language understanding.In contrast to existing span detection question answering data sets, NLQuAD has non-factoid questions that are not answerable by a short span of text and demanding multiple-sentence descriptive answers and opinions.We show the limitation of the F1 score for evaluation of long answers and introduce Intersection over Union (IoU), which measures position-sensitive overlap between the predicted and the target answer spans.To establish baseline performances, we compare BERT, RoBERTa, and Longformer models.Experimental results and human evaluations show that Longformer outperforms the other architectures, but results are still far behind a human upper bound, leaving substantial room for improvements.NLQuAD's samples exceed the input limitation of most pretrained Transformer-based models, encouraging future research on long sequence language models. 1 Amir Soleimani, Christof Monz, Marcel Worring |
EACL | 2 |
| 2021 | Conversations with Search Engines: SERP-based Conversational Response GenerationabstractIn this article, we address the problem of answering complex information needs by conducting conversations with search engines , in the sense that users can express their queries in natural language and directly receive the information they need from a short system response in a conversational manner. Recently, there have been some attempts towards a similar goal, e.g., studies on Conversational Agent s (CAs) and Conversational Search (CS). However, they either do not address complex information needs in search scenarios or they are limited to the development of conceptual frameworks and/or laboratory-based user studies. We pursue two goals in this article: (1) the creation of a suitable dataset, the Search as a Conversation (SaaC) dataset, for the development of pipelines for conversations with search engines, and (2) the development of a state-of-the-art pipeline for conversations with search engines, Conversations with Search Engines (CaSE), using this dataset. SaaC is built based on a multi-turn conversational search dataset, where we further employ workers from a crowdsourcing platform to summarize each relevant passage into a short, conversational response. CaSE enhances the state-of-the-art by introducing a supporting token identification module and a prior-aware pointer generator, which enables us to generate more accurate responses. We carry out experiments to show that CaSE is able to outperform strong baselines. We also conduct extensive analyses on the SaaC dataset to show where there is room for further improvement beyond CaSE. Finally, we release the SaaC dataset and the code for CaSE and all models used for comparison to facilitate future research on this topic. Pengjie Ren, Zhumin Chen, Zhaochun Ren, Evangelos Kanoulas, Christof Monz, Maarten de Rijke |
ACM Trans. Inf. Syst. | 5 |
| 2020 | RefNet: A Reference-Aware Network for Background Based ConversationabstractExisting conversational systems tend to generate generic responses. Recently, Background Based Conversation (BBCs) have been introduced to address this issue. Here, the generated responses are grounded in some background information. The proposed methods for BBCs are able to generate more informative responses, however, they either cannot generate natural responses or have difficulties in locating the right background information. In this paper, we propose a Reference-aware Network (RefNet) to address both issues. Unlike existing methods that generate responses token by token, RefNet incorporates a novel reference decoder that provides an alternative way to learn to directly select a semantic unit (e.g., a span containing complete semantic information) from the background. Experimental results show that RefNet significantly outperforms state-of-the-art methods in terms of both automatic and human evaluations, indicating that RefNet can generate more appropriate and human-like responses. Chuan Meng, Pengjie Ren, Zhumin Chen, Christof Monz, Jun Ma 0001, Maarten de Rijke |
AAAI | 4 |
| 2020 | Thinking Globally, Acting Locally: Distantly Supervised Global-to-Local Knowledge Selection for Background Based ConversationabstractBackground Based Conversation (BBCs) have been introduced to help conversational systems avoid generating overly generic responses. In a BBC, the conversation is grounded in a knowledge source. A key challenge in BBCs is Knowledge Selection (KS): given a conversational context, try to find the appropriate background knowledge (a text fragment containing related facts or comments, etc.) based on which to generate the next response. Previous work addresses KS by employing attention and/or pointer mechanisms. These mechanisms use a local perspective, i.e., they select a token at a time based solely on the current decoding state. We argue for the adoption of a global perspective, i.e., pre-selecting some text fragments from the background knowledge that could help determine the topic of the next response. We enhance KS in BBCs by introducing a Global-to-Local Knowledge Selection (GLKS) mechanism. Given a conversational context and background knowledge, we first learn a topic transition vector to encode the most likely text fragments to be used in the next response, which is then used to guide the local KS at each decoding timestamp. In order to effectively learn the topic transition vector, we propose a distantly supervised learning schema. Experimental results show that the GLKS model significantly outperforms state-of-the-art methods in terms of both automatic and human evaluation. More importantly, GLKS achieves this without requiring any extra annotations, which demonstrates its high degree of scalability. Pengjie Ren, Zhumin Chen, Christof Monz, Jun Ma 0001, Maarten de Rijke |
AAAI | 3 |
| 2020 | Optimizing Transformer for Low-Resource Neural Machine TranslationabstractLanguage pairs with limited amounts of parallel data, also known as low-resource languages, remain a challenge for neural machine translation.While the Transformer model has achieved significant improvements for many language pairs and has become the de facto mainstream architecture, its capability under low-resource conditions has not been fully investigated yet.Our experiments on different subsets of the IWSLT14 training data show that the effectiveness of Transformer under low-resource conditions is highly dependent on the hyper-parameter settings.Our experiments show that using an optimized Transformer for low-resource conditions improves the translation quality up to 7.3 BLEU points compared to using the Transformer default settings. Ali Araabi, Christof Monz |
COLING | 2 |
| 2020 | Retrospective and Prospective Mixture-of-Generators for Task-Oriented Dialogue Response Generation
Jiahuan Pei, Pengjie Ren, Christof Monz, Maarten de Rijke |
ECAI | 3 |
| 2020 | BERT for Evidence Retrieval and Claim Verification
Amir Soleimani, Christof Monz, Marcel Worring |
ECIR (2) | 2 |
| 2019 | Improving Neural Machine Translation Using Noisy Parallel Data through Distillation
Praveen Dakwale, Christof Monz |
MTSummit (1) | 2 |
| 2019 | An Intrinsic Nearest Neighbor Analysis of Neural Machine Translation Architectures
Hamidreza Ghader, Christof Monz |
MTSummit (1) | 2 |
| 2019 | Improving Neural Response Diversity with Frequency-Aware Cross-Entropy LossabstractSequence-to-Sequence (Seq2Seq) models have achieved encouraging performance on the dialogue response generation task. However, existing Seq2Seq-based response generation methods suffer from a low-diversity problem: they frequently generate generic responses, which make the conversation less interesting. In this paper, we address the low-diversity problem by investigating its connection with model over-confidence reflected in predicted distributions. Specifically, we first analyze the influence of the commonly used Cross-Entropy (CE) loss function, and find that the CE loss function prefers high-frequency tokens, which results in low-diversity responses. We then propose a Frequency-Aware Cross-Entropy (FACE) loss function that improves over the CE loss function by incorporating a weighting mechanism conditioned on token frequency. Extensive experiments on benchmark datasets show that the FACE loss function is able to substantially improve the diversity of existing state-of-the-art Seq2Seq response generation methods, in terms of both automatic and human evaluations. Shaojie Jiang, Pengjie Ren, Christof Monz, Maarten de Rijke |
WWW | 3 |
| 2018 | Back-Translation Sampling by Targeting Difficult Words in Neural Machine TranslationabstractNeural Machine Translation has achieved state-of-the-art performance for several language pairs using a combination of parallel and synthetic data.Synthetic data is often generated by back-translating sentences randomly sampled from monolingual data using a reverse translation model.While backtranslation has been shown to be very effective in many cases, it is not entirely clear why.In this work, we explore different aspects of back-translation, and show that words with high prediction loss during training benefit most from the addition of synthetic data.We introduce several variations of sampling strategies targeting difficult-to-predict words using prediction losses and frequencies of words.In addition, we also target the contexts of difficult words and sample sentences that are similar in context.Experimental results for the WMT news translation task show that our method improves translation quality by up to 1.7 and 1.2 BLEU points over back-translation using random sampling for GermanÑEnglish and EnglishÑGerman, respectively. Marzieh Fadaee, Christof Monz |
EMNLP | 2 |
| 2018 | The importance of Being Recurrent for Modeling Hierarchical StructureabstractRecent work has shown that recurrent neural networks (RNNs) can implicitly capture and exploit hierarchical information when trained to solve common natural language processing tasks (Blevins et al., 2018) such as language modeling (Linzen et al., 2016;Gulordava et al., 2018) and neural machine translation (Shi et al., 2016).In contrast, the ability to model structured data with non-recurrent neural networks has received little attention despite their success in many NLP tasks (Gehring et al., 2017;Vaswani et al., 2017).In this work, we compare the two architectures-recurrent versus non-recurrent-with respect to their ability to model hierarchical structure and find that recurrency is indeed important for this purpose.The code and data used in our experiments is available at https://github.com/ ketranm/fan_vs_rnn Ke M. Tran, Arianna Bisazza, Christof Monz |
EMNLP | 3 |
| 2018 | Examining the Tip of the Iceberg: A Data Set for Idiom Translation
Marzieh Fadaee, Arianna Bisazza, Christof Monz |
LREC | 3 |
| 2018 | Evaluation of Machine Translation Performance Across Multiple Genres and Languages
Marlies van der Wees, Arianna Bisazza, Christof Monz |
LREC | 3 |
| 2017 | Dynamic Data Selection for Neural Machine TranslationabstractIntelligent selection of training data has proven a successful technique to simultaneously increase training efficiency and translation performance for phrase-based machine translation (PBMT).With the recent increase in popularity of neural machine translation (NMT), we explore in this paper to what extent and how NMT can also benefit from data selection.While state-of-the-art data selection (Axelrod et al., 2011) consistently performs well for PBMT, we show that gains are substantially lower for NMT.Next, we introduce dynamic data selection for NMT, a method in which we vary the selected subset of training data between different training epochs.Our experiments show that the best results are achieved when applying a technique we call gradual fine-tuning, with improvements up to +2.6 BLEU over the original data selection approach and up to +3.1 BLEU over a general baseline. Marlies van der Wees, Arianna Bisazza, Christof Monz |
EMNLP | 3 |
| 2017 | What does Attention in Neural Machine Translation Pay Attention to?abstractAttention in neural machine translation provides the possibility to encode relevant parts of the source sentence at each translation step. As a result, attention is considered to be an alignment model as well. However, there is no work that specifically studies attention and provides analysis of what is being learned by attention models. Thus, the question still remains that how attention is similar or different from the traditional alignment. In this paper, we provide detailed analysis of attention and compare it to traditional alignment. We answer the question of whether attention is only capable of modelling translational equivalent or it captures more information. We show that attention is different from alignment in some cases and is capturing useful information other than alignments. Hamidreza Ghader, Christof Monz |
IJCNLP(1) | 2 |
| 2017 | Fine-Tuning for Neural Machine Translation with Limited Degradation across In- and Out-of-Domain Data
Praveen Dakwale, Christof Monz |
MTSummit (1) | 2 |
| 2016 | Ensemble Learning for Multi-Source Neural Machine TranslationabstractIn this paper we describe and evaluate methods to perform ensemble prediction in neural machine translation (NMT). We compare two methods of ensemble set induction: sampling parameter initializations for an NMT system, which is a relatively established method in NMT (Sutskever et al., 2014), and NMT systems translating from different source languages into the same target language, i.e., multi-source ensembles, a method recently introduced by Firat et al. (2016). We are motivated by the observation that for different language pairs systems make different types of mistakes. We propose several methods with different degrees of parameterization to combine individual predictions of NMT systems so that they mutually compensate for each other’s mistakes and improve overall performance. We find that the biggest improvements can be obtained from a context-dependent weighting scheme for multi-source ensembles. This result offers stronger support for the linguistic motivation of using multi-source ensembles than previous approaches. Evaluation is carried out for German and French into English translation. The best multi-source ensemble method achieves an improvement of up to 2.2 BLEU points over the strongest single-source ensemble baseline, and a 2 BLEU improvement over a multi-source ensemble baseline. Ekaterina Garmash, Christof Monz |
COLING | 2 |
| 2016 | Measuring the Effect of Conversational Aspects on Machine Translation QualityabstractResearch in statistical machine translation (SMT) is largely driven by formal translation tasks, while translating informal text is much more challenging. In this paper we focus on SMT for the informal genre of dialogues, which has rarely been addressed to date. Concretely, we investigate the effect of dialogue acts, speakers, gender, and text register on SMT quality when translating fictional dialogues. We first create and release a corpus of multilingual movie dialogues annotated with these four dialogue-specific aspects. When measuring translation performance for each of these variables, we find that BLEU fluctuations between their categories are often significantly larger than randomly expected. Following this finding, we hypothesize and show that SMT of fictional dialogues benefits from adaptation towards dialogue acts and registers. Finally, we find that male speakers are harder to translate and use more vulgar language than female speakers, and that vulgarity is often not preserved during translation. Marlies van der Wees, Arianna Bisazza, Christof Monz |
COLING | 3 |
| 2016 | Recurrent Memory Networks for Language ModelingabstractRecurrent Neural Networks (RNNs) have obtained excellent result in many natural language processing (NLP) tasks.However, understanding and interpreting the source of this success remains a challenge.In this paper, we propose Recurrent Memory Network (RMN), a novel RNN architecture, that not only amplifies the power of RNN but also facilitates our understanding of its internal functioning and allows us to discover underlying patterns in data.We demonstrate the power of RMN on language modeling and sentence completion tasks.On language modeling, RMN outperforms Long Short-Term Memory (LSTM) network on three large German, Italian, and English dataset.Additionally we perform indepth analysis of various linguistic dimensions that RMN captures.On Sentence Completion Challenge, for which it is essential to capture sentence coherence, our RMN obtains 69.2% accuracy, surpassing the previous state of the art by a large margin. 1 Ke M. Tran, Arianna Bisazza, Christof Monz |
HLT-NAACL | 3 |
| 2016 | State of the art in statistical methods for language and speech processingabstractRecent years have seen rapid growth in the deployment of statistical methods for computational language and speech processing. The current popularity of such methods can be traced to the convergence of several factors, including the increasing amount of data now accessible, sustained advances in computing power and storage capabilities, and ongoing improvements in machine learning algorithms. The purpose of this contribution is to review the state of the art in both areas, point out the top trends in statistical modelling across a wide range of problems, and identify their most salient characteristics. The paper concludes with some prognostications regarding the likely impact on the field going forward. Jerome R. Bellegarda, Christof Monz |
Comput. Speech Lang. | 2 |
| 2015 | Bilingual Structured Language Models for Statistical Machine TranslationabstractThis paper describes a novel target-side syntactic language model for phrase-based statistical machine translation, bilingual structured language model.Our approach represents a new way to adapt structured language models (Chelba and Jelinek, 2000) to statistical machine translation, and a first attempt to adapt them to phrasebased statistical machine translation.We propose a number of variations of the bilingual structured language model and evaluate them in a series of rescoring experiments.Rescoring of 1000-best translation lists produces statistically significant improvements of up to 0.7 BLEU over a strong baseline for Chinese-English, but does not yield improvements for Arabic-English. Ekaterina Garmash, Christof Monz |
EMNLP | 2 |
| 2015 | A distributed inflection model for translating into morphologically rich languages
Ke M. Tran, Arianna Bisazza, Christof Monz |
MTSummit | 3 |
| 2014 | Class-Based Language Modeling for Translating into Morphologically Rich Languages
Arianna Bisazza, Christof Monz |
COLING | 2 |
| 2014 | Maximizing Component Quality in Bilingual Word-Aligned SegmentationsabstractGiven a pair of source and target language sentences which are translations of each other with known word alignments between them, we extract bilingual phrase-level segmentations of such a pair. This is done by identifying two appropriate measures that assess the quality of phrase segments, one on the monolingual level for both language sides, and one on the bilingual level. The monolingual measure is based on the notion of partition refinements and the bilingual measure is based on structural properties of the graph that represents phrase segments and word alignments. These two measures are incorporated in a basic adaptation of the Cross-Entropy method for the purpose of extracting an N-best list of bilingual phrase-level segmentations. A straight-forward application of such lists in Statistical Machine Translation (SMT) yields a conservative phrase pair extraction method that reduces phrase-table sizes by 90% with insignificant loss in translation quality. Spyros Martzoukos, Christof Monz, Christophe Costa Florêncio |
EACL | 2 |
| 2014 | Dependency-Based Bilingual Language Models for Reordering in Statistical Machine TranslationabstractThis paper presents a novel approach to improve reordering in phrase-based ma-chine translation by using richer, syntac-tic representations of units of bilingual language models (BiLMs). Our method to include syntactic information is simple in implementation and requires minimal changes in the decoding algorithm. The approach is evaluated in a series of Arabic-English and Chinese-English translation experiments. The best models demon-strate significant improvements in BLEU and TER over the phrase-based baseline, as well as over the lexicalized BiLM by Niehues et al. (2011). Further improve-ments of up to 0.45 BLEU for Arabic-English and up to 0.59 BLEU for Chinese-English are obtained by combining our de-pendency BiLM with a lexicalized BiLM. An improvement of 0.98 BLEU is ob-tained for Chinese-English in the setting of an increased distortion limit. 1 Ekaterina Garmash, Christof Monz |
EMNLP | 2 |
| 2014 | Word Translation Prediction for Morphologically Rich Languages with Bilingual Neural NetworksabstractTranslating into morphologically rich lan-guages is a particularly difficult problem in machine translation due to the high de-gree of inflectional ambiguity in the tar-get language, often only poorly captured by existing word translation models. We present a general approach that exploits source-side contexts of foreign words to improve translation prediction accuracy. Our approach is based on a probabilistic neural network which does not require lin-guistic annotation nor manual feature en-gineering. We report significant improve-ments in word translation prediction accu-racy for three morphologically rich target languages. In addition, preliminary results for integrating our approach into a large-scale English-Russian statistical machine translation system show small but statisti-cally significant improvements in transla-tion quality. 1 Ke M. Tran, Arianna Bisazza, Christof Monz |
EMNLP | 3 |
| 2012 | User Edits Classification Using Document Revision Histories
Amit Bronner, Christof Monz |
EACL | 2 |
| 2012 | Power-Law Distributions for Paraphrases Extracted from Bilingual Corpora
Spyros Martzoukos, Christof Monz |
EACL | 2 |
| 2012 | Adaptation of Statistical Machine Translation Model for Cross-Lingual Information Retrieval in a Service Context
Vassilina Nikoulina, Bogomil Kovachev, Nikolaos Lagos, Christof Monz |
EACL | 4 |
| 2011 | Statistical Machine Translation with Local Language Models
Christof Monz |
EMNLP | 1 |
| 2011 | Syntactic discriminative language model rerankers for statistical machine translationabstractThis article describes a method that successfully exploits syntactic features for n-best translation candidate reranking using perceptrons. We motivate the utility of syntax by demonstrating the superior performance of parsers over n-gram language models in differentiating between Statistical Machine Translation output and human translations. Our approach uses discriminative language modelling to rerank the n-best translations generated by a statistical machine translation system. The performance is evaluated for Arabic-to-English translation using NIST’s MT-Eval benchmarks. While deep features extracted from parse trees do not consistently help, we show how features extracted from a shallow Part-of-Speech annotation layer outperform a competitive baseline and a state-of-the-art comparative reranking approach, leading to significant BLEU improvements on three different test sets. Simon Carter, Christof Monz |
Mach. Transl. | 2 |
| 2011 | Machine learning for query formulation in question answeringabstractAbstract Research on question answering dates back to the 1960s but has more recently been revisited as part of TREC's evaluation campaigns, where question answering is addressed as a subarea of information retrieval that focuses on specific answers to a user's information need. Whereas document retrieval systems aim to return the documents that are most relevant to a user's query, question answering systems aim to return actual answers to a users question. Despite this difference, question answering systems rely on information retrieval components to identify documents that contain an answer to a user's question. The computationally more expensive answer extraction methods are then applied only to this subset of documents that are likely to contain an answer. As information retrieval methods are used to filter the documents in the collection, the performance of this component is critical as documents that are not retrieved are not analyzed by the answer extraction component. The formulation of queries that are used for retrieving those documents has a strong impact on the effectiveness of the retrieval component. In this paper, we focus on predicting the importance of terms from the original question. We use model tree machine learning techniques in order to assign weights to query terms according to their usefulness for identifying documents that contain an answer. Term weights are learned by inspecting a large number of query formulation variations and their respective accuracy in identifying documents containing an answer. Several linguistic features are used for building the models, including part-of-speech tags, degree of connectivity in the dependency parse tree of the question, and ontological information. All of these features are extracted automatically by using several natural language processing tools. Incorporating the learned weights into a state-of-the-art retrieval system results in statistically significant improvements in identifying answer-bearing documents. Christof Monz |
Nat. Lang. Eng. | 1 |
| 2009 | Automatic Single-Document Key Fact Extraction from Newswire Articles
Itamar Kastner, Christof Monz |
EACL | 2 |
| 2009 | Decoding by Dynamic Chunking for Statistical Machine Translation
Sirvan Yahyaei, Christof Monz |
MTSummit | 2 |
| 2009 | A comparison of retrieval-based hierarchical clustering approaches to person name disambiguationabstractThis paper describes a simple clustering approach to person name disambiguation of retrieved documents. The methods are based on standard IR concepts and do not require any task-specific features. We compare different term-weighting and indexing methods and evaluate their performance against the Web People Search task (WePS). Despite their simplicity these approaches achieve very competitive performance. Christof Monz, Wouter Weerkamp |
SIGIR | 1 |
| 2009 | Symbolic-to-statistical hybridization: extending generation-heavy machine translationabstractThe last few years have witnessed an increasing interest in hybridizing surface-based statistical approaches and rule-based symbolic approaches to machine translation (MT). Much of that work is focused on extending statistical MT systems with symbolic knowledge and components. In the brand of hybridization discussed here, we go in the opposite direction: adding statistical bilingual components to a symbolic system. Our base system is Generation-heavy machine translation (GHMT), a primarily symbolic asymmetrical approach that addresses the issue of Interlingual MT resource poverty in source-poor/target-rich language pairs by exploiting symbolic and statistical target-language resources. GHMT’s statistical components are limited to target-language models, which arguably makes it a simple form of a hybrid system . We extend the hybrid nature of GHMT by adding statistical bilingual components. We also describe the details of retargeting it to Arabic–English MT. The morphological richness of Arabic brings several challenges to the hybridization task. We conduct an extensive evaluation of multiple system variants. Our evaluation shows that this new variant of GHMT—a primarily symbolic system extended with monolingual and bilingual statistical components—has a higher degree of grammaticality than a phrase-based statistical MT system, where grammaticality is measured in terms of correct verb-argument realization and long-distance dependency translation. Nizar Habash, Bonnie J. Dorr, Christof Monz |
Mach. Transl. | 3 |
| 2008 | Applying Maximum Entropy to Known-Item Email Retrieval
Sirvan Yahyaei, Christof Monz |
ECIR | 2 |
| 2007 | Model Tree Learning for Query Term Weighting in Question Answering
Christof Monz |
ECIR | 1 |
| 2007 | Task-based evaluation of text summarization using Relevance Prediction
Stacy Hobson, Bonnie J. Dorr, Christof Monz, Richard M. Schwartz |
Inf. Process. Manag. | 3 |
| 2005 | Iterative translation disambiguation for cross-language information retrievalabstractFinding a proper distribution of translation probabilities is one of the most important factors impacting the effectiveness of a cross-language information retrieval system. In this paper we present a new approach that computes translation probabilities for a given query by using only a bilingual dictionary and a monolingual corpus in the target language. The algorithm combines term association measures with an iterative machine learning approach based on expectation maximization. Our approach considers only pairs of translation candidates and is therefore less sensitive to data-sparseness issues than approaches using higher n-grams. The learned translation probabilities are used as query term weights and integrated into a vector-space retrieval system. Results for English-German cross-lingual retrieval show substantial improvements over a baseline using dictionary lookup without term weighting. Christof Monz, Bonnie J. Dorr |
SIGIR | 1 |
| 2004 | Monolingual Document Retrieval for European Languages
Vera Hollink, Jaap Kamps, Christof Monz, Maarten de Rijke |
Inf. Retr. | 3 |
| 2003 | Document Retrieval in the Context of Question Answering
Christof Monz |
ECIR | 1 |
| 2002 | Document understanding for a broad class of documents
Marco Aiello 0001, Christof Monz, Leon Todoran |
Int. J. Document Anal. Recognit. | 2 |
| 1999 | A Tableau Calculus for Pronoun Resolution
Christof Monz, Maarten de Rijke |
TABLEAUX | 1 |
| 1998 | Dynamic Semantics and Underspecification
Christof Monz |
ECAI | 1 |
| 1998 | A Tableaux Calculus for Ambiguous Quantification
Christof Monz, Maarten de Rijke |
TABLEAUX | 1 |