EDBT 2026 Demo / reviewers in the wild / expert
Rachel Bawden
dblp:146/4432
· DBLP profile ↗
30ranked-venue papers
9as first author
23since 2021 · last 2026
0000-0001-9553-1768ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 30 · 9 first-author · 23 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MetaDocEval: A Contrastive Framework for Evaluating Machine Translation Metrics at the Document-LevelabstractRecent advances in neural machine translation (MT) have spurred increased interest in evaluating translations beyond the sentence level, making it possible to assess discourse-level phenomena related to coherence and consistency. While existing metrics can be applied to multi-sentence spans, it remains unclear whether their scores truly capture document-level quality. We introduce MetaDocEval, an automatic contrastive test set for evaluating MT metrics across three language pairs (en–fr, en–es, en–de) when applied at the document-level. It targets a range of discourse-level phenomena and potential problems linked to translation at the document level. To evaluate how metrics behave as a function of context size, we apply them under a sliding-window protocol, varying the input from single sentences up to full documents. Our experiments show that no current metric genuinely captures document-level coherence: reference-based metrics overfit lexical overlap, reference+source metrics gain little from added context, reference-free encoders show brief context sensitivity before degrading on longer spans, and LLM-based scorers collapse beyond short inputs. A key finding is that reference access can be actively harmful for detecting discourse-level errors. Using short windows (≈ 3 sentences) offers the best trade-off between discourse error detection and score dilution. Nicolas Dahan, Rachel Bawden, François Yvon |
EAMT (1) | 2 |
| 2026 | When the Gold Standard Isn't Necessarily Standard: Challenges of Evaluating the Translation of User-Generated ContentabstractUser-generated content (UGC) is characterised by frequent use of non-standard language, from spelling errors to expressive choices such as slang, character repetitions, and emojis. This makes evaluating UGC translation challenging: what counts as a "good" translation depends on the desired standardness level of the output. To explore this, we examine the human translation guidelines of four UGC datasets, and derive a taxonomy of twelve non-standard phenomena and five translation actions (NORMALISE, COPY, TRANSFER, OMIT, CENSOR). Our analysis reveals notable differences in how UGC is treated, resulting in a spectrum of standardness in reference translations. We show that translation scores of large language models are highly sensitive to prompts with explicit UGC translation instructions, and that they improve when they align with the dataset guidelines. We argue that fair evaluation requires both models and metrics to be aware of translation guidelines. Finally, we call for clear guidelines during dataset creation and for the development of controllable, guideline-aware evaluation frameworks for UGC translation. Lydia Nishimwe, Benoît Sagot, Rachel Bawden |
EAMT (1) | 3 |
| 2026 | The MaTOS Pipeline for the Translation of Scientific Abstracts on the HAL PlatformabstractEnglish dominates scientific publishing, which disadvantages researchers who are not native English speakers, especially those in the earlier stages of their careers. Being able to write and engage with scientific content written in their own language would clearly facilitate scientific production. The MaTOS project (Machine Translation for Open Science) seeks to reduce these barriers by developing machine translation tools for scientific documents in English and French. This article presents the design of the MaTOS pipeline for the HAL platform to automatically translate article abstracts, with author validation, to increase the number of bilingual abstracts on the platform. We also report preliminary experiments comparing translation of sentence, three-sentence chunks, and whole abstracts, evaluated using quality estimation metrics. Panagiotis Tsolakis, Ziqian Peng, Laurent Romary, François Yvon, Rachel Bawden |
EAMT (2) | 5 |
| 2026 | Hindsight Quality Prediction Experiments in Multi-Candidate Human-Post-Edited Machine TranslationabstractInternational audience Malik Marmonier, Benoît Sagot, Rachel Bawden |
LREC | 3 |
| 2026 | ForumOccitania: A Corpus of User-Generated Content for Multiple Occitan VarietiesabstractAccepted at LREC 2026 Oriane Nédey, Juliette Janes, Rachel Bawden, Thibault Clérice, Benoît Sagot |
LREC | 3 |
| 2025 | AFRIDOC-MT: Document-level MT Corpus for African LanguagesabstractJesujoba Oluwadara Alabi, Israel Abebe Azime, Miaoran Zhang, Cristina España-Bonet, Rachel Bawden, Dawei Zhu, David Ifeoluwa Adelani, Clement Oyeleke Odoje, Idris Akinade, Iffat Maab, Davis David, Shamsuddeen Hassan Muhammad, Neo Putini, David O. Ademuyiwa, Andrew Caines, Dietrich Klakow. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Jesujoba O. Alabi, Israel Abebe Azime, Miaoran Zhang, Cristina España-Bonet, Rachel Bawden, David Ifeoluwa Adelani, Clement Odoje, Idris Akinade, Iffat Maab, Davis David, Shamsuddeen Hassan Muhammad, Neo Putini, David O. Ademuyiwa, Andrew Caines, Dietrich Klakow |
EMNLP | 5 |
| 2025 | Explicit Learning and the LLM in Machine TranslationabstractThis study explores an LLM's ability to learn new languages using explanations found in a grammar book-a process we term "explicit learning."To rigorously assess this ability, we design controlled translation experiments between English and constructed languages generated-through specific cryptographic means-from Latin or French.Contrary to previous studies, our results demonstrate that LLMs do possess a measurable capacity for explicit learning.This ability, however, diminishes as the complexity of the linguistic phenomena to be learned increases.Supervised fine-tuning on ad hoc chains of thought significantly enhances LLM performance but struggles to generalize to typologically novel or more complex linguistic features.These findings point to the need for more diverse training sets and alternative fine-tuning strategies to further improve explicit learning by LLMs, benefiting low-resource languages typically described in grammar books but lacking extensive corpora. Malik Marmonier, Rachel Bawden, Benoît Sagot |
EMNLP | 2 |
| 2025 | MaTOS: Machine Translation for Open ScienceabstractThis paper is a short presentation of MaTOS, a project focusing on the automatic translation of scholarly documents. Its main aims are threefold: (a) to develop resources (term lists and corpora) for high-quality machine translation; (b) to study methods for translating complete, structured documents in a cohesive and consistent manner; (c) to propose novel metrics to evaluate machine translation in technical domains. Publications and resources are available on the project web site: https://anr-matos.gihub.io. Rachel Bawden, Maud Bénard, José Cornejo Cárcamo, Nicolas Dahan, Manon Delorme, Mathilde Huguin, Natalie Kübler, Paul Lerner, Alexandra Mestivier, Joachim Minder, Jean-François Nominé, Ziqian Peng, Laurent Romary, Panagiotis Tsolakis, Lichao Zhu, François Yvon |
MTSummit (2) | 1 |
| 2025 | Investigating Length Issues in Document-level Machine TranslationabstractTransformer architectures are increasingly effective at processing and generating very long chunks of texts, opening new perspectives for document-level machine translation (MT). In this work, we challenge the ability of MT systems to handle texts comprising up to several thousands of tokens. We design and implement a new approach designed to precisely measure the effect of length increments on MT outputs. Our experiments with two representative architectures unambiguously show that (a) translation performance decreases with the length of the input text; (b) the position of sentences within the document matters and translation quality is higher for sentences occurring earlier in a document. We further show that manipulating the distribution of document lengths and of positional embeddings only marginally mitigates such problems. Our results suggest that even though document-level MT is computationally feasible, it does not yet match the performance of sentence-based MT. Ziqian Peng, Rachel Bawden, François Yvon |
MTSummit (1) | 2 |
| 2024 | When Your Cousin Has the Right Connections: Unsupervised Bilingual Lexicon Induction for Related Data-Imbalanced LanguagesabstractMost existing approaches for unsupervised bilingual lexicon induction (BLI) depend on good quality static or contextual embeddings requiring large monolingual corpora for both languages. However, unsupervised BLI is most likely to be useful for low-resource languages (LRLs), where large datasets are not available. Often we are interested in building bilingual resources for LRLs against related high-resource languages (HRLs), resulting in severely imbalanced data settings for BLI. We first show that state-of-the-art BLI methods in the literature exhibit near-zero performance for severely data-imbalanced language pairs, indicating that these settings require more robust techniques. We then present a new method for unsupervised BLI between a related LRL and HRL that only requires inference on a masked language model of the HRL, and demonstrate its effectiveness on truly low-resource languages Bhojpuri and Magahi (with <5M monolingual tokens each), against Hindi. We further present experiments on (mid-resource) Marathi and Nepali to compare approach performances by resource range, and release our resulting lexicons for five low-resource Indic languages: Bhojpuri, Magahi, Awadhi, Braj, and Maithili, against Hindi. Niyati Bafna, Cristina España-Bonet, Josef van Genabith, Benoît Sagot, Rachel Bawden |
LREC/COLING | 5 |
| 2024 | Making Sentence Embeddings Robust to User-Generated ContentabstractNLP models have been known to perform poorly on user-generated content (UGC), mainly because it presents a lot of lexical variations and deviates from the standard texts on which most of these models were trained. In this work, we focus on the robustness of LASER, a sentence embedding model, to UGC data. We evaluate this robustness by LASER’s ability to represent non-standard sentences and their standard counterparts close to each other in the embedding space. Inspired by previous works extending LASER to other languages and modalities, we propose RoLASER, a robust English encoder trained using a teacher-student approach to reduce the distances between the representations of standard and UGC sentences. We show that with training only on standard and synthetic UGC-like data, RoLASER significantly improves LASER’s robustness to both natural and artificial UGC data by achieving up to 2x and 11x better scores. We also perform a fine-grained analysis on artificial UGC data and find that our model greatly outperforms LASER on its most challenging UGC phenomena such as keyboard typos and social media abbreviations. Evaluation on downstream tasks shows that RoLASER performs comparably to or better than LASER on standard data, while consistently outperforming it on UGC data. Lydia Nishimwe, Benoît Sagot, Rachel Bawden |
LREC/COLING | 3 |
| 2024 | Translate your Own: a Post-Editing Experiment in the NLP domainabstractThe improvements in neural machine translation make translation and post-editing pipelines ever more effective for a wider range of applications. In this paper, we evaluate the effectiveness of such a pipeline for the translation of scientific documents (limited here to article abstracts). Using a dedicated interface, we collect, then analyse the post-edits of approximately 350 abstracts (English→French) in the Natural Language Processing domain for two groups of post-editors: domain experts (academics encouraged to post-edit their own articles) on the one hand and trained translators on the other. Our results confirm that such pipelines can be effective, at least for high-resource language pairs. They also highlight the difference in the post-editing strategy of the two subgroups. Finally, they suggest that working on term translation is the most pressing issue to improve fully automatic translations, but that in a post-editing setup, other error types can be equally annoying for post-editors. Rachel Bawden, Ziqian Peng, Maud Bénard, Éric Villemonte de la Clergerie, Raphaël Esamotunu, Mathilde Huguin, Natalie Kübler, Alexandra Mestivier, Mona Michelot, Laurent Romary, Lichao Zhu, François Yvon |
EAMT (1) | 1 |
| 2024 | Tree of Problems: Improving structured problem solving with compositionalityabstractLarge Language Models (LLMs) have demonstrated remarkable performance across multiple tasks through in-context learning.For complex reasoning tasks that require step-by-step thinking, Chain-of-Thought (CoT) prompting has given impressive results, especially when combined with self-consistency.Nonetheless, some tasks remain particularly difficult for LLMs to solve.Tree of Thoughts (ToT) and Graph of Thoughts (GoT) emerged as alternatives, dividing the complex problem into paths of subproblems.In this paper, we propose Tree of Problems (ToP), a simpler version of ToT, which we hypothesise can work better for complex tasks that can be divided into identical subtasks.Our empirical results show that our approach outperforms ToT and GoT, and in addition performs better than CoT on complex reasoning tasks.All code for this paper is publicly available here: https://github. com/ArmelRandy/tree-of-problems. Armel Zebaze, Benoît Sagot, Rachel Bawden |
EMNLP | 3 |
| 2023 | Tackling Ambiguity with Images: Improved Multimodal Machine Translation and Contrastive EvaluationabstractOne of the major challenges of machine translation (MT) is ambiguity, which can in some cases be resolved by accompanying context such as images.However, recent work in multimodal MT (MMT) has shown that obtaining improvements from images is challenging, limited not only by the difficulty of building effective cross-modal representations, but also by the lack of specific evaluation and training data.We present a new MMT approach based on a strong text-only MT model, which uses neural adapters, a novel guided self-attention mechanism and which is jointly trained on both visually-conditioned masking and MMT.We also introduce CoMMuTE, a Contrastive Multilingual Multimodal Translation Evaluation set of ambiguous sentences and their possible translations, accompanied by disambiguating images corresponding to each translation.Our approach obtains competitive results compared to strong text-only models on standard English→French, English→German and English→Czech benchmarks and outperforms baselines and state-of-the-art MMT systems by a large margin on our contrastive test set.Our code 1 and CoMMuTE 2 are freely available.8 mBART is pretrained on CC25 (Wenzek et al., 2020). Matthieu Futeral, Cordelia Schmid, Ivan Laptev, Benoît Sagot, Rachel Bawden |
ACL (1) | 5 |
| 2023 | Investigating the Translation Performance of a Large Multilingual Language Model: the Case of BLOOMabstractThe NLP community recently saw the release of a new large open-access multilingual language model, BLOOM (BigScience et al., 2022) covering 46 languages. We focus on BLOOM’s multilingual ability by evaluating its machine translation performance across several datasets (WMT, Flores-101 and DiaBLa) and language pairs (high- and low-resourced). Our results show that 0-shot performance suffers from overgeneration and generating in the wrong language, but this is greatly improved in the few-shot setting, with very good results for a number of language pairs. We study several aspects including prompt design, model sizes, cross-lingual transfer and the use of discursive context. Rachel Bawden, François Yvon |
EAMT | 1 |
| 2023 | Investigating Lexical Sharing in Multilingual Machine Translation for Indian LanguagesabstractMultilingual language models have shown impressive cross-lingual transfer ability across a diverse set of languages and tasks. To improve the cross-lingual ability of these models, some strategies include transliteration and finer-grained segmentation into characters as opposed to subwords. In this work, we investigate lexical sharing in multilingual machine translation (MT) from Hindi, Gujarati, Nepali into English. We explore the trade-offs that exist in translation performance between data sampling and vocabulary size, and we explore whether transliteration is useful in encouraging cross-script generalisation. We also verify how the different settings generalise to unseen languages (Marathi and Bengali). We find that transliteration does not give pronounced improvements and our analysis suggests that our multilingual MT models trained on original scripts are already robust to cross-script differences even for relatively low-resource languages. Sonal Sannigrahi, Rachel Bawden |
EAMT | 2 |
| 2022 | Multitask Prompted Training Enables Zero-Shot Task Generalization
Victor Sanh, Albert Webson, Colin Raffel, Stephen H. Bach, Lintang Sutawika, Zaid Alyafeai, Antoine Chaffin, Arnaud Stiegler, Arun Raja, Manan Dey, M Saiful Bari, Canwen Xu, Urmish Thakker, Shanya Sharma Sharma, Eliza Szczechla, Taewoon Kim 0002, Gunjan Chhablani, Nihal V. Nayak, Debajyoti Datta, Mike Tian-Jian Jiang, Matteo Manica, Sheng Shen 0001, Harshit Pandey, Rachel Bawden, Thomas Wang, Trishala Neeraj, Jos Rozen, Abheesht Sharma, Andrea Santilli, Thibault Févry, Jason Alan Fries, Ryan Teehan, Teven Le Scao, Stella Biderman, Leo Gao, Thomas Wolf 0008, Alexander M. Rush |
ICLR | 27 |
| 2022 | Automatic Normalisation of Early Modern FrenchabstractSpelling normalisation is a useful step in the study and analysis of historical language texts, whether it is manual analysis by experts or automatic analysis using downstream natural language processing (NLP) tools. Not only does it help to homogenise the variable spelling that often exists in historical texts, but it also facilitates the use of off-the-shelf contemporary NLP tools, if contemporary spelling conventions are used for normalisation. We present FREEMnorm, a new benchmark for the normalisation of Early Modern French (from the 17th century) into contemporary French and provide a thorough comparison of three different normalisation methods: ABA, an alignment-based approach and MT-approaches, (both statistical and neural), including extensive parameter searching, which is often missing in the normalisation literature. Rachel Bawden, Jonathan Poinhos, Eleni Kogkitsidou, Philippe Gambette, Benoît Sagot, Simon Gabay |
LREC | 1 |
| 2022 | Complex Labelling and Similarity Prediction in Legal Texts: Automatic Analysis of France's Court of Cassation RulingsabstractDetecting divergences in the applications of the law (where the same legal text is applied differently by two rulings) is an important task. It is the mission of the French Cour de Cassation. The first step in the detection of divergences is to detect similar cases, which is currently done manually by experts. They rely on summarised versions of the rulings (syntheses and keyword sequences), which are currently produced manually and are not available for all rulings. There is also a high degree of variability in the keyword choices and the level of granularity used. In this article, we therefore aim to provide automatic tools to facilitate the search for similar rulings. We do this by (i) providing automatic keyword sequence generation models, which can be used to improve the coverage of the analysis, and (ii) providing measures of similarity based on the available texts and augmented with predicted keyword sequences. Our experiments show that the predictions improve correlations of automatically obtained similarities against our specially colelcted human judgments of similarity. Thibault Charmet, Inès Cherichi, Matthieu Allain, Urszula Czerwinska, Amaury Fouret, Benoît Sagot, Rachel Bawden |
LREC | 7 |
| 2022 | From FreEM to D'AlemBERT: a Large Corpus and a Language Model for Early Modern Frenchabstractanguage models for historical states of language are becoming increasingly important to allow the optimal digitisation and analysis of old textual sources. Because these historical states are at the same time more complex to process and more scarce in the corpora available, this paper presents recent efforts to overcome this difficult situation. These efforts include producing a corpus, creating the model, and evaluating it with an NLP task currently used by scholars in other ongoing projects. Simon Gabay, Pedro Ortiz Suarez, Alexandre Bartz, Alix Chagué, Rachel Bawden, Philippe Gambette, Benoît Sagot |
LREC | 5 |
| 2022 | Survey of Low-Resource Machine TranslationabstractAbstract We present a survey covering the state of the art in low-resource machine translation (MT) research. There are currently around 7,000 languages spoken in the world and almost all language pairs lack significant resources for training machine translation models. There has been increasing interest in research addressing the challenge of producing useful translation models when very little translated training data is available. We present a summary of this topical research field and provide a description of the techniques evaluated by researchers in several recent shared tasks in low-resource MT. Barry Haddow, Rachel Bawden, Antonio Valerio Miceli Barone, Jindrich Helcl, Alexandra Birch |
Comput. Linguistics | 2 |
| 2021 | Few-shot learning through contextual data augmentationabstractMachine translation (MT) models used in industries with constantly changing topics, such as translation or news agencies, need to adapt to new data to maintain their performance over time.Our aim is to teach a pre-trained MT model to translate previously unseen words accurately, based on very few examples.We propose (i) an experimental setup allowing us to simulate novel vocabulary appearing in human-submitted translations, and (ii) corresponding evaluation metrics to compare our approaches.We extend a data augmentation approach using a pre-trained language model to create training examples with similar contexts for novel words.We compare different fine-tuning and data augmentation approaches and show that adaptation on the scale of one to five examples is possible.Combining data augmentation with randomly selected training sentences leads to the highest BLEU score and accuracy improvements.Impressively, with only 1 to 5 examples, our model reports better accuracy scores than a reference system trained with on average 313 parallel examples. Farid Arthaud, Rachel Bawden, Alexandra Birch |
EACL | 2 |
| 2021 | Understanding Dialogue: Language Use and Social InteractionabstractUnderstanding Dialogue: Language Use and Social Interaction represents a departure from classic theories in psycholinguistics and cognitive sciences; instead of taking as a starting point the isolated speech of an individual that can be extended to accommodate dialogue, a primary focus is put on developing a model adapted to dialogue itself, bearing in mind important aspects of dialogue as an activity with a heavily cooperative component. As a researcher of natural language processing with a background in linguistics, I find highly intriguing the possibilities provided by the dialogue model presented. Although the book does not itself touch upon the potential for automated dialogue, I am inevitably writing this review from the point of view of a computational linguist with these aspects in mind.Building on numerous previous works, including many of the authors’ own studies and theories, Understanding Dialogue presents the shared workspace framework, a framework for understanding not just dialogue but cooperative activities in general, of which dialogue is viewed as a subtype. Based on Bratman’s (1992) concept of shared cooperative activity, the framework provides a joint environment with which interlocutors can interact, both by contributing to the space (with actions or utterances for example), and by perceiving and processing their own or the other participants’ productions. The authors do not limit their work to linguistic communication: Many of their examples, particularly at the beginning of the book, are non-linguistic (e.g., hand shaking, dancing a tango, playing singles tennis); others are primarily physical, but will most likely also involve linguistic communication (such as jointly constructing flat-pack furniture); and others are purely linguistic (e.g., suggesting which restaurant to go to for lunch).The notion of alignment is highly important to this framework both from a linguistic and non-linguistic perspective, and is one of the main inspirations of the book, having previously been presented in Toward a Mechanistic Theory of Dialogue by the same authors. As individuals interact via the joint space, alignment concerns the equivalence in their representations at a conceptual level, with respect to their goals and relevant props in the shared environment (dialogue model alignment) and linguistic representations shared in the workspace (linguistic alignment). Roughly speaking, in this second (linguistic) case, this may for instance correspond to whether or not the individuals have the same representation of the utterance in terms of phonetics (were the sounds perceived correctly?) or in terms of lexical semantics (do they understand the same reference by the word uttered?). From here can be explained a number of different dialogue behaviors linked to the quest for alignment and the resolution of misalignment should it occur.The book is structured in four main parts, preceded by an Introduction presenting the challenges of dialogue and the main ideas behind the framework. The focus of the book is clearly stated from the beginning as being dialogue first, in a rejection of models that seek to study language primarily from a monologic point of view. As the authors point out, the notion of alignment underpinning the framework involves by its very nature multiple participants and therefore dialogic interactions must be studied in their own right. I shall provide only a brief summary of the four parts here, highlighting some components that in my view are key to the model, without however covering all themes, which would require a far more extensive description.Part I introduces the basis of the shared workspace framework as applied to activities with a cooperative component and then specifically to dialogue. The basic sender-receiver framework is quickly rejected, as it lacks the ability to represent certain key ele- ments of cooperative activities, such as allowing for feedback and representing an environment that is common to the participants. The shared workspace framework is then introduced, along with the four important characteristics of cooperative joint aspect systems that can be successfully illustrated with it: alignment (mentioned above), simulation (the representation of an activity without actually going through with it), prediction (the anticipation of participants’ behaviors), and synchrony (concerning the timing of behaviors in a joint activity), elements that are first studied in the context of joint activities in general (Chapter 3), before being reviewed specifically for dialogue (Chapter 4).Part II is dedicated to the aforementioned concept of alignment, fundamental to the framework of cooperative activity. The chapters in this section look at the distinction between the different levels at which alignment can occur, the processes involved, and the consequences of alignment, such as participants uttering similar linguistic productions. Another important notion introduced in this part is that of the meta-representation of alignment, which represents the participants’ belief about how aligned they are, which has inevitable consequences on how they then plan and implement their contributions.Part III continues with the theme of alignment but turns to aspects involving the efficiency of communication: succinctness of formulation (Chapter 8) and how we time our contributions (Chapter 9). Particularly interesting is the role of commentaries, which are contributions providing some sort of feedback on the alignment of participants and which can therefore affect the participants’ meta-representation of alignment. There is an important distinction between positive and negative commentaries, positive commentaries (such as “uh huh” in English) providing feedback that the speaker is aligned, therefore enabling the participants to meta-represent alignment, and negative ones (such as “huh?”) indicating a misalignment, but then enabling the participants to recover from that it. These commentaries contribute to the succinctness of dialogue and to maximizing the efficiency of joint participation by indicating meta-alignment. Finally, Chapter 9 discusses the notion of “speaking in good time,” related to the necessarily sequential nature of dialogue and the importance of timing, including the effects of different speech rates and the natural adaptation that occurs between interlocutors.Part IV looks beyond the main theme of dialogue to other forms of conversation, including multiparty conversations and collectives, exploring the possible roles of the different participants, and how this relates back to alignment and their contribution to the shared workspace. Also mentioned is monologue and the challenges that it poses with respect to the primary and more natural form of language communication that is dialogue. The final chapter introduces how the shared workspace can be augmented by adding props, illustrations, and recordings and by using alternative communicative tools, such as text messages and social media, which come with their own constraints with respect to the access they allow to the shared workspace.The description of the framework is thorough and well exemplified, with a continuity in the use of examples throughout the book. A repetition and embellishment of schemas helps to keep track of how the new additions from each chapter fit into the framework. I found some of the descriptions a little wordy, particularly because of the reiteration of definitions and motivations, and in the minutely detailed illustration of examples. However, from the point of view of pedagogy, this could be seen as adding clarity, particularly for the reader who decides to focus on particular chapters rather than reading the book from cover to cover. In my opinion, the book will be highly accessible to all readers, even those who have limited background on the topic, and the authors take care to make it clear how their framework and definitions agree with or differ from previous works.For me, there remain two main areas that could have been worthy of further exploration within the scope of this book. The first is the effect of cultural and linguistic differences. The authors do address the topic in Chapter 11, but in comparison with the detail afforded to the description of the framework, this subject remains rather lacking, with only a short section touching on it, under the the title of Social Norms and Joint Planning. The authors cite an interesting study by Fujii (2012) on the differences between American and Japanese speakers in terms of their use of language to foster alignment. However, this teaser does not lead on to a deeper discussion about cross-cultural differences as explained in terms of the concepts used in this framework. The second topic is the link to sign languages, which would appear to link more than perfectly with the shared workspace framework and yet is not mentioned by the authors.There is little doubt that the framework is an important step in modeling dialogue from a psycholinguistic perspective. As a researcher in natural language processing, I would be excited to see the the possibilities for this framework in a computational setting for automated dialogue, something that the authors mention in their conclusion. They evoke the failure of current chatbots such as Siri and Alexa to effectively dialogue due to their inability to provide commentary (e.g., in the context of an ambiguous question) and to meta-represent alignment (i.e., have an opinion on whether the representations of the dialogue participants are the same). They suggest that this framework could help provide the solution to the current disruptions in communication we meet when interacting with these systems. I therefore look forward to seeing what progress can be made from this point of view. Rachel Bawden |
Comput. Linguistics | 1 |
| 2020 | Document-level Neural MT: A Systematic ComparisonabstractIn this paper we provide a systematic comparison of existing and new document-level neural machine translation solutions. As part of this comparison, we introduce and evaluate a document-level variant of the recently proposed Star Transformer architecture. In addition to using the traditional metric BLEU, we report the accuracy of the models in handling anaphoric pronoun translation as well as coherence and cohesion using contrastive test sets. Finally, we report the results of human evaluation in terms of Multidimensional Quality Metrics (MQM) and analyse the correlation of the results obtained by the automatic metrics with human judgments. António V. Lopes, M. Amin Farajian, Rachel Bawden, André F. T. Martins |
EAMT | 3 |
| 2020 | Document Sub-structure in Neural Machine TranslationabstractCurrent approaches to machine translation (MT) either translate sentences in isolation, disregarding the context they appear in, or model context at the level of the full document, without a notion of any internal structure the document may have. In this work we consider the fact that documents are rarely homogeneous blocks of text, but rather consist of parts covering different topics. Some documents, such as biographies and encyclopedia entries, have highly predictable, regular structures in which sections are characterised by different topics. We draw inspiration from Louis and Webber (2014) who use this information to improve statistical MT and transfer their proposal into the framework of neural MT. We compare two different methods of including information about the topic of the section within which each sentence is found: one using side constraints and the other using a cache-based model. We create and release the data on which we run our experiments - parallel corpora for three language pairs (Chinese-English, French-English, Bulgarian-English) from Wikipedia biographies, which we extract automatically, preserving the boundaries of sections within the articles. Radina Dobreva, Rachel Bawden |
LREC | 3 |
| 2019 | Global Under-Resourced Media Translation (GoURMET)
Alexandra Birch, Barry Haddow, Ivan Titov 0001, Antonio Valerio Miceli Barone, Rachel Bawden, Felipe Sánchez-Martínez, Mikel L. Forcada, Miquel Esplà-Gomis, Víctor M. Sánchez-Cartagena, Juan Antonio Pérez-Ortiz, Wilker Aziz, Andrew Secker, Peggy van der Kreeft |
MTSummit (2) | 5 |
| 2018 | Evaluating Discourse Phenomena in Neural Machine TranslationabstractRachel Bawden, Rico Sennrich, Alexandra Birch, Barry Haddow. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Rachel Bawden, Rico Sennrich, Alexandra Birch, Barry Haddow |
NAACL-HLT | 1 |
| 2017 | Machine Translation, it's a question of style, innit? The case of English tag questionsabstractIn this paper, we address the problem of generating English tag questions (TQs) (e.g. it is, isn't it?) in Machine Translation (MT).We propose a post-edition solution, formulating the problem as a multiclass classification task.We present (i) the automatic annotation of English TQs in a parallel corpus of subtitles and (ii) an approach using a series of classifiers to predict TQ forms, which we use to post-edit state-of-the-art MT outputs.Our method provides significant improvements in English TQ translation when translating from Czech, French and German, in turn improving the fluidity, naturalness, grammatical correctness and pragmatic coherence of MT output. Rachel Bawden |
EMNLP | 1 |
| 2016 | Boosting for Efficient Model Selection for Syntactic ParsingabstractWe present an efficient model selection method using boosting for transition-based constituency parsing. It is designed for exploring a high-dimensional search space, defined by a large set of feature templates, as for example is typically the case when parsing morphologically rich languages. Our method removes the need to manually define heuristic constraints, which are often imposed in current state-of-the-art selection methods. Our experiments for French show that the method is more efficient and is also capable of producing compact, state-of-the-art models. Rachel Bawden, Benoît Crabbé |
COLING | 1 |
| 2014 | Correcting and Validating Syntactic Dependency in the Spoken French Treebank Rhapsodie
Rachel Bawden, Marie-Amélie Botalla, Kim Gerdes, Sylvain Kahane |
LREC | 1 |