VLDB 2026 Research / reviewers in the wild / expert
Pierrette Bouillon
dblp:98/6521
· DBLP profile ↗
49ranked-venue papers
12as first author
10since 2021 · last 2025
0000-0002-8854-6360ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 46 · 12 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Towards High-Quality LLM-Based Data for French Spontaneous Speech Simplification: an Exo-Refinement ApproachabstractInternational audience Lucia Ormaechea Grijalba, Nikos Tsourakis, Pierrette Bouillon, Benjamin Lecouteux, Didier Schwab |
INTERSPEECH | 3 |
| 2025 | Arabizi vs LLMs: Can the Genie Understand the Language of Aladdin?abstractIn an era of rapid technological advancements, communication continues to evolve as new linguistic phenomena emerge. Among these is Arabizi, a hybrid form of Arabic that incorporates Latin characters and numbers to represent the spoken dialects of Arab communities. Arabizi is Widely used on social media and allows people to communicate in an informal and dynamic way, but it poses significant challenges for machine translation due to its lack of formal structure and deeply embedded cultural nuances. This case study is motivated by a growing need to translate Arabizi for gisting purpose. It evaluates the capacity of different LLMs’ to decode and translate Arabizi, focusing on multiple Arabic dialects that have rarely been studied up until now. Using a combination of human evaluators and automatic metrics, this research project investigates the model’s performance in translating Arabizi into both Modern Standard Arabic and English. Key questions explored include which dialects are translated most effectively and whether translations into English surpass those into Arabic. Perla Al Almaoui, Pierrette Bouillon, Simon Hengchen |
MTSummit (2) | 2 |
| 2024 | A Concept Based Approach for Translation of Medical Dialogues into PictographsabstractPictographs have been found to improve patient comprehension of medical information or instructions. However, tools to produce pictograph representations from natural language are still scarce. In this contribution we describe a system that automatically translates French speech into pictographs to enable diagnostic interviews in emergency settings, thereby providing a tool to overcome the language barrier or provide support in Augmentative and Alternative Communication (AAC) contexts. Our approach is based on a semantic gloss that serves as pivot between spontaneous language and pictographs, with medical concepts represented using the UMLS ontology. In this study we evaluate different available pre-trained models fine-tuned on artificial data to translate French into this semantic gloss. On unseen data collected in real settings, consisting of questions and instructions by physicians, the best model achieves an F0.5 score of 86.7. A complementary human evaluation of the semantic glosses differing from the reference shows that 71% of these would be usable to transmit the intended meaning. Finally, a human evaluation of the pictograph sequences derived from the gloss reveals very few additions, omissions or order issues (<3%), suggesting that the gloss as designed is well suited as a pivot for translation into pictographs. Johanna Gerlach, Pierrette Bouillon, Jonathan Mutal, Hervé Spechbach |
LREC/COLING | 2 |
| 2024 | Jargon: A Suite of Language Models and Evaluation Tasks for French Specialized DomainsabstractPretrained Language Models (PLMs) are the de facto backbone of most state-of-the-art NLP systems. In this paper, we introduce a family of domain-specific pretrained PLMs for French, focusing on three important domains: transcribed speech, medicine, and law. We use a transformer architecture based on efficient methods (LinFormer) to maximise their utility, since these domains often involve processing long documents. We evaluate and compare our models to state-of-the-art models on a diverse set of tasks and datasets, some of which are introduced in this paper. We gather the datasets into a new French-language evaluation benchmark for these three domains. We also compare various training configurations: continued pretraining, pretraining from scratch, as well as single- and multi-domain pretraining. Extensive domain-specific experiments show that it is possible to attain competitive downstream performance even when pre-training with the approximative LinFormer attention mechanism. For full reproducibility, we release the models and pretraining data, as well as contributed datasets. Vincent Segonne, Aidan Mannion, Laura Cristina Alonzo Canul, Alexandre Audibert, Cécile Macaire, Adrien Pupier, Yongxin Zhou 0004, Mathilde Aguiar, Felix Herron, Magali Norré, Massih-Reza Amini, Pierrette Bouillon, Iris Eshkol-Taravella, Emmanuelle Esperança-Rodier, Thomas François, Lorraine Goeuriot, Jérôme Goulian, Mathieu Lafourcade, Benjamin Lecouteux, François Portet, Fabien Ringeval, Vincent Vandeghinste, Maximin Coavoux, Marco Dinarelli, Didier Schwab |
LREC/COLING | 13 |
| 2024 | RCnum: A Semantic and Multilingual Online Edition of the Geneva Council Registers from 1545 to 1550abstractThe RCnum project is funded by the Swiss National Science Foundation and aims at producing a multilingual and semantically rich online edition of the Registers of Geneva Council from 1545 to 1550. Combining multilingual NLP, history and paleography, this collaborative project will clear hurdles inherent to texts manually written in 16th century Middle French while allowing for easy access and interactive consultation of these archives. Pierrette Bouillon, Christophe Chazalon, Sandra Coram-Mekkey, Gilles Falquet, Johanna Gerlach, Stéphane Marchand-Maillet, Laurent Moccozet, Jonathan Mutal, Raphaël Rubino, Marco Sorbi |
EAMT (2) | 1 |
| 2024 | Post-editors as Gatekeepers of Lexical and Syntactic Diversity: Comparative Analysis of Human Translation and Post-editing in Professional SettingsabstractThis paper presents a comparative analysis between human translation (HT) and post-edited machine translation (PEMT) from a lexical and syntactic perspective to verify whether the tendency of neural machine translation (NMT) systems to produce lexically and syntactically poorer translations shines through after post-editing (PE). The analysis focuses on three datasets collected in professional contexts containing translations from English into French and German into French. Through a comparison of word translation entropy (HTRa) scores, we observe a lower degree of lexical diversity in PEMT compared to HT. Additionally, metrics of syntactic equivalence indicate that PEMT is more likely to mirror the syntactic structure of the source text in contrast to HT. By incorporating raw machine translation (MT) output into our analysis, we underline the important role post-editors play in adding lexical and syntactic diversity to MT output. Our findings provide relevant input for MT users and decision-makers in language services as well as for MT and PE trainers and advisers. Lise Volkart, Pierrette Bouillon |
EAMT (1) | 2 |
| 2023 | PROPICTO: Developing Speech-to-Pictograph Translation Systems to Enhance Communication AccessibilityabstractPROPICTO is a project funded by the French National Research Agency and the Swiss National Science Foundation, that aims at creating Speech-to-Pictograph translation systems, with a special focus on French as an input language. By developing such technologies, we intend to enhance communication access for non-French speaking patients and people with cognitive impairments. Lucia Ormaechea Grijalba, Pierrette Bouillon, Maximin Coavoux, Emmanuelle Esperança-Rodier, Johanna Gerlach, Jérôme Goulian, Benjamin Lecouteux, Cécile Macaire, Jonathan Mutal, Magali Norré, Adrien Pupier, Didier Schwab |
EAMT | 2 |
| 2022 | The PASSAGE project : Standard German Subtitling of Swiss German TV contentabstractWe present the PASSAGE project, which aims at automatic Standard German subtitling of Swiss German TV content. This is achieved in a two step process, beginning with ASR to produce a normalised transcription, followed by translation into Standard German. We focus on the second step, for which we explore different approaches and contribute aligned corpora for future research. Pierrette Bouillon, Johanna Gerlach, Jonathan Mutal, Marianne Starlander |
EAMT | 1 |
| 2022 | Studying Post-Editese in a Professional Context: A Pilot StudyabstractThe past few years have seen the multiplication of studies on post-editese, following the massive adoption of post-editing in professional translation workflows. These studies mainly rely on the comparison of post-edited machine translation and human translation on artificial parallel corpora. By contrast, we investigate here post-editese on comparable corpora of authentic translation jobs for the language direction English into French. We explore commonly used scores and also proposes the use of a novel metric. Our analysis shows that post-edited machine translation is not only lexically poorer than human translation, but also less dense and less varied in terms of translation solutions. It also tends to be more prolific than human translation for our language direction. Finally, our study highlights some of the challenges of working with comparable corpora in post-editese research. Lise Volkart, Pierrette Bouillon |
EAMT | 2 |
| 2022 | Standard German Subtitling of Swiss German TV content: the PASSAGE ProjectabstractIn Switzerland, two thirds of the population speak Swiss German, a primarily spoken language with no standardised written form. It is widely used on Swiss TV, for example in news reports, interviews or talk shows, and subtitles are required for people who cannot understand this spoken language. This paper focuses on the task of automatic Standard German subtitling of spoken Swiss German, and more specifically on the translation of a normalised Swiss German speech recognition result into Standard German suitable for subtitles. Our contribution consists of a comparison of different statistical and deep learning MT systems for this task and an aligned corpus of normalised Swiss German and Standard German subtitles. Results of two evaluations, automatic and human, show that the systems succeed in improving the content, but are currently not capable of producing entirely correct Standard German. Jonathan Mutal, Pierrette Bouillon, Johanna Gerlach, Veronika Haberkorn |
LREC | 2 |
| 2020 | Re-design of the Machine Translation Training Tool (MT3)abstractWe believe that machine translation (MT) must be introduced to translation students as part of their training, in preparation for their professional life. In this paper we present a new version of the tool called MT3, which builds on and extends a joint effort undertaken by the Faculty of Languages of the University of Córdoba and Faculty of Translation and Interpreting of the University of Geneva to develop an open-source web platform to teach MT to translation students. We also report on a pilot experiment with the goal of testing the viability of using MT^3 in an MT course. The pilot let us identify areas for improvement and collect students’ feedback about the tool’s usability. Paula Estrella, Emiliano Cuenca, Laura Bruno, Jonathan Mutal, Sabrina Girletti, Lise Volkart, Pierrette Bouillon |
EAMT | 7 |
| 2020 | Ellipsis Translation for a Medical Speech to Speech Translation SystemabstractIn diagnostic interviews, elliptical utterances allow doctors to question patients in a more efficient and economical way. However, literal translation of such incomplete utterances is rarely possible without affecting communication. Previous studies have focused on automatic ellipsis detection and resolution, but only few specifically address the problem of automatic translation of ellipsis. In this work, we evaluate four different approaches to translate ellipsis in medical dialogues in the context of the speech to speech translation system BabelDr. We also investigate the impact of training data, using an under-sampling method and data with elliptical utterances in context. Results show that the best model is able to translate 88% of elliptical utterances. Jonathan Mutal, Johanna Gerlach, Pierrette Bouillon, Hervé Spechbach |
EAMT | 3 |
| 2019 | Surveying the potential of using speech technologies for post-editing purposes in the context of international organizations: What do professional translators think?
Jeevanthi Liyanapathirana, Pierrette Bouillon, Bartolomé Mesa-Lao |
MTSummit (2) | 2 |
| 2019 | Monolingual backtranslation in a medical speech translation system for diagnostic interviews - a NMT approach
Jonathan Mutal, Pierrette Bouillon, Johanna Gerlach, Paula Estrella, Hervé Spechbach |
MTSummit (2) | 2 |
| 2018 | Integrating MT at Swiss Post's Language Service: preliminary resultsabstractThis paper presents the preliminary results of an ongoing academia-industry collaboration that aims to integrate MT into the workflow of Swiss Post’s Language Service. We describe the evaluations carried out to select an MT tool (commercial or open-source) and assess the suitability of machine translation for post-editing in Swiss Post’s various subject areas and language pairs. The goal of this first phase is to provide recommendations with regard to the tool, language pair and most suitable domain for implementing MT. Pierrette Bouillon, Sabrina Girletti, Paula Estrella, Jonathan Mutal, Martina Bellodi, Beatrice Bircher |
EAMT | 1 |
| 2018 | Developing a New Swiss Research Centre for Barrier-Free CommunicationabstractThe project ‘Proposal and Implementation of a Swiss Research Centre for Barrier-free Communication’ (BFC) is a four-year project (2017–2020) funded by the Rectors’ Conference of Swiss Higher Education Institutions (swissuniversities).1 Its purpose is to ensure that individuals with a visual or hearing disability, people with a temporary cognitive impairment and speakers without sufficient knowledge of local languages can communicate and enjoy barrier-free access to information in all spheres of life, with a special focus on higher education. Pierrette Bouillon, Silvia Rodríguez Vázquez, Irene Strasly |
EAMT | 1 |
| 2017 | A Robust Medical Speech-to-Speech/Speech-to-Sign Phraselator
Farhia Ahmed, Pierrette Bouillon, Chelle Destefano, Johanna Gerlach, Sonia Halimi, Angela Hooper, Manny Rayner, Hervé Spechbach, Irene Strasly, Nikos Tsourakis |
INTERSPEECH | 2 |
| 2016 | In Memoriam: Susan ArmstrongabstractShe had a fundamental role in the founding and successful development of SIGDAT.Susan arrived in Switzerland from the United States in 1978 to work at the University of Lausanne in the German Department teaching literature and linguistics, as well as pursuing her interest in computers and language as translator and consultant to Logitech.She then moved to Geneva to work at the ISSCO research institute where she participated in a number of European research projects related to natural language processing (NLP) and machine translation, also working with corpora in translation Pierrette Bouillon, Paola Merlo, Gertjan van Noord, Mike Rosner |
Comput. Linguistics | 1 |
| 2015 | The ACCEPT Academic Portal: Bringing Together Pre-editing, MT and Post-editing into a Learning Environment
Pierrette Bouillon, Johanna Gerlach, Asheesh Gulati, Victoria Porro, Violeta Seretan |
EAMT | 1 |
| 2014 | A Large-Scale Evaluation of Pre-editing Strategies for Improving User-Generated Content Translation
Violeta Seretan, Pierrette Bouillon, Johanna Gerlach |
LREC | 2 |
| 2014 | Applying Accessibility-Oriented Controlled Language (CL) Rules to Improve Appropriateness of Text Alternatives for Images: an Exploratory Study
Silvia Rodríguez Vázquez, Pierrette Bouillon, Anton Bolfing |
LREC | 2 |
| 2012 | Annotating Qualia Relations in Italian and French Complex Nominals
Pierrette Bouillon, Elisabetta Jezek, Chiara Melloni, Aurélie Picton |
LREC | 1 |
| 2012 | Evaluating Appropriateness Of System Responses In A Spoken CALL Game
Manny Rayner, Pierrette Bouillon, Johanna Gerlach |
LREC | 2 |
| 2010 | A Bootstrapped Interlingua-Based SMT Architecture
Manny Rayner, Paula Estrella, Pierrette Bouillon |
EAMT | 3 |
| 2010 | A Multilingual CALL Game Based on Speech Translation
Manny Rayner, Pierrette Bouillon, Nikos Tsourakis, Johanna Gerlach, Maria Georgescul, Yukie Nakao, Claudia Baur |
LREC | 2 |
| 2010 | Examining the Effects of Rephrasing User Input on Two Mobile Spoken Language Systems
Nikos Tsourakis, Agnes Lisowska Masson, Manny Rayner, Pierrette Bouillon |
LREC | 4 |
| 2010 | Spoken Language Understanding via Supervised Learning and Linguistically Motivated Features
Maria Georgescul, Manny Rayner, Pierrette Bouillon |
NLDB | 3 |
| 2009 | Technology in Translator Training and tools for translators
Pierrette Bouillon, Marianne Starlander |
MTSummit | 1 |
| 2008 | Almost Flat Functional Semantics for Speech Translation
Manny Rayner, Pierrette Bouillon, Beth Ann Hockey, Yukie Nakao |
COLING | 2 |
| 2008 | Comparing two different bidirectional versions of the limited-domain medical spoken language translator MedSLT
Marianne Starlander, Pierrette Bouillon, Glenn Flores, Manny Rayner, Nikos Tsourakis |
EAMT | 2 |
| 2008 | Developing Non-European Translation Pairs in a Medium-Vocabulary Medical Speech Translation System
Pierrette Bouillon, Sonia Halimi, Yukie Nakao, Kyoko Kanzaki, Hitoshi Isahara, Nikos Tsourakis, Marianne Starlander, Beth Ann Hockey, Manny Rayner |
LREC | 1 |
| 2008 | Building Mobile Spoken Dialogue Applications Using Regulus
Nikos Tsourakis, Maria Georgescul, Pierrette Bouillon, Manny Rayner |
LREC | 3 |
| 2008 | Discriminative learning using linguistic features to rescore n-best speech hypothesesabstractWe describe how we were able to improve the accuracy of a medium-vocabulary spoken dialog system by rescoring the list of n-best recognition hypotheses using a combination of acoustic, syntactic, semantic and discourse information. The non-acoustic features are extracted from different intermediate processing results produced by the natural language processing module, and automatically filtered. We apply discriminative support vector learning designed for re-ranking, using both word error rate and semantic error rate as ranking target value, and evaluating using five-fold cross-validation; to show robustness of our method, confidence intervals for word and semantic error rates are computed via bootstrap sampling. The reduction in semantic error rate, from 19% to 11%, is statistically significant at 0.01 level. Maria Georgescul, Manny Rayner, Pierrette Bouillon, Nikos Tsourakis |
SLT | 3 |
| 2006 | REGULUS: A Generic Multilingual Open Source Platform for Grammar-Based Speech Applications
Manny Rayner, Pierrette Bouillon, Beth Ann Hockey, Nikos Chatzichrisafis |
LREC | 2 |
| 2005 | A generic multi-lingual open source platform for limited-domain medical speech translation
Pierrette Bouillon, Manny Rayner, Nikos Chatzichrisafis, Beth Ann Hockey, Marianne Santaholma, Marianne Starlander, Yukie Nakao, Kyoko Kanzaki, Hitoshi Isahara |
EAMT | 1 |
| 2005 | A methodology for comparing grammar-based and robust approaches to speech understandingabstractWe present a series of experiments designed to compare grammar-based and robust approaches to speech understanding, performed in the context of an Open Source medical speech translation system. We used two versions of the system, one grammar-based and one robust, trained off the same training data, and evaluated them on test data collected using both versions of the system. The experiments were constructed so as to avoid several methodological problems which occurred in earlier work reported in the literature. We found that the grammarbased version gave significantly better results than the robust version, with the difference increasing as subjects became more familiar with the system's coverage. The rate of improvement in subject performance was positively affected by providing them with an intelligent online help system. Manny Rayner, Pierrette Bouillon, Nikos Chatzichrisafis, Beth Ann Hockey, Marianne Santaholma, Marianne Starlander, Hitoshi Isahara, Kyoko Kanzaki, Yukie Nakao |
INTERSPEECH | 2 |
| 2005 | Practicing Controlled Language through a Help System integrated into the Medical Speech Translation System (MedSLT)abstractIn this paper, we present evidence that providing users of a speech to speech translation system for emergency diagnosis (MedSLT) with a tool that helps them to learn the coverage greatly improves their success in using the system. In MedSLT, the system uses a grammar-based recogniser that provides more predictable results to the translation component. The help module aims at addressing the lack of robustness inherent in this type of approach. It takes as input the result of a robust statistical recogniser that performs better for out-of-coverage data and produces a list of in-coverage example sentences. These examples are selected from a defined list using a heuristic that prioritises sentences maximising the number of N-grams shared with those extracted from the recognition result. Marianne Starlander, Pierrette Bouillon, Nikos Chatzichrisafis, Marianne Santaholma, Manny Rayner, Beth Ann Hockey, Hitoshi Isahara, Kyoko Kanzaki, Yukie Nakao |
MTSummit | 2 |
| 2004 | Methodology For Building Thematic Indexes In Medicine For French
Yalina Alphonse, Pierrette Bouillon |
LREC | 2 |
| 2004 | Automatisation of the Activity of Term Collection in Different Languages
Bruno Cartoni, Pierrette Bouillon, Yalina Alphonse, Sabine Lehmann |
LREC | 2 |
| 2004 | Semi-Automatic Derivation of a French Lexicon from CLIPS
Nilda Ruimy, Pierrette Bouillon, Bruno Cartoni |
LREC | 2 |
| 2003 | Learning Semantic Lexicons from a Part-of-Speech and Semantically Tagged Corpus Using Inductive Logic Programming
Vincent Claveau, Pascale Sébillot, Cécile Fabre, Pierrette Bouillon |
J. Mach. Learn. Res. | 4 |
| 2002 | From Resources to Applications. Designing the Multilingual ISLE Lexical Entry
Sue Atkins, Núria Bel, Francesca Bertagna, Pierrette Bouillon, Nicoletta Calzolari, Christiane Fellbaum, Ralph Grishman, Alessandro Lenci, Catherine Macleod, Martha Palmer, Gregor Thurmair, Marta Villegas, Antonio Zampolli |
LREC | 4 |
| 2002 | Acquisition of Qualia Elements from Corpora - Evaluation of a Symbolic Learning Method
Pierrette Bouillon, Vincent Claveau, Cécile Fabre, Pascale Sébillot |
LREC | 1 |
| 2000 | Medical document anonymization with a semantic lexicon
Patrick Ruch, Robert H. Baud, Anne-Marie Rassinoux, Pierrette Bouillon, Gilbert Robert |
AMIA | 4 |
| 1999 | MEDTAG: tag-like semantics for medical document indexing
Patrick Ruch, Judith C. Wagner, Pierrette Bouillon, Robert H. Baud, Anne-Marie Rassinoux, Jean-Raoul Scherrer |
AMIA | 3 |
| 1996 | Mental State Adjectives: the Perspective of Generative Lexicon
Pierrette Bouillon |
COLING | 1 |
| 1994 | On the Proper Role of Coercion in Semantic Typing
James Pustejovsky, Pierrette Bouillon |
COLING | 2 |
| 1994 | Semantic Lexicons: The Cornerstone for Lexical Choice in Natural Language Generation
Evelyne Viegas, Pierrette Bouillon |
INLG | 2 |
| 1992 | Une représentation sémantique et un système de transfert pour une traduction de haute qualité. Présentation de projet Avec démonstration sur machines SUN
Katharina Boesefeldt, Pierrette Bouillon |
COLING | 2 |