Laurent Romary

dblp:21/4228 · DBLP profile ↗
← Back
55ranked-venue papers
4as first author
10since 2021 · last 2026
0000-0002-0756-0508ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 51 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 2 first-authorDatabases, data management, data science and information retrieval · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 2Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 The MaTOS Pipeline for the Translation of Scientific Abstracts on the HAL Platform
abstract
English dominates scientific publishing, which disadvantages researchers who are not native English speakers, especially those in the earlier stages of their careers. Being able to write and engage with scientific content written in their own language would clearly facilitate scientific production. The MaTOS project (Machine Translation for Open Science) seeks to reduce these barriers by developing machine translation tools for scientific documents in English and French. This article presents the design of the MaTOS pipeline for the HAL platform to automatically translate article abstracts, with author validation, to increase the number of bilingual abstracts on the platform. We also report preliminary experiments comparing translation of sentence, three-sentence chunks, and whole abstracts, evaluated using quality estimation metrics.
Panagiotis Tsolakis, Ziqian Peng, Laurent Romary, François Yvon, Rachel Bawden
EAMT (2)3
2025 MaTOS: Machine Translation for Open Science
abstract
This paper is a short presentation of MaTOS, a project focusing on the automatic translation of scholarly documents. Its main aims are threefold: (a) to develop resources (term lists and corpora) for high-quality machine translation; (b) to study methods for translating complete, structured documents in a cohesive and consistent manner; (c) to propose novel metrics to evaluate machine translation in technical domains. Publications and resources are available on the project web site: https://anr-matos.gihub.io.
Rachel Bawden, Maud Bénard, José Cornejo Cárcamo, Nicolas Dahan, Manon Delorme, Mathilde Huguin, Natalie Kübler, Paul Lerner, Alexandra Mestivier, Joachim Minder, Jean-François Nominé, Ziqian Peng, Laurent Romary, Panagiotis Tsolakis, Lichao Zhu, François Yvon
MTSummit (2)13
2024 On Modelling Corpus Citations in Computational Lexical Resources
abstract
In this article we look at how two different standards for lexical resources, TEI and OntoLex, deal with corpus citations in lexicons. We will focus on how corpus citations in retrodigitised dictionaries can be modelled using each of the two standards since this provides us with a suitably challenging use case. After looking at the structure of an example entry from a legacy dictionary, we examine the two approaches offered by the two different standards by outlining an encoding for the example entry using both of them (note that this article features the first extended discussion of how the Frequency Attestation and Corpus (FrAC) module of OntoLex deals with citations). After comparing the two approaches and looking at the advantages and disadvantages of both, we argue for a combination of both. In the last part of the article we discuss different ways of doing this, giving our preference for a strategy which makes use of RDFa.
Anas Fahad Khan, Maxim Ionov, Christian Chiarcos, Laurent Romary, Gilles Sérasset, Besim Kabashi
LREC/COLING4
2024 Conversational Grounding: Annotation and Analysis of Grounding Acts and Grounding Units
abstract
Successful conversations often rest on common understanding, where all parties are on the same page about the information being shared. This process, known as conversational grounding, is crucial for building trustworthy dialog systems that can accurately keep track of and recall the shared information. The proficiencies of an agent in grounding the conveyed information significantly contribute to building a reliable dialog system. Despite recent advancements in dialog systems, there exists a noticeable deficit in their grounding capabilities. Traum (Traum, 1995) provided a framework for conversational grounding introducing Grounding Acts and Grounding Units, but substantial progress, especially in the realm of Large Language Models, remains lacking. To bridge this gap, we present the annotation of two dialog corpora employing Grounding Acts, Grounding Units, and a measure of their degree of grounding. We discuss our key findings during the annotation and also provide a baseline model to test the performance of current Language Models in categorizing the grounding acts of the dialogs. Our work aims to provide a useful resource for further research in making conversations with machines better understood and more reliable in natural day-to-day collaborative dialogs.
Biswesh Mohapatra, Seemab Hassan, Laurent Romary, Justine Cassell
LREC/COLING3
2024 Translate your Own: a Post-Editing Experiment in the NLP domain
abstract
The improvements in neural machine translation make translation and post-editing pipelines ever more effective for a wider range of applications. In this paper, we evaluate the effectiveness of such a pipeline for the translation of scientific documents (limited here to article abstracts). Using a dedicated interface, we collect, then analyse the post-edits of approximately 350 abstracts (English→French) in the Natural Language Processing domain for two groups of post-editors: domain experts (academics encouraged to post-edit their own articles) on the one hand and trained translators on the other. Our results confirm that such pipelines can be effective, at least for high-resource language pairs. They also highlight the difference in the post-editing strategy of the two subgroups. Finally, they suggest that working on term translation is the most pressing issue to improve fully automatic translations, but that in a post-editing setup, other error types can be equally annoying for post-editors.
Rachel Bawden, Ziqian Peng, Maud Bénard, Éric Villemonte de la Clergerie, Raphaël Esamotunu, Mathilde Huguin, Natalie Kübler, Alexandra Mestivier, Mona Michelot, Laurent Romary, Lichao Zhu, François Yvon
EAMT (1)10
2024 Evaluating the Effectiveness of Large Language Models in Establishing Conversational Grounding
abstract
Conversational grounding, vital for building effective dialogue between people and between people and dialogue systems, involves ensuring a mutual understanding of shared information.Despite its importance, there has been limited research on this aspect of conversation in recent years, especially after the advent of Large Language Models (LLMs).Previous studies have highlighted the shortcomings of some pre-trained language models in conversational grounding.However, most testing for conversational grounding capabilities involves human evaluations that are costly and time-consuming.This has led to a lack of testing across multiple models of varying sizes, a critical need given the rapid rate of new model releases.This gap in research becomes more significant considering recent advances in language models, which have led to new emergent capabilities.In this paper, we evaluate the performance of LLMs in various aspects of conversational grounding and analyze why some models perform better than others.We demonstrate a direct correlation between the size of the pre-training dataset, size of the model and conversational grounding abilities, suggesting that they have independently acquired some pragmatic capabilities from larger pre-training datasets.Finally, we propose ways to enhance the capabilities of the models that lag in our tests.
Biswesh Mohapatra, Manav Nitin Kapadnis, Laurent Romary, Justine Cassell
EMNLP3
2023 ISO LMF 24613-6: A Revised Syntax Semantics Module for the Lexical Markup Framework
Francesca Frontini, Laurent Romary, Anas Fahad Khan
LDK2
2022 Towards a Cleaner Document-Oriented Multilingual Crawled Corpus
abstract
The need for large corpora raw corpora has dramatically increased in recent years with the introduction of transfer learning and semi-supervised learning methods to Natural Language Processing. And while there have been some recent attempts to manually curate the amount of data necessary to train large language models, the main way to obtain this data is still through automatic web crawling. In this paper we take the existing multilingual web corpus OSCAR and its pipeline Ungoliant that extracts and classifies data from Common Crawl at the line level, and propose a set of improvements and automatic annotations in order to produce a new document-oriented version of OSCAR that could prove more suitable to pre-train large generative language models as well as hopefully other applications in Natural Language Processing and Digital Humanities.
Julien Abadji, Pedro Ortiz Suarez, Laurent Romary, Benoît Sagot
LREC3
2022 BERTrade: Using Contextual Embeddings to Parse Old French
abstract
The successes of contextual word embeddings learned by training large-scale language models, while remarkable, have mostly occurred for languages where significant amounts of raw texts are available and where annotated data in downstream tasks have a relatively regular spelling. Conversely, it is not yet completely clear if these models are also well suited for lesser-resourced and more irregular languages. We study the case of Old French, which is in the interesting position of having relatively limited amount of available raw text, but enough annotated resources to assess the relevance of contextual word embedding models for downstream NLP tasks. In particular, we use POS-tagging and dependency parsing to evaluate the quality of such models in a large array of configurations, including models trained from scratch from small amounts of raw text and models pre-trained on other languages but fine-tuned on Medieval French data.
Loïc Grobol, Mathilde Regnault, Pedro Ortiz Suarez, Benoît Sagot, Laurent Romary, Benoît Crabbé
LREC5
2021 Arabic factoid Question-Answering system for Islamic sciences using normalized corpora
abstract
Factoid Question-Answering (QA) Systems were developed to provide an accurate answer to a factoid question expressed in a natural language. The prime knowledge resources for most of factoid QA systems are online databases. However, the unstructured information in these resources rises the complexity of the information retrieval task. In this paper, we aim to develop an Arabic QA system for factoid questions specialized in Islamic sciences as prophetic tradition (Hadith), Hadith narrator and Quran interpretation (Tafsir). In fact, many questions in the Islamic research fields are focusing on Tafsir and Hadith text. In addition, a number of those interrogations focus on the chain of narrators who transmitted the Hadith. Besides, analyzing the narrator profile is considered an important and enquired task in hadith science. Furthermore, the proposed QA system is based on a normalized database specified in Text Encoding Initiative (TEI) standard. To achieve this, we propose a method composed of three phases to retrieve an accurate answer for the user question. The first phase is the question analysis. The second is the information search. The third phase is the answer processing. Besides, we implant the proposed QA system with a graphic interface allowing the interaction with the user. Finally, we experiment our prototype on 100 questions in the theme of Hadith, narrator and Tafsir text. The QA system succeeds to generate accurate responses for 92% of the entered questions.
Hajer Maraoui, Kais Haddar, Laurent Romary
KES3
2020 CamemBERT: a Tasty French Language Model
abstract
Louis Martin, Benjamin Muller, Pedro Javier Ortiz Suárez, Yoann Dupont, Laurent Romary, Éric de la Clergerie, Djamé Seddah, Benoît Sagot. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020.
Louis Martin, Benjamin Muller, Pedro Ortiz Suarez, Yoann Dupont, Laurent Romary, Éric Villemonte de la Clergerie, Djamé Seddah, Benoît Sagot
ACL5
2020 A Monolingual Approach to Contextualized Word Embeddings for Mid-Resource Languages
abstract
We use the multilingual OSCAR corpus, extracted from Common Crawl via language classification, filtering and cleaning, to train monolingual contextualized word embeddings (ELMo) for five mid-resource languages.We then compare the performance of OSCARbased and Wikipedia-based ELMo embeddings for these languages on the part-ofspeech tagging and parsing tasks.We show that, despite the noise in the Common-Crawlbased OSCAR data, embeddings trained on OSCAR perform much better than monolingual embeddings trained on Wikipedia.They actually equal or improve the current state of the art in tagging and parsing for all five languages.In particular, they also improve over multilingual Wikipedia-based contextual embeddings (multilingual BERT), which almost always constitutes the previous state of the art, thereby showing that the benefit of a larger, more diverse corpus surpasses the crosslingual benefit of multilingual embedding architectures.
Pedro Ortiz Suarez, Laurent Romary, Benoît Sagot
ACL2
2020 Modelling Etymology in LMF/TEI: The Grande Dicionário Houaiss da Língua Portuguesa Dictionary as a Use Case
abstract
In this article we will introduce two of the new parts of the new multi-part version of the Lexical Markup Framework (LMF) ISO standard, namely part 3 of the standard (ISO 24613-3), which deals with etymological and diachronic data, and Part 4 (ISO 24613-4), which consists of a TEI serialisation of all of the prior parts of the model. We will demonstrate the use of both standards by describing the LMF encoding of a small number of examples taken from a sample conversion of the reference Portuguese dictionary Grande Dicionário Houaiss da Língua Portuguesa, part of a broader experiment comprising the analysis of different, heterogeneously encoded, Portuguese lexical resources. We present the examples in the Unified Modelling Language (UML) and also in a couple of cases in TEI.
Anas Fahad Khan, Laurent Romary, Ana Salgado, Jack Bowers, Mohamed Khemakhem, Toma Tasovac
LREC2
2020 Establishing a New State-of-the-Art for French Named Entity Recognition
abstract
The French TreeBank developed at the University Paris 7 is the main source of morphosyntactic and syntactic annotations for French. However, it does not include explicit information related to named entities, which are among the most useful information for several natural language processing tasks and applications. Moreover, no large-scale French corpus with named entity annotations contain referential information, which complement the type and the span of each mention with an indication of the entity it refers to. We have manually annotated the French TreeBank with such information, after an automatic pre-annotation step. We sketch the underlying annotation guidelines and we provide a few figures about the resulting annotations.
Pedro Ortiz Suarez, Yoann Dupont, Benjamin Muller, Laurent Romary, Benoît Sagot
LREC4
2019 Automatic Identification and Normalisation of Physical Measurements in Scientific Literature
abstract
We present Grobid-quantities, an open-source application for extracting and normalising measurements from scientific and patent literature. Tools of this kind, aiming to understand and make unstructured information accessible, represent the building blocks for large-scale Text and Data Mining (TDM) systems. Grobid-quantities is a module built on top of Grobid [6] [13], a machine learning framework for parsing and structuring PDF documents. Designed to process large quantities of data, it provides a robust implementation accessible in batch mode or via a REST API. The machine learning engine architecture follows the cascade approach, where each model is specialised in the resolution of a specific task. The models are trained using CRF (Conditional Random Field) algorithm [12] for extracting quantities (atomic values, intervals and lists), units (such as length, weight) and different value representations (numeric, alphabetic or scientific notation). Identified measurements are normalised according to the International System of Units (SI). Thanks to its stable recall and reliable precision, Grobid-quantities has been integrated as the measurement-extraction engine in various TDM projects, such as Marve (Measurement Context Extraction from Text), for extracting semantic measurements and meaning in Earth Science [10]. At the National Institute for Materials Science in Japan (NIMS), it is used in an ongoing project to discover new superconducting materials. Normalised materials characteristics (such as critical temperature, pressure) extracted from scientific literature are a key resource for materials informatics (MI) [9].
Luca Foppiano, Laurent Romary, Masashi Ishii, Mikiko Tanifuji
DocEng2
2016 Algebraic Specification for Interoperability Between Data Formats: Application on Arabic Lexical Data
Malek Lhioui, Kais Haddar, Laurent Romary
CICLing (1)3
2016 TermITH-Eval: a French Standard-Based Resource for Keyphrase Extraction Evaluation
Adrien Bougouin, Sabine Barreaux, Laurent Romary, Florian Boudin, Béatrice Daille
LREC3
2014 Natural Language Processing for Historical Texts Michael Piotrowski (Leibniz Institute of European History) Morgan & Claypool (Synthesis Lectures on Human Language Technologies, edited by Graeme Hirst, volume 17), 2012, ix+157 pp; paperbound, ISBN 978-1608459469
abstract
The publication of a scholarly book is always the conjunction of an author's desire (or need) to disseminate their experience and knowledge and the interest or expectations of a potential community of readers to gain benefit from the publication itself. Michael Piotrowski has indeed managed to optimize this relation by bringing to the public a compendium of information that I think has been heavily awaited by many scholars having to deal with corpora of historical texts. The book covers most topics related to the acquisition, encoding, and annotation of historical textual data, seen from the point of view of their linguistic content. As such, it does not address issues related, for instance, to scholarly editions of these texts, but conveys a wealth of information on the various aspects where recent developments in language technology may help digital humanities projects to be aware of the current state of the art in the field.Still, the book is not an encyclopedic description of such technologies. It is based on the experience acquired by the author within the corpus development projects he has been involved in, and reflects in particular the specific topics on which he has made more in-depth explorations. It is thus written more as a series of returns on experience than a systematic resource to which one would want to return after its initial reading.The book is organized as a series of nine short chapters.In the first two (very short) chapters, the author presents the general scope of the book and provides an overview of the reasons why natural language processing (NLP) has such an entrenched position in digital humanities at large and the study of historical text in particular. Citing several prominent projects and corpus initiatives that have taken place in the last few decades, Piotrowski defends the thesis, which I share, that a deep understanding of textual documents requires some basic knowledge of language processing methods and techniques. Chapter 2 in particular (“NLP and Digital Humanities”) could be read as an autonomous position paper, which, independently of the following chapters, presents the current landscape of infrastructural initiatives and scholarly projects that shape this convergence between the two fields.Chapter 3 (“Spelling in Historical Texts,” pp. 11–23) describes the various issues related to spelling variations in historical text. It shows how difficult it may be to deal with both diachronic (e.g., in comparison to modern standardized spellings) and synchronic (degree of stabilization of historical spellings) variations, especially in the context of the uncertainty brought about by the transcription process itself. This is particularly true for historical manuscripts and Piotrowski goes deeply into this, showing some concrete examples of the kind of hurdles that a scholar may fall into. This is the kind of short introduction I would recommend for anyone, in particular students, wanting to gain a first understanding in the domain of historical spelling.Chapter 4 is the longest chapter in the book (“Acquiring Historical Texts,” pp. 25–52) and covers various aspects of the digitization workflow that needs to be set up to create a corpus of historical texts. The chapter is quite difficult to read as a single unit because of its intrinsic heterogeneity. Indeed, it covers quite a wide range of topics: presentation of existing digitization projects worldwide, technical issues related to scanning, comparison of various optical character recognition systems for various types of scripts, the potential role of lexical resources, crowdsourcing for optical character recognition (OCR) post-processing, and manual or semi-automatic keying. Getting an overview of the various topics is even more difficult because of the way the author has followed his own personal experience, and alternates between general considerations and in-depth presentations of concrete results. Pages 34–40, for instance, is one single subsection on the comparison of OCR outputs that goes into so much detail that it breaks the continuity of the argument, although in itself this subsection could be really interesting for a specialized reader. This chapter illustrates the point that the content of this book would benefit from being published in a more modern and open setting.Data representation aspects are covered in Chapter 5 (“Text Encoding and Annotation Schemes,” pp. 53–68), which tackles two specific issues, namely, character and document encoding. On these two, the author presents what could be considered best practices. For character encoding, the book rightly focuses on the advantages that the move towards ISO 10646/Unicode has brought to the community. The corresponding sub-section actually covers three different aspects: It first makes an extensive presentation of the history of character encoding standards (from ASCII/ISO 646 to Unicode/ISO 10646), it provides insights into the current coverage and encoding principles (e.g., UTF-8 vs. UTF-16) of ISO 10646, and finally, it focuses on the specific difficulties occurring in historical texts both from the point of view of legacy ASCII-based transcription languages and the management of characters that are not present in Unicode. Although well documented, these three topics should have been more clearly separated so that readers interested in one or the other could directly refer to it. This is a typical case where, given the great expertise of the author on the subject, I can imagine the corresponding texts being published on-line as separate entries in a blog. The second half of the chapter focuses on the role of the Text Encoding Initiative (TEI) guidelines for the transcription and encoding of historical text. It covers the various representation levels that may be concerned (metadata, text structure, surface annotation) and insists on the current difficulty of linking current NLP tools to TEI encoded documents. Although this is indeed still an issue in general, it might have been interesting to refer to standards (ISO 24611– MAF) and initiatives (Textgrid core encoding at the token level; the TXM platform for text mining) that have started to provide concrete sustainable answers to the issue.The following chapter (“Handling Spelling Variations,” pp. 69–84), provides a series of short studies describing possible methods for dealing with OCR errors or spelling variations as described in Chapter 3. Independent of the fact that I find it strange to see the two chapters set quite far from one another, Chapter 6 distinguishes itself by its profound heterogeneity. Whereas several sections do have the most appropriate level of detail and topicality for historical texts (in particular those on canonicalization), some sections seem to be completely off topic (Section 6.2, “Edit Distance,” describes what I would consider as background knowledge for such a book). It is all the more disappointing that the author shows here a very high level of expertise and, as in the case of Chapter 3, I would strongly recommend the reading of the relevant sections to newcomers in the field.In contrast with the previous chapter, Chapter 7 (“NLP Tools for Historical Languages,” pp. 85–100) is more coherent and focused. It mainly addresses the morpho-syntactic analysis of historical text and presents, through concrete deployment scenarios, possible methods to constrain the appropriate parsers, in a context where hardly any existing tools can be simply re-used. The chapter is very well documented and refers to most of the relevant initiatives in the domain of morphology for historical text, at least on the European scene. This focus may also be misleading because recent work on named entity recognition on historical texts are not at all mentioned and are probably, to my view, one of the most promising direction for enhanced digital scholarship.The last chapter (“Historical Corpora,” pp. 101–116) is a compendium, sorted by language, of the major historical corpora available worldwide. It shows the dynamic that currently exists in the community and is an essential background resource to both understanding who is active in maintaining historical corpora and discerning the most relevant resources. The chapter as a whole provides an interesting “historical” perspective on the progress made by most text-based projects in using the TEI guidelines as their reference standard. It seems quite difficult now to imagine an initiative which would not take TEI for granted, and would not build inside the TEI framework. On another issue, namely, copyright, Piotrowski also provides an interesting analysis on the difficulty of re-using old editions which have been recently re-edited on paper, and thus fall into some publisher's copyright restrictions. The conclusion could have been a little tougher here though, and probably should have recommended putting a hold on any paper publication of historical sources by a private publisher unless it is guaranteed that the electronic material can be used freely, under an appropriate open license.As a whole, the book leaves the reader with a mixed feeling of enthusiasm and disappointment. Enthusiasm, because the content is so rich that it should serve as background reference (and indeed be quoted) for any further work on the creation, management, and curation of historical corpora. Still, I cannot help thinking that the editorial setting as a book is not the most appropriate setting for such content. The variety of topics that are addressed as well as the heterogeneous level of detail provided through the different chapters would benefit from a more fragmented treatment. Indeed, this would be the perfect content for a series of blog entries (for instance, in a scholarly blog such as those on the hypotheses.org platform) which in turn would allow an interested reader to discover exactly the topics they want information about and cite the corresponding entries. With the bibliography in Zotero and relevant pointers to the corresponding on-line corpora or tools, I could imagine the resulting content soon becoming one of the most cited on-line resources. I am sure the author would gain more visibility in doing so than having the material hidden on a library shelf or behind a paywall. Not knowing the exact copyright transfer agreement associated with the book, I cannot judge if it is too late for the author to think in these terms, but this could be a lesson for scholars who are now planning to write such an introductory publication. Is the book still the best medium?
Laurent Romary
Comput. Linguistics1
2012 Collaborative Machine Translation Service for Scientific texts
Patrik Lambert, Jean Senellart, Laurent Romary, Holger Schwenk, Florian Zipser, Patrice Lopez, Frédéric Blain
EACL3
2010 Towards an ISO Standard for Dialogue Act Annotation
Harry Bunt, Jan Alexandersson, Jean Carletta, Jae-Woong Choe, Alex Chengyu Fang, Kôiti Hasida, Kiyong Lee, Volha Petukhova, Andrei Popescu-Belis, Laurent Romary, Claudia Soria, David R. Traum
LREC10
2010 MLIF : A Metamodel to Represent and Exchange Multilingual Textual Information
Samuel Cruz-Lara, Gil Francopoulo, Laurent Romary, Nasredine Semmar
LREC3
2010 GRISP: A Massive Multilingual Terminological Database for Scientific and Technical Domains
Patrice Lopez, Laurent Romary
LREC2
2010 ISO-TimeML: An International Standard for Semantic Annotation
James Pustejovsky, Kiyong Lee, Harry Bunt, Laurent Romary
LREC4
2008 Foundation of a Component-based Flexible Registry for Language Resources and Technology
Daan Broeder, Thierry Declerck, Erhard W. Hinrichs, Stelios Piperidis, Laurent Romary, Nicoletta Calzolari, Peter Wittenburg
LREC5
2006 Representing Linguistic Corpora and Their Annotations
Nancy Ide, Laurent Romary
LREC2
2006 An API for accessing the Data Category Registry
Marc Kemps-Snijders, Julien Ducret, Laurent Romary, Peter Wittenburg
LREC3
2006 A Lexicalized Tree-Adjoining Grammar for Vietnamese
Hong Phuong Le, Nguyên Thi Minh Huyên, Laurent Romary, Azim Roussanaly
LREC3
2006 Metadata Profile in the ISO Data Category Registry
Freddy Offenga, Daan Broeder, Peter Wittenburg, Julien Ducret, Laurent Romary
LREC5
2006 Foundations of Modern Language Resource Archives
Peter Wittenburg, Daan Broeder, Wolfgang Klein, Stephen C. Levinson, Laurent Romary
LREC5
2004 A Large Metadata Domain of Language Resources
Daan Broeder, Thierry Declerck, Laurent Romary, Markus Uneson, Sven Strömqvist, Peter Wittenburg
LREC3
2004 Standardization in Multimodal Content Representation: Some Methodological Issues
Harry Bunt, Laurent Romary
LREC2
2004 The French MEDIA/EVALDA Project: the Evaluation of the Understanding Capability of Spoken Language Dialogue Systems
Laurence Devillers, Hélène Bonneau-Maynard, Sophie Rosset, Patrick Paroubek, Kevin McTait, Djamel Mostefa, Khalid Choukri, Laurent Charnay, Caroline Bousquet-Vernhettes, Nadine Vigouroux, Frédéric Béchet, Laurent Romary, Jean-Yves Antoine, Jeanne Villaneau, Myriam Vergnes, Jérôme Goulian
LREC12
2004 A Registry of Standard Data Categories for Linguistic Annotation
Nancy Ide, Laurent Romary
LREC2
2004 Multimodal Meaning Representation for Generic Dialogue Systems Architectures
Frédéric Landragin, Alexandre Denis 0002, Annalisa Ricci, Laurent Romary
LREC4
2004 Towards an International Standard on Feature Structure Representation
Kiyong Lee, Lou Burnard, Laurent Romary, Éric Villemonte de la Clergerie, Thierry Declerck, Syd Bauman, Harry Bunt, Lionel Clément, Tomaz Erjavec, Azim Roussanaly, Claude Roux
LREC3
2004 Developping Tools and Building Linguistic Resources for Vietnamese Morpho-syntactic Processing
Thanh Bon Nguyen, Nguyên Thi Minh Huyên, Laurent Romary, Xuân Luong Vu
LREC3
2004 Online Evaluation of Coreference Resolution
Andrei Popescu-Belis, Loïs Rigouste, Susanne Salmon-Alt, Laurent Romary
LREC4
2004 Experiments on Building Language Resources for Multi-Modal Dialogue Systems
Laurent Romary, Amalia Todirascu, David Langlois 0001
LREC1
2004 Towards a Reference Annotation Framework
Susanne Salmon-Alt, Laurent Romary
LREC2
2004 International standard for a linguistic annotation framework
abstract
This paper describes the Linguistic Annotation Framework under development within ISO TC37 SC4 WG1. The Linguistic Annotation Framework is intended to serve as a basis for harmonizing existing language resources as well as developing new ones.
Nancy Ide, Laurent Romary
Nat. Lang. Eng.2
2003 SYSTRAN new generation: the XML translation workflow
abstract
Customization of Machine Translation (MT) is a prerequisite for corporations to adopt the technology. It is therefore important but nonetheless challenging. Ongoing implementation proves that XML is an excellent exchange device between MT modules that efficiently enables interaction between the user and the processes to reach highly granulated structure-based customization. Accomplished through an innovative approach called the SYSTRAN Translation Stylesheet, this method is coherent with the current evolution of the “authoring process”. As a natural progression, the next stage in the customization process is the integration of MT in a multilingual tool kit designed for the “authoring process”.
Jean Senellart, Christian Boitet, Laurent Romary
MTSummit3
2002 Referring to Objects with Spoken and Haptic Modalities
abstract
The gesture input modality considered in multimodal dialogue systems is mainly reduced to pointing or manipulating actions. With an approach based on spontaneous character of the communication, the treatment of such actions involves many processes. Without constraints, the user may use gesture in association with speech, and may exploit visual context peculiarities, guiding her/his articulation of gesture trajectories and her/his choice of words. Semantic interpretation of multimodal utterances also becomes a complex problem, taking into account varieties of referring expressions, varieties of gestural trajectories, structural parameters from the visual context, and also directives from a specific task. Following the spontaneous approach, we propose to give maximal understanding capabilities to dialogue systems, to ensure that various interaction modes must be taken into account. Considering the development of haptic sense devices (such as PHANToM) which increase the capabilities of sensations, particularly tactile and kinesthetic, we propose to explore a new domain of research concerning the integration of haptic gesture into multimodal dialogue systems, in terms of its possible associations with speech for object reference and manipulation. We focus on the compatibility between haptic gesture and multimodal reference models, and on the consequences of processing this new modality on intelligent system architectures, which has been sufficiently studied from a semantic point of view.
Frédéric Landragin, Nadia Bellalem, Laurent Romary
ICMI3
2002 LREP: A Language Repository Exchange Protocol
Daan Broeder, Peter Wittenburg, Thierry Declerck, Laurent Romary
LREC4
2002 Standards for Language Resources
Nancy Ide, Laurent Romary
LREC2
2002 Towards Reusable NLP Components
Amalia Todirascu, Eric Kow, Laurent Romary
LREC3
2002 Vulcain - An Ontology-Based Information Extraction System
Amalia Todirascu, Laurent Romary, Dalila Bekhouche
NLDB2
2001 A Common Framework for Syntactic Annotation
abstract
It is widely recognized that the proliferation of annotation schemes runs counter to the need to re-use language resources, and that standards for linguistic annotation are becoming increasingly mandatory. To answer this need, we have developed a representation framework comprised of an abstract model for a variety of different annotation types (e.g., morpho-syntactic tagging, syntactic annotation, co-reference annotation, etc.), which can be instantiated in different ways depending on the annotator s approach and goals. In this paper we provide an overview of our representation framework and demonstrate its applicability to syntactic annotation. We show how the framework can contribute to comparative evaluation and merging of parser output and diverse syntactic annotation schemes.
Nancy Ide, Laurent Romary
ACL2
2000 XCES: An XML-based Encoding Standard for Linguistic Corpora
Nancy Ide, Patrice Bonhomme, Laurent Romary
LREC3
1999 A Contextual Analysis of Referring Gestures
abstract
Colloque avec actes et comité de lecture.
Frederic Wolff, Laurent Romary
IUI2
1998 Marking- up multiple views of a text: discourse and reference
Dan Cristea, Nancy Ide, Laurent Romary
LREC3
1998 East meets West: multilingual resources in a European context
Tomaz Erjavec, Ann Lawson, Laurent Romary
LREC3
1994 Frames, a unified model for the representation of reference and space in a man-machine dialogue
abstract
International audience
Daniel Schang, Laurent Romary
ICSLP2
1991 References in a multimodal dialogue: towards a unified processing
abstract
International audience
Bertrand Gaiffe, Laurent Romary, Jean-Marie Pierrel
EUROSPEECH2
1989 Should an oral dialogue system be modular?
abstract
International audience
Laurent Romary, Jean-Marie Pierrel
EUROSPEECH1
1989 The use of the Dempster-Shafer rule in the lexical component of a man-machine oral dialogue system
Laurent Romary, Jean-Marie Pierrel
Speech Commun.1