Serge Sharoff

dblp:49/607 · DBLP profile ↗
← Back
48ranked-venue papers
13as first author
10since 2021 · last 2026
0000-0002-4877-0210ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 45 · 12 first-author · 10 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 To Predict or Not to Predict? Towards Reliable Uncertainty Estimation in the Presence of Noise
Nouran Khallaf, Serge Sharoff
LREC2
2026 How Much Noise Can BERT Handle? Insights from Multilingual Sentence Difficulty Detection
Nouran Khallaf, Serge Sharoff
LREC2
2025 BERT-based Classical Arabic Poetry Authorship Attribution
abstract
This study introduces a novel computational approach to authorship attribution (AA) in Arabic poetry, using the entire Classical Arabic Poetry corpus for the first time and offering a direct analysis of real cases of misattribution. AA in Arabic poetry has been a significant issue since the 9th century, particularly due to the loss of pre-Islamic poetry and the misattribution of post-Islamic works to earlier poets. While previous research has predominantly employed qualitative methods, this study uses computational techniques to address these challenges. The corpus was scraped from online sources and enriched with manually curated Date of Death (DoD) information to overcome the problematic traditional sectioning. Additionally, we applied Embedded Topic Modeling (ETM) to label each poem with its topic contributions, further enhancing the dataset’s value. An ensemble model based on CAMeLBERT was developed and tested across three dimensions: topic, number of poets, and number of training examples. After parameter optimization, the model achieved F1 scores ranging from 0.97 to 1.0. The model was also applied to four pre-Islamic misattribution cases, producing results consistent with historical and literary studies.
Lama Alqurashi, Serge Sharoff, Janet Watson, Jacob Blakesley
COLING2
2025 Controlling Out-of-Domain Gaps in LLMs for Genre Classification and Generated Text Detection
abstract
This study demonstrates that the modern generation of Large Language Models (LLMs, such as GPT-4) suffers from the same out-of-domain (OOD) performance gap observed in prior research on pre-trained Language Models (PLMs, such as BERT). We demonstrate this across two non-topical classification tasks: (1) genre classification and (2) generated text detection. Our results show that when demonstration examples for In-Context Learning (ICL) come from one domain (e.g., travel) and the system is tested on another domain (e.g., history), classification performance declines significantly. To address this, we introduce a method that controls which predictive indicators are used and which are excluded during classification. For the two tasks studied here, this ensures that topical features are omitted, while the model is guided to focus on stylistic rather than content-based attributes. This approach reduces the OOD gap by up to 20 percentage points in a few-shot setup. Straightforward Chain-of-Thought (CoT) methods, used as the baseline, prove insufficient, while our approach consistently enhances domain transfer performance.
Dmitri Roussinov, Serge Sharoff, Nadezhda Puchnina
COLING2
2024 Enhancing Image-to-Text Generation in Radiology Reports through Cross-modal Multi-Task Learning
abstract
Image-to-text generation involves automatically generating descriptive text from images and has applications in medical report generation. However, traditional approaches often exhibit a semantic gap between visual and textual information. In this paper, we propose a multi-task learning framework to leverage both visual and non-imaging data for generating radiology reports. Along with chest X-ray images, 10 additional features comprising numeric, binary, categorical, and text data were incorporated to create a unified representation. The model was trained to generate text, predict the degree of patient severity, and identify medical findings. Multi-task learning, especially with text generation prioritisation, improved performance over single-task baselines across language generation metrics. The framework also mitigated overfitting in auxiliary tasks compared to single-task models. Qualitative analysis showed logically coherent narratives and accurate identification of findings, though some repetition and disjointed phrasing remained. This work demonstrates the benefits of multi-modal, multi-task learning for image-to-text generation applications.
Nurbanu Aksoy, Nishant Ravikumar, Serge Sharoff
LREC/COLING3
2024 Quantifying the Contribution of MWEs and Polysemy in Translation Errors for English-Igbo MT
abstract
In spite of recent successes in improving Machine Translation (MT) quality overall, MT engines require a large amount of resources, which leads to markedly lower quality for lesser-resourced languages. This study explores the case of translation from English into Igbo, a very low resource language spoken by about 45 million speakers. With the aim of improving MT quality in this scenario, we investigate methods for guided detection of critical/harmful MT errors, more specifically those caused by non-compositional multi-word expressions and polysemy. We have designed diagnostic tests for these cases and applied them to collections of medical texts from CDC, Cochrane, NCDC, NHS and WHO.
Adaeze Ohuoba, Serge Sharoff, Callum Walker
EAMT (1)2
2022 Applying Natural Annotation and Curriculum Learning to Named Entity Recognition for Under-Resourced Languages
abstract
Current practices in building new NLP models for low-resourced languages rely either on Machine Translation of training sets from better resourced languages or on cross-lingual transfer from them. Still we can see a considerable performance gap between the models originally trained within better resourced languages and the models transferred from them. In this study we test the possibility of (1) using natural annotation to build synthetic training sets from resources not initially designed for the target downstream task and (2) employing curriculum learning methods to select the most suitable examples from synthetic training sets. We test this hypothesis across seven Slavic languages and across three curriculum learning strategies on Named Entity Recognition as the downstream task. We also test the possibility of fine-tuning the synthetic resources to reflect linguistic properties, such as the grammatical case and gender, both of which are important for the Slavic languages. We demonstrate the possibility to achieve the mean F1 score of 0.78 across the three basic entities types for Belarusian starting from zero resources in comparison to the baseline of 0.63 using the zero-shot transfer from English. For comparison, the English model trained on the original set achieves the mean F1-score of 0.75. The experimental results are available from https://github.com/ValeraLobov/SlavNER
Valeriy Lobov, Alexandra Ivoylova, Serge Sharoff
COLING3
2022 BERTology for Machine Translation: What BERT Knows about Linguistic Difficulties for Translation
abstract
Pre-trained transformer-based models, such as BERT, have shown excellent performance in most natural language processing benchmark tests, but we still lack a good understanding of the linguistic knowledge of BERT in Neural Machine Translation (NMT). Our work uses syntactic probes and Quality Estimation (QE) models to analyze the performance of BERT’s syntactic dependencies and their impact on machine translation quality, exploring what kind of syntactic dependencies are difficult for NMT engines based on BERT. While our probing experiments confirm that pre-trained BERT “knows” about syntactic dependencies, its ability to recognize them often decreases after fine-tuning for NMT tasks. We also detect a relationship between syntactic dependencies in three languages and the quality of their translations, which shows which specific syntactic dependencies are likely to be a significant cause of low-quality translations.
Yuqian Dai, Marc de Kamps, Serge Sharoff
LREC3
2022 Estimating Confidence of Predictions of Individual Classifiers and TheirEnsembles for the Genre Classification Task
abstract
Genre identification is a kind of non-topic text classification. The main difference between this task and topic classification is that genre, unlike topic, usually cannot be expressed just by some keywords and is defined as a functional space. Neural models based on pre-trained transformers, such as BERT or XLM-RoBERTa, demonstrate SOTA results in many NLP tasks, including non-topical classification. However, in many cases, their downstream application to very large corpora, such as those extracted from social media, can lead to unreliable results because of dataset shifts, when some raw texts do not match the profile of the training set. To mitigate this problem, we experiment with individual models as well as with their ensembles. To evaluate the robustness of all models we use a prediction confidence metric, which estimates the reliability of a prediction in the absence of a gold standard label. We can evaluate robustness via the confidence gap between the correctly classified texts and the misclassified ones on a labeled test corpus, higher gaps make it easier to identify whether a text is classified correctly. Our results show that for all of the classifiers tested in this study, there is a confidence gap, but for the ensembles, the gap is wider, meaning that ensembles are more robust than their individual models.
Mikhail Lepekhin, Serge Sharoff
LREC2
2022 Multimodal Pipeline for Collection of Misinformation Data from Telegram
abstract
The paper presents the outcomes of AI-COVID19, our project aimed at better understanding of misinformation flow about COVID-19 across social media platforms. The specific focus of the study reported in this paper is on collecting data from Telegram groups which are active in promotion of COVID-related misinformation. Our corpus collected so far contains around 28 million words, from almost one million messages. Given that a substantial portion of misinformation flow in social media is spread via multimodal means, such as images and video, we have also developed a mechanism for utilising such channels via producing automatic transcripts for videos and automatic classification for images into such categories as memes, screenshots of posts and other kinds of images. The accuracy of the image classification pipeline is around 87%.
Jose Sosa, Serge Sharoff
LREC2
2020 Recognizing Semantic Relations: Attention-Based Transformers vs. Recurrent Models
Dmitri Roussinov, Serge Sharoff, Nadezhda Puchnina
ECIR (1)2
2020 Recognizing Semantic Relations by Combining Transformers and Fully Connected Models
abstract
Automatically recognizing an existing semantic relation (e.g. “is a”, “part of”, “property of”, “opposite of” etc.) between two words (phrases, concepts, etc.) is an important task affecting many NLP applications and has been subject of extensive experimentation and modeling. Current approaches to automatically telling if a relation exists between two given concepts X and Y can be grouped into two types: 1) those modeling word-paths connecting X and Y in text and 2) those modeling distributional properties of X and Y separately, not necessary in the proximity to each other. Here, we investigate how both types can be improved and combined. We suggest a distributional approach that is based on an attention-based transformer. We have also developed a novel word path model that combines useful properties of a convolutional network with a fully connected language model. While our transformer-based approach works better, both our models significantly outperform the state-of-the-art within their classes of approaches. We also demonstrate that combining the two approaches results in additional gains since they use somewhat different data sources.
Dmitri Roussinov, Serge Sharoff, Nadezhda Puchnina
LREC2
2020 Know thy Corpus! Robust Methods for Digital Curation of Web corpora
abstract
This paper proposes a novel framework for digital curation of Web corpora in order to provide robust estimation of their parameters, such as their composition and the lexicon. In recent years language models pre-trained on large corpora emerged as clear winners in numerous NLP tasks, but no proper analysis of the corpora which led to their success has been conducted. The paper presents a procedure for robust frequency estimation, which helps in establishing the core lexicon for a given corpus, as well as a procedure for estimating the corpus composition via unsupervised topic models and via supervised genre classification of Web pages. The results of the digital curation study applied to several Web-derived corpora demonstrate their considerable differences. First, this concerns different frequency bursts which impact the core lexicon obtained from each corpus. Second, this concerns the kinds of texts they contain. For example, OpenWebText contains considerably more topical news and political argumentation in comparison to ukWac or Wikipedia. The tools and the results of analysis have been released.
Serge Sharoff
LREC1
2020 Sentence Level Human Translation Quality Estimation with Attention-based Neural Networks
abstract
This paper explores the use of Deep Learning methods for automatic estimation of quality of human translations. Automatic estimation can provide useful feedback for translation teaching, examination and quality control. Conventional methods for solving this task rely on manually engineered features and external knowledge. This paper presents an end-to-end neural model without feature engineering, incorporating a cross attention mechanism to detect which parts in sentence pairs are most relevant for assessing quality. Another contribution concerns oprediction of fine-grained scores for measuring different aspects of translation quality, such as terminological accuracy or idiomatic writing. Empirical results on a large human annotated dataset show that the neural model outperforms feature-based methods significantly. The dataset and the tools are available.
Serge Sharoff
LREC2
2020 Finding next of kin: Cross-lingual embedding spaces for related languages
abstract
Abstract Some languages have very few NLP resources, while many of them are closely related to better-resourced languages. This paper explores how the similarity between the languages can be utilised by porting resources from better- to lesser-resourced languages. The paper introduces a way of building a representation shared across related languages by combining cross-lingual embedding methods with a lexical similarity measure which is based on the weighted Levenshtein distance. One of the outcomes of the experiments is a Panslavonic embedding space for nine Balto-Slavonic languages. The paper demonstrates that the resulting embedding space helps in such applications as morphological prediction, named-entity recognition and genre classification.
Serge Sharoff
Nat. Lang. Eng.1
2018 Language adaptation experiments via cross-lingual embeddings for related languages
Serge Sharoff
LREC1
2018 Cross-lingual Terminology Extraction for Translation Quality Estimation
Yuze Gao, Yue Zhang 0004, Serge Sharoff
LREC4
2018 Investigating the Influence of Bilingual MWU on Trainee Translation Quality
Serge Sharoff
LREC2
2018 A Multilingual Dataset for Evaluating Parallel Sentence Extraction from Comparable Corpora
Pierre Zweigenbaum, Serge Sharoff, Reinhard Rapp
LREC2
2016 Adam Kilgarriff's Legacy to Computational Linguistics and Beyond
Roger Evans, Alexander F. Gelbukh, Gregory Grefenstette, Patrick Hanks, Milos Jakubícek, Diana McCarthy, Martha Palmer, Ted Pedersen, Michael Rundell, Pavel Rychlý, Serge Sharoff, David Tugwell
CICLing (1)11
2016 MoBiL: A Hybrid Feature Set for Automatic Human Translation Quality Assessment
Serge Sharoff, Bogdan Babych
LREC2
2016 Preface
abstract
After several decades of work on rule-based machine translation (MT) where linguists try to manually encode their knowledge about language, the time around 1990 brought a paradigm change towards automatic systems which try to learn how to translate by looking at large collections of high-quality sample translations as produced by professional translators. The first such attempts were called example- or analogy-based translation, and somewhat later the so-called statistical approach to MT was introduced. Both can be subsumed under the label data-driven approaches to MT. It took about 10 years until these self-learning systems became serious competitors of the traditional rule-based systems, and by now some of the most successful MT systems, such as Google Translate and Moses, are based on the statistical approach.
Reinhard Rapp, Serge Sharoff, Pierre Zweigenbaum
Nat. Lang. Eng.2
2016 Recent advances in machine translation using comparable corpora
abstract
Abstract This paper highlights some of the recent developments in the field of machine translation using comparable corpora. We start by updating previous definitions of comparable corpora and then look at bilingual versions of continuous vector space models. Recently, neural networks have been used to obtain latent context representations with only few dimensions which are often called word embeddings. These promising new techniques cannot only be applied to parallel but also to comparable corpora. Subsequent sections of the paper discuss work specifically targeting at machine translation using comparable corpora, as well as work dealing with the extraction of parallel segments from comparable corpora. Finally, we give an overview on the design and the results of a recent shared task on measuring document comparability across languages.
Reinhard Rapp, Serge Sharoff, Pierre Zweigenbaum
Nat. Lang. Eng.2
2015 Web Corpus Construction Roland Schäfer and Felix Bildhauer (Freie Universität Berlin) Morgan & Claypool (Synthesis Lectures on Human Language Technologies, edited by Graeme Hirst, volume 22), 2013, 145 pages, paper-bound, ISBN 9781608459834, doi: 10.2200/S00508ED1V01Y201305HLT022
abstract
The Web is the main source of data in modern computational linguistics. Other volumes in the same series, for example, Introductions to Opinion Mining (Liu 2012) and Semi-supervised Machine Learning (Søgaard 2013), start their problem statements by referring to data from the Web. This volume starts its own introduction by praising Web corpora for their size, ease of construction, and availability as a source of new text types. A random check of papers from the most recent ACL meeting also shows that the majority of them use Web data in one way or another. Our field definitely needs a comprehensive overview and a DIY manual for the task of constructing a corpus from the Web. This book is, to the best of my knowledge, the first attempt at providing such an overview.The book consists of an introduction and four chapters outlining the four main steps of Web corpus construction. They include: “Data Collection” (Chapter 2), “Basic Corpus Cleaning” (Chapter 3), “Linguistic Processing” (Chapter 4), and “Corpus Evaluation” (Chapter 5).Chapter 2 provides a very useful outline of the main properties of the Web and the crawling strategies. The chapter starts with an overview of a large-scale study of Web connectivity from Baeza-Yates, Castillo, and Efthimiadis (2007), listing various parameters of connectivity for a range of Top-Level Domains. However, there is little discussion of the implications for the corpus development task; for example, does the difference of the in-degree parameter of the Web pages from Chile and the UK have any implications for the Web corpora crawled from those domains? The chapter then proceeds to another important topic, which concerns the parameters of crawling; for example, the crawl bias and the number of seeds, and their influence on the final corpus. Section 2.4.1 illustrates the problems with the crawl bias by an example of deWac, a large commonly used corpus of German (Baroni et al. 2009). The second most frequent proper name bigram in this corpus is found to be Falun Gong. However, more analysis into the nature of the bias should have been beneficial. It is less likely to be related to the PageRank bias, the main bias discussed in Section 2.4.2. Other most frequent bigrams from deWac are not presented in the book, but it is interesting to note that the fourth place in it is occupied by Hartz IV, and the tenth place by Digital Eyes. This suggests that the bias comes from frequency spikes (i.e, a large number of instances collected from a small number of Web sites). Another shortcoming of this chapter is that nothing is said specifically about obtaining data from such resources as Twitter or Facebook, which need access via APIs rather than direct crawling.Chapter 3 introduces methods for basic cleaning of the corpus content, such as processing of text formats (primarily HTML tags), language identification, boilerplate removal, and deduplication. Such low-level tasks are not considered to be glamorous from the view of computational linguistics, but they are extremely important for making Web-derived corpora usable (Baroni et al. 2008). The introduction offered in this chapter is reasonably complete, with good explanations of the sources of problems as well as with suggestions for the tools to be used in each task. An important bit which is missing in this chapter concerns the suggestions for choosing a particular cleaning pipeline. Although the choice indeed depends on the purposes of corpus collection, an indication of which pipeline suits which purpose is desirable.Chapter 4 is devoted to basic steps for linguistic processing of Web corpora, such as tokenization, POS tagging, and lemmatization, as well as orthographic normalization. Even though the processing pipeline is roughly the same for all NLP tasks, it becomes harder for Web corpora because they exhibit greater diversity in comparison with more homogeneous text collections (e.g., WSJ texts). Web texts are also considerably noisier, in the sense of containing nonstandard linguistic expressions, which are likely to be a challenge to the tools trained on more standard texts. The chapter presents some interesting case studies—in particular, the sources of POS tagging errors and non-standard orthography.Chapter 5 describes ways for evaluating and comparing corpora. It gives examples of checking for word and sentence length and for sentence-level duplication. It also introduces methods for comparing frequency lists. Like other chapters it includes many interesting observations, such as the methods for extrinsic evaluation of corpora. However, the chapter does not address many issues important for corpus evaluation and comparison. Given that the previous chapters introduced a number of pipelines and corpora, this chapter would have been an ideal place to illustrate all the aspects of the pipelines by evaluating them in a consistent way. There are occasional references to this goal, such as the frequency lists of French nouns in Section 5.3.1, but this particular comparison is fairly impressionistic, and it concludes with a declaration of basic similarity of the underlying corpora. Does this mean that the crawling, cleaning, and linguistic processing pipelines do not matter? In any case, not even an impressionistic comparison of the pipelines is performed for other evaluation methods. Some illustrations are also not informative (e.g., Table 5.1.1 shows two frequency lists with the identical ranks for their words, which leads to the trivial rank correlation value of 1). The chapter contains a single paragraph devoted to composition of Web corpora. Given the size of such corpora, their evaluation crucially depends on understanding what has been crawled. The task has been approached by a number of models, such as supervised and semi-supervised classification, clustering, topic modeling, and so forth, which should have been included in the discussion. The discussion does contain a relevant reference to Mehler, Sharoff, and Santini (2010), which surveys approaches to the genres of the Web, but other aspects of corpus composition need to be addressed, too.Overall, it is very useful to have a book that introduces all the aspects of Web corpus construction in a single volume with coherent presentation. The volume under review does cover the entire range of topics relevant to Web corpus construction and illustrates them via numerous examples. I would recommend it to students just starting their corpus development experiments. As for the drawbacks of the volume, there is a need to improve the structure of argumentation for the next edition. Bits of information are sometimes introduced in an incomplete way and re-introduced again in subsequent sections. For examples, two tools for crawling are discussed towards the end of Section 2.3.3, while more tools are mentioned as the discussion of crawling strategies progresses. Chapter 1 starts with a fairly random list of non-Web corpora, whereas an overview of the book structure is confined to a short paragraph. Often, frustratingly little information is provided besides an annotated bibliography, rather than a presentation of the relevant methods and issues. In some cases this is accompanied with a statement that “covering this topic is beyond the scope of this volume,” even if the nature of the problem and the solutions could have been easily explained in a one-page summary. Another minor concern is an (understandable) emphasis on the tools and corpora developed by the authors, primarily on their German corpus.I have to admit ambivalence in my final verdict: The book is a useful introduction to an important topic, but it definitely warrants a new edition, which eliminates the shortcomings of the current one.
Serge Sharoff
Comput. Linguistics1
2014 Designing and Evaluating a Reliable Corpus of Web Genres via Crowd-Sourcing
Noushin Rezapour Asheghi, Serge Sharoff, Katja Markert
LREC2
2012 Identifying Word Translations from Comparable Documents Without a Seed Lexicon
Reinhard Rapp, Serge Sharoff, Bogdan Babych
LREC2
2010 Fine-Grained Genre Classification Using Structural Learning Algorithms
Zhi-Li Wu, Katja Markert, Serge Sharoff
ACL3
2010 The Web Library of Babel: evaluating genre collections
Serge Sharoff, Zhi-Li Wu, Katja Markert
LREC1
2010 Advanced Corpus Solutions for Humanities Researchers
Anthony Hartley, Serge Sharoff, Paul Stephenson
PACLIC3
2010 Using an integrated feature set to generalize and justify the Chinese-to-English transferring rule of the 'ZHE' aspect
abstract
In machine translation (MT) practice, there is an urgent need for constructing a set of Chinese-to-English aspect transferring rules to define the transferring conditions. The integrated feature set was used to generalize and justify the Chinese-to-English transferring rule of the ‘ZHE’ aspect (ZHE Rule). A ZHE classification model was built in this study. The impacts of each set of temporal, lexical aspectual, and syntactic features, and their integrated impacts, on the accuracy of the ZHE Rule were tested. Over 600 misclassified corpus sentences were manually examined. A 10-fold cross-validation was used with a decision tree algorithm. The main results are: (1) The ZHE Rule was generalized and justified to have a higher accuracy under the two metrics: the precision rate and the areas under the receiver operating characteristic curve (AUC). (2) The temporal, lexical aspectual, and syntactic feature sets have an integrated contribution to the accuracy of the ZHE Rule. The syntactic and temporal features have an impact on ZHE aspect derivations, while the lexical aspectual features are not predictive of ZHE aspect derivation. (3) While associated with active verbs, the ZHE aspect can denote a perfective situation. This study suggests that the temporal and syntactic features are the predictive ZHE aspect classification features and that the ZHE Rule with an overall precision rate of 80.1% is accurate enough to be further explored in MT practice. The machine learning method, decision tree, can be applied to the automatic aspect transferring in MT research and aspectual interpretations in linguistic research.
Yun-hua Qu, Tian-jiong Tao, Serge Sharoff, Narisong Jin, Ruoyuan Gao, Yu-Ting Yang, Cheng-zhi Xu
J. Zhejiang Univ. Sci. C3
2009 Evaluation-Guided Pre-Editing of Source Text: Improving MT-Tractability of Light Verb Constructions
Bogdan Babych, Anthony Hartley, Serge Sharoff
EAMT3
2008 Generalising Lexical Translation Strategies for MT Using Comparable Corpora
Bogdan Babych, Serge Sharoff, Anthony Hartley
LREC2
2008 Cleaneval: a Competition for Cleaning Web Pages
Marco Baroni, Francis Chantree, Adam Kilgarriff, Serge Sharoff
LREC4
2008 Corpus-Based Tools for Computer-Assisted Acquisition of Reading Abilities in Cognate Languages
Svitlana Kurella, Serge Sharoff, Anthony Hartley
LREC2
2008 Designing and Evaluating a Russian Tagset
Serge Sharoff, Mikhail Kopotev, Tomaz Erjavec, Anna Feldman, Dagmar Divjak
LREC1
2007 Assisting Translators in Indirect Lexical Transfer
Bogdan Babych, Anthony Hartley, Serge Sharoff, Olga Mudraya
ACL3
2007 Translating from under-resourced languages: comparing direct transfer against pivot translation
Bogdan Babych, Anthony Hartley, Serge Sharoff
MTSummit3
2006 Using Comparable Corpora to Solve Problems Difficult for Human Translators
Serge Sharoff, Bogdan Babych, Anthony Hartley
ACL1
2006 ASSIST: Automated Semantic Assistance for Translators
Serge Sharoff, Bogdan Babych, Paul Rayson, Olga Mudraya, Scott Piao
EACL1
2006 Using Richly Annotated Trilingual Language Resources for Acquiring Reading Skills in a Foreign Language
Dragoç Ciobanu, Tony Hartley, Serge Sharoff
LREC3
2006 A Uniform Interface to Large-Scale Linguistic Resources
Serge Sharoff
LREC1
2006 Using collocations from comparable corpora to find translation equivalents
Serge Sharoff, Bogdan Babych, Anthony Hartley
LREC1
2004 Towards Basic Categories for Describing Properties of Texts in a Corpus
Serge Sharoff
LREC1
2002 Meaning as use: exploitation of aligned corpora for the contrastive study of lexical semantics
Serge Sharoff
LREC1
2001 Concordancing for parallel spoken language corpora
abstract
Concordancing is one of the oldest corpus analysis tools, especially for written corpora. In NLP concordancing appears intraining of speech-recognition system. Additionally, comparative studies of different languages result in parallel corpora. Concordancing for these corpora in a NLP context is a new approach. We propose to combine these fields of interest for a multi-purpose concordance for Spoken Language Data, opening the opportunity of combining corpus-linguistic and NLP methods resulting in a broader empirical basis for NLP research. Theoretic models for audio-concordances are discussed. Principles of the structure and design of a parallel audio concordance are given, coding by means of XML to ensure reusability and flexibility, using time stamps for referencing from annotations to the signal.
Dafydd Gibbon, Thorsten Trippel, Serge Sharoff
INTERSPEECH3
2000 Multilinguality in a Text Generation System For Three Slavic Languages
Geert-Jan M. Kruijff, Elke Teich, John A. Bateman, Ivana Kruijff-Korbayová, Hana Skoumalová, Serge Sharoff, Elena G. Sokolova, Tony Hartley, Kamenka Staykova, Jiri Hana
COLING6
2000 Resources for Multilingual Text Generation in Three Slavic Languages
John A. Bateman, Elke Teich, Geert-Jan M. Kruijff, Ivana Kruijff-Korbayová, Serge Sharoff, Hana Skoumalová
LREC5
1999 Register-domain Separation as a Methodology for Development of Natural Language Interfaces to Databases
Serge Sharoff, Vlad Zhigalov
INTERACT1