VLDB 2026 Research / reviewers in the wild / expert
Serge Sharoff
dblp:49/607
· DBLP profile ↗
48ranked-venue papers
13as first author
10since 2021 · last 2026
0000-0002-4877-0210ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 45 · 12 first-author · 10 since 2021Databases, data management, data science and information retrieval · 1Graphics, computer vision, multimedia, augmented reality and games · 1Human-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | To Predict or Not to Predict? Towards Reliable Uncertainty Estimation in the Presence of Noise
Nouran Khallaf, Serge Sharoff |
LREC | 2 |
| 2026 | How Much Noise Can BERT Handle? Insights from Multilingual Sentence Difficulty Detection
Nouran Khallaf, Serge Sharoff |
LREC | 2 |
| 2025 | BERT-based Classical Arabic Poetry Authorship AttributionabstractThis study introduces a novel computational approach to authorship attribution (AA) in Arabic poetry, using the entire Classical Arabic Poetry corpus for the first time and offering a direct analysis of real cases of misattribution. AA in Arabic poetry has been a significant issue since the 9th century, particularly due to the loss of pre-Islamic poetry and the misattribution of post-Islamic works to earlier poets. While previous research has predominantly employed qualitative methods, this study uses computational techniques to address these challenges. The corpus was scraped from online sources and enriched with manually curated Date of Death (DoD) information to overcome the problematic traditional sectioning. Additionally, we applied Embedded Topic Modeling (ETM) to label each poem with its topic contributions, further enhancing the dataset’s value. An ensemble model based on CAMeLBERT was developed and tested across three dimensions: topic, number of poets, and number of training examples. After parameter optimization, the model achieved F1 scores ranging from 0.97 to 1.0. The model was also applied to four pre-Islamic misattribution cases, producing results consistent with historical and literary studies. Lama Alqurashi, Serge Sharoff, Janet Watson, Jacob Blakesley |
COLING | 2 |
| 2025 | Controlling Out-of-Domain Gaps in LLMs for Genre Classification and Generated Text DetectionabstractThis study demonstrates that the modern generation of Large Language Models (LLMs, such as GPT-4) suffers from the same out-of-domain (OOD) performance gap observed in prior research on pre-trained Language Models (PLMs, such as BERT). We demonstrate this across two non-topical classification tasks: (1) genre classification and (2) generated text detection. Our results show that when demonstration examples for In-Context Learning (ICL) come from one domain (e.g., travel) and the system is tested on another domain (e.g., history), classification performance declines significantly. To address this, we introduce a method that controls which predictive indicators are used and which are excluded during classification. For the two tasks studied here, this ensures that topical features are omitted, while the model is guided to focus on stylistic rather than content-based attributes. This approach reduces the OOD gap by up to 20 percentage points in a few-shot setup. Straightforward Chain-of-Thought (CoT) methods, used as the baseline, prove insufficient, while our approach consistently enhances domain transfer performance. Dmitri Roussinov, Serge Sharoff, Nadezhda Puchnina |
COLING | 2 |
| 2024 | Enhancing Image-to-Text Generation in Radiology Reports through Cross-modal Multi-Task LearningabstractImage-to-text generation involves automatically generating descriptive text from images and has applications in medical report generation. However, traditional approaches often exhibit a semantic gap between visual and textual information. In this paper, we propose a multi-task learning framework to leverage both visual and non-imaging data for generating radiology reports. Along with chest X-ray images, 10 additional features comprising numeric, binary, categorical, and text data were incorporated to create a unified representation. The model was trained to generate text, predict the degree of patient severity, and identify medical findings. Multi-task learning, especially with text generation prioritisation, improved performance over single-task baselines across language generation metrics. The framework also mitigated overfitting in auxiliary tasks compared to single-task models. Qualitative analysis showed logically coherent narratives and accurate identification of findings, though some repetition and disjointed phrasing remained. This work demonstrates the benefits of multi-modal, multi-task learning for image-to-text generation applications. Nurbanu Aksoy, Nishant Ravikumar, Serge Sharoff |
LREC/COLING | 3 |
| 2024 | Quantifying the Contribution of MWEs and Polysemy in Translation Errors for English-Igbo MTabstractIn spite of recent successes in improving Machine Translation (MT) quality overall, MT engines require a large amount of resources, which leads to markedly lower quality for lesser-resourced languages. This study explores the case of translation from English into Igbo, a very low resource language spoken by about 45 million speakers. With the aim of improving MT quality in this scenario, we investigate methods for guided detection of critical/harmful MT errors, more specifically those caused by non-compositional multi-word expressions and polysemy. We have designed diagnostic tests for these cases and applied them to collections of medical texts from CDC, Cochrane, NCDC, NHS and WHO. Adaeze Ohuoba, Serge Sharoff, Callum Walker |
EAMT (1) | 2 |
| 2022 | Applying Natural Annotation and Curriculum Learning to Named Entity Recognition for Under-Resourced LanguagesabstractCurrent practices in building new NLP models for low-resourced languages rely either on Machine Translation of training sets from better resourced languages or on cross-lingual transfer from them. Still we can see a considerable performance gap between the models originally trained within better resourced languages and the models transferred from them. In this study we test the possibility of (1) using natural annotation to build synthetic training sets from resources not initially designed for the target downstream task and (2) employing curriculum learning methods to select the most suitable examples from synthetic training sets. We test this hypothesis across seven Slavic languages and across three curriculum learning strategies on Named Entity Recognition as the downstream task. We also test the possibility of fine-tuning the synthetic resources to reflect linguistic properties, such as the grammatical case and gender, both of which are important for the Slavic languages. We demonstrate the possibility to achieve the mean F1 score of 0.78 across the three basic entities types for Belarusian starting from zero resources in comparison to the baseline of 0.63 using the zero-shot transfer from English. For comparison, the English model trained on the original set achieves the mean F1-score of 0.75. The experimental results are available from https://github.com/ValeraLobov/SlavNER Valeriy Lobov, Alexandra Ivoylova, Serge Sharoff |
COLING | 3 |
| 2022 | BERTology for Machine Translation: What BERT Knows about Linguistic Difficulties for TranslationabstractPre-trained transformer-based models, such as BERT, have shown excellent performance in most natural language processing benchmark tests, but we still lack a good understanding of the linguistic knowledge of BERT in Neural Machine Translation (NMT). Our work uses syntactic probes and Quality Estimation (QE) models to analyze the performance of BERT’s syntactic dependencies and their impact on machine translation quality, exploring what kind of syntactic dependencies are difficult for NMT engines based on BERT. While our probing experiments confirm that pre-trained BERT “knows” about syntactic dependencies, its ability to recognize them often decreases after fine-tuning for NMT tasks. We also detect a relationship between syntactic dependencies in three languages and the quality of their translations, which shows which specific syntactic dependencies are likely to be a significant cause of low-quality translations. Yuqian Dai, Marc de Kamps, Serge Sharoff |
LREC | 3 |
| 2022 | Estimating Confidence of Predictions of Individual Classifiers and TheirEnsembles for the Genre Classification TaskabstractGenre identification is a kind of non-topic text classification. The main difference between this task and topic classification is that genre, unlike topic, usually cannot be expressed just by some keywords and is defined as a functional space. Neural models based on pre-trained transformers, such as BERT or XLM-RoBERTa, demonstrate SOTA results in many NLP tasks, including non-topical classification. However, in many cases, their downstream application to very large corpora, such as those extracted from social media, can lead to unreliable results because of dataset shifts, when some raw texts do not match the profile of the training set. To mitigate this problem, we experiment with individual models as well as with their ensembles. To evaluate the robustness of all models we use a prediction confidence metric, which estimates the reliability of a prediction in the absence of a gold standard label. We can evaluate robustness via the confidence gap between the correctly classified texts and the misclassified ones on a labeled test corpus, higher gaps make it easier to identify whether a text is classified correctly. Our results show that for all of the classifiers tested in this study, there is a confidence gap, but for the ensembles, the gap is wider, meaning that ensembles are more robust than their individual models. Mikhail Lepekhin, Serge Sharoff |
LREC | 2 |
| 2022 | Multimodal Pipeline for Collection of Misinformation Data from TelegramabstractThe paper presents the outcomes of AI-COVID19, our project aimed at better understanding of misinformation flow about COVID-19 across social media platforms. The specific focus of the study reported in this paper is on collecting data from Telegram groups which are active in promotion of COVID-related misinformation. Our corpus collected so far contains around 28 million words, from almost one million messages. Given that a substantial portion of misinformation flow in social media is spread via multimodal means, such as images and video, we have also developed a mechanism for utilising such channels via producing automatic transcripts for videos and automatic classification for images into such categories as memes, screenshots of posts and other kinds of images. The accuracy of the image classification pipeline is around 87%. Jose Sosa, Serge Sharoff |
LREC | 2 |
| 2020 | Recognizing Semantic Relations: Attention-Based Transformers vs. Recurrent Models
Dmitri Roussinov, Serge Sharoff, Nadezhda Puchnina |
ECIR (1) | 2 |
| 2020 | Recognizing Semantic Relations by Combining Transformers and Fully Connected ModelsabstractAutomatically recognizing an existing semantic relation (e.g. “is a”, “part of”, “property of”, “opposite of” etc.) between two words (phrases, concepts, etc.) is an important task affecting many NLP applications and has been subject of extensive experimentation and modeling. Current approaches to automatically telling if a relation exists between two given concepts X and Y can be grouped into two types: 1) those modeling word-paths connecting X and Y in text and 2) those modeling distributional properties of X and Y separately, not necessary in the proximity to each other. Here, we investigate how both types can be improved and combined. We suggest a distributional approach that is based on an attention-based transformer. We have also developed a novel word path model that combines useful properties of a convolutional network with a fully connected language model. While our transformer-based approach works better, both our models significantly outperform the state-of-the-art within their classes of approaches. We also demonstrate that combining the two approaches results in additional gains since they use somewhat different data sources. Dmitri Roussinov, Serge Sharoff, Nadezhda Puchnina |
LREC | 2 |
| 2020 | Know thy Corpus! Robust Methods for Digital Curation of Web corporaabstractThis paper proposes a novel framework for digital curation of Web corpora in order to provide robust estimation of their parameters, such as their composition and the lexicon. In recent years language models pre-trained on large corpora emerged as clear winners in numerous NLP tasks, but no proper analysis of the corpora which led to their success has been conducted. The paper presents a procedure for robust frequency estimation, which helps in establishing the core lexicon for a given corpus, as well as a procedure for estimating the corpus composition via unsupervised topic models and via supervised genre classification of Web pages. The results of the digital curation study applied to several Web-derived corpora demonstrate their considerable differences. First, this concerns different frequency bursts which impact the core lexicon obtained from each corpus. Second, this concerns the kinds of texts they contain. For example, OpenWebText contains considerably more topical news and political argumentation in comparison to ukWac or Wikipedia. The tools and the results of analysis have been released. Serge Sharoff |
LREC | 1 |
| 2020 | Sentence Level Human Translation Quality Estimation with Attention-based Neural NetworksabstractThis paper explores the use of Deep Learning methods for automatic estimation of quality of human translations. Automatic estimation can provide useful feedback for translation teaching, examination and quality control. Conventional methods for solving this task rely on manually engineered features and external knowledge. This paper presents an end-to-end neural model without feature engineering, incorporating a cross attention mechanism to detect which parts in sentence pairs are most relevant for assessing quality. Another contribution concerns oprediction of fine-grained scores for measuring different aspects of translation quality, such as terminological accuracy or idiomatic writing. Empirical results on a large human annotated dataset show that the neural model outperforms feature-based methods significantly. The dataset and the tools are available. Serge Sharoff |
LREC | 2 |
| 2020 | Finding next of kin: Cross-lingual embedding spaces for related languagesabstractAbstract Some languages have very few NLP resources, while many of them are closely related to better-resourced languages. This paper explores how the similarity between the languages can be utilised by porting resources from better- to lesser-resourced languages. The paper introduces a way of building a representation shared across related languages by combining cross-lingual embedding methods with a lexical similarity measure which is based on the weighted Levenshtein distance. One of the outcomes of the experiments is a Panslavonic embedding space for nine Balto-Slavonic languages. The paper demonstrates that the resulting embedding space helps in such applications as morphological prediction, named-entity recognition and genre classification. Serge Sharoff |
Nat. Lang. Eng. | 1 |
| 2018 | Language adaptation experiments via cross-lingual embeddings for related languages
Serge Sharoff |
LREC | 1 |
| 2018 | Cross-lingual Terminology Extraction for Translation Quality Estimation
Yuze Gao, Yue Zhang 0004, Serge Sharoff |
LREC | 4 |
| 2018 | Investigating the Influence of Bilingual MWU on Trainee Translation Quality
Serge Sharoff |
LREC | 2 |
| 2018 | A Multilingual Dataset for Evaluating Parallel Sentence Extraction from Comparable Corpora
Pierre Zweigenbaum, Serge Sharoff, Reinhard Rapp |
LREC | 2 |
| 2016 | Adam Kilgarriff's Legacy to Computational Linguistics and Beyond
Roger Evans, Alexander F. Gelbukh, Gregory Grefenstette, Patrick Hanks, Milos Jakubícek, Diana McCarthy, Martha Palmer, Ted Pedersen, Michael Rundell, Pavel Rychlý, Serge Sharoff, David Tugwell |
CICLing (1) | 11 |
| 2016 | MoBiL: A Hybrid Feature Set for Automatic Human Translation Quality Assessment
Serge Sharoff, Bogdan Babych |
LREC | 2 |
| 2016 | PrefaceabstractAfter several decades of work on rule-based machine translation (MT) where linguists try to manually encode their knowledge about language, the time around 1990 brought a paradigm change towards automatic systems which try to learn how to translate by looking at large collections of high-quality sample translations as produced by professional translators. The first such attempts were called example- or analogy-based translation, and somewhat later the so-called statistical approach to MT was introduced. Both can be subsumed under the label data-driven approaches to MT. It took about 10 years until these self-learning systems became serious competitors of the traditional rule-based systems, and by now some of the most successful MT systems, such as Google Translate and Moses, are based on the statistical approach. Reinhard Rapp, Serge Sharoff, Pierre Zweigenbaum |
Nat. Lang. Eng. | 2 |
| 2016 | Recent advances in machine translation using comparable corporaabstractAbstract This paper highlights some of the recent developments in the field of machine translation using comparable corpora. We start by updating previous definitions of comparable corpora and then look at bilingual versions of continuous vector space models. Recently, neural networks have been used to obtain latent context representations with only few dimensions which are often called word embeddings. These promising new techniques cannot only be applied to parallel but also to comparable corpora. Subsequent sections of the paper discuss work specifically targeting at machine translation using comparable corpora, as well as work dealing with the extraction of parallel segments from comparable corpora. Finally, we give an overview on the design and the results of a recent shared task on measuring document comparability across languages. Reinhard Rapp, Serge Sharoff, Pierre Zweigenbaum |
Nat. Lang. Eng. | 2 |
| 2015 | Web Corpus Construction Roland Schäfer and Felix Bildhauer (Freie Universität Berlin) Morgan & Claypool (Synthesis Lectures on Human Language Technologies, edited by Graeme Hirst, volume 22), 2013, 145 pages, paper-bound, ISBN 9781608459834, doi: 10.2200/S00508ED1V01Y201305HLT022abstractThe Web is the main source of data in modern computational linguistics. Other volumes in the same series, for example, Introductions to Opinion Mining (Liu 2012) and Semi-supervised Machine Learning (Søgaard 2013), start their problem statements by referring to data from the Web. This volume starts its own introduction by praising Web corpora for their size, ease of construction, and availability as a source of new text types. A random check of papers from the most recent ACL meeting also shows that the majority of them use Web data in one way or another. Our field definitely needs a comprehensive overview and a DIY manual for the task of constructing a corpus from the Web. This book is, to the best of my knowledge, the first attempt at providing such an overview.The book consists of an introduction and four chapters outlining the four main steps of Web corpus construction. They include: “Data Collection” (Chapter 2), “Basic Corpus Cleaning” (Chapter 3), “Linguistic Processing” (Chapter 4), and “Corpus Evaluation” (Chapter 5).Chapter 2 provides a very useful outline of the main properties of the Web and the crawling strategies. The chapter starts with an overview of a large-scale study of Web connectivity from Baeza-Yates, Castillo, and Efthimiadis (2007), listing various parameters of connectivity for a range of Top-Level Domains. However, there is little discussion of the implications for the corpus development task; for example, does the difference of the in-degree parameter of the Web pages from Chile and the UK have any implications for the Web corpora crawled from those domains? The chapter then proceeds to another important topic, which concerns the parameters of crawling; for example, the crawl bias and the number of seeds, and their influence on the final corpus. Section 2.4.1 illustrates the problems with the crawl bias by an example of deWac, a large commonly used corpus of German (Baroni et al. 2009). The second most frequent proper name bigram in this corpus is found to be Falun Gong. However, more analysis into the nature of the bias should have been beneficial. It is less likely to be related to the PageRank bias, the main bias discussed in Section 2.4.2. Other most frequent bigrams from deWac are not presented in the book, but it is interesting to note that the fourth place in it is occupied by Hartz IV, and the tenth place by Digital Eyes. This suggests that the bias comes from frequency spikes (i.e, a large number of instances collected from a small number of Web sites). Another shortcoming of this chapter is that nothing is said specifically about obtaining data from such resources as Twitter or Facebook, which need access via APIs rather than direct crawling.Chapter 3 introduces methods for basic cleaning of the corpus content, such as processing of text formats (primarily HTML tags), language identification, boilerplate removal, and deduplication. Such low-level tasks are not considered to be glamorous from the view of computational linguistics, but they are extremely important for making Web-derived corpora usable (Baroni et al. 2008). The introduction offered in this chapter is reasonably complete, with good explanations of the sources of problems as well as with suggestions for the tools to be used in each task. An important bit which is missing in this chapter concerns the suggestions for choosing a particular cleaning pipeline. Although the choice indeed depends on the purposes of corpus collection, an indication of which pipeline suits which purpose is desirable.Chapter 4 is devoted to basic steps for linguistic processing of Web corpora, such as tokenization, POS tagging, and lemmatization, as well as orthographic normalization. Even though the processing pipeline is roughly the same for all NLP tasks, it becomes harder for Web corpora because they exhibit greater diversity in comparison with more homogeneous text collections (e.g., WSJ texts). Web texts are also considerably noisier, in the sense of containing nonstandard linguistic expressions, which are likely to be a challenge to the tools trained on more standard texts. The chapter presents some interesting case studies—in particular, the sources of POS tagging errors and non-standard orthography.Chapter 5 describes ways for evaluating and comparing corpora. It gives examples of checking for word and sentence length and for sentence-level duplication. It also introduces methods for comparing frequency lists. Like other chapters it includes many interesting observations, such as the methods for extrinsic evaluation of corpora. However, the chapter does not address many issues important for corpus evaluation and comparison. Given that the previous chapters introduced a number of pipelines and corpora, this chapter would have been an ideal place to illustrate all the aspects of the pipelines by evaluating them in a consistent way. There are occasional references to this goal, such as the frequency lists of French nouns in Section 5.3.1, but this particular comparison is fairly impressionistic, and it concludes with a declaration of basic similarity of the underlying corpora. Does this mean that the crawling, cleaning, and linguistic processing pipelines do not matter? In any case, not even an impressionistic comparison of the pipelines is performed for other evaluation methods. Some illustrations are also not informative (e.g., Table 5.1.1 shows two frequency lists with the identical ranks for their words, which leads to the trivial rank correlation value of 1). The chapter contains a single paragraph devoted to composition of Web corpora. Given the size of such corpora, their evaluation crucially depends on understanding what has been crawled. The task has been approached by a number of models, such as supervised and semi-supervised classification, clustering, topic modeling, and so forth, which should have been included in the discussion. The discussion does contain a relevant reference to Mehler, Sharoff, and Santini (2010), which surveys approaches to the genres of the Web, but other aspects of corpus composition need to be addressed, too.Overall, it is very useful to have a book that introduces all the aspects of Web corpus construction in a single volume with coherent presentation. The volume under review does cover the entire range of topics relevant to Web corpus construction and illustrates them via numerous examples. I would recommend it to students just starting their corpus development experiments. As for the drawbacks of the volume, there is a need to improve the structure of argumentation for the next edition. Bits of information are sometimes introduced in an incomplete way and re-introduced again in subsequent sections. For examples, two tools for crawling are discussed towards the end of Section 2.3.3, while more tools are mentioned as the discussion of crawling strategies progresses. Chapter 1 starts with a fairly random list of non-Web corpora, whereas an overview of the book structure is confined to a short paragraph. Often, frustratingly little information is provided besides an annotated bibliography, rather than a presentation of the relevant methods and issues. In some cases this is accompanied with a statement that “covering this topic is beyond the scope of this volume,” even if the nature of the problem and the solutions could have been easily explained in a one-page summary. Another minor concern is an (understandable) emphasis on the tools and corpora developed by the authors, primarily on their German corpus.I have to admit ambivalence in my final verdict: The book is a useful introduction to an important topic, but it definitely warrants a new edition, which eliminates the shortcomings of the current one. Serge Sharoff |
Comput. Linguistics | 1 |
| 2014 | Designing and Evaluating a Reliable Corpus of Web Genres via Crowd-Sourcing
Noushin Rezapour Asheghi, Serge Sharoff, Katja Markert |
LREC | 2 |
| 2012 | Identifying Word Translations from Comparable Documents Without a Seed Lexicon
Reinhard Rapp, Serge Sharoff, Bogdan Babych |
LREC | 2 |
| 2010 | Fine-Grained Genre Classification Using Structural Learning Algorithms
Zhi-Li Wu, Katja Markert, Serge Sharoff |
ACL | 3 |
| 2010 | The Web Library of Babel: evaluating genre collections
Serge Sharoff, Zhi-Li Wu, Katja Markert |
LREC | 1 |
| 2010 | Advanced Corpus Solutions for Humanities Researchers
Anthony Hartley, Serge Sharoff, Paul Stephenson |
PACLIC | 3 |
| 2010 | Using an integrated feature set to generalize and justify the Chinese-to-English transferring rule of the 'ZHE' aspectabstractIn machine translation (MT) practice, there is an urgent need for constructing a set of Chinese-to-English aspect transferring rules to define the transferring conditions. The integrated feature set was used to generalize and justify the Chinese-to-English transferring rule of the ‘ZHE’ aspect (ZHE Rule). A ZHE classification model was built in this study. The impacts of each set of temporal, lexical aspectual, and syntactic features, and their integrated impacts, on the accuracy of the ZHE Rule were tested. Over 600 misclassified corpus sentences were manually examined. A 10-fold cross-validation was used with a decision tree algorithm. The main results are: (1) The ZHE Rule was generalized and justified to have a higher accuracy under the two metrics: the precision rate and the areas under the receiver operating characteristic curve (AUC). (2) The temporal, lexical aspectual, and syntactic feature sets have an integrated contribution to the accuracy of the ZHE Rule. The syntactic and temporal features have an impact on ZHE aspect derivations, while the lexical aspectual features are not predictive of ZHE aspect derivation. (3) While associated with active verbs, the ZHE aspect can denote a perfective situation. This study suggests that the temporal and syntactic features are the predictive ZHE aspect classification features and that the ZHE Rule with an overall precision rate of 80.1% is accurate enough to be further explored in MT practice. The machine learning method, decision tree, can be applied to the automatic aspect transferring in MT research and aspectual interpretations in linguistic research. Yun-hua Qu, Tian-jiong Tao, Serge Sharoff, Narisong Jin, Ruoyuan Gao, Yu-Ting Yang, Cheng-zhi Xu |
J. Zhejiang Univ. Sci. C | 3 |
| 2009 | Evaluation-Guided Pre-Editing of Source Text: Improving MT-Tractability of Light Verb Constructions
Bogdan Babych, Anthony Hartley, Serge Sharoff |
EAMT | 3 |
| 2008 | Generalising Lexical Translation Strategies for MT Using Comparable Corpora
Bogdan Babych, Serge Sharoff, Anthony Hartley |
LREC | 2 |
| 2008 | Cleaneval: a Competition for Cleaning Web Pages
Marco Baroni, Francis Chantree, Adam Kilgarriff, Serge Sharoff |
LREC | 4 |
| 2008 | Corpus-Based Tools for Computer-Assisted Acquisition of Reading Abilities in Cognate Languages
Svitlana Kurella, Serge Sharoff, Anthony Hartley |
LREC | 2 |
| 2008 | Designing and Evaluating a Russian Tagset
Serge Sharoff, Mikhail Kopotev, Tomaz Erjavec, Anna Feldman, Dagmar Divjak |
LREC | 1 |
| 2007 | Assisting Translators in Indirect Lexical Transfer
Bogdan Babych, Anthony Hartley, Serge Sharoff, Olga Mudraya |
ACL | 3 |
| 2007 | Translating from under-resourced languages: comparing direct transfer against pivot translation
Bogdan Babych, Anthony Hartley, Serge Sharoff |
MTSummit | 3 |
| 2006 | Using Comparable Corpora to Solve Problems Difficult for Human Translators
Serge Sharoff, Bogdan Babych, Anthony Hartley |
ACL | 1 |
| 2006 | ASSIST: Automated Semantic Assistance for Translators
Serge Sharoff, Bogdan Babych, Paul Rayson, Olga Mudraya, Scott Piao |
EACL | 1 |
| 2006 | Using Richly Annotated Trilingual Language Resources for Acquiring Reading Skills in a Foreign Language
Dragoç Ciobanu, Tony Hartley, Serge Sharoff |
LREC | 3 |
| 2006 | A Uniform Interface to Large-Scale Linguistic Resources
Serge Sharoff |
LREC | 1 |
| 2006 | Using collocations from comparable corpora to find translation equivalents
Serge Sharoff, Bogdan Babych, Anthony Hartley |
LREC | 1 |
| 2004 | Towards Basic Categories for Describing Properties of Texts in a Corpus
Serge Sharoff |
LREC | 1 |
| 2002 | Meaning as use: exploitation of aligned corpora for the contrastive study of lexical semantics
Serge Sharoff |
LREC | 1 |
| 2001 | Concordancing for parallel spoken language corporaabstractConcordancing is one of the oldest corpus analysis tools, especially for written corpora. In NLP concordancing appears intraining of speech-recognition system. Additionally, comparative studies of different languages result in parallel corpora. Concordancing for these corpora in a NLP context is a new approach. We propose to combine these fields of interest for a multi-purpose concordance for Spoken Language Data, opening the opportunity of combining corpus-linguistic and NLP methods resulting in a broader empirical basis for NLP research. Theoretic models for audio-concordances are discussed. Principles of the structure and design of a parallel audio concordance are given, coding by means of XML to ensure reusability and flexibility, using time stamps for referencing from annotations to the signal. Dafydd Gibbon, Thorsten Trippel, Serge Sharoff |
INTERSPEECH | 3 |
| 2000 | Multilinguality in a Text Generation System For Three Slavic Languages
Geert-Jan M. Kruijff, Elke Teich, John A. Bateman, Ivana Kruijff-Korbayová, Hana Skoumalová, Serge Sharoff, Elena G. Sokolova, Tony Hartley, Kamenka Staykova, Jiri Hana |
COLING | 6 |
| 2000 | Resources for Multilingual Text Generation in Three Slavic Languages
John A. Bateman, Elke Teich, Geert-Jan M. Kruijff, Ivana Kruijff-Korbayová, Serge Sharoff, Hana Skoumalová |
LREC | 5 |
| 1999 | Register-domain Separation as a Methodology for Development of Natural Language Interfaces to Databases
Serge Sharoff, Vlad Zhigalov |
INTERACT | 1 |