VLDB 2026 Research / reviewers in the wild / expert
Constantin Orasan
dblp:83/316
· DBLP profile ↗
49ranked-venue papers
18as first author
11since 2021 · last 2025
0000-0003-2067-8890ORCID · reported
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 47 · 18 first-author · 10 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Prompt-based Explainable Quality Estimation for English-MalayalamabstractThe aim of this project was to curate data for the English-Malayalam language pair for the tasks of Quality Estimation (QE) and Automatic Post-Editing (APE) of Machine Translation. Whilst the primary aim of the project was to create a dataset for a low-resource language pair, we plan to use this dataset to investigate different zero-shot and few-shot prompting strategies including chain-of-thought, towards a unified explainable QE-APE framework. Archchana Sindhujan, Diptesh Kanojia, Constantin Orasan |
MTSummit (2) | 3 |
| 2024 | Product Retrieval and Ranking for Alphanumeric QueriesabstractThis talk addresses the challenge of improving user experience on e-commerce platforms by enhancing product ranking relevant to user's search queries. Queries such as S2716DG consist of alphanumeric characters where a letter or number can signify important detail for the product/model. Speaker describes recent research where we curate samples from existing datasets at eBay, manually annotated with buyer-centric relevance scores, and centrality scores which reflect how well the product title matches the user's intent. We introduce a User-intent Centrality Optimization (UCO) approach for existing models, which optimizes for the user intent in semantic product search. To that end, we propose a dual-loss based optimization to handle hard negatives, i.e., product titles that are semantically relevant but do not reflect the user's intent. Our contributions include curating a challenging evaluation set and implementing UCO, resulting in significant improvements in product ranking efficiency, observed for different evaluation metrics. Our work aims to ensure that the most buyer-centric titles for a query are ranked higher, thereby, enhancing the user experience on e-commerce platforms. Hadeel Saadany, Swapnil Bhosale, Samarth Agrawal, Constantin Orasan, Diptesh Kanojia |
CIKM | 5 |
| 2024 | Linking Judgement Text to Court Hearing Videos: UK Supreme Court as a Case StudyabstractOne the most important archived legal material in the UK is the video recordings of Supreme Court hearings and their corresponding judgements. The impact of Supreme Court published material extends far beyond the parties involved in any given case as it provides landmark rulings on points of law of the greatest public and constitutional importance. Typically, transcripts of legal hearings are lengthy, making it time-consuming for legal professionals to analyse crucial arguments. This study focuses on summarising the second phase of a collaborative research-industrial project aimed at creating an automatic tool designed to connect sections of written judgements with relevant moments in Supreme Court hearing videos, streamlining access to critical information. Acting as a User-Interface (UI) platform, the tool enhances access to justice by pinpointing significant moments in the videos, aiding in comprehension of the final judgement. We make available the initial dataset of judgement-hearing pairs for legal Information Retrieval research, and elucidate our use of AI generative technology to enhance it. Additionally, we demonstrate how fine-tuning GPT text embeddings to our dataset optimises accuracy for an automated linking system tailored to the legal domain. Hadeel Saadany, Constantin Orasan, Sophie Walker, Catherine Breslin |
LREC/COLING | 2 |
| 2024 | Character-level Language Models for Abbreviation and Long-form Detection
Leonardo Zilio, Shenbin Qian, Diptesh Kanojia, Constantin Orasan |
LREC/COLING | 4 |
| 2024 | Evaluating Machine Translation for Emotion-loaded User Generated Content (TransEval4Emo-UGC)abstractThis paper presents a dataset for evaluating the machine translation of emotion-loaded user generated content. It contains human-annotated quality evaluation data and post-edited reference translations. The dataset is available at our GitHub repository. Shenbin Qian, Constantin Orasan, Félix do Carmo, Diptesh Kanojia |
EAMT (2) | 2 |
| 2024 | What do Large Language Models Need for Machine Translation Evaluation?abstractShenbin Qian, Archchana Sindhujan, Minnie Kabra, Diptesh Kanojia, Constantin Orasan, Tharindu Ranasinghe, Fred Blain. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Shenbin Qian, Archchana Sindhujan, Minnie Kabra, Diptesh Kanojia, Constantin Orasan, Tharindu Ranasinghe, Frédéric Blain |
EMNLP | 5 |
| 2023 | Evaluation of Chinese-English Machine Translation of Emotion-Loaded Microblog Texts: A Human Annotated Dataset for the Quality Assessment of Emotion TranslationabstractIn this paper, we focus on how current Machine Translation (MT) engines perform on the translation of emotion-loaded texts by evaluating outputs from Google Translate according to a framework proposed in this paper. We propose this evaluation framework based on the Multidimensional Quality Metrics (MQM) and perform detailed error analyses of the MT outputs. From our analysis, we observe that about 50% of MT outputs are erroneous in preserving emotions. After further analysis of the erroneous examples, we find that emotion carrying words and linguistic phenomena such as polysemous words, negation, abbreviation etc., are common causes for these translation errors. Shenbin Qian, Constantin Orasan, Félix do Carmo, Qiuliang Li, Diptesh Kanojia |
EAMT | 2 |
| 2023 | Analysing Mistranslation of Emotions in Multilingual Tweets by Online MT ToolsabstractIt is common for websites that contain User-Generated Text (UGT) to provide an automatic translation option to reach out to their linguistically diverse users. In such scenarios, the process of translating the users’ emotions is entirely automatic with no human intervention, neither for post-editing, nor for accuracy checking. In this paper, we assess whether automatic translation tools can be a successful real-life utility in transferring emotion in multilingual tweets. Our analysis shows that the mistranslation of the source tweet can lead to critical errors where the emotion is either completely lost or flipped to an opposite sentiment. We identify linguistic phenomena specific to Twitter data which pose a challenge in translation of emotions and show how frequent these features are in different language pairs. We also show that commonly-used quality metrics can lend false confidence in the performance of online MT tools specifically when the source emotion is distorted in telegraphic messages such as tweets. Hadeel Saadany, Constantin Orasan, Rocio Caro Quintana, Félix do Carmo, Leonardo Zilio |
EAMT | 2 |
| 2022 | A Semi-Automated Live Interlingual Communication Workflow Featuring Intralingual Respeaking: Evaluation and BenchmarkingabstractIn this paper, we present a semi-automated workflow for live interlingual speech-to-text communication which seeks to reduce the shortcomings of existing ASR systems: a human respeaker works with a speaker-dependent speech recognition software (e.g., Dragon Naturally Speaking) to deliver punctuated same-language output of superior quality than obtained using out-of-the-box automatic speech recognition of the original speech. This is fed into a machine translation engine (the EU’s eTranslation) to produce live-caption ready text. We benchmark the quality of the output against the output of best-in-class (human) simultaneous interpreters working with the same source speeches from plenary sessions of the European Parliament. To evaluate the accuracy and facilitate the comparison between the two types of output, we use a tailored annotation approach based on the NTR model (Romero-Fresco and Pöchhacker, 2017). We find that the semi-automated workflow combining intralingual respeaking and machine translation is capable of generating outputs that are similar in terms of accuracy and completeness to the outputs produced in the benchmarking workflow, although the small scale of our experiment requires caution in interpreting this result. Tomasz Korybski, Elena Davitti, Constantin Orasan, Sabine Braun |
LREC | 3 |
| 2022 | PLOD: An Abbreviation Detection Dataset for Scientific DocumentsabstractThe detection and extraction of abbreviations from unstructured texts can help to improve the performance of Natural Language Processing tasks, such as machine translation and information retrieval. However, in terms of publicly available datasets, there is not enough data for training deep-neural-networks-based models to the point of generalising well over data. This paper presents PLOD, a large-scale dataset for abbreviation detection and extraction that contains 160k+ segments automatically annotated with abbreviations and their long forms. We performed manual validation over a set of instances and a complete automatic validation for this dataset. We then used it to generate several baseline models for detecting abbreviations and long forms. The best models achieved an F1-score of 0.92 for abbreviations and 0.89 for detecting their corresponding long forms. We release this dataset along with our code and all the models publicly at https://github.com/surrey-nlp/PLOD-AbbreviationDetection Leonardo Zilio, Hadeel Saadany, Diptesh Kanojia, Constantin Orasan |
LREC | 5 |
| 2022 | Biographical Semi-Supervised Relation Extraction DatasetabstractExtracting biographical information from online documents is a popular research topic among the information extraction (IE) community. Various natural language processing (NLP) techniques such as text classification, text summarisation and relation extraction are commonly used to achieve this. Among these techniques, RE is the most common since it can be directly used to build biographical knowledge graphs. RE is usually framed as a supervised machine learning (ML) problem, where ML models are trained on annotated datasets. However, there are few annotated datasets for RE since the annotation process can be costly and time-consuming. To address this, we developedBiographical, the first semi-supervised dataset for RE. The dataset, which is aimed towards digital humanities (DH) and historical research, is automatically compiled by aligning sentences from Wikipedia articles with matching structured data from sources including Pantheon and Wikidata. By exploiting the structure of Wikipedia articles and robust named entity recognition (NER), we match information with relatively high precision in order to compile annotated relation pairs for ten different relations that are important in the DH domain. Furthermore, we demonstrate the effectiveness of the dataset by training a state-of-the-art neural model to classify relation pairs, and evaluate it on a manually annotated gold standard set.Biographical is primarily aimed at training neural models for RE within the domain of digital humanities and history, but as we discuss at the end of this paper, it can be useful for other purposes as well. Alistair Plum, Tharindu Ranasinghe, Spencer Jones, Constantin Orasan, Ruslan Mitkov |
SIGIR | 4 |
| 2020 | TransQuest: Translation Quality Estimation with Cross-lingual TransformersabstractRecent years have seen big advances in the field of sentence-level quality estimation (QE), largely as a result of using neural-based architectures.However, the majority of these methods work only on the language pair they are trained on and need retraining for new language pairs.This process can prove difficult from a technical point of view and is usually computationally expensive.In this paper we propose a simple QE framework based on cross-lingual transformers, and we use it to implement and evaluate two different neural architectures.Our evaluation shows that the proposed methods achieve state-of-the-art results outperforming current open-source quality estimation frameworks when trained on datasets from WMT.In addition, the framework proves very useful in transfer learning settings, especially when dealing with low-resourced languages, allowing us to obtain very competitive results. Tharindu Ranasinghe, Constantin Orasan, Ruslan Mitkov |
COLING | 2 |
| 2020 | Intelligent Translation Memory Matching and Retrieval with Sentence EncodersabstractMatching and retrieving previously translated segments from the Translation Memory is a key functionality in Translation Memories systems. However this matching and retrieving process is still limited to algorithms based on edit distance which we have identified as a major drawback in Translation Memories systems. In this paper, we introduce sentence encoders to improve matching and retrieving process in Translation Memories systems - an effective and efficient solution to replace edit distance-based algorithms. Tharindu Ranasinghe, Constantin Orasan, Ruslan Mitkov |
EAMT | 2 |
| 2019 | Identifying signs of syntactic complexity for rule-based sentence simplificationabstractAbstract This article presents a new method to automatically simplify English sentences. The approach is designed to reduce the number of compound clauses and nominally bound relative clauses in input sentences. The article provides an overview of a corpus annotated with information about various explicit signs of syntactic complexity and describes the two major components of a sentence simplification method that works by exploiting information on the signs occurring in the sentences of a text. The first component is a sign tagger which automatically classifies signs in accordance with the annotation scheme used to annotate the corpus. The second component is an iterative rule-based sentence transformation tool. Exploiting the sign tagger in conjunction with other NLP components, the sentence transformation tool automatically rewrites long sentences containing compound clauses and nominally bound relative clauses as sequences of shorter single-clause sentences. Evaluation of the different components reveals acceptable performance in rewriting sentences containing compound clauses but less accuracy when rewriting sentences containing nominally bound relative clauses. A detailed error analysis revealed that the major sources of error include inaccurate sign tagging, the relatively limited coverage of the rules used to rewrite sentences, and an inability to discriminate between various subtypes of clause coordination. Despite this, the system performed well in comparison with two baselines. This finding was reinforced by automatic estimations of the readability of system output and by surveys of readers’ opinions about the accuracy, accessibility, and meaning of this output. Richard Evans 0002, Constantin Orasan |
Nat. Lang. Eng. | 2 |
| 2019 | Automatic summarisation: 25 years OnabstractAbstract Automatic text summarisation is a topic that has been receiving attention from the research community from the early days of computational linguistics, but it really took off around 25 years ago. This article presents the main developments from the last 25 years. It starts by defining what a summary is and how its definition changed over time as a result of the interest in processing new types of documents. The article continues with a brief history of the field and highlights the main challenges posed by the evaluation of summaries. The article finishes with some thoughts about the future of the field. Constantin Orasan |
Nat. Lang. Eng. | 1 |
| 2017 | Word from the editors
Constantin Orasan, Marcello Federico |
Mach. Transl. | 1 |
| 2016 | Semantic Textual Similarity in Quality Estimation
Hannah Béchara, Carla Parra Escartín, Constantin Orasan, Lucia Specia |
EAMT | 3 |
| 2016 | The first Automatic Translation Memory Cleaning Shared Task
Eduard Barbu, Carla Parra Escartín, Luisa Bentivogli, Matteo Negri, Marco Turchi, Constantin Orasan, Marcello Federico |
Mach. Transl. | 6 |
| 2016 | Improving translation memory matching and retrieval using paraphrases
Rohit Gupta 0007, Constantin Orasan, Marcos Zampieri, Mihaela Vela, Josef van Genabith, Ruslan Mitkov |
Mach. Transl. | 2 |
| 2016 | Word from the editors
Constantin Orasan, Marcello Federico |
Mach. Transl. | 1 |
| 2015 | Can Translation Memories afford not to use paraphrasing?
Rohit Gupta 0007, Constantin Orasan, Marcos Zampieri, Mihaela Vela, Josef van Genabith |
EAMT | 2 |
| 2015 | ReVal: A Simple and Effective Machine Translation Evaluation Metric Based on Recurrent Neural NetworksabstractMany state-of-the-art Machine Translation (MT) evaluation metrics are complex, involve extensive external resources (e.g. for paraphrasing) and require tuning to achieve best results.We present a simple alternative approach based on dense vector spaces and recurrent neural networks (RNNs), in particular Long Short Term Memory (LSTM) networks.For WMT-14, our new metric scores best for two out of five language pairs, and overall best and second best on all language pairs, using Spearman and Pearson correlation, respectively.We also show how training data is computed automatically from WMT ranks data. Rohit Gupta 0007, Constantin Orasan, Josef van Genabith |
EMNLP | 2 |
| 2014 | Incorporating paraphrasing in translation memory matching and retrieval
Rohit Gupta 0007, Constantin Orasan |
EAMT | 2 |
| 2014 | Densification: Semantic document analysis using WikipediaabstractAbstract This paper proposes a new method for semantic document analysis: densification, which identifies and ranks Wikipedia pages relevant to a given document. Although there are similarities with established tasks such as wikification and entity linking, the method does not aim for strict disambiguation of named entity mentions. Instead, densification uses existing links to rank additional articles that are relevant to the document, a form of explicit semantic indexing that enables higher-level semantic retrieval procedures that can be beneficial for a wide range of NLP applications. Because a gold standard for densification evaluation does not exist, a study is carried out to investigate the level of agreement achievable by humans, which questions the feasibility of creating an annotated data set. As a result, a semi-supervised approach is employed to develop a two-stage densification system: filtering unlikely candidate links and then ranking the remaining links. In a first evaluation experiment, Wikipedia articles are used to automatically estimate the performance in terms of recall. Results show that the proposed densification approach outperforms several wikification systems. A second experiment measures the impact of integrating the links predicted by the densification system into a semantic question answering (QA) system that relies on Wikipedia links to answer complex questions. Densification enables the QA system to find twice as many additional answers than when using a state-of-the-art wikification system. Iustin Dornescu, Constantin Orasan |
Nat. Lang. Eng. | 2 |
| 2012 | Annotating Near-Identity from Coreference Disagreements
Marta Recasens, Maria Antònia Martí, Constantin Orasan |
LREC | 3 |
| 2012 | CLCM - A Linguistic Resource for Effective Simplification of Instructions in the Crisis Management Domain and its Evaluations
Irina P. Temnikova, Constantin Orasan, Ruslan Mitkov |
LREC | 2 |
| 2012 | Interactive Multi-Modal Question-Answering Antal van den Bosch* and Gosse Bouma‡ (editors) (*Tilburg University and ‡University of Groningen) Berlin: Springer (Theory and Applications of Natural Language Processing series, edited by Eduard Hovy), 2011, xii+279 pp; hardbound, ISBN 978-3-642-17524-4, $124.00; e-book, ISBN 978-3-642-17525-1; paperbound, $24.95 or €24.95 to members of subscribing institutionsabstractProcessing and presentation of multimodal information was one of the important directions pursued by researchers in the areas related to information processing and management in the first decade of this century (Stock and Zancanaro 2005; Maragos, Potamianos, and Gros 2008; Lalanne et al. 2009). The Interactive Multimodal Information eXtraction (IMIX) Programme, a research program that ran between 2004 and 2009 and was funded by the Netherlands Organisation for Scientific Research (NWO), adhered to this direction of research. This book contains a collection of articles describing research carried out in the IMIX Programme. Given the large scale of the program, the book covers only parts of it, arguably the most important ones: question answering, (spoken) dialogue systems, and human–machine interaction.The book is organized into four parts and an epilogue. The first part introduces the IMIX Programme and the demonstrator developed by it. The main purpose of the program was to bring together research groups from the Netherlands to build an interactive multimodal question answering (QA) system that is able to answer general encyclopedic medical questions. The IMIX Programme funded seven individual projects that worked in a common field and contributed to a common demonstrator. The fact that these projects ran largely independently is also apparent from the book because there are few links between its chapters.The architecture of the demonstrator is presented in the second article of the book, „The IMIX Demonstrator” (Dennis Hofs, Boris van Schooten, and Rieks op den Akker). The demonstrator showed users a fully functional system and allowed them to ask questions using text, speech, and gestures. The answers produced by the system were presented in the form of text, speech, or images, and could be used in follow-up questions. The article features a detailed description of the architecture, as well as several screenshots and diagrams; these can be useful to researchers who want to find out more about the demonstrator. In addition to the technical details, there is an interesting discussion about the role of demonstrators in large projects and problems that need to be addressed when building them. I think this brief discussion could be very useful for anyone involved in a medium or large project that includes several research groups and needs to build demonstrators.The second part of the book focuses on dialog managers and covers most of the interaction discussed in this book. First, the Vidiam (DIAlogue Management and the VIsual channel) project is described in the article entitled „Corpus-Based Development of a Dialogue Manager for Multimodal Question Answering” (Van Schooten and Op den Akker). In addition to the corpora built in the project and the dialog manager developed on the basis of these corpora, the article also contains a very good discussion about how it is possible to integrate a dialog manager with a QA engine as a way of developing an interactive QA system. I am not aware of any other articles that contain all the information presented here in one place and in such detail. The second article in this part, „Multidimensional Dialogue Management” (Simon Keizer, Harry Bunt, and Volha Petukhova), is more theoretical and presents a dialog manager built using the framework of Dynamic Interpretation Theory (Bunt 2000) which is able to both interpret and generate utterances using dialog acts. The article also presents briefly the way in which this dialog manager was integrated in the IMIX demonstrator.In my opinion, the editors of the book could have chosen a better title for the third part of the book: „Fusing Text, Speech, and Images.” Both articles in this part present work done in the IMOGEN (Interactive Multimodal Output GENeration) project,1 one of the subprojects embedded in the IMIX Programme that focused on producing multimodal presentations that combine text, speech, and graphics. Only the first article focuses on the multimodal aspect of the project, however. The other one discusses only text processing. The article „Experiments in Multimodal Information Presentation” (Charlotte van Hooijdonk, Wauter Bosma, Emiel Krahmer, Alfons Maes, and Mariët Theune) presents three experiments for finding the appropriate way of combining text and images when answering questions from the medical domain. In one of these experiments, the multimodal answers are produced automatically. The other article, „Text-to-Text Generation for Question Answering” (Bosma, Erwin Marsi, Krahmer, and Theune), discusses sentence fusion and could fit very well in a book dedicated to text summarization, as the method presented there is tested not only on data specific to IMIX, but also on the DUC 2005 data.2The fourth and the largest part of the book is „Text Analysis for Question Answering.” It contains five articles, none of which describe a full QA system. Instead, as the title suggests, they focus on various ways of processing texts that can help with answering questions. One common feature of these articles is that they describe methods to extract entities or relations between entities from texts. Most of the articles also briefly discuss how this information is used in QA systems.Most of the methods described in the fourth part of the book are now widely used in computational linguistics, but when they were proposed a few years ago many of them were rather innovative. For brevity, I give only a succinct indication of the methods presented in the articles. „Automatic Extraction of Medical Term Variants from Multilingual pllel Translations” (Lonneke van der Plas, Jörg Tiedemann, and Ismail Fahmi) describes how to acquire medical terms and their variants from parallel corpora. „Relation Extraction for Open and Closed Domain Question Answering” (Bouma, Fahmi, and Jori Mur) shows how it is possible to extract relations between entities using dependency paths in a large collection of newspaper articles and in a much smaller and closed domain corpus of medical documents. A sequence labeling method for entity recognition is presented in the article „Constraint-Satisfaction Inference for Entity Recognition” (Sander Canisius, Antal van den Bosch, and Walter Daelemans). Large newspaper corpora and the Web are used in „Extraction of Hypernymy Information from Texts” (Erik Tjong Kim Sang, Katja Hofmann, and Maarten de Rijke) to determine hypernymy relations between entities. The last article in the fourth part, „Towards a Discourse-Driven Taxonomic Inference Model” (Piroska Lendvai) looks at how the structure of discourse can be used for knowledge discovery from encyclopedic texts. All the articles are well written and could be very interesting for researchers working on information extraction.The book finishes with an epilogue written by three members of the international review panel of IMIX (Eduard Hovy, Jon Oberlander, and Norbert Reithinger) who give a very good overview of the project, providing information that is not covered in any other article of the book. For example, it expands on the multimodal research carried out in the program and presents some details from the point of view of project management. An objective evaluation of the overall program is also included.The book is interesting and I enjoyed reading it. I have to point out, however, that the research presented here is rather old. The IMIX Programme effectively ended in 2008, so it can be argued that most of the articles refer to work that is more than 5 years old. The authors of the epilogue praise the researchers involved in IMIX for the large number of publications they produced. This means that most of the information presented in the book was already published in one form or another somewhere else. Despite this, the book compiles in one place information about the IMIX Program which otherwise could take a while to collect.When I started reading the book, I expected to find more about interactive multi-modal question answering. Each of these topics is presented individually, but with the exception of the article about the IMIX demonstrator, they are not discussed as a whole. I was particularly disappointed by how little space was dedicated to multimodal processing.The articles in the book are written by different groups of authors and they are more or less stand-alone. To achieve this, they all present brief background information about the IMIX Programme. Despite the extra space required for it and the overlap between the information presented in the articles, this is not necessarily bad because it means that researchers who do not have the time to read the whole book can focus on only the articles that are most relevant for them.The potential readers of this book are likely to be researchers interested in the processing of Dutch texts. Researchers in question answering, dialog processing, and information extraction would also benefit from the book. Constantin Orasan |
Comput. Linguistics | 1 |
| 2011 | The QALL-ME Framework: A specifiable-domain multilingual Question Answering architecture
Óscar Ferrández, Christian Spurk, Milen Kouylekov, Iustin Dornescu, Sergio Ferrández, Matteo Negri, Rubén Izquierdo, David Tomás 0001, Constantin Orasan, Günter Neumann, Bernardo Magnini, José Luis Vicedo González |
J. Web Semant. | 9 |
| 2008 | The QALL-ME Benchmark: a Multilingual Resource of Annotated Spoken Requests for Question Answering
Elena Cabrio, Milen Kouylekov, Bernardo Magnini, Matteo Negri, Laura Hasler, Constantin Orasan, David Tomás 0001, José Luis Vicedo González, Günter Neumann, Corinna Weber |
LREC | 6 |
| 2008 | Evaluation of a Cross-lingual Romanian-English Multi-document Summariser
Constantin Orasan, Oana Andreea Chiorean |
LREC | 1 |
| 2008 | Anaphora Resolution Exercise: an Overview
Constantin Orasan, Dan Cristea, Ruslan Mitkov, António Branco |
LREC | 1 |
| 2008 | Development and Alignment of a Domain-Specific Ontology for Question Answering
Shiyan Ou, Viktor Pekar 0001, Constantin Orasan, Christian Spurk, Matteo Negri |
LREC | 3 |
| 2007 | NP Animacy Identification for Anaphora ResolutionabstractIn anaphora resolution for English, animacy identification can play an integral role in the application of agreement restrictions between pronouns and candidates, and as a result, can improve the accuracy of anaphora resolution systems. In this paper, two methods for animacy identification are proposed and evaluated using intrinsic and extrinsic measures. The first method is a rule-based one which uses information about the unique beginners in WordNet to classify NPs on the basis of their animacy. The second method relies on a machine learning algorithm which exploits a WordNet enriched with animacy information for each sense. The effect of word sense disambiguation on the two methods is also assessed. The intrinsic evaluation reveals that the machine learning method reaches human levels of performance. The extrinsic evaluation demonstrates that animacy identification can be beneficial in anaphora resolution, especially in the cases where animate entities are identified with high precision. Constantin Orasan, Richard Evans 0002 |
J. Artif. Intell. Res. | 1 |
| 2006 | NPs for Events: Experiments in Coreference Annotation
Laura Hasler, Constantin Orasan, Karin Naumann |
LREC | 2 |
| 2006 | Computer-aided summarisation - what the user really wants
Constantin Orasan, Laura Hasler |
LREC | 1 |
| 2006 | Transferring Coreference Chains through Word Alignment
Oana Postolache, Dan Cristea, Constantin Orasan |
LREC | 3 |
| 2005 | Automatic Annotation of Corpora for Text Summarisation: A Comparative Study
Constantin Orasan |
CICLing | 1 |
| 2005 | Building a WSD module within an MT system to enable interactive resolution in the user's source language
Constantin Orasan, Ted Marshall, Robert Clark, Le An Ha, Ruslan Mitkov |
EAMT | 1 |
| 2004 | A Comparison of Summarisation Methods Based on Term Specificity Estimation
Constantin Orasan, Viktor Pekar 0001, Laura Hasler |
LREC | 1 |
| 2004 | Annotation of Anaphoric Expressions in an Aligned Bilingual Corpus
Agnès Tutin, Meriam Haddara, Ruslan Mitkov, Constantin Orasan |
LREC | 4 |
| 2003 | CAST: A computer-aided summarisation tool
Constantin Orasan, Ruslan Mitkov, Laura Hasler |
EACL | 1 |
| 2003 | How to build a QA system in your back-garden: application for Romanian
Constantin Orasan, Doina Tatar, Gabriela Serban Czibula, Dana Lupsa, Adrian Onet |
EACL | 1 |
| 2002 | A New, Fully Automatic Version of Mitkov's Knowledge-Poor Pronoun Resolution Method
Ruslan Mitkov, Richard Evans 0002, Constantin Orasan |
CICLing | 3 |
| 2002 | Bilingual alignment of anaphoric expressions
Rafael Muñoz 0001, Ruslan Mitkov, Manuel Palomar, Jesús Peral Cortés, Richard Evans 0002, Lidia Moreno, Constantin Orasan, Maximiliano Saiz-Noeda, Antonio Ferrández Rodríguez, Catalina Barbu, Patricio Martínez-Barco, Armando Suárez |
LREC | 7 |
| 2002 | Building annotated resources for automatic text summarisation
Constantin Orasan |
LREC | 1 |
| 2002 | Assessing the difficulty of finding people in texts
Constantin Orasan, Richard Evans 0002 |
LREC | 1 |
| 2002 | A corpus-based investigation of junk emails
Constantin Orasan, Ramesh Krishnamurthy |
LREC | 1 |
| 2000 | CLinkA A Coreferential Links Annotator
Constantin Orasan |
LREC | 1 |
| 2000 | An Open Architecture for the Construction and Administration of Corpora
Constantin Orasan, Ramesh Krishnamurthy |
LREC | 1 |