Marek Rei

dblp:136/9233 · DBLP profile ↗
← Back
34ranked-venue papers
9as first author
14since 2021 · last 2026
0000-0002-0911-1087ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 31 · 9 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 AgentCoMa: A Compositional Benchmark Mixing Commonsense and Mathematical Reasoning in Real-World Scenarios
abstract
Lisa Alazraki, Lihu Chen, Ana Brassard, Joe Stacey, Hossein A. Rahmani, Marek Rei. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Lisa Alazraki, Lihu Chen, Ana Brassard, Joe Stacey, Hossein A. Rahmani, Marek Rei
ACL (1)6
2025 DiffuseDef: Improved Robustness to Adversarial Attacks via Iterative Denoising
abstract
Pretrained language models have significantly advanced performance across various natural language processing tasks.However, adversarial attacks continue to pose a critical challenge to systems built using these models, as they can be exploited with carefully crafted adversarial texts.Inspired by the ability of diffusion models to predict and reduce noise in computer vision, we propose a novel and flexible adversarial defense method for language classification tasks, DiffuseDef 1 , which incorporates a diffusion layer as a denoiser between the encoder and the classifier.The diffusion layer is trained on top of the existing classifier, ensuring seamless integration with any model in a plug-and-play manner.During inference, the adversarial hidden state is first combined with sampled noise, then denoised iteratively and finally ensembled to produce a robust text representation.By integrating adversarial training, denoising, and ensembling techniques, we show that DiffuseDef improves over existing adversarial defense methods and achieves stateof-the-art performance against common blackbox and white-box adversarial attacks.
Zhenhao Li 0003, Huichi Zhou, Marek Rei, Lucia Specia
ACL (1)3
2025 No Need for Explanations: LLMs can implicitly learn from mistakes in-context
abstract
Showing incorrect answers to Large Language Models (LLMs) is a popular strategy to improve their performance in reasoning-intensive tasks.It is widely assumed that, in order to be helpful, the incorrect answers must be accompanied by comprehensive rationales, explicitly detailing where the mistakes are and how to correct them.However, in this work we present a counterintuitive finding: we observe that LLMs perform better in math reasoning tasks when these rationales are eliminated from the context and models are left to infer on their own what makes an incorrect answer flawed.This approach also substantially outperforms chainof-thought prompting in our evaluations.These results are consistent across LLMs of different sizes and varying reasoning abilities.To gain an understanding of why LLMs learn from mistakes more effectively without explicit corrective rationales, we perform a thorough analysis, investigating changes in context length and answer diversity between different prompting strategies, and their effect on performance.We also examine evidence of overfitting to the in-context rationales when these are provided, and study the extent to which LLMs are able to autonomously infer high-quality corrective rationales given only incorrect answers as input.We find evidence that, while incorrect answers are more beneficial for LLM learning than additional diverse correct answers, explicit corrective rationales over-constrain the model, thus limiting those benefits.
Lisa Alazraki, Maximilian Mozes, Jon Ander Campos, Yi Chern Tan, Marek Rei, Max Bartolo
EMNLP5
2025 Reverse Engineering Human Preferences with Reinforcement Learning
abstract
The capabilities of Large Language Models (LLMs) are routinely evaluated by other LLMs trained to predict human preferences. This framework—known as *LLM-as-a-judge*—is highly scalable and relatively low cost. However, it is also vulnerable to malicious exploitation, as LLM responses can be tuned to overfit the preferences of the judge. Previous work shows that the answers generated by a candidate-LLM can be edited *post hoc* to maximise the score assigned to them by a judge-LLM. In this study, we adopt a different approach and use the signal provided by judge-LLMs as a reward to adversarially tune models that generate text preambles designed to boost downstream performance. We find that frozen LLMs pipelined with these models attain higher LLM-evaluation scores than existing frameworks. Crucially, unlike other frameworks which intervene directly on the model's response, our method is virtually undetectable. We also demonstrate that the effectiveness of the tuned preamble generator transfers when the candidate-LLM and the judge-LLM are replaced with models that are not used during training. These findings raise important questions about the design of more reliable LLM-as-a-judge evaluation settings. They also demonstrate that human preferences can be reverse engineered effectively, by pipelining LLMs to optimise upstream preambles via reinforcement learning—an approach that could find future applications in diverse tasks and domains beyond adversarial attacks.
Lisa Alazraki, Yi Chern Tan, Jon Ander Campos, Maximilian Mozes, Marek Rei, Max Bartolo
NeurIPS5
2025 The "Question Neighbourhood" Approach for Systematic Evaluation of Code-Generating LLMs
abstract
We present the concept of aquestion neighbourhoodfor systematically evaluating instruction-tuned large language models (LLMs) for code generation via a new benchmark, Turbulence. Turbulence consists of a large set of natural languagequestion templates, each of which is a programming problem, parameterised so that it can be asked in many different forms. Each question template has an associatedtest oraclethat judges whether a code solution returned by an LLM is correct. Thus, from a single question template, it is possible to ask an LLM aneighbourhoodof very similar programming questions, and assess the correctness of the result returned for each question. This allows gaps in an LLM’s code generation abilities to be identified, includinganomalieswhere the LLM correctly solvesmanyquestions in a neighbourhood but fails for particular parameter instantiations. We present experiments against 22 state-of-the-art proprietary and open-source LLMs, each at two temperature configurations. Our evaluation is based on three complementary scores: accuracy score, correctness-potential score, and consistent-correctness score. Our findings show that, across the board, Turbulence is able to reveal cases where LLMs do not behave in a correct and consistent manner, highlighting gaps in their reasoning ability. This goes beyond merely highlighting that LLMs sometimes produce wrong code (which is no surprise): by systematically identifying cases where LLMs are able to solve some problems in a neighbourhood but do not manage to generalise to solve the whole neighbourhood, our method provides detailed insight into the behavioural characteristics of current code-generating LLMs. We present data and examples that shed light on the kinds of mistakes that LLMs make when they return incorrect code results.
Shahin Honarvar, Marek Rei, Alastair F. Donaldson
IEEE Trans. Software Eng.2
2024 Atomic Inference for NLI with Generated Facts as Atoms
abstract
With recent advances, neural models can achieve human-level performance on various natural language tasks.However, there are no guarantees that any explanations from these models are faithful, i.e. that they reflect the inner workings of the model.Atomic inference overcomes this issue, providing interpretable and faithful model decisions.This approach involves making predictions for different components (or atoms) of an instance, before using interpretable and deterministic rules to derive the overall prediction based on the individual atom-level predictions.We investigate the effectiveness of using LLM-generated facts as atoms, decomposing Natural Language Inference premises into lists of facts.While directly using generated facts in atomic inference systems can result in worse performance, with 1) a multi-stage fact generation process, and 2) a training regime that incorporates the facts, our fact-based method outperforms other approaches. 1 Logical Rules for TrainingInstance
Joe Stacey, Pasquale Minervini, Haim Dubossarsky, Oana-Maria Camburu, Marek Rei
EMNLP5
2024 Did the Neurons Read your Book? Document-level Membership Inference for Large Language Models
Matthieu Meeus, Shubham Jain 0006, Marek Rei, Yves-Alexandre de Montjoye
USENIX Security Symposium3
2023 Modelling Temporal Document Sequences for Clinical ICD Coding
abstract
Past studies on the ICD coding problem focus on predicting clinical codes primarily based on the discharge summary.This covers only a small fraction of the notes generated during each hospital stay and leaves potential for improving performance by analysing all the available clinical notes.We propose a hierarchical transformer architecture that uses text across the entire sequence of clinical notes in each hospital stay for ICD coding, and incorporates embeddings for text metadata such as their position, time, and type of note.While using all clinical notes increases the quantity of data substantially, superconvergence can be used to reduce training costs.We evaluate the model on the MIMIC-III dataset.Our model exceeds the prior state-of-the-art when using only discharge summaries as input, and achieves further performance improvements when all clinical notes are used as input.
Clarence Boon Liang Ng, Marek Rei
EACL3
2022 Supervising Model Attention with Human Explanations for Robust Natural Language Inference
abstract
Natural Language Inference (NLI) models are known to learn from biases and artefacts within their training data, impacting how well they generalise to other unseen datasets. Existing de-biasing approaches focus on preventing the models from learning these biases, which can result in restrictive models and lower performance. We instead investigate teaching the model how a human would approach the NLI task, in order to learn features that will generalise better to previously unseen examples. Using natural language explanations, we supervise the model’s attention weights to encourage more attention to be paid to the words present in the explanations, significantly improving model performance. Our experiments show that the in-distribution improvements of this method are also accompanied by out-of-distribution improvements, with the supervised models learning from features that generalise better to other NLI datasets. Analysis of the model indicates that human explanations encourage increased attention on the important words, with more attention paid to words in the premise and less attention paid to punctuation and stopwords.
Joe Stacey, Yonatan Belinkov, Marek Rei
AAAI3
2022 Memorisation versus Generalisation in Pre-trained Language Models
abstract
State-of-the-art pre-trained language models have been shown to memorise facts and perform well with limited amounts of training data.To gain a better understanding of how these models learn, we study their generalisation and memorisation capabilities in noisy and low-resource scenarios.We find that the training of these models is almost unaffected by label noise and that it is possible to reach near-optimal results even on extremely noisy datasets.However, our experiments also show that they mainly learn from high-frequency patterns and largely fail when tested on lowresource tasks such as few-shot learning and rare entity recognition.To mitigate such limitations, we propose an extension based on prototypical networks that improves performance in low-resource named entity recognition tasks.
Michael Tänzer, Sebastian Ruder, Marek Rei
ACL (1)3
2022 Probing for targeted syntactic knowledge through grammatical error detection
abstract
Targeted studies testing knowledge of subjectverb agreement (SVA) indicate that pre-trained language models encode syntactic information.We assert that if models robustly encode subject-verb agreement, they should be able to identify when agreement is correct and when it is incorrect.To that end, we propose grammatical error detection as a diagnostic probe to evaluate token-level contextual representations for their knowledge of SVA.We evaluate contextual representations at each layer from five pre-trained English language models: BERT, XLNET, GPT-2, ROBERTA, and ELEC-TRA.We leverage public annotated training data from both English second language learners and Wikipedia edits, and report results on manually crafted stimuli for subject-verb agreement.We find that masked language models linearly encode information relevant to the detection of SVA errors, while the autoregressive models perform on par with our baseline.However, we also observe a divergence in performance when probes are trained on different training sets, and when they are evaluated on different syntactic constructions, suggesting the information pertaining to SVA error detection is not robustly encoded.
Christopher Bryant 0001, Andrew Caines, Marek Rei, Paula Buttery
CoNLL4
2022 Logical Reasoning with Span-Level Predictions for Interpretable and Robust NLI Models
abstract
Current Natural Language Inference (NLI) models achieve impressive results, sometimes outperforming humans when evaluating on indistribution test sets.However, as these models are known to learn from annotation artefacts and dataset biases, it is unclear to what extent the models are learning the task of NLI instead of learning from shallow heuristics in their training data.We address this issue by introducing a logical reasoning framework for NLI, creating highly transparent model decisions that are based on logical rules.Unlike prior work, we show that improved interpretability can be achieved without decreasing the predictive accuracy.We almost fully retain performance on SNLI, while also identifying the exact hypothesis spans that are responsible for each model prediction.Using the e-SNLI human explanations, we verify that our model makes sensible decisions at a span level, despite not using any span labels during training.We can further improve model performance and span-level decisions by using the e-SNLI explanations during training.Finally, our model is more robust in a reduced data setting.When training with only 1,000 examples, out-of-distribution performance improves on the MNLI matched and mismatched validation sets by 13% and 16% relative to the baseline.Training with fewer observations yields further improvements, both in-distribution and out-ofdistribution.
Joe Stacey, Pasquale Minervini, Haim Dubossarsky, Marek Rei
EMNLP4
2022 Guiding Visual Question Generation
abstract
Nihir Vedd, Zixu Wang, Marek Rei, Yishu Miao, Lucia Specia. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Nihir Vedd, Marek Rei, Yishu Miao, Lucia Specia
NAACL-HLT3
2021 How Metaphors Impact Political Discourse: A Large-Scale Topic-Agnostic Study Using Neural Metaphor Detection
Vinodkumar Prabhakaran, Marek Rei, Ekaterina Shutova
ICWSM2
2020 Multidirectional Associative Optimization of Function-Specific Word Representations
abstract
We present a neural framework for learning associations between interrelated groups of words such as the ones found in Subject-Verb-Object (SVO) structures.Our model induces a joint function-specific word vector space, where vectors of e.g.plausible SVO compositions lie close together.The model retains information about word group membership even in the joint space, and can thereby effectively be applied to a number of tasks reasoning over the SVO structure.We show the robustness and versatility of the proposed framework by reporting state-of-the-art results on the tasks of estimating selectional preference and event similarity.The results indicate that the combinations of representations learned with our task-independent model outperform task-specific architectures from prior work, while reducing the number of parameters by up to 95%.
Daniela Gerz, Ivan Vulic, Marek Rei, Roi Reichart, Anna Korhonen
ACL3
2020 Verbal Multiword Expressions for Identification of Metaphor
abstract
Metaphor is a linguistic device in which a concept is expressed by mentioning another.Identifying metaphorical expressions, therefore, requires a non-compositional understanding of semantics.Multiword Expressions (MWEs), on the other hand, are linguistic phenomena with varying degrees of semantic opacity and their identification poses a challenge to computational models.This work is the first attempt at analysing the interplay of metaphor and MWEs processing through the design of a neural architecture whereby classification of metaphors is enhanced by informing the model of the presence of MWEs.To the best of our knowledge, this is the first "MWE-aware" metaphor identification system paving the way for further experiments on the complex interactions of these phenomena.The results and analyses show that this proposed architecture reach state-of-the-art on two different established metaphor datasets.
Omid Rohanian, Marek Rei, Shiva Taslimipoor, Le An Ha
ACL2
2020 Grammatical error detection in transcriptions of spoken English
abstract
We describe the collection of transcription corrections and grammatical error annotations for the CROWDED Corpus of spoken English monologues on business topics.The corpus recordings were crowdsourced from native speakers of English and learners of English with German as their first language.The new transcriptions and annotations are obtained from different crowdworkers: we analyse the 1108 new crowdworker submissions and propose that they can be used for automatic transcription post-editing and grammatical error correction for speech.To further explore the data we train grammatical error detection models with various configurations including pretrained and contextual word representations as input, additional features and auxiliary objectives, and extra training data from written error-annotated corpora.We find that a model concatenating pre-trained and contextual word representations as input performs best, and that additional information does not lead to further performance gains.
Andrew Caines, Christian Bentz, Kate M. Knill, Marek Rei, Paula Buttery
COLING4
2020 Seeing Both the Forest and the Trees: Multi-head Attention for Joint Classification on Different Compositional Levels
abstract
In natural languages, words are used in association to construct sentences.It is not words in isolation, but the appropriate combination of hierarchical structures that conveys the meaning of the whole sentence.Neural networks can capture expressive language features; however, insights into the link between words and sentences are difficult to acquire automatically.In this work, we design a deep neural network architecture that explicitly wires lower and higher linguistic components; we then evaluate its ability to perform the same task at different hierarchical levels.Settling on broad text classification tasks, we show that our model, MHAL, learns to simultaneously solve them at different levels of granularity by fluidly transferring knowledge between hierarchies.Using a multi-head attention mechanism to tie the representations between single words and full sentences, MHAL systematically outperforms equivalent models that are not incentivized towards developing compositional representations.Moreover, we demonstrate that, with the proposed architecture, the sentence information flows naturally to individual words, allowing the model to behave like a sequence labeller (which is a lower, word-level task) even without any word supervision, in a zero-shot fashion.
Miruna Pislar, Marek Rei
COLING2
2020 Grammatical Error Correction in Low Error Density Domains: A New Benchmark and Analyses
abstract
Evaluation of grammatical error correction (GEC) systems has primarily focused on essays written by non-native learners of English, which however is only part of the full spectrum of GEC applications.We aim to broaden the target domain of GEC and release CWEB, a new benchmark for GEC consisting of website text generated by English speakers of varying levels of proficiency.Website data is a common and important domain that contains far fewer grammatical errors than learner essays, which we show presents a challenge to stateof-the-art GEC systems.We demonstrate that a factor behind this is the inability of systems to rely on a strong internal language model in low error density domains.We hope this work shall facilitate the development of opendomain GEC models that generalize to different topics and genres.
Simon Flachs, Ophélie Lacroix, Helen Yannakoudakis, Marek Rei, Anders Søgaard
EMNLP (1)4
2019 Jointly Learning to Label Sentences and Tokens
abstract
Learning to construct text representations in end-to-end systems can be difficult, as natural languages are highly compositional and task-specific annotated datasets are often limited in size. Methods for directly supervising language composition can allow us to guide the models based on existing knowledge, regularizing them towards more robust and interpretable representations. In this paper, we investigate how objectives at different granularities can be used to learn better language representations and we propose an architecture for jointly learning to label sentences and tokens. The predictions at each level are combined together using an attention mechanism, with token-level labels also acting as explicit supervision for composing sentence-level representations. Our experiments show that by learning to perform these tasks jointly on multiple levels, the model achieves substantial improvements for both sentence classification and sequence labeling.
Marek Rei, Anders Søgaard
AAAI1
2019 Modelling the interplay of metaphor and emotion through multitask learning
abstract
Verna Dankers, Marek Rei, Martha Lewis, Ekaterina Shutova. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Verna Dankers, Marek Rei, Martha Lewis, Ekaterina Shutova
EMNLP/IJCNLP (1)2
2019 Semi-Supervised Bootstrapping of Dialogue State Trackers for Task-Oriented Modelling
abstract
Bo-Hsiang Tseng, Marek Rei, Paweł Budzianowski, Richard Turner, Bill Byrne, Anna Korhonen. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Bo-Hsiang Tseng, Marek Rei, Pawel Budzianowski, Richard E. Turner, William J. Byrne, Anna Korhonen
EMNLP/IJCNLP (1)2
2018 Sequence Classification with Human Attention
abstract
Learning attention functions requires large volumes of data, but many NLP tasks simulate human behavior, and in this paper, we show that human attention really does provide a good inductive bias on many attention functions in NLP.Specifically, we use estimated human attention derived from eyetracking corpora to regularize attention functions in recurrent neural networks.We show substantial improvements across a range of tasks, including sentiment analysis, grammatical error detection, and detection of abusive language.
Maria Barrett, Joachim Bingel, Nora Hollenstein, Marek Rei, Anders Søgaard
CoNLL4
2018 Zero-Shot Sequence Labeling: Transferring Knowledge from Sentences to Tokens
abstract
Can attention-or gradient-based visualization techniques be used to infer token-level labels for binary sequence tagging problems, using networks trained only on sentence-level labels?We construct a neural network architecture based on soft attention, train it as a binary sentence classifier and evaluate against tokenlevel annotation on four different datasets.Inferring token labels from a network provides a method for quantitatively evaluating what the model is learning, along with generating useful feedback in assistance systems.Our results indicate that attention-based methods are able to predict token-level labels more accurately, compared to gradient-based methods, sometimes even rivaling the supervised oracle network.
Marek Rei, Anders Søgaard
NAACL-HLT1
2018 Variable Typing: Assigning Meaning to Variables in Mathematical Text
abstract
Yiannos Stathopoulos, Simon Baker, Marek Rei, Simone Teufel. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Yiannos Stathopoulos, Simon Baker, Marek Rei, Simone Teufel
NAACL-HLT3
2017 Semi-supervised Multitask Learning for Sequence Labeling
abstract
We propose a sequence labeling framework with a secondary training objective, learning to predict surrounding words for every word in the dataset.This language modeling objective incentivises the system to learn general-purpose patterns of semantic and syntactic composition, which are also useful for improving accuracy on different sequence labeling tasks.The architecture was evaluated on a range of datasets, covering the tasks of error detection in learner texts, named entity recognition, chunking and POS-tagging.The novel language modeling objective provided consistent performance improvements on every benchmark, without requiring any additional annotated or unannotated data.
Marek Rei
ACL (1)1
2017 Grasping the Finer Point: A Supervised Similarity Network for Metaphor Detection
abstract
The ubiquity of metaphor in our everyday communication makes it an important problem for natural language understanding.Yet, the majority of metaphor processing systems to date rely on handengineered features and there is still no consensus in the field as to which features are optimal for this task.In this paper, we present the first deep learning architecture designed to capture metaphorical composition.Our results demonstrate that it outperforms the existing approaches in the metaphor identification task.
Marek Rei, Luana Bulat, Douwe Kiela, Ekaterina Shutova
EMNLP1
2017 Neural Sequence-Labelling Models for Grammatical Error Correction
abstract
We propose an approach to N -best list reranking using neural sequence-labelling models.We train a compositional model for error detection that calculates the probability of each token in a sentence being correct or incorrect, utilising the full sentence as context.Using the error detection model, we then re-rank the N best hypotheses generated by statistical machine translation systems.Our approach achieves state-of-the-art results on error correction for three different datasets, and it has the additional advantage of only using a small set of easily computed features that require no linguistic input.
Helen Yannakoudakis, Marek Rei, Øistein E. Andersen, Zheng Yuan 0003
EMNLP2
2016 Automatic Text Scoring Using Neural Networks
abstract
Automated Text Scoring (ATS) provides a cost-effective and consistent alternative to human marking.However, in order to achieve good performance, the predictive features of the system need to be manually engineered by human experts.We introduce a model that forms word representations by learning the extent to which specific words contribute to the text's score.Using Long-Short Term Memory networks to represent the meaning of texts, we demonstrate that a fully automated framework is able to achieve excellent results over similar approaches.In an attempt to make our results more interpretable, and inspired by recent advances in visualizing neural networks, we introduce a novel method for identifying the regions of the text that the model has found more discriminative.
Dimitrios Alikaniotis, Helen Yannakoudakis, Marek Rei
ACL (1)3
2016 Compositional Sequence Labeling Models for Error Detection in Learner Writing
abstract
In this paper, we present the first experiments using neural network models for the task of error detection in learner writing.We perform a systematic comparison of alternative compositional architectures and propose a framework for error detection based on bidirectional LSTMs.Experiments on the CoNLL-14 shared task dataset show the model is able to outperform other participants on detecting errors in learner writing.Finally, the model is integrated with a publicly deployed self-assessment system, leading to performance comparable to human annotators.
Marek Rei, Helen Yannakoudakis
ACL (1)1
2016 Attending to Characters in Neural Sequence Labeling Models
abstract
Sequence labeling architectures use word embeddings for capturing similarity, but suffer when handling previously unseen or rare words. We investigate character-level extensions to such models and propose a novel architecture for combining alternative word representations. By using an attention mechanism, the model is able to dynamically decide how much information to use from a word- or character-level component. We evaluated different architectures on a range of sequence labeling datasets, and character-level extensions were found to improve performance on every benchmark. In addition, the proposed attention-based architecture delivered the best results even with a smaller number of trainable parameters.
Marek Rei, Gamal K. O. Crichton, Sampo Pyysalo
COLING1
2015 Online Representation Learning in Recurrent Neural Language Models
abstract
We investigate an extension of continuous online learning in recurrent neural network language models.The model keeps a separate vector representation of the current unit of text being processed and adaptively adjusts it after each prediction.The initial experiments give promising results, indicating that the method is able to increase language modelling accuracy, while also decreasing the parameters needed to store the model along with the computation required at each step.
Marek Rei
EMNLP1
2014 Looking for Hyponyms in Vector Space
abstract
The task of detecting and generating hyponyms is at the core of semantic understanding of language, and has numerous practical applications. We investigate how neural network embeddings perform on this task, compared to dependency-based vector space models, and evaluate a range of similarity measures on hyponym generation. A new asymmetric similarity measure and a combination approach are described, both of which significantly improve precision. We release three new datasets of lexical vector representations trained on the BNC and our evaluation dataset for hyponym generation.
Marek Rei, Ted Briscoe
CoNLL1
2013 Parser lexicalisation through self-learning
Marek Rei, Ted Briscoe
HLT-NAACL1