VLDB 2026 Research / reviewers in the wild / expert
Katrin Erk
dblp:23/556
· DBLP profile ↗
55ranked-venue papers
16as first author
9since 2021 · last 2026
0000-0002-2888-6961ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 50 · 14 first-author · 8 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 1 since 2021Theory of computation · 3 · 2 first-authorGraphics, computer vision, multimedia, augmented reality and games · 2Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
21 papers |
Information extraction and text analysis · 57% Knowledge representation and reasoning · 12% Language models and text generation · 11% | |
| Databases, data mining, and information retrieval
2 papers |
Information retrieval · 97% Data mining · 3% |
Topics — the 30 heaviest of 49, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval
evidence retrieval |
1.0 | 1 | 2026 | Uncertainty-Aware Web-Conditioned Scientific Fact-Checking · WWW 2026 |
Information retrieval
fact-checking |
1.0 | 1 | 2026 | Uncertainty-Aware Web-Conditioned Scientific Fact-Checking · WWW 2026 |
Information retrieval › fact-checking
scientific fact-checking |
1.0 | 1 | 2026 | Uncertainty-Aware Web-Conditioned Scientific Fact-Checking · WWW 2026 |
Machine learning › Trustworthy machine learning
interpretability |
0.7 | 1 | 2023 | A Method for Studying Semantic Construal in Grammatical Constructions with Interpretable Contextual Embedding Spaces · ACL (1) 2023 |
Natural language and speech › Information extraction and text analysis › event analysis
event understanding |
0.6 | 1 | 2022 | POQue: Asking Participant-specific Outcome Questions for a Deeper Understanding of Complex Events · EMNLP 2022 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › pragmatics
pragmatic language understanding |
0.4 | 1 | 2020 | Help! Need Advice on Identifying Advice · EMNLP (1) 2020 |
Natural language and speech › Language models and text generation
pre-trained language model |
0.4 | 1 | 2020 | Attending to Entities for Better Text Understanding · AAAI 2020 |
Natural language and speech › Information extraction and text analysis
lexical semantics |
0.4 | 5 | 2016 | Relations such as Hypernymy: Identifying and Exploiting Hearst Patterns in Distributional Vectors for Lexical Entailment · EMNLP 2016 A Structured Vector Space Model for Word Meaning in Context · EMNLP 2008 Graded Word Sense Assignment · EMNLP 2009 |
Natural language and speech › Information extraction and text analysis › discourse analysis
discourse parsing |
0.4 | 1 | 2019 | Evaluating Discourse in Structured Text Representations · ACL (1) 2019 |
Natural language and speech › Information extraction and text analysis › discourse analysis
discourse structure |
0.4 | 1 | 2019 | Evaluating Discourse in Structured Text Representations · ACL (1) 2019 |
Natural language and speech › Information extraction and text analysis
event extraction |
0.4 | 1 | 2019 | Query-focused Scenario Construction · EMNLP/IJCNLP (1) 2019 |
Natural language and speech › Information extraction and text analysis › semantic role labeling
implicit argument prediction |
0.4 | 1 | 2019 | Implicit Argument Prediction as Reading Comprehension · AAAI 2019 |
Natural language and speech › Question answering and dialogue systems
reading comprehension modeling |
0.4 | 1 | 2019 | Implicit Argument Prediction as Reading Comprehension · AAAI 2019 |
Robotics › Autonomous driving
scenario generation |
0.4 | 1 | 2019 | Query-focused Scenario Construction · EMNLP/IJCNLP (1) 2019 |
Natural language and speech › Information extraction and text analysis
semantic role labeling |
0.4 | 1 | 2019 | Implicit Argument Prediction as Reading Comprehension · AAAI 2019 |
Machine learning › Deep learning architectures and training › attention mechanism
structured attention |
0.4 | 1 | 2019 | Evaluating Discourse in Structured Text Representations · ACL (1) 2019 |
Natural language and speech › Information extraction and text analysis
text classification |
0.4 | 1 | 2019 | Evaluating Discourse in Structured Text Representations · ACL (1) 2019 |
Natural language and speech › Information extraction and text analysis
narrative understanding |
0.3 | 1 | 2018 | Picking Apart Story Salads · EMNLP 2018 |
Natural language and speech › Information extraction and text analysis
claim decomposition |
0.3 | 1 | 2026 | Uncertainty-Aware Web-Conditioned Scientific Fact-Checking · WWW 2026 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › semantic representation › semantic relations
hypernymy detection |
0.2 | 1 | 2016 | Relations such as Hypernymy: Identifying and Exploiting Hearst Patterns in Distributional Vectors for Lexical Entailment · EMNLP 2016 |
Natural language and speech › Information extraction and text analysis › textual entailment
lexical entailment |
0.2 | 1 | 2016 | Relations such as Hypernymy: Identifying and Exploiting Hearst Patterns in Distributional Vectors for Lexical Entailment · EMNLP 2016 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
semantic representation |
0.2 | 1 | 2023 | A Method for Studying Semantic Construal in Grammatical Constructions with Interpretable Contextual Embedding Spaces · ACL (1) 2023 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › probabilistic reasoning
probabilistic logic |
0.2 | 1 | 2014 | Probabilistic Soft Logic for Semantic Textual Similarity · ACL (1) 2014 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › probabilistic reasoning › probabilistic logic
probabilistic soft logic |
0.2 | 1 | 2014 | Probabilistic Soft Logic for Semantic Textual Similarity · ACL (1) 2014 |
Natural language and speech › Information extraction and text analysis › text similarity › semantic similarity
semantic textual similarity |
0.2 | 1 | 2014 | Probabilistic Soft Logic for Semantic Textual Similarity · ACL (1) 2014 |
Natural language and speech › Information extraction and text analysis
word sense disambiguation |
0.2 | 2 | 2009 | Graded Word Sense Assignment · EMNLP 2009 Investigations on Word Senses and Word Usages · ACL/IJCNLP 2009 |
Natural language and speech › Language models and text generation
grammar induction |
0.1 | 1 | 2011 | Simple Unsupervised Grammar Induction from Raw Text with Cascaded Finite State Models · ACL 2011 |
Natural language and speech › Language models and text generation › grammar induction
unsupervised grammar induction |
0.1 | 1 | 2011 | Simple Unsupervised Grammar Induction from Raw Text with Cascaded Finite State Models · ACL 2011 |
Natural language and speech › Information extraction and text analysis › relation extraction
event relation extraction |
0.1 | 1 | 2019 | Query-focused Scenario Construction · EMNLP/IJCNLP (1) 2019 |
Natural language and speech › Information extraction and text analysis › sequence labeling
part-of-speech tagging |
0.1 | 1 | 2010 | Crouching Dirichlet, Hidden Markov Model: Unsupervised POS Tagging with Context Local Tag Generation · EMNLP 2010 |
Methods — techniques the papers use, named apart from their topics
uncertainty calibration · 2.0embedding alignment · 2.0abstention · 2.0psycholinguistic feature norms · 0.7contextual embeddings · 0.7crowdsourcing · 0.6annotation · 0.6self-attention · 0.4pre-trained language model · 0.4coreference supervision · 0.4global coherence modeling · 0.3bag-of-words similarity · 0.3underspecification · 0.0constraint language for lambda structures · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Uncertainty-Aware Web-Conditioned Scientific Fact-CheckingabstractScientific fact-checking is vital for assessing claims in specialized domains such as biomedicine and materials science, yet existing systems often hallucinate or apply inconsistent reasoning, especially when verifying technical, compositional claims against an evidence snippet under source and cost/latency constraints. We present a pipeline centered on atomic predicate–argument decomposition and calibrated, uncertainty-gated corroboration: atomic facts are aligned to local snippets via embeddings, verified by a compact evidence-grounded checker, and only facts with uncertain support trigger domain-restricted web search over authoritative sources. The system supports both binary and tri-valued classification where it predicts labels from Supported, Refuted, NEI for three-way tasks. We evaluate under two regimes, Context-Only (no web) and Context+Web (uncertainty-gated web corroboration); when retrieved evidence conflicts with the provided context, we abstain with NEI rather than overriding the context. On multiple benchmarks, our framework surpasses the strongest benchmarks. In our experiments, web corroboration was invoked for only a minority of atomic facts on average, indicating that external evidence is consulted selectively under calibrated uncertainty rather than routinely. Overall, coupling atomic granularity with calibrated, uncertainty-gated corroboration yields more interpretable and context-conditioned verification, making the approach well-suited to high-stakes, single-document settings that demand traceable rationales, predictable cost/latency, and conservative abstention under domain and evidence grounding constraints. Ashwin Vinod, Katrin Erk |
WWW | 2 |
| 2025 | Reinforcement learning produces efficient case-marking systems
Sasha Boguraev, Katrin Erk, Kyle Mahowald, James W. Shearer, Stephen Wechsler |
CogSci | 2 |
| 2024 | To Learn or Not to Learn: Replaced Token Detection for Learning the Meaning of NegationabstractState-of-the-art language models perform well on a variety of language tasks, but they continue to struggle with understanding negation cues in tasks like natural language inference (NLI). Inspired by Hossain et al. (2020), who show under-representation of negation in language model pretraining datasets, we experiment with additional pretraining with negation data for which we introduce two new datasets. We also introduce a new learning strategy for negation building on ELECTRA’s (Clark et al., 2020) replaced token detection objective. We find that continuing to pretrain ELECTRA-Small’s discriminator leads to substantial gains on a variant of RTE (Recognizing Textual Entailment) with additional negation. On SNLI (Stanford NLI) (Bowman et al., 2015), there are no gains due to the extreme under-representation of negation in the data. Finally, on MNLI (Multi-NLI) (Williams et al., 2018), we find that performance on negation cues is primarily stymied by neutral-labeled examples. Gunjan Bhattarai, Katrin Erk |
LREC/COLING | 2 |
| 2024 | Adjusting Interpretable Dimensions in Embedding Space with Human JudgmentsabstractKatrin Erk, Marianna Apidianaki. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Katrin Erk, Marianna Apidianaki |
NAACL-HLT | 1 |
| 2024 | X-PARADE: Cross-Lingual Textual Entailment and Information Divergence across ParagraphsabstractJuan Rodriguez, Katrin Erk, Greg Durrett. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Juan Diego Rodriguez, Katrin Erk, Greg Durrett |
NAACL-HLT | 2 |
| 2023 | A Method for Studying Semantic Construal in Grammatical Constructions with Interpretable Contextual Embedding SpacesabstractWe study semantic construal in grammatical constructions using large language models.First, we project contextual word embeddings into three interpretable semantic spaces, each defined by a different set of psycholinguistic feature norms.We validate these interpretable spaces and then use them to automatically derive semantic characterizations of lexical items in two grammatical constructions: nouns in subject or object position within the same sentence, and the AANN construction (e.g., 'a beautiful three days').We show that a word in subject position is interpreted as more agentive than the very same word in object position, and that the nouns in the AANN construction are interpreted as more measurement-like than when in the canonical alternation.Our method can probe the distributional meaning of syntactic constructions at a templatic level, abstracted away from specific lexemes. Gabriella Chronis, Kyle Mahowald, Katrin Erk |
ACL (1) | 3 |
| 2023 | An analysis of property inference methodsabstractAbstract Property inference involves predicting properties for a word from its distributional representation. We focus on human-generated resources that link words to their properties and on the task of predicting these properties for unseen words. We introduce the use of label propagation, a semi-supervised machine learning approach, for this task and, in the first systematic study of models for this task, find that label propagation achieves state-of-the-art results. For more variety in the kinds of properties tested, we introduce two new property datasets. Alex Rosenfeld, Katrin Erk |
Nat. Lang. Eng. | 2 |
| 2022 | POQue: Asking Participant-specific Outcome Questions for a Deeper Understanding of Complex EventsabstractKnowledge about outcomes is critical for complex event understanding but is hard to acquire.We show that by pre-identifying a participant in a complex event, crowdworkers are able to (1) infer the collective impact of salient events that make up the situation, (2) annotate the volitional engagement of participants in causing the situation, and (3) ground the outcome of the situation in state changes of the participants.By creating a multi-step interface and a careful quality control strategy, we collect a high quality annotated dataset of 8K short newswire narratives and ROCStories with high inter-annotator agreement (0.74-0.96 weighted Fleiss Kappa).Our dataset, POQue (Participant Outcome Questions), enables the exploration and development of models that address multiple aspects of semantic understanding.Experimentally, we show that current language models lag behind human performance in subtle ways through our task formulations that target abstract and specific comprehension of a complex event, its outcome, and a participant's influence over the event culmination. Sai Vallurupalli, Sayontan Ghosh, Katrin Erk, Niranjan Balasubramanian, Francis Ferraro |
EMNLP | 3 |
| 2021 | Did they answer? Subjective acts and intents in conversational discourseabstractElisa Ferracane, Greg Durrett, Junyi Jessy Li, Katrin Erk. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Elisa Ferracane, Greg Durrett, Junyi Jessy Li, Katrin Erk |
NAACL-HLT | 4 |
| 2020 | Attending to Entities for Better Text UnderstandingabstractRecent progress in NLP witnessed the development of large-scale pre-trained language models (GPT, BERT, XLNet, etc.) based on Transformer (Vaswani et al. 2017), and in a range of end tasks, such models have achieved state-of-the-art results, approaching human performance. This clearly demonstrates the power of the stacked self-attention architecture when paired with a sufficient number of layers and a large amount of pre-training data. However, on tasks that require complex and long-distance reasoning where surface-level cues are not enough, there is still a large gap between the pre-trained models and human performance. Strubell et al. (2018) recently showed that it is possible to inject knowledge of syntactic structure into a model through supervised self-attention. We conjecture that a similar injection of semantic knowledge, in particular, coreference information, into an existing model would improve performance on such complex problems. On the LAMBADA (Paperno et al. 2016) task, we show that a model trained from scratch with coreference as auxiliary supervision for self-attention outperforms the largest GPT-2 model, setting the new state-of-the-art, while only containing a tiny fraction of parameters compared to GPT-2. We also conduct a thorough analysis of different variants of model architectures and supervision configurations, suggesting future directions on applying similar techniques to other problems. Pengxiang Cheng 0001, Katrin Erk |
AAAI | 2 |
| 2020 | Leveraging WordNet Paths for Neural Hypernym PredictionabstractWe formulate the problem of hypernym prediction as a sequence generation task, where the sequences are taxonomy paths in WordNet.Our experiments with encoder-decoder models show that training to generate taxonomy paths can improve the performance of direct hypernym prediction.As a simple but powerful model, the hypo2path model achieves state-of-the-art performance, outperforming the best benchmark by 4.11 points in hit-at-one (H@1). Yejin Cho, Juan Diego Rodriguez, Katrin Erk |
COLING | 4 |
| 2020 | When is a bishop not like a rook? When it's like a rabbi! Multi-prototype BERT embeddings for estimating semantic relationshipsabstractThis paper investigates contextual language models, which produce token representations, as a resource for lexical semantics at the word or type level.We construct multi-prototype word embeddings from bert-base-uncased (Devlin et al., 2018).These embeddings retain contextual knowledge that is critical for some type-level tasks, while being less cumbersome and less subject to outlier effects than exemplar models.Similarity and relatedness estimation, both type-level tasks, benefit from this contextual knowledge, indicating the context-sensitivity of these processes.BERT's token level knowledge also allows the testing of a type-level hypothesis about lexical abstractness, demonstrating the relationship between token-level behavior and type-level concreteness ratings.Our findings provide important insight into the interpretability of BERT: layer 7 approximates semantic similarity, while the final layer (11) approximates relatedness. Gabriella Chronis, Katrin Erk |
CoNLL | 2 |
| 2020 | Help! Need Advice on Identifying AdviceabstractHumans use language to accomplish a wide variety of tasks - asking for and giving advice being one of them. In online advice forums, advice is mixed in with non-advice, like emotional support, and is sometimes stated explicitly, sometimes implicitly. Understanding the language of advice would equip systems with a better grasp of language pragmatics; practically, the ability to identify advice would drastically increase the efficiency of advice-seeking online, as well as advice-giving in natural language generation systems. We present a dataset in English from two Reddit advice forums - r/AskParents and r/needadvice - annotated for whether sentences in posts contain advice or not. Our analysis reveals rich linguistic phenomena in advice discourse. We present preliminary models showing that while pre-trained language models are able to capture advice better than rule-based systems, advice identification is challenging, and we identify directions for future research. Comments: To be presented at EMNLP 2020. Venkata Subrahmanyan Govindarajan, Benjamin T. Chen, Rebecca Warholic, Katrin Erk, Junyi Jessy Li |
EMNLP (1) | 4 |
| 2019 | Implicit Argument Prediction as Reading ComprehensionabstractImplicit arguments, which cannot be detected solely through syntactic cues, make it harder to extract predicate-argument tuples. We present a new model for implicit argument prediction that draws on reading comprehension, casting the predicate-argument tuple with the missing argument as a query. We also draw on pointer networks and multi-hop computation. Our model shows good performance on an argument cloze task as well as on a nominal implicit argument prediction task. Pengxiang Cheng 0001, Katrin Erk |
AAAI | 2 |
| 2019 | Evaluating Discourse in Structured Text RepresentationsabstractDiscourse structure is integral to understanding a text and is helpful in many NLP tasks.Learning latent representations of discourse is an attractive alternative to acquiring expensive labeled discourse data.Liu and Lapata (2018) propose a structured attention mechanism for text classification that derives a tree over a text, akin to an RST discourse tree.We examine this model in detail, and evaluate on additional discourse-relevant tasks and datasets, in order to assess whether the structured attention improves performance on the end task and whether it captures a text's discourse structure.We find the learned latent trees have little to no structure and instead focus on lexical cues; even after obtaining more structured trees with proposed model modifications, the trees are still far from capturing discourse structure when compared to discourse dependency trees from an existing discourse parser.Finally, ablation studies show the structured attention provides little benefit, sometimes even hurting performance.1 Elisa Ferracane, Greg Durrett, Junyi Jessy Li, Katrin Erk |
ACL (1) | 4 |
| 2019 | Query-focused Scenario ConstructionabstractThe news coverage of events often contains not one but multiple incompatible accounts of what happened. We develop a query-based system that extracts compatible sets of events (scenarios) from such data, formulated as one-class clustering. Our system incrementally evaluates each event’s compatibility with already selected events, taking order into account. We use synthetic data consisting of article mixtures for scalable training and evaluate our model on a new human-curated dataset of scenarios about real-world news topics. Stronger neural network models and harder synthetic training settings are both important to achieve high performance, and our final scenario construction system substantially outperforms baselines based on prior work. Su Wang 0001, Greg Durrett, Katrin Erk |
EMNLP/IJCNLP (1) | 3 |
| 2018 | Picking Apart Story SaladsabstractDuring natural disasters and conflicts, information about what happened is often confusing, messy, and distributed across many sources.We would like to be able to automatically identify relevant information and assemble it into coherent narratives of what happened.To make this task accessible to neural models, we introduce Story Salads, mixtures of multiple documents that can be generated at scale.By exploiting the Wikipedia hierarchy, we can generate salads that exhibit challenging inference problems.Story salads give rise to a novel, challenging clustering task, where the objective is to group sentences from the same narratives.We demonstrate that simple bag-of-words similarity clustering falls short on this task and that it is necessary to take into account global context and coherence. Su Wang 0001, Eric Holgate, Greg Durrett, Katrin Erk |
EMNLP | 4 |
| 2018 | Implicit Argument Prediction with Event KnowledgeabstractPengxiang Cheng, Katrin Erk. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Pengxiang Cheng 0001, Katrin Erk |
NAACL-HLT | 2 |
| 2018 | Deep Neural Models of Semantic ShiftabstractAlex Rosenfeld, Katrin Erk. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Alex Rosenfeld, Katrin Erk |
NAACL-HLT | 2 |
| 2017 | Distributional Modeling on a Diet: One-shot Word Learning from Text OnlyabstractWe test whether distributional models can do one-shot learning of definitional properties from text only. Using Bayesian models, we find that first learning overarching structure in the known data, regularities in textual contexts and in properties, helps one-shot learning, and that individual context items can be highly informative. Su Wang 0001, Stephen Roller, Katrin Erk |
IJCNLP(1) | 3 |
| 2016 | Relations such as Hypernymy: Identifying and Exploiting Hearst Patterns in Distributional Vectors for Lexical EntailmentabstractWe consider the task of predicting lexical entailment using distributional vectors.We perform a novel qualitative analysis of one existing model which was previously shown to only measure the prototypicality of word pairs.We find that the model strongly learns to identify hypernyms using Hearst patterns, which are well known to be predictive of lexical relations.We present a novel model which exploits this behavior as a method of feature extraction in an iterative procedure similar to Principal Component Analysis.Our model combines the extracted features with the strengths of other proposed models in the literature, and matches or outperforms prior work on multiple data sets. Stephen Roller, Katrin Erk |
EMNLP | 2 |
| 2016 | PIC a Different Word: A Simple Model for Lexical Substitution in ContextabstractThe Lexical Substitution task involves selecting and ranking lexical paraphrases for a target word in a given sentential context.We present PIC, a simple measure for estimating the appropriateness of substitutes in a given context.PIC outperforms another simple, comparable model proposed in recent work, especially when selecting substitutes from the entire vocabulary.Analysis shows that PIC improves over baselines by incorporating frequency biases into predictions. Stephen Roller, Katrin Erk |
HLT-NAACL | 2 |
| 2016 | Representing Meaning with a Combination of Logical and Distributional ModelsabstractNLP tasks differ in the semantic information they require, and at this time no single semantic representation fulfills all requirements. Logic-based representations characterize sentence structure, but do not capture the graded aspect of meaning. Distributional models give graded similarity ratings for words and phrases, but do not capture sentence structure in the same detail as logic-based approaches. It has therefore been argued that the two are complementary. We adopt a hybrid approach that combines logical and distributional semantics using probabilistic logic, specifically Markov Logic Networks. In this article, we focus on the three components of a practical system:11) Logical representation focuses on representing the input problems in probabilistic logic; 2) knowledge base construction creates weighted inference rules by integrating distributional information with other sources; and 3) probabilistic inference involves solving the resulting MLN inference problems efficiently. To evaluate our approach, we use the task of textual entailment, which can utilize the strengths of both logic-based and distributional representations. In particular we focus on the SICK data set, where we achieve state-of-the-art results. We also release a lexical entailment data set of 10,213 rules extracted from the SICK data set, which is a valuable resource for evaluating lexical entailment systems.2 Islam Beltagy, Stephen Roller, Pengxiang Cheng 0001, Katrin Erk, Raymond J. Mooney |
Comput. Linguistics | 4 |
| 2016 | Word Sense Clustering and ClusterabilityabstractWord sense disambiguation and the related field of automated word sense induction traditionally assume that the occurrences of a lemma can be partitioned into senses. But this seems to be a much easier task for some lemmas than others. Our work builds on recent work that proposes describing word meaning in a graded fashion rather than through a strict partition into senses; in this article we argue that not all lemmas may need the more complex graded analysis, depending on their partitionability. Although there is plenty of evidence from previous studies and from the linguistics literature that there is a spectrum of partitionability of word meanings, this is the first attempt to measure the phenomenon and to couple the machine learning literature on clusterability with word usage data used in computational linguistics. We propose to operationalize partitionability as clusterability, a measure of how easy the occurrences of a lemma are to cluster. We test two ways of measuring clusterability: (1) existing measures from the machine learning literature that aim to measure the goodness of optimal k-means clusterings, and (2) the idea that if a lemma is more clusterable, two clusterings based on two different “views” of the same data points will be more congruent. The two views that we use are two different sets of manually constructed lexical substitutes for the target lemma, on the one hand monolingual paraphrases, and on the other hand translations. We apply automatic clustering to the manual annotations. We use manual annotations because we want the representations of the instances that we cluster to be as informative and “clean” as possible. We show that when we control for polysemy, our measures of clusterability tend to correlate with partitionability, in particular some of the type-(1) clusterability measures, and that these measures outperform a baseline that relies on the amount of overlap in a soft clustering. Diana McCarthy, Marianna Apidianaki, Katrin Erk |
Comput. Linguistics | 3 |
| 2014 | Probabilistic Soft Logic for Semantic Textual SimilarityabstractProbabilistic Soft Logic (PSL) is a re-cently developed framework for proba-bilistic logic. We use PSL to combine logical and distributional representations of natural-language meaning, where distri-butional information is represented in the form of weighted inference rules. We ap-ply this framework to the task of Seman-tic Textual Similarity (STS) (i.e. judg-ing the semantic similarity of natural-language sentences), and show that PSL gives improved results compared to a pre-vious approach based on Markov Logic Networks (MLNs) and a purely distribu-tional approach. 1 Islam Beltagy, Katrin Erk, Raymond J. Mooney |
ACL (1) | 2 |
| 2014 | Inclusive yet Selective: Supervised Distributional Hypernymy Detection
Stephen Roller, Katrin Erk, Gemma Boleda |
COLING | 2 |
| 2014 | What Substitutes Tell Us - Analysis of an "All-Words" Lexical Substitution CorpusabstractWe present the first large-scale English "allwords lexical substitution" corpus.The size of the corpus provides a rich resource for investigations into word meaning.We investigate the nature of lexical substitute sets, comparing them to WordNet synsets.We find them to be consistent with, but more fine-grained than, synsets.We also identify significant differences to results for paraphrase ranking in context reported for the SEMEVAL lexical substitution data.This highlights the influence of corpus construction approaches on evaluation results. Gerhard Kremer, Katrin Erk, Sebastian Padó, Stefan Thater |
EACL | 2 |
| 2013 | Measuring Word Meaning in ContextabstractWord sense disambiguation (WSD) is an old and important task in computational linguistics that still remains challenging, to machines as well as to human annotators. Recently there have been several proposals for representing word meaning in context that diverge from the traditional use of a single best sense for each occurrence. They represent word meaning in context through multiple paraphrases, as points in vector space, or as distributions over latent senses. New methods of evaluating and comparing these different representations are needed. In this paper we propose two novel annotation schemes that characterize word meaning in context in a graded fashion. In WSsim annotation, the applicability of each dictionary sense is rated on an ordinal scale. Usim annotation directly rates the similarity of pairs of usages of the same lemma, again on a scale. We find that the novel annotation schemes show good inter-annotator agreement, as well as a strong correlation with traditional single-sense annotation and with annotation of multiple lexical paraphrases. Annotators make use of the whole ordinal scale, and give very fine-grained judgments that “mix and match” senses for each individual usage. We also find that the Usim ratings obey the triangle inequality, justifying models that treat usage similarity as metric. There has recently been much work on grouping senses into coarse-grained groups. We demonstrate that graded WSsim and Usim ratings can be used to analyze existing coarse-grained sense groupings to identify sense groups that may not match intuitions of untrained native speakers. In the course of the comparison, we also show that the WSsim ratings are not subsumed by any static sense grouping. Katrin Erk, Diana McCarthy, Nicholas Gaylord |
Comput. Linguistics | 1 |
| 2013 | An inference-based model of word meaning in context as a paraphrase distributionabstractGraded models of word meaning in context characterize the meaning of individual usages (occurrences) without reference to dictionary senses. We introduce a novel approach that frames the task of computing word meaning in context as a probabilistic inference problem. The model represents the meaning of a word as a probability distribution over potential paraphrases, inferred using an undirected graphical model. Evaluated on paraphrasing tasks, the model achieves state-of-the-art performance. Taesun Moon, Katrin Erk |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2011 | Simple Unsupervised Grammar Induction from Raw Text with Cascaded Finite State Models
Elias Ponvert, Jason Baldridge, Katrin Erk |
ACL | 3 |
| 2010 | Crouching Dirichlet, Hidden Markov Model: Unsupervised POS Tagging with Context Local Tag Generation
Taesun Moon, Katrin Erk, Jason Baldridge |
EMNLP | 2 |
| 2010 | A Flexible, Corpus-Driven Model of Regular and Inverse Selectional PreferencesabstractWe present a vector space–based model for selectional preferences that predicts plausibility scores for argument headwords. It does not require any lexical resources (such as WordNet). It can be trained either on one corpus with syntactic annotation, or on a combination of a small semantically annotated primary corpus and a large, syntactically analyzed generalization corpus. Our model is able to predict inverse selectional preferences, that is, plausibility scores for predicates given argument heads. We evaluate our model on one NLP task (pseudo-disambiguation) and one cognitive task (prediction of human plausibility judgments), gauging the influence of different parameters and comparing our model against other model classes. We obtain consistent benefits from using the disambiguation and semantic role information provided by a semantically tagged primary corpus. As for parameters, we identify settings that yield good performance across a range of experimental conditions. However, frequency remains a major influence of prediction quality, and we also identify more robust parameter settings suitable for applications with many infrequent items. Katrin Erk, Sebastian Padó, Ulrike Padó |
Comput. Linguistics | 1 |
| 2009 | Investigations on Word Senses and Word Usages
Katrin Erk, Diana McCarthy, Nicholas Gaylord |
ACL/IJCNLP | 1 |
| 2009 | Representing words as regions in vector space
Katrin Erk |
CoNLL | 1 |
| 2009 | Graded Word Sense Assignment
Katrin Erk, Diana McCarthy |
EMNLP | 1 |
| 2009 | Unsupervised morphological segmentation and clustering with document boundaries
Taesun Moon, Katrin Erk, Jason Baldridge |
EMNLP | 2 |
| 2008 | A Structured Vector Space Model for Word Meaning in Context
Katrin Erk, Sebastian Padó |
EMNLP | 1 |
| 2007 | A Simple, Similarity-based Model for Selectional Preferences
Katrin Erk |
ACL | 1 |
| 2007 | Flexible, Corpus-Based Modelling of Human Plausibility Judgements
Sebastian Padó, Ulrike Padó, Katrin Erk |
EMNLP-CoNLL | 3 |
| 2007 | Dominance constraints in stratified context unification
Katrin Erk, Joachim Niehren |
Inf. Process. Lett. | 1 |
| 2006 | SALTO - A Versatile Multi-Level Annotation Tool
Aljoscha Burchardt, Katrin Erk, Anette Frank, Andrea Kowalski, Sebastian Padó |
LREC | 2 |
| 2006 | The SALSA Corpus: a German Corpus Resource for Lexical Semantics
Aljoscha Burchardt, Katrin Erk, Anette Frank, Andrea Kowalski, Sebastian Padó |
LREC | 2 |
| 2006 | Shalmaneser - A Toolchain For Shallow Semantic Parsing
Katrin Erk, Sebastian Padó |
LREC | 1 |
| 2006 | Unknown word sense detection as outlier detection
Katrin Erk |
HLT-NAACL | 1 |
| 2004 | Towards an LFG Syntax-Semantics Interface for Frame Semantics Annotation
Anette Frank, Katrin Erk |
CICLing | 2 |
| 2004 | Semantic Role Labelling With Chunk Sequences
Ulrike Padó, Katrin Erk, Sebastian Padó, Detlef Prescher |
CoNLL | 2 |
| 2004 | A Powerful and Versatile XML Format for Representing Role-semantic Annotation
Katrin Erk, Sebastian Padó |
LREC | 1 |
| 2004 | Querying Both Time-aligned and Hierarchical Corpora with NXT Search
Ulrich Heid, Holger Voormann, Jan-Torsten Milde, Ulrike Gut, Katrin Erk, Sebastian Padó |
LREC | 5 |
| 2003 | Towards a Resource for Lexical Semantics: A Large German Corpus with Extensive Semantic AnnotationabstractWe describe the ongoing construction of a large, semantically annotated corpus resource as reliable basis for the large-scale acquisition of word-semantic information, e.g. the construction of domain-independent lexica. The backbone of the annotation are semantic roles in the frame semantics paradigm. We report experiences and evaluate the annotated data from the first project stage. On this basis, we discuss the problems of vagueness and ambiguity in semantic annotation. Katrin Erk, Andrea Kowalski, Sebastian Padó, Manfred Pinkal |
ACL | 1 |
| 2003 | Well-Nested Parallelism Constraints for Ellipsis Resolution
Katrin Erk, Joachim Niehren |
EACL | 1 |
| 2001 | Underspecified Beta ReductionabstractFor ambiguous sentences, traditional semantics construction produces large numbers of higher-order formulas, which must then be -reduced individually.Underspecified versions can produce compact descriptions of all readings, but it is not known how to perform -reduction on these descriptions.We show how to do this using -reduction constraints in the constraint language for -structures (CLLS). Manuel Bodirsky, Katrin Erk, Alexander Koller, Joachim Niehren |
ACL | 2 |
| 2001 | Beta Reduction Constraints
Manuel Bodirsky, Katrin Erk, Alexander Koller, Joachim Niehren |
RTA | 2 |
| 2000 | Parallelism Constraints
Katrin Erk, Joachim Niehren |
RTA | 1 |
| 1999 | Simulating Boolean circuits by finite splicingabstractAs a computational model to be simulated in a DNA computing context, Boolean circuits are especially interesting because of their parallelism. Simulations in concrete biochemical computing settings have been given by Ogihara and Ray (1996) and Amos and Dunne (1997). In this paper, we show how to simulate Boolean circuits by finite splicing systems, an abstract model of enzymatic recombination. We argue that using an abstract model of DNA computation as a basis leads to simulations of greater clarity and generality. In our construction, the running time of the simulating system is proportional to the depth, and the use of material is proportional to the size of the Boolean circuit simulated. However, the rules of the simulating splicing system depend on the size of the Boolean circuit, but not on the connectives used. Katrin Erk |
CEC | 1 |
| 1998 | Defeasible logic graphs: I. Theory
Donald Nute, Katrin Erk |
Decis. Support Syst. | 2 |