VLDB 2026 Research / reviewers in the wild / expert
Katja Markert
dblp:70/2476
· DBLP profile ↗
45ranked-venue papers
6as first author
5since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 42 · 6 first-author · 5 since 2021Databases, data management, data science and information retrieval · 5Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
14 papers |
Language models and text generation · 54% Information extraction and text analysis · 24% Trustworthy machine learning · 16% | |
| Databases, data mining, and information retrieval
3 papers |
Data mining · 58% Information retrieval · 42% |
Topics — the 25 heaviest of 27, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation › retrieval-augmented generation
knowledge conflict resolution |
1.0 | 1 | 2026 | Whose Facts Win? LLM Source Preferences under Knowledge Conflicts · ACL (1) 2026 |
Natural language and speech › Language models and text generation
retrieval-augmented generation |
1.0 | 1 | 2026 | Whose Facts Win? LLM Source Preferences under Knowledge Conflicts · ACL (1) 2026 |
Machine learning › Trustworthy machine learning
robustness |
1.0 | 1 | 2026 | Whose Facts Win? LLM Source Preferences under Knowledge Conflicts · ACL (1) 2026 |
Natural language and speech › Language models and text generation › text summarization
sentence compression |
0.4 | 1 | 2020 | Discrete Optimization for Unsupervised Sentence Summarization with Word-Level Extraction · ACL 2020 |
Natural language and speech › Language models and text generation
text summarization |
0.4 | 1 | 2020 | Discrete Optimization for Unsupervised Sentence Summarization with Word-Level Extraction · ACL 2020 |
Natural language and speech › Language models and text generation › text summarization
unsupervised summarization |
0.4 | 1 | 2020 | Discrete Optimization for Unsupervised Sentence Summarization with Word-Level Extraction · ACL 2020 |
Natural language and speech › Information extraction and text analysis › document analysis › scholarly text analysis
citation analysis |
0.3 | 1 | 2017 | Fine Grained Citation Span for References in Wikipedia · EMNLP 2017 |
Natural language and speech › Information extraction and text analysis › word sense disambiguation
metonymy resolution |
0.3 | 4 | 2012 | Local and Global Context for Supervised and Unsupervised Metonymy Resolution · EMNLP-CoNLL 2012 Syntactic Features and Word Similarity for Supervised Metonymy Resolution · ACL 2003 Understanding metonymies in discourse · Artif. Intell. 2002 |
Information retrieval › text summarization › temporal summarization
timeline summarization |
0.2 | 1 | 2015 | Joint Graphical Models for Date Selection in Timeline Summarization · ACL (1) 2015 |
Data mining › text mining
text classification |
0.2 | 2 | 2017 | Fine-Grained Genre Classification Using Structural Learning Algorithms · ACL 2010 Fine Grained Citation Span for References in Wikipedia · EMNLP 2017 |
Natural language and speech › Information extraction and text analysis › coreference resolution
bridging resolution |
0.2 | 1 | 2014 | A Rule-Based System for Unrestricted Bridging Resolution: Recognizing Bridging Anaphora and Finding Links to Antecedents · EMNLP 2014 |
Computer vision › Segmentation and scene understanding
context modeling |
0.1 | 1 | 2012 | Local and Global Context for Supervised and Unsupervised Metonymy Resolution · EMNLP-CoNLL 2012 |
Natural language and speech › Information extraction and text analysis › discourse analysis
discourse parsing |
0.1 | 1 | 2011 | Modelling Discourse Relations for Arabic · EMNLP 2011 |
Natural language and speech › Information extraction and text analysis › discourse analysis › discourse processing
discourse relations |
0.1 | 1 | 2011 | Modelling Discourse Relations for Arabic · EMNLP 2011 |
Natural language and speech › Information extraction and text analysis › coreference resolution
anaphora resolution |
0.1 | 3 | 2014 | A Rule-Based System for Unrestricted Bridging Resolution: Recognizing Bridging Anaphora and Finding Links to Antecedents · EMNLP 2014 Using the Web in Machine Learning for Other-Anaphora Resolution · EMNLP 2003 On the Interaction of Metonymies and Anaphora · IJCAI (2) 1997 |
Data mining › text mining › text classification
genre classification |
0.1 | 1 | 2010 | Fine-Grained Genre Classification Using Structural Learning Algorithms · ACL 2010 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › semantic representation
semantic relations |
0.1 | 1 | 2009 | A Comparison of Windowless and Window-Based Computational Association Measures as Predictors of Syntagmatic Human Associations · EMNLP 2009 |
Natural language and speech › Information extraction and text analysis › lexical semantics
word association |
0.1 | 1 | 2009 | A Comparison of Windowless and Window-Based Computational Association Measures as Predictors of Syntagmatic Human Associations · EMNLP 2009 |
Natural language and speech › Information extraction and text analysis
word sense disambiguation |
0.1 | 2 | 2003 | Syntactic Features and Word Similarity for Supervised Metonymy Resolution · ACL 2003 Metonymy Resolution as a Classification Task · EMNLP 2002 |
Machine learning › Probabilistic and Bayesian machine learning › structured models
graphical models |
0.1 | 1 | 2015 | Joint Graphical Models for Date Selection in Timeline Summarization · ACL (1) 2015 |
Natural language and speech › Information extraction and text analysis
natural language semantics |
0.0 | 1 | 2002 | Understanding metonymies in discourse · Artif. Intell. 2002 |
Programming languages and type systems
language design |
0.0 | 1 | 1999 | Lean Semantic Interpretation · IJCAI 1999 |
Natural language and speech › Information extraction and text analysis
web information extraction |
0.0 | 1 | 2003 | Using the Web in Machine Learning for Other-Anaphora Resolution · EMNLP 2003 |
Natural language and speech › Information extraction and text analysis
discourse analysis |
0.0 | 1 | 2002 | Understanding metonymies in discourse · Artif. Intell. 2002 |
Natural language and speech › Information extraction and text analysis
named entity processing |
0.0 | 1 | 2002 | Metonymy Resolution as a Classification Task · EMNLP 2002 |
Methods — techniques the papers use, named apart from their topics
synthetic source evaluation · 1.0repetition bias mitigation · 1.0sequence classification · 0.6joint graphical model · 0.4semantic similarity · 0.4language modeling · 0.4discrete optimization · 0.4rule-based system · 0.2collective classification · 0.2cascaded classification · 0.2structural learning algorithms · 0.1lean · 0.0
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Whose Facts Win? LLM Source Preferences under Knowledge ConflictsabstractAs large language models (LLMs) are more frequently used in retrieval-augmented generation pipelines, it is increasingly relevant to study their behavior under knowledge conflicts.Thus far, the role of the source of the retrieved information has gone unexamined.We address this gap with a novel framework to investigate how source preferences affect LLM resolution of inter-context knowledge conflicts in English, motivated by interdisciplinary research on credibility.By using synthetic sources, we study preferences for different types of sources without inheriting the biases of specific real-world sources.With a comprehensive, tightly-controlled evaluation of 13 open-weight LLMs, we find that LLMs prefer institutionallycorroborated information (e.g., government or newspaper sources) over information from people and social media.However, these source preferences can be reversed by simply repeating information from less credible sources.To mitigate repetition effects and maintain consistent preferences, we propose a novel method that reduces repetition bias by up to 79.2%, while also maintaining at least 72.5% of original preferences.We release all data and code to encourage future work on credibility and source preferences in knowledge-intensive NLP. Jakob Schuster, Vagrant Gautam, Katja Markert |
ACL (1) | 3 |
| 2022 | How to Find Strong Summary Coherence Measures? A Toolbox and a Comparative Study for Summary Coherence Measure EvaluationabstractAutomatically evaluating the coherence of summaries is of great significance both to enable cost-efficient summarizer evaluation and as a tool for improving coherence by selecting high-scoring candidate summaries. While many different approaches have been suggested to model summary coherence, they are often evaluated using disparate datasets and metrics. This makes it difficult to understand their relative performance and identify ways forward towards better summary coherence modelling. In this work, we conduct a large-scale investigation of various methods for summary coherence modelling on an even playing field. Additionally, we introduce two novel analysis measures, intra-system correlation and bias matrices, that help identify biases in coherence measures and provide robustness against system-level confounders. While none of the currently available automatic coherence measures are able to assign reliable coherence scores to system summaries across all evaluation metrics, large-scale language models fine-tuned on self-supervised tasks show promising results, as long as fine-tuning takes into account that they need to generalize across different summary lengths. Julius Steen, Katja Markert |
COLING | 2 |
| 2022 | Biographically Relevant Tweets - a New Dataset, Linguistic Analysis and Classification ExperimentsabstractWe present a new dataset comprising tweets for the novel task of detecting biographically relevant utterances. Biographically relevant utterances are all those utterances that reveal some persistent and non-trivial information about the author of a tweet, e.g. habits, (dis)likes, family status, physical appearance, employment information, health issues etc. Unlike previous research we do not restrict biographical relevance to a small fixed set of pre-defined relations. Next to classification experiments employing state-of-the-art classifiers to establish strong baselines for future work, we carry out a linguistic analysis that compares the predictiveness of various high-level features. We also show that the task is different from established tasks, such as aspectual classification or sentiment analysis. Michael Wiegand, Rebecca Wilm, Katja Markert |
COLING | 3 |
| 2022 | The Chinese Causative-Passive Homonymy Disambiguation: an adversarial Dataset for NLI and a Probing TaskabstractThe disambiguation of causative-passive homonymy (CPH) is potentially tricky for machines, as the causative and the passive are not distinguished by the sentences’ syntactic structure. By transforming CPH disambiguation to a challenging natural language inference (NLI) task, we present the first Chinese Adversarial NLI challenge set (CANLI). We show that the pretrained transformer model RoBERTa, fine-tuned on an existing large-scale Chinese NLI benchmark dataset, performs poorly on CANLI. We also employ Word Sense Disambiguation as a probing task to investigate to what extent the CPH feature is captured in the model’s internal representation. We find that the model’s performance on CANLI does not correspond to its internal representation of CPH, which is the crucial linguistic ability central to the CANLI dataset. CANLI is available on Hugging Face Datasets (Lhoest et al., 2021) at https://huggingface.co/datasets/sxu/CANLI Katja Markert |
LREC | 2 |
| 2021 | How to Evaluate a Summarizer: Study Design and Statistical Analysis for Manual Linguistic Quality EvaluationabstractManual evaluation is essential to judge progress on automatic text summarization.However, we conduct a survey on recent summarization system papers that reveals little agreement on how to perform such evaluation studies.We conduct two evaluation experiments on two aspects of summaries' linguistic quality (coherence and repetitiveness) to compare Likert-type and ranking annotations and show that best choice of evaluation method can vary from one aspect to another.In our survey, we also find that study parameters such as the overall number of annotators and distribution of annotators to annotation items are often not fully reported and that subsequent statistical analysis ignores grouping factors arising from one annotator judging multiple summaries.Using our evaluation experiments, we show that the total number of annotators can have a strong impact on study power and that current statistical analysis methods can inflate type I error rates up to eight-fold.In addition, we highlight that for the purpose of system comparison the current practice of eliciting multiple judgements per summary leads to less powerful and reliable annotations given a fixed study budget. Julius Steen, Katja Markert |
EACL | 2 |
| 2020 | Discrete Optimization for Unsupervised Sentence Summarization with Word-Level ExtractionabstractAutomatic sentence summarization produces a shorter version of a sentence, while preserving its most important information.A good summary is characterized by language fluency and high information overlap with the source sentence.We model these two aspects in an unsupervised objective function, consisting of language modeling and semantic similarity metrics.We search for a high-scoring summary by discrete optimization.Our proposed method achieves a new state-of-the art for unsupervised sentence summarization according to ROUGE scores.Additionally, we demonstrate that the commonly reported ROUGE F1 metric is sensitive to summary length.Since this is unwillingly exploited in recent work, we emphasize that future evaluation should explicitly group summarization systems by output length brackets.1 Raphael Schumann, Lili Mou, Olga Vechtomova, Katja Markert |
ACL | 5 |
| 2020 | Context in Informational Bias DetectionabstractInformational bias is bias conveyed through sentences or clauses that provide tangential, speculative or background information that can sway readers' opinions towards entities.By nature, informational bias is context-dependent, but previous work on informational bias detection has not explored the role of context beyond the sentence.In this paper, we explore four kinds of context for informational bias in English news articles: neighboring sentences, the full article, articles on the same event from other news publishers, and articles from the same domain (but potentially different events).We find that integrating event context improves classification performance over a very strong baseline.In addition, we perform the first error analysis of models on this task.We find that the best-performing context-inclusive model outperforms the baseline on longer sentences, and sentences from politically centrist articles. Esther van den Berg, Katja Markert |
COLING | 2 |
| 2020 | An analysis of language models for metaphor recognitionabstractWe conduct a linguistic analysis of recent metaphor recognition systems, all of which are based on language models.We show that their performance, although reaching high F-scores, has considerable gaps from a linguistic perspective.First, they perform substantially worse on unconventional metaphors than on conventional ones.Second, they struggle with handling rarer word types.These two findings together suggest that a large part of the systems' success is due to optimising the disambiguation of conventionalised, metaphoric word senses for specific words instead of modelling general properties of metaphors.As a positive result, the systems show increasing capabilities to recognise metaphoric readings of unseen words if synonyms or morphological variations of these words have been seen before, leading to enhanced generalisation beyond word sense disambiguation. Arthur Neidlein, Philip Wiesenbach, Katja Markert |
COLING | 3 |
| 2020 | Doctor Who? Framing Through Names and Titles in GermanabstractEntity framing is the selection of aspects of an entity to promote a particular viewpoint towards that entity. We investigate entity framing of political figures through the use of names and titles in German online discourse, enhancing current research in entity framing through titling and naming that concentrates on English only. We collect tweets that mention prominent German politicians and annotate them for stance. We find that the formality of naming in these tweets correlates positively with their stance. This confirms sociolinguistic observations that naming and titling can have a status-indicating function and suggests that this function is dominant in German tweets mentioning political figures. We also find that this status-indicating function is much weaker in tweets from users that are politically left-leaning than in tweets by right-leaning users. This is in line with observations from moral psychology that left-leaning and right-leaning users assign different importance to maintaining social hierarchies. Esther van den Berg, Katharina Korfhage, Josef Ruppenhofer, Michael Wiegand, Katja Markert |
LREC | 5 |
| 2020 | Dataset Reproducibility and IR Methods in Timeline SummarizationabstractTimeline summarization (TLS) generates a dated overview of real-world events based on event-specific corpora. The two standard datasets for this task were collected using Google searches for news reports on given events. Not only is this IR method not reproducible at different search times, it also uses components (such as document popularity) that are not always available for any large news corpus. It is unclear how TLS algorithms fare when provided with event corpora collected with varying IR methods. We therefore construct event-specific corpora from a large static background corpus, the newsroom dataset, using differing, relatively simple IR methods based on raw text alone. We show that the choice of IR method plays a crucial role in the performance of various TLS algorithms. A weak TLS algorithm can even match a stronger one by employing a stronger IR method in the data collection phase. Furthermore, the results of TLS systems are often highly sensitive to additional sentence filtering. We consequently advocate for integrating IR into the development of TLS systems and having a common static background corpus for evaluation of TLS systems. Leo Born, Maximilian Bacher, Katja Markert |
LREC | 3 |
| 2018 | Distinguishing affixoid formations from compoundsabstractWe study German affixoids, a type of morpheme in between affixes and free stems. Several properties have been associated with them – increased productivity; a bleached semantics, which is often evaluative and/or intensifying and thus of relevance to sentiment analysis; and the existence of a free morpheme counterpart – but not been validated empirically. In experiments on a new data set that we make available, we put these key assumptions from the morphological literature to the test and show that despite the fact that affixoids generate many low-frequency formations, we can classify these as affixoid or non-affixoid instances with a best F1-score of 74%. Josef Ruppenhofer, Michael Wiegand, Rebecca Wilm, Katja Markert |
COLING | 4 |
| 2018 | A Temporally Sensitive Submodularity Framework for Timeline SummarizationabstractTimeline summarization (TLS) creates an overview of long-running events via dated daily summaries for the most important dates.TLS differs from standard multi-document summarization (MDS) in the importance of date selection, interdependencies between summaries of different dates and by having very short summaries compared to the number of corpus documents.However, we show that MDS optimization models using submodular functions can be adapted to yield wellperforming TLS models by designing objective functions and constraints that model the temporal dimension inherent in TLS.Importantly, these adaptations retain the elegance and advantages of the original MDS models (clear separation of features and inference, performance guarantees and scalability, little need for supervision) that current TLS-specific models lack.An open-source implementation of the framework and all models described in this paper is available online. Sebastian Martschat, Katja Markert |
CoNLL | 2 |
| 2018 | Unrestricted Bridging ResolutionabstractIn contrast to identity anaphors, which indicate coreference between a noun phrase and its antecedent, bridging anaphors link to their antecedent(s) via lexico-semantic, frame, or encyclopedic relations. Bridging resolution involves recognizing bridging anaphors and finding links to antecedents. In contrast to most prior work, we tackle both problems. Our work also follows a more wide-ranging definition of bridging than most previous work and does not impose any restrictions on the type of bridging anaphora or relations between anaphor and antecedent. We create a corpus (ISNotes) annotated for information status (IS), bridging being one of the IS subcategories. The annotations reach high reliability for all categories and marginal reliability for the bridging subcategory. We use a two-stage statistical global inference method for bridging resolution. Given all mentions in a document, the first stage, bridging anaphora recognition, recognizes bridging anaphors as a subtask of learning fine-grained IS. We use a cascading collective classification method where (i) collective classification allows us to investigate relations among several mentions and autocorrelation among IS classes and (ii) cascaded classification allows us to tackle class imbalance, important for minority classes such as bridging. We show that our method outperforms current methods both for IS recognition overall as well as for bridging, specifically. The second stage, bridging antecedent selection, finds the antecedents for all predicted bridging anaphors. We investigate the phenomenon of semantically or syntactically related bridging anaphors that share the same antecedent, a phenomenon we call sibling anaphors. We show that taking sibling anaphors into account in a joint inference model improves antecedent selection performance. In addition, we develop semantic and salience features for antecedent selection and suggest a novel method to build the candidate antecedent list for an anaphor, using the discourse scope of the anaphor. Our model outperforms previous work significantly. Yufang Hou 0001, Katja Markert, Michael Strube 0001 |
Comput. Linguistics | 2 |
| 2017 | Fine Grained Citation Span for References in WikipediaabstractVerifiability is one of the core editing principles in Wikipedia, editors being encouraged to provide citations for the added content.For a Wikipedia article, determining the citation span of a citation, i.e. what content is covered by a citation, is important as it helps decide for which content citations are still missing.We are the first to address the problem of determining the citation span in Wikipedia articles.We approach this problem by classifying which textual fragments in an article are covered by a citation.We propose a sequence classification approach where for a paragraph and a citation, we determine the citation span at a finegrained level.We provide a thorough experimental evaluation and compare our approach against baselines adopted from the scientific domain, where we show improvement for all evaluation metrics. Besnik Fetahu, Katja Markert, Avishek Anand |
EMNLP | 2 |
| 2017 | RussianFlu-DE: A German Corpus for a Historical Epidemic with Temporal Annotation
Tran Van Canh, Katja Markert, Wolfgang Nejdl |
TPDL | 2 |
| 2017 | Headlines Matter: Using Headlines to Predict the Popularity of News Articles on Twitter and Facebook
Alicja Piotrkowicz, Vania Dimitrova, Jahna Otterbacher, Katja Markert |
ICWSM | 4 |
| 2016 | Finding News Citations for WikipediaabstractAn important editing policy in Wikipedia is to provide citations for added statements in Wikipedia pages, where statements can be arbitrary pieces of text, ranging from a sentence to a paragraph. In many cases citations are either outdated or missing altogether. Besnik Fetahu, Katja Markert, Wolfgang Nejdl, Avishek Anand |
CIKM | 2 |
| 2016 | Harvesting Training Images for Fine-Grained Object Categories Using Visual Descriptions
Josiah Wang, Katja Markert, Mark Everingham |
ECIR | 2 |
| 2015 | Joint Graphical Models for Date Selection in Timeline SummarizationabstractGiang Tran, Eelco Herder, Katja Markert. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Giang Binh Tran, Eelco Herder, Katja Markert |
ACL (1) | 3 |
| 2015 | Automated News Suggestions for Populating Wikipedia Entity PagesabstractWikipedia entity pages are a valuable source of information for direct consumption and for knowledge-base construction, update and maintenance. Facts in these entity pages are typically supported by references. Recent studies show that as much as 20% of the references are from online news sources. However, many entity pages are incomplete even if relevant information is already available in existing news articles. Even for the already present references, there is often a delay between the news article publication time and the reference time. In this work, we therefore look at Wikipedia through the lens of news and propose a novel news-article suggestion task to improve news coverage in Wikipedia, and reduce the lag of newsworthy references. Our work finds direct application, as a precursor, to Wikipedia page generation and knowledge-base acceleration tasks that rely on relevant and high quality input sources. Besnik Fetahu, Katja Markert, Avishek Anand |
CIKM | 2 |
| 2014 | A Rule-Based System for Unrestricted Bridging Resolution: Recognizing Bridging Anaphora and Finding Links to AntecedentsabstractBridging resolution plays an important role in establishing (local) entity coherence.This paper proposes a rule-based approach for the challenging task of unrestricted bridging resolution, where bridging anaphors are not limited to definite NPs and semantic relations between anaphors and their antecedents are not restricted to meronymic relations.The system consists of eight rules which target different relations based on linguistic insights.Our rule-based system significantly outperforms a reimplementation of a previous rule-based system (Vieira and Poesio, 2000).Furthermore, it performs better than a learning-based approach which has access to the same knowledge resources as the rule-based system.Additionally, incorporating the rules and more features into the learning-based system yields a minor improvement over the rule-based system. Yufang Hou 0001, Katja Markert, Michael Strube 0001 |
EMNLP | 2 |
| 2014 | Designing and Evaluating a Reliable Corpus of Web Genres via Crowd-Sourcing
Noushin Rezapour Asheghi, Serge Sharoff, Katja Markert |
LREC | 3 |
| 2013 | Cascading Collective Classification for Bridging Anaphora Recognition using a Rich Linguistic Feature SetabstractRecognizing bridging anaphora is difficult due to the wide variation within the phenomenon, the resulting lack of easily identifiable surface markers and their relative rarity.We develop linguistically motivated discourse structure, lexico-semantic and genericity detection features and integrate these into a cascaded minority preference algorithm that models bridging recognition as a subtask of learning finegrained information status (IS).We substantially improve bridging recognition without impairing performance on other IS classes. Yufang Hou 0001, Katja Markert, Michael Strube 0001 |
EMNLP | 2 |
| 2013 | Global Inference for Bridging Anaphora Resolution
Yufang Hou 0001, Katja Markert, Michael Strube 0001 |
HLT-NAACL | 2 |
| 2012 | Collective Classification for Fine-grained Information Status
Katja Markert, Yufang Hou 0001, Michael Strube 0001 |
ACL (1) | 1 |
| 2012 | Local and Global Context for Supervised and Unsupervised Metonymy Resolution
Vivi Nastase, Alex Judea, Katja Markert, Michael Strube 0001 |
EMNLP-CoNLL | 3 |
| 2011 | Modelling Discourse Relations for Arabic
Amal Alsaif, Katja Markert |
EMNLP | 2 |
| 2010 | Fine-Grained Genre Classification Using Structural Learning Algorithms
Zhi-Li Wu, Katja Markert, Serge Sharoff |
ACL | 2 |
| 2010 | The Leeds Arabic Discourse Treebank: Annotating Discourse Connectives for Arabic
Amal Alsaif, Katja Markert |
LREC | 2 |
| 2010 | The Web Library of Babel: evaluating genre collections
Serge Sharoff, Zhi-Li Wu, Katja Markert |
LREC | 3 |
| 2010 | Word Sense Subjectivity for Cross-lingual Lexical Substitution
Fangzhong Su, Katja Markert |
HLT-NAACL | 2 |
| 2009 | Learning Models for Object Recognition from Natural Language DescriptionsabstractWe investigate the task of learning models for visual object recognition from natural language descriptions alone. The approach contributes to the recognition of fine-grain object categories, such as animal and plant species, where it may be difficult to collect many images for training, but where textual descriptions of visual attributes are readily available. As an example we tackle recognition of butterfly species, learning models from descriptions in an online nature guide. We propose natural language processing methods for extracting salient visual attributes from these descriptions to use as ‘templates ’ for the object categories, and apply vision methods to extract corresponding attributes from test images. A generative model is used to connect textual terms in the learnt templates to visual attributes. We report experiments comparing the performance of humans and the proposed method on a dataset of ten butterfly categories. 1 Josiah Wang, Katja Markert, Mark Everingham |
BMVC | 2 |
| 2009 | A Comparison of Windowless and Window-Based Computational Association Measures as Predictors of Syntagmatic Human Associations
Justin Washtell, Katja Markert |
EMNLP | 2 |
| 2009 | Subjectivity Recognition on Word Senses via Semi-supervised Mincuts
Fangzhong Su, Katja Markert |
HLT-NAACL | 2 |
| 2008 | From Words to Senses: A Case Study of Subjectivity Recognition
Fangzhong Su, Katja Markert |
COLING | 2 |
| 2005 | Comparing Knowledge Sources for Nominal Anaphora ResolutionabstractWe compare two ways of obtaining lexical knowledge for antecedent selection in other-anaphora and definite noun phrase coreference. Specifically, we compare an algorithm that relies on links encoded in the manually created lexical hierarchy WordNet and an algorithm that mines corpora by means of shallow lexico-semantic patterns. As corpora we use the British National Corpus (BNC), as well as the Web, which has not been previously used for this task. Our results show that (a) the knowledge encoded in WordNet is often insufficient, especially for anaphor' antecedent relations that exploit subjective or context-dependent knowledge; (b) for other-anaphora, the Web-based method outperforms the WordNet-based method; (c) for definite NP coreference, the Web-based method yields results comparable to those obtained using WordNet over the whole data set and outperforms the WordNet-based method on subsets of the data set; (d) in both case studies, the BNC-based method is worse than the other methods because of data sparseness. Thus, in our studies, the Web-based method alleviated the lexical knowledge gap often encountered in anaphora resolution and handled examples with context-dependent relations between anaphor and antecedent. Because it is inexpensive and needs no hand-modeling of lexical knowledge, it is a promising knowledge source to integrate into anaphora resolution systems. Katja Markert, Malvina Nissim |
Comput. Linguistics | 1 |
| 2003 | Syntactic Features and Word Similarity for Supervised Metonymy ResolutionabstractWe present a supervised machine learning algorithm for metonymy resolution, which exploits the similarity between examples of conventional metonymy. We show that syntactic head-modifier relations are a high precision feature for metonymy recognition but suffer from data sparseness. We partially overcome this problem by integrating a thesaurus and introducing simpler grammatical features, thereby preserving precision and increasing recall. Our algorithm generalises over two levels of contextual similarity. Resulting inferences exceed the complexity of inferences undertaken in word sense disambiguation. We also compare automatic and manual methods for syntactic feature extraction. Malvina Nissim, Katja Markert |
ACL | 2 |
| 2003 | Using the Web in Machine Learning for Other-Anaphora Resolution
Natalia N. Modjeska, Katja Markert, Malvina Nissim |
EMNLP | 2 |
| 2002 | Metonymy Resolution as a Classification TaskabstractWe reformulate metonymy resolution as a classification task. This is motivated by the regularity of metonymic readings and makes general classification and word sense disambiguation methods available for metonymy resolution. We then present a case study for location names, presenting both a corpus of location names annotated for metonymy as well as experiments with a supervised classification algorithm on this corpus. We especially explore the contribution of features used in word sense disambiguation to metonymy resolution. Katja Markert, Malvina Nissim |
EMNLP | 1 |
| 2002 | Towards a Corpus Annotated for Metonymies: the Case of Location Names
Katja Markert, Malvina Nissim |
LREC | 1 |
| 2002 | Understanding metonymies in discourse
Katja Markert, Udo Hahn |
Artif. Intell. | 1 |
| 1999 | Lean Semantic Interpretation
Martin Romacker, Udo Hahn, Katja Markert |
IJCAI | 3 |
| 1997 | On the Interaction of Metonymies and Anaphora
Katja Markert, Udo Hahn |
IJCAI (2) | 1 |
| 1996 | Bridging Textual Ellipses
Udo Hahn, Michael Strube 0001, Katja Markert |
COLING | 3 |
| 1996 | A Conceptual Reasoning Approach to Textual Ellipsis
Udo Hahn, Katja Markert, Michael Strube 0001 |
ECAI | 2 |