Katharina Kann

dblp:182/1923 · DBLP profile ↗
← Back
29ranked-venue papers
11as first author
14since 2021 · last 2024
0000-0002-9342-9927ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 26 · 11 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 first-authorHuman-computer interaction and ubiquitous computing · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2024 Prompting as Panacea? A Case Study of In-Context Learning Performance for Qualitative Coding of Classroom Dialog
Ananya Ganesh, Chelsea Chandler, Sidney K. D'Mello, Martha Palmer, Katharina Kann
EDM5
2023 Ethical Considerations for Machine Translation of Indigenous Languages: Giving a Voice to the Speakers
abstract
In recent years machine translation has become very successful for high-resource language pairs.This has also sparked new interest in research on the automatic translation of lowresource languages, including Indigenous languages.However, the latter are deeply related to the ethnic and cultural groups that speak (or used to speak) them.The data collection, modeling and deploying machine translation systems thus result in new ethical questions that must be addressed.Motivated by this, we first survey the existing literature on ethical considerations for the documentation, translation, and general natural language processing for Indigenous languages.Afterward, we conduct and analyze an interview study to shed light on the positions of community leaders, teachers, and language activists regarding ethical concerns for the automatic translation of their languages.Our results show that the inclusion, at different degrees, of native speakers and community members is vital to performing better and more ethical research on Indigenous languages.
Manuel Mager, Elisabeth Mager, Katharina Kann, Ngoc Thang Vu
ACL (1)3
2023 Navigating Wanderland: Highlighting Off-Task Discussions in Classrooms
Ananya Ganesh, Michael Alan Chang, Rachel Dickler, Michael Regan, Jon Z. Cai, Kristin Wright-Bettner, James Pustejovsky, James H. Martin, Jeffrey Flanigan, Martha Palmer, Katharina Kann
AIED11
2023 Meeting the Needs of Low-Resource Languages: The Value of Automatic Alignments via Pretrained Models
abstract
Abteen Ebrahimi, Arya D. McCarthy, Arturo Oncevay, John E. Ortega, Luis Chiruzzo, Gustavo Giménez-Lugo, Rolando Coto-Solano, Katharina Kann. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023.
Abteen Ebrahimi, Arya McCarthy, Arturo Oncevay, John E. Ortega, Luis Chiruzzo, Gustavo Giménez Lugo, Rolando Coto-Solano, Katharina Kann
EACL8
2023 Emerging Challenges in Personalized Medicine: Assessing Demographic Effects on Biomedical Question Answering Systems
abstract
Sagi Shaier, Kevin Bennett, Lawrence Hunter, Katharina Kann. Proceedings of the 13th International Joint Conference on Natural Language Processing and the 3rd Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Sagi Shaier, Kevin Bennett, Lawrence Hunter, Katharina Kann
IJCNLP (1)4
2023 A Comparative Analysis of Automatic Speech Recognition Errors in Small Group Classroom Discourse
abstract
In collaborative learning environments, effective intelligent learning systems need to accurately analyze and understand the collaborative discourse between learners (i.e., group modeling) to provide adaptive support. We investigate how automatic speech recognition (ASR) errors influence discourse models of small group collaboration in noisy real-world classrooms. Our dataset consisted of 30 students recorded by consumer off-the-shelf microphones (Yeti Blue) while engaging in dyadic- and triadic- collaborative learning in a multi-day STEM curriculum unit. We found that two state-of-the-art ASR systems (Google Speech and OpenAI Whisper) yielded very high word error rates (0.822, 0.847) but very different profiles of error with Google being more conservative, rejecting 38% of utterances instead of 12% for Whisper. Next, we examined how these ASR errors influenced down-stream small group modeling based on pre-trained large language models for three tasks: Abstract Meaning Representation parsing (AMRParsing), on-task/off-task detection (OnTask), and Accountable Productive Talk prediction (TalkMove). As expected, models trained on clean human transcripts yielded degraded performance on all three tasks, measured by the transfer ratio (TR). However, the TR of the specific sentence-level AMRParsing task (.39 - .62) was much lower than that of the abstract discourse-level OnTask (.63- .94) and TalkMove tasks (.64-.72). Furthermore, different training strategies that incorporated ASR transcripts alone or as augmentations of human transcripts increased accuracy for the discourse-level tasks (OnTask and TalkMove) but not AMRParsing. Simulation experiments suggested that the models were tolerant of missing utterances in the dialog context, and that jointly improving ASR accuracy on important word classes (e.g., verbs and nouns) can improve performance across all tasks. Overall, our results provide insights into how different types of NLP-based tasks might be tolerant of ASR errors under extremely noisy conditions and provide suggestions for how to improve accuracy in small group modeling settings for a more equitable, engaging, and adaptive collaborative learning environment.
Jie Cao 0010, Ananya Ganesh, Jon Z. Cai, Rosy Southwell, Margaret Perkoff, Michael Regan, Katharina Kann, James H. Martin, Martha Palmer, Sidney K. D'Mello
UMAP7
2022 AmericasNLI: Evaluating Zero-shot Natural Language Understanding of Pretrained Multilingual Models in Truly Low-resource Languages
abstract
Abteen Ebrahimi, Manuel Mager, Arturo Oncevay, Vishrav Chaudhary, Luis Chiruzzo, Angela Fan, John Ortega, Ricardo Ramos, Annette Rios, Ivan Vladimir Meza Ruiz, Gustavo Giménez-Lugo, Elisabeth Mager, Graham Neubig, Alexis Palmer, Rolando Coto-Solano, Thang Vu, Katharina Kann. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022.
Abteen Ebrahimi, Manuel Mager, Arturo Oncevay, Vishrav Chaudhary, Luis Chiruzzo, Angela Fan, John E. Ortega, Ricardo Ramos, Annette Rios, Iván V. Meza, Gustavo Giménez Lugo, Elisabeth Mager, Graham Neubig, Alexis Palmer, Rolando Coto-Solano, Ngoc Thang Vu, Katharina Kann
ACL (1)17
2022 Match the Script, Adapt if Multilingual: Analyzing the Effect of Multilingual Pretraining on Cross-lingual Transferability
abstract
Pretrained multilingual models enable zeroshot learning even for unseen languages, and that performance can be further improved via adaptation prior to finetuning.However, it is unclear how the number of pretraining languages influences a model's zero-shot learning for languages unseen during pretraining.To fill this gap, we ask the following research questions: (1) How does the number of pretraining languages influence zero-shot performance on unseen target languages?( 2) Does the answer to that question change with model adaptation?(3) Do the findings for our first question change if the languages used for pretraining are all related?Our experiments on pretraining with related languages indicate that choosing a diverse set of languages is crucial.Without model adaptation, surprisingly, increasing the number of pretraining languages yields better results up to adding related languages, after which performance plateaus.In contrast, with model adaptation via continued pretraining, pretraining on a larger number of languages often gives further improvement, suggesting that model adaptation is crucial to exploit additional pretraining languages.1
Yoshinari Fujinuma, Jordan L. Boyd-Graber, Katharina Kann
ACL (1)3
2022 A Major Obstacle for NLP Research: Let's Talk about Time Allocation!
abstract
The field of natural language processing (NLP) has grown over the last few years: conferences have become larger, we have published an incredible amount of papers, and state-of-the-art research has been implemented in a large variety of customer-facing products.However, this paper argues that we have been less successful than we should have been and reflects on where and how the field fails to tap its full potential.Specifically, we demonstrate that, in recent years, subpar time allocation has been a major obstacle for NLP research.We outline multiple concrete problems together with their negative consequences and, importantly, suggest remedies to improve the status quo.We hope that this paper will be a starting point for discussions around which common practices are -or are not -beneficial for NLP research.
Katharina Kann, Shiran Dudy, Arya McCarthy
EMNLP1
2022 A Comprehensive Comparison of Neural Networks as Cognitive Models of Inflection
abstract
Neural networks have long been at the center of a debate around the cognitive mechanism by which humans process inflectional morphology.This debate has gravitated into NLP by way of the question: Are neural networks a feasible account for human behavior in morphological inflection?We address that question by measuring the correlation between human judgments and neural network probabilities for unknown word inflections.We test a larger range of architectures than previously studied on two important tasks for the cognitive processing debate: English past tense, and German number inflection.We find evidence that the Transformer may be a better account of human behavior than LSTMs on these datasets, and that LSTM features known to increase inflection accuracy do not always result in more human-like behavior.
Adam Wiemerslage, Shiran Dudy, Katharina Kann
EMNLP3
2021 How to Adapt Your Pretrained Multilingual Model to 1600 Languages
abstract
Abteen Ebrahimi, Katharina Kann. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Abteen Ebrahimi, Katharina Kann
ACL/IJCNLP (1)2
2021 Coloring the Black Box: What Synesthesia Tells Us about Character Embeddings
abstract
In contrast to their word-or sentence-level counterparts, character embeddings are still poorly understood.We aim at closing this gap with an in-depth study of English character embeddings.For this, we use resources from research on grapheme-color synesthesia -a neuropsychological phenomenon where letters are associated with colors -, which give us insight into which characters are similar for synesthetes and how characters are organized in color space.Comparing 10 different character embeddings, we ask: How similar are character embeddings to a synesthete's perception of characters?And how similar are character embeddings extracted from different models?We find that LSTMs agree with humans more than transformers.Comparing across tasks, grapheme-to-phoneme conversion results in the most human-like character embeddings.Finally, ELMo embeddings differ from both humans and other models.
Katharina Kann, Mauro M. Monsalve-Mercado
EACL1
2021 CLiMP: A Benchmark for Chinese Language Model Evaluation
abstract
Linguistically informed analyses of language models (LMs) contribute to the understanding and improvement of these models.Here, we introduce the corpus of Chinese linguistic minimal pairs (CLiMP), which can be used to investigate what knowledge Chinese LMs acquire.CLiMP consists of sets of 1,000 minimal pairs (MPs) for 16 syntactic contrasts in Mandarin, covering 9 major Mandarin linguistic phenomena.The MPs are semiautomatically generated, and human agreement with the labels in CLiMP is 95.8%.We evaluate 11 different LMs on CLiMP, covering n-grams, LSTMs, and Chinese BERT.We find that classifier-noun agreement and verb complement selection are the phenomena that models generally perform best at.However, models struggle the most with the bǎ construction, binding, and filler-gap dependencies.Overall, Chinese BERT achieves an 81.8% average accuracy, while the performances of LSTMs and 5-grams are only moderately above chance level.
Beilei Xiang, Changbing Yang, Alex Warstadt, Katharina Kann
EACL5
2021 The World of an Octopus: How Reporting Bias Influences a Language Model's Perception of Color
abstract
Recent work has raised concerns about the inherent limitations of text-only pretraining.In this paper, we first demonstrate that reporting bias, the tendency of people to not state the obvious, is one of the causes of this limitation, and then investigate to what extent multimodal training can mitigate this issue.To accomplish this, we 1) generate the Color Dataset (CoDa), a dataset of human-perceived color distributions for 521 common objects; 2) use CoDa to analyze and compare the color distribution found in text, the distribution captured by language models, and a human's perception of color; and 3) investigate the performance differences between text-only and multimodal models on CoDa.Our results show that the distribution of colors that a language model recovers correlates more strongly with the inaccurate distribution found in text than with the ground-truth, supporting the claim that reporting bias negatively impacts and inherently limits text-only training.We then demonstrate that multimodal models can leverage their visual training to mitigate these effects, providing a promising avenue for future research.* *Email has no accent, but includes the hyphen. 1 In this paper, we use LM to refer to both causal LMs as well as masked LMs.
Cory Paik, Stephane Aroca-Ouellette, Alessandro Roncone, Katharina Kann
EMNLP (1)4
2020 Learning to Learn Morphological Inflection for Resource-Poor Languages
abstract
We propose to cast the task of morphological inflection—mapping a lemma to an indicated inflected form—for resource-poor languages as a meta-learning problem. Treating each language as a separate task, we use data from high-resource source languages to learn a set of model parameters that can serve as a strong initialization point for fine-tuning on a resource-poor target language. Experiments with two model architectures on 29 target languages from 3 families show that our suggested approach outperforms all baselines. In particular, it obtains a 31.7% higher absolute accuracy than a previously proposed cross-lingual transfer model and outperforms the previous state of the art by 1.7% absolute accuracy on average over languages.
Katharina Kann, Samuel R. Bowman, Kyunghyun Cho
AAAI1
2020 Weakly Supervised POS Taggers Perform Poorly on Truly Low-Resource Languages
abstract
Part-of-speech (POS) taggers for low-resource languages which are exclusively based on various forms of weak supervision – e.g., cross-lingual transfer, type-level supervision, or a combination thereof – have been reported to perform almost as well as supervised ones. However, weakly supervised POS taggers are commonly only evaluated on languages that are very different from truly low-resource languages, and the taggers use sources of information, like high-coverage and almost error-free dictionaries, which are likely not available for resource-poor languages. We train and evaluate state-of-the-art weakly supervised POS taggers for a typologically diverse set of 15 truly low-resource languages. On these languages, given a realistic amount of resources, even our best model gets only less than half of the words right. Our results highlight the need for new and different approaches to POS tagging for truly low-resource languages.
Katharina Kann, Ophélie Lacroix, Anders Søgaard
AAAI1
2020 Unsupervised Morphological Paradigm Completion
abstract
We propose the task of unsupervised morphological paradigm completion.Given only raw text and a lemma list, the task consists of generating the morphological paradigms, i.e., all inflected forms, of the lemmas.From a natural language processing (NLP) perspective, this is a challenging unsupervised task, and high-performing systems have the potential to improve tools for low-resource languages or to assist linguistic annotators.From a cognitive science perspective, this can shed light on how children acquire morphological knowledge.We further introduce a system for the task, which generates morphological paradigms via the following steps: (i) EDIT TREE retrieval, (ii) additional lemma retrieval, (iii) paradigm size discovery, and (iv) inflection generation.We perform an evaluation on 14 typologically diverse languages.Our system outperforms trivial baselines with ease and, for some languages, even obtains a higher accuracy than minimally supervised systems. 1 Sé vigilante y confirma las otras cosas que están para morir , porque no he hallado tus obras bien acabadas delante de Dios .Acuérdate , pues , de lo que has recibido y oído ; guárdalo y arrepiéntete , pues si no velas vendré sobre ti como ladrón y no sabrás a qué hora vendré sobre ti .El vencedor será vestido de vestiduras blancas , y no borraré su nombre del libro de la vida , y confesaré su nombre delante de mi Padre y delante de sus ángeles .El que tiene oído , oiga lo que el Espíritu dice a las iglesias .
Huiming Jin, Liwei Cai, Yihui Peng, Chen Xia, Arya McCarthy, Katharina Kann
ACL6
2020 Intermediate-Task Transfer Learning with Pretrained Language Models: When and Why Does It Work?
abstract
Yada Pruksachatkun, Jason Phang, Haokun Liu, Phu Mon Htut, Xiaoyi Zhang, Richard Yuanzhe Pang, Clara Vania, Katharina Kann, Samuel R. Bowman. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020.
Yada Pruksachatkun, Jason Phang, Haokun Liu, Phu Mon Htut, Richard Yuanzhe Pang, Clara Vania, Katharina Kann, Samuel R. Bowman
ACL8
2020 Acrostic Poem Generation
abstract
We propose a new task in the area of computational creativity: acrostic poem generation in English.Acrostic poems are poems that contain a hidden message; typically, the first letter of each line spells out a word or short phrase.We define the task as a generation task with multiple constraints: given an input word, 1) the initial letters of each line should spell out the provided word, 2) the poem's semantics should also relate to it, and 3) the poem should conform to a rhyming scheme.We further provide a baseline model for the task, which consists of a conditional neural language model in combination with a neural rhyming model.Since no dedicated datasets for acrostic poem generation exist, we create training data for our task by first training a separate topic prediction model on a small set of topic-annotated poems and then predicting topics for additional poems.Our experiments show that the acrostic poems generated by our baseline are received well by humans and do not lose much quality due to the additional constraints.Last, we confirm that poems generated by our model are indeed closely related to the provided prompts, and that pretraining on Wikipedia can boost performance.1 https://poemhunter.com/poem-topics 2 Extending our method to longer poems is straightforward.
Rajat Agarwal, Katharina Kann
EMNLP (1)2
2020 Tackling the Low-resource Challenge for Canonical Segmentation
abstract
Canonical morphological segmentation consists of dividing words into their standardized morphemes.Here, we are interested in approaches for the task when training data is limited.We compare model performance in a simulated low-resource setting for the highresource languages German, English, and Indonesian to experiments on new datasets for the truly low-resource languages Popoluca and Tepehua.We explore two new models for the task, borrowing from the closely related area of morphological generation: an LSTM pointer-generator and a sequence-to-sequence model with hard monotonic attention trained with imitation learning.We find that, in the low-resource setting, the novel approaches outperform existing ones on all languages by up to 11.4% accuracy.However, while accuracy in emulated low-resource scenarios is over 50% for all languages, for the truly lowresource languages Popoluca and Tepehua, our best model only obtains 37.4% and 28.4% accuracy, respectively.Thus, we conclude that canonical segmentation is still a challenging task for low-resource languages.
Manuel Mager, Özlem Çetinoglu, Katharina Kann
EMNLP (1)3
2020 IGT2P: From Interlinear Glossed Texts to Paradigms
abstract
An intermediate step in the linguistic analysis of an under-documented language is to find and organize inflected forms that are attested in natural speech.From this data, linguists generate unseen inflected word forms in order to test hypotheses about the language's inflectional patterns and to complete inflectional paradigm tables.To get the data linguists spend many hours manually creating interlinear glossed texts (IGTs).We introduce a new task that speeds this process and automatically generates new morphological resources for natural language processing systems: IGTto-paradigms (IGT2P).IGT2P generates entire morphological paradigms from IGT input.We show that existing morphological reinflection models can solve the task with 21% to 64% accuracy, depending on the language.We further find that (i) having a language expert spend only a few hours cleaning the noisy IGT data improves performance by as much as 21 percentage points, and (ii) POS tags, which are generally considered a necessary part of NLP morphological reinflection input, have no effect on the accuracy of the models considered here.
Sarah R. Moeller, Changbing Yang, Katharina Kann, Mans Hulden
EMNLP (1)4
2019 Probing for Semantic Classes: Diagnosing the Meaning Content of Word Embeddings
abstract
Word embeddings typically represent different meanings of a word in a single conflated vector.Empirical analysis of embeddings of ambiguous words is currently limited by the small size of manually annotated resources and by the fact that word senses are treated as unrelated individual concepts.We present a large dataset based on manual Wikipedia annotations and word senses, where word senses from different words are related by semantic classes.This is the basis for novel diagnostic tests for an embedding's content: we probe word embeddings for semantic classes and analyze the embedding space by classifying embeddings into semantic classes.Our main findings are: (i) Information about a sense is generally represented well in a single-vector embedding -if the sense is frequent.(ii) A classifier can accurately predict whether a word is single-sense or multi-sense, based only on its embedding.(iii) Although rare senses are not well represented in single-vector embeddings, this does not have negative impact on an NLP application whose performance depends on frequent senses.
Yadollah Yaghoobzadeh, Katharina Kann, Timothy J. Hazen, Eneko Agirre, Hinrich Schütze
ACL (1)2
2019 Towards Realistic Practices In Low-Resource Natural Language Processing: The Development Set
abstract
Katharina Kann, Kyunghyun Cho, Samuel R. Bowman. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Katharina Kann, Kyunghyun Cho, Samuel R. Bowman
EMNLP/IJCNLP (1)1
2018 Sentence-Level Fluency Evaluation: References Help, But Can Be Spared!
abstract
Motivated by recent findings on the probabilistic modeling of acceptability judgments, we propose syntactic log-odds ratio (SLOR), a normalized language model score, as a metric for referenceless fluency evaluation of natural language generation output at the sentence level.We further introduce WPSLOR, a novel WordPiece-based version, which harnesses a more compact language model.Even though word-overlap metrics like ROUGE are computed with the help of hand-written references, our referenceless methods obtain a significantly higher correlation with human fluency scores on a benchmark dataset of compressed sentences.Finally, we present ROUGE-LM, a reference-based metric which is a natural extension of WPSLOR to the case of available references.We show that ROUGE-LM yields a significantly higher correlation with human judgments than all baseline metrics, including WPSLOR on its own.
Katharina Kann, Sascha Rothe, Katja Filippova
CoNLL1
2018 Neural Transductive Learning and Beyond: Morphological Generation in the Minimal-Resource Setting
abstract
Neural state-of-the-art sequence-to-sequence (seq2seq) models often do not perform well for small training sets.We address paradigm completion, the morphological task of, given a partial paradigm, generating all missing forms.We propose two new methods for the minimalresource setting: (i) Paradigm transduction: Since we assume only few paradigms available for training, neural seq2seq models are able to capture relationships between paradigm cells, but are tied to the idiosyncracies of the training set.Paradigm transduction mitigates this problem by exploiting the input subset of inflected forms at test time.(ii) Source selection with high precision (SHIP): Multi-source models which learn to automatically select one or multiple sources to predict a target inflection do not perform well in the minimal-resource setting.SHIP is an alternative to identify a reliable source if training data is limited.On a 52-language benchmark dataset, we outperform the previous state of the art by up to 9.71% absolute accuracy.
Katharina Kann, Hinrich Schütze
EMNLP1
2018 Fortification of Neural Morphological Segmentation Models for Polysynthetic Minimal-Resource Languages
abstract
Katharina Kann, Jesus Manuel Mager Hois, Ivan Vladimir Meza-Ruiz, Hinrich Schütze. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018.
Katharina Kann, Manuel Mager, Iván V. Meza, Hinrich Schütze
NAACL-HLT1
2017 One-Shot Neural Cross-Lingual Transfer for Paradigm Completion
abstract
We present a novel cross-lingual transfer method for paradigm completion, the task of mapping a lemma to its inflected forms, using a neural encoder-decoder model, the state of the art for the monolingual task.We use labeled data from a high-resource language to increase performance on a lowresource language.In experiments on 21 language pairs from four different language families, we obtain up to 58% higher accuracy than without transfer and show that even zero-shot and one-shot learning are possible.We further find that the degree of language relatedness strongly influences the ability to transfer morphological knowledge.
Katharina Kann, Ryan Cotterell, Hinrich Schütze
ACL (1)1
2017 Neural Multi-Source Morphological Reinflection
abstract
We explore the task of multi-source morphological reinflection, which generalizes the standard, single-source version.The input consists of (i) a target tag and (ii) multiple pairs of source form and source tag for a lemma.The motivation is that it is beneficial to have access to more than one source form since different source forms can provide complementary information, e.g., different stems.We further present a novel extension to the encoder-decoder recurrent neural architecture, consisting of multiple encoders, to better solve the task.We show that our new architecture outperforms single-source reinflection models and publish our dataset for multi-source morphological reinflection to facilitate future research.
Katharina Kann, Ryan Cotterell, Hinrich Schütze
EACL (1)1
2016 Neural Morphological Analysis: Encoding-Decoding Canonical Segments
abstract
Canonical morphological segmentation aims to divide words into a sequence of standardized segments.In this work, we propose a character-based neural encoderdecoder model for this task.Additionally, we extend our model to include morphemelevel and lexical information through a neural reranker.We set the new state of the art for the task improving previous results by up to 21% accuracy.Our experiments cover three languages: English, German and Indonesian.RR ED Joint WFST UB error en .19(.01) .25 (.01) 0.27 (.02) 0.63 (.01) .06(.01) de .20 (.01) .26(.02) 0.41 (.03) 0.74 (.01) .04(.01) id .05(.01) .09(.01) 0.10 (.01) 0.71 (.01) .02(.01) edit en .21(.02) .47(.02) 0.98 (.34) 1.35 (.01) .10(.02) de .29 (.02) .51(.03) 1.01 (.07) 4.24 (.20) .06(.01) id .05(.00) .12(.01) 0.15 (.02) 2.13 (.01) .02(.01) F1
Katharina Kann, Ryan Cotterell, Hinrich Schütze
EMNLP1