EDBT 2026 Demo / reviewers in the wild / expert
Marzena Karpinska
dblp:203/5020
· DBLP profile ↗
13ranked-venue papers
3as first author
13since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 3 first-author · 13 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AI use in American newspapers is widespread, uneven, and rarely disclosedabstractJenna Russell, Marzena Karpinska, Destiny Akinode, James Zhou, Katherine Thai, Bradley Emi, Max Spero, Mohit Iyyer. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Jenna Russell, Marzena Karpinska, Destiny Akinode, James Zhou, Katherine Thai, Bradley Emi, Max Spero, Mohit Iyyer |
ACL (1) | 2 |
| 2025 | CaLMQA: Exploring culturally specific long-form question answering across 23 languagesabstractShane Arora, Marzena Karpinska, Hung-Ting Chen, Ipsita Bhattacharjee, Mohit Iyyer, Eunsol Choi. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Shane Arora, Marzena Karpinska, Hung-Ting Chen, Ipsita Bhattacharjee, Mohit Iyyer, Eunsol Choi |
ACL (1) | 2 |
| 2025 | People who frequently use ChatGPT for writing tasks are accurate and robust detectors of AI-generated textabstractIn this paper, we study how well humans can detect text generated by commercial LLMs (GPT-4o, Claude, o1). We hire annotators to read 300 non-fiction English articles, label them as either human-written or AI-generated, and provide paragraph-length explanations for their decisions. Our experiments show that annotators who frequently use LLMs for writing tasks excel at detecting AI-generated text, even without any specialized training or feedback. In fact, the majority vote among five such “expert” annotators misclassifies only 1 of 300 articles, significantly outperforming most commercial and open-source detectors we evaluated even in the presence of evasion tactics like paraphrasing and humanization. Qualitative analysis of the experts’ free-form explanations shows that while they rely heavily on specific lexical clues (‘AI vocabulary’), they also pick up on more complex phenomena within the text (e.g., formality, originality, clarity) that are challenging to assess for automatic detectors. We release our annotated dataset and code to spur future research into both human and automated detection of AI-generated text. Jenna Russell, Marzena Karpinska, Mohit Iyyer |
ACL (1) | 2 |
| 2025 | An Interdisciplinary Approach to Human-Centered Machine TranslationabstractMarine Carpuat, Omri Asscher, Kalika Bali, Luisa Bentivogli, Fred Blain, Lynne Bowker, Monojit Choudhury, Hal Daumé Iii, Kevin Duh, Ge Gao, Alvin C Grissom II, Marzena Karpinska, Elaine C Khoong, William D. Lewis, Andre Martins, Mary Nurminen, Douglas W. Oard, Maja Popovic, Michel Simard, François Yvon. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Marine Carpuat, Omri Asscher, Kalika Bali, Luisa Bentivogli, Frédéric Blain, Lynne Bowker, Monojit Choudhury, Hal Daumé III, Kevin Duh, Ge Gao 0001, Alvin Grissom II, Marzena Karpinska, Elaine C. Khoong, William D. Lewis, André F. T. Martins, Mary Nurminen, Douglas W. Oard, Maja Popovic, Michel Simard, François Yvon |
EMNLP | 12 |
| 2025 | Does quantization affect models' performance on long-context tasks?abstractLarge language models (LLMs) now support context windows exceeding 128K tokens, but this comes with significant memory requirements and high inference latency.Quantization can mitigate these costs, but may degrade performance.In this work, we present the first systematic evaluation of quantized LLMs on tasks with long inputs (≥64K tokens) and long-form outputs.Our evaluation spans 9.7K test examples, five quantization methods (FP8, GPTQ-int8, AWQ-int4, GPTQ-int4, BNB-nf4), and five models (Llama-3.1 8B and 70B; Qwen-2.5 7B, 32B, and 72B).We find that, on average, 8-bit quantization preserves accuracy (~0.8% drop), whereas 4-bit methods lead to substantial losses, especially for tasks involving longcontext inputs (drops of up to 59%).This degradation tends to worsen when the input is in a language other than English.Crucially, the effects of quantization depend heavily on the quantization method, model, and task.For instance, while Qwen-2.5 72B remains robust under BNB-nf4, Llama-3.1 70B experiences a 32% performance drop on the same task.These findings highlight the importance of a careful, task-specific evaluation before deploying quantized LLMs, particularly in long-context scenarios and for languages other than English. github.com/molereddy/long-context-quantization FP8 GPTQ-int8 AWQ-int4 GPTQ-int4 BNB-nf4Ruler Anmol Mekala, Anirudh Atmakuru, Yixiao Song, Marzena Karpinska, Mohit Iyyer |
EMNLP | 4 |
| 2025 | OWL: Probing Cross-Lingual Recall of Memorized Texts via World LiteratureabstractAlisha Srivastava, Emir Kaan Korukluoglu, Minh Nhat Le, Duyen Tran, Chau Minh Pham, Marzena Karpinska, Mohit Iyyer. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Alisha Srivastava, Emir Korukluoglu, Minh Nhat Le, Duyen Tran, Chau Minh Pham, Marzena Karpinska, Mohit Iyyer |
EMNLP | 6 |
| 2024 | NarrativeTime: Dense Temporal Annotation on a TimelineabstractFor the past decade, temporal annotation has been sparse: only a small portion of event pairs in a text was annotated. We present NarrativeTime, the first timeline-based annotation framework that achieves full coverage of all possible TLINKs. To compare with the previous SOTA in dense temporal annotation, we perform full re-annotation of the classic TimeBankDense corpus (American English), which shows comparable agreement with a signigicant increase in density. We contribute TimeBankNT corpus (with each text fully annotated by two expert annotators), extensive annotation guidelines, open-source tools for annotation and conversion to TimeML format, and baseline results. Anna Rogers, Marzena Karpinska, Vladislav Lialin, Gregory Smelkov, Anna Rumshisky |
LREC/COLING | 2 |
| 2024 | One Thousand and One Pairs: A "novel" challenge for long-context language modelsabstractSynthetic long-context LLM benchmarks (e.g., "needle-in-the-haystack") test only surfacelevel retrieval capabilities, but how well can long-context LLMs retrieve, synthesize, and reason over information across book-length inputs?We address this question by creating NOCHA, a dataset of 1,001 minimally different pairs of true and false claims about 67 recentlypublished English fictional books, written by human readers of those books.In contrast to existing long-context benchmarks, our annotators confirm that the largest share of pairs in NOCHA require global reasoning over the entire book to verify.Our experiments show that while human readers easily perform this task, it is enormously challenging for all ten long-context LLMs that we evaluate: no open-weight model performs above random chance (despite their strong performance on synthetic benchmarks), while GPT-4O achieves the highest accuracy at 55.8%.Further analysis reveals that (1) on average, models perform much better on pairs that require only sentence-level retrieval vs. global reasoning; (2) model-generated explanations for their decisions are often inaccurate even for correctly-labeled claims; and (3) models perform substantially worse on speculative fiction books that contain extensive world-building.The methodology proposed in NOCHA allows for the evolution of the benchmark dataset and the easy analysis of future models.TRUE TRUE Niema takes her students out to world's end and tells them that the fog kills anything it touches, a statement backed by her extensive research into the fog.When Niema takes her students out to world's end and tells them that the fog kills anything it touches, she is intentionally lying to them. Marzena Karpinska, Katherine Thai, Kyle Lo, Tanya Goyal, Mohit Iyyer |
EMNLP | 1 |
| 2023 | Paraphrasing evades detectors of AI-generated text, but retrieval is an effective defenseabstractThe rise in malicious usage of large language models, such as fake content creation and academic plagiarism, has motivated the development of approaches that identify AI-generated text, including those based on watermarking or outlier detection. However, the robustness of these detection algorithms to paraphrases of AI-generated text remains unclear. To stress test these detectors, we build a 11B parameter paraphrase generation model (DIPPER) that can paraphrase paragraphs, condition on surrounding context, and control lexical diversity and content reordering. Paraphrasing text generated by three large language models (including GPT3.5-davinci-003) with DIPPER successfully evades several detectors, including watermarking, GPTZero, DetectGPT, and OpenAI's text classifier. For example, DIPPER drops detection accuracy of DetectGPT from 70.3% to 4.6% (at a constant false positive rate of 1%), without appreciably modifying the input semantics.
To increase the robustness of AI-generated text detection to paraphrase attacks, we introduce a simple defense that relies on retrieving semantically-similar generations and must be maintained by a language model API provider. Given a candidate text, our algorithm searches a database of sequences previously generated by the API, looking for sequences that match the candidate text within a certain threshold. We empirically verify our defense using a database of 15M generations from a fine-tuned T5-XXL model and find that it can detect 80% to 97% of paraphrased generations across different settings while only classifying 1% of human-written sequences as AI-generated. We open-source our models, code and data. Kalpesh Krishna, Yixiao Song, Marzena Karpinska, John Wieting, Mohit Iyyer |
NeurIPS | 3 |
| 2022 | Revisiting Statistical Laws of Semantic Shift in Romance CognatesabstractThis article revisits statistical relationships across Romance cognates between lexical semantic shift and six intra-linguistic variables, such as frequency and polysemy. Cognates are words that are derived from a common etymon, in this case, a Latin ancestor. Despite their shared etymology, some cognate pairs have experienced semantic shift. The degree of semantic shift is quantified using cosine distance between the cognates’ corresponding word embeddings. In the previous literature, frequency and polysemy have been reported to be correlated with semantic shift; however, the understanding of their effects needs revision because of various methodological defects. In the present study, we perform regression analysis under improved experimental conditions, and demonstrate a genuine negative effect of frequency and positive effect of polysemy on semantic shift. Furthermore, we reveal that morphologically complex etyma are more resistant to semantic shift and that the cognates that have been in use over a longer timespan are prone to greater shift in meaning. These findings add to our understanding of the historical process of semantic change. Yoshifumi Kawasaki, Maëlys Salingre, Marzena Karpinska, Hiroya Takamura, Ryo Nagata |
COLING | 3 |
| 2022 | DEMETR: Diagnosing Evaluation Metrics for TranslationabstractWhile machine translation evaluation metrics based on string overlap (e.g., BLEU) have their limitations, their computations are transparent: the BLEU score assigned to a particular candidate translation can be traced back to the presence or absence of certain words.The operations of newer learned metrics (e.g., BLEURT, COMET), which leverage pretrained language models to achieve higher correlations with human quality judgments than BLEU, are opaque in comparison.In this paper, we shed light on the behavior of these learned metrics by creating DEMETR, a diagnostic dataset with 31K English examples (translated from 10 source languages) for evaluating the sensitivity of MT evaluation metrics to 35 different linguistic perturbations spanning semantic, syntactic, and morphological error categories.All perturbations were carefully designed to form minimal pairs with the actual translation (i.e., differ in only one aspect).We find that learned metrics perform substantially better than string-based metrics on DEMETR.Additionally, learned metrics differ in their sensitivity to various phenomena (e.g., BERTSCORE is sensitive to untranslated words but relatively insensitive to gender manipulation, while COMET is much more sensitive to word repetition than to aspectual changes).We publicly release DEMETR to spur more informed future development of machine translation evaluation metrics 1 . Marzena Karpinska, Nishant Raj, Katherine Thai, Yixiao Song, Mohit Iyyer |
EMNLP | 1 |
| 2022 | Exploring Document-Level Literary Machine Translation with Parallel Paragraphs from World LiteratureabstractLiterary translation is a culturally significant task, but it is bottlenecked by the small number of qualified literary translators relative to the many untranslated works published around the world.Machine translation (MT) holds potential to complement the work of human translators by improving both training procedures and their overall efficiency.Literary translation is less constrained than more traditional MT settings since translators must balance meaning equivalence, readability, and critical interpretability in the target language.This property, along with the complex discourse-level context present in literary texts, also makes literary MT more challenging to computationally model and evaluate.To explore this task, we collect a dataset (PAR3) of non-English language novels in the public domain, each aligned at the paragraph level to both human and automatic English translations.Using PAR3, we discover that expert literary translators prefer reference human translations over machinetranslated paragraphs at a rate of 84%, while state-of-the-art automatic MT metrics do not correlate with those preferences.The experts note that MT outputs contain not only mistranslations, but also discourse-disrupting errors and stylistic inconsistencies.To address these problems, we train a post-editing model whose output is preferred over normal MT output at a rate of 69% by experts.We publicly release PAR3 to spur future research into literary MT. 1 Katherine Thai, Marzena Karpinska, Kalpesh Krishna, Bill Ray, Moira Inghilleri, John Wieting, Mohit Iyyer |
EMNLP | 2 |
| 2021 | The Perils of Using Mechanical Turk to Evaluate Open-Ended Text GenerationabstractRecent text generation research has increasingly focused on open-ended domains such as story and poetry generation.Because models built for such tasks are difficult to evaluate automatically, most researchers in the space justify their modeling choices by collecting crowdsourced human judgments of text quality (e.g., Likert scores of coherence or grammaticality) from Amazon Mechanical Turk (AMT).In this paper, we first conduct a survey of 45 open-ended text generation papers and find that the vast majority of them fail to report crucial details about their AMT tasks, hindering reproducibility.We then run a series of story evaluation experiments with both AMT workers and English teachers and discover that even with strict qualification filters, AMT workers (unlike teachers) fail to distinguish between model-generated text and human-generated references.We show that AMT worker judgments improve when they are shown model-generated output alongside human-generated references, which enables the workers to better calibrate their ratings.Finally, interviews with the English teachers provide deeper insights into the challenges of the evaluation process, particularly when rating model-generated text. Marzena Karpinska, Nader Akoury, Mohit Iyyer |
EMNLP (1) | 1 |