Piotr Sawicki 0001

dblp:61/3044-1 · DBLP profile ↗
← Back
6ranked-venue papers
4as first author
6since 2021 · last 2025
0009-0004-0973-4892ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 5 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Can Large Language Models Outperform Non-Experts in Poetry Evaluation? A Comparative Study Using the Consensual Assessment Technique
abstract
This study adapts the Consensual Assessment Technique (CAT) for Large Language Models (LLMs), introducing a novel methodology for poetry evaluation.Using a 90-poem dataset with a ground truth based on publication venue, we demonstrate that this approach allows LLMs to significantly surpass the performance of non-expert human judges.Our method, which leverages forced-choice ranking within small, randomized batches, enabled Claude-3-Opus to achieve a Spearman's Rank Correlation of 0.87 with the ground truth, dramatically outperforming the best human nonexpert evaluation (SRC = 0.38).The LLM assessments also exhibited high inter-rater reliability, underscoring the methodology's robustness.These findings establish that LLMs, when guided by a comparative framework, can be effective and reliable tools for assessing poetry, paving the way for their broader application in other creative domains.
Piotr Sawicki 0001, Marek Grzes, Dan Brown 0001, Fabrício Góes
EMNLP1
2023 Is GPT-4 Good Enough to Evaluate Jokes?
Fabrício Góes, Piotr Sawicki 0001, Marek Grzes, Marco Volpe 0001, Dan Brown 0001
ICCC2
2023 Pushing GPT's Creativity to Its Limits: Alternative Uses and Torrance Tests
Fabrício Góes, Piotr Sawicki 0001, Marek Grzes, Marco Volpe 0001, Jacob Watson
ICCC2
2023 Bits of Grass: Does GPT already know how to write like Whitman?
Piotr Sawicki 0001, Marek Grzes, Fabrício Góes, Dan Brown 0001, Max Peeperkorn, Aisha Khatun
ICCC1
2023 On the power of special-purpose GPT models to create and evaluate new poetry in old styles
Piotr Sawicki 0001, Marek Grzes, Fabrício Góes, Anna Jordanous, Dan Brown 0001, Simona Paraskevopoulou, Max Peeperkorn, Aisha Khatun
ICCC1
2022 Training GPT-2 to represent two Romantic-era authors: challenges, evaluations and pitfalls
Piotr Sawicki 0001, Marek Grzes, Anna Jordanous, Dan Brown 0001, Max Peeperkorn
ICCC1