EDBT 2026 Demo / reviewers in the wild / expert
Tuhin Chakrabarty
dblp:227/2812
· DBLP profile ↗
25ranked-venue papers
14as first author
19since 2021 · last 2026
0009-0003-9120-4526ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 10 first-author · 15 since 2021Human-computer interaction and ubiquitous computing · 4 · 4 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Can Good Writing Be Generative? Expert-Level AI Writing Emerges through Fine-Tuning on High Quality Books
Tuhin Chakrabarty, Paramveer S. Dhillon |
CHI | 1 |
| 2025 | Can AI writing be salvaged? Mitigating Idiosyncrasies and Improving Human-AI Alignment in the Writing Process through Edits
Tuhin Chakrabarty, Philippe Laban, Chien-Sheng Wu |
CHI | 1 |
| 2025 | How do Humans and Language Models Reason About Creativity? A Comparative Analysis
Antonio Laverghetta, Tuhin Chakrabarty, Tom Hope, Jimmy Pronchick, Krupa Bhawsar, Roger E. Beaty |
CogSci | 2 |
| 2025 | Understanding Figurative Meaning through Explainable Visual EntailmentabstractArkadiy Saakyan, Shreyas Kulkarni, Tuhin Chakrabarty, Smaranda Muresan. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Arkadiy Saakyan, Tuhin Chakrabarty, Smaranda Muresan |
NAACL (Long Papers) | 3 |
| 2024 | Creativity Support in the Age of Large Language Models: An Empirical Study Involving Professional WritersabstractThe development of large language models (LLMs) capable of following instructions and engaging in conversational interactions has led to increased interest in their use across various support tools. We investigate the effectiveness of contemporary LLMs in assisting professional writers via an empirical user study (n=30). The design of our collaborative writing interface is grounded in the cognitive process model of writing [17]. This allows writers to obtain model help in each of the three non-linear cognitive activities in the writing process: planning, translating and reviewing. Participants write short fiction/non-fiction with model help and are subsequently asked to submit a post-completion survey to provide qualitative feedback on the potential and pitfalls of LLMs as writing collaborators. Upon analyzing the writer-LLM interactions, we find that while seeking help across all three types of cognitive activities, writers find LLMs more helpful in translation and reviewing. Our findings from analyzing both the interactions and the survey responses highlight future research directions in creative writing assistance using LLMs. Tuhin Chakrabarty, Vishakh Padmakumar, Faeze Brahman, Smaranda Muresan |
Creativity & Cognition | 1 |
| 2024 | Art or Artifice? Large Language Models and the False Promise of CreativityabstractResearchers have argued that large language models (LLMs) exhibit high-quality writing capabilities from blogs to stories. However, evaluating objectively the creativity of a piece of writing is challenging. Inspired by the Torrance Test of Creative Thinking (TTCT) [64], which measures creativity as a process, we use the Consensual Assessment Technique [3] and propose Torrance Test of Creative Writing (TTCW) to evaluate creativity as product. TTCW consists of 14 binary tests organized into the original dimensions of Fluency, Flexibility, Originality, and Elaboration. We recruit 10 creative writers and implement a human assessment of 48 stories written either by professional authors or LLMs using TTCW. Our analysis shows that LLM-generated stories pass 3-10X less TTCW tests than stories written by professionals. In addition, we explore the use of LLMs as assessors to automate the TTCW evaluation, revealing that none of the LLMs positively correlate with the expert assessments. Tuhin Chakrabarty, Philippe Laban, Divyansh Agarwal, Smaranda Muresan, Chien-Sheng Wu |
CHI | 1 |
| 2024 | Connecting the Dots: Evaluating Abstract Reasoning Capabilities of LLMs Using the New York Times Connections Word GameabstractThe New York Times Connections game has emerged as a popular and challenging pursuit for word puzzle enthusiasts.We collect 438 Connections games to evaluate the performance of state-of-the-art large language models (LLMs) against expert and novice human players.Our results show that even the bestperforming LLM, Claude 3.5 Sonnet, which has otherwise shown impressive reasoning abilities on a wide variety of benchmarks, can only fully solve 18% of the games.Novice and expert players perform better than Claude 3.5 Sonnet, with expert human players significantly outperforming it.We create a taxonomy of the knowledge types required to successfully cluster and categorize words in the Connections game.We find that while LLMs perform relatively well on categorizing words based on semantic relations they struggle with other types of knowledge such as Encyclopedic Knowledge, Multiword Expressions or knowledge that combines both Word Form and Meaning.Our results establish the New York Times Connections game as a challenging benchmark for evaluating abstract reasoning capabilities in AI systems. Prisha Samadarshi, Mariam Mustafa, Anushka Kulkarni, Raven Rothkopf, Tuhin Chakrabarty, Smaranda Muresan |
EMNLP | 5 |
| 2023 | NORMSAGE: Multi-Lingual Multi-Cultural Norm Discovery from Conversations On-the-FlyabstractKnowledge of norms is needed to understand and reason about acceptable behavior in human communication and interactions across sociocultural scenarios.Most computational research on norms has focused on a single culture, and manually built datasets, from nonconversational settings.We address these limitations by proposing a new framework, NORMSAGE 1 , to automatically extract culturespecific norms from multi-lingual conversations.NORMSAGE uses GPT-3 prompting to 1) extract candidate norms directly from conversations and 2) provide explainable selfverification to ensure correctness and relevance.Comprehensive empirical results show the promise of our approach to extract highquality culture-aware norms from multi-lingual conversations (English and Chinese), across several quality metrics.Further, our relevance verification can be extended to assess the adherence and violation of any norm with respect to a conversation on-the-fly, along with textual explanation.NORMSAGE achieves an AUC of 94.6% in this grounding setup, with generated explanations matching human-written quality.𝐚) 𝐈𝐧𝐢𝐭𝐢𝐚𝐥 𝐃𝐢𝐬𝐜𝐨𝐯𝐞𝐫𝒚: 𝒅𝒗𝒓(⋅) irrelevant contradict entail Correctness Verdict ( ! 𝑪 𝒗 ): Correctness Explanation ( ! 𝑪 𝒆 ): Yes, honesty is the foundation of trust, and strong family relationships are built on trust. Yi R. Fung 0001, Tuhin Chakrabarty, Owen Rambow, Smaranda Muresan, Heng Ji 0001 |
EMNLP | 2 |
| 2022 | Multitask Instruction-based Prompting for Fallacy RecognitionabstractFallacies are used as seemingly valid arguments to support a position and persuade the audience about its validity.Recognizing fallacies is an intrinsically difficult task both for humans and machines.Moreover, a big challenge for computational models lies in the fact that fallacies are formulated differently across the datasets with differences in the input format (e.g., question-answer pair, sentence with fallacy fragment), genre (e.g., social media, dialogue, news), as well as types and number of fallacies (from 5 to 18 types per dataset).To move towards solving the fallacy recognition task, we approach these differences across datasets as multiple tasks and show how instruction-based prompting in a multitask setup based on the T5 model improves the results against approaches built for a specific dataset such as T5, BERT or GPT-3.We show the ability of this multitask prompting approach to recognize 28 unique fallacies across domains and genres and study the effect of model size and prompt choice by analyzing the per-class (i.e., fallacy type) results.Finally, we analyze the effect of annotation quality on model performance, and the feasibility of complementing this approach with external knowledge. Tariq Alhindi, Tuhin Chakrabarty, Elena Musi, Smaranda Muresan |
EMNLP | 2 |
| 2022 | Help me write a Poem - Instruction Tuning as a Vehicle for Collaborative Poetry WritingabstractRecent work in training large language models (LLMs) to follow natural language instructions has opened up exciting opportunities for natural language interface design.Building on the prior success of LLMs in the realm of computerassisted creativity, we aim to study if LLMs can improve the quality of user-generated content through collaboration.We present CoPoet, a collaborative poetry writing system.In contrast to auto-completing a user's text, CoPoet is controlled by user instructions that specify the attributes of the desired text, such as Write a sentence about 'love' or Write a sentence ending in 'fly'.The core component of our system is a language model fine-tuned on a diverse collection of instructions for poetry writing.Our model is not only competitive with publicly available LLMs trained on instructions (InstructGPT), but is also capable of satisfying unseen compositional instructions.A study with 15 qualified crowdworkers shows that users successfully write poems with CoPoet on diverse topics ranging from Monarchy to Climate change.Further, the collaboratively written poems are preferred by third-party evaluators over those written without the system. 1 Tuhin Chakrabarty, Vishakh Padmakumar, He He 0001 |
EMNLP | 1 |
| 2022 | FLUTE: Figurative Language Understanding through Textual ExplanationsabstractFigurative language understanding has been recently framed as a recognizing textual entailment (RTE) task (a.k.a.natural language inference, or NLI).However, similar to classical RTE/NLI datasets, the current benchmarks suffer from spurious correlations and annotation artifacts.To tackle this problem, work on NLI has built explanation-based datasets such as e-SNLI, allowing us to probe whether language models are right for the right reasons.Yet no such data exists for figurative language, making it harder to assess genuine understanding of such expressions.To address this issue, we release FLUTE, a dataset of 9,000 figurative NLI instances with explanations, spanning four categories: Sarcasm, Simile, Metaphor, and Idioms.We collect the data through a model-in-the-loop framework based on GPT-3, crowd workers, and expert annotators.We show how utilizing GPT-3 in conjunction with human annotators (novices and experts) can aid in scaling up the creation of datasets even for such complex linguistic phenomena as figurative language.The baseline performance of the T5 model fine-tuned on FLUTE shows that our dataset can bring us a step closer to developing models that understand figurative language through textual explanations. Tuhin Chakrabarty, Arkadiy Saakyan, Debanjan Ghosh, Smaranda Muresan |
EMNLP | 1 |
| 2022 | Fine-tuned Language Models are Continual LearnersabstractRecent work on large language models relies on the intuition that most natural language processing tasks can be described via natural language instructions and that models trained on these instructions show strong zero-shot performance on several standard datasets.However, these models even though impressive still perform poorly on a wide range of tasks outside of their respective training and evaluation sets.To address this limitation, we argue that a model should be able to keep extending its knowledge and abilities, without forgetting previous skills.In spite of the limited success of Continual Learning we show that Fine-tuned Language Models can be continual learners.We empirically investigate the reason for this success and conclude that Continual Learning emerges from self-supervision pre-training.Our resulting model Continual-T0 (CT0) is able to learn 8 new diverse language generation tasks, while still maintaining good performance on previous tasks, spanning in total 70 datasets.Finally, we show that CT0 is able to combine instructions in ways it was never trained for, demonstrating some level of instruction compositionality. Thomas Scialom, Tuhin Chakrabarty, Smaranda Muresan |
EMNLP | 2 |
| 2022 | It's not Rocket Science: Interpreting Figurative Language in NarrativesabstractAbstract Figurative language is ubiquitous in English. Yet, the vast majority of NLP research focuses on literal language. Existing text representations by design rely on compositionality, while figurative language is often non- compositional. In this paper, we study the interpretation of two non-compositional figurative languages (idioms and similes). We collected datasets of fictional narratives containing a figurative expression along with crowd-sourced plausible and implausible continuations relying on the correct interpretation of the expression. We then trained models to choose or generate the plausible continuation. Our experiments show that models based solely on pre-trained language models perform substantially worse than humans on these tasks. We additionally propose knowledge-enhanced models, adopting human strategies for interpreting figurative language types: inferring meaning from the context and relying on the constituent words’ literal meanings. The knowledge-enhanced models improve the performance on both the discriminative and generative tasks, further bridging the gap from human performance. Tuhin Chakrabarty, Yejin Choi 0001, Vered Shwartz |
Trans. Assoc. Comput. Linguistics | 1 |
| 2021 | COVID-Fact: Fact Extraction and Verification of Real-World Claims on COVID-19 PandemicabstractArkadiy Saakyan, Tuhin Chakrabarty, Smaranda Muresan. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Arkadiy Saakyan, Tuhin Chakrabarty, Smaranda Muresan |
ACL/IJCNLP (1) | 2 |
| 2021 | Metaphor Generation with Conceptual MappingsabstractKevin Stowe, Tuhin Chakrabarty, Nanyun Peng, Smaranda Muresan, Iryna Gurevych. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Kevin Stowe, Tuhin Chakrabarty, Nanyun Peng 0001, Smaranda Muresan, Iryna Gurevych |
ACL/IJCNLP (1) | 2 |
| 2021 | Don't Go Far Off: An Empirical Study on Neural Poetry TranslationabstractDespite constant improvements in machine translation quality, automatic poetry translation remains a challenging problem due to the lack of open-sourced parallel poetic corpora, and to the intrinsic complexities involved in preserving the semantics, style and figurative nature of poetry.We present an empirical investigation for poetry translation along several dimensions: 1) size and style of training data (poetic vs. non-poetic), including a zeroshot setup; 2) bilingual vs. multilingual learning; and 3) language-family-specific models vs. mixed-language-family models.To accomplish this, we contribute a parallel dataset of poetry translations for several language pairs.Our results show that multilingual fine-tuning on poetic text significantly outperforms multilingual fine-tuning on non-poetic text that is 35X larger in size, both in terms of automatic metrics (BLEU, BERTScore, COMET) and human evaluation metrics such as faithfulness (meaning and poetic style).Moreover, multilingual fine-tuning on poetic data outperforms bilingual fine-tuning on poetic data. Tuhin Chakrabarty, Arkadiy Saakyan, Smaranda Muresan |
EMNLP (1) | 1 |
| 2021 | Implicit Premise Generation with Discourse-aware Commonsense Knowledge ModelsabstractEnthymemes are defined as arguments where a premise or conclusion is left implicit.We tackle the task of generating the implicit premise in an enthymeme, which requires not only an understanding of the stated conclusion and premise, but also additional inferences that could depend on commonsense knowledge.The largest available dataset for enthymemes (Habernal et al., 2018) consists of 1.7k samples, which is not large enough to train a neural text generation model.To address this issue, we take advantage of a similar task and dataset: Abductive reasoning in narrative text (Bhagavatula et al., 2020).However, we show that simply using a state-of-the-art seq2seq model fine-tuned on this data might not generate meaningful implicit premises associated with the given enthymemes.We demonstrate that encoding discourse-aware commonsense during fine-tuning improves the quality of the generated implicit premises and outperforms all other baselines both in automatic and human evaluations on three different datasets. Tuhin Chakrabarty, Aadit Trivedi, Smaranda Muresan |
EMNLP (1) | 1 |
| 2021 | ENTRUST: Argument Reframing with Language Models and EntailmentabstractTuhin Chakrabarty, Christopher Hidey, Smaranda Muresan. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Tuhin Chakrabarty, Christopher Hidey, Smaranda Muresan |
NAACL-HLT | 1 |
| 2021 | MERMAID: Metaphor Generation with Symbolism and Discriminative DecodingabstractTuhin Chakrabarty, Xurui Zhang, Smaranda Muresan, Nanyun Peng. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Tuhin Chakrabarty, Xurui Zhang, Smaranda Muresan, Nanyun Peng 0001 |
NAACL-HLT | 1 |
| 2020 | R^3: Reverse, Retrieve, and Rank for Sarcasm Generation with Commonsense KnowledgeabstractWe propose an unsupervised approach for sarcasm generation based on a non-sarcastic input sentence. Our method employs a retrieve-and-edit framework to instantiate two major characteristics of sarcasm: reversal of valence and semantic incongruity with the context, which could include shared commonsense or world knowledge between the speaker and the listener. While prior works on sarcasm generation predominantly focus on context incongruity, we show that combining valence reversal and semantic incongruity based on the commonsense knowledge generates sarcasm of higher quality. Human evaluation shows that our system generates sarcasm better than humans 34% of the time, and better than a reinforced hybrid baseline 90% of the time. Tuhin Chakrabarty, Debanjan Ghosh, Smaranda Muresan, Nanyun Peng 0001 |
ACL | 1 |
| 2020 | DeSePtion: Dual Sequence Prediction and Adversarial Examples for Improved Fact-CheckingabstractChristopher Hidey, Tuhin Chakrabarty, Tariq Alhindi, Siddharth Varia, Kriste Krstovski, Mona Diab, Smaranda Muresan. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. Christopher Hidey, Tuhin Chakrabarty, Tariq Alhindi, Siddharth Varia, Kriste Krstovski, Mona T. Diab, Smaranda Muresan |
ACL | 2 |
| 2020 | Generating similes effortlessly like a Pro: A Style Transfer Approach for Simile GenerationabstractLiterary tropes, from poetry to stories, are at the crux of human imagination and communication.Figurative language, such as a simile, goes beyond plain expressions to give readers new insights and inspirations.We tackle the problem of simile generation.Generating a simile requires proper understanding for effective mapping of properties between two concepts.To this end, we first propose a method to automatically construct a parallel corpus by transforming a large number of similes collected from Reddit to their literal counterpart using structured common sense knowledge.We then fine-tune a pretrained sequence to sequence model, BART (Lewis et al., 2019), on the literal-simile pairs to generate novel similes given a literal sentence.Experiments show that our approach generates 88% novel similes that do not share properties with the training data.Human evaluation on an independent set of literal statements shows that our model generates similes better than two literary experts 37% 1 of the times, and three baseline systems including a recent metaphor generation model 71% 2 of the times when compared pairwise.3 We also show how replacing literal sentences with similes from our best model in machine generated stories improves evocativeness and leads to better acceptance by human judges.* The research was conducted when the author was at USC/ISI.1 We average 32.6% and 41.3% for 2 humans. 2 We average 82% ,63% and 68% for three baselines.3 The simile in the title is generated by our best model.Input: Generating similes effortlessly, output: Generating similes like a Pro. Tuhin Chakrabarty, Smaranda Muresan, Nanyun Peng 0001 |
EMNLP (1) | 1 |
| 2020 | Content Planning for Neural Story Generation with Aristotelian RescoringabstractLong-form narrative text generated from large language models manages a fluent impersonation of human writing, but only at the local sentence level, and lacks structure or global cohesion.We posit that many of the problems of story generation can be addressed via highquality content planning, and present a system that focuses on how to learn good plot structures to guide story generation.We utilize a plot-generation language model along with an ensemble of rescoring models that each implement an aspect of good story-writing as detailed in Aristotle's Poetics.We find that stories written with our more principled plotstructure are both more relevant to a given prompt and higher quality than baselines that do not content plan, or that plan in an unprincipled way. 1 Seraphina Goldfarb-Tarrant, Tuhin Chakrabarty, Ralph M. Weischedel, Nanyun Peng 0001 |
EMNLP (1) | 2 |
| 2019 | AMPERSAND: Argument Mining for PERSuAsive oNline DiscussionsabstractTuhin Chakrabarty, Christopher Hidey, Smaranda Muresan, Kathy McKeown, Alyssa Hwang. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Tuhin Chakrabarty, Christopher Hidey, Smaranda Muresan, Kathy McKeown, Alyssa Hwang |
EMNLP/IJCNLP (1) | 1 |
| 2019 | Discourse Relation Prediction: Revisiting Word Pairs with Convolutional NetworksabstractWord pairs across argument spans have been shown to be effective for predicting the discourse relation between them.We propose an approach to distill knowledge from word pairs for discourse relation classification with convolutional neural networks by incorporating joint learning of implicit and explicit relations.Our novel approach of representing the input as word pairs achieves state-of-theart results on four-way classification of both implicit and explicit relations as well as one of the binary classification tasks.For explicit relation prediction, we achieve around 20% error reduction on the four-way task.At the same time, compared to a two-layered Bi-LSTM-CRF model, our model is able to achieve these results with half the number of learnable parameters and approximately half the amount of training time. Siddharth Varia, Christopher Hidey, Tuhin Chakrabarty |
SIGdial | 3 |