Lucia Donatelli

dblp:225/6227 · DBLP profile ↗
← Back
17ranked-venue papers
3as first author
14since 2021 · last 2026
0000-0002-5974-7454ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 17 · 3 first-author · 14 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Green Bots versus Red Bots: Evaluating Large Language Models for Simulating Persuasion Dynamics in Online Influence Campaigns
Majd Eddin Al Ali, Filip Mihai Muntean, Lucia Donatelli, Jurriaan van Diggelen
LREC3
2025 Trustworthy AI Psychotherapy: Multi-Agent LLM Workflow for Counseling and Explainable Mental Disorder Diagnosis
abstract
LLM-based agents have emerged as transformative tools capable of executing complex tasks through iterative planning and action, achieving significant advancements in understanding and addressing user needs. Yet, their effectiveness remains limited in specialized domains such as mental health diagnosis, where they underperform compared to general applications. Current approaches to integrating diagnostic capabilities into LLMs rely on scarce, highly sensitive mental health datasets, which are challenging to acquire. These methods also fail to emulate clinicians' proactive inquiry skills, lack multi-turn conversational comprehension, and struggle to align outputs with expert clinical reasoning. To address these gaps, we propose DSM5AgentFlow, the first LLM-based agent workflow designed to autonomously generate DSM-5 Level-1 diagnostic questionnaires. By simulating therapist-client dialogues with specific client profiles, the framework delivers transparent, step-by-step disorder predictions, producing explainable and trustworthy results. This workflow serves as a complementary tool for mental health diagnosis, ensuring adherence to ethical and legal standards. Through comprehensive experiments, we evaluate leading LLMs across three critical dimensions: conversational realism, diagnostic accuracy, and explainability. Our datasets and implementations are fully open-sourced.
Mithat Can Ozgun, Jiahuan Pei, Koen V. Hindriks, Lucia Donatelli, Qingzhi Liu
CIKM4
2025 Vision-and-Language Training Helps Deploy Taxonomic Knowledge but Does Not Fundamentally Alter It
abstract
Does vision-and-language (VL) training change the linguistic representations of language models in meaningful ways? In terms of downstream task performance on text-only tasks, most results in the literature have shown marginal differences. In this work, we start from the hypothesis that the domain in which VL training could have a significant effect is lexical-conceptual knowledge, in particular its taxonomic organization. Through comparing minimal pairs of text-only LMs and their VL-trained counterparts, we first show that the VL models often outperform their text-only counterparts on a text-only question-answering task that requires taxonomic understanding of concepts mentioned in the questions. Using an array of targeted behavioral and representational analyses, we show that the LMs and VLMs do not differ significantly in terms of their taxonomic knowledge itself, but they differ in how they represent questions that contain concepts in a taxonomic relation vs. a non-taxonomic relation. This implies that the taxonomic knowledge itself does not change substantially through additional VL training, but VL training does improve the deployment of this knowledge in the context of a specific task, even when the presentation of the task is purely linguistic.
Yulu Qin, Dheeraj Varghese, Adam Dahlgren Lindström, Lucia Donatelli, Kanishka Misra, Najoung Kim
NeurIPS4
2024 More frequent verbs are associated with more diverse valency frames: Efficient principles at the lexicon-grammar interface
abstract
A substantial body of work has provided evidence that the lexicons of natural languages are organized to support efficient communication.However, existing work has largely focused on word-internal properties, such as Zipf's observation that more frequent words are optimized in form to minimize communicative cost.Here, we investigate the hypothesis that efficient lexicon organization is also reflected in valency, or the combinations and orders of additional words and phrases a verb selects for in a sentence.We consider two measures of valency diversity for verbs: valency frame count (VFC), the number of distinct frames associated with a verb, and valency frame entropy (VFE), the average information content of frame selection associated with a verb.Using data from 79 languages, we provide evidence that more frequent verbs are associated with a greater diversity of valency frames, suggesting that the organization of valency is consistent with communicative efficiency principles.We discuss our findings in relation to classical findings such as Zipf's meaning-frequency law and the principle of least effort, as well as implications for theories of valency and communicative efficiency principles. 1
Lucia Donatelli, Michael Hahn 0001
ACL (1)2
2024 SPOTTER: A Framework for Investigating Convention Formation in a Visually Grounded Human-Robot Reference Task
abstract
Linguistic conventions that arise in dialogue reflect common ground and can increase communicative efficiency. Social robots that can understand these conventions and the process by which they arise have the potential to become efficient communication partners. Nevertheless, it is unclear how robots can engage in convention formation when presented with both familiar and new information. We introduce an adaptable game platform, SPOTTER, to study the dynamics of convention formation for visually grounded referring expressions in both human-human and human-robot interaction. Specifically, we seek to elicit convention forming for members of an inner circle of well-known individuals in the common ground, as opposed to individuals from an outer circle, who are unfamiliar. We release an initial corpus of 5000 utterances from two exploratory pilot experiments in Dutch. Different from previous work focussing on human-human interaction, we find that referring expressions for both familiar and unfamiliar individuals maintain their length throughout human-robot interaction. Stable conventions are formed, although these conventions can be impacted by distracting outer circle individuals. With our distinction between familiar and unfamiliar, we create a contrastive operationalization of common ground, which aids research into convention formation.
Jaap Kruijt, Peggy van Minkelen, Lucia Donatelli, Piek Vossen, Elly A. Konijn, Thomas Baier 0007
LREC/COLING3
2024 Encoding Gesture in Multimodal Dialogue: Creating a Corpus of Multimodal AMR
abstract
Abstract Meaning Representation (AMR) is a general-purpose meaning representation that has become popular for its clear structure, ease of annotation and available corpora, and overall expressiveness. While AMR was designed to represent sentence meaning in English text, recent research has explored its adaptation to broader domains, including documents, dialogues, spatial information, cross-lingual tasks, and gesture. In this paper, we present an annotated corpus of multimodal (speech and gesture) AMR in a task-based setting. Our corpus is multilayered, containing temporal alignments to both the speech signal and to descriptions of gesture morphology. We also capture coreference relationships across modalities, enabling fine-grained analysis of how the semantics of gesture and natural language interact. We discuss challenges that arise when identifying cross-modal coreference and anaphora, as well as in creating and evaluating multimodal corpora in general. Although we find AMR’s abstraction away from surface form (in both language and gesture) occasionally too coarse-grained to capture certain cross-modal interactions, we believe its flexibility allows for future work to fill in these gaps. Our corpus and annotation guidelines are available at https://github.com/klai12/encoding-gesture-multimodal-dialogue.
Kenneth Lai, Richard Brutti, Lucia Donatelli, James Pustejovsky
LREC/COLING3
2024 SCOUT: A Situated and Multi-Modal Human-Robot Dialogue Corpus
abstract
We introduce the Situated Corpus Of Understanding Transactions (SCOUT), a multi-modal collection of human-robot dialogue in the task domain of collaborative exploration. The corpus was constructed from multiple Wizard-of-Oz experiments where human participants gave verbal instructions to a remotely-located robot to move and gather information about its surroundings. SCOUT contains 89,056 utterances and 310,095 words from 278 dialogues averaging 320 utterances per dialogue. The dialogues are aligned with the multi-modal data streams available during the experiments: 5,785 images and 30 maps. The corpus has been annotated with Abstract Meaning Representation and Dialogue-AMR to identify the speaker’s intent and meaning within an utterance, and with Transactional Units and Relations to track relationships between utterances to reveal patterns of the Dialogue Structure. We describe how the corpus and its annotations have been used to develop autonomous human-robot systems and enable research in open questions of how humans speak to robots. We release this corpus to accelerate progress in autonomous, situated, human-robot dialogue, especially in the context of navigation tasks where details about the environment need to be discovered.
Stephanie M. Lukin, Claire Bonial, Matthew Marge, Taylor Hudson, Cory J. Hayes, Kimberly A. Pollard, Anthony Baker, Ashley Foots, Ron Artstein, Felix Gervits, Mitchell Abrams, Cassidy Henry, Lucia Donatelli, Anton Leuski, Susan G. Hill, David R. Traum, Clare R. Voss
LREC/COLING13
2024 A Corpus of German Abstract Meaning Representation (DeAMR)
abstract
We present the first comprehensive set of guidelines for German Abstract Meaning Representation (Deutsche AMR, DeAMR) along with an annotated corpus of 400 DeAMR. Taking English AMR (EnAMR) as our starting point, we propose significant adaptations to faithfully represent the structure and semantics of German, focusing particularly on verb frames, compound words, and modality. We validate our annotation through inter-annotator agreement and further evaluate our corpus with a comparison of structural divergences between EnAMR and DeAMR on parallel sentences, replicating previous work that finds both cases of cross-lingual structural alignment and cases of meaningful linguistic divergence. Finally, we fine-tune state-of-the-art multi-lingual and cross-lingual AMR parsers on our corpus and find that, while our small corpus is insufficient to produce quality output, there is a need to continue develop and evaluate against gold non-English AMR data.
Christoph Otto, Jonas Groschwitz, Alexander Koller, Xiulin Yang, Lucia Donatelli
LREC/COLING5
2024 ConvoFusion: Multi-Modal Conversational Diffusion for Co-Speech Gesture Synthesis
abstract
Gestures play a key role in human communication. Recent methods for co-speech gesture generation, while managing to generate beat-aligned motions, struggle generating gestures that are semantically aligned with the utterance. Compared to beat gestures that align naturally to the audio signal, semantically coherent gestures require modeling the complex interactions between the language and human motion, and can be controlled by focusing on certain words. Therefore, we present ConvoFusion, a diffusion-based approach for multi-modal gesture synthesis, which can not only generate gestures based on multi-modal speech inputs, but can also facilitate controllability in gesture synthesis. Our method proposes two guidance objectives that allow the users to modulate the impact of different conditioning modalities (e.g. audio vs text) as well as to choose certain words to be emphasized during gesturing. Our method is versatile in that it can be trained either for generating monologue gestures or even the conversational gestures. To further advance the research on multi-party interactive gestures, the DndGroup Gesture dataset is released, which contains 6 hours of gesture data showing 5 people interacting with one another. We compare our method with several recent works and demonstrate effectiveness of our method on a variety of tasks. We urge the reader to watch our supplementary video at our webpage.
Muhammad Hamza Mughal, Rishabh Dabral, Ikhsanul Habibie, Lucia Donatelli, Marc Habermann, Christian Theobalt
CVPR4
2024 Cultural Adaptation of Recipes
abstract
Abstract Building upon the considerable advances in Large Language Models (LLMs), we are now equipped to address more sophisticated tasks demanding a nuanced understanding of cross-cultural contexts. A key example is recipe adaptation, which goes beyond simple translation to include a grasp of ingredients, culinary techniques, and dietary preferences specific to a given culture. We introduce a new task involving the translation and cultural adaptation of recipes between Chinese- and English-speaking cuisines. To support this investigation, we present CulturalRecipes, a unique dataset composed of automatically paired recipes written in Mandarin Chinese and English. This dataset is further enriched with a human-written and curated test set. In this intricate task of cross-cultural recipe adaptation, we evaluate the performance of various methods, including GPT-4 and other LLMs, traditional machine translation, and information retrieval techniques. Our comprehensive analysis includes both automatic and human evaluation metrics. While GPT-4 exhibits impressive abilities in adapting Chinese recipes into English, it still lags behind human expertise when translating English recipes into Chinese. This underscores the multifaceted nature of cultural adaptations. We anticipate that these insights will significantly contribute to future research on culturally aware language models and their practical application in culturally diverse contexts.
Yong Cao 0001, Yova Kementchedjhieva, Ruixiang Cui, Antonia Karamolegkou, Li Zhou 0010, Megan Dare, Lucia Donatelli, Daniel Hershcovich
Trans. Assoc. Comput. Linguistics7
2023 AMR Parsing is Far from Solved: GrAPES, the Granular AMR Parsing Evaluation Suite
abstract
Message from the General Chair I am happy to welcome you to EMNLP-2023 in Singapore!Like EMNLP-2021, EMNLP-2022, and other ACL-related meetings, we decided to host EMNLP-2023 as another hybrid conference having both in-person and virtual presentations and participants.We are not sure how long this style of our meetings will last.However, we have already accustomed to this style of conferences, which has its own advantages, while it causes a heavy burden to those organizing such events.The past one year has been a terrific and thrilling year since the advent of ChatGPT and other Large Language Models.Any people having access to those models has posed a big impact on people's impression about AI and has started to give them a feeling of fear.People now can do not only natural conversation with AI but also conduct various natural language tasks using our own languages.We now know it is difficult to guarantee that Large Language Models produce honest and harmless outputs.We have found that good prompting, demonstrations and complex ones like the Chain of Thought prompting draw out or enhance the emergent abilities of Large Language Models.However, we still don't know precisely why and how such in-context learning works.This year's EMNLP highlights a theme track, "Large Language Models and the Future of NLP."I hope we can see enthusiastic discussions and innovative ideas will be presented in EMNLP-2023.One big trial is that the Program Chairs decided to use OpenReview as the cradle of the main conference papers, for making reviews and author responses publicly available.The motivation and effects of this trial will be explained by the PC Chairs.Another important trial is to rent out the Universal Studio Singapore for our Social Event.I hope everyone will enjoy this event.EMNLP-2023 is the biggest conference ever in the SIGDAT history.Organizing such a big event is very difficult.As the General Chair, the most important and difficult task is to organize all the committees by a group of enthusiastic and talented people.I was very fortunate to be able to collect great committee members.Without such a wonderful group of colleagues, it almost has been impossible to make this great event happen.I would like to send my sincere thanks to all the members of our organization teams.Here, I only list the chairs by names, but I also like to send gratitude from my heart to all the people involved in EMNLP-2023, including keynote speakers, panelists, workshop organizers, tutorial tutors, senior area chairs, area chairs, reviewers, volunteers, sponsors, the Underline team, and all of you attending EMNLP-2023 in-person or virtually.• The program chairs -Houda Bouamor, Juan Pino, and Kalika Bali -who made a number of innovations and handled a huge number of submitted papers.I cannot help but be grateful for their tireless work.• The Local Chair and the Local Team -Haizhou Li the Chair organized and lead a wonderful group of people.While I cannot name every one of them, weekly meetings with the team members including related Chairs made our communication smooth and worked as a good time-keeper.For the remaining committee chairs, I only list them by names, as I cannot give all my gratitude only with short messages.
Jonas Groschwitz, Shay B. Cohen, Lucia Donatelli, Meaghan Fowlie
EMNLP3
2023 SLOG: A Structural Generalization Benchmark for Semantic Parsing
abstract
The goal of compositional generalization benchmarks is to evaluate how well models generalize to new complex linguistic expressions.Existing benchmarks often focus on lexical generalization, the interpretation of novel lexical items in syntactic structures familiar from training.Structural generalization tasks, where a model needs to interpret syntactic structures that are themselves unfamiliar from training, are often underrepresented, resulting in overly optimistic perceptions of how well models can generalize.We introduce SLOG, a semantic parsing dataset that extends COGS (Kim and Linzen, 2020) with 17 structural generalization cases.In our experiments, the generalization accuracy of Transformer models, including pretrained ones, only reaches 40.6%, while a structure-aware parser only achieves 70.8%.These results are far from the near-perfect accuracy existing models achieve on COGS, demonstrating the role of SLOG in foregrounding the large discrepancy between models' lexical and structural generalization capacities.
Bingzhi Li, Lucia Donatelli, Alexander Koller, Tal Linzen, Yuekun Yao, Najoung Kim
EMNLP2
2022 Abstract Meaning Representation for Gesture
abstract
This paper presents Gesture AMR, an extension to Abstract Meaning Representation (AMR), that captures the meaning of gesture. In developing Gesture AMR, we consider how gesture form and meaning relate; how gesture packages meaning both independently and in interaction with speech; and how the meaning of gesture is temporally and contextually determined. Our case study for developing Gesture AMR is a focused human-human shared task to build block structures. We develop an initial taxonomy of gesture act relations that adheres to AMR’s existing focus on predicate-argument structure while integrating meaningful elements unique to gesture. Pilot annotation shows Gesture AMR to be more challenging than standard AMR, and illustrates the need for more work on representation of dialogue and multimodal meaning. We discuss challenges of adapting an existing meaning representation to non-speech-based modalities and outline several avenues for expanding Gesture AMR.
Richard Brutti, Lucia Donatelli, Kenneth Lai, James Pustejovsky
LREC2
2021 Aligning Actions Across Recipe Graphs
abstract
Recipe texts are an idiosyncratic form of instructional language that pose unique challenges for automatic understanding.One challenge is that a cooking step in one recipe can be explained in another recipe in different words, at a different level of abstraction, or not at all.Previous work has annotated correspondences between recipe instructions at the sentence level, often glossing over important correspondences between cooking steps across recipes.We present a novel and fully-parsed English recipe corpus, ARA (Aligned Recipe Actions), which annotates correspondences between individual actions across similar recipes with the goal of capturing information implicit for accurate recipe understanding.We represent this information in the form of recipe graphs, and we train a neural model for predicting correspondences on ARA.We find that substantial gains in accuracy can be obtained by taking fine-grained structural information about the recipes into account.
Lucia Donatelli, Theresa Schmidt, Debanjali Biswas, Arne Köhn, Fangzhou Zhai, Alexander Koller
EMNLP (1)1
2020 Normalizing Compositional Structures Across Graphbanks
abstract
The emergence of a variety of graph-based meaning representations (MRs) has sparked an important conversation about how to adequately represent semantic structure.MRs exhibit structural differences that reflect different theoretical and design considerations, presenting challenges to uniform linguistic analysis and cross-framework semantic parsing.Here, we ask the question of which design differences between MRs are meaningful and semantically-rooted, and which are superficial.We present a methodology for normalizing discrepancies between MRs at the compositional level (Lindemann et al., 2019), finding that we can normalize the majority of divergent phenomena using linguistically-grounded rules.Our work significantly increases the match in compositional structure between MRs and improves multi-task learning (MTL) in a low-resource setting, serving as a proof of concept for future broad-scale cross-MR normalization.
Lucia Donatelli, Jonas Groschwitz, Matthias Lindemann, Alexander Koller, Pia Weißenhorn
COLING1
2020 A Two-Level Interpretation of Modality in Human-Robot Dialogue
abstract
We analyze the use and interpretation of modal expressions in a corpus of situated human-robot dialogue and ask how to effectively represent these expressions for automatic learning and dynamic interpretation in context.We present a two-level annotation scheme for modality that captures both content and intent, integrating a logic-based, semantic representation and a task-oriented, pragmatic representation that maps to our robot's capabilities.Data from our annotation task reveals that the interpretation of modal expressions in human-robot dialogue is quite diverse, yet highly constrained by the physical environment and asymmetrical speaker/addressee relationship.We sketch a formal model of human-robot common ground in which modality can be grounded and dynamically interpreted relative to speaker role, temporal constraints, and physical environment.
Lucia Donatelli, Kenneth Lai, James Pustejovsky
COLING1
2020 Dialogue-AMR: Abstract Meaning Representation for Dialogue
abstract
This paper describes a schema that enriches Abstract Meaning Representation (AMR) in order to provide a semantic representation for facilitating Natural Language Understanding (NLU) in dialogue systems. AMR offers a valuable level of abstraction of the propositional content of an utterance; however, it does not capture the illocutionary force or speaker’s intended contribution in the broader dialogue context (e.g., make a request or ask a question), nor does it capture tense or aspect. We explore dialogue in the domain of human-robot interaction, where a conversational robot is engaged in search and navigation tasks with a human partner. To address the limitations of standard AMR, we develop an inventory of speech acts suitable for our domain, and present “Dialogue-AMR”, an enhanced AMR that represents not only the content of an utterance, but the illocutionary force behind it, as well as tense and aspect. To showcase the coverage of the schema, we use both manual and automatic methods to construct the “DialAMR” corpus—a corpus of human-robot dialogue annotated with standard AMR and our enriched Dialogue-AMR schema. Our automated methods can be used to incorporate AMR into a larger NLU pipeline supporting human-robot dialogue.
Claire Bonial, Lucia Donatelli, Mitchell Abrams, Stephanie M. Lukin, Stephen Tratz, Matthew Marge, Ron Artstein, David R. Traum, Clare R. Voss
LREC2