Luciana Benotti

dblp:22/1403 · DBLP profile ↗
← Back
17ranked-venue papers
10as first author
7since 2021 · last 2025
0000-0001-7456-4333ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 8 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 5 · 2 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Vision and language · 42% Trustworthy machine learning · 32% Information extraction and text analysis · 11%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computing education · 100%

Topics — the 7 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
fairness
0.912025
HESEIA: A community-based dataset for evaluating social biases in large language models, co-designed in real school settings in Latin America · EMNLP 2025
Machine learning › Trustworthy machine learning › fairness › fairness evaluation
social bias evaluation
0.912025
HESEIA: A community-based dataset for evaluating social biases in large language models, co-designed in real school settings in Latin America · EMNLP 2025
Computer vision › Vision and language › visual question answering
multilingual visual question answering
0.812024
CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark · NeurIPS 2024
Computer vision › Vision and language
multimodal benchmark
0.812024
CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark · NeurIPS 2024
Computer vision › Vision and language
visual question answering
0.812024
CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark · NeurIPS 2024
Natural language and speech › Question answering and dialogue systems
visual dialog
0.512021
Region under Discussion for visual dialog · EMNLP (1) 2021
Natural language and speech › Language models and text generation
large language model
0.312025
HESEIA: A community-based dataset for evaluating social biases in large language models, co-designed in real school settings in Latin America · EMNLP 2025

Methods — techniques the papers use, named apart from their topics

dataset co-design · 1.7multimodal large language model benchmarking · 0.8manual classification · 0.6multimodal transformer · 0.5interpretable representation · 0.5
YearPublicationVenuePosition
2025 HESEIA: A community-based dataset for evaluating social biases in large language models, co-designed in real school settings in Latin America
abstract
Guido Ivetta, Marcos J Gomez, Sofía Martinelli, Pietro Palombini, M Emilia Echeveste, Nair Carolina Mazzeo, Beatriz Busaniche, Luciana Benotti. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025.
Guido Ivetta, Marcos J. Gomez, Sofía Martinelli, Pietro Palombini, Maria Emilia Echeveste, Nair Carolina Mazzeo, Beatriz Busaniche, Luciana Benotti
EMNLP8
2024 Your Stereotypical Mileage May Vary: Practical Challenges of Evaluating Biases in Multiple Languages and Cultural Contexts
abstract
Warning: This paper contains explicit statements of offensive stereotypes which may be upsetting The study of bias, fairness and social impact in Natural Language Processing (NLP) lacks resources in languages other than English. Our objective is to support the evaluation of bias in language models in a multilingual setting. We use stereotypes across nine types of biases to build a corpus containing contrasting sentence pairs, one sentence that presents a stereotype concerning an underadvantaged group and another minimally changed sentence, concerning a matching advantaged group. We build on the French CrowS-Pairs corpus and guidelines to provide translations of the existing material into seven additional languages. In total, we produce 11,139 new sentence pairs that cover stereotypes dealing with nine types of biases in seven cultural contexts. We use the final resource for the evaluation of relevant monolingual and multilingual masked language models. We find that language models in all languages favor sentences that express stereotypes in most bias categories. The process of creating a resource that covers a wide range of language types and cultural settings highlights the difficulty of bias evaluation, in particular comparability across languages and contexts.
Karën Fort, Laura Alonso Alemany, Luciana Benotti, Julien Bezançon, Claudia Borg, Marthese Borg, Yongjian Chen, Fanny Ducel, Yoann Dupont, Guido Ivetta, Margot Mieskes, Marco Naguib, Yuyan Qian, Matteo Radaelli, Wolfgang Schmeisser-Nieto, Emma Raimundo Schulz, Thiziri Saci, Sarah Saidi, Javier Torroba Marchante, Shilin Xie, Sergio E. Zanotto, Aurélie Névéol
LREC/COLING3
2024 CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark
abstract
Visual Question Answering~(VQA) is an important task in multimodal AI, which requires models to understand and reason on knowledge present in visual and textual data. However, most of the current VQA datasets and models are primarily focused on English and a few major world languages, with images that are Western-centric. While recent efforts have tried to increase the number of languages covered on VQA datasets, they still lack diversity in low-resource languages. More importantly, some datasets extend the text to other languages, either via translation or some other approaches, but usually keep the same images, resulting in narrow cultural representation. To address these limitations, we create CVQA, a new Culturally-diverse Multilingual Visual Question Answering benchmark dataset, designed to cover a rich set of languages and regions, where we engage native speakers and cultural experts in the data collection process. CVQA includes culturally-driven images and questions from across 28 countries in four continents, covering 26 languages with 11 scripts, providing a total of 9k questions. We benchmark several Multimodal Large Language Models (MLLMs) on CVQA, and we show that the dataset is challenging for the current state-of-the-art models. This benchmark will serve as a probing evaluation suite for assessing the cultural bias of multimodal models and hopefully encourage more research efforts towards increasing cultural awareness and linguistic diversity in this field.
Chenyang Lyu, Haryo Akbarianto Wibowo, Santiago Góngora, Aishik Mandal, Sukannya Purkayastha, Jesús-Germán Ortiz-Barajas, Emilio Villa-Cueva, Jinheon Baek, Soyeong Jeong, Injy Hamed, Zheng Wei Lim, Paula Mónica Silva, Jocelyn Dunstan, Mélanie Jouitteau, David Le Meur, Joan Nwatu, Ganzorig Batnasan, Munkh-Erdene Otgonbold, Munkhjargal Gochoo, Guido Ivetta, Luciana Benotti, Laura Alonso Alemany, Hernán Maina, Jiahui Geng, Tiago Timponi Torrent, Frederico Belcavello, Marcelo Viridiano, Jan Christian Blaise Cruz, Dan John Velasco, Oana Ignat, Zara Burzo, Chenxi Whitehouse, Artem Abzaliev, Teresa Clifford, Grainne Caulfield, Teresa Lynn, Christian Salamea Palacios, Vladimir Araujo, Yova Kementchedjhieva, Mihail Mihaylov, Israel Abebe Azime, Henok Biadglign Ademtew, Bontu Fufa Balcha, Naome A. Etori, David Ifeoluwa Adelani, Rada Mihalcea, Atnafu Lambebo Tonja, Maria Camila Buitrago Cabrera, Gisela Vallejo, Holy Lovenia, Ruochen Zhang 0001, Marcos Estecha-Garitagoitia, Mario Rodríguez-Cantelar, Toqeer Ehsan, Rendi Chevi, Muhammad Farid Adilazuarda, Ryandito Diandaru, Samuel Cahyawijaya, Fajri Koto, Tatsuki Kuribayashi, Haiyue Song, Aditya Khandavally, Thanmay Jayakumar, Raj Dabre, Mohamed Fazli Mohamed Imam, Kumaranage Ravindu Yasas Nagasinghe, Alina Dragonetti, Luis Fernando D'Haro, Olivier Niyomugisha, Jay Gala, Pranjal A. Chitale, Fauzan Farooqui, Thamar Solorio, Alham Fikri Aji
NeurIPS23
2022 Ethics consideration sections in natural language processing papers
abstract
In this paper we present the results of a manual classification of all ethical consideration sections for ACL 2021.We also compare how many papers had an ethics consideration section per track and per world region in ACL 2021.We classified papers according to the ethical issues covered (research benefits, potential harms, and vulnerable groups affected) and whether the paper was marked as requiring ethics review by at least one reviewer.Moreover, we discuss recurring obstacles we have observed (highlighting some interesting texts we found along the way) and conclude with three suggestions.We think that this paper may be useful for anyone who needs to write -or review -an ethics section and would like to get an overview of what others have done.
Luciana Benotti, Patrick Blackburn
EMNLP1
2021 Grounding as a Collaborative Process
abstract
Collaborative grounding is a fundamental aspect of human-human dialog which allows people to negotiate meaning.In this paper we argue that it is missing from current deep learning approaches to dialog and interactive systems.Our central point is that making mistakes and being able to recover from them collaboratively is a key ingredient in grounding meaning.We illustrate the pitfalls of being unable to ground collaboratively, discuss what can be learned from the language acquisition and dialog systems literature, and reflect on how to move forward.
Luciana Benotti, Patrick Blackburn
EACL1
2021 Region under Discussion for visual dialog
abstract
Visual Dialog is assumed to require the dialog history to generate correct responses during a dialog.However, it is not clear from previous work how dialog history is needed for visual dialog.In this paper we define what it means for visual questions to require dialog history and we propose a methodology for identifying them.We release a subset of the Guesswhat?! questions for which their dialog history completely changes their responses.We propose a novel interpretable representation that visually grounds dialog history: the Region under Discussion.It constrains the image's spatial features according to a semantic representation of the history inspired by the information structure notion of Question under Discussion.We evaluate the architecture on task-specific multimodal models and the visual transformer model LXMERT and show that there is still room for improvement.
Mauricio Mazuecos, Franco M. Luque, Hernán Maina, Thomas Vadora, Luciana Benotti
EMNLP (1)6
2021 A recipe for annotating grounded clarifications
abstract
In order to interpret the communicative intents of an utterance, it needs to be grounded in something that is outside of language; that is, grounded in world modalities.In this paper we argue that dialogue clarification mechanisms make explicit the process of interpreting the communicative intents of the speaker's utterances by grounding them in the various modalities in which the dialogue is situated.This paper frames dialogue clarification mechanisms as an understudied research problem and a key missing piece in the giant jigsaw puzzle of natural language understanding.We discuss both the theoretical background and practical challenges posed by this problem, and propose a recipe for obtaining grounding annotations.We conclude by highlighting ethical issues that need to be addressed in future work.
Luciana Benotti, Patrick Blackburn
NAACL-HLT1
2019 Text-based Programming in Elementary School: A Comparative Study of Programming Abilities in Children with and without Block-based Experience
abstract
This paper describes an elementary school intervention to teach a text-based programming language to 10-11 year old students. We compare students with no previous programming experience with students with 3 semesters of experience with a block-based programming language. We analyze students' performance and learning based on detailed logs in an online programming platform and on multiple choice tests. Although both groups have a similar percentage of syntactical errors, the experienced group showed a better performance on exam scores and a lower number of test case errors. These findings suggest that, 10-11 year old students benefit from block-based experience when learning a new text-based programming language.
Marcos J. Gomez, Marco Moresi, Luciana Benotti
ITiCSE3
2018 The Effect of a Web-based Coding Tool with Automatic Feedback on Students' Performance and Perceptions
abstract
In this paper we do three things. First, we describe a web-based coding tool that is open-source, publicly available and provides formative feedback and assessment. Second, we compare several metrics on student performance in courses that use the tool versus courses that do not use it when learning to program in Haskell. We find that the dropout rates are significantly lower in those courses that use the tool at two different universities. Finally we apply the technology acceptance model to analyse students perceptions.
Luciana Benotti, Federico Aloi, Franco Bulgarelli, Marcos J. Gomez
SIGCSE1
2017 Modeling the clarification potential of instructions: Predicting clarification requests and other reactions
Luciana Benotti, Patrick Blackburn
Comput. Speech Lang.1
2016 Lessons Learned on Computer Science Teachers Professional Development
abstract
This paper describes an introductory Computer Science (CS) Professional Development (PD) course for K-12 teachers in Argentina that integrates pedagogical content knowledge and teacher classroom practice. We analyzed teachers' learning of what CS entails and the implementation of inquirybased programming lessons in their schools. Based on pre and post teachers surveys and classroom observations, we found that most teachers learned about the CS object of study and about fundamental programming concepts such as conditionals, loops, variables, etc. Teachers were more likely to replicate the same activities they experienced during PD workshops in their classrooms than to produce their own. Teachers who had a previous background on CS provided in-depth explanations of CS concepts to their students while other teachers superficially introduced the content knowledge. We describe PD activities and characteristics that could explain teachers' learning and incorporation of programming lessons. Findings imply that a PD program that integrates pedagogical content knowledge and teachers classroom practice can effectively improve inquiry-based CS teaching, but may be insufficient preparation for teachers with no previous background on CS.
María Cecilia Martínez, Marcos J. Gomez, Marco Moresi, Luciana Benotti
ITiCSE4
2015 A Comparison of Preschool and Elementary School Children Learning Computer Science Concepts through a Multilanguage Robot Programming Platform
abstract
This paper describes a school intervention to teach fundamental Computer Science (CS) concepts to 3-11 year old students with a multilanguage robot programming platform (using drag and drop, Python and C++ languages) in Argentina. We analyze students' performance and learning process based on multiple choice test and classroom observations. Data show that all students can intuitively learn sequence, conditional, loops and parameters and that girls performed slightly better than boys. Older students can easily combine these concepts to write a program. The multilanguage platform promotes student spontaneous exploration of more sophisticated CS concepts and languages. These findings imply that introducing CS in mandatory schooling from an inquiry based approach is both achievable and beneficial.
María Cecilia Martínez, Marcos J. Gomez, Luciana Benotti
ITiCSE3
2014 Engaging high school students using chatbots
abstract
Chatbots have been used in different scenarios for getting people interested in CS for decades. However, their potential for teaching basic concepts and their engaging effect has not been measured. In this paper we present a software platform called Chatbot designed to foster engagement while teaching basic CS concepts such as variables, conditionals and finite state automata, among others. We carried out two experiences using Chatbot and the well known platform Alice: 1) an online nation-wide competition, and 2) an in-class 15-lesson pilot course in 2 high schools. Data shows that retention and girl interest are higher with Chatbot than with Alice, indicating student engagement.
Luciana Benotti, María Cecilia Martínez, Fernando Schapachnik
ITiCSE1
2014 Interpreting Natural Language Instructions Using Language, Vision, and Behavior
abstract
We define the problem of automatic instruction interpretation as follows. Given a natural language instruction, can we automatically predict what an instruction follower, such as a robot, should do in the environment to follow that instruction? Previous approaches to automatic instruction interpretation have required either extensive domain-dependent rule writing or extensive manually annotated corpora. This article presents a novel approach that leverages a large amount of unannotated, easy-to-collect data from humans interacting in a game-like environment. Our approach uses an automatic annotation phase based on artificial intelligence planning, for which two different annotation strategies are compared: one based on behavioral information and the other based on visibility information. The resulting annotations are used as training data for different automatic classifiers. This algorithm is based on the intuition that the problem of interpreting a situated instruction can be cast as a classification problem of choosing among the actions that are possible in the situation. Classification is done by combining language, vision, and behavior information. Our empirical analysis shows that machine learning classifiers achieve 77% accuracy on this task on available English corpora and 74% on similar German corpora. Finally, the inclusion of human feedback in the interpretation process is shown to boost performance to 92% for the English corpus and 90% for the German corpus.
Luciana Benotti, Tessa A. Lau, Martin Villalba
ACM Trans. Interact. Intell. Syst.1
2011 Giving instructions in virtual environments by corpus based selection
Luciana Benotti, Alexandre Denis 0002
SIGDIAL Conference1
2010 Negotiating causal implicatures
Luciana Benotti, Patrick Blackburn
SIGDIAL Conference1
2009 Clarification Potential of Instructions
Luciana Benotti
SIGDIAL Conference1