Simone Borsci

dblp:10/8236 · DBLP profile ↗
← Back
10ranked-venue papers
8as first author
5since 2021 · last 2024
0000-0002-3591-3577ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 8 · 8 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 2
YearPublicationVenuePosition
2024 Human AI conversational systems: when humans and machines start to chat
Simone Borsci, Alan Chamberlain, Elena Nichele, Mads Bødker, Tommaso Turchi
Pers. Ubiquitous Comput.1
2024 Re-examining the chatBot Usability Scale (BUS-11) to assess user experience with customer relationship management chatbots
abstract
Abstract Intelligent systems, such as chatbots, are likely to strike new qualities of UX that are not covered by instruments validated for legacy human–computer interaction systems. A new validated tool to evaluate the interaction quality of chatbots is the chatBot Usability Scale (BUS) composed of 11 items in five subscales. The BUS-11 was developed mainly from a psychometric perspective, focusing on ranking people by their responses and also by comparing designs’ properties (designometric). In this article, 3186 observations (BUS-11) on 44 chatbots are used to re-evaluate the inventory looking at its factorial structure, and reliability from the psychometric and designometric perspectives. We were able to identify a simpler factor structure of the scale, as previously thought. With the new structure, the psychometric and the designometric perspectives coincide, with good to excellent reliability. Moreover, we provided standardized scores to interpret the outcomes of the scale. We conclude that BUS-11 is a reliable and universal scale, meaning that it can be used to rank people and designs, whatever the purpose of the research.
Simone Borsci, Martin Schmettow
Pers. Ubiquitous Comput.1
2023 Ciao AI: the Italian adaptation and validation of the Chatbot Usability Scale
abstract
Abstract Chatbot-based tools are becoming pervasive in multiple domains from commercial websites to rehabilitation applications. Only recently, an eleven-item satisfaction inventory was developed (the ChatBot Usability Scale, BUS-11) to help designers in the assessment process of their systems. The BUS-11 has been validated in multiple contexts and languages, i.e., English, German, Dutch, and Spanish. This scale forms a solid platform enabling designers to rapidly assess chatbots both during and after the design process. The present work aims to adapt and validate the BUS-11 inventory in Italian. A total of 1360 questionnaires were collected which related to a total of 10 Italian chatbot-based systems using the BUS-11 inventory and also using the lite version of the Usability Metrics for User eXperience for convergent validity purposes. The Italian version of the BUS-11 was adapted in terms of the wording of one item, and a Multi-Group Confirmatory Factorial Analysis was performed to establish the factorial structure of the scale and compare the effects of the wording adaptation. Results indicate that the adapted Italian version of the scale matches the expected factorial structure of the original scale. The Italian BUS-11 is highly reliable (Cronbach alpha: 0.921), and it correlates to other measures of satisfaction (e.g., UMUX-Lite, τb = 0.67; p < .001) by also offering specific insights regarding the chatbots’ characteristics. The Italian BUS-11 can be confidently used by chatbot designers to assess the satisfaction of their users during formative or summative tests.
Simone Borsci, Elisa Prati, Alessio Malizia, Martin Schmettow, Alan Chamberlain, Stefano Federici
Pers. Ubiquitous Comput.1
2023 A confirmatory factorial analysis of the Chatbot Usability Scale: a multilanguage validation
abstract
Abstract The Bot Usability Scale (BUS) is a standardised tool to assess and compare the satisfaction of users after interacting with chatbots to support the development of usable conversational systems. The English version of the 15-item BUS scale (BUS-15) was the result of an exploratory factorial analysis; a confirmatory factorial analysis tests the replicability of the initial model and further explores the properties of the scale aiming to optimise this tool seeking for the stability of the original model, the potential reduction of items, and testing multiple language versions of the scale. BUS-15 and the usability metrics for user experience (UMUX-LITE), used here for convergent validity purposes, were translated from English to Spanish, German, and Dutch. A total of 1292 questionnaires were completed in multiple languages; these were collected from 209 participants interacting with an overall pool of 26 chatbots. BUS-15 was acceptably reliable; however, a shorter and more reliable solution with 11 items (BUS-11) emerged from the data. The satisfaction ratings obtained with the translated version of BUS-11 were not significantly different from the original version in English, suggesting that the BUS-11 could be used in multiple languages. The results also suggested that the age of participants seems to affect the evaluation when using the scale, with older participants significantly rating the chatbots as less satisfactory, when compared to younger participants. In line with the expectations, based on reliability, BUS-11 positively correlates with UMUX-LITE scale. The new version of the scale (BUS-11) aims to facilitate the evaluation with chatbots, and its diffusion could help practitioners to compare the performances and benchmark chatbots during the product assessment stage. This tool could be a way to harmonise and enable comparability in the field of human and conversational agent interaction.
Simone Borsci, Martin Schmettow, Alessio Malizia, Alan Chamberlain, Frank van der Velde
Pers. Ubiquitous Comput.1
2022 The Chatbot Usability Scale: the Design and Pilot of a Usability Scale for Interaction with AI-Based Conversational Agents
abstract
Abstract Standardised tools to assess a user’s satisfaction with the experience of using chatbots and conversational agents are currently unavailable. This work describes four studies, including a systematic literature review, with an overall sample of 141 participants in the survey (experts and novices), focus group sessions and testing of chatbots to (i) define attributes to assess the quality of interaction with chatbots and (ii) the designing and piloting a new scale to measure satisfaction after the experience with chatbots. Two instruments were developed: (i) A diagnostic tool in the form of a checklist (BOT-Check). This tool is a development of previous works which can be used reliably to check the quality of a chatbots experience in line with commonplace principles. (ii) A 15-item questionnaire (BOT Usability Scale, BUS-15) with estimated reliability between .76 and .87 distributed in five factors. BUS-15 strongly correlates with UMUX-LITE by enabling designers to consider a broader range of aspects usually not considered in satisfaction tools for non-conversational agents, e.g. conversational efficiency and accessibility, quality of the chatbot’s functionality and so on. Despite the convincing psychometric properties, BUS-15 requires further testing and validation. Designers can use it as a tool to assess products, thus building independent databases for future evaluation of its reliability, validity and sensitivity.
Simone Borsci, Alessio Malizia, Martin Schmettow, Frank van der Velde, Gunay Tariverdiyeva, Divyaa Balaji, Alan Chamberlain
Pers. Ubiquitous Comput.1
2020 Preliminary Results of a Systematic Review: Quality Assessment of Conversational Agents (Chatbots) for People with Disabilities or Special Needs
Maria Laura De Filippis, Stefano Federici, Maria Laura Mele, Simone Borsci, Marco Bracalenti, Giancarlo Gaudino, Antonello Cocco, Massimo Amendola, Emilio Simonetti
ICCHP (1)4
2019 Shaking the usability tree: why usability is not a dead end, and a constructive way forward
abstract
A recent contribution to the ongoing debate concerning the concept of usability and its measures proposed that usability reached a dead end – i.e. a construct unable to provide stable results and to unify scientific knowledge. Extensive commentaries rejected the conclusion that researchers need to look for alternative constructs to measure the quality of interaction. Nevertheless, several practitioners involved in this international debate asked for a constructive way to move forward the usability practice. In fact, two key issues of the usability field were identified in this debate: (i) knowledge fragmentation in the scientific community, and (ii) the unstable relationship among the usability metrics. We recognise both the importance and impact of these key issues, although, in line with others, we may not agree with the conclusion that the usability is a dead end. Under the light of the international debate, this work discusses the strengths and weaknesses of usability construct and its application. Our discussion focuses on identifying alternative explanations to the issues and to suggest mitigation strategies, which may be considered the starting point to move forward the usability field. However, scientific community actions will be needed to implement these mitigation strategies and to harmonise the usability practice.
Simone Borsci, Stefano Federici, Alessio Malizia, Maria Laura De Filippis
Behav. Inf. Technol.1
2015 Assessing User Satisfaction in the Era of User Experience: Comparison of the SUS, UMUX, and UMUX-LITE as a Function of Product Experience
abstract
Nowadays, practitioners extensively apply quick and reliable scales of user satisfaction as part of their user experience analyses to obtain well-founded measures of user satisfaction within time and budget constraints. However, in the human–computer interaction literature the relationship between the outcomes of standardized satisfaction scales and the amount of product usage has been only marginally explored. The few studies that have investigated this relationship have typically shown that users who have interacted more with a product have higher satisfaction. The purpose of this article was to systematically analyze the variation in outcomes of three standardized user satisfaction scales (SUS, UMUX, UMUX-LITE) when completed by users who had spent different amounts of time with a website. In two studies, the amount of interaction was manipulated to assess its effect on user satisfaction. Measurements of the three scales were strongly correlated and their outcomes were significantly affected by the amount of interaction time. Notably, the SUS acted as a unidimensional scale when administered to people who had less product experience but was bidimensional when administered to users with more experience. Previous findings of similar magnitudes for the SUS and UMUX-LITE (after adjustment) were replicated but did not show the previously reported similarities of magnitude for the SUS and the UMUX. Results strongly encourage further research to analyze the relationships of the three scales with levels of product exposure. Recommendations for practitioners and researchers in the use of the questionnaires are also provided.
Simone Borsci, Stefano Federici, Silvia Bacci, Michela Gnaldi, Francesco Bartolucci
Int. J. Hum. Comput. Interact.1
2013 Reviewing and Extending the Five-User Assumption: A Grounded Procedure for Interaction Evaluation
abstract
The debate concerning how many participants represents a sufficient number for interaction testing is well-established and long-running, with prominent contributions arguing that five users provide a good benchmark when seeking to discover interaction problems. We argue that adoption of five users in this context is often done with little understanding of the basis for, or implications of, the decision. We present an analysis of relevant research to clarify the meaning of the five-user assumption and to examine the way in which the original research that suggested it has been applied. This includes its blind adoption and application in some studies, and complaints about its inadequacies in others. We argue that the five-user assumption is often misunderstood, not only in the field of Human-Computer Interaction, but also in fields such as medical device design, or in business and information applications. The analysis that we present allows us to define a systematic approach for monitoring the sample discovery likelihood, in formative and summative evaluations, and for gathering information in order to make critical decisions during the interaction testing, while respecting the aim of the evaluation and allotted budget. This approach -- which we call the Grounded Procedure -- is introduced and its value argued.
Simone Borsci, Robert D. Macredie, Julie Barnett, Jennifer L. Martin, Jasna Kuljis, Terry Young
ACM Trans. Comput. Hum. Interact.1
2010 Beyond a Visuocentric Way of a Visual Web Search Clustering Engine: The Sonification of WhatsOnWeb
Maria Laura Mele, Stefano Federici, Simone Borsci, Giuseppe Liotta
ICCHP (1)3