VLDB 2026 Research / reviewers in the wild / expert
Martin Schmettow
dblp:43/1364
· DBLP profile ↗
11ranked-venue papers
5as first author
4since 2021 · last 2024
0000-0003-0240-4087ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 8 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-authorArtificial intelligence and machine learning · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | Re-examining the chatBot Usability Scale (BUS-11) to assess user experience with customer relationship management chatbotsabstractAbstract Intelligent systems, such as chatbots, are likely to strike new qualities of UX that are not covered by instruments validated for legacy human–computer interaction systems. A new validated tool to evaluate the interaction quality of chatbots is the chatBot Usability Scale (BUS) composed of 11 items in five subscales. The BUS-11 was developed mainly from a psychometric perspective, focusing on ranking people by their responses and also by comparing designs’ properties (designometric). In this article, 3186 observations (BUS-11) on 44 chatbots are used to re-evaluate the inventory looking at its factorial structure, and reliability from the psychometric and designometric perspectives. We were able to identify a simpler factor structure of the scale, as previously thought. With the new structure, the psychometric and the designometric perspectives coincide, with good to excellent reliability. Moreover, we provided standardized scores to interpret the outcomes of the scale. We conclude that BUS-11 is a reliable and universal scale, meaning that it can be used to rank people and designs, whatever the purpose of the research. Simone Borsci, Martin Schmettow |
Pers. Ubiquitous Comput. | 2 |
| 2023 | Ciao AI: the Italian adaptation and validation of the Chatbot Usability ScaleabstractAbstract Chatbot-based tools are becoming pervasive in multiple domains from commercial websites to rehabilitation applications. Only recently, an eleven-item satisfaction inventory was developed (the ChatBot Usability Scale, BUS-11) to help designers in the assessment process of their systems. The BUS-11 has been validated in multiple contexts and languages, i.e., English, German, Dutch, and Spanish. This scale forms a solid platform enabling designers to rapidly assess chatbots both during and after the design process. The present work aims to adapt and validate the BUS-11 inventory in Italian. A total of 1360 questionnaires were collected which related to a total of 10 Italian chatbot-based systems using the BUS-11 inventory and also using the lite version of the Usability Metrics for User eXperience for convergent validity purposes. The Italian version of the BUS-11 was adapted in terms of the wording of one item, and a Multi-Group Confirmatory Factorial Analysis was performed to establish the factorial structure of the scale and compare the effects of the wording adaptation. Results indicate that the adapted Italian version of the scale matches the expected factorial structure of the original scale. The Italian BUS-11 is highly reliable (Cronbach alpha: 0.921), and it correlates to other measures of satisfaction (e.g., UMUX-Lite, τb = 0.67; p < .001) by also offering specific insights regarding the chatbots’ characteristics. The Italian BUS-11 can be confidently used by chatbot designers to assess the satisfaction of their users during formative or summative tests. Simone Borsci, Elisa Prati, Alessio Malizia, Martin Schmettow, Alan Chamberlain, Stefano Federici |
Pers. Ubiquitous Comput. | 4 |
| 2023 | A confirmatory factorial analysis of the Chatbot Usability Scale: a multilanguage validationabstractAbstract The Bot Usability Scale (BUS) is a standardised tool to assess and compare the satisfaction of users after interacting with chatbots to support the development of usable conversational systems. The English version of the 15-item BUS scale (BUS-15) was the result of an exploratory factorial analysis; a confirmatory factorial analysis tests the replicability of the initial model and further explores the properties of the scale aiming to optimise this tool seeking for the stability of the original model, the potential reduction of items, and testing multiple language versions of the scale. BUS-15 and the usability metrics for user experience (UMUX-LITE), used here for convergent validity purposes, were translated from English to Spanish, German, and Dutch. A total of 1292 questionnaires were completed in multiple languages; these were collected from 209 participants interacting with an overall pool of 26 chatbots. BUS-15 was acceptably reliable; however, a shorter and more reliable solution with 11 items (BUS-11) emerged from the data. The satisfaction ratings obtained with the translated version of BUS-11 were not significantly different from the original version in English, suggesting that the BUS-11 could be used in multiple languages. The results also suggested that the age of participants seems to affect the evaluation when using the scale, with older participants significantly rating the chatbots as less satisfactory, when compared to younger participants. In line with the expectations, based on reliability, BUS-11 positively correlates with UMUX-LITE scale. The new version of the scale (BUS-11) aims to facilitate the evaluation with chatbots, and its diffusion could help practitioners to compare the performances and benchmark chatbots during the product assessment stage. This tool could be a way to harmonise and enable comparability in the field of human and conversational agent interaction. Simone Borsci, Martin Schmettow, Alessio Malizia, Alan Chamberlain, Frank van der Velde |
Pers. Ubiquitous Comput. | 2 |
| 2022 | The Chatbot Usability Scale: the Design and Pilot of a Usability Scale for Interaction with AI-Based Conversational AgentsabstractAbstract Standardised tools to assess a user’s satisfaction with the experience of using chatbots and conversational agents are currently unavailable. This work describes four studies, including a systematic literature review, with an overall sample of 141 participants in the survey (experts and novices), focus group sessions and testing of chatbots to (i) define attributes to assess the quality of interaction with chatbots and (ii) the designing and piloting a new scale to measure satisfaction after the experience with chatbots. Two instruments were developed: (i) A diagnostic tool in the form of a checklist (BOT-Check). This tool is a development of previous works which can be used reliably to check the quality of a chatbots experience in line with commonplace principles. (ii) A 15-item questionnaire (BOT Usability Scale, BUS-15) with estimated reliability between .76 and .87 distributed in five factors. BUS-15 strongly correlates with UMUX-LITE by enabling designers to consider a broader range of aspects usually not considered in satisfaction tools for non-conversational agents, e.g. conversational efficiency and accessibility, quality of the chatbot’s functionality and so on. Despite the convincing psychometric properties, BUS-15 requires further testing and validation. Designers can use it as a tool to assess products, thus building independent databases for future evaluation of its reliability, validity and sensitivity. Simone Borsci, Alessio Malizia, Martin Schmettow, Frank van der Velde, Gunay Tariverdiyeva, Divyaa Balaji, Alan Chamberlain |
Pers. Ubiquitous Comput. | 3 |
| 2017 | An extended protocol for usability validation of medical devices: Research design and reference model
Martin Schmettow, Raphaela Schnittker, Jan Maarten Schraagen |
J. Biomed. Informatics | 1 |
| 2016 | Linking card sorting to browsing performance - are congruent municipal websites more efficient to use?abstractCard sorting is a method for eliciting mental models and is frequently used for creating efficient website navigation structures. The present studies set out to validate card sorting by linking browsing performance to the degree of match between the mental model and the navigation structure. First, a card sorting study was conducted (n = 27) to elicit users’ mental model of municipal websites. Second, performance was measured for a number of search tasks with varying degrees of congruence with users’ mental model (n = 50). Analysis by linear mixed-effect models suggests that the match between mental model and website structure has no effect on browsing performance. We discuss possible reasons and consequences of the failure to validate card sorting for designing navigation structures of informational websites. Martin Schmettow, Jan Sommer |
Behav. Inf. Technol. | 1 |
| 2015 | A Semantic Map for Evaluating Creativity
Frank van der Velde, Roger A. Wolf, Martin Schmettow, Deniece S. Nazareth |
ICCC | 3 |
| 2015 | Tutorial: Modern Regression Techniques for HCI Researchers
Martin Schmettow |
INTERACT (4) | 1 |
| 2014 | Is an accelerating robot perceived as energetic or as gaining in speed?abstractPrevious studies have found that basic movement characteristics of a robot influence the emotional attributes people perceive independent of the embodiment of the motion (e.g. iCat vs. Roomba). Here, with a very simple LEGO robot, we replicate these associations between levels of acceleration and curvature and the extent to which positive and negative emotions are attributed. Importantly, we also show that these associations might not be valid. Prior to the emotional questionnaires participants were asked neutral questions on what they deemed relevant observations pertaining to the different robot motions. Only 3% of the remarks coincided with the emotional terms found in the questionnaires. HRI researchers interested in what people attribute to robot motion should be mindful of participant heuristics and experimenter biases. We provide some suggestions how to create experiments that are robust against these biases. Matthijs L. Noordzij, Martin Schmettow, Melle R. Lorijn |
HRI | 2 |
| 2013 | With how many users should you test a medical infusion pump? Sampling strategies for usability tests on high-risk systems
Martin Schmettow, Wendy Vos, Jan Maarten Schraagen |
J. Biomed. Informatics | 1 |
| 2008 | Introducing item response theory for measuring usability inspection processesabstractUsability evaluation methods have a long history of research. Latest contributions significantly raised the validity of method evaluation studies. But there is still a measurement model lacking that incorporates the relevant factors for inspection performance and accounts for the probabilistic nature of the process. This paper transfers a modern probabilistic approach from psychometric research, known as the Item Response Theory, to the domain of measuring usability evaluation processes. The basic concepts, assumptions and several advanced procedures are introduced and related to the domain of usability inspection. The practical use of the approach is exemplified in three scenarios from research and practice. These are also made available as simulation programs. Martin Schmettow, Wolfgang Vietze |
CHI | 1 |