Niels van Berkel

dblp:139/3792 · DBLP profile ↗
← Back
93ranked-venue papers
20as first author
67since 2021 · last 2026
0000-0001-5106-7692ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 80 · 17 first-author · 57 since 2021Software engineering, systems software and programming languages · 5 · 5 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Systems, architecture and hardware · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Who Gets to Interpret the Workout? User Tensions With AI-Generated Fitness Feedback
abstract
Fitness tracking platforms increasingly integrate generative AI to interpret activity data, such as Strava’s Athlete Intelligence. These integrations raise questions about how athletes engage with AI-supported fitness self-tracking. We analyzed 297 Reddit threads and 5,692 comments from r/Strava following the company’s launch of AI features to examine user reactions to AI-generated fitness feedback. Our findings revealed four recurring tensions: (1) numerical evaluation versus contextual understanding; (2) isolated session summaries versus ongoing training narratives; (3) a fixed AI tone versus diverse emotional states; and (4) a single AI voice versus different athletic types. Across these tensions, users resisted AI feedback that constrained interpretations of their own lived experiences. These findings shed light on the implicit challenges of integrating AI into self-tracking platforms. We conclude with implications for the design of AI-supported self-tracking systems that preserve interpretive openness and user agency.
Sujay Shalawadi, Joel Wester, Samuel Rhys Cox, Niels van Berkel
DIS4
2026 Polite But Boring? Trade-offs Between Engagement and Psychological Reactance to Chatbot Feedback Styles
abstract
As conversational agents become increasingly common in behaviour change interventions, understanding optimal feedback delivery mechanisms becomes increasingly important. However, choosing a style that both lessens psychological reactance (perceived threats to freedom) while simultaneously eliciting feelings of surprise and engagement represents a complex design problem. We explored how three different feedback styles: Direct, Politeness, and Verbal Leakage (slips or disfluencies to reveal a desired behaviour) affect user perceptions and behavioural intentions. Matching expectations from literature, the Direct chatbot led to lower behavioural intentions and higher reactance, while the Politeness chatbot evoked higher behavioural intentions and lower reactance. However, Politeness was also seen as unsurprising and unengaging by participants. In contrast, Verbal Leakage evoked reactance, yet also elicited higher feelings of surprise, engagement, and humour. These findings highlight that effective feedback requires navigating trade-offs between user reactance and engagement, with novel approaches such as Verbal Leakage offering promising alternative design opportunities.
Samuel Rhys Cox, Joel Wester, Niels van Berkel
CHI3
2026 Chaplains' Reflections on the Design and Usage of AI for Conversational Care
abstract
Despite growing recognition that responsible AI requires domain knowledge, current work on conversational AI primarily draws on clinical expertise that prioritises diagnosis and intervention. However, much of everyday emotional support needs occur in non-clinical contexts, and therefore requires different conversational approaches. We examine how chaplains, who guide individuals through personal crises, grief, and reflection, perceive and engage with conversational AI. We recruited eighteen chaplains to build AI chatbots. While some chaplains viewed chatbots with cautious optimism, the majority expressed limitations of chatbots’ ability to support everyday well-being. Our analysis reveals how chaplains perceive their pastoral care duties and areas where AI chatbots fall short, along the themes of Listening, Connecting, Carrying, and Wanting. These themes resonate with the idea of attunement, recently highlighted as a relational lens for understanding the delicate experiences care technologies provide. This perspective informs chatbot design aimed at supporting well-being in non-clinical contexts.
Joel Wester, Samuel Rhys Cox, Henning Pohl, Niels van Berkel
CHI4
2026 Crowd-Powered Discovery of Mental Health Self-Care Techniques in Higher Education
abstract
Mental disorders deteriorate well-being in many ways. Clinical mental health care is insufficient and struggles to keep up with the growing demand for services in many parts of the world. Human–Computer Interaction researchers have focused on building different interactive systems that support mental well-being. In our work, we focus on collecting, peer-assessing, and assisting in the discovery of mental health self-care techniques through an online tool. Our target demographic is the higher education community. The implemented tool can offer new self-care techniques to its users, and our user studies validate its usefulness in capturing and offering valuable mental health care information. Further, we study how source disclosure affects the perception of the discovered techniques, finding no major differences in the measured variables. Finally, we discuss how user-generated self-care suggestions in this context warrants caution and the implications of our study.
Andy Alorwu, Niels van Berkel, Joonas Moilanen, Parsa Sharmila, Niloofar Meftahi, Koji Yatani, Aku Visuri, Simo Hosio
ACM Trans. Comput. Hum. Interact.2
2025 Prompt Machine: A Tangible Generative AI Tool for Supporting Children's Learning and Literacy
abstract
Figure 1: The Prompt Machine, a tangible learning tool for integrated AI in education.(A) Pupils write assignments.(B) Pupils input their written assignments.(C) Pupils modify their texts using tangible prompt cubes.(D) Pupils receive physical print outs of modified texts.(E) Teachers facilitate reflections and discussions about modified texts with pupils.
Martin V. A. Lindrup, Rune Møberg Jacobsen, Joel Wester, Niels van Berkel, Dimitrios Raptis, Peter Axel Nielsen
Conference on Designing Interactive Systems4
2025 General Practitioners' Perspectives on a Pre-Consultation Chatbot for Shared Decision-Making
abstract
General practitioner (GP) consultations are the typical starting point for a patient's healthcare journey.Here, GPs aim to support and inform patients to enable a shared decision-making process.In this work we explore how an interactive chatbot, designed to prepare patients for their GP consultation, is perceived by GPs to impact patient consultations, patient-GP interaction, and their work.We conducted an in-depth evaluation and interview with 15 GPs from 12 different practices.Our findings provide insights into common challenges in shared decision-making, GP perspectives on the role of chatbots in preparing patients, and how chatbot technology could impact and transform general practice.Finally, we reflect on patient and GP agency in shared decision-making and the impact of technology on this complex relationship.
Mana Samiee, Joel Wester, Rune Møberg Jacobsen, Michael Skovdal Rathleff, Niels van Berkel
Conference on Designing Interactive Systems5
2025 Beyond Productivity: Rethinking the Impact of Creativity Support Tools
abstract
Figure 1: Measures used in (n=173) empirical studies of CSTs from a survey of 10 years of ACM DL literature.Measures of User Experience with CSTs were most prevalent (90%), followed by measures of Creative Artefact Quality (54%), and measures of User-Centric Benefits were least prevalent (15%).
Samuel Rhys Cox, Helena Bøjer Djernæs, Niels van Berkel
Creativity & Cognition3
2025 Coordination Mechanisms in AI Development: Practitioner Experiences on Integrating UX Activities
abstract
Software development relies on collaboration and alignment between a variety of roles, including software developers and user experience designers. The increasing focus on artificial intelligence in today's development projects has given rise to new challenges in this collaboration. We extend previous work on the process of designing human-AI systems by analysing collaborative practices between UX designers and AI developers through Mintzberg's theory on coordination mechanisms. We conducted 15 in-depth interviews with UX designers and AI developers currently working on AI projects. We contribute by identifying how coordination mechanisms impact the UX design process when developing AI systems , inter-team (a)symmetries in power relations, and a growing need for tools and cross-disciplinary knowledge to support these collaborative efforts. In particular, we outline the risks of coordinating AI development work through the standardisation of output and skills in separately organised UX and AI development teams. CCS Concepts • Software and its engineering → Collaboration in software development; • Human-centered computing → Collaborative and social computing.
Anders Bruun, Niels van Berkel, Dimitrios Raptis, Effie Lai-Chong Law
CHI2
2025 Visual Augmentations for Ultrasound Assessment Training of Medical Students
abstract
Ultrasound assessments are key in assessing traumatic injuries to the human body during urgent medical emergencies. Obtaining proficiency in conducting ultrasound assessments is challenging, and relies on hands-on, individually instructed training provided by a scarce number of ultrasound experts. We investigate how to support medical students’ learning of ultrasound assessment through visual augmentations. By enhancing the learning process, we seek to support medical students in reaching higher proficiency in ultrasound assessments. We followed an ultrasound assessment course to identify the primary challenges faced by medical students learning to conduct ultrasound assessments. Based on our findings, we designed four distinct visual augmentations in collaboration with a course educator that guide students in achieving better ultrasound image quality. We evaluated these visual augmentations in a mixed-method study with 15 medical students. Our findings provide insights on the use of digital technology in supporting clinical training, and the possibilities of bridging existing training practices.
Helena Bøjer Djernæs, Rune Møberg Jacobsen, Simo Hosio, Niels van Berkel
CHI4
2025 Chatbots for Data Collection in Surveys: A Comparison of Four Theory-Based Interview Probes
abstract
Surveys are a widespread method for collecting data at scale, but their rigid structure often limits the depth of qualitative insights obtained. While interviews naturally yield richer responses, they are challenging to conduct across diverse locations and large participant pools. To partially bridge this gap, we investigate the potential of using LLM-based chatbots to support qualitative data collection through interview probes embedded in surveys. We assess four theory-based interview probes: descriptive, idiographic, clarifying, and explanatory. Through a split-plot study design (N=64), we compare the probes' impact on response quality and user experience across three key stages of HCI research: exploration, requirements gathering, and evaluation. Our results show that probes facilitate the collection of high-quality survey data, with specific probes proving effective at different research stages. We contribute practical and methodological implications for using chatbots as research tools to enrich qualitative data collection.
Rune Møberg Jacobsen, Samuel Rhys Cox, Carla F. Griggio, Niels van Berkel
CHI4
2025 Challenging Futures: Using Chatbots to Reflect on Aging and Dementia
abstract
Intertemporal reflection, flexibly thinking forward and backward in time, is vital for one's future planning.Yet, cultivating intertemporal reflection about encountering difficult futures, e.g., developing a progressive cognitive condition like dementia, can be challenging.We assessed people's attitudes towards dementia following conversing with a chatbot presented as either neurotypical or simulating dementia symptoms.While neither the chatbot's presentation nor the framing of participants' future selves impacted attitudes toward dementia, it influenced participants' experiences.When framed as future selves, the chatbot evoked a strong emotional connection, leading to reflection on aging, particularly with the chatbot simulating dementia symptoms.Participants interacting with the chatbot framed as a stranger with simulated symptoms often felt frustrated, especially when they had a task-oriented mindset.Chatbots can be promising tools for prompting reflections on challenging futures, such as dementia, although their effectiveness varies due to the tensions between simulated cognitive decline and expectations for effective communication.
Rucha Khot, Teis Arets, Joel Wester, Franziska Burger, Niels van Berkel, Rens Brankaert, Wijnand A. IJsselsteijn, Minha Lee
CHI5
2025 Enhancing Self-Efficacy in Health Self-Examination through Conversational Agent's Encouragement
abstract
Health self-examination, such as checking for changes to skin moles, is key to identifying potential negative changes to one's body. A major barrier to initiating a self-examination is a perceived lack of confidence or knowledge. In this study, we use a 2 × 2 between-subjects design to evaluate the effect of an AI conversational agent (CA) on participant self-efficacy and trust. We manipulated both participants' perceived skill in self-examination (based on prior perceived Success vs. Failure) and the CA's verbal persuasions (Encouraging vs. Neutral), with participants asked to complete a series of skin self-assessment tasks. Our findings show that participants' self-efficacy increased when exposed to encouraging CA persuasion. Additionally, we observed that an encouraging CA significantly increased participants' trust scores in perceived benevolence compared to a neutral-sounding CA. Our results inform the design of CAs to support users' independent self-examination.
Naja Kathrine Kollerup Als, Maria-Theresa Bahodi, Samuel Rhys Cox, Niels van Berkel
CHI4
2025 Before It Falls: Supporting Drone Fleet Management Through Battery Visualizations
Maria-Theresa Bahodi, Nathan Lau, Niels van Berkel, Kasper Andreas Rømer Grøntved, Mikael B. Skov, Timothy Merritt 0001
INTERACT (2)3
2025 Clinical needs and preferences for AI-based explanations in clinical simulation training
abstract
Medical training is a key element in maintaining and improving today's healthcare standards. Given the nature of medical work, students must master not only theory but also develop their hands-on abilities and skills in clinical practice. Medical simulators play an increasing role in supporting the active learning of these students due to their ability to present a large variety of tasks allowing students to train and experiment indefinitely without causing any patient harm. While the criticality of explainable AI systems has been extensively discussed in the literature, the medical training context presents unique user needs for explanations. In this paper, we explore the potential gap of current limitations within simulation-based training, and the role Artificial Intelligence (AI) holds in supporting the needs of medical students in training. Through contextual inquiries and interviews with clinicians in training (N = 9) and subsequent validation with medical experts (N = 4), we obtain an understanding of the shortcomings in current simulation-based training and offer recommendations for future AI-driven training. Our results stress the need for continuous and actionable feedback that resembles the interaction between clinical supervisor and resident in real-world training scenarios while adjusting training material to the residents' skills and prior performance.
Naja Kathrine Kollerup Als, Stine S. Johansen, Martin Grønnebæk Tolsgaard, Mikkel Lønborg Friis, Mikael B. Skov, Niels van Berkel
Behav. Inf. Technol.6
2025 Using LLMs for self-care: User and counsellor perspectives
abstract
People are increasingly relying on technology for self-care, including, more recently, seeking help through conversational interfaces driven by large language models (LLMs). Yet, how interaction with LLMs has impacted people’s self-care processes is not well understood. Therefore, we collected 405 user stories posted on Reddit about using LLMs for self-care. We identified four key themes on how people use LLMs for this purpose: Letting go , Finding comfort , Building up , and Reflecting on . We interviewed twelve counsellors to capture their perspectives on this practice, given their professional expertise and understanding of healthy self-care practices. Our results show that counsellors recognised several benefits, such as using LLMs as stepping stones or springboards towards improved self-care. They also highlighted several areas of concern, such as unintended consequences that might negatively affect users. We discuss the dissonance around how the early adopters of LLMs appropriate this technology to care for themselves, how counsellors see such usage, and outline implications of using LLMs as a technology for self-care. • We present an analysis of 405 LLM self-care stories. • We used stories representative of themes in interviews with diverse counsellors. • We discuss user and counsellor perspectives on LLMs as self-care technologies.
Joel Wester, Sander de Jong, Henning Pohl, Niels van Berkel
Int. J. Hum. Comput. Stud.4
2025 From Reflection to Action: Enhancing Workplace Well-Being Through Digital Solutions
abstract
Abstract Despite the widely acknowledged importance of well-being, our well-being can regularly be under pressure from external sources. Work is often attributed as a source of stress and dissatisfaction, so, unsurprisingly, extensive efforts are made to measure and improve our well-being in this context. This paper examines opportunities to better design supportive digital solutions through two complementary studies. In the first study, we present a longitudinal assessment of a well-being-focused self-report application deployed in two organizations. Through an analysis of one year of application usage across 219 users, we find both established and novel patterns of application usage and well-being evaluation. While prior work has highlighted substantial dropout rates and daily well-being fluctuations that peak in the morning and early evening, our results highlight that substantial breaks in usage are common, suggesting that users choose to engage with well-being applications mainly when they need them. In the second study, we expand on the topic of well-being reflection at work and the use of technology for this purpose. Through a survey involving 100 participants, we identify current practices in increasing well-being at work, obstacles to sharing and discussing mental well-being states, opportunities for digital well-being solutions and reflections on transparency and communication. Our combined results highlight opportunities for HCI research and practice to address the ongoing challenges of maintaining well-being in today’s work environments.
Niels van Berkel, Aku Visuri, Sujay Shalawadi, Madeleine R. Evans, Benjamin Tag, Simo Hosio
Interact. Comput.1
2025 Impact of Explanation Techniques and Representations on Users' Comprehension and Confidence in Explainable AI
abstract
Local explainability, an important sub-field of eXplainable AI, focuses on describing the decisions of AI models for individual use cases by providing the underlying relationships between a model's inputs and outputs. While the machine learning community has made substantial progress in improving explanation accuracy and completeness, these explanations are rarely evaluated by the final users. In this paper, we evaluate the impact of various explanation and representation techniques on users' comprehension and confidence. Through a user study on two different domains, we assessed three commonly used local explanation techniques—feature-attribution, rule-based, and counterfactual—and explored how their visual representation—graphical or text-based—influences users' comprehension and trust. Our results show that the choice of explanation technique primarily affects user comprehension, whereas the graphical representation impacts user confidence.
Julien Delaunay, Luis Galárraga, Christine Largouët, Niels van Berkel
Proc. ACM Hum. Comput. Interact.4
2025 Cognitive Forcing for Better Decision-Making: Reducing Overreliance on AI Systems Through Partial Explanations
abstract
In AI-assisted decision-making, explanations aim to enhance transparency and user trust but can also lead to negligence. In two separate studies, we explore the use of partial explanations to activate cognitive forcing and increase user engagement. In Study I (N = 264), we present participants with weighted graphs and ask them to identify the shortest paths. In Study II (N = 210), participants correct spelling and grammar mistakes in short text segments. In both studies, we provide a solution suggestion accompanied by either no explanation, a full explanation, or a partial explanation. Our results show that partial explanations reduce overreliance on incorrect AI suggestions, performing significantly better than the baseline but not as well as full explanations. Individuals with a high need for cognition benefit more from AI explanations and consequently perform better. Our work suggests that partial explanations can be valuable in domains where reducing overreliance on AI is critical, like medical diagnosis. It also underscores the need to consider explanation effectiveness across different task difficulties, a factor often overlooked in contemporary human-AI studies.
Sander de Jong, Ville Paananen, Benjamin Tag, Niels van Berkel
Proc. ACM Hum. Comput. Interact.4
2025 Cognitive performance measurements and the impact of sleep quality using wearable and mobile sensors
abstract
Abstract Human cognitive performance affects a wide range of aspects of our daily lives. Numerous factors influence our cognitive performance, and cognitive performance in turn impacts our capabilities. Partial sleep deprivation in particular negatively affects vigilance, a key factor in many work tasks. Sleep in general plays a large role in physiological recovery and our capability to perform mental tasks. In this work, we focus on two research questions. First, we investigate how fluctuations in sleep quality influence cognitive vigilance. Second, we study how smartphone typing can be leveraged as a continuous measurement for cognitive vigilance and can thus be an indicator of decline in cognitive capabilities and sleep quality. We report on a 2-month field study in which we collected cognitive performance data using the Psychomotor Vigilance Task (PVT), mobile keyboard typing metrics from participants’ personal smartphones, and sleep quality metrics through a wearable sleep-tracking ring. Our findings highlight that individual sleep metrics such as night-time heart rate, sleep latency, sleep timing, sleep restfulness, and overall sleep quantity significantly influence vigilance. Long sleep latencies can reduce reaction times up to 30 ms, abnormal sleep durations up to 20 ms, and night-time awake time up to 10 ms. Heart rate is a well-known indicator of recovery quality, and improvements in both heart rate and heart rate variability (HRV) show positive variations of 15–20 ms in reaction test performance. To expand the current research on cognitive computing, we introduce smartphone typing metrics as a proxy or a complementary method for continuous passive measurement of cognitive vigilance and report on statistically significant correlations in PVT performance and typing speed and error rates. Together, our findings contribute to ubiquitous computing via a longitudinal case study with a novel wearable device, the resulting findings on the association between sleep and cognitive function, and the introduction of smartphone keyboard typing as a proxy of cognitive function.
Aku Visuri, Heli Koskimäki, Niels van Berkel, Andy Alorwu, Ella Peltonen, Saeed Abdullah, Simo Hosio
Pers. Ubiquitous Comput.3
2024 How Can I Signal You To Trust Me: Investigating AI Trust Signalling in Clinical Self-Assessments
abstract
Individuals are increasingly interested in and responsible for assessing their own health. This study evaluates a fictional AI dermatologist for assistance in the self-assessment of moles. Building on the Signalling Theory, we tested the effect of textual descriptions provided by a virtual dermatologist, as manipulated across ‘Ability’, ‘Integrity, ’ and ‘Benevolence’, along with the clinical assessment, ‘benign’ or ‘malignant’, affect users’ trust in the aforementioned trust pillars. Our study (N = 40) follows a 2 (Ability low/high) × 2 (Integrity low/high) × 2 (Benevolence low/high) × 2 (mole assessment benign/malignant) within-subject factorial design. Our results demonstrate that we can successfully influence perceptions of ability and benevolence by manipulating the corresponding aspects of trust but not perceived integrity. Further, in the case of a malignant assessment, participants’ perception of trust increased across all aspects. Our results provide insights into the design of AI support systems for sensitive use cases, such as clinical self-assessments.
Naja Kathrine Kollerup Als, Joel Wester, Mikael B. Skov, Niels van Berkel
Conference on Designing Interactive Systems4
2024 Manual, Hybrid, and Automatic Privacy Covers for Smart Home Cameras
abstract
Smart home cameras (SHCs) offer convenience and security to users, but also cause greater privacy concerns than other sensors due to constant collection and processing of sensitive data. Moreover, privacy perceptions may differ between primary users and other users at home. To address these issues, we developed three physical cover prototypes for SHCs: Manual, Hybrid, and Automatic, based on design criteria of observability, understandability, and tangibility. With 90 SHC users, we ran an online survey using video vignettes of the prototypes. We evaluated how the physical covers alleviated privacy concerns by measuring perceived creepiness and trustworthiness. Our results show that the physical covers were well received, even though primary SHC users valued always-on surveillance. We advocate for the integration of physical covers into future SHC designs, emphasizing their potential to establish a shared understanding of surveillance status. Additionally, we provide design recommendations to support this proposition.
Sujay Shalawadi, Christopher Getschmann, Niels van Berkel, Florian Echtler
Conference on Designing Interactive Systems3
2024 "As an AI language model, I cannot": Investigating LLM Denials of User Requests
abstract
Users ask large language models (LLMs) to help with their homework, for lifestyle advice, or for support in making challenging decisions. Yet LLMs are often unable to fulfil these requests, either as a result of their technical inabilities or policies restricting their responses. To investigate the effect of LLMs denying user requests, we evaluate participants’ perceptions of different denial styles. We compare specific denial styles (baseline, factual, diverting, and opinionated) across two studies, respectively focusing on LLM’s technical limitations and their social policy restrictions. Our results indicate significant differences in users’ perceptions of the denials between the denial styles. The baseline denial, which provided participants with brief denials without any motivation, was rated significantly higher on frustration and significantly lower on usefulness, appropriateness, and relevance. In contrast, we found that participants generally appreciated the diverting denial style. We provide design recommendations for LLM denials that better meet peoples’ denial expectations.
Joel Wester, Tim Schrills, Henning Pohl, Niels van Berkel
CHI4
2024 Assessing Cognitive and Social Awareness among Group Members in AI-assisted Collaboration
abstract
Successful collaboration in computer-mediated teams requires awareness among group members of each other’s knowledge, skills, and goals. Large Language Models (LLMs) can play a mediating role in establishing and maintaining this awareness among group members. In an in-situ study, we explored the impact of an LLM-based chatbot on cognitive and social group awareness through a distributed text-based group task. We instructed participants (N = 48) to complete a travel-planning task in sixteen groups of three, with each member given conflicting goals. Each chat was complemented by a chatbot that could be asked for assistance. Through a survey and semi-structured interview, we gained insight into participants’ deliberations on the task and the chatbot’s role. We found that the chatbot’s presence helped increase group awareness as users are forced to clearly and transparently formulate their intentions when prompting the chatbot. The chatbot’s ability to provide suggestions that compromise between user goals based on the chat history helped participants reach a consensus. We present implications for the design of chatbots for collaborative settings.
Sander de Jong, Joel Wester, Tim Schrills, Kristina Skjødt Secher, Carla F. Griggio, Niels van Berkel
MUM6
2024 From Voice to Value: Leveraging AI to Enhance Spoken Online Reviews on the Go
abstract
Online reviews help people make better decisions. Review platforms usually depend on typed input, where leaving a good review requires significant effort because users must carefully organize and articulate their thoughts. This may discourage users from leaving comprehensive and high-quality reviews, especially when they are on the go. To address this challenge, we developed Vocalizer, a mobile application that enables users to provide reviews through voice input, with enhancements from a large language model (LLM). In a longitudinal study, we analysed user interactions with the app, focusing on AI-driven features that help refine and improve reviews. Our findings show that users frequently utilized the AI agent to add more detailed information to their reviews. We also show how interactive AI features can improve users self-efficacy and willingness to share reviews online. Finally, we discuss the opportunities and challenges of integrating AI assistance into review-writing systems.
Kavindu Ravishan, Dániel Szabó, Niels van Berkel, Aku Visuri, Chi-Lan Yang, Koji Yatani, Simo Hosio
MUM3
2024 Monetary valuation of personal health data in the wild
abstract
The value of personal health data continues to be a debated topic in HCI and society more broadly. We investigate the monetary value people attach to their health data. Using a custom mobile app for 14 days with 55 participants, we collected health data (sleep duration, sleep quality, pain intensity, wake-up times) and a daily monetary data valuation using a reverse second-price auction. Participants bid to sell their data to a for-profit company, the government, or academia. Our findings indicate that people value their data differently based on who is buying. We also show that people are interested in monetizing their personal health data despite privacy and data protection concerns. The presented study helps us understand the data value landscape and paves way to a healthier data-driven future where people may benefit more from their own contributions, either in monetary or other forms.
Andy Alorwu, Niels van Berkel, Aku Visuri, Sharadhi Alape Suryanarayana, Takuya Yoshihiro, Simo Hosio
Int. J. Hum. Comput. Stud.2
2024 Impact of interaction technique in interactive data visualisations: A study on lookup, comparison, and relation-seeking tasks
abstract
This paper presents an analysis of different interaction techniques used in interactive data visualisations to support end-users in visual analytics tasks. Our selection of interaction techniques is based on prior work and consists of the interaction techniques SELECT, EXPLORE, RECONFIGURE, ENCODE, FILTER, ABSTRACT/ELABORATE, and CONNECT. Through a within-subject study, we assessed participants’ abilities to utilise these techniques when faced with three distinct types of data-driven tasks; lookup, comparison, and Relation-seeking. Our research investigates the impact of these interaction techniques on the correctness, confidence, perceived difficulty, and cognitive load of N = 80 self-identified data scientists and N = 80 non-experts. We find that interaction technique significantly impacts answer correctness and participant confidence. Participants performed best across those interaction techniques that allow for information that is deemed least relevant to be concealed, which is reflected in lower intrinsic and extraneous cognitive load. Interestingly, participants’ expertise affected their confidence but not their accuracy. Our results provide insights useful for a more targeted and informed design and usage of interactive data visualisations.
Niels van Berkel, Benjamin Tag, Rune Møberg Jacobsen, Daniel Russo 0002, Helen C. Purchase, Daniel Buschek
Int. J. Hum. Comput. Stud.1
2024 Generative AI in Software Engineering Must Be Human-Centered: The Copenhagen Manifesto
Daniel Russo 0002, Sebastian Baltes, Niels van Berkel, Paris Avgeriou, Fabio Calefato, Beatriz Cabrero-Daniel, Gemma Catolino, Jürgen Cito, Neil A. Ernst, Thomas Fritz 0001, Hideaki Hata, Reid Holmes, Maliheh Izadi, Foutse Khomh, Mikkel Baun Kjærgaard, Grischa Liebel, Alberto Lluch-Lafuente, Stefano Lambiase, Walid Maalej, Gail C. Murphy, Nils Brede Moe, Gabrielle O'Brien, Elda Paja, Mauro Pezzè, John Stouby Persson, Rafael Prikladnicki, Paul Ralph, Martin P. Robillard, Thiago Rocha Silva, Klaas-Jan Stol, Margaret-Anne D. Storey, Viktoria Stray, Paolo Tell, Christoph Treude, Bogdan Vasilescu
J. Syst. Softw.3
2024 PACMHCI, VI, MHCI, September 2024 Editorial
abstract
Welcome to this issue of the Proceedings of the ACM on Human-Computer Interaction, which brings together contributions from the Mobile Human-Computer Interaction (MHCI) community. This issue showcases innovations in research focused on mobile, wearable, and personal devices. Research in MHCI encompasses both technical innovations and social considerations. Mobile technologies accompany us throughout our daily lives: at home, at work, in traffic, and out in the wild. However, while mobile technologies provide the capacity for constant connectivity, they also pose risks of harm and unintended consequences. There has never been a more pressing time to debate and explore the meaning of digital culture, and to understand how mobile technologies can and should be used to connect us, our data, and mobile applications and services in meaningful ways - true to the conference theme of 'Connecting Cultures'.
Marion Koelle, Niels van Berkel, Jenny Waycott
Proc. ACM Hum. Comput. Interact.2
2024 Effect of Explanation Conceptualisations on Trust in AI-assisted Credibility Assessment
abstract
As misinformation increasingly proliferates on social media platforms, it has become crucial to explore how to best convey automated news credibility assessments to end-users, and foster trust in fact-checking AIs. In this paper, we investigate how model-agnostic, natural language explanations influence trust and reliance on a fact-checking AI. We construct explanations from four Conceptualisation Validations (CVs) - namely consensual, expert, internal (logical), and empirical - which are foundational units of evidence that humans utilise to validate and accept new information. Our results show that providing explanations significantly enhances trust in AI, even in a fact-checking context where influencing pre-existing beliefs is often challenging, with different CVs causing varying degrees of reliance. We find consensual explanations to be the least influential, with expert, internal, and empirical explanations exerting twice as much influence. However, we also find that users could not discern whether the AI directed them towards the truth, highlighting the dual nature of explanations to both guide and potentially mislead. Further, we uncover the presence of automation bias and aversion during collaborative fact-checking, indicating how users' previously established trust in AI can moderate their reliance on AI judgements. We also observe the manifestation of a 'boomerang'/backfire effect often seen in traditional corrections to misinformation, with individuals who perceive AI as biased or untrustworthy doubling down and reinforcing their existing (in)correct beliefs when challenged by the AI. We conclude by presenting nuanced insights into the dynamics of user behaviour during AI-based fact-checking, offering important lessons for social media platforms.
Saumya Pareek, Niels van Berkel, Eduardo Velloso, Jorge Gonçalves 0001
Proc. ACM Hum. Comput. Interact.2
2024 Facing LLMs: Robot Communication Styles in Mediating Health Information between Parents and Young Adults
abstract
Young adults may feel embarrassed when disclosing sensitive information to their parents, while parents might similarly avoid sharing sensitive aspects of their lives with their children. How to design interactive interventions that are sensitive to the needs of both younger and older family members in mediating sensitive information remains an open question. In this paper, we explore the integration of large language models (LLMs) with social robots. Specifically, we use GPT-4 to adapt different Robot Communication Styles (RCS) for a social robot mediator designed to elicit self-disclosure and mediate health information between parents and young adults living apart. We design and compare four literature-informed RCS: three LLM-adapted (Humorous, Self-deprecating, and Persuasive) and one manually created (Human-scripted), and assess participant perceptions of Likeability, Usefulness, Helpfulness, Relatedness, and Interpersonal Closeness . Through an online experiment with 183 participants, we assess the RCS across two groups: adults with children (Parents) and young adults without children (Young Adults). Our results indicate that both Parents and Young Adults favoured the Human-scripted and Self-deprecating RCS as compared to the other two RCS. The Self-deprecating RCS furthermore led to increased relatedness as compared to the Humorous RCS. Our qualitative findings reveal challenges people have in disclosing health information to family members, and who normally assumes the role of family facilitator-two areas in which social robots can play a key role. The findings offer insights for integrating LLMs with social robots in health-mediation and other contexts involving the sharing of sensitive information.
Joel Wester, Bhakti Moghe, Katie Winkle, Niels van Berkel
Proc. ACM Hum. Comput. Interact.4
2024 "This Chatbot Would Never...": Perceived Moral Agency of Mental Health Chatbots
abstract
Despite repeated reports of socially inappropriate and dangerous chatbot behaviour, chatbots are increasingly used as mental health services in providing support for young people. In sensitive settings as such, the notion of perceived moral agency (PMA) is crucial, given its critical role in human-human interactions. In this paper, we investigate the role of PMA in human-chatbot interactions. Specifically, we seek to understand how PMA influence the perception of trust, likeability, and perceived safety of chatbots for mental health across two distinct age groups. We conduct an online experiment(N = 279)to evaluate chatbots with low and high PMA as targeted towards teenagers and adults. Our results indicate increased trust, likeability, and perceived safety in mental health chatbots displaying high PMA. A qualitative analysis revealed four themes, assessing participants' expectations of mental health chatbots in general, as well as targeted towards teenagers: Anthropomorphism, Warmth, Sensitivity, and Appearance manifestation. We show that PMA plays a crucial role in influencing the perceptions of chatbots and provide recommendations for designing socially appropriate mental health chatbots.
Joel Wester, Henning Pohl, Simo Hosio, Niels van Berkel
Proc. ACM Hum. Comput. Interact.4
2024 Collaborating with Bots and Automation on OpenStreetMap
abstract
OpenStreetMap (OSM) is a large online community where users collaborate to map the world. In addition to manual edits, the OSM mapping database is regularly modified by bots and automated edits. In this article, we seek to better understand how people and bots interact and conflict with each other. We start by analysing over 15 years of mailing list discussions related to bots and automated edits. From this data, we uncover five themes, including how automation results in power differentials between users and how community ideals of consensus clash with the realities of bot use. Subsequently, we surveyed OSM contributors on their experiences with bots and automated edits. We present findings about the current escalation and review mechanisms, as well as the lack of appropriate tools for evaluating and discussing bots. We discuss how OSM and similar communities could use these findings to better support collaboration between humans and bots.
Niels van Berkel, Henning Pohl
ACM Trans. Comput. Hum. Interact.1
2024 Understanding Developers Well-Being and Productivity: A 2-year Longitudinal Analysis during the COVID-19 Pandemic
abstract
The COVID-19 pandemic has brought significant and enduring shifts in various aspects of life, including increased flexibility in work arrangements. In a longitudinal study, spanning 24 months with six measurement points from April 2020 to April 2022, we explore changes in well-being, productivity, social contacts, and needs of software engineers during this time. Our findings indicate systematic changes in various variables. For example, well-being and quality of social contacts increased while emotional loneliness decreased as lockdown measures were relaxed. Conversely, people’s boredom and productivity remained stable. Furthermore, a preliminary investigation into the future of work at the end of the pandemic revealed a consensus among developers for a preference of hybrid work arrangements. We also discovered that prior job changes and low job satisfaction were consistently linked to intentions to change jobs if current work conditions do not meet developers’ needs. This highlights the need for software organizations to adapt to various work arrangements to remain competitive employers. Building upon our findings and the existing literature, we introduce the Integrated Job Demands-Resources and Self-Determination (IJARS) Model as a comprehensive framework to explain the well-being and productivity of software engineers during the COVID-19 pandemic.
Daniel Russo 0002, Paul H. P. Hanel, Niels van Berkel
ACM Trans. Softw. Eng. Methodol.3
2024 Understanding Developers Well-being and Productivity: A 2-year Longitudinal Analysis during the COVID-19 Pandemic - RCR Report
abstract
The artifact accompanying the paper “Understanding Developers Well-Being and Productivity: A 2-year Longitudinal Analysis during the COVID-19 Pandemic” provides a comprehensive set of tools, data, and scripts that were utilized in the longitudinal study. Spanning 24 months, from April 2020 to April 2022, the study delves into the shifts in well-being, productivity, social contacts, needs, and several other variables of software engineers during the COVID-19 pandemic. The artifact facilitates the reproduction of the study’s findings, offering a deeper insight into the systematic changes observed in various variables, such as well-being, quality of social contacts, and emotional loneliness. By providing access to the evidence-generating mechanisms and the generated data, the artifact ensures transparency and reproducibility and allows researchers to use our rich dataset to test their own research question. This Replicated Computational Results report aims to detail the contents of the artifact, its relevance to the main paper, and guidelines for its effective utilization.
Daniel Russo 0002, Paul H. P. Hanel, Niels van Berkel
ACM Trans. Softw. Eng. Methodol.3
2023 "If I Had All the Time in the World": Ophthalmologists' Perceptions of Anchoring Bias Mitigation in Clinical AI Support
abstract
Clinical needs and technological advances have resulted in increased use of Artificial Intelligence (AI) in clinical decision support. However, such support can introduce new and amplify existing cognitive biases. Through contextual inquiry and interviews, we set out to understand the use of an existing AI support system by ophthalmologists. We identified concerns regarding anchoring bias and a misunderstanding of the AI’s capabilities. Following, we evaluated clinicians’ perceptions of three bias mitigation strategies as integrated into their existing decision support system. While clinicians recognised the danger of anchoring bias, we identified a concern around the impact of bias mitigation on procedure time. Our participants were divided in their expectations of any positive impact on diagnostic accuracy, stemming from varying reliance on the decision support. Our results provide insights into the challenges of integrating bias mitigation into AI decision support.
Anne Kathrine Petersen Bach, Trine Munch Nørgaard, Jens Christian Brok, Niels van Berkel
CHI4
2023 "You've Got a Friend in Me": A Formal Understanding of the Critical Friend Agent
abstract
State-of-the-art intelligent and interactive agents, such as Alexa or Siri, often present overly conforming behaviour during interactions with humans. This can result in a misalignment between end-user expectations and agent behaviour. To overcome this barrier in human-AI interactions, we introduce the Critical Friend (CF), a conceptual idea that guides critical behaviour in human-human interactions. We present our results as a formal understanding that can be described through description logic and utilised for reasoning capabilities, enabling implementations of the CF as an intelligent interactive agent.
Joel Wester, Andreas Brännström, Juan Carlos Nieves, Niels van Berkel
HAI4
2023 A Longitudinal Analysis of Real-World Self-report Data
Niels van Berkel, Sujay Shalawadi, Madeleine R. Evans, Aku Visuri, Simo Hosio
INTERACT (3)1
2023 A Review on Mood Assessment Using Smartphones
Zhanna Sarsenbayeva, Charlie Fleming, Benjamin Tag, Anusha Withana, Niels van Berkel, Alistair Lee McEwan
INTERACT (2)5
2023 4th Crowd Science Workshop - CANDLE: Collaboration of Humans and Learning Algorithms for Data Labeling
abstract
Crowdsourcing has been used to produce impactful and large-scale datasets for Machine Learning and Artificial Intelligence (AI), such as ImageNET, SuperGLUE, etc. Since the rise of crowdsourcing in early 2000s, the AI community has been studying its computational, system design, and data-centric aspects at various angles. We welcome the studies on developing and enhancing of crowdworker-centric tools, that offer task matching, requester assessment, instruction validation, among other topics. We are also interested in exploring methods that leverage the integration of crowdworkers to improve the recognition and performance of the machine learning models. Thus, we invite studies that focus on shipping active learning techniques, methods for joint learning from noisy data and from crowds, novel approaches for crowd-computer interaction, repetitive task automation, and role separation between humans and machines. Moreover, we invite works on designing and applying such techniques in various domains, including e-commerce and medicine.
Dmitry Ustalov, Saiph Savage, Niels van Berkel, Yang Liu 0018
WSDM3
2023 Satisfaction and performance of software developers during enforced work from home in the COVID-19 pandemic
abstract
Following the onset of the COVID-19 pandemic and subsequent lockdowns, the daily lives of software engineers were heavily disrupted as they were abruptly forced to work remotely from home. To better understand and contrast typical working days in this new reality with work in pre-pandemic times, we conducted one exploratory ( N = 192) and one confirmatory study ( N = 290) with software engineers recruited remotely. Specifically, we build on self-determination theory to evaluate whether and how specific activities are associated with software engineers’ satisfaction and productivity. To explore the subject domain, we first ran a two-wave longitudinal study. We found that the time software engineers spent on specific activities (e.g., coding, bugfixing, helping others) while working from home was similar to pre-pandemic times. Also, the amount of time developers spent on each activity was unrelated to their general well-being, perceived productivity, and other variables such as basic needs. Our confirmatory study found that activity-specific variables (e.g., how much autonomy software engineers had during coding) do predict activity satisfaction and productivity but not by activity-independent variables such as general resilience or a good work-life balance. Interestingly, we found that satisfaction and autonomy were significantly higher when software engineers were helping others and lower when they were bugfixing. Finally, we discuss implications for software engineers, management, and researchers. In particular, active company policies to support developers’ need for autonomy, relatedness, and competence appear particularly effective in a WFH context.
Daniel Russo 0002, Paul H. P. Hanel, Seraphina Altnickel, Niels van Berkel
Empir. Softw. Eng.4
2023 The methodology of studying fairness perceptions in Artificial Intelligence: Contrasting CHI and FAccT
abstract
The topic of algorithmic fairness is of increasing importance to the Human–Computer Interaction research community following accumulating concerns regarding the use and deployment of Artificial Intelligence-based systems. How we conduct research on algorithmic fairness directly influences our inferences and conclusions regarding algorithmic fairness. To better understand the methodological decisions of studies focused on people’s perceptions of algorithmic fairness, we systematic analysed relevant papers from the CHI and FAccT conferences. We identified 200 relevant papers published between 1993 and 2022 and assessed their study design, participant sample, and geographical location of participants and authors. Our results highlight that studies are predominantly cross-sectional, cover a wide range of participant roles, and that both authors and participants are primarily from the United States. Based on these findings, we reflect on the potential pitfalls and shortcomings in how the community studies algorithmic fairness.
Niels van Berkel, Zhanna Sarsenbayeva, Jorge Gonçalves 0001
Int. J. Hum. Comput. Stud.1
2023 Exploring crowdsourced self-care techniques: A study on Parkinson's disease
abstract
Living with Parkinson’s Disease introduces a range of significant challenges into one’s daily life. While medical interventions exist to overcome some of these challenges, patient self-care techniques often form an essential complement to the treatments recommended by medical doctors. Knowledge on these self-care techniques often originates from those living with Parkinson’s themselves or their close caregivers, as they have the knowledge and experience required to assess self-care techniques. This so-called ‘patient knowledge’ is usually exchanged in peer meetings or discussion forums. Although vital to the Parkinson’s Disease community, this information is often difficult to access due to its unstructured format and the difficulty of navigating through online forums. We present an online tool that allows for contributing, assessing, and finally discovering Parkinson’s Disease self-care techniques. The custom discovery tool was populated with self-care knowledge by over 300 people with Parkinson’s and dozens of their carers, spanning areas such as daily well-being and using assistive equipment. Then, we invited patients to explore the discover features in a smaller scale trial. While well-received, our deployment highlighted several challenges that we further discuss in this paper. Overall, our study contributes to crowdsourced digital health solutions and provides both design and research implications to this challenging domain with a vulnerable user group.
Elina Kuosmanen, Eetu Huusko, Niels van Berkel, Francisco Nunes, Julio Vega, Jorge Gonçalves 0001, Mohamed Khamis, Augusto Esteves, Denzil Ferreira, Simo Hosio
Int. J. Hum. Comput. Stud.3
2023 Mapping 20 years of accessibility research in HCI: A co-word analysis
Zhanna Sarsenbayeva, Niels van Berkel, Danula Hettiachchi, Benjamin Tag, Eduardo Velloso, Jorge Gonçalves 0001, Vassilis Kostakos
Int. J. Hum. Comput. Stud.2
2023 AWARE-Light: a smartphone tool for experience sampling and digital phenotyping
Niels van Berkel, Simon D'Alfonso, Rio Kurnia Susanto, Denzil Ferreira, Vassilis Kostakos
Pers. Ubiquitous Comput.1
2023 Measurements, Algorithms, and Presentations of Reality: Framing Interactions with AI-Enabled Decision Support
abstract
Bringing AI technology into clinical practice has proved challenging for system designers and medical professionals alike. The academic literature has, for example, highlighted the dangers of black-box decision-making and biased datasets. Furthermore, end-users’ ability to validate a system’s performance often disappears following the introduction of AI decision-making. We present the MAP model to understand and describe the three stages through which medical observations are interpreted and handled by AI systems. These stages are Measurement, in which information is gathered and converted into data points that can be stored and processed; Algorithm, in which computational processes transform the collected data; and Presentation, where information is returned to the user for interpretation. For each stage, we highlight possible challenges that need to be overcome to develop Human-Centred AI systems. We illuminate our MAP model through complementary case studies on colonoscopy practice and dementia diagnosis, providing examples of the challenges encountered in real-world settings. By defining Human-AI interaction across these three stages, we untangle some of the inherent complexities in designing AI technology for clinical decision-making, and aim to overcome misalignment between medical end-users and AI researchers and developers.
Niels van Berkel, Maura Bellio, Mikael B. Skov, Ann Blandford
ACM Trans. Comput. Hum. Interact.1
2023 Near-infrared Imaging for Information Embedding and Extraction with Layered Structures
abstract
Non-invasive inspection and imaging techniques are used to acquire non-visible information embedded in samples. Typical applications include medical imaging, defect evaluation, and electronics testing. However, existing methods have specific limitations, including safety risks (e.g., X-ray), equipment costs (e.g., optical tomography), personnel training (e.g., ultrasonography), and material constraints (e.g., terahertz spectroscopy). Such constraints make these approaches impractical for everyday scenarios. In this article, we present a method that is low-cost and practical for non-invasive inspection in everyday settings. Our prototype incorporates a miniaturized near-infrared spectroscopy scanner driven by a computer-controlled 2D-plotter. Our work presents a method to optimize content embedding, as well as a wavelength selection algorithm to extract content without human supervision. We show that our method can successfully extract occluded text through a paper stack of up to 16 pages. In addition, we present a deep-learning-based image enhancement model that can further improve the image quality and simultaneously decompose overlapping content. Finally, we demonstrate how our method can be generalized to different inks and other layered materials beyond paper. Our approach enables a wide range of content embedding applications, including chipless information embedding, physical secret sharing, 3D print evaluations, and steganography.
Weiwei Jiang 0001, Difeng Yu, Chaofan Wang 0001, Zhanna Sarsenbayeva, Niels van Berkel, Jorge Gonçalves 0001, Vassilis Kostakos
ACM Trans. Graph.5
2022 Characterising Soundscape Research in Human-Computer Interaction
abstract
‘Soundscapes’ are an increasingly active topic in Human-Computer Interaction (HCI) and interaction design. From mapping acoustic environments through sound recordings to designing compositions as interventions, soundscapes appear as a recurring theme across a wide body of HCI research. Based on this growing interest, now is the time to explore the types of studies in which soundscapes provide a valuable lens to HCI research. In this paper, we review papers from conferences sponsored or co-sponsored by the ACM Special Interest Group on Computer-Human Interaction in which the term ’soundscape’ occurs. We analyse a total of 235 papers to understand the role of soundscapes as a research focus and identify untapped opportunities for soundscape research within HCI. We identify two common soundscape conceptualisations: (1) Acoustic environments and (2) Compositions, and describe what characterises studies into each concept and the hybrid forms that also occur. On the basis of this, we carve out a foundation for future soundscape research in HCI as a methodological anchor to form a common ground and support this growing research interest. Finally, we offer five recommendations for further research into soundscapes within HCI.
Stine S. Johansen, Niels van Berkel, Jonas Fritsch
Conference on Designing Interactive Systems2
2022 Method for Appropriating the Brief Implicit Association Test to Elicit Biases in Users
abstract
Implicit tendencies and cognitive biases play an important role in how information is perceived and processed, a fact that can be both utilised and exploited by computing systems. The Implicit Association Test (IAT) has been widely used to assess people’s associations of target concepts with qualitative attributes, such as the likelihood of being hired or convicted depending on race, gender, or age. The condensed version–the Brief IAT–aims to implicit biases by measuring the reaction time to concept classifications. To use this measure in HCI research, however, we need a way to construct and validate target concepts, which tend to quickly evolve and depend on geographical and cultural interpretations. In this paper, we introduce and evaluate a new method to appropriate the BIAT using crowdsourcing to measure people’s leanings on polarising topics. We present a web-based tool to test participants’ bias on custom themes, where self-assessments often fail. We validated our approach with 14 domain experts and assessed the fit of crowdsourced test construction. Our method allows researchers of different domains to create and validate bias tests that can be geographically tailored and updated over time. We discuss how our method can be applied to surface implicit user biases and run studies where cognitive biases may impede reliable results.
Tilman Dingler, Benjamin Tag, David A. Eccles, Niels van Berkel, Vassilis Kostakos
CHI4
2022 Do You See What I Hear? - Peripheral Absolute and Relational Visualisation Techniques for Sound Zones
abstract
Sound zone technology allows multiple simultaneous sound experiences for multiple people in the same room without interference. However, given the inherent invisible and intangible nature of sound zones, it is unclear how to communicate the position and size of sound zones to users. This paper compares two visualisation techniques; absolute visualisation, relational visualisation, as well as a baseline condition without visualisations. In a within-subject experiment (N = 33), we evaluated these techniques for effectiveness and efficiency across four representative tasks. Our findings show that the absolute and relational visualisation techniques increase effectiveness in multi-user tasks but not in single-user tasks. The efficiency for all tasks was improved using visualisations. We discuss the potential of visualisations for sound zones and highlight future research opportunities for sound zone interaction.
Rune Møberg Jacobsen, Niels van Berkel, Mikael B. Skov, Stine S. Johansen, Jesper Kjeldskov
CHI2
2022 Quantifying Synthesis and Fusion and their Impact on Machine Translation
abstract
Arturo Oncevay, Duygu Ataman, Niels Van Berkel, Barry Haddow, Alexandra Birch, Johannes Bjerva. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022.
Arturo Oncevay, Duygu Ataman, Niels van Berkel, Barry Haddow, Alexandra Birch, Johannes Bjerva
NAACL-HLT3
2022 Tangible Self-Report Devices: Accuracy and Resolution of Participant Input
abstract
Tangible input has been explored as a means for participants to self-report experiences while minimising disruption and allowing for discrete data collection. However, the accuracy of these tangible devices has not been studied systematically. We compared six input techniques, including slider, slider with resistance, capacitive touch slider, squeeze, rotary knob, and joystick, to understand their accuracy and resolution profile. Each of these wireless devices was designed in a similar form factor and intended to be operated discretely with one hand. We assessed input accuracy and participant perceptions across devices through a controlled lab study (N = 20), highlighting diverging limits to the accuracy of the input technique and possible explanations for the differences in resolution. Our results indicate that participant accuracy was highest using a slider, and lowest using a squeeze-based input. We discuss the suitability and challenges of discreet tangible self-report techniques, and highlight open research questions for future work.
Niels van Berkel, Timothy Merritt 0001, Anders Bruun, Mikael B. Skov
TEI1
2022 Human-centred artificial intelligence: a contextual morality perspective
abstract
The emergence of big data combined with the technical developments in Artificial Intelligence has enabled novel opportunities for autonomous and continuous decision support. While initial work has begun to explore how human morality can inform the decision making of future Artificial Intelligence applications, these approaches typically consider human morals as static and immutable. In this work, we present an initial exploration of the effect of context on human morality from a Utilitarian perspective. Through an online narrative transportation study, in which participants are primed with either a positive story, a negative story or a control condition (N = 82), we collect participants' perceptions on technology that has to deal with moral judgment in changing contexts. Based on an in-depth qualitative analysis of participant responses, we contrast participant perceptions to related work on Fairness, Accountability and Transparency. Our work highlights the importance of contextual morality for Artificial Intelligence and identifies opportunities for future work through a FACT-based (Fairness, Accountability, Context and Transparency) perspective.
Niels van Berkel, Benjamin Tag, Jorge Gonçalves 0001, Simo Hosio
Behav. Inf. Technol.1
2022 Quantifying determinants of social conformity in an online debating website
Senuri Wijenayake, Niels van Berkel, Vassilis Kostakos, Jorge Gonçalves 0001
Int. J. Hum. Comput. Stud.2
2022 How Does Sleep Tracking Influence Your Life?: Experiences from a Longitudinal Field Study with a Wearable Ring
abstract
A new generation of wearable devices now enable end-users to keep track of their sleep patterns. This paper reports on a longitudinal study of 82 participants who used a state-of-the-art sleep tracking ring for an average of 65 days. We conducted interviews and questionnaires to understand changes to their lifestyle, their perceptions of the tracked information and sleep, and the overall experience of using an unobtrusive sleep tracking device. Our results indicate that such a device is suitable for long-term sleep tracking and helpful in identifying detrimental lifestyle elements that hinder sleep quality. However, tracking one's sleep can also introduce stress or physical discomfort, potentially leading to adverse outcomes. We discuss these findings in light of related work and highlight the near-term research directions that the rapid commoditisation of sleep tracking technology enables.
Elina Kuosmanen, Aku Visuri, Saba Kheirinejad, Niels van Berkel, Heli Koskimäki, Denzil Ferreira, Simo Hosio
Proc. ACM Hum. Comput. Interact.4
2022 Crowdsourcing sensitive data using public displays - opportunities, challenges, and considerations
abstract
Abstract Interactive public displays are versatile two-way interfaces between the digital world and passersby. They can convey information and harvest purposeful data from their users. Surprisingly little work has exploited public displays for collecting tagged data that might be useful beyond a single application. In this work, we set to fill this gap and present two studies: (1) a field study where we investigated collecting biometrically tagged video-selfies using public kiosk-sized screens, and (2) an online narrative transportation study that further elicited rich qualitative insights on key emerging aspects from the first study. In the first study, a 61-day deployment resulted in 199 video-selfies with consent to leverage the videos in any non-profit research. The field study indicates that people are willing to donate even highly sensitive data about themselves in public. The subsequent online narrative transportation study provides a deeper understanding of a variety of issues arising from the first study that can be leveraged in the future design of such systems. The two studies combined in this article pave the way forward towards a vision where volunteers can, should they so choose, ethically and serendipitously help unleash advances in data-driven areas such as computer vision and machine learning in health care.
Andy Alorwu, Niels van Berkel, Jorge Gonçalves 0001, Jonas Oppenlaender, Miguel Bordallo López, Mahalakshmy Seetharaman, Simo Hosio
Pers. Ubiquitous Comput.2
2022 Initial Responses to False Positives in AI-Supported Continuous Interactions: A Colonoscopy Case Study
abstract
The use of artificial intelligence (AI) in clinical support systems is increasing. In this article, we focus on AI support for continuous interaction scenarios. A thorough understanding of end-user behaviour during these continuous human-AI interactions, in which user input is sustained over time and during which AI suggestions can appear at any time, is still missing. We present a controlled lab study involving 21 endoscopists and an AI colonoscopy support system. Using a custom-developed application and an off-the-shelf videogame controller, we record participants’ navigation behaviour and clinical assessment across 14 endoscopic videos. Each video is manually annotated to mimic an AI recommendation, being either true positive or false positive in nature. We find that time between AI recommendation and clinical assessment is significantly longer for incorrect assessments. Further, the type of medical content displayed significantly affects decision time. Finally, we discover that the participant’s clinical role plays a large part in the perception of clinical AI support systems. Our study presents a realistic assessment of the effects of imperfect and continuous AI support in a clinical scenario.
Niels van Berkel, Jeremy Opie, Omer F. Ahmad, Laurence B. Lovat, Danail Stoyanov, Ann Blandford
ACM Trans. Interact. Intell. Syst.1
2021 User Trust in Assisted Decision-Making Using Miniaturized Near-Infrared Spectroscopy
abstract
We investigate the use of a miniaturized Near-Infrared Spectroscopy (NIRS) device in an assisted decision-making task. We consider the real-world scenario of determining whether food contains gluten, and we investigate how end-users interact with our NIRS detection device to ultimately make this judgment. In particular, we explore the effects of different nutrition labels and representations of confidence on participants’ perception and trust. Our results show that participants tend to be conservative in their judgment and are willing to trust the device in the absence of understandable label information. We further identify strategies to increase user trust in the system. Our work contributes to the growing body of knowledge on how NIRS can be mass-appropriated for everyday sensing tasks, and how to enhance the trustworthiness of assisted decision-making systems.
Weiwei Jiang 0001, Zhanna Sarsenbayeva, Niels van Berkel, Chaofan Wang 0001, Difeng Yu, Jing Wei 0002, Jorge Gonçalves 0001, Vassilis Kostakos
CHI3
2021 Assessing MyData Scenarios: Ethics, Concerns, and the Promise
abstract
Public controversies around the unethical use of personal data are increasing, spotlighting data ethics as an increasingly important field of study. MyData is a related emerging vision that emphasizes individuals’ control of their personal data. In this paper, we investigate people’s perceptions of various data management scenarios by measuring the perceived ethicality and level of felt concern concerning the scenarios. We deployed a set of 96 unique scenarios to an online crowdsourcing platform for assessment and invited a representative sample of the participants to a second-stage questionnaire about the MyData vision and its potential in the field of healthcare. Our results provide a timely investigation into how topical data-related practices affect the perceived ethicality and the felt concern. The questionnaire analysis reveals great potential in the MyData vision. Through the combined quantitative and qualitative results, we contribute to the field of data ethics.
Andy Alorwu, Saba Kheirinejad, Niels van Berkel, Marianne Kinnula, Denzil Ferreira, Aku Visuri, Simo Hosio
CHI3
2021 Effect of Information Presentation on Fairness Perceptions of Machine Learning Predictors
abstract
The uptake of artificial intelligence-based applications raises concerns about the fairness and transparency of AI behaviour. Consequently, the Computer Science community calls for the involvement of the general public in the design and evaluation of AI systems. Assessing the fairness of individual predictors is an essential step in the development of equitable algorithms. In this study, we evaluate the effect of two common visualisation techniques (text-based and scatterplot) and the display of the outcome information (i.e., ground-truth) on the perceived fairness of predictors. Our results from an online crowdsourcing study (N = 80) show that the chosen visualisation technique significantly alters people’s fairness perception and that the presented scenario, as well as the participant’s gender and past education, influence perceived fairness. Based on these results we draw recommendations for future work that seeks to involve non-experts in AI fairness evaluations.
Niels van Berkel, Jorge Gonçalves 0001, Daniel Russo 0002, Simo Hosio, Mikael B. Skov
CHI1
2021 E-Scooter Sustainability - A Clash of Needs, Perspectives, and Experiences
Maria Kjærup, Mikael B. Skov, Niels van Berkel
INTERACT (3)3
2021 Rainmaker: A Tangible Work-Companion for the Personal Office Space
abstract
Routines are an important element of day-to-day work life, supporting people in structuring their day around required tasks. Effectively managing these routines is, however, experienced as challenging by many – an issue further amplified by the current work from home lockdown measures. In this paper we present Rainmaker, a tangible device to support people in their working life in the context of their own homes. We evaluate and iterate on our prototype through two qualitative studies, spanning respectively three days (N = 11) and 15 days (N = 2). Our results highlight the perceived advantages of the use of a primarily physical rather than digital tool for work support, allowing users to stay focused on their tasks and reflect on their work achievements. We present lessons for future work in this area and publicly release the software and hardware used in the construction of Rainmaker.
Sujay Shalawadi, Anas Alnayef, Niels van Berkel, Jesper Kjeldskov, Florian Echtler
MobileHCI3
2021 Predictors of well-being and productivity among software professionals during the COVID-19 pandemic - a longitudinal study
abstract
The COVID-19 pandemic has forced governments worldwide to impose movement restrictions on their citizens. Although critical to reducing the virus’ reproduction rate, these restrictions come with far-reaching social and economic consequences. In this paper, we investigate the impact of these restrictions on an individual level among software engineers who were working from home. Although software professionals are accustomed to working with digital tools, but not all of them remotely, in their day-to-day work, the abrupt and enforced work-from-home context has resulted in an unprecedented scenario for the software engineering community. In a two-wave longitudinal study ( N = 192), we covered over 50 psychological, social, situational, and physiological factors that have previously been associated with well-being or productivity. Examples include anxiety, distractions, coping strategies, psychological and physical needs, office set-up, stress, and work motivation. This design allowed us to identify the variables that explained unique variance in well-being and productivity. Results include (1) the quality of social contacts predicted positively, and stress predicted an individual’s well-being negatively when controlling for other variables consistently across both waves; (2) boredom and distractions predicted productivity negatively; (3) productivity was less strongly associated with all predictor variables at time two compared to time one, suggesting that software engineers adapted to the lockdown situation over time; and (4) longitudinal analyses did not provide evidence that any predictor variable causal explained variance in well-being and productivity. Overall, we conclude that working from home was per se not a significant challenge for software engineers. Finally, our study can assess the effectiveness of current work-from-home and general well-being and productivity support guidelines and provides tailored insights for software professionals.
Daniel Russo 0002, Paul H. P. Hanel, Seraphina Altnickel, Niels van Berkel
Empir. Softw. Eng.4
2021 Designing Visual Markers for Continuous Artificial Intelligence Support: A Colonoscopy Case Study
abstract
Colonoscopy, the visual inspection of the large bowel using an endoscope, offers protection against colorectal cancer by allowing for the detection and removal of pre-cancerous polyps. The literature on polyp detection shows widely varying miss rates among clinicians, with averages ranging around 22%--27%. While recent work has considered the use of AI support systems for polyp detection, how to visualise and integrate these systems into clinical practice is an open question. In this work, we explore the design of visual markers as used in an AI support system for colonoscopy. Supported by the gastroenterologists in our team, we designed seven unique visual markers and rendered them on real-life patient video footage. Through an online survey targeting relevant clinical staff ( N = 36), we evaluated these designs and obtained initial insights and understanding into the way in which clinical staff envision AI to integrate in their daily work-environment. Our results provide concrete recommendations for the future deployment of AI support systems in continuous, adaptive scenarios.
Niels van Berkel, Omer F. Ahmad, Danail Stoyanov, Laurence B. Lovat, Ann Blandford
ACM Trans. Comput. Heal.1
2021 Modeling interaction as a complex system
abstract
Researchers in Human-Computer Interaction typically rely on experiments to assess the causal effects of experimental conditions on variables of interest. Although this classic approach can be very useful, it offers little help in tackling questions of causality in the kind of data that are increasingly common in HCI – capturing user behavior ‘in the wild.’ To analyze such data, model-based regressions such as cross-lagged panel models or vector autoregressions can be used, but these require parametric assumptions about the structural form of effects among the variables. To overcome some of the limitations associated with experiments and model-based regressions, we adopt and extend ‘empirical dynamic modelling’ methods from ecology that lend themselves to conceptualizing multiple users’ behavior as complex nonlinear dynamical systems. Extending a method known as ‘convergent cross mapping’ or CCM, we show how to make causal inferences that do not rely on experimental manipulations or model-based regressions and, by virtue of being non-parametric, can accommodate data emanating from complex nonlinear dynamical systems. By using this approach for multiple users, which we call ‘multiple convergent cross mapping’ or MCCM, researchers can achieve a better understanding of the interactions between users and technology – by distinguishing causality from correlation – in real-world settings.
Niels van Berkel, Simon Dennis 0001, Michael Zyphur, Jinjing Li, Andrew Heathcote, Vassilis Kostakos
Hum. Comput. Interact.1
2021 Information flow and cognition affect each other: Evidence from digital learning
abstract
In the context of learning systems, identifying causal relationships among information presented to the user, their behavior and cognitive effort required/exerted to understand and perform a task is key to building effective learning experiences, and to maintain engagement in learning processes. An unexplored question is whether our interaction with presented information affects our cognitive effort (and behaviour), or vice-versa. We investigate causal relationship between information presented and cognitive effort (and behaviour) in the context of two separate studies (N = 40, N = 98), and study the effect of instruction (active/passive task). We utilize screen-recordings and eye-tracking data to investigate the relationship among these variables. To investigate the causal relationships among the different measurements, we use Granger’s causality. Further, we propose a new method to combine two time-series from multiple participants for detecting causal relationships. Our results indicate that information presentation drives user focus size (behaviour), and that cognitive load (a measure of cognitive effort exerted) drives information presentation. This relationship is also moderated by instruction type and performance-level (high/low). We draw implications for design of educational material and learning technologies.
Kshitij Sharma, Katerina Mangaroska, Niels van Berkel, Michail N. Giannakos, Vassilis Kostakos
Int. J. Hum. Comput. Stud.3
2021 Understanding usage style transformation during long-term smartwatch use
abstract
Abstract Despite large investments in smartwatch development, the market growth remains smaller than forecasted. The purpose of smartwatch use remains unclear, indicated by the lack of large-scale adoption. Thus, we aim to better understand the early adoption and everyday smartwatch use. We investigate a diverse usage data of smartwatches logged over a period of up to 14 months from 79 individuals between December 2015 and March 2017, one of the largest wearable datasets collected. First, we identify both explorative and accepted behaviours that users exhibit and further investigate how the individual usage traits and features differ between the two categories. Our analysis offers an insightful perspective on how smartwatch use evolves organically. Our results improve our shared understanding of smartwatch use and users adapting their use of smartwatch over time to match the capabilities of the technology by validating numerous findings from previous literature.
Aku Visuri, Niels van Berkel, Jorge Gonçalves 0001, Reza Rawassizadeh, Denzil Ferreira, Vassilis Kostakos
Pers. Ubiquitous Comput.2
2021 Making Appearances: How Robots Should Approach People
abstract
To prepare for a future in which robots are more commonplace, it is important to know what robot behaviors people find socially normative. Previous work suggests that for robots to be accepted by people, the robot should adhere to the prevalent social norms, such as those related to approaching people. However, we do not expect that socially normative approach behaviors for robots can be translated on a one-on-one basis from people to robots, because currently robots have unique and different features to humans, including (but not limited to) wheels, sounds, and shapes. The two studies presented in this article go beyond the state-of-the-art and focus on socially normative approach behaviors for robots. In the first study, we compared people’s responses to violations of personal space done by robots compared to people. In the second study, we explored what features (sound, size, speed) of a robot approaching people have an effect on acceptance. Findings indicate that people are more lenient toward violations of a social norm by a robot as compared to a person. Also, we found that robots can use their unique features to mitigate the negative effects of norm violations by communicating intent.
Michiel Joosse, Manja Lohse, Niels van Berkel, Aziez Sardar, Vanessa Evers
ACM Trans. Hum. Robot Interact.3
2020 "Hi! I am the Crowd Tasker" Crowdsourcing through Digital Voice Assistants
abstract
Inspired by the increasing prevalence of digital voice assistants, we demonstrate the feasibility of using voice interfaces to deploy and complete crowd tasks. We have developed Crowd Tasker, a novel system that delivers crowd tasks through a digital voice assistant. In a lab study, we validate our proof-of-concept and show that crowd task performance through a voice assistant is comparable to that of a web interface for voice-compatible and voice-based crowd tasks for native English speakers. We also report on a field study where participants used our system in their homes. We find that crowdsourcing through voice can provide greater flexibility to crowd workers by allowing them to work in brief sessions, enabling multi-tasking, and reducing the time and effort required to initiate tasks. We conclude by proposing a set of design guidelines for the creation of crowd tasks for voice and the development of future voice-based crowdsourcing systems.
Danula Hettiachchi, Zhanna Sarsenbayeva, Fraser Allison, Niels van Berkel, Tilman Dingler, Gabriele Marini, Vassilis Kostakos, Jorge Gonçalves 0001
CHI4
2020 Does Smartphone Use Drive our Emotions or vice versa? A Causal Analysis
abstract
In this paper, we demonstrate the existence of a bidirectional causal relationship between smartphone application use and user emotions. In a two-week long in-the-wild study with 30 participants we captured 502,851 instances of smartphone application use in tandem with corresponding emotional data from facial expressions. Our analysis shows that while in most cases application use drives user emotions, multiple application categories exist for which the causal effect is in the opposite direction. Our findings shed light on the relationship between smartphone use and emotional states. We furthermore discuss the opportunities for research and practice that arise from our findings and their potential to support emotional well-being.
Zhanna Sarsenbayeva, Gabriele Marini, Niels van Berkel, Chu Luo, Weiwei Jiang 0001, Kangning Yang, Greg Wadley, Tilman Dingler, Vassilis Kostakos, Jorge Gonçalves 0001
CHI3
2020 Overcoming compliance bias in self-report studies: A cross-study analysis
Niels van Berkel, Jorge Gonçalves 0001, Simo Hosio, Zhanna Sarsenbayeva, Eduardo Velloso, Vassilis Kostakos
Int. J. Hum. Comput. Stud.1
2020 Human accuracy in mobile data collection
Niels van Berkel, Jorge Gonçalves 0001, Katarzyna Wac, Simo Hosio, Anna Louise Cox
Int. J. Hum. Comput. Stud.1
2020 Dimensions of ecological validity for usability evaluations in clinical settings
Niels van Berkel, Matthew J. Clarkson, Guofang Xiao, Eren Dursun, Moustafa Allam, Brian R. Davidson, Ann Blandford
J. Biomed. Informatics1
2020 CrowdCog: A Cognitive Skill based System for Heterogeneous Task Assignment and Recommendation in Crowdsourcing
abstract
While crowd workers typically complete a variety of tasks in crowdsourcing platforms, there is no widely accepted method to successfully match workers to different types of tasks. Researchers have considered using worker demographics, behavioural traces, and prior task completion records to optimise task assignment. However, optimum task assignment remains a challenging research problem due to limitations of proposed approaches, which in turn can have a significant impact on the future of crowdsourcing. We present 'CrowdCog', an online dynamic system that performs both task assignment and task recommendations, by relying on fast-paced online cognitive tests to estimate worker performance across a variety of tasks. Our work extends prior work that highlights the effect of workers' cognitive ability on crowdsourcing task performance. Our study, deployed on Amazon Mechanical Turk, involved 574 workers and 983 HITs that span across four typical crowd tasks (Classification, Counting, Transcription, and Sentiment Analysis). Our results show that both our assignment method and recommendation method result in a significant performance increase (5% to 20%) as compared to a generic or random task assignment. Our findings pave the way for the use of quick cognitive tests to provide robust recommendations and assignments to crowd workers.
Danula Hettiachchi, Niels van Berkel, Vassilis Kostakos, Jorge Gonçalves 0001
Proc. ACM Hum. Comput. Interact.2
2020 Quantifying the Effect of Social Presence on Online Social Conformity
abstract
Social conformity occurs when individuals in group settings change their personal opinion to be in agreement with the majority's position. While recent literature frequently reports on conformity in online group settings, the causes for online conformity are yet to be fully understood. This study aims to understand how social presencei.e., the sense of being connected to others via mediated communication, influences conformity among individuals placed in online groups while answering subjective and objective questions. Acknowledging its multifaceted nature, we investigate three aspects of online social presence: user representation (generic vs.user-specific avatars), interactivity (discussion vs.no discussion ), and response visibility (public vs.private ). Our results show an overall conformity rate of 30% and main effects from task objectivity, group size difference between the majority and the minority, and self-confidence on personal answer. Furthermore, we observe an interaction effect between interactivity and response visibility, such that conformity is highest in the presence of peer discussion and public responses, and lowest when these two elements are absent. We conclude with a discussion on the implications of our findings in designing online group settings, accounting for the effects of social presence on conformity.
Senuri Wijenayake, Niels van Berkel, Vassilis Kostakos, Jorge Gonçalves 0001
Proc. ACM Hum. Comput. Interact.2
2019 Context-Informed Scheduling and Analysis: Improving Accuracy of Mobile Self-Reports
abstract
Mobile self-reports are a popular technique to collect participant labelled data in the wild. While literature has focused on increasing participant compliance to self-report questionnaires, relatively little work has assessed response accuracy. In this paper, we investigate how participant context can affect response accuracy and help identify strategies to improve the accuracy of mobile self-report data. In a 3-week study we collect over 2,500 questionnaires containing both verifiable and non-verifiable questions. We find that response accuracy is higher for questionnaires that arrive when the phone is not in ongoing or very recent use. Furthermore, our results show that long completion times are an indicator of a lower accuracy. Using contextual mechanisms readily available on smartphones, we are able to explain up to 13% of the variance in participant accuracy. We offer actionable recommendations to assist researchers in their future deployments of mobile self-report studies.
Niels van Berkel, Jorge Gonçalves 0001, Peter Koval, Simo Hosio, Tilman Dingler, Denzil Ferreira, Vassilis Kostakos
CHI1
2019 Effect of Cognitive Abilities on Crowdsourcing Task Performance
Danula Hettiachchi, Niels van Berkel, Simo Hosio, Vassilis Kostakos, Jorge Gonçalves 0001
INTERACT (1)2
2019 Effect of Ambient Light on Mobile Interaction
Zhanna Sarsenbayeva, Niels van Berkel, Weiwei Jiang 0001, Danula Hettiachchi, Vassilis Kostakos, Jorge Gonçalves 0001
INTERACT (3)2
2019 Improving Experience Sampling with Multi-view User-driven Annotation Prediction
abstract
A fundamental challenge in real-time labelling of activity data is user burden. The Experience Sampling Method (ESM) is widely used to obtain such labels for sensor data. However, in an in-situ deployment, it is not feasible to expect users to precisely label the start and end time of each event or activity. For this reason, time-point based experience sampling (without an actual start and end time) is prevalent. We present a framework that applies multi-instance and semi-supervised learning techniques to perform to predict user annotations from multiple mobile sensor data streams. Our proposed framework estimates users' annotations in ESM-based studies progressively, via an interactive pipeline of co-training and active learning. We evaluate our work using data collected from an in-the-wild data collection.
Jonathan Liono, Flora D. Salim, Niels van Berkel, Vassilis Kostakos, A. K. Qin 0001
PerCom3
2019 Effect of experience sampling schedules on response rate and recall accuracy of objective self-reports
Niels van Berkel, Jorge Gonçalves 0001, Lauri Lovén, Denzil Ferreira, Simo Hosio, Vassilis Kostakos
Int. J. Hum. Comput. Stud.1
2019 Understanding smartphone notifications' user interactions and content importance
Aku Visuri, Niels van Berkel, Tadashi Okoshi, Jorge Gonçalves 0001, Vassilis Kostakos
Int. J. Hum. Comput. Stud.2
2019 Crowdsourcing Perceptions of Fair Predictors for Machine Learning: A Recidivism Case Study
abstract
The increased reliance on algorithmic decision-making in socially impactful processes has intensified the calls for algorithms that are unbiased and procedurally fair. Identifying fair predictors is an essential step in the construction of equitable algorithms, but the lack of ground-truth in fair predictor selection makes this a challenging task. In our study, we recruit 90 crowdworkers to judge the inclusion of various predictors for recidivism. We divide participants across three conditions with varying group composition. Our results show that participants were able to make informed decisions on predictor selection. We find that agreement with the majority vote is higher when participants are part of a more diverse group. The presented workflow, which provides a scalable and practical approach to reach a diverse audience, allows researchers to capture participants' perceptions of fairness in private while simultaneously allowing for structured participant discussion.
Niels van Berkel, Jorge Gonçalves 0001, Danula Hettiachchi, Senuri Wijenayake, Ryan Kelly 0001, Vassilis Kostakos
Proc. ACM Hum. Comput. Interact.1
2019 Measuring the Effects of Gender on Online Social Conformity
abstract
Social conformity occurs when an individual changes their behaviour in line with the majority's expectations. Although social conformity has been investigated in small group settings, the effect of gender - of both the individual and the majority/minority - is not well understood in online settings. Here we systematically investigate the impact of groups' gender composition on social conformity in online settings. We use an online quiz in which participants submit their answers and confidence scores, both prior to and following the presentation of peer answers that are dynamically fabricated. Our results show an overall conformity rate of 39%, and a significant effect of gender that manifests in a number of ways: gender composition of the majority, the perceived nature of the question, participant gender, visual cues of the system, and final answer correctness. We conclude with a discussion on the implications of our findings in designing online group settings, accounting for the effects of gender on conformity.
Senuri Wijenayake, Niels van Berkel, Vassilis Kostakos, Jorge Gonçalves 0001
Proc. ACM Hum. Comput. Interact.2
2019 Energy-efficient prediction of smartphone unlocking
Chu Luo, Aku Visuri, Simon Klakegg, Niels van Berkel, Zhanna Sarsenbayeva, Antti Möttönen, Jorge Gonçalves 0001, Theodoros Anagnostopoulos, Denzil Ferreira, Huber Flores, Eduardo Velloso, Vassilis Kostakos
Pers. Ubiquitous Comput.4
2018 Crowdsourcing Treatments for Low Back Pain
abstract
Low back pain (LBP) is a globally common condition with no silver bullet solutions. Further, the lack of therapeutic consensus causes challenges in choosing suitable solutions to try. In this work, we crowdsourced knowledge bases on LBP treatments. The knowledge bases were used to rank and offer best-matching LBP treatments to end users. We collected two knowledge bases: one from clinical professionals and one from non-professionals. Our quantitative analysis revealed that non-professional end users perceived the best treatments by both groups as equally good. However, the worst treatments by non-professionals were clearly seen as inferior to the lowest ranking treatments by professionals. Certain treatments by professionals were also perceived significantly differently by non-professionals and professionals themselves. Professionals found our system handy for self-reflection and for educating new patients, while non-professionals appreciated the reliable decision support that also respected the non-professional opinion.
Simo Hosio, Jaro Karppinen, Esa-Pekka Takala, Jani Takatalo, Jorge Gonçalves 0001, Niels van Berkel, Shin'ichi Konomi, Vassilis Kostakos
CHI6
2018 Facilitating Collocated Crowdsourcing on Situated Displays
abstract
Online crowdsourcing enables the distribution of work to a global labor force as small and often repetitive tasks. Recently, situated crowdsourcing has emerged as a complementary enabler to elicit labor in specific locations and from specific crowds. Teamwork in online crowdsourcing has been recently shown to increase the quality of output, but teamwork in situated crowdsourcing remains unexplored. We set out to fill this gap. We present a generic crowdsourcing platform that supports situated teamwork and provide experiences from a laboratory study that focused on comparing traditional online crowdsourcing to situated team-based crowdsourcing. We built a crowdsourcing desk that hosts three networked terminal displays. The displays run our custom team-driven crowdsourcing platform that was used to investigate collocated crowdsourcing in small teams. In addition to analyzing quantitative data, we provide findings based on questionnaires, interviews, and observations. We highlight 1) emerging differences between traditional and collocated crowdsourcing, 2) the collaboration strategies that teams exhibited in collocated crowdsourcing, and 3) that a priori team familiarity does not significantly affect collocated interaction in crowdsourcing. The approach we introduce is a novel multi-display crowdsourcing setup that supports collocated labor teams and along with the reported study makes specific contributions to situated crowdsourcing research.
Simo Hosio, Jorge Gonçalves 0001, Niels van Berkel, Simon Klakegg, Shin'ichi Konomi, Vassilis Kostakos
Hum. Comput. Interact.3
2017 Towards Commoditised Near Infrared Spectroscopy
abstract
Near Infrared Spectroscopy (NIRS) is a sensing technique in which near infrared light is transmitted into a sample, followed by light absorbance measurements at various wavelengths. This technique enables the inference of the inner chemical composition of the scanned sample, and therefore can be used to identify or classify objects. In this paper, we describe how to facilitate the use of NIRS by non- expert users in everyday settings. Our work highlights the key challenges of placing NIRS devices in the hands of non-experts. We develop a system to mitigate these challenges, and evaluate it in a user study. We show how NIRS technology can be successfully utilised by untrained users in an unsupervised manner through a special enclosure and an accompanying smartphone app. Finally, we discuss potential future developments of commoditised NIRS.
Simon Klakegg, Jorge Gonçalves 0001, Niels van Berkel, Chu Luo, Simo Hosio, Vassilis Kostakos
Conference on Designing Interactive Systems3
2017 Quantifying Sources and Types of Smartwatch Usage Sessions
abstract
We seek to quantify smartwatch use, and establish differences and similarities to smartphone use. Our analysis considers use traces from 307 users that include over 2.8 million notifications and 800,000 screen usage events, and we compare our findings to previous work that quantifies smartphone use. The results show that smartwatches are used more briefly and more frequently throughout the day, with half the sessions lasting less than 5 seconds. Interaction with notifications is similar across both types of devices, both in terms of response times and preferred application types. We also analyse the differences between our smartwatch dataset and a dataset aggregated from four previously conducted smartphone studies. The similarities and differences between smartwatch and smartphone use suggest effect on usage that go beyond differences in form factor.
Aku Visuri, Zhanna Sarsenbayeva, Niels van Berkel, Jorge Gonçalves 0001, Reza Rawassizadeh, Vassilis Kostakos, Denzil Ferreira
CHI3
2017 Understanding elderly care: a field-study for designing future homes
abstract
While the population is aging the role of information and communication technology (ICT) has grown in elderly care. This development has brought versatile ICT-related supportive systems to professionals and laymen working with aging people. The current study analyzed how professionals in elderly care perceived their workflow challenges before new ICT is developed and implemented to support their work. The results of this study are set to inform the design of a novel ICT system for a sheltered care home.
Hanna-Leena Huttunen, Simon Klakegg, Niels van Berkel, Aku Visuri, Denzil Ferreira, Raija Halonen
iiWAS3
2017 Predicting interruptibility for manual data collection: a cluster-based user model
abstract
Previous work suggests that Quantified-Self applications can retain long-term usage with motivational methods. These methods often require intermittent attention requests with manual data input. This may cause unnecessary burden to the user, leading to annoyance, frustration and possible application abandonment. We designed a novel method that uses on-screen alert dialogs to transform recurrent smartphone usage sessions into moments of data contributions and evaluate how accurately machine learning can reduce unintended interruptions. We collected sensor data from 48 participants during a 4-week long deployment and analysed how personal device usage can be considered in scheduling data inputs. We show that up to 81.7% of user interactions with the alert dialogs can be accurately predicted using user clusters, and up to 75.5% of unintended interruptions can be prevented and rescheduled. Our approach can be leveraged by applications that require self-reports on a frequent basis and may provide a better longitudinal QS experience.
Aku Visuri, Niels van Berkel, Chu Luo, Jorge Gonçalves 0001, Denzil Ferreira, Vassilis Kostakos
MobileHCI2
2017 Tapping Task Performance on Smartphones in Cold Temperature
abstract
We present a study that quantifies the effect of cold temperature on smartphone input performance, particularly on tapping tasks. Our results show that smartphone input performance decreases when completing tapping tasks in cold temperatures. We show that colder temperature is associated with lower throughput and less accurate performance when using the phone in both one-handed and two-handed operations. We also demonstrate that colder temperature is related to higher error rate when using the phone in one-handed operation only, but not two-handed. Finally, we identify a number of design recommendations from the literature that can be considered as a countermeasure to poorer smartphone input performance in completing tapping tasks in cold temperature.
Jorge Gonçalves 0001, Zhanna Sarsenbayeva, Niels van Berkel, Chu Luo, Simo Hosio, Sirkka Rissanen, Hannu Rintamäki, Vassilis Kostakos
Interact. Comput.3
2016 A Systematic Assessment of Smartphone Usage Gaps
abstract
Researchers who analyse smartphone usage logs often make the assumption that users who lock and unlock their phone for brief periods of time (e.g., less than a minute) are continuing the same "session" of interaction. However, this assumption is not empirically validated, and in fact different studies apply different arbitrary thresholds in their analysis. To validate this assumption, we conducted a field study where we collected user-labelled activity data through ESM and sensor logging. Our results indicate that for the majority of instances where users return to their smartphone, i.e., unlock their device, they in fact begin a new session as opposed to continuing a previous one. Our findings suggest that the commonly used approach of ignoring brief standby periods is not reliable, but optimisation is possible. We therefore propose various metrics related to usage sessions and evaluate various machine learning approaches to classify gaps in usage.
Niels van Berkel, Chu Luo, Theodoros Anagnostopoulos, Denzil Ferreira, Jorge Gonçalves 0001, Simo Hosio, Vassilis Kostakos
CHI1
2016 Monetary Assessment of Battery Life on Smartphones
abstract
Research claims that users value the battery life of their smartphones, but no study to date has attempted to quantify battery value and how this value changes according to users' current context and needs. Previous work has quantified the monetary value that smartphone users place on their data (e.g., location), but not on battery life. Here we present a field study and methodology for systematically measuring the monetary value of smartphone battery life, using a reverse second-price sealed-bid auction protocol. Our results show that the prices for the first and last 10% battery segments differ substantially. Our findings also quantify the tradeoffs that users consider in relation to battery, and provide a monetary model that can be used to measure the value of apps and enable fair ad-hoc sharing of smartphone resources.
Simo Hosio, Denzil Ferreira, Jorge Gonçalves 0001, Niels van Berkel, Chu Luo, Muzamil Ahmed, Huber Flores, Vassilis Kostakos
CHI4
2013 The influence of approach speed and functional noise on users' perception of a robot
abstract
How a robot approaches a person greatly determines the interaction that follows. This is particularly relevant when the person has never interacted with the robot before. In human communication, we exchange a multitude of multimodal signals to communicate our intent while we approach others. However, most robots do not have the capabilities to produce such signals and easily communicate their intent. In this paper we propose to communicate intent when a robot approaches a person through functional noise and approach speed. Both were manipulated in a between-subjects experiment (N=40) either slowly increasing at the start of the approach and slowly decreasing when the robot reached the human or maximized at the start and abruptly stopped at the end of the approach. We analyzed questionnaires and video data from the interaction and found that particularly functional noise that in-/decreased in volume was helpful to communicate the robot's intent but only in congruence with an in-/decreasing velocity.
Manja Lohse, Niels van Berkel, Betsy van Dijk, Michiel Joosse, Daphne E. Karreman, Vanessa Evers
IROS2