Florian Kunneman

dblp:133/9549 · also Florian A. Kunneman · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
7since 2021 · last 2027
0000-0002-1932-3200ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 5 · 5 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
YearPublicationVenuePosition
2027 Crossing margins: Intersectional users' ethical concerns about software
abstract
Abstract Many modern software applications present numerous ethical concerns due to conflicts between users’ values and companies’ priorities. Intersectional communities, those with multiple marginalized identities, are disproportionately affected by these ethical issues, leading to legal, financial, and reputational consequences for software companies, as well as real-world harm for intersectional users. Historically, the voices of intersectional communities have been systematically marginalized and excluded from contributing their unique perspectives to software design, perpetuating software-related ethical concerns. This work aims to fill the gap in research on intersectional users’ software-related perspectives and provide software practitioners with a methodology for analyzing intersectional voices in software ethics discourse. We collected 36,777 posts from over 700 intersectional subreddits discussing software applications and utilized large language models to identify ethical concerns in these posts. We then applied regression models with counterfactual analysis to examine how intersectional identity dimensions shape the amplification or suppression of ethical concern expression across software genres, and conducted a time-series analysis to examine how concern expression varies over time in relation to real-world events. As a case study in the social media domain, we further demonstrate how identified ethical concerns can be prioritized to surface issues warranting timely developer attention, validated against survey-derived ground truth. Together, these analyses form the basis of a nascent feedback-driven framework for assessing whether software systems are meeting the needs of intersectional users.
Lauren Olson, Tom P. Humbert, Ricarda Anna-Lena Fischer, Bob Westerveld, Florian Kunneman, Emitza Guzman
Empir. Softw. Eng.5
2025 What Can You Say to a Robot? Capability Communication Leads to More Natural Conversations
abstract
When encountering a robot in the wild, it is not inherently clear to human users what the robot's capabilities are. When encountering misunderstandings or problems in spoken interaction, robots often just apologize and move on, without additional effort to make sure the user understands what happened. We set out to compare the effect of two speech based capability communication strategies (proactive, reactive) to a robot without such a strategy, in regard to the user's rating of and their behavior during the interaction. For this, we conducted an in-person user study with 120 participants who had three speech-based interactions with a social robot in a restaurant setting. Our results suggest that users preferred the robot communicating its capabilities proactively and adjusted their behavior in those interactions, using a more conversational interaction style while also enjoying the interaction more.
Merle M. Reimann, Koen V. Hindriks, Florian Kunneman, Catharine Oertel, Gabriel Skantze, Iolanda Leite
HRI3
2025 Whose voices are heard? Gender disparities in platform-facilitated discrimination and content moderation
abstract
Historically, cisgender men have maintained systemic social, cultural, and political privilege over other genders. Online discrimination serves as a mechanism for reinforcing this dominance. Content moderation plays a crucial role in shaping online experiences, yet the ways it may perpetuate or mitigate discrimination remain underexplored. This study examines how content moderation and discussions of discrimination vary across gendered online communities, with a focus on identifying differential impacts by gender group. We analyzed 124 subreddits spanning three gender groups—women, men, and gender minorities (GM). The analysis included manual annotation of 1,535 posts and machine learning classification of an additional 6,613 posts to assess the prevalence of user discussions regarding online discrimination and content moderation. Women were most likely to report top-down moderation issues, such as bans and content removal, while GM users engaged more frequently with general moderation concerns. Time series analysis revealed that complaints about content moderation have increased over time, with the steepest rise among women users. These patterns demonstrate that moderation policies and enforcement impact gender groups differently. Our findings highlight the need for improvements in software engineering and user experience design for content moderation tools. Enhancing transparency, promoting equity, and enabling more user-driven moderation experiences are critical steps toward protecting marginalized groups against discrimination online.
Lauren Olson, Ricarda Anna-Lena Fischer, Tom P. Humbert, Florian Kunneman, Emitza Guzman
Inf. Softw. Technol.4
2024 A Near-Real-Time Processing Ego Speech Filtering Pipeline Designed for Speech Interruption During Human-Robot Interaction
abstract
With current state-of-the-art (SOTA) automatic speech recognition (ASR) systems, it is not possible to transcribe overlapping speech audio streams separately. Consequently, when these ASR systems are used as part of a social robot like Pepper for interaction with a human, it is common practice to close the robot’s microphone while it is talking itself. This prevents the human users to interrupt the robot, which limits speech-based human-robot interaction. To enable a more natural interaction which allows for such interruptions, we propose an audio processing pipeline for filtering out robot’s ego speech using only a single-channel microphone. This pipeline takes advantage of the possibility to feed the robot ego speech signal, generated by a text-to-speech API, as training data into a machine learning model. The proposed pipeline combines a convolutional neural network and spectral subtraction to extract overlapping human speech from the audio recorded by the robot-embedded microphone. When evaluating on a held-out test set, we find that this pipeline outperforms our previous approach to this task, as well as SOTA target speech extraction systems that were retrained on the same dataset. We have also integrated the proposed pipeline into a lightweight robot software development framework to make it available for broader use. As a step towards demonstrating the feasibility of deploying our pipeline, we use this framework to evaluate the effectiveness of the pipeline in a small lab-based feasibility pilot using the social robot Pepper. Our results show that when participants interrupt the robot, the pipeline can extract the participant’s speech from one-second streaming audio buffers received by the robot-embedded single-channel microphone, hence in near-real time.
Yue Li 0044, Florian Kunneman, Koen V. Hindriks
RO-MAN2
2024 A Survey on Dialogue Management in Human-robot Interaction
abstract
As social robots see increasing deployment within the general public, improving the interaction with those robots is essential. Spoken language offers an intuitive interface for the human–robot interaction (HRI), with dialogue management (DM) being a key component in those interactive systems. Yet, to overcome current challenges and manage smooth, informative, and engaging interaction, a more structural approach to combining HRI and DM is needed. In this systematic review, we analyze the current use of DM in HRI and focus on the type of dialogue manager used, its capabilities, evaluation methods, and the challenges specific to DM in HRI. We identify the challenges and current scientific frontier related to the DM approach, interaction domain, robot appearance, physical situatedness, and multimodality.
Merle M. Reimann, Florian Kunneman, Catharine Oertel, Koen V. Hindriks
ACM Trans. Hum. Robot Interact.2
2023 Predicting Interaction Quality Aspects Using Level-Based Scores for Conversational Agents
abstract
In order to improve human-agent interaction, it is essential to have good measures of interaction quality. We define interaction quality based on multiple aspects, including usability, likability and perceived conversation quality as subjective measures, and interaction length, completion rate and frequency of unrecognized utterances as objective measures. Determining necessary improvements to a conversational agent is a non-trivial task, because it is difficult to infer from an evaluation of the agent as a whole, which aspects of the agent need to be improved to raise the interaction quality. In this paper, we propose a scoring system for task-oriented conversational agents to predict aspects of interaction quality and to guide an iterative improvement process. Our scoring system does not provide a single score, but leverages structural features of the dialogue management approach and assigns a score on three levels: the utterance, dialogue move, and genre level. Using the agent's scores on separate levels to predict the interaction quality allows making targeted improvements to the conversational agent. In order to evaluate our scoring system, we apply it over the course of multiple crowdsourcing pilot studies, using a recipe recommendation agent. We evaluate the obtained scores in regard to their ability to predict selected objective and subjective interaction quality aspects, as well as their suitability for making informed decisions about necessary improvements.
Merle M. Reimann, Catharine Oertel, Florian Kunneman, Koen V. Hindriks
IVA3
2023 A Semi-Real-Time Method for Social Robots to Detect and Locate Overlapping Speech Events
abstract
It is useful for a social robot to detect and locate users based on their speech. Notable challenges hampering the effective localization of a speaker are background noise and overlapping speech. Convolutional Neural Networks (CNNs) have shown to yield good performance on locating single speakers on a curated dataset, but to a lesser extent in scenarios with two speakers. In addition, their computational cost is still too high for a timely reaction in real-world settings. We build on the current state-of-the-art CNN approach, and propose several improvements for distinguishing multiple speakers by time-alignment in the input representation and reducing computational costs by considerably shortening the input audio blocks. We evaluate this approach on an existing dataset with blocks of noisy and overlapping speech recorded in rooms of different sizes, predicting the number of active speech events and their azimuth locations. The results show that our approach outperforms other approaches in locating two speakers and is considerably faster than the best-performing alternative approach. The time-domain information in the input representation was found essential for predicting the location of the signal source.
Yue Li 0044, Koen V. Hindriks, Florian Kunneman
RO-MAN3
2018 Aspect-based summarization of pros and cons in unstructured product reviews
abstract
We developed three systems for generating pros and cons summaries of product reviews. Automating this task eases the writing of product reviews, and offers readers quick access to the most important information. We compared SynPat, a system based on syntactic phrases selected on the basis of valence scores, against a neural-network-based system trained to map bag-of-words representations of reviews directly to pros and cons, and the same neural system trained on clusters of word-embedding encodings of similar pros and cons. We evaluated the systems in two ways: first on held-out reviews with gold-standard pros and cons, and second by asking human annotators to rate the systems’ output on relevance and completeness. In the second evaluation, the gold-standard pros and cons were assessed along with the system output. We find that the human-generated summaries are not deemed as significantly more relevant or complete than the SynPat systems; the latter are scored higher than the human-generated summaries on a precision metric. The neural approaches yield a lower performance in the human assessment, and are outperformed by the baseline.
Florian Kunneman, Sander Wubben, Antal van den Bosch, Emiel Krahmer
COLING1
2016 Open-domain extraction of future events from Twitter
abstract
Abstract Explicit references on Twitter to future events can be leveraged to feed a fully automatic monitoring system of real-world events. We describe a system that extracts open-domain future events from the Twitter stream. It detects future time expressions and entity mentions in tweets, clusters tweets together that overlap in these mentions above certain thresholds, and summarizes these clusters into event descriptions that can be presented to users of the system. Terms for the event description are selected in an unsupervised fashion.1We evaluated the system on a month of Dutch tweets, by showing the top-250 ranked events found in this month to human annotators. Eighty per cent of the candidate events were indeed assessed as being an event by at least three out of four human annotators, while all four annotators regarded sixty-three per cent as a real event. An added component to complement event descriptions with additional terms was not assessed better than the original system, due to the occasional addition of redundant terms. Comparing the found events to gold-standard events from maintained calendars on the Web mentioned in at least five tweets, the system yields a recall-at-250 of 0.20 and a recall based on all retrieved events of 0.40.
Florian Kunneman, Antal van den Bosch
Nat. Lang. Eng.1
2015 Signaling sarcasm: From hyperbole to hashtag
abstract
To avoid a sarcastic message being understood in its unintended literal meaning, in microtexts such as messages on Twitter.com sarcasm is often explicitly marked with a hashtag such as ‘#sarcasm’. We collected a training corpus of about 406 thousand Dutch tweets with hashtag synonyms denoting sarcasm. Assuming that the human labeling is correct (annotation of a sample indicates that about 90% of these tweets are indeed sarcastic), we train a machine learning classifier on the harvested examples, and apply it to a sample of a day’s stream of 2.25 million Dutch tweets. Of the 353 explicitly marked tweets on this day, we detect 309 (87%) with the hashtag removed. We annotate the top of the ranked list of tweets most likely to be sarcastic that do not have the explicit hashtag. 35% of the top-250 ranked tweets are indeed sarcastic. Analysis indicates that the use of hashtags reduces the further use of linguistic markers for signaling sarcasm, such as exclamations and intensifiers. We hypothesize that explicit markers such as hashtags are the digital extralinguistic equivalent of non-verbal expressions that people employ in live interaction when conveying sarcasm. Checking the consistency of our finding in a language from another language family, we observe that in French the hashtag ‘#sarcasme’ has a similar polarity switching function, be it to a lesser extent.
Florian Kunneman, Christine Liebrecht, Margot van Mulken, Antal van den Bosch
Inf. Process. Manag.1