Söhnke Benedikt Fischedick

dblp:324/2329 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
8since 2021 · last 2026
0000-0001-8447-0584ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 4 · 4 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Exploring Mediated Communication with Older Adults: Comparing AR Avatars, Telepresence Robots, and Face-to-Face Interaction
abstract
Older adults, a growing demographic, face an increased risk of experiencing loneliness and are less exposed to emerging communication technologies. Augmented reality (AR) avatars and telepresence robots have been proposed as tools to foster social connection, yet their suitability for older users remains underexplored. We present an exploratory study with ten healthy older adults who engaged in both conversational and spatial collaboration tasks using AR avatar-mediated communication, robot-mediated communication, and face-to-face interaction. We collected self-reported measures of co-presence, social presence, closeness, uncanny valley, preferences, and open feedback. Our findings suggest that telepresence robots enhanced co-presence, while avatars were valued for their expressivity and humanlike qualities. Task type influenced co-presence in spatial collaboration only during communication using the telepresence robot. Other measures, such as social presence and closeness, were unaffected by task type or representation. While neither technology outperformed face-to-face interaction, both were positively received, underscoring their potential to address the social needs of older adults and highlighting the importance of enhancing nonverbal expressivity, particularly nonverbal cues in mediated communication. Ultimately, our results contribute to the fundamental understanding of mediated communication with older adults, motivating further empirical work to confirm and extend these findings.
Stephanie Arevalo, Jakob Hartbrich, Florian Weidner, Melisa Conde, Veronika Mikhailova, Felix Immohr, Söhnke Benedikt Fischedick, Bea Vorhof, Christoph Gerhardt, Kay Richter, Christian Kunert, Nicola Döring, Horst-Michael Groß, Wolfgang Broll, Alexander Raake
IMX7
2026 "The Robot Should Be Programmed for Me": User Tests Evaluating a Telepresence Robot for the Social Integration of Older Adults
abstract
Telepresence robots that allow communication between older adults and their remotely located social contacts can foster social integration. The present laboratory test study explores older adults’ successful use of a telepresence robot (Research Question 1 [RQ1]), as well as their perceived enjoyment (RQ2), perceived ease of use (RQ3), perceived usefulness (RQ4), perceived social presence (RQ5), and intention to use (RQ6) a telepresence robot for robot-mediated communication (RMC). Semi-structured interviews, observations, and questionnaires were applied with a group of N = 14 older adults living in Germany. Participants completed a navigational task (as remote users) and an interpersonal communication task (as local users). Results show older adults used the telepresence robot successfully (RQ1) during the tasks. Furthermore, in interviews, older adults described their perceived enjoyment (RQ2), perceived ease of use (RQ3), and perceived usefulness (RQ4) during RMC as generally high. Perceived social presence (RQ5) during RMC was generally described as high, with RMC being considered a viable substitute when face-to-face communication is not possible. Finally, only two participants (2/14) had no intention to use (RQ6) a telepresence robot in the long term. Future design recommendations are provided, such as adapting the telepresence robot’s functions to older adults physical, psychological, and social conditions.
Melisa Conde, Söhnke Benedikt Fischedick, Kay Richter, Stephanie Arevalo, Horst-Michael Groß, Alexander Raake, Nicola Döring
ACM Trans. Hum. Robot Interact.2
2025 Efficient Prediction of Dense Visual Embeddings via Distillation and RGB-D Transformers
abstract
In domestic environments, robots require a comprehensive understanding of their surroundings to interact effectively and intuitively with untrained humans. In this paper, we propose DVEFormer – an efficient RGB-D Transformer-based approach that predicts dense text-aligned visual embeddings (DVE) via knowledge distillation. Instead of directly performing classical semantic segmentation with fixed predefined classes, our method uses teacher embeddings from Alpha-CLIP to guide our efficient student model DVEFormer in learning fine-grained pixel-wise embeddings. While this approach still enables classical semantic segmentation, e.g., via linear probing, it further enables flexible text-based querying and other applications, such as creating comprehensive 3D maps. Evaluations on common indoor datasets demonstrate that our approach achieves competitive performance while meeting real-time requirements, operating at 26.3FPS for the full model and 77.0FPS for a smaller variant on an NVIDIA Jetson AGX Orin. Additionally, we show qualitative results that highlight the effectiveness and possible use cases in real-world applications. Overall, our method serves as a drop-in replacement for traditional segmentation approaches while enabling flexible natural-language querying and seamless integration into 3D mapping pipelines for mobile robotics.
Söhnke Benedikt Fischedick, Daniel Seichter, Benedict Stephan, Robin Schmidt, Horst-Michael Groß
IROS1
2025 Robot, Avatar, or Human: The Impact of Partner Representation and Task on the Communication Experience
abstract
Avatars and telepresence robots have long received attention for remote communication. However, the specific nature of their physicality, expressiveness, and mobility may affect their usefulness for different tasks. This work compares using an avatar (presented in augmented reality) and a telepresence robot to Face-to-Face (F2F) communication during different communication tasks: free conversation, negotiation, and referential communication with movement. We conducted a user study (split-plot design, N=54) with the type of representation of the conversational partner as the within variable and the communication task as the between variable. Our results show that the type of task, especially referential communication with movement, influenced the perceived attention to nonverbal cues and closeness. Generally, gestures and body movements received the least focus with telepresence robots. Gestures in avatars and F2F drew similar attention, which we attribute to the avatar's tracking fidelity. Gaze received less attention in both avatar- and robot-mediated communication compared to F2F, while facial expressions on the robot's screen heightened attention compared to avatars. These findings advance the fundamental understanding of mediated communication and support researchers and practitioners in shaping the design of communication applications beyond today's video calls.
Stephanie Arevalo, Jakob Hartbrich, Florian Weidner, Söhnke Benedikt Fischedick, Christoph Gerhardt, Kay Richter, Christian Kunert, Bea Vorhof, Horst-Michael Groß, Wolfgang Broll, Alexander Raake
Proc. ACM Hum. Comput. Interact.4
2024 An exploratory study on the impact of varying levels of robot control on presence in robot-mediated communication
abstract
Telepresence robots can enhance communication experiences by providing a sense of physical presence, embodiment and may evoke co-presence. In spite of that, telepresence robots have not made it fully to consumer markets. In this paper, we investigate how different levels of controlling a telepresence robot (teleoperation, shared control, and no control) influence presence. To this aim, we conducted a study (N=45) where participants were evenly distributed to one of the robot control conditions. The task involved navigating an unknown room and listening to stories told by a person co-located with the robot. We collected subjective impressions of presence using the temple presence inventory and performed a thematic content analysis on a post-experiment interview. Our results suggest nuances in perceived presence under different levels of robot control after performing a thematic content analysis. Copresence can be experienced during teleoperation and shared control, and teleoperation may evoke negative sentiments if it does not provide enough spatial information during navigation. However, our results did not point to significant differences in spatial or social presence. We consider that these findings encourage further discussions on how presence is perceived in robot-mediated communication.
Stephanie Arevalo, Söhnke Benedikt Fischedick, Chenayo Diao, Kay Richter, Horst-Michael Groß, Alexander Raake
RO-MAN2
2023 Efficient Multi-Task Scene Analysis with RGB-D Transformers
abstract
Scene analysis is essential for enabling autonomous systems, such as mobile robots, to operate in real-world environments. However, obtaining a comprehensive understanding of the scene requires solving multiple tasks, such as panoptic segmentation, instance orientation estimation, and scene classification. Solving these tasks given limited computing and battery capabilities on mobile platforms is challenging. To address this challenge, we introduce an efficient multi-task scene analysis approach, called EMSAFormer, that uses an RGB-D Transformer-based encoder to simultaneously perform the aforementioned tasks. Our approach builds upon the previously published EMSANet. However, we show that the dual CNN-based encoder of EMSANet can be replaced with a single Transformer-based encoder. To achieve this, we investigate how information from both RGB and depth data can be effectively incorporated in a single encoder. To accelerate inference on robotic hardware, we provide a custom NVIDIA TensorRT extension enabling highly optimization for our EMSAFormer approach. Through extensive experiments on the commonly used indoor datasets NYUv2, SUNRGB-D, and ScanNet, we show that our approach achieves state-of-the-art performance while still enabling inference with up to 39.1 FPS on an NVIDIA Jetson AGX Orin 32 GB.
Söhnke Benedikt Fischedick, Daniel Seichter, Robin Schmidt, Leonard Rabes, Horst-Michael Groß
IJCNN1
2023 PanopticNDT: Efficient and Robust Panoptic Mapping
abstract
As the application scenarios of mobile robots are getting more complex and challenging, scene understanding becomes increasingly crucial. A mobile robot that is supposed to operate autonomously in indoor environments must have precise knowledge about what objects are present, where they are, what their spatial extent is, and how they can be reached; i.e., information about free space is also crucial. Panoptic mapping is a powerful instrument providing such information. However, building 3D panoptic maps with high spatial resolution is challenging on mobile robots, given their limited computing capabilities. In this paper, we propose PanopticNDT – an efficient and robust panoptic mapping approach based on occupancy normal distribution transform (NDT) mapping. We evaluate our approach on the publicly available datasets Hypersim and ScanNetV2. The results reveal that our approach can represent panoptic information at a higher level of detail than other state-of-the-art approaches while enabling real-time panoptic mapping on mobile robots. Finally, we prove the real-world applicability of PanopticNDT with qualitative results in a domestic application.
Daniel Seichter, Benedict Stephan, Söhnke Benedikt Fischedick, Steffen Müller 0001, Leonard Rabes, Horst-Michael Groß
IROS3
2022 Efficient Multi-Task RGB-D Scene Analysis for Indoor Environments
abstract
Semantic scene understanding is essential for mobile agents acting in various environments. Although semantic segmentation already provides a lot of information, details about individual objects as well as the general scene are missing but required for many real-world applications. However, solving multiple tasks separately is expensive and cannot be accomplished in real time given limited computing and battery capabilities on a mobile platform. In this paper, we propose an efficient multi-task approach for RGB-D scene analysis (EMSANet) that simultaneously performs semantic and instance segmentation (panoptic segmentation), instance orientation estimation, and scene classification. We show that all tasks can be accomplished using a single neural network in real time on a mobile platform without diminishing performance - by contrast, the individual tasks are able to benefit from each other. In order to evaluate our multi-task approach, we extend the annotations of the common RGB-D indoor datasets NYUv2 and SUNRGB-D for instance segmentation and orientation estimation. To the best of our knowledge, we are the first to provide results in such a comprehensive multi-task setting for indoor scene analysis on NYUv2 and SUNRGB-D.
Daniel Seichter, Söhnke Benedikt Fischedick, Mona Köhler, Horst-Michael Groß
IJCNN2