Artem V. Belopolsky

dblp:06/7227 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
4since 2021 · last 2025
0000-0003-3606-2134ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 4 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Beware of the Tablet: A Dominant Distractor in Human-Robot Interaction
abstract
The present study aims at investigating how humans engage with common communication modalities—speech, tablet, and gesture—when interacting with a humanoid robot. To explore this, we designed a live interaction experiment using a congruence paradigm, where participants engaged with a robot presenting two out of three modalities simultaneously: one as the primary cue and the other as a distracting cue. We measured participants’ task performance (response time, error rate) and fixation distribution (fixation count and duration proportions) across different roles (primary, distracting, neither) and areas of interest (face, tablet, gesture). Additionally, we compared fixation patterns between the performance and baseline phases. Our findings reveal that while the tablet is the most effective modality for task engagement, it also serves as a strong attentional distractor, dominating gaze allocation regardless of its informational value. This underscores the importance of carefully balancing tablet integration in HRI design. Notably, our results demonstrate that gaze patterns alone do not fully reveal attentional focus, emphasizing the need to consider both overt and covert cognitive processes in multimodal HRI. These insights provide valuable guidelines for designing more effective and engaging human-robot interactions.
Linlin Cheng, Artem V. Belopolsky, Mark de Bruijn, Koen V. Hindriks
IROS2
2025 Evaluating Appearance-Based Gaze Pattern for Human-Robot Interaction
abstract
Appearance-based gaze estimation, an accessible and unobtrusive alternative to eye tracking, has advanced significantly, yet their adoption in human-robot interaction (HRI) remains limited. A key barrier is the lack of clarity on how they compare to high-precision eye trackers. To address this, we evaluate this method against eye-tracker glasses in an HRI setting using calibration and attention detection tasks. We assess performance across different cameras (4K and robot’s built-in camera) and participant conditions (with and without glasses). Results show that the 4K camera and participants without glasses yield higher accuracy and precision. With a simple offset correction, this method achieves comparable performance to eye-tracker glasses for average gaze pattern but struggles with detecting gaze patterns over time. It also demonstrates potential for real-time robot attention detection. We conclude that appearance-based gaze estimation is a viable, cost-effective alternative to traditional eye tracking in HRI, particularly for average gaze pattern detection.
Linlin Cheng, Koen V. Hindriks, Mark De Bruijn, Artem V. Belopolsky
RO-MAN4
2024 Automating Gaze Target Annotation in Human-Robot Interaction
abstract
Identifying gaze targets in videos of human-robot interaction is useful for measuring engagement. In practice, this requires manually annotating for a fixed set of objects that a participant is looking at in a video, which is very time-consuming. To address this issue, we propose an annotation pipeline for automating this effort. In this work, we focus on videos in which the objects looked at do not move. As input for the proposed pipeline, we therefore only need to annotate object bounding boxes for the first frame of each video. The benefit, moreover, of manually annotating these frames is that we can also draw bounding boxes for objects outside of it, which enables estimating gaze targets in videos where not all objects are visible. A second issue that we address is that the models used for automating the pipeline annotate individual video frames. In practice, however, manual annotation is done at the event level for video segments instead of single frames. Therefore, we also introduce and investigate several variants of algorithms for aggregating frame-level to event-level annotations, which are used in the last step in our annotation pipeline.We compare two versions of our pipeline: one that uses a state-of-the-art gaze estimation model (GEM) and a second one using a state-of-the-art target detection model (TDM). Our results show that both versions successfully automate the annotation, but the GEM pipeline performs slightly (≈10%) better for videos where not all objects are visible. Analysis of our aggregation algorithm, moreover, shows that there is no need for manual video segmentation because a fixed time interval for segmentation yields very similar results. We conclude that the proposed pipeline can be used to automate almost all of the annotation effort.
Linlin Cheng, Koen V. Hindriks, Artem V. Belopolsky
RO-MAN3
2023 Boundary Conditions for Human Gaze Estimation on A Social Robot using State-of-the-Art Models
abstract
Appearance-based methods are a promising solution for gaze estimation, as they eliminate the need for additional devices and calibration. This makes them particularly well-suited for human-robot interaction (HRI) research. However, until recently their performance was under par compared to traditional eye-trackers. Recent breakthroughs have been made with the release of two large-scale datasets with a wide range of gaze directions (Gaze360 and ETH-XGaze) and the accompanying state-of-the-art deep neural networks (L2CS and ETH). In this paper, we systematically evaluate the performance of these two appearance-based models on a social robot. In our setup, we vary the distance from the robot (1-3m) and camera resolution (640*480 and 3840*2160) and analyze the performance in terms of accuracy and precision. We find that the L2CS model trained on the Gaze360 dataset combined with a 4K camera achieves the best performance on the 2 m and 3 m distances. We show that a simple offset correction on pitch and yaw can further increase the accuracy and precision by 18.6% and 9.6% respectively. We conclude that for a range up to 3 m appearance-based gaze estimation models provide a promising approach for application in HRI research.
Linlin Cheng, Artem V. Belopolsky, Koen V. Hindriks
RO-MAN2