EDBT 2026 Demo / reviewers in the wild / expert
Heike Brock
dblp:205/7288
· DBLP profile ↗
19ranked-venue papers
6as first author
13since 2021 · last 2025
0000-0002-9530-4714ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 4 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 6 · 2 first-author · 3 since 2021Systems, architecture and hardware · 5 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | HypCAD: Geometry-Enhanced Hyperbolic Contrastive Learning for CAD Model RetrievalabstractRetrieving CAD models for real-world object scans enhances object-level mapping, providing a nuanced spatial understanding crucial for precise interactions in robotics or mixed reality. Commonly, CAD model retrieval is performed by matching features learned in Euclidean space. However, learning discriminative features in Euclidean space faces significant challenges, primarily due to its flat nature and the wide variety of CAD models with different levels of detail. To address the limitations of Euclidean space and improve CAD model retrieval, this paper introduces HypCAD, a contrastive learning framework in hyperbolic space. We present a novel geometry-enhanced hyperbolic distance and utilize a three-component contrastive learning loss to learn hyperbolic feature representations for the CAD model retrieval task. We demonstrate HypCAD’s superior retrieval accuracy through comparisons with baseline contrastive learning methods on both the synthetic ShapeNet dataset and the real-world Scan2CAD dataset. Adam Misik, Driton Salihu, Heike Brock, Eckehard G. Steinbach |
ICASSP | 4 |
| 2025 | EQUR: Equivariant Uncertainty Quantification and Refinement for Point Cloud RegistrationabstractPoint cloud registration is a crucial task for robotics and mixed reality applications, serving as a foundational component for problems such as 3D reconstruction and localization. Arbitrary poses and real-world artifacts, including noise and occlusions, increase registration uncertainty and limit the performance of current point cloud registration algorithms. This paper proposes a novel, sampling-free uncertainty quantification and refinement method for point cloud registration, termed EQUR. To consistently predict uncertainty with high robustness, we extract equivariant point features, from which we regress an uncertainty score, enabling robust quantification of registration uncertainty. Subsequently, we leverage the estimated registration uncertainty as an auxiliary input to enhance the prediction of transformation refinement terms. We employ an introspective learning strategy to train EQUR based on the errors of a baseline registration model. Through quantitative and qualitative analyses on synthetic ShapeNet and real-world ScanObjectNN datasets, we showcase the effectiveness of EQUR, demonstrating both high accuracies in uncertainty quantification and uncertainty-aided refinement of point cloud registration. Adam Misik, Driton Salihu, Xiaoang Zhang, Heike Brock, Eckehard G. Steinbach |
ICIP | 4 |
| 2024 | NeRF-Feat: 6D Object Pose Estimation using Feature RenderingabstractObject Pose Estimation is a crucial component in robotic grasping and augmented reality. Learning based approaches typically require training data from a highly accurate CAD model or labeled training data acquired using a complex setup. We address this by learning to estimate pose from weakly labeled data without a known CAD model. We propose to use a NeRF to learn object shape implicitly which is later used to learn view-invariant features in conjunction with CNN using a contrastive loss. While NeRF helps in learning features that are view-consistent, CNN ensures that the learned features respect symmetry. During inference, CNN is used to predict view-invariant features which can be used to establish correspondences with the implicit 3d model in NeRF. The correspondences are then used to estimate the pose in the reference frame of NeRF. Our approach can also handle symmetric objects unlike other approaches using a similar training setup. Specifically, we learn viewpoint invariant, discriminative features using NeRF which are later used for pose estimation. We evaluated our approach on LM, LM-Occlusion, and T-Less dataset and achieved benchmark accuracy despite using weakly labeled data. Shishir Reddy Vutukur, Heike Brock, Benjamin Busam, Tolga Birdal, Andreas Hutter, Slobodan Ilic |
3DV | 2 |
| 2024 | HEGN: Hierarchical Equivariant Graph Neural Network for 9DoF Point Cloud RegistrationabstractGiven its wide application in robotics, point cloud registration is a widely researched topic. Conventional methods aim to find a rotation and translation that align two point clouds in 6 degrees of freedom (DoF). However, certain tasks in robotics, such as category-level pose estimation, involve non-uniformly scaled point clouds, requiring a 9DoF transform for accurate alignment. We propose HEGN, a novel equivariant graph neural network for 9DoF point cloud registration. HEGN utilizes equivariance to rotation, translation, and scaling to estimate the transformation without relying on point correspondences. Based on graph representations for both point clouds, we extract equivariant node features aggregated in their local, cross-, and global context. In addition, we introduce a novel node pooling mechanism that leverages the cross-context importance of nodes to pool the graph representation. By repeating the feature extraction and node pooling, we obtain a graph hierarchy. Finally, we determine rotation and translation by aligning equivariant features aggregated over the graph hierarchy. To estimate scaling, we leverage scale information in the vector norm of the equivariant features. We evaluate the effectiveness of HEGN through experiments with the synthetic ModelNet40 dataset and the real-world ScanObjectNN dataset. The results show the superior performance of HEGN in 9DoF point cloud registration and its competitive performance in conventional 6DoF point cloud registration. Adam Misik, Driton Salihu, Heike Brock, Eckehard G. Steinbach |
ICRA | 4 |
| 2023 | COCCA: Point Cloud Completion through Cad Cross-Attentionabstract3D scene- and object-level scans typically result in sparse and incomplete point clouds. Since dense point clouds of high quality are essential for the 3D reconstruction process, a promising approach is to improve the scan quality by point cloud completion. In this paper, we present COCCA, an extension of point cloud completion networks for scan-to-CAD use cases. The proposed extension is based on cross-attention of features extracted from a scan with rotation-, translation-, and scale-invariant features extracted from a sampled CAD point cloud. With the proposed cross-attention operation, we improve the learning of scan features and the subsequent decoding to a complete shape. We demonstrate the effectiveness of COCCA on the ShapeNet dataset in quantitative and qualitative experiments. COCCA improves the overall completion performance of point cloud completion networks by up to 11.8% for Chamfer Distance and up to 2.2% for F-Score. Our qualitative experiments visualize how COCCA completes point clouds with higher geometric detail. In addition, we demonstrate how completion by COCCA improves the point cloud registration task required for scan-to-CAD alignment. Adam Misik, Driton Salihu, Heike Brock, Eckehard G. Steinbach |
ICIP | 3 |
| 2022 | Making The Unknown More Certain: A Stacked Ensemble Classifier for Open Gesture Recognition with a Social RobotabstractWe introduce a novel stacked ensemble classifier for the unconstrained recognition of known and unknown gestural input data in nonverbal communication with a social robot. The architecture utilizes three separate CNNs of different expected data input size and combines their output predictions to a unified estimate. Analysis shows that in comparison to a single CNN architecture, the combined estimate reduces prediction confidence values for unknown gestural movement segments, making the system able to identify unknown data input with higher certainty under both laboratory and real environment conditions. In a human-robot interaction experiment, we are able to improve unknown class detection accuracy by up to 40% under maintained or equal known class recognition performance, and hence considerably enhance the overall robustness of the recognition system. Heike Brock, Randy Gomez |
ICASSP | 1 |
| 2022 | Developing The Bottom-up Attentional System of A Social RobotabstractThis paper describes the development of a 3- stage signalling framework to trigger a social robot's bottom- up reactive behavior inspired by a biological model. In the first stage, low-level firing of stimuli due to external sources is constructed through perception grounding. This is followed by a saliency classifier which fires-up high level salient signals that require attention and are used to trigger the robot's reactive behavior. The whole framework evolves primarily on the knowledge ontology that defines the characteristics of the social robot and the querying mechanism that correlates the perceived stimuli with the ontology to trigger the reactive behavior. We evaluated the performance of our system with timing metrics and we achieved good results for our application. Randy Gomez, Álvaro Páez, Yu Fang 0007, Serge Thill, Luis Merino, Eric Nichols, Keisuke Nakamura, Heike Brock |
ICRA | 8 |
| 2022 | Affective Behavior Learning for Social Robot Haru with Implicit Evaluative FeedbackabstractWe propose a human-in-the-loop reinforcement learning mechanism to help robots learn emotional behavior. Unlike the previous methods of providing explicit feedback via pressing keyboard buttons or mouse clicks, we provide a more natural way for ordinary people to train social robots how to perform social tasks according to their preferences - facial expressions. The whole experiment is carried out on the desktop robot Haru, which is mainly used for the research of emotion and empathy participation. Our experimental results show that through learning from implicit feedback of facial features, Haru can quickly understand and dynamically adapt to individual preferences, and obtain a similar performance to learning from explicit feedback. In addition, we observe that the recognition error of human feedback will cause a “temporary regress” of the robot's learning performance, which is more obvious at the beginning of the training process. This phenomenon is shown to be correlated with the accuracy of recognizing negative implicit feedback. Hui Wang 0141, Jinying Lin, Yurii Vasylkiv, Heike Brock, Keisuke Nakamura, Randy Gomez, Bo He 0002, Guangliang Li |
IROS | 5 |
| 2021 | Automating Behavior Selection for Affective Telepresence RobotabstractThe tabletop robot Haru, used for affective telepresence research, enables a teleoperator to communicate affects from a distance. The robot’s expressiveness offers myriad ways of communicating affects through the execution of emotive routines. The teleoperator reacts to input modalities such as the user’s facial expression, gestures and speech-based intent as perceived by the robot’s perception system. However, due to the sheer number of routines to select from, the task of choosing the appropriate or the most preferred routine is becoming cumbersome. In this paper, we propose a human-in-the-loop reinforcement learning mechanism in which an agent learns the teleoperator’s selection preference as a function of the input modalities and aids the routine selection process by narrowing it to n-best optimal choices. Our experimental results show that with only a few number of interactions from the teleoperator, the system can learn to recommend optimal routine behaviors for all perceived modalities, which greatly reduces the workload of the teleoperator. Yurii Vasylkiv, Guangliang Li, Eleanor Sandry, Heike Brock, Keisuke Nakamura, Pourang Irani, Randy Gomez |
ICRA | 5 |
| 2021 | Personalization of Human-Robot Gestural Communication through Voice Interaction GroundingabstractIn this paper we develop a gestural communication perception system for a social robot companion that is able to autonomously learn novel gestures on-the-fly. The system constantly tracks human gestural activities with a camera and evaluates the performed gestures under an open-set assumption. This allows for the identification of unknown gestures. Once detected, the system stores motion sequences of the novel gesture class and employs a dialogue interaction with the human to automatically label the unknown gesture. Subsequently, the gestural model is updated, grounding the unknown gesture through dialog interaction. In our experiment, we evaluate a neural network with varying threshold values for the open gesture recognition with unknown detection. Results show that the general classifier reaches an accuracy of more than 83%, and an f1-score of 0.79 in an open-ended scenario. The method is furthermore tested in a first in-lab interaction setting, which shows the system usability and its potential for future personalized human-robot gestural communication. Heike Brock, Randy Gomez |
IROS | 1 |
| 2021 | Developing an Engagement-Aware System for the Detection of Unfocused InteractionabstractWe introduce a perception system for social robots that is able to detect a person’s engagement in an interaction from nonverbal cues independently of principal user activity. This was achieved by the introduction of a set of proxemics, body posture and attention features relevant for human-human interaction. The features were extracted from RGB-D image data of a single Kinect and utilized to train two separate machine learning models. Multiple system configurations and feature combinations were tested, and their impact on the detection of user engagement evaluated. Combining all features, our perception system reaches an F1-score of 81% when estimating an observed person’s interaction intent through binary classification. Regression of a user’s level of availability deviates from the given ground truth values by 13.27% on average. Finally, a prototype was implemented which is able to simultaneously run both previous estimates in real-time using a shared feature vector. In the following, the proposed system shall be used to design robots whose behavior shows their awareness of user engagement. Marvin Brenner, Heike Brock, Andreas Stiegler, Randy Gomez |
RO-MAN | 2 |
| 2021 | Exploring Affective Storytelling with an Embodied AgentabstractIn this paper, we explore the storytelling potential of a robot. We exploit the use of creative contents that maximize the embodied communication affordance of the empathic robot Haru. We identify the elements in storytelling such as narration, agency, engagement and education and synthesized these into the robot. Through effective design we investigated the possible answers that could leverage the limitations and the challenges in developing storytelling applications through a robotic medium. Our preliminary findings show that the use of an embodied agent such as a robot in storytelling only has meaning when its communicative affordance (i.e. embodiment, expressiveness, and other modalities) is tapped, adding new dimension to the experience. Otherwise, traditional storytelling delivery (e.g. tablet) without the use of embodiment will suffice. Hence, robots need to be performers rather than just mere props in storytelling. Randy Gomez, Deborah Szapiro, Kerl Galindo, Luis Merino, Heike Brock, Keisuke Nakamura, Yu Fang 0007, Eric Nichols |
RO-MAN | 5 |
| 2021 | Shaping Affective Robot Haru's Reactive ResponseabstractWe describe a method of teaching a robot its empathic behavioural response from its interaction with people. We used the input modalities such as relative spatial information, facial expressions, body gestures and speech information as perception input that triggers the robot’s empathic response. First, we bootstrap the training through a pre-learning mechanism in which training is conducted by users who know the robotic system. This phase provides simulation-based training using a simple graphical user interface to simulate the input, rewards and correction feedback. In the second phase, we developed an online learning scheme for naive users to personalize their robot further, building on top of the bootstrapped model. Here, we developed a natural user interface that enables natural human-robot interaction via the suite of sensors that allows the users to provide evaluative feedback during the interaction with the robot. We evaluated the system and our results show that bootstrapping is an efficient tool to hasten the robot’s learning while online learning provided some form of personalization in the real environment with naive users. Yurii Vasylkiv, Guangliang Li, Heike Brock, Keisuke Nakamura, Pourang Irani, Randy Gomez |
RO-MAN | 4 |
| 2020 | Robust Real-Time Hand Gestural Recognition for Non-Verbal Communication with Tabletop Robot HaruabstractIn this paper, we present our work in close-distance non-verbal communication with tabletop robot Haru through hand gestural interaction. We implemented a novel hand gestural understanding system by training a machine-learning architecture for real-time hand gesture recognition with the Leap Motion. The proposed system is activated based on the velocity of a user's palm and index finger movement, and subsequently labels the detected movement segments under an early classification scheme. Our system is able to combine multiple gesture labels for recognition of consecutive gestures without clear movement boundaries. System evaluation is conducted on data simulating real human-robot interaction conditions, taking into account relevant performance variables such as movement style, timing and posture. Our results show robustness in hand gesture classification performance under variant conditions. We furthermore examine system behavior under sequential data input, paving the way towards seamless and natural real-time close-distance hand-gestural communication in the future. Heike Brock, Selma Sabanovic, Keisuke Nakamura, Randy Gomez |
RO-MAN | 1 |
| 2020 | Learning Three-dimensional Skeleton Data from Sign Language VideoabstractData for sign language research is often difficult and costly to acquire. We therefore present a novel pipeline able to generate motion three-dimensional (3D) skeleton data from single-camera sign language videos only. First, three recurrent neural networks are learned to infer the three-dimensional position data of body, face, and finger joints for a high resolution of the signer’s skeleton. Subsequently, the angular displacements of all joints over time are estimated using inverse kinematics and mapped to a virtual sign avatar for animation. Last, the generated data are evaluated in detail, including a sign language recognition and sign language synthesis scenario. Utilizing a neural word classifier trained on real motion capture data, we reliably classify word segments built from our newly generated position data with similar accuracy as motion capture data (absolute difference 3.8%). Furthermore, qualitative evaluation of sign animations shows that the avatar performs natural movements that are comprehensible and resemble animations created with original motion capture data. Heike Brock, Felix Law, Kazuhiro Nakadai, Yuji Nagashima |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2019 | Learning Motion Disfluencies for Automatic Sign Language SegmentationabstractWe introduce a novel technique for the automatic detection of word boundaries within continuous sentence expressions in Japanese Sign Language from three-dimensional body joint positions. First, the flow of signed sentence data within a temporal neighborhood is determined utilizing the spatial correlations between line segments of inter-joint pairs. Next, a frame-wise binary random forest classifier is trained to distinguish word and non-word frame content based on the extracted spatio-temporal features. The output of the classifier is used to propose an automatic word synthesis that achieves reliable and accurate sentence segmentation with an average frame-wise F1 score of 0.89. Evaluation with a baseline data set furthermore shows that the proposed approach can easily be adapted to distinguish between motion transitions and motion primitives for a coarse-action domain. Iva Farag, Heike Brock |
ICASSP | 2 |
| 2018 | To animate or anime-te?: Investigating sign avatar comprehensibilityabstractIn this study, we investigated the effect of avatar design on the perception of sign language animation display. Signed sentence expressions were comparatively evaluated using a natural and an anime-style avatar model of three different clothing colors and patterns to determine each design parameter's influence on comprehensibility, naturalness and user affinity. Results show that plain settled color clothing was perceived as most favorable, and that the natural avatar was clearly preferred over the anime-style avatar as an informant of the signed sentence content. Heike Brock, Shigeaki Nishina, Kazuhiro Nakadai |
IVA | 1 |
| 2018 | Deep JSLC: A Multimodal Corpus Collection for Data-driven Generation of Japanese Sign Language Expressions
Heike Brock, Kazuhiro Nakadai |
LREC | 1 |
| 2018 | Data-driven development of Virtual Sign Language Communication AgentsabstractEngaging deaf and hearing people in common discussions requires interfaces to help them understand each other, such as robot agents that translate spoken language into Sign Language (SL) expressions and vice-versa. However, the recognition and generation of signed sentences is a complex task of high dimensionality that cannot be solved in sufficient quality yet. Thus, it is necessary to develop new technologies of improved performances. The sequence to sequence neural network model, traditionally used for machine translation, is adapted to the above two tasks by treating a SL sequence as a multi-dimensional sentence. We defined an encoding of the SL annotations and conducted experiments on the network structure to define a most accurate translation model. This study proves the network trainable and possibly applicable in real-life with an extended dataset, which shall be tested for deployment in virtual translation assistants in the following. Agathe Balayn, Heike Brock, Kazuhiro Nakadai |
RO-MAN | 2 |