Mårten Björkman

dblp:55/6781 · DBLP profile ↗
← Back
60ranked-venue papers
10as first author
29since 2021 · last 2025
0000-0003-0579-3372ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 50 · 10 first-author · 23 since 2021Systems, architecture and hardware · 22 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 5 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 11 · 9 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 6 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2025 REFLEX Dataset: A Multimodal Dataset of Human Reactions to Robot Failures and Explanations
abstract
This work presents REFLEX: Robotic Explanations to FaiLures and Human EXpressions, a comprehensive multimodal dataset capturing human reactions to robot failures and subsequent explanations in collaborative settings. It aims to facilitate research into human-robot interaction dynamics, addressing the need to study reactions to both initial failures and explanations, as well as the evolution of these reactions in long-term interactions. By providing rich, annotated data on human responses to different types of failures, explanation levels, and explanation varying strategies, the dataset contributes to the development of more robust, adaptive, and satisfying robotic systems capable of maintaining positive relationships with human collaborators, even during challenges like repeated failures.
Parag Khanna, Andreas Naoum, Elmira Yadollahi, Mårten Björkman, Christian Smith
HRI4
2025 Human-Aligned Image Models Improve Visual Decoding from the Brain
abstract
Decoding visual images from brain activity has significant potential for advancing brain-computer interaction and enhancing the understanding of human perception. Recent approaches align the representation spaces of images and brain activity to enable visual decoding. In this paper, we introduce the use of human-aligned image encoders to map brain signals to images. We hypothesize that these models more effectively capture perceptual attributes associated with the rapid visual stimuli presentations commonly used in visual brain data recording experiments. Our empirical results support this hypothesis, demonstrating that this simple modification improves image retrieval accuracy by up to 21\% compared to state-of-the-art methods. Comprehensive experiments confirm consistent performance improvements across diverse EEG architectures, image encoders, alignment methods, participants, and brain imaging modalities.
Nona Rajabi, Antônio H. Ribeiro, Miguel Vasco, Farzaneh Taleb, Mårten Björkman, Danica Kragic
ICML5
2025 Automatic Behavior Tree Expansion with LLMs for Robotic Manipulation
abstract
Robotic systems for manipulation tasks are increasingly expected to be easy to configure for new tasks or unpredictable environments, while keeping a transparent policy that is readable and verifiable by humans. We propose the method BEhavior TRee eXPansion with Large Language Models (BETR-XP-LLM) to dynamically and automatically expand and configure Behavior Trees as policies for robot control. The method utilizes an LLM to resolve errors outside the task planner's capabilities, both during planning and execution. We show that the method is able to solve a variety of tasks and failures and permanently update the policy to handle similar problems in the future.
Jonathan Styrud, Matteo Iovino, Mikael Norrlöf, Mårten Björkman, Christian Smith
ICRA4
2025 Domain Randomization for Object Detection in Manufacturing Applications Using Synthetic Data: A Comprehensive Study
abstract
This paper addresses key aspects of domain randomization in generating synthetic data for manufacturing object detection applications. To this end, we present a comprehensive data generation pipeline that reflects different factors: object characteristics, background, illumination, camera settings, and post-processing. We also introduce the Synthetic Industrial Parts Object Detection dataset (SIP15-OD) consisting of 15 objects from three industrial use cases under varying environments as a test bed for the study, while also employing an industrial dataset publicly available for robotic applications. In our experiments, we present more abundant results and insights into the feasibility as well as challenges of sim-toreal object detection. In particular, we identified material properties, rendering methods, post-processing, and distractors as important factors. Our method, leveraging these, achieves top performance on the public dataset with Yolov8 models trained exclusively on synthetic data; mAP@50 scores of 96.4% for the robotics dataset, and 94.1%, 99.5%, and 95.3% across three of the SIP15-OD use cases, respectively. The results showcase the effectiveness of the proposed domain randomization, potentially covering the distribution close to real data for the applications.
Xiaomeng Zhu 0003, Jacob Henningsson, Duruo Li, Pär Mårtensson, Lars Hanson, Mårten Björkman, Atsuto Maki
ICRA6
2025 Adapting Robot's Explanation for Failures Based on Observed Human Behavior in Human-Robot Collaboration
abstract
This work aims to interpret human behavior to anticipate potential user confusion when a robot provides explanations for failure, allowing the robot to adapt its explanations for more natural and efficient collaboration. Using a dataset [1] that included facial emotion detection, eye gaze estimation, and gestures from 55 participants in a user study [2], we analyzed how human behavior changed in response to different types of failures and varying explanation levels. Our goal is to assess whether human collaborators are ready to accept less detailed explanations without inducing confusion. We formulate a data-driven predictor to predict human confusion during robot failure explanations. We also propose and evaluate a mechanism, based on the predictor, to adapt the explanation level according to observed human behavior. The promising results from this evaluation indicate the potential of this research in adapting a robot’s explanations for failures to enhance the collaborative experience.
Andreas Naoum, Parag Khanna, Elmira Yadollahi, Mårten Björkman, Christian Smith
IROS4
2025 YCB-Handovers Dataset: Analyzing Object Weight Impact on Human Handovers to Adapt Robotic Handover Motion
abstract
This paper introduces the YCB-Handovers dataset, capturing motion data of 2771 human-human handovers with varying object weights. The dataset aims to bridge a gap in human-robot collaboration research, providing insights into the impact of object weight in human handovers and readiness cues for intuitive robotic motion planning. The underlying dataset for object recognition and tracking is the YCB (Yale-CMU-Berkeley) Object and Model Set, which is an established standard dataset used in algorithms for robotic manipulation, including grasping and carrying objects. The YCB-Handovers dataset incorporates human motion patterns in handovers, making it applicable for data-driven, human-inspired models aimed at weight-sensitive motion planning and adaptive robotic behaviors. This dataset covers an extensive range of weights, allowing for a more robust study of handover behavior and weight variation. Some objects also require careful handovers, highlighting contrasts with standard handovers. We also provide a detailed analysis of the object’s weight impact on the human reaching motion in these handovers.
Parag Khanna, Karen Jane Dsouza, Mårten Björkman, Christian Smith
RO-MAN4
2025 Mind Meets Robots: A Review of EEG-Based Brain-Robot Interaction Systems
abstract
Brain-robot interaction (BRI) empowers individuals to control (semi-)automated machines through brain activity, either passively or actively. In the past decade, BRI systems have advanced significantly, primarily leveraging electroencephalogram (EEG) signals. This article presents an up-to-date review of 87 curated studies published between 2018 and 2023, identifying the research landscape of EEG-based BRI systems. The review consolidates methodologies, interaction modes, application contexts, system evaluation, existing challenges, and future directions in this domain. Based on our analysis, we propose a BRI system model comprising three entities: Brain, Robot, and Interaction, depicting their internal relationships. We especially examine interaction modes between human brains and robots, an aspect not yet fully explored. Within this model, we scrutinize and classify current research, extract insights, highlight challenges, and offer recommendations for future studies. Our findings provide a structured design space for human-robot interaction (HRI), informing the development of more efficient BRI frameworks.
Yuchong Zhang 0001, Nona Rajabi, Farzaneh Taleb, Andrii Matviienko, Yong Ma 0003, Mårten Björkman, Danica Kragic
Int. J. Hum. Comput. Interact.6
2025 SmartTBD: Smart Tracking for Resource-constrained Object Detection
abstract
With the growing demand for video analysis on mobile devices, object tracking has demonstrated to be a suitable assistance to object detection under the Tracking-By-Detection (TBD) paradigm for reducing computational overhead and power demands. However, performing TBD with fixed hyper-parameters leads to computational inefficiency and ignores perceptual dynamics, as fixed setups tend to run suboptimally, given the variability of scenarios. In this article, we propose SmartTBD, a scheduling strategy for TBD based on multi-objective optimization of accuracy-latency metrics. SmartTBD is a novel deep reinforcement learning based scheduling architecture that computes appropriate TBD configurations in video sequences to improve the speed and detection accuracy. This involves a challenging optimization problem due to the intrinsic relation between the video characteristics and the TBD performance. Therefore, we leverage video characteristics, frame information, and the past TBD results to drive the optimization problem. Our approach surpasses baselines with fixed TBD configurations and recent research, achieving accuracy comparable to pure detection while significantly reducing latency. Moreover, it enables performance analysis of tracking and detection in diverse scenarios. The method is proven to be generalizable and highly practical in common video analytics datasets on resource-constrained devices.
Shihang Zhou, Alejandra C. Hernández, Clara Gómez, Mårten Björkman
ACM Trans. Embed. Comput. Syst.5
2024 Scalable Motion Style Transfer with Constrained Diffusion Generation
abstract
Current training of motion style transfer systems relies on consistency losses across style domains to preserve contents, hindering its scalable application to a large number of domains and private data. Recent image transfer works show the potential of independent training on each domain by leveraging implicit bridging between diffusion models, with the content preservation, however, limited to simple data patterns. We address this by imposing biased sampling in backward diffusion while maintaining the domain independence in the training stage. We construct the bias from the source domain keyframes and apply them as the gradient of content constraints, yielding a framework with keyframe manifold constraint gradients (KMCGs). Our validation demonstrates the success of training separate models to transfer between as many as ten dance motion styles. Comprehensive experiments find a significant improvement in preserving motion contents in comparison to baseline and ablative diffusion-based style transfer models. In addition, we perform a human study for a subjective assessment of the quality of generated dance motions. The results validate the competitiveness of KMCGs.
Yi Yu 0001, Hang Yin 0001, Danica Kragic, Mårten Björkman
AAAI5
2024 Can Transformers Smell Like Humans?
abstract
The human brain encodes stimuli from the environment into representations that form a sensory perception of the world. Despite recent advances in understanding visual and auditory perception, olfactory perception remains an under-explored topic in the machine learning community due to the lack of large-scale datasets annotated with labels of human olfactory perception. In this work, we ask the question of whether pre-trained transformer models of chemical structures encode representations that are aligned with human olfactory perception, i.e., can transformers smell like humans? We demonstrate that representations encoded from transformers pre-trained on general chemical structures are highly aligned with human olfactory perception. We use multiple datasets and different types of perceptual representations to show that the representations encoded by transformer models are able to predict: (i) labels associated with odorants‌‌ provided by experts; (ii) continuous ratings provided by human participants with respect to pre-defined descriptors; and (iii) similarity ratings between odorants provided by human participants. Finally, we evaluate the extent to which this alignment is associated with physicochemical features of odorants known to be relevant for olfactory decoding.
Farzaneh Taleb, Miguel Vasco, Antônio H. Ribeiro, Mårten Björkman, Danica Kragic
NeurIPS4
2023 TD-GEM: Text-Driven Garment Editing Mapper
Reza Dadfar, Sanaz Sabzevari, Mårten Björkman, Danica Kragic
BMVC3
2023 On the Lipschitz Constant of Deep Networks and Double Descent
Matteo Gamba, Hossein Azizpour, Mårten Björkman
BMVC3
2023 How do Humans take an Object from a Robot: Behavior changes observed in a User Study
abstract
To facilitate human-robot interaction and gain human trust, a robot should recognize and adapt to changes in human behavior. This work documents different human behaviors observed while taking objects from an interactive robot in an experimental study, categorized across two dimensions: pull force applied and handedness. We also present the changes observed in human behavior upon repeated interaction with the robot to take various objects.
Parag Khanna, Elmira Yadollahi, Iolanda Leite, Mårten Björkman, Christian Smith
HAI4
2023 Learning Continuous Normalizing Flows For Faster Convergence To Target Distribution via Ascent Regularizations
Sihao Ding 0002, Yiannis Karayiannidis, Mårten Björkman
ICLR4
2023 A Multimodal Data Set of Human Handovers with Design Implications for Human-Robot Handovers
abstract
Handovers are basic yet sophisticated motor tasks performed seamlessly by humans. They are among the most common activities in our daily lives and social environments. This makes mastering the art of handovers critical for a social and collaborative robot. In this work, we present an experimental study that involved human-human handovers by 13 pairs, i.e., 26 participants. We record and explore multiple features of handovers amongst humans aimed at inspiring handovers amongst humans and robots. With this work, we further create and publish a novel data set of 8672 handovers, which includes human motion tracking and the handover-forces. We further analyze the effect of object weight and the role of visual sensory input in human-human handovers, as well as possible design implications for robots. As a proof of concept, the data set was used for creating a human-inspired datadriven strategy for robotic grip release in handovers, which was demonstrated to result in better robot to human handovers.
Parag Khanna, Mårten Björkman, Christian Smith
RO-MAN2
2023 Effects of Explanation Strategies to Resolve Failures in Human-Robot Collaboration
abstract
DH Despite significant improvements in robot capabilities, they are likely to fail in human-robot collaborative tasks due to high unpredictability in human environments and varying human expectations. In this work, we explore the role of explanation of failures by a robot in a human-robot collaborative task. We present a user study incorporating common failures in collaborative tasks with human assistance to resolve the failure. In the study, a robot and a human work together to fill a shelf with objects. Upon encountering a failure, the robot explains the failure and the resolution to overcome the failure, either through handovers or humans completing the task. The study is conducted using different levels of robotic explanation based on the failure action, failure cause, and action history, and different strategies in providing the explanation over the course of repeated interaction. Our results show that the success in resolving the failures is not only a function of the level of explanation but also the type of failures. Furthermore, while novice users rate the robot higher overall in terms of their satisfaction with the explanation, their satisfaction is not only a function of the robot’s explanation level at a certain round but also the prior information they received from the robot.
Parag Khanna, Elmira Yadollahi, Mårten Björkman, Iolanda Leite, Christian Smith
RO-MAN3
2023 Detecting the Intention of Object Handover in Human-Robot Collaborations: An EEG Study
abstract
Human-robot collaboration (HRC) relies on smooth and safe interactions. In this paper, we focus on the human-to-robot handover scenario, where the robot acts as a taker. We investigate the feasibility of detecting the intention of a human-to-robot handover action through the analysis of electroencephalogram (EEG) signals. Our study confirms that temporal patterns in EEG signals provide information about motor planning and can be leveraged to predict the likelihood of an individual executing a motor task with an average accuracy of 94.7%. We also suggest the effectiveness of the time-frequency features of EEG signals in the final second prior to the movement for distinguishing between handover action and other actions. Furthermore, we classify human intentions for different tasks based on time-frequency representations of pre-movement EEG signals and achieve an average accuracy of 63.5% for contrasting every two tasks against each other. The result encourages the possibility of using EEG signals to detect human handover intention in HRC tasks.
Nona Rajabi, Parag Khanna, Sumeyra Demir Kanik, Elmira Yadollahi, Miguel Vasco, Mårten Björkman, Christian Smith, Danica Kragic
RO-MAN6
2023 Controllable Motion Synthesis and Reconstruction with Autoregressive Diffusion Models
abstract
Data-driven and controllable human motion synthesis and prediction are active research areas with various applications in interactive media and social robotics. Challenges remain in these fields for generating diverse motions given past observations and dealing with imperfect poses. This paper introduces MoDiff, an autoregressive probabilistic diffusion model over motion sequences conditioned on control contexts of other modalities. Our model integrates a cross-modal Transformer encoder and a Transformer-based decoder, which are found effective in capturing temporal correlations in motion and control modalities. We also introduce a new data dropout method based on the diffusion forward process to provide richer data representations and robust generation. We demonstrate the superior performance of MoDiff in controllable motion synthesis for locomotion with respect to two baselines and show the benefits of diffusion data dropout for robust synthesis and reconstruction of high-fidelity motion close to recorded data.
Ruibo Tu, Hang Yin 0001, Danica Kragic, Hedvig Kjellström, Mårten Björkman
RO-MAN6
2023 Diffusion-Based Time Series Data Imputation for Cloud Failure Prediction at Microsoft 365
abstract
Ensuring reliability in large-scale cloud systems like Microsoft 365 is crucial. Cloud failures, such as disk and node failure, threaten service reliability, causing service interruptions and financial loss. Existing works focus on failure prediction and proactively taking action before failures happen. However, they suffer from poor data quality, like data missing in model training and prediction, which limits performance. In this paper, we focus on enhancing data quality through data imputation by the proposed Diffusion+, a sample-efficient diffusion model, to impute the missing data efficiently conditioned on the observed data. Experiments with industrial datasets and application practice show that our model contributes to improving the performance of downstream failure prediction.
Fangkai Yang, Lu Wang 0029, Pu Zhao 0004, Bo Liu 0006, Bo Qiao 0001, Mårten Björkman, Saravan Rajmohan, Qingwei Lin, Dongmei Zhang 0001
ESEC/SIGSOFT FSE10
2023 Dance Style Transfer with Cross-modal Transformer
abstract
We present CycleDance, a dance style transfer system to transform an existing motion clip in one dance style to a motion clip in another dance style while attempting to preserve motion context of the dance. Our method extends an existing CycleGAN architecture for modeling audio sequences and integrates multimodal transformer encoders to account for music context. We adopt sequence length-based curriculum learning to stabilize training. Our approach captures rich and long-term intra-relations between motion frames, which is a common challenge in motion transfer and synthesis work. We further introduce new metrics for gauging transfer strength and content preservation in the context of dance movements. We perform an extensive ablation study as well as a human study including 30 participants with 5 or more years of dance experience. The results demonstrate that CycleDance generates realistic movements with the target style, significantly outperforming the baseline CycleGAN on naturalness, transfer strength, and content preservation.1
Hang Yin 0001, Kim Baraka, Danica Kragic, Mårten Björkman
WACV5
2023 Multimodal dance style transfer
abstract
Abstract This paper first presents CycleDance, a novel dance style transfer system that transforms an existing motion clip in one dance style into a motion clip in another dance style while attempting to preserve the motion context of the dance. CycleDance extends existing CycleGAN architectures with multimodal transformer encoders to account for the music context. We adopt a sequence length-based curriculum learning strategy to stabilize training. Our approach captures rich and long-term intra-relations between motion frames, which is a common challenge in motion transfer and synthesis work. Building upon CycleDance, we further propose StarDance, which enables many-to-many mappings across different styles using a single generator network. Additionally, we introduce new metrics for gauging transfer strength and content preservation in the context of dance movements. To evaluate the performance of our approach, we perform an extensive ablation study and a human study with 30 participants, each with 5 or more years of dance experience. Our experimental results show that our approach can generate realistic movements with the target style, outperforming the baseline CycleGAN and its variants on naturalness, transfer strength, and content preservation. Our proposed approach has potential applications in choreography, gaming, animation, and tool development for artistic and scientific innovations in the field of dance.
Hang Yin 0001, Kim Baraka, Danica Kragic, Mårten Björkman
Mach. Vis. Appl.5
2022 Are All Linear Regions Created Equal?
abstract
The number of linear regions has been studied as a proxy of complexity for ReLU networks. However, the empirical success of network compression techniques like pruning and knowledge distillation, suggest that in the overparameterized setting, linear regions density might fail to capture the effective nonlinearity. In this work, we propose an efficient algorithm for discovering linear regions and use it to investigate the effectiveness of density in capturing the nonlinearity of trained VGGs and ResNets on CIFAR-10 and CIFAR-100. We contrast the results with a more principled nonlinearity measure based on function variation, highlighting the shortcomings of linear regions density. Furthermore, interestingly, our measure of nonlinearity clearly correlates with model-wise deep double descent, connecting reduced test error with reduced nonlinearity, and increased local similarity of linear regions.
Matteo Gamba, Adrian Chmielewski-Anders, Josephine Sullivan, Hossein Azizpour, Mårten Björkman
AISTATS5
2022 IL-GAN: Rare Sample Generation via Incremental Learning in GANs
abstract
Industry 4.0 imposes strict requirements on the fifth generation of wireless systems (5G), such as high reliability, high availability, and low latency. Guaranteeing such requirements implies that system failures should occur with an extremely low probability. However, some applications (e.g., training a reinforcement learning algorithm to operate in highly reliable systems or rare event simulations) require access to a broad range of observed failures and extreme values, preferably in a short time. In this paper, we propose IL-GAN, an alternative training framework for generative adversarial networks (GANs), which leverages incremental learning (IL) to enable the generation to learn the tail behavior of the distribution using only a few samples. We validate the proposed IL-GAN with data from 5G simulations on a factory automation scenario and real measurements gathered from various video streaming platforms. Our evaluations show that, compared to the state-of-the-art, our solution can significantly improve the learning and generation performance, not only for the tail distribution but also for the rest of the distribution.
Jón R. Baldvinsson, Milad Ganjalizadeh, Abdulrahman Alabbasi, Mårten Björkman, Amir Hossein Payberah
GLOBECOM4
2022 Combining Planning and Learning of Behavior Trees for Robotic Assembly
abstract
Industrial robots can solve tasks in controlled environments, but modern applications require robots able to operate also in unpredictable surroundings. An increasingly popular reactive policy architecture in robotics is Behavior Trees (BTs) but as other architectures, programming time drives cost and limits flexibility. The two main branches of algorithms to generate policies automatically, automated planning and machine learning, both have their own drawbacks and have not previously been combined for generation of BTs. We propose a method for creating BTs by combining these branches, inserting the result of an automated planner into the population of a Genetic Programming algorithm. Experiments confirm that the proposed method performs well on a variety of robotic assembly problems and outperforms the base methods used separately. We also show that this high level learning of Behavior Trees can be transferred to a real system without further training.
Jonathan Styrud, Matteo Iovino, Mikael Norrlöf, Mårten Björkman, Christian Smith
ICRA4
2022 Training and Evaluation of Deep Policies Using Reinforcement Learning and Generative Models
abstract
We present a data-efficient framework for solving sequential decision-making problems which exploits the combination of reinforcement learning (RL) and latent variable generative models. The framework, called GenRL, trains deep policies by introducing an action latent variable such that the feed-forward policy search can be divided into two parts: (i) training a sub-policy that outputs a distribution over the action latent variable given a state of the system, and (ii) unsupervised training of a generative model that outputs a sequence of motor actions conditioned on the latent action variable. GenRL enables safe exploration and alleviates the data-inefficiency problem as it exploits prior knowledge about valid sequences of motor actions. Moreover, we provide a set of measures for evaluation of generative models such that we are able to predict the performance of the RL policy training prior to the actual training on a physical robot. We experimentally determine the characteristics of generative models that have most influence on the performance of the final policy training on two robotics tasks: shooting a hockey puck and throwing a basketball. Furthermore, we empirically demonstrate that GenRL is the only method which can safely and efficiently solve the robotics tasks compared to two state-of-the-art RL methods.
Ali Ghadirzadeh, Petra Poklukar, Karol Arndt, Chelsea Finn, Ville Kyrki, Danica Kragic, Mårten Björkman
J. Mach. Learn. Res.7
2022 In Memoriam: Jan-Olof Eklundh
Atsuto Maki, Danica Kragic, Hedvig Kjellström, Hossein Azizpour, Josephine Sullivan, Mårten Björkman, Patric Jensfelt, Stefan Carlsson, Tony Lindeberg, Yngve Sundblad
IEEE Trans. Pattern Anal. Mach. Intell.6
2021 Monte Carlo Filtering Objectives
Sihao Ding 0002, Yiannis Karayiannidis, Mårten Björkman
IJCAI4
2021 Bayesian Meta-Learning for Few-Shot Policy Adaptation Across Robotic Platforms
abstract
Reinforcement learning methods can achieve significant performance but require a large amount of training data collected on the same robotic platform. A policy trained with expensive data is rendered useless after making even a minor change to the robot hardware. In this paper, we address the challenging problem of adapting a policy, trained to perform a task, to a novel robotic hardware platform given only few demonstrations of robot motion trajectories on the target robot. We formulate it as a few-shot meta-learning problem where the goal is to find a meta-model that captures the common structure shared across different robotic platforms such that data-efficient adaptation can be performed. We achieve such adaptation by introducing a learning framework consisting of a probabilistic gradient-based meta-learning algorithm that models the uncertainty arising from the few-shot setting with a low-dimensional latent variable. We experimentally evaluate our framework on a simulated reaching and a real-robot picking task using 400 simulated robots generated by varying the physical parameters of an existing set of robotic platforms. Our results show that the proposed method can successfully adapt a trained policy to different robotic platforms with novel physical parameters and the superiority of our meta-learning algorithm compared to state-of-the-art methods for the introduced few-shot policy adaptation problem.
Ali Ghadirzadeh, Xi Chen 0051, Petra Poklukar, Chelsea Finn, Mårten Björkman, Danica Kragic
IROS5
2021 Graph-based Normalizing Flow for Human Motion Generation and Reconstruction
abstract
Data-driven approaches for modeling human skeletal motion have found various applications in interactive media and social robotics. Challenges remain in these fields for generating high-fidelity samples and robustly reconstructing motion from imperfect input data, due to e.g. missed marker detection. In this paper, we propose a probabilistic generative model to synthesize and reconstruct long horizon motion sequences conditioned on past information and control signals, such as the path along which an individual is moving. Our method adapts the existing work MoGlow by introducing a new graph-based model. The model leverages the spatial-temporal graph convolutional network (ST-GCN) to effectively capture the spatial structure and temporal correlation of skeletal motion data at multiple scales. We evaluate the models on a mixture of motion capture datasets of human locomotion with foot-step and bone-length analysis. The results demonstrate the advantages of our model in reconstructing missing markers and achieving comparable results on generating realistic future poses. When the inputs are imperfect, our model shows improvements on robustness of generation.
Hang Yin 0001, Danica Kragic, Mårten Björkman
RO-MAN4
2020 Group Behavior Recognition Using Attention- and Graph-Based Neural Networks
Fangkai Yang, Tetsunari Inamura, Mårten Björkman, Christopher Peters 0001
ECAI4
2020 Adversarial Feature Training for Generalizable Robotic Visuomotor Control
abstract
Deep reinforcement learning (RL) has enabled training action-selection policies, end-to-end, by learning a function which maps image pixels to action outputs. However, it's application to visuomotor robotic policy training has been limited because of the challenge of large-scale data collection when working with physical hardware. A suitable visuomotor policy should perform well not just for the task-setup it has been trained for, but also for all varieties of the task, including novel objects at different viewpoints surrounded by task-irrelevant objects. However, it is impractical for a robotic setup to sufficiently collect interactive samples in a RL framework to generalize well to novel aspects of a task. In this work, we demonstrate that by using adversarial training for domain transfer, it is possible to train visuomotor policies based on RL frameworks, and then transfer the acquired policy to other novel task domains. We propose to leverage the deep RL capabilities to learn complex visuomotor skills for uncomplicated task setups, and then exploit transfer learning to generalize to new task domains provided only still images of the task in the target domain. We evaluate our method on two real robotic tasks, picking and pouring, and compare it to a number of prior works, demonstrating its superiority.
Xi Chen 0051, Ali Ghadirzadeh, Mårten Björkman, Patric Jensfelt
ICRA3
2020 Amortized Variational Inference for Road Friction Estimation
abstract
Road friction estimation concerns inference of the coefficient between the tire and road surface to facilitate active safety features. Current state-of-the-art methods lack generalization capability to cope with different tire characteristics and models are restricted when using Bayesian inference in estimation while recent supervised learning methods lack uncertainty prediction on estimates. This paper introduces variational inference to approximate intractable posterior of friction estimates and learns an amortized variational inference model from tire measurement data to facilitate probabilistic estimation while sustaining the flexibility of tire models. As a by-product, a probabilistic tire model can be learned jointly with friction estimator model. Experiments on simulated and field test data show that the learned friction estimator provides accurate estimates with robust uncertainty measures in a wide range of tire excitation levels. Meanwhile, the learned tire model reflects well-studied tire characteristics from field test data.
Sihao Ding 0002, L. Srikar Muppirisetty, Yiannis Karayiannidis, Mårten Björkman
IV5
2020 Impact of Trajectory Generation Methods on Viewer Perception of Robot Approaching Group Behaviors
abstract
Mobile robots that approach free-standing conversational groups to join them should behave in a safe and socially-acceptable way. Existing trajectory generation methods focus on collision avoidance with pedestrians, and the models that generate approach behaviors into groups are evaluated in simulation. However, it is challenging to generate approach and join trajectories that avoid collisions with group members while also ensuring that they do not invoke feelings of discomfort. In this paper, we conducted an experiment to examine the impact of three trajectory generation methods for a mobile robot to approach groups from multiple directions: a Wizard-of-Oz (WoZ) method, a procedural social-aware navigation model (PM) and a novel generative adversarial model imitating human approach behaviors (IL). Measures also compared two camera viewpoints and static versus quasi-dynamic groups. The latter refers to a group whose members change orientation and position throughout the approach task, even though the group entity remains static in the environment. This represents a more realistic but challenging scenario for the robot. We evaluate three methods with objective measurements and subjective measurements from viewer perception, and results show that WoZ and IL have comparable performance, and both perform better than PM under most conditions.
Fangkai Yang, Mårten Björkman, Christopher Peters 0001
RO-MAN3
2019 Meta-Learning for Multi-objective Reinforcement Learning
abstract
Multi-objective reinforcement learning (MORL) is the generalization of standard reinforcement learning (RL) approaches to solve sequential decision making problems that consist of several, possibly conflicting, objectives. Generally, in such formulations, there is no single optimal policy which optimizes all the objectives simultaneously, and instead, a number of policies has to be found each optimizing a preference of the objectives. In this paper, we introduce a novel MORL approach by training a meta-policy, a policy simultaneously trained with multiple tasks sampled from a task distribution, for a number of randomly sampled Markov decision processes (MDPs). In other words, the MORL is framed as a meta-learning problem, with the task distribution given by a distribution over the preferences. We demonstrate that such a formulation results in a better approximation of the Pareto optimal solutions in terms of both the optimality and the computational efficiency. We evaluated our method on obtaining Pareto optimal policies using a number of continuous control problems with high degrees of freedom.
Xi Chen 0051, Ali Ghadirzadeh, Mårten Björkman, Patric Jensfelt
IROS3
2018 Deep Reinforcement Learning to Acquire Navigation Skills for Wheel-Legged Robots in Complex Environments
abstract
Mobile robot navigation in complex and dynamic environments is a challenging but important problem. Reinforcement learning approaches fail to solve these tasks efficiently due to reward sparsities, temporal complexities and high-dimensionality of sensorimotor spaces which are inherent in such problems. We present a novel approach to train action policies to acquire navigation skills for wheel-legged robots using deep reinforcement learning. The policy maps height-map image observations to motor commands to navigate to a target position while avoiding obstacles. We propose to acquire the multifaceted navigation skill by learning and exploiting a number of manageable navigation behaviors. We also introduce a domain randomization technique to improve the versatility of the training samples. We demonstrate experimentally a significant improvement in terms of data-efficiency, success rate, robustness against irrelevant sensory data, and also the quality of the maneuver skills.
Xi Chen 0051, Ali Ghadirzadeh, John Folkesson, Mårten Björkman, Patric Jensfelt
IROS4
2017 Deep predictive policy training using reinforcement learning
abstract
Skilled robot task learning is best implemented by predictive action policies due to the inherent latency of sensorimotor processes. However, training such predictive policies is challenging as it involves finding a trajectory of motor activations for the full duration of the action. We propose a data-efficient deep predictive policy training (DPPT) framework with a deep neural network policy architecture which maps an image observation to a sequence of motor activations. The architecture consists of three sub-networks referred to as the perception, policy and behavior super-layers. The perception and behavior super-layers force an abstraction of visual and motor data trained with synthetic and simulated training samples, respectively. The policy super-layer is a small subnetwork with fewer parameters that maps data in-between the abstracted manifolds. It is trained for each task using methods for policy search reinforcement learning. We demonstrate the suitability of the proposed architecture and learning framework by training predictive policies for skilled object grasping and ball throwing on a PR2 robot. The effectiveness of the method is illustrated by the fact that these tasks are trained using only about 180 real robot attempts with qualitative terminal rewards.
Ali Ghadirzadeh, Atsuto Maki, Danica Kragic, Mårten Björkman
IROS4
2016 Self-learning and adaptation in a sensorimotor framework
abstract
We present a general framework to autonomously achieve the task of finding a sequence of actions that result in a desired state. Autonomy is acquired by learning sensorimotor patterns of a robot, while it is interacting with its environment.
Ali Ghadirzadeh, Judith Bütepage, Danica Kragic, Mårten Björkman
ICRA4
2016 A sensorimotor reinforcement learning framework for physical Human-Robot Interaction
abstract
Modeling of physical human-robot collaborations is generally a challenging problem due to the unpredictive nature of human behavior. To address this issue, we present a data-efficient reinforcement learning framework which enables a robot to learn how to collaborate with a human partner. The robot learns the task from its own sensorimotor experiences in an unsupervised manner. The uncertainty in the interaction is modeled using Gaussian processes (GP) to implement a forward model and an action-value function. Optimal action selection given the uncertain GP model is ensured by Bayesian optimization. We apply the framework to a scenario in which a human and a PR2 robot jointly control the ball position on a plank based on vision and force/torque data. Our experimental results show the suitability of the proposed method in terms of fast and data-efficient model learning, optimal action selection under uncertainty and equal role sharing between the partners.
Ali Ghadirzadeh, Judith Bütepage, Atsuto Maki, Danica Kragic, Mårten Björkman
IROS5
2015 A sensorimotor approach for self-learning of hand-eye coordination
abstract
This paper presents a sensorimotor contingencies (SMC) based method to fully autonomously learn to perform hand-eye coordination. We divide the task into two visuomotor subtasks, visual fixation and reaching, and implement these on a PR2 robot assuming no prior information on its kinematic model. Our contributions are three-fold: i) grounding a robot in the environment by exploiting SMCs in the action planning system, which eliminates the need for prior knowledge of the kinematic or dynamic models of the robot; ii) using a forward model to search for proper actions to solve the task by minimizing a cost function, instead of training a separate inverse model, to speed up training; iii) encoding 3D spatial positions of a target object based on the robot's joint positions, thus avoiding calibration with respect to an external coordinate system. The method is capable of learning the task of hand-eye coordination from scratch by less than 20 sensory-motor pairs that are iteratively generated at real-time speed. In order to examine the robustness of the method while dealing with nonlinear image distortions, we apply a so-called retinal mapping image deformation to the input images. Experimental results show the successfulness of the method even under considerable image deformations.
Ali Ghadirzadeh, Atsuto Maki, Mårten Björkman
IROS3
2014 Learning visual forward models to compensate for self-induced image motion
abstract
Predicting the sensory consequences of an agent's own actions is considered an important skill for intelligent behavior. In terms of vision, so-called visual forward models can be applied to learn such predictions. This is no trivial task given the high-dimensionality of sensory data and complex action spaces. In this work, we propose to learn the visual consequences of changes in pan and tilt of a robotic head using a visual forward model based on Gaussian processes and SURF correspondences. This is done without any assumptions on the kinematics of the system or requirements on calibration. The proposed method is compared to an earlier work using accumulator-based correspondences and Radial Basis function networks. We also show the feasibility of the proposed method for detection of independent motion using a moving camera system. By comparing the predicted and actual captured images, image motion due to the robot's own actions and motion caused by moving external objects can be distinguished. Results show the proposed method to be preferable from the earlier method in terms of both prediction errors and ability to detect independent motion.
Ali Ghadirzadeh, Gert Kootstra, Atsuto Maki, Mårten Björkman
RO-MAN4
2014 Detecting, segmenting and tracking unknown objects using multi-label MRF inference
Mårten Björkman, Niklas Bergström, Danica Kragic
Comput. Vis. Image Underst.1
2013 Enhancing visual perception of shape through tactile glances
abstract
Object shape information is an important parameter in robot grasping tasks. However, it may be difficult to obtain accurate models of novel objects due to incomplete and noisy sensory measurements. In addition, object shape may change due to frequent interaction with the object (cereal boxes, etc). In this paper, we present a probabilistic approach for learning object models based on visual and tactile perception through physical interaction with an object. Our robot explores unknown objects by touching them strategically at parts that are uncertain in terms of shape. The robot starts by using only visual features to form an initial hypothesis about the object shape, then gradually adds tactile measurements to refine the object model. Our experiments involve ten objects of varying shapes and sizes in a real setup. The results show that our method is capable of choosing a small number of touches to construct object models similar to real object shapes and to determine similarities among acquired models.
Mårten Björkman, Yasemin Bekiroglu, Virgile Hogman, Danica Kragic
IROS1
2013 Interactive object classification using sensorimotor contingencies
abstract
Understanding and representing objects and their function is a challenging task. Objects we manipulate in our daily activities can be described and categorized in various ways according to their properties or affordances, depending also on our perception of those. In this work, we are interested in representing the knowledge acquired through interaction with objects, describing these in terms of action-effect relations, i.e. sensorimotor contingencies, rather than static shape or appearance representations. We demonstrate how a robot learns sensorimotor contingencies through pushing using a probabilistic model. We show how functional categories can be discovered and how entropy-based action selection can improve object classification.
Virgile Hogman, Mårten Björkman, Danica Kragic
IROS2
2012 YES - YEt another object segmentation: Exploiting camera movement
abstract
We address the problem of object segmentation in image sequences where no a-priori knowledge of objects is assumed. We take advantage of robots' ability to move, gathering multiple images of the scene. Our approach starts by extracting edges, uses a polar domain representation and performs integration over time based on a simple dilation operation. The proposed system can be used for providing reliable initial segmentation of unknown objects in scenes of varying complexity, allowing for recognition, categorization or physical interaction with the objects. The experimental evaluation on both self-captured and a publicly available dataset shows the efficiency and stability of the proposed method.
Lazaros Nalpantidis, Mårten Björkman, Danica Kragic
IROS2
2011 Scene Understanding through Autonomous Interactive Perception
Niklas Bergström, Carl Henrik Ek, Mårten Björkman, Danica Kragic
ICVS3
2011 Generating object hypotheses in natural scenes through human-robot interaction
abstract
We propose a method for interactive modeling of objects and object relations based on real-time segmentation of video sequences. In interaction with a human, the robot can perform multi-object segmentation through principled modeling of physical constraints. The key contribution is an efficient multi-labeling framework, that allows object modeling and disambiguation in natural scenes. Object modeling and labeling is done in a real-time segmentation system, to which hypotheses and constraints denoting relations between objects can be added incrementally. Through instructions such as key presses or spoken words, a scene can be segmented in regions corresponding to multiple physical objects. The approach solves some of the difficult problems related to disambiguation of objects merged due to their direct physical contact. Results show that even a limited set of simple interactions with a human operator can substantially improve segmentation results.
Niklas Bergström, Mårten Björkman, Danica Kragic
IROS2
2010 Active 3D Segmentation through Fixation of Previously Unseen Objects
abstract
We present an approach for active segmentation based on integration of several cues.It serves as a framework for generation of object hypotheses of previously unseen objectsin natural scenes. Using an approximate Expectation-Maximisation method, the appearance,3D shape and size of objects are modelled in an iterative manner, with fixation usedfor unsupervised initialisation. To better cope with situations where an object is hard tosegregate from the surface it is placed on, a flat surface model is added to the typical twohypotheses used in classical figure-ground segmentation. The framework is further extendedto include modelling over time, in order to exploit temporal consistency for bettersegmentation and to facilitate tracking.
Mårten Björkman, Danica Kragic
BMVC1
2010 Active 3D scene segmentation and detection of unknown objects
abstract
We present an active vision system for segmentation of visual scenes based on integration of several cues. The system serves as a visual front end for generation of object hypotheses for new, previously unseen objects in natural scenes. The system combines a set of foveal and peripheral cameras where, through a stereo based fixation process, object hypotheses are generated. In addition to considering the segmentation process in 3D, the main contribution of the paper is integration of different cues in a temporal framework and improvement of initial hypotheses over time.
Mårten Björkman, Danica Kragic
ICRA1
2010 Strategies for multi-modal scene exploration
abstract
We propose a method for multi-modal scene exploration where initial object hypothesis formed by active visual segmentation are confirmed and augmented through haptic exploration with a robotic arm. We update the current belief about the state of the map with the detection results and predict yet unknown parts of the map with a Gaussian Process. We show that through the integration of different sensor modalities, we achieve a more complete scene model. We also show that the prediction of the scene structure leads to a valid scene representation even if the map is not fully traversed. Furthermore, we propose different exploration strategies and evaluate them both in simulation and on our robotic platform.
Jeannette Bohg, Matthew Johnson-Roberson, Mårten Björkman, Danica Kragic
IROS3
2010 Attention-based active 3D point cloud segmentation
abstract
In this paper we present a framework for the segmentation of multiple objects from a 3D point cloud. We extend traditional image segmentation techniques into a full 3D representation. The proposed technique relies on a state-of-the-art min-cut framework to perform a fully 3D global multi-class labeling in a principled manner. Thereby, we extend our previous work in which a single object was actively segmented from the background. We also examine several seeding methods to bootstrap the graphical model-based energy minimization and these methods are compared over challenging scenes. All results are generated on real-world data gathered with an active vision robotic head. We present quantitive results over aggregate sets as well as visual results on specific examples.
Matthew Johnson-Roberson, Jeannette Bohg, Mårten Björkman, Danica Kragic
IROS3
2008 Integration of Visual and Shape Attributes for Object Action Complexes
Kai Huebner, Mårten Björkman, Babak Rasolzadeh, Martina Schmidt, Danica Kragic
ICVS2
2006 A Framework for Vision Based bearing only 3D SLAM
abstract
This paper presents a framework for 3D vision based bearing only SLAM using a single camera, an interesting setup for many real applications due to its low cost. The focus in is on the management of the features to achieve real-time performance in extraction, matching and loop detection. For matching image features to map landmarks a modified, rotationally variant SIFT descriptor is used in combination with a Harris-Laplace detector. To reduce the complexity in the map estimation while maintaining matching performance only a few, high quality, image features are used for map landmarks. The rest of the features are used for matching. The framework has been combined with an EKF implementation for SLAM. Experiments performed in indoor environments are presented. These experiments demonstrate the validity and effectiveness of the approach. In particular they show how the robot is able to successfully match current image features to the map when revisiting an area
Patric Jensfelt, Danica Kragic, John Folkesson, Mårten Björkman
ICRA4
2006 Strategies for Object Manipulation using Foveal and Peripheral Vision
abstract
Visual feedback is used extensively in robotics and application areas range from human-robot interaction to object grasping and manipulation. There have been a number of examples of how to develop different components required by the above applications and very few general vision systems capable of performing a variety of tasks. In this paper, we concentrate on vision strategies for robotic manipulation tasks in a domestic environment. In particular, given fetch-and-carry type of tasks, the issues related to the whole detect-approach-grasp loop are considered. We deal with the problem of flexibility and robustness by using monocular and binocular visual cues and their integration. We demonstrate real-time disparity estimation, object recognition and pose estimation. We also show how a combination of foveal and peripheral vision system can be combined in order to provide a wide, low resolution and narrow, high resolution field of view.
Danica Kragic, Mårten Björkman
ICVS2
2005 Foveated Figure-Ground Segmentation and Its Role in Recognition
abstract
Figure-ground segmentation and recognition are two interrelated processes. In this paper we present a method for foveated segmentation and evaluate it in the context of a binocular real-time recognition system. Segmentation is solved as a binary labeling problem using priors derived from the results ofa simplistic disparity method. Doing so we are able to cope with situations when the disparity range is very wide, situations that has rarely been considered, but appear frequently for narrow-field camera sets. Segmentation and recognition are then integrated into a system able to locate, attend to and recognise objects in typical cluttered indoor scenes. Finally, we try to answer two questions: is recognition really helped by segmentation and what is the benefit of multiple cues for recognition?
Mårten Björkman, Jan-Olof Eklundh
BMVC1
2004 Attending, Foveating and Recognizing Objects in Real World Scenes
abstract
Recognition in cluttered real world scenes is a challenging problem. To find a particular object of interest within a reasonable time, a wide field of view is preferable. However, as we will show with practical experiments, robust recognition is easier if the object is foveated and subtends a considerable partof the visual field. In this paper a binocular system able to overcome these two conflicting requirements will be presented. The system consists of two sets of cameras, a wide field pair and a foveal one. From disparities a number of object hypotheses are generated. An attentional process based on hue and 3D size guides the foveal cameras towards the most salient regions. With the object foveated and segmented in 3D, recognition is performed using scale invariant features. The system is fully automised and runs at real-time speed.
Mårten Björkman, Jan-Olof Eklundh
BMVC1
2004 Combination of Foveal and Peripheral Vision for Object Recognition and Pose Estimation
abstract
In this paper, we present a real-time vision system that integrates a number of algorithms using monocular and binocular cues to achieve robustness in realistic settings, for tasks such as object recognition, tracking and pose estimation. The system consists of two sets of binocular cameras; a peripheral set for disparity based attention and a foveal one for higher level processes. Thus the conflicting requirements of a wide field of view and high resolution can be overcome. One important property of the system is that the step from task specification through object recognition to pose estimation is completely automatic, combining both appearance and geometric models. Experimental evaluation is performed in a realistic indoor environment with occlusions, clutter, changing lighting and background conditions.
Mårten Björkman, Danica Kragic
ICRA1
2002 Real-Time Epipolar Geometry Estimation of Binocular Stereo Heads
abstract
Stereo is an important cue for visually guided robots. While moving around in the world, such a robot can use dynamic fixation to overcome limitations in image resolution and field of view. In this paper, a binocular stereo system capable of dynamic fixation is presented. The external calibration is performed continuously taking temporal consistency into consideration, greatly simplifying the process. The essential matrix, which is estimated in real-time, is used to describe the epipolar geometry. It will be shown, how outliers can be identified and excluded from the calculations. An iterative approach based on a differential model of the optical flow, commonly used in structure from motion, is also presented and tested towards the essential matrix. The iterative method will be shown to be superior in terms of both computational speed and robustness, when the vergence angles are less than about 15. For larger angles, the differential model is insufficient and the essential matrix is preferably used instead.
Mårten Björkman, Jan-Olof Eklundh
IEEE Trans. Pattern Anal. Mach. Intell.1
2000 A Real-Time System for Epipolar Geometry and Ego-Motion Estimation
abstract
A visually guided mobile platform acting in a dynamic environment needs information, delivered from cues such as stereo and motion. In this paper a binocular stereo system, capable of dynamic vergence and ego-motion estimation, will be presented. The ability to dynamically verge the cameras is essential for visual tasks such as tracking, recognition and manipulation, since the computation of 3D geometry will be considerably simplified at fixation. We will show that the epipolar geometry of the presented system can be estimated in real-time. The number of degrees of freedom has been minimized, without affecting the flexibility of the system. Reconstructed 3D data is used in order to determine the ego-motion of the platform, at a very low computational cost. Finally, it will be shown how disparity maps and estimated ego-motion can facilitate the localization of independently moving objects.
Mårten Björkman, Jan-Olof Eklundh
CVPR1
1999 Real-Time Epipolar Geometry Estimation and Disparity
abstract
Many visual tasks, such as recognition and visually guided manipulation, benefit from the use of binocular stereo with dynamic vergence. In order to achieve this the epipolar geometry of the two cameras has to be known. In this article we present a system capable of estimating the epipolar geometry of a stereo-head and calculating the binocular disparities in real-time producing dense disparity maps that, for example, can be used for figure-ground segmentation. The number of degrees of freedom has been minimized, without affecting the flexibility of the stereo-head, speeding up the process considerably.
Mårten Björkman, Jan-Olof Eklundh
ICCV1
1997 Reducing the Read-Miss Penalty for Flat COMA Protocols
abstract
In flat cache-only memory architectures (COMA), an attraction-memory miss must first interrogate a directory before a copy of the requested data can be located, which often involves three network traversals. By keeping track of the identity of a potential holder of the copy–called a hint–one network traversal can be saved which reduces the read penalty. We have evaluated the reduction of the read-miss penalty provided by hints using detailed architectural simulations and four benchmark applications. The results show that a previously proposed protocol using hints can actually make the read-miss penalty larger because when the hint is not correct, an extra network traversal is needed. This has motivated us to study a new protocol, using hints, that simultaneously sends a request to the potential holder and to the directory. This protocol reduces the read-miss penalty for all applications but the protocol complexity does not seem to justify the performance improvement.
Fredrik Dahlgren, Per Stenström, Mårten Björkman
Comput. J.3