Hang Yin 0001

dblp:81/7626-1 · DBLP profile ↗
← Back
22ranked-venue papers
4as first author
16since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 4 first-author · 13 since 2021Systems, architecture and hardware · 9 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Towards Safe Reinforcement Learning with Reduced Conservativeness: A Case Study on Drone Flight Control
abstract
Incorporating formal methods into reinforcement learning (RL) has the potential to result in the best of both worlds, combining the robustness of formal guarantees with the adaptability and learning capabilities of RL, though careful design is needed to balance safety and exploration. In this work, we propose a framework to mitigate this loss of exploration while still allowing for the safety of the system to be ensured. Specifically, we introduce a less restrictive method that can reduce the conservativeness of formal methods by refining a disturbance model using online collected data and it evaluates the safety of a learning-based controller, using computationally efficient zonotopic reachability analysis for the safety analysis to facilitate a real-time implementation. We validate the framework in a real-world drone flight through a canyon, where the drone is subjected to unknown external disturbances and the framework is tasked with learning those disturbances online and adjusting the safety guarantees accordingly. The results show that the framework enables a less restrictive online training of learning-based controllers without compromising the safety of the system.
Loizos Hadjiloizou, Michael C. Welle, Hang Yin 0001, Danica Kragic
IROS3
2024 Scalable Motion Style Transfer with Constrained Diffusion Generation
abstract
Current training of motion style transfer systems relies on consistency losses across style domains to preserve contents, hindering its scalable application to a large number of domains and private data. Recent image transfer works show the potential of independent training on each domain by leveraging implicit bridging between diffusion models, with the content preservation, however, limited to simple data patterns. We address this by imposing biased sampling in backward diffusion while maintaining the domain independence in the training stage. We construct the bias from the source domain keyframes and apply them as the gradient of content constraints, yielding a framework with keyframe manifold constraint gradients (KMCGs). Our validation demonstrates the success of training separate models to transfer between as many as ten dance motion styles. Comprehensive experiments find a significant improvement in preserving motion contents in comparison to baseline and ablative diffusion-based style transfer models. In addition, we perform a human study for a subjective assessment of the quality of generated dance motions. The results validate the competitiveness of KMCGs.
Yi Yu 0001, Hang Yin 0001, Danica Kragic, Mårten Björkman
AAAI3
2024 ADAPT: AI-Driven Artefact Purging Technique for IMU Based Motion Capture
abstract
Abstract While IMU based motion capture offers a cost‐effective alternative to premium camera‐based systems, it often falls short in matching the latter's realism. Common distortions, such as self‐penetrating body parts, foot skating, and floating, limit the usability of these systems, particularly for high‐end users. To address this, we employed reinforcement learning to train an AI agent that mimics erroneous sample motion. Since our agent operates within a simulated environment, it inherently avoids generating these distortions since it must adhere to the laws of physics. Impressively, the agent manages to mimic the sample motions while preserving their distinctive characteristics. We assessed our method's efficacy across various types of input data, showcasing an ideal blend of artefact‐laden IMU‐based data with high‐grade optical motion capture data. Furthermore, we compared the configuration of observation and action spaces with other implementations, pinpointing the most suitable configuration for our purposes. All our models underwent rigorous evaluation using a spectrum of quantitative metrics complemented by a qualitative review. These evaluations were performed using a benchmark dataset of IMU‐based motion data from actors not included in the training data.
Paul Schreiner, Rasmus Netterstrøm, Hang Yin 0001, Sune Darkner, Kenny Erleben
Comput. Graph. Forum3
2023 Learning Geometric Representations of Objects via Interaction
Alfredo Reichlin, Giovanni Luca Marchetti, Hang Yin 0001, Anastasiia Varava, Danica Kragic
ECML/PKDD (4)3
2023 Controllable Motion Synthesis and Reconstruction with Autoregressive Diffusion Models
abstract
Data-driven and controllable human motion synthesis and prediction are active research areas with various applications in interactive media and social robotics. Challenges remain in these fields for generating diverse motions given past observations and dealing with imperfect poses. This paper introduces MoDiff, an autoregressive probabilistic diffusion model over motion sequences conditioned on control contexts of other modalities. Our model integrates a cross-modal Transformer encoder and a Transformer-based decoder, which are found effective in capturing temporal correlations in motion and control modalities. We also introduce a new data dropout method based on the diffusion forward process to provide richer data representations and robust generation. We demonstrate the superior performance of MoDiff in controllable motion synthesis for locomotion with respect to two baselines and show the benefits of diffusion data dropout for robust synthesis and reconstruction of high-fidelity motion close to recorded data.
Ruibo Tu, Hang Yin 0001, Danica Kragic, Hedvig Kjellström, Mårten Björkman
RO-MAN3
2023 Dance Style Transfer with Cross-modal Transformer
abstract
We present CycleDance, a dance style transfer system to transform an existing motion clip in one dance style to a motion clip in another dance style while attempting to preserve motion context of the dance. Our method extends an existing CycleGAN architecture for modeling audio sequences and integrates multimodal transformer encoders to account for music context. We adopt sequence length-based curriculum learning to stabilize training. Our approach captures rich and long-term intra-relations between motion frames, which is a common challenge in motion transfer and synthesis work. We further introduce new metrics for gauging transfer strength and content preservation in the context of dance movements. We perform an extensive ablation study as well as a human study including 30 participants with 5 or more years of dance experience. The results demonstrate that CycleDance generates realistic movements with the target style, significantly outperforming the baseline CycleGAN on naturalness, transfer strength, and content preservation.1
Hang Yin 0001, Kim Baraka, Danica Kragic, Mårten Björkman
WACV2
2023 Multimodal dance style transfer
abstract
Abstract This paper first presents CycleDance, a novel dance style transfer system that transforms an existing motion clip in one dance style into a motion clip in another dance style while attempting to preserve the motion context of the dance. CycleDance extends existing CycleGAN architectures with multimodal transformer encoders to account for the music context. We adopt a sequence length-based curriculum learning strategy to stabilize training. Our approach captures rich and long-term intra-relations between motion frames, which is a common challenge in motion transfer and synthesis work. Building upon CycleDance, we further propose StarDance, which enables many-to-many mappings across different styles using a single generator network. Additionally, we introduce new metrics for gauging transfer strength and content preservation in the context of dance movements. To evaluate the performance of our approach, we perform an extensive ablation study and a human study with 30 participants, each with 5 or more years of dance experience. Our experimental results show that our approach can generate realistic movements with the target style, outperforming the baseline CycleGAN and its variants on naturalness, transfer strength, and content preservation. Our proposed approach has potential applications in choreography, gaming, animation, and tool development for artistic and scientific innovations in the field of dance.
Hang Yin 0001, Kim Baraka, Danica Kragic, Mårten Björkman
Mach. Vis. Appl.2
2023 Enabling Visual Action Planning for Object Manipulation Through Latent Space Roadmap
abstract
In this article, we present a framework for visual action planning of complex manipulation tasks with high-dimensional state spaces, focusing on manipulation of deformable objects. We propose a latent space roadmap (LSR) for task planning, which is a graph-based structure globally capturing the system dynamics in a low-dimensional latent space. Our framework consists of the following three parts. First, a mapping module (MM) that maps observations is given in the form of images into a structured latent space extracting the respective states as well as generates observations from the latent states. Second, the LSR, which builds and connects clusters containing similar states in order to find the latent plans between start and goal states, extracted by MM. Third, the action proposal module that complements the latent plan found by the LSR with the corresponding actions. We present a thorough investigation of our framework on simulated box stacking and rope/box manipulation tasks, and a folding task executed on a real robot.
Martina Lippi, Petra Poklukar, Michael C. Welle, Anastasia Varava, Hang Yin 0001, Alessandro Marino, Danica Kragic
IEEE Trans. Robotics5
2022 Geometric Multimodal Contrastive Representation Learning
abstract
Learning representations of multimodal data that are both informative and robust to missing modalities at test time remains a challenging problem due to the inherent heterogeneity of data obtained from different channels. To address it, we present a novel Geometric Multimodal Contrastive (GMC) representation learning method consisting of two main components: i) a two-level architecture consisting of modality-specific base encoders, allowing to process an arbitrary number of modalities to an intermediate representation of fixed dimensionality, and a shared projection head, mapping the intermediate representations to a latent representation space; ii) a multimodal contrastive loss function that encourages the geometric alignment of the learned representations. We experimentally demonstrate that GMC representations are semantically rich and achieve state-of-the-art performance with missing modality information on three different learning problems including prediction and reinforcement learning tasks.
Petra Poklukar, Miguel Vasco, Hang Yin 0001, Francisco S. Melo, Ana Paiva 0001, Danica Kragic
ICML3
2022 Back to the Manifold: Recovering from Out-of-Distribution States
abstract
Learning from previously collected datasets of expert data offers the promise of acquiring robotic policies without unsafe and costly online explorations. However, a major challenge is a distributional shift between the states in the training dataset and the ones visited by the learned policy at the test time. While prior works mainly studied the distribution shift caused by the policy during the offline training, the problem of recovering from out-of-distribution states at the deployment time is not very well studied yet. We alleviate the distributional shift at the deployment time by introducing a recovery policy that brings the agent back to the training manifold whenever it steps out of the in-distribution states, e.g., due to an external perturbation. The recovery policy relies on an approximation of the training data density and a learned equivariant mapping that maps visual observations into a latent space in which translations correspond to the robot actions. We demonstrate the effectiveness of the proposed method through several manipulation experiments on a real robotic platform. Our results show that the recovery policy enables the agent to complete tasks while the behavioral cloning alone fails because of the distributional shift problem.
Alfredo Reichlin, Giovanni Luca Marchetti, Hang Yin 0001, Ali Ghadirzadeh, Danica Kragic
IROS3
2022 Consensus-based Normalizing-Flow Control: A Case Study in Learning Dual-Arm Coordination
abstract
We develop two consensus-based learning algorithms for multi-robot systems applied on complex tasks involving collision constraints and force interactions, such as the cooperative peg-in-hole placement. The proposed algorithms integrate multi-robot distributed consensus and normalizing-flow-based reinforcement learning. The algorithms guarantee the stability and the consensus of the multi-robot system's generalized variables in a transformed space. This transformed space is obtained via a diffeomorphic transformation parameterized by normalizing-flow models that the algorithms use to train the underlying task, learning hence skillful, dexterous trajectories required for the task accomplishment. We validate the proposed algorithms by parameterizing reinforcement learning policies, demonstrating efficient cooperative learning, and strong generalization of dual-arm assembly skills in a dynamics-engine simulator.
Hang Yin 0001, Christos K. Verginis, Danica Kragic
IROS1
2022 Embedding Koopman Optimal Control in Robot Policy Learning
abstract
Embedding an optimization process has been explored for imposing efficient and flexible policy structures. Existing work often build upon nonlinear optimization with explicitly iteration steps, making policy inference prohibitively expensive for online learning and real-time control. Our approach embeds a linear-quadratic-regulator (LQR) formulation with a Koopman representation, thus exhibiting the tractability from a closed-form solution and richness from a non-convex neural network. We use a few auxiliary objectives and reparameterization to enforce optimality conditions of the policy that can be easily integrated to standard gradient-based learning. Our approach is shown to be effective for learning policies rendering an optimality structure and efficient reinforcement learning, including simulated pendulum control, 2D and 3D walking, and manipulation for both rigid and deformable objects. We also demonstrate real world application in a robot pivoting task.
Hang Yin 0001, Michael C. Welle, Danica Kragic
IROS1
2022 Leveraging hierarchy in multimodal generative models for effective cross-modality inference
Miguel Vasco, Hang Yin 0001, Francisco S. Melo, Ana Paiva 0001
Neural Networks2
2021 Learning Stable Normalizing-Flow Control for Robotic Manipulation
abstract
Reinforcement Learning (RL) of robotic manipulation skills, despite its impressive successes, stands to benefit from incorporating domain knowledge from control theory. One of the most important properties that is of interest is control stability. Ideally, one would like to achieve stability guarantees while staying within the framework of state-of-the-art deep RL algorithms. Such a solution does not exist in general, especially one that scales to complex manipulation tasks. We contribute towards closing this gap by introducing normalizing-flow control structure, that can be deployed in any latest deep RL algorithms. While stable exploration is not guaranteed, our method is designed to ultimately produce deterministic controllers with provable stability. In addition to demonstrating our method on challenging contact-rich manipulation tasks, we also show that it is possible to achieve considerable exploration efficiency–reduced state space coverage and actuation efforts– without losing learning efficiency.
Shahbaz Abdul Khader, Hang Yin 0001, Pietro Falco, Danica Kragic
ICRA2
2021 Graph-based Task-specific Prediction Models for Interactions between Deformable and Rigid Objects
abstract
Capturing scene dynamics and predicting the future scene state is challenging but essential for robotic manipulation tasks, especially when the scene contains both rigid and deformable objects. In this work, we contribute a simulation environment and generate a novel dataset for task-specific manipulation, involving interactions between rigid objects and a deformable bag. The dataset incorporates a rich variety of scenarios including different object sizes, object numbers and manipulation actions. We approach dynamics learning by proposing an object-centric graph representation and two modules which are Active Prediction Module (APM) and Position Prediction Module (PPM) based on graph neural networks with an encode-process-decode architecture. At the inference stage, we build a two-stage model based on the learned modules for single time step prediction. We combine modules with different prediction horizons into a mixed-horizon model which addresses long-term prediction. In an ablation study, we show the benefits of the two-stage model for single time step prediction and the effectiveness of the mixed-horizon model for long-term prediction tasks. Supplementary material is available at https://github.com/wengzehang/deformable_rigid_interaction_prediction
Zehang Weng, Fabian Paus, Anastasiia Varava, Hang Yin 0001, Tamim Asfour, Danica Kragic
IROS4
2021 Graph-based Normalizing Flow for Human Motion Generation and Reconstruction
abstract
Data-driven approaches for modeling human skeletal motion have found various applications in interactive media and social robotics. Challenges remain in these fields for generating high-fidelity samples and robustly reconstructing motion from imperfect input data, due to e.g. missed marker detection. In this paper, we propose a probabilistic generative model to synthesize and reconstruct long horizon motion sequences conditioned on past information and control signals, such as the path along which an individual is moving. Our method adapts the existing work MoGlow by introducing a new graph-based model. The model leverages the spatial-temporal graph convolutional network (ST-GCN) to effectively capture the spatial structure and temporal correlation of skeletal motion data at multiple scales. We evaluate the models on a mixture of motion capture datasets of human locomotion with foot-step and bone-length analysis. The results demonstrate the advantages of our model in reconstructing missing markers and achieving comparable results on generating realistic future poses. When the inputs are imperfect, our model shows improvements on robustness of generation.
Hang Yin 0001, Danica Kragic, Mårten Björkman
RO-MAN2
2020 In-Hand Manipulation of Objects with Unknown Shapes
abstract
This work addresses the problem of changing grasp configurations on objects with an unknown shape through in-hand manipulation. Our approach leverages shape priors, learned as deep generative models, to infer novel object shapes from partial visual sensing. The Dexterous Manipulation Graph method is extended to build incrementally and account for object shape uncertainty when planning a sequence of manipulation actions. We show that our approach successfully solves in-hand manipulation tasks with unknown objects, and demonstrate the validity of these solutions with robot experiments.
Silvia Cruciani, Hang Yin 0001, Danica Kragic
ICRA2
2020 Latent Space Roadmap for Visual Action Planning of Deformable and Rigid Object Manipulation
abstract
We present a framework for visual action planning of complex manipulation tasks with high-dimensional state spaces such as manipulation of deformable objects. Planning is performed in a low-dimensional latent state space that embeds images. We define and implement a Latent Space Roadmap (LSR) which is a graph-based structure that globally captures the latent system dynamics. Our framework consists of two main components: a Visual Foresight Module (VFM) that generates a visual plan as a sequence of images, and an Action Proposal Network (APN) that predicts the actions between them. We show the effectiveness of the method on a simulated box stacking task as well as a T-shirt folding task performed with a real robot.
Martina Lippi, Petra Poklukar, Michael C. Welle, Anastasiia Varava, Hang Yin 0001, Alessandro Marino, Danica Kragic
IROS5
2018 Do Children Perceive Whether a Robotic Peer is Learning or Not?
abstract
Social robots are being used to create better educational scenarios, thereby fostering children»s learning. In the work presented here, we describe an autonomous social robot that was designed to enhance children»s handwriting skills. Exploiting the benefits of the learning-by-teaching method, the system provides a scenario in which a child acts as a teacher and corrects the handwriting difficulties of the robotic agent. To explore the children»s perception towards this social robot and the effect on their learning, we have conducted a multi-session study with children that compared two contrasting competencies in the robot: 'learning' vs 'non-learning' and presented as two conditions in the study. The results suggest that the children learned more in the learning condition compared with the non-learning condition and their learning gains seem to be affected by their perception of the robot. The results did not lead to any significant differences in the children»s perception of the robot in the first two weeks of interaction. However, by the end of the 4th week, the results changed. The children in the learning condition gave significantly higher writing ability and overall performance scores to the robot compared with the non-learning condition. In addition, the change in the robot»s learning capabilities did not show to affect their perceived intelligence, likability and friendliness towards it.
Shruti Chandra, Raul Benites Paradeda, Hang Yin 0001, Pierre Dillenbourg, Rui Prada, Ana Paiva 0001
HRI3
2017 Associate Latent Encodings in Learning from Demonstrations
abstract
We contribute a learning from demonstration approach for robots to acquire skills from multi-modal high-dimensional data. Both latent representations and associations of different modalities are proposed to be jointly learned through an adapted variational auto-encoder. The implementation and results are demonstrated in a robotic handwriting scenario, where the visual sensory input and the arm joint writing motion are learned and coupled. We show the latent representations successfully construct a task manifold for the observed sensor modalities. Moreover, the learned associations can be exploited to directly synthesize arm joint handwriting motion from an image input in an end-to-end manner. The advantages of learning associative latent encodings are further highlighted with the examples of inferring upon incomplete input images. A comparison with alternative methods demonstrates the superiority of the present approach in these challenging tasks.
Hang Yin 0001, Francisco S. Melo, Aude Billard, Ana Paiva 0001
AAAI1
2016 Synthesizing Robotic Handwriting Motion by Learning from Human Demonstrations
Hang Yin 0001, Patrícia Alves-Oliveira, Francisco S. Melo, Aude Billard, Ana Paiva 0001
IJCAI1
2014 Learning object-level impedance control for robust grasping and dexterous manipulation
abstract
Object-level impedance control is of great importance for object-centric tasks, such as robust grasping and dexterous manipulation. Despite the recent progress on this topic, how to specify the desired object impedance for a given task remains an open issue. In this paper, we decompose the object's impedance into two complementary components-the impedance for stable grasping and impedance for object manipulation. Then, we present a method to learn the desired object's manipulation impedance (stiffness) using data obtained from human demonstration. The approach is validated in two tasks, for robust grasping of a wine glass and for inserting a bulb, using the 16 degrees of freedom Allegro Hand mounted with the SynTouch tactile sensors.
Miao Li 0002, Hang Yin 0001, Kenji Tahara, Aude Billard
ICRA2