Houjian Yu

dblp:216/3376 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
7since 2021 · last 2025
0000-0001-8869-5078ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Systems, architecture and hardware · 4 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Computer networks · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Routing Manipulation of Deformable Linear Object Using Reinforcement Learning and Diffusion Policy
abstract
Tasks involving deformable linear objects (DLOs) are prevalent in daily life but pose significant challenges due to their infinite degrees of freedom and underactuated nature. Frequent contact between DLOs and surrounding objects with unknown physical parameters, such as friction, further complicates their manipulation. Performing tasks like routing ropes through a hole requires gentle yet robust manipulation, making it particularly challenging. Previous research has not adequately addressed general DLO manipulation tasks that involve intensive contact, especially in environments with rough surfaces. This paper presents a robust and delicate manipulation learning approach for the DLO routing task, leveraging reinforcement learning (RL) and diffusion policy. First, reinforcement learning agents are trained separately for rope insertion and pulling. During training, the agents are encouraged to minimize rope tension throughout task execution in environments with randomized friction to achieve delicate motion. Next, the rollouts from these agents are collected as expert demonstrations to train a diffusion policy. Our approach generates delicate motions to prevent the rope from being damaged or getting stuck on rough surfaces while remaining robust against environmental disturbances. Please refer to our project page: https://lmeee.github.io/DLOPull/
Mingen Li, Houjian Yu, Changhyun Choi
ICRA2
2025 A Parameter-Efficient Tuning Framework for Language-Guided Object Grounding and Robot Grasping
abstract
The language-guided robot grasping task requires a robot agent to integrate multimodal information from both visual and linguistic inputs to predict actions for target-driven grasping. While recent approaches utilizing Multimodal Large Language Models (MLLMs) have shown promising results, their extensive computation and data demands limit the feasibility of local deployment and customization. To address this, we propose a novel CLIP-based [1] multimodal parameter-efficient tuning (PET) framework designed for three language-guided object grounding and grasping tasks: (1) Referring Expression Segmentation (RES), (2) Referring Grasp Synthesis (RGS), and (3) Referring Grasp Affordance (RGA). Our approach introduces two key innovations: a bi-directional vision-language adapter that aligns multimodal inputs for pixel-level language understanding and a depth fusion branch that incorporates geometric cues to facilitate robot grasping predictions. Experiment results demonstrate superior performance in the RES object grounding task compared with existing CLIP-based full-model tuning or PET approaches. In the RGS and RGA tasks, our model not only effectively interprets object attributes based on simple language descriptions but also shows strong potential for comprehending complex spatial reasoning scenarios, such as multiple identical objects present in the workspace. Project page: https://z.umn.edu/etog-etrg
Houjian Yu, Mingen Li, Alireza Rezazadeh, Yang Yang 0083, Changhyun Choi
ICRA1
2025 InvSlotGNN: Unsupervised Discovery of Viewpoint Invariant Multiobject Representations and Visual Dynamics
abstract
Learning multiobject dynamics purely from visual data is challenging due to the need for robust object representations that can be learned through robot interactions. In previous work (Rezazadeh et al., 2023), we introduced two novel architectures: SlotTransport for discovering object-centric representations from singleview RGB images, referred to as slots, and SlotGNN for predicting scene dynamics from singleview RGB images and robot interactions using the discovered slots. This article introduces InvSlotGNN, a novel framework for learning multiview slot discovery and dynamics that are invariant to the camera viewpoint. First, we demonstrate that SlotTransport can be trained on multiview data such that a single model discovers temporally aligned, object-centric representations from a wide range of different camera angles. These slots bind to objects from various viewpoints, even under occlusion or absence. Next, we introduce InvSlotGNN, an extension of SlotGNN, that learns multiobject dynamics invariant to the camera angle and predicts the future state from observations taken by uncalibrated cameras. InvSlotGNN learns a graph representation of the scene using the slots from SlotTransport and performs relational and spatial reasoning to predict the future state of the scene for arbitrary viewpoints, conditioned on robot actions. We demonstrate the effectiveness of SlotTransport in learning multiview object-centric features that accurately encode visual and positional information. Furthermore, we highlight the accuracy of InvSlotGNN in downstream robotic tasks, including long-horizon prediction and multiobject rearrangement. Finally, with minimal real data, our framework robustly predicts slots and their dynamics in real-world multiview scenarios.
Alireza Rezazadeh, Houjian Yu, Karthik Desingh, Changhyun Choi
IEEE Trans. Robotics2
2024 Attribute-Based Robotic Grasping With Data-Efficient Adaptation
abstract
Robotic grasping is one of the most fundamental robotic manipulation tasks and has been the subject of extensive research. However, swiftly teaching a robot to grasp a novel target object in clutter remains challenging. This paper attempts to address the challenge by leveraging object attributes that facilitate recognition, grasping, and rapid adaptation to new domains. In this work, we present an end-to-end encoder-decoder network to learn attribute-based robotic grasping with data-efficient adaptation capability. We first pre-train the end-to-end model with a variety of basic objects to learn generic attribute representation for recognition and grasping. Our approach fuses the embeddings of a workspace image and a query text using a gated-attention mechanism and learns to predict instance grasping affordances. To train the joint embedding space of visual and textual attributes, the robot utilizes object persistence before and after grasping. Our model is self-supervised in a simulation that only uses basic objects of various colors and shapes but generalizes to novel objects in new environments. To further facilitate generalization, we propose two adaptation methods, adversarial adaption and one-grasp adaptation. Adversarial adaptation regulates the image encoder using augmented data of unlabeled images, whereas one-grasp adaptation updates the overall end-to-end model using augmented data from one grasp trial. Both adaptation methods are data-efficient and considerably improve instance grasping performance. Experimental results in both simulation and the real world demonstrate that our approach achieves over 81% instance grasping success rate on unknown objects, which outperforms several baselines by large margins. Supplementary material is available athttps://z.umn.edu/attr-grasp.
Yang Yang 0083, Houjian Yu, Xibai Lou, Yuanhao Liu 0003, Changhyun Choi
IEEE Trans. Robotics2
2023 Adversarial Object Rearrangement in Constrained Environments with Heterogeneous Graph Neural Networks
abstract
Adversarial object rearrangement in the real world (e.g., previously unseen or oversized items in kitchens and stores) could benefit from understanding task scenes, which inherently entail heterogeneous components such as current objects, goal objects, and environmental constraints. The semantic relationships among these components are distinct from each other and crucial for multi-skilled robots to perform efficiently in everyday scenarios. We propose a hierarchical robotic manipulation system that learns the underlying relationships and maximizes the collaborative power of its diverse skills (e.g., PICK-PLACE, PUSH) for rearranging adversarial objects in constrained environments. The high-level coordinator employs a heterogeneous graph neural network (HetGNN), which reasons about the current objects, goal objects, and environmental constraints; the low-level 3D Convolutional Neural Network-based actors execute the action primitives. Our approach is trained entirely in simulation, and achieved an average success rate of 87.88% and a planning cost of 12.82 in real-world experiments, surpassing all baseline methods. Supplementary material is available at https://sites.google.com/umn.edu/versatile-rearrangement.
Xibai Lou, Houjian Yu, Ross Worobel, Yang Yang 0083, Changhyun Choi
IROS2
2023 IOSG: Image-Driven Object Searching and Grasping
abstract
When robots retrieve specific objects from cluttered scenes, such as home and warehouse environments, the target objects are often partially occluded or completely hidden. Robots are thus required to search, identify a target object, and successfully grasp it. Preceding works have relied on pre-trained object recognition or segmentation models to find the target object. However, such methods require laborious manual annotations to train the models and even fail to find novel target objects. In this paper, we propose an Image-driven Object Searching and Grasping (IOSG) approach where a robot is provided with the reference image of a novel target object and tasked to find and retrieve it. We design a Target Similarity Network that generates a probability map to infer the location of the novel target. IOSG learns a hierarchical policy; the high-level policy predicts the subtask type, whereas the low-level policies, explorer and coordinator, generate effective push and grasp actions. The explorer is responsible for searching the target object when it is hidden or occluded by other objects. Once the target object is found, the coordinator conducts target-oriented pushing and grasping to retrieve the target from the clutter. The proposed pipeline is trained with full self-supervision in simulation and applied to a real environment. Our model achieves a 96.0% and 94.5% task success rate on coordination and exploration tasks in simulation respectively, and 85.0% success rate on a real robot for the search-and-grasp task. Please refer to our project page for more information: https://z.umn.edu/iosg.
Houjian Yu, Xibai Lou, Yang Yang 0083, Changhyun Choi
IROS1
2022 Self-supervised Interactive Object Segmentation Through a Singulation-and-Grasping Approach
Houjian Yu, Changhyun Choi
ECCV (39)1
2018 Trajectory-Based Reliable Content Distribution in D2D-Based Cooperative Vehicular Networks: A Coalition Formation Approach
abstract
In this paper, we investigate how to achieve reliable content distribution in device-to-device (D2D) based cooperative vehicular networks by combining big data based vehicle trajectory prediction with coalition formation game based resource allocation. Firstly, vehicle trajectory is predicted based on global positioning system (GPS) and geographic information system (GIS) data, which is critical for finding reliable and longlasting vehicle connections. Then, the determination of content distribution groups with different lifetimes is formulated as a coalition formation game. We model the utility function based on the minimization of average network delay to guarantee the end-to-end quality of service (QoS), which is transferable to the individual payoff of each coalition member according to its contribution. The merge and split process is implemented iteratively based on preference relations, and the final partition is proved to converge to a Nash- stable equilibrium. Finally, we evaluate the proposed algorithm based on real-world map and realistic vehicular traffic.
Zhenyu Zhou 0001, Houjian Yu, Chen Xu 0002, Shahid Mumtaz, Jonathan Rodriguez 0001, Muhammad Tariq 0001
ICC3
2018 Dependable Content Distribution in D2D-Based Cooperative Vehicular Networks: A Big Data-Integrated Coalition Game Approach
abstract
Driven by the evolutionary development of automobile industry and cellular technologies, dependable vehicular connectivity has become essential to realize future intelligent transportation systems (ITS). In this paper, we investigate how to achieve dependable content distribution in device-to-device (D2D)-based cooperative vehicular networks by combining big data-based vehicle trajectory prediction with coalition formation game-based resource allocation. First, vehicle trajectory is predicted based on global positioning system and geographic information system data, which is critical for finding reliable and long-lasting vehicle connections. Then, the determination of content distribution groups with different lifetimes is formulated as a coalition formation game. We model the utility function based on the minimization of average network delay, which is transferable to the individual payoff of each coalition member according to its contribution. The merge and split process is implemented iteratively based on preference relations, and the final partition is proved to converge to a Nash-stable equilibrium. Finally, we evaluate the proposed algorithm based on real-world map and realistic vehicular traffic. Numerical results demonstrate that the proposed algorithm can achieve superior performance in terms of average network delay and content distribution efficiency compared with the other heuristic schemes.
Zhenyu Zhou 0001, Houjian Yu, Chen Xu 0002, Yan Zhang 0002, Shahid Mumtaz, Jonathan Rodriguez 0001
IEEE Trans. Intell. Transp. Syst.2