Yubing Bai

dblp:261/4426 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
6since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 4 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Robot navigation and mapping · 71% Graph learning · 21% Trustworthy machine learning · 3%

Topics — the 8 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Robotics › Robot navigation and mapping
object goal navigation
2.642025
HOZ++: Versatile Hierarchical Object-to-Zone Graph for Object Navigation · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Layout-based Causal Inference for Object Navigation · CVPR 2023
Generative Meta-Adversarial Network for Unseen Object Navigation · ECCV (39) 2022
Robotics › Robot navigation and mapping › robot mapping
cognitive map
0.912025
HOZ++: Versatile Hierarchical Object-to-Zone Graph for Object Navigation · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Machine learning › Graph learning › graph neural network
hierarchical graph
0.912025
HOZ++: Versatile Hierarchical Object-to-Zone Graph for Object Navigation · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Robotics › Robot navigation and mapping
embodied navigation
0.612022
Generative Meta-Adversarial Network for Unseen Object Navigation · ECCV (39) 2022
Machine learning › Graph learning
graph neural network
0.512021
ION: Instance-level Object Navigation · ACM Multimedia 2021
Robotics › Robot navigation and mapping › object goal navigation
visual object navigation
0.512021
ION: Instance-level Object Navigation · ACM Multimedia 2021
Machine learning › Trustworthy machine learning › robustness › adversarial robustness
adversarial training
0.212022
Generative Meta-Adversarial Network for Unseen Object Navigation · ECCV (39) 2022
Machine learning › Reinforcement learning › deep reinforcement learning
deep reinforcement learning for navigation
0.112021
Hierarchical Object-to-Zone Graph for Object Navigation · ICCV 2021

Methods — techniques the papers use, named apart from their topics

reinforcement learning · 0.9modular navigation · 0.9total direct effect · 0.7causal inference · 0.7KL divergence · 0.7meta-learning · 0.6adversarial learning · 0.6online learning · 0.5graph-based planning · 0.5deep reinforcement learning · 0.5
YearPublicationVenuePosition
2025 HOZ++: Versatile Hierarchical Object-to-Zone Graph for Object Navigation
abstract
The goal of object navigation task is to reach the expected objects using visual information in unseen environments. Previous works typically implement deep models as agents that are trained to predict actions based on visual observations. Despite extensive training, agents often fail to make wise decisions when navigating in unseen environments toward invisible targets. In contrast, humans demonstrate a remarkable talent to navigate toward targets even in unseen environments. This superior capability is attributed to the cognitive map in the hippocampus, which enables humans to recall past experiences in similar situations and anticipate future occurrences during navigation. It is also dynamically updated with new observations from unseen environments. The cognitive map equips humans with a wealth of prior knowledge, significantly enhancing their navigation capabilities. Inspired by human navigation mechanisms, we propose the Hierarchical Object-to-Zone (HOZ++) graph, which encapsulates the regularities among objects, zones, and scenes. The HOZ++ graph helps the agent to identify the current zone and the target zone, and computes an optimal path between them, then selects the next zone along the path as the guidance for the agent. Moreover, the HOZ++ graph continuously updates based on real-time observations in new environments, thereby enhancing its adaptability to new environments. Our HOZ++ graph is versatile and can be integrated into existing methods, including end-to-end RL and modular methods. Our method is evaluated across four simulators, including AI2-THOR, RoboTHOR, Gibson, and Matterport 3D. Additionally, we build a realistic environment to evaluate our method in the real world. Experimental results demonstrate the effectiveness and efficiency of our proposed method.
Sixian Zhang, Xinhang Song, Xinyao Yu 0002, Yubing Bai, Xinlong Guo, Shuqiang Jiang
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 Layout-based Causal Inference for Object Navigation
abstract
Previous works for ObjectNav task attempt to learn the association (e.g. relation graph) between the visual inputs and the goal during training. Such association contains the prior knowledge of navigating in training environments, which is denoted as the experience. The experience performs a positive effect on helping the agent infer the likely location of the goal when the layout gap between the unseen environments of the test and the prior knowledge obtained in training is minor. However, when the layout gap is significant, the experience exerts a negative effect on navigation. Motivated by keeping the positive effect and removing the negative effect of the experience, we propose the layout-based soft Total Direct Effect (L-sTDE) framework based on the causal inference to adjust the prediction of the navigation policy. In particular, we propose to calculate the layout gap which is defined as the KL divergence between the posterior and the prior distribution of the object layout. Then the sTDE is proposed to appropriately control the effect of the experience based on the layout gap. Experimental results on AI2THOR, RoboTHOR, and Habitat demonstrate the effectiveness of our method. The code is available at https://github.com/sx-zhang/Layout-based-sTDE.git.
Sixian Zhang, Xinhang Song, Yubing Bai, Xinyao Yu 0002, Shuqiang Jiang
CVPR4
2023 Long-Short Term Policy for Visual Object Navigation
abstract
The goal of visual object navigation for an agent is to find the target objects accurately. Recent works mainly focus on the feature of embedding, attempting to learn better features with different variants, such as object distribution and graph representations. However, some typical navigation problems in complex environments, such as partially known and obstacle problems, may not be effectively addressed by previous feature embedding methods. In this paper, we propose a framework with a long-short objective policy, where the hidden states are classified according to the navigation objectives at that moment and separately rewarded. Specifically, we consider two objectives: the long-term objective is to go closer to the target, and the short-term objective is for obstacle avoidance and exploration. To alleviate the effect of long-term and short-term alternation, we build a state memory and propose an adjustment gate to update the state memory. Finally, all past hidden states are reweighted and combined for action prediction with an action-boosting gate. Experimental results on RoboTHOR show that the proposed method can significantly outperform the state-of-the-art.
Yubing Bai, Xinhang Song, Sixian Zhang, Shuqiang Jiang
IROS1
2022 Generative Meta-Adversarial Network for Unseen Object Navigation
Sixian Zhang, Xinhang Song, Yubing Bai, Shuqiang Jiang
ECCV (39)4
2021 Hierarchical Object-to-Zone Graph for Object Navigation
abstract
The goal of object navigation is to reach the expected objects according to visual information in the unseen environments. Previous works usually implement deep models to train an agent to predict actions in real-time. However, in the unseen environment, when the target object is not in egocentric view, the agent may not be able to make wise decisions due to the lack of guidance. In this paper, we propose a hierarchical object-to-zone (HOZ) graph to guide the agent in a coarse-to-fine manner, and an online-learning mechanism is also proposed to update HOZ according to the real-time observation in new environments. In particular, the HOZ graph is composed of scene nodes, zone nodes and object nodes. With the pre-learned HOZ graph, the real-time observation and the target goal, the agent can constantly plan an optimal path from zone to zone. In the estimated path, the next potential zone is regarded as sub-goal, which is also fed into the deep reinforcement learning model for action prediction. Our methods are evaluated on the AI2-Thor simulator. In addition to widely used evaluation metrics SR and SPL, we also propose a new evaluation metric of SAE that focuses on the effective action rate. Experimental results demonstrate the effectiveness and efficiency of our proposed method. The code is available at https://github.com/sx-zhang/HOZ.git.
Sixian Zhang, Xinhang Song, Yubing Bai, Yakui Chu, Shuqiang Jiang
ICCV3
2021 ION: Instance-level Object Navigation
abstract
Visual object navigation is a fundamental task in Embodied AI. Previous works focus on the category-wise navigation, in which navigating to any possible instance of target object category is considered a success. Those methods may be effective to find the general objects. However, it may be more practical to navigate to the specific instance in our real life, since our particular requirements are usually satisfied with specific instances rather than all instances of one category. How to navigate to the specific instance has been rarely researched before and is typically challenging to current works. In this paper, we introduce a new task of Instance Object Navigation (ION), where instance-level descriptions of targets are provided and instance-level navigation is required. In particular, multiple types of attributes such as colors, materials and object references are involved in the instance-level descriptions of the targets. In order to allow the agent to maintain the ability of instance navigation, we propose a cascade framework with Instance-Relation Graph (IRG) based navigator and instance grounding module. To specify the different instances of the same object categories, we construct instance-level graph instead of category-level one, where instances are regarded as nodes, encoded with the representation of colors, materials and locations (bounding boxes). During navigation, the detected instances can activate corresponding nodes in IRG, which are updated with graph convolutional neural network (GCNN). The final instance prediction is obtained with the grounding module by selecting the candidates (instances) with maximum probability (a joint probability of category, color and material, obtained by corresponding regressors with softmax). For the task evaluation, we build a benchmark for instance-level object navigation on AI2-Thor simulator, where over 27,735 object instance descriptions and navigation groundtruth are automatically obtained through the interaction with the simulator. The proposed model outperforms the baseline in instance-level metrics, showing that our proposed graph model can guide instance object navigation, as well as leaving promising room for further improvement. The project is available at https://github.com/LWJ312/ION.
Xinhang Song, Yubing Bai, Sixian Zhang, Shuqiang Jiang
ACM Multimedia3