Yingbo Tang

dblp:331/0128 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
9since 2021 · last 2026
0009-0001-1657-256XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 NavA³: Understanding Any Instruction, Navigating Anywhere, Finding Anything
abstract
Lingfeng Zhang, Xiaoshuai Hao, Yingbo Tang, Haoxiang Fu, Xinyu Zheng, Pengwei Wang, Zhongyuan Wang, Wenbo Ding, Shanghang Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Xiaoshuai Hao, Yingbo Tang, Haoxiang Fu, Pengwei Wang 0005, Zhongyuan Wang 0006, Wenbo Ding 0001, Shanghang Zhang
ACL (1)3
2026 Diff-COPE: Diffusion-Based Category-Level 6D Object Pose Estimation
abstract
Category-level 6D object pose estimation has gained increasing attention in applications of robotic manipulation, augmented reality, and scene understanding, due to its ability to generalize to unseen instances within the same category. However, existing methods struggle with handling the intra-class shape variations as they either adopt mean shape as priors, or build the associations among different instances without explicit category-shared information. To address this problem, a novel category-level object pose estimation method based on diffusion model is proposed, which utilize the generative ability of diffusion model to refine a sparse categorical representation. In contrast to existing dense correspondence-based methods, our method employs a set of keypoints provided by learnable queries to represent object shape, enabling better categorical representation of different instances by focusing on the representative object components. The keypoints are then refined through a forward diffusion process and a reverse denoising process conditioned on category information. This allows the flexible adaption to various instances, especially for those that deviate from the mean shape within the same category. On this basis, a geometric-semantic feature fusion module is presented to enhance keypoint feature representation. By integrating the geometric information from point cloud with the high-level semantics from RGB image using a two-branch attention mechanism, the keypoint feature is enriched and deeply combined, which facilitates the subsequent pose estimation. Extensive experiments on the REAL275 dataset, the CAMERA25 dataset, and real-world complex scenarios demonstrated the effectiveness of proposed method.
Yingbo Tang, Zhiqiang Cao 0002, Peiyu Guan, Xurong Gong, Junzhi Yu 0001
IEEE Trans. Circuits Syst. Video Technol.1
2026 Embodied Spatial Affordance: Spatial-Aware Affordance Learning for Embodied Navigation and Manipulation
abstract
Embodied navigation and manipulation are fundamental capabilities for embodied agents operating in physical environments. A key challenge in this process is understanding the spatial context and the affordances of the environment, which involves recognizing how objects can be interacted with (object affordance) and identifying suitable locations for movement and object placement (free space affordance). While Vision-Language Models (VLMs) have shown promise in high-level task planning, their ability to translate reasoning into precise executable actions remains limited, particularly in image-based spatial understanding and precise affordance localization-a critical gap in image processing for robotics. To bridge this gap, we propose EspA, a novel image-to-keypoint model that leverages spatial-aware affordance learning to predict actionable affordances directly from 2D image inputs. Built on a hierarchical vision-language architecture, EspA jointly reasons about object affordances and free space affordances, enabling pixel-level localization of both types of interactions. Crucially, EspA translates language instructions into precise 2D affordance keypoints from observed images, which are then projected into 3D actionable coordinates using depth information. To support this unified affordance reasoning, we introduce the Embodied Spatial Affordance (ESA) dataset, which captures both object-centric interactions and free space contexts. By jointly modeling these affordances in a shared representation space, EspA overcomes the limitations of prior works that treat them independently. The dataset's fine-grained annotations enable our model to learn the intricate relationship between object functionality and spatial feasibility, significantly enhancing the spatial understanding in embodied tasks. Extensive experimental results demonstrate that EspA outperforms existing state-of-the-art Vision-Language Models (VLMs), both open-source and closed-source, in object and free space affordance prediction. Furthermore, it exhibits superior performance in real-world embodied navigation and manipulation experiments. Our work advances the field of image-based spatial reasoning by providing a scalable solution for translating high-level instructions into low-level actionable affordances. We believe this work paves the way for more robust and versatile embodied agents capable of effectively interacting with complex environments. The dataset, benchmark, and evaluation code will be publicly available to facilitate future research. Project website: https://embodied-spatial-affordance.github.io/.
Xiaoshuai Hao, Yingbo Tang, Long Chen 0015, Wei Zhou 0021, Jungong Han, Wenbo Ding 0001, Xiao-Ping Zhang 0002
IEEE Trans. Image Process.2
2026 Synergistic Prompting for Complementarity and Consistency in Incomplete Multi-View Clustering
abstract
Incomplete multi-view clustering (IMVC) aims to partition unlabeled multi-view data into semantically coherent groups, even when certain views are missing due to sensor failures, data collection constraints, or privacy concerns. Despite advancements in deep IMVC methods, two critical challenges remain unresolved: (i) the lack of explicit mechanisms to model cross-view complementarity and (ii) the absence of principled strategies to ensure global semantic consistency across views. To address these challenges, we propose SP-IMVC, a novel Synergistic Prompting framework that jointly models complementarity and consistency under view incompleteness. Specifically, we introduce two types of learnable prompts: the Cross-View Complementary Prompt (CVCP), which aggregates auxiliary representations from available views to enrich the semantics of the current view and mitigate information loss; and the Latent Anchor Prompt (LAP), which utilizes a global anchor prompt pool to provide adaptive semantic priors that promote globally consistent representations. These prompts are optimized jointly within a unified architecture to achieve synergistic prompting of cross-view complementarity and global semantic consistency. Extensive experiments on six public benchmarks demonstrate that SP-IMVC consistently outperforms 14 state-of-the-art IMVC approaches, particularly in scenarios with high missing-view ratios, validating the effectiveness and robustness of our synergistic prompt-guided clustering framework. The code will be released to facilitate future research.
Xiaoshuai Hao, Yingbo Tang, Peng Hao 0003, Yunfeng Diao, Guangyin Jin, Yu Liu 0023
IEEE Trans. Image Process.3
2025 AffordGrasp: In-Context Affordance Reasoning for Open-Vocabulary Task-Oriented Grasping in Clutter
abstract
Inferring the affordance of an object and grasping it in a task-oriented manner is crucial for robots to successfully complete manipulation tasks. Affordance indicates where and how to grasp an object by taking its functionality into account, serving as the foundation for effective task-oriented grasping. However, current task-oriented methods often depend on extensive training data that is confined to specific tasks and objects, making it difficult to generalize to novel objects and complex scenes. In this paper, we introduce AffordGrasp, a novel open-vocabulary grasping framework that leverages the reasoning capabilities of vision-language models (VLMs) for in-context affordance reasoning. Unlike existing methods that rely on explicit task and object specifications, our approach infers tasks directly from implicit user instructions, enabling more intuitive and seamless human-robot interaction in everyday scenarios. Building on the reasoning outcomes, our framework identifies task-relevant objects and grounds their part-level affordances using a visual grounding module. This allows us to generate task-oriented grasp poses precisely within the affordance regions of the object, ensuring both functional and context-aware robotic manipulation. Extensive experiments demonstrate that AffordGrasp achieves state-of-the-art performance in both simulation and real-world scenarios, highlighting the effectiveness of our method. We believe our approach advances robotic manipulation techniques and contributes to the broader field of embodied AI. Project website: https://eqcy.github.io/affordgrasp/.
Yingbo Tang, Shuaike Zhang, Xiaoshuai Hao, Pengwei Wang 0004, Jianlong Wu, Zhongyuan Wang 0006, Shanghang Zhang
IROS1
2025 RoboAfford: A Dataset and Benchmark for Enhancing Object and Spatial Affordance Learning in Robot Manipulation
abstract
Robot manipulation is a fundamental capability of embodied intelligence, enabling effective robot interactions with the physical world. In robotic manipulation tasks, predicting precise grasping positions and object placement is essential. Achieving this requires object recognition to localize target object, predicting object affordances for interaction and spatial affordances for optimal arrangement. While Vision-Language Models (VLMs) provide insights for high-level task planning and scene understanding, they often struggle to predict precise action positions, such as functional grasp points and spatial placements. This limitation stems from the lack of annotations for object and spatial affordance data in their training datasets. To address this gap, we introduce RoboAfford , a novel large-scale dataset designed to enhance object and spatial affordance learning in robot manipulation. Our dataset comprises 819,987 images paired with 1.9 million question answering (QA) annotations, covering three critical tasks: object affordance recognition to identify objects based on attributes and spatial relationships, object affordance prediction to pinpoint functional grasping parts, and spatial affordance localization to identify free space for placement. Complementing this dataset, we propose RoboAfford-Eval , a comprehensive benchmark for assessing affordance-aware prediction in real-world scenarios, featuring 338 meticulously annotated samples across the same three tasks. Extensive experimental results reveal the deficiencies of existing VLMs in affordance learning, while fine-tuning on the RoboAfford dataset significantly enhances their affordance prediction in robot manipulation, validating the dataset's effectiveness. The dataset, benchmark and evaluation code will be made publicly available to facilitate future research. Project website: https://roboafford-dataset.github.io/.
Yingbo Tang, Yinuo Zhao, Xiaoshuai Hao
ACM Multimedia1
2025 Video-CoT: A Comprehensive Dataset for Spatiotemporal Understanding of Videos Based on Chain-of-Thought
abstract
Video content comprehension is essential for various applications, ranging from video analysis to interactive systems. Despite advancements in large-scale vision-language models (VLMs), these models often struggle to capture the nuanced, spatiotemporal details essential for thorough video analysis. To address this gap, we introduce Video-CoT, a groundbreaking dataset designed to enhance spatiotemporal understanding using Chain-of-Thought(CoT) methodologies. Video-CoT contains 192,000 fine-grained spatiotemporal question-answer pairs and 23,000 high-quality CoT-annotated samples, providing a solid foundation for evaluating spatiotemporal understanding in video comprehension. Addition- ally, we provide a comprehensive benchmark for assessing these tasks, with each task featuring 750 images and tailored evaluation metrics. Our extensive experiments reveal that current VLMs face significant challenges in achieving satisfactory performance, high- lighting the difficulties of effective spatiotemporal understanding. Overall, the Video-CoT dataset and benchmark open new avenues for research in multimedia understanding and support future innovations in intelligent systems requiring advanced video analysis capabilities. By making these resources publicly available, we aim to encourage further exploration in this critical area. Project website: https://video-cot.github.io/ .
Xiaoshuai Hao, Yingbo Tang, Pengwei Wang 0005, Zhongyuan Wang 0006, Hongxuan Ma, Shanghang Zhang
ACM Multimedia3
2024 Semi-Supervised Few-Shot Object Detection via Adaptive Pseudo Labeling
abstract
Few-shot object detection (FSOD) aims to detect novel objects with limited annotated examples. Mainstream methods suffer from the data scarcity of novel classes with insufficient intra-class variations, which makes the trained model biased to base classes. Actually, there are massive unlabeled novel instances in the base dataset and their adequate utilization will enhance the discriminability of model to novel classes. This paper proposes a semi-supervised few-shot object detection method, which utilizes a teacher model and a pre-trained few-shot object detector to guide the learning of a student model through adaptive pseudo labeling. In particular, a class-adaptive threshold filtering (CATF) strategy is designed to deal with the class-imbalance problem of pseudo labels. And for each novel class, the threshold to select valuable pseudo labels is determined by quantile statistics of the confidence score distribution of pseudo labels. Furthermore, the pre-trained detector and the teacher model are associated with the preliminary CATF and in-depth CATF, respectively, and then the pseudo labels from the two-stream CATF are fused to provide supervisions. In this way, the knowledge of these two models is exploited, which improves the quality of pseudo labels. Under these supervisions, the student model is trained and the teacher model is correspondingly updated through parameters sharing, thus forming a positive feedback to improve the performance of both models. Besides, an attention module is integrated to the teacher and student models to enhance the feature representation of novel instances. The validations on PASCAL VOC and MS COCO show the effectiveness of the proposed method.
Yingbo Tang, Zhiqiang Cao 0002, Yuequan Yang, Jierui Liu, Junzhi Yu 0001
IEEE Trans. Circuits Syst. Video Technol.1
2022 Category-Level 6D Object Pose Estimation With Structure Encoder and Reasoning Attention
abstract
Category-level 6D object pose estimation has gained popularity and it is still challenging due to the diversity of different instances within the same category. In this paper, a novel category-level 6D object pose estimation framework with structure encoder and reasoning attention is proposed. A structure autoencoder is introduced to mine the shared structure features in the color images within the same category, via a distinct learning strategy that recovers the image of another instance but with the most similar pose to the input. On this basis, a reasoning attention decoder and full connected layers are stacked to form a rotation prediction network, where the structure features and 3D shape features are integrated and projected to a semantic space. The semantic space includes observed patterns and learnable patterns, which are better learned by adding a shortcut connection branch parallel to reasoning attention decoder with gradient decouple. Further reasoning based on these patterns endows the decoder with powerful feature representation. Without 3D object models, the proposed method models the attributes of category implicitly in the semantic space and better performance of 6D object pose estimation is guaranteed by reasoning on this space. The effectiveness of the proposed method is verified by the results on public datasets and actual experiments.
Jierui Liu, Zhiqiang Cao 0002, Yingbo Tang, Xilong Liu, Min Tan 0001
IEEE Trans. Circuits Syst. Video Technol.3