Bo Zhao 0037

dblp:94/4810-37 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
3D vision · 54% Vision and language · 12% Deep learning architectures and training · 10%

Topics — the 10 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › 3D vision
depth perception
0.912025
SpatialBot: Precise Spatial Understanding with Vision Language Models · ICRA 2025
Computer vision › 3D vision
spatial understanding
0.912025
SpatialBot: Precise Spatial Understanding with Vision Language Models · ICRA 2025
Computer vision › Vision and language
vision-language model
0.912025
SpatialBot: Precise Spatial Understanding with Vision Language Models · ICRA 2025
Computer vision › 3D vision › object pose estimation
6d object pose estimation
0.812024
Omni6DPose: A Benchmark and Model for Universal 6D Object Pose Estimation and Tracking · ECCV (75) 2024
Machine learning › Transfer learning and domain adaptation
distribution matching
0.812024
Real-Fake: Effective Training Data Synthesis Through Distribution Matching · ICLR 2024
Computer vision › 3D vision › object pose estimation
object pose tracking
0.812024
Omni6DPose: A Benchmark and Model for Universal 6D Object Pose Estimation and Tracking · ECCV (75) 2024
Computer vision › 3D vision › pose estimation
pose estimation benchmark
0.812024
Omni6DPose: A Benchmark and Model for Universal 6D Object Pose Estimation and Tracking · ECCV (75) 2024
Machine learning › Generative modeling
synthetic data generation
0.812024
Real-Fake: Effective Training Data Synthesis Through Distribution Matching · ICLR 2024
Machine learning › Deep learning architectures and training › data-centric deep learning
training data generation
0.812024
Real-Fake: Effective Training Data Synthesis Through Distribution Matching · ICLR 2024
Computer vision › Image recognition and object detection
image classification
0.212024
Real-Fake: Effective Training Data Synthesis Through Distribution Matching · ICLR 2024

Methods — techniques the papers use, named apart from their topics

vision-language model · 0.9synthetic data augmentation · 0.8object tracking · 0.8distribution matching · 0.86d pose estimation · 0.8
YearPublicationVenuePosition
2025 SpatialBot: Precise Spatial Understanding with Vision Language Models
abstract
Vision Language Models (VLMs) have achieved impressive performance in 2D image understanding; however, they still struggle with spatial understanding, which is fundamental to embodied AI. In this paper, we propose SpatialBot, a model designed to enhance spatial understanding by utilizing both RGB and depth images. To train VLMs for depth perception, we introduce the SpatialQA and SpatialQA$\boldsymbol{E}$datasets, which include multi-level depth-related questions spanning various scenarios and embodiment tasks. SpatialBench is also developed to comprehensively evaluate VLMs' spatial understanding capabilities across different levels. Extensive experiments on our spatial-understanding benchmark, general VLM benchmarks, and embodied AI tasks demonstrate the remarkable improvements offered by SpatialBot. The model, code, and datasets are available at https://github.com/BAAI-DCAI/SpatialBot.
Wenxiao Cai, Iaroslav Ponomarenko, Jianhao Yuan, Xiaoqi Li 0020, Wankou Yang, Hao Dong 0003, Bo Zhao 0037
ICRA7
2024 Omni6DPose: A Benchmark and Model for Universal 6D Object Pose Estimation and Tracking
Jiyao Zhang, Weiyao Huang, Mingdong Wu, Bo Zhao 0037, Hao Dong 0003
ECCV (75)7
2024 Real-Fake: Effective Training Data Synthesis Through Distribution Matching
abstract
Synthetic training data has gained prominence in numerous learning tasks and scenarios, offering advantages such as dataset augmentation, generalization evaluation, and privacy preservation. Despite these benefits, the efficiency of synthetic data generated by current methodologies remains inferior when training advanced deep models exclusively, limiting its practical utility. To address this challenge, we analyze the principles underlying training data synthesis for supervised learning and elucidate a principled theoretical framework from the distribution-matching perspective that explicates the mechanisms governing synthesis efficacy. Through extensive experiments, we demonstrate the effectiveness of our synthetic data across diverse image classification tasks, both as a replacement for and augmentation to real datasets, while also benefits such as out-of-distribution generalization, privacy preservation, and scalability. Specifically, we achieve 70.9% top1 classification accuracy on ImageNet1K when training solely with synthetic data equivalent to 1 × the original real data size, which increases to 76.0% when scaling up to 10 × synthetic data.
Jianhao Yuan, Jie Zhang 0050, Shuyang Sun, Philip Torr 0001, Bo Zhao 0037
ICLR5