Bu Jin

dblp:326/8464 · DBLP profile ↗
← Back
7ranked-venue papers
3as first author
7since 2021 · last 2026
0000-0001-7577-2177ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021Systems, architecture and hardware · 5 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 HiPrompt: Tuning-free Higher-Resolution Generation with Hierarchical MLLM Prompts
abstract
Abstract The potential for higher-resolution image generation using pretrained diffusion models is immense. However, these models often struggle with object repetition and structural artifacts especially when scaling to 4K resolution and beyond. Our analysis reveals that causes the problem, a single prompt for the generation of multiple scales provides insufficient efficacy. To address this, we propose HiPrompt, a new tuning-free solution that tackles the above problems by introducing hierarchical prompts. The hierarchical prompts provide both global and local semantic guidance. Specifically, the global prompt captures overall scene semantics from user input, while local guidance comes from patch-wise descriptions generated by MLLMs to refine regional structures and textures. Furthermore, during inverse denoising, noise is decomposed into low- and high-frequency components, each conditioned on different prompt levels, facilitating prompt-guided denoising under hierarchical semantic guidance. It further allows the generation to focus more on local spatial regions and ensures the generated images maintain coherent local and global semantics, structures, and textures with high definition. Extensive experiments demonstrate that HiPrompt outperforms state-of-the-art works in higher-resolution image generation, significantly reducing object repetition and enhancing structural quality. The demo and code can be found on the project website: https://liuxinyv.github.io/HiPrompt/ .
Yingqing He, Lanqing Guo, Bu Jin, Chi-Min Chan, Wei Xue 0002, Wenhan Luo, Yike Guo
Int. J. Comput. Vis.5
2025 AVD2: Accident Video Diffusion for Accident Video Description
abstract
Traffic accidents present complex challenges for autonomous driving, often featuring unpredictable scenarios that hinder accurate system interpretation and responses. Nonetheless, prevailing methodologies fall short in elucidating the causes of accidents and proposing preventive measures due to the paucity of training data specific to accident scenarios. In this work, we introduce AVD2 (Accident Video Diffusion for Accident Video Description), a novel framework that enhances accident scene understanding by generating accident videos that aligned with detailed natural language descriptions and reasoning, resulting in the contributed EMM-AU (Enhanced Multi-Modal Accident Video Understanding) dataset. Empirical results reveal that the integration of the EMM-AU dataset establishes state-of-the-art performance across both automated metrics and human evaluations, markedly advancing the domains of accident analysis and prevention. Project resources are available at https://an-answer-tree.github.io
Cheng Li 0066, Keyuan Zhou, Mingqiao Zhuang, Huan-ang Gao, Bu Jin, Hao Zhao 0002
ICRA7
2025 PosePilot: Steering Camera Pose for Generative World Models with Self-supervised Depth
abstract
Recent advancements in autonomous driving (AD) systems have highlighted the potential of world models in achieving robust and generalizable performance across both ordinary and challenging driving conditions. However, a key challenge remains: precise and flexible camera pose control, which is crucial for accurate viewpoint transformation and realistic simulation of scene dynamics. In this paper, we introduce PosePilot, a lightweight yet powerful framework that significantly enhances camera pose controllability in generative world models. Drawing inspiration from self-supervised depth estimation, PosePilot leverages structure-from-motion principles to establish a tight coupling between camera pose and video generation. Specifically, we incorporate self-supervised depth and pose readouts, allowing the model to infer depth and relative camera motion directly from video sequences. These outputs drive pose-aware frame warping, guided by a photometric warping loss that enforces geometric consistency across synthesized frames. To further refine camera pose estimation, we introduce a reverse warping step and a pose regression loss, improving viewpoint precision and adaptability. Extensive experiments on autonomous driving and general-domain video datasets demonstrate that PosePilot significantly enhances structural understanding and motion reasoning in both diffusion-based and auto-regressive world models. By steering camera pose with self-supervised depth, PosePilot sets a new benchmark for pose controllability, enabling physically consistent, reliable viewpoint synthesis in generative world models.
Bu Jin, Weize Li 0001, Baihan Yang, Zhenxin Zhu, Junpeng Jiang, Huan-ang Gao, Kun Zhan, Hengtong Hu, Xueyang Zhang, Peng Jia 0007, Hao Zhao 0002
IROS1
2024 TOD3Cap: Towards 3D Dense Captioning in Outdoor Scenes
Bu Jin, Yupeng Zheng, Pengfei Li 0007, Weize Li 0001, Yuhang Zheng 0004, Sujie Hu, Zhijie Yan, Kun Zhan, Peng Jia 0007, Xiaoxiao Long, Hao Zhao 0002
ECCV (18)1
2024 MonoOcc: Digging into Monocular Semantic Occupancy Prediction
abstract
Monocular Semantic Occupancy Prediction aims to infer the complete 3D geometry and semantic information of scenes from only 2D images. It has garnered significant attention, particularly due to its potential to enhance the 3D perception of autonomous vehicles. However, existing methods rely on a complex cascaded framework with relatively limited information to restore 3D scenes, including a dependency on supervision solely on the whole network’s output, single-frame input, and the utilization of a small backbone. These challenges, in turn, hinder the optimization of the framework and yield inferior prediction results, particularly concerning smaller and long-tailed objects. To address these issues, we propose MonoOcc. In particular, we (i) improve the monocular occupancy prediction framework by proposing an auxiliary semantic loss as supervision to the shallow layers of the framework and an image-conditioned cross-attention module to refine voxel features with visual clues, and (ii) employ a distillation module that transfers temporal information and richer knowledge from a larger image backbone to the monocular semantic occupancy prediction framework with low cost of hardware. With these advantages, our method yields state-of-the-art performance on the camera-based SemanticKITTI Scene Completion benchmark. Codes and models can be accessed at https://github.com/ucaszyp/MonoOcc.
Yupeng Zheng, Xiang Li 0205, Pengfei Li 0007, Yuhang Zheng 0004, Bu Jin, Chengliang Zhong, Xiaoxiao Long, Hao Zhao 0002
ICRA5
2023 ADAPT: Action-aware Driving Caption Transformer
abstract
End-to-end autonomous driving has great potential in the transportation industry. However, the lack of transparency and interpretability of the automatic decision-making process hinders its industrial adoption in practice. There have been some early attempts to use attention maps or cost volume for better model explainability which is difficult for ordinary passengers to understand. To bridge the gap, we propose an end-to-end transformer-based architecture, ADAPT (Action-aware Driving cAPtion Transformer), which provides user-friendly natural language narrations and reasoning for each decision making step of autonomous vehicular control and action. ADAPT jointly trains both the driving caption task and the vehicular control prediction task, through a shared video representation. Experiments on BDD-X (Berkeley DeepDrive eXplanation) dataset demonstrate state-of-the-art performance of the ADAPT framework on both automatic metrics and human evaluation. To illustrate the feasibility of the proposed framework in real-world applications, we build a novel deployable system that takes raw car videos as input and outputs the action narrations and reasoning in real time. The code, models and data are available at https://github.com/jxbbb/ADAPT.
Bu Jin, Yupeng Zheng, Pengfei Li 0007, Hao Zhao 0002, Yuhang Zheng 0004, Guyue Zhou
ICRA1
2023 STEPS: Joint Self-supervised Nighttime Image Enhancement and Depth Estimation
abstract
Self-supervised depth estimation draws a lot of attention recently as it can promote the 3D sensing capa-bilities of self-driving vehicles. However, it intrinsically relies upon the photometric consistency assumption, which hardly holds during nighttime. Although various supervised night-time image enhancement methods have been proposed, their generalization performance in challenging driving scenarios is not satisfactory. To this end, we propose the first method that jointly learns a nighttime image enhancer and a depth estimator, without using ground truth for either task. Our method tightly entangles two self-supervised tasks using a newly proposed uncertain pixel masking strategy. This strategy originates from the observation that nighttime images not only suffer from underexposed regions but also from overexposed regions. By fitting a bridge-shaped curve to the illumination map distribution, both regions are suppressed and two tasks are bridged naturally. We benchmark the method on two established datasets: nuScenes and RobotCar and demonstrate state-of-the-art performance on both of them. Detailed ablations also reveal the mechanism of our proposal. Last but not least, to mitigate the problem of sparse ground truth of existing datasets, we provide a new photo-realistically enhanced nighttime dataset based upon CARLA. It brings meaningful new challenges to the community. Codes, data, and models are available at https://github.com/ucaszyp/STEPS.
Yupeng Zheng, Chengliang Zhong, Pengfei Li 0007, Huan-ang Gao, Yuhang Zheng 0004, Bu Jin, Ling Wang 0001, Hao Zhao 0002, Guyue Zhou, Dongbin Zhao
ICRA6