Dubing Chen

dblp:319/3113 · DBLP profile ↗
← Back
10ranked-venue papers
6as first author
10since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 6 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 6 first-author · 7 since 2021
YearPublicationVenuePosition
2026 Multimodal Large Language Models for Multi-Subject In-Context Image Generation
abstract
Recent advances in text-to-image (T2I) generation have enabled visually coherent image synthesis from descriptions, but generating images containing multiple given subjects remains challenging.As the number of reference identities increases, existing methods often suffer from subject missing and semantic drift.To address this problem, we propose MU-SIC, the first MLLM specifically designed for MUlti-Subject In-Context image generation.To overcome the data scarcity, we introduce an automatic and scalable data generation pipeline that eliminates the need for manual annotation.Furthermore, we enhance the model's understanding of multi-subject semantic relationships through a vision chain-of-thought (CoT) mechanism, guiding step-by-step reasoning from subject images to semantics and generation.To mitigate identity entanglement and manage visual complexity, we develop a novel semantics-driven spatial layout planning method and demonstrate its test-time scalability.By incorporating complex subject images during training, we improve the model's capacity for chained reasoning.In addition, we curate MSIC, a new benchmark tailored for multi-subject in-context generation.Experimental results demonstrate that MUSIC significantly surpasses other methods in both multiand single-subject scenarios.
Yucheng Zhou 0001, Dubing Chen, Jianbing Shen
ACL (1)2
2025 Rethinking Temporal Fusion with a Unified Gradient Descent View for 3D Semantic Occupancy Prediction
abstract
We present GDFusion, a temporal fusion method for vision-based 3D semantic occupancy prediction (VisionOcc). GDFusion opens up the underexplored aspects of temporal fusion within the VisionOcc framework, focusing on both temporal cues and fusion strategies. It systematically examines the entire VisionOcc pipeline, identifying three fundamental yet previously overlooked temporal cues: scene-level consistency, motion calibration, and geometric complementation. These cues capture diverse facets of temporal evolution and make distinct contributions across various modules in the VisionOcc framework. To effectively fuse temporal signals across heterogeneous representations, we propose a novel fusion strategy by reinterpreting the formulation of vanilla RNNs. This reinterpretation leverages gradient descent on features to unify the integration of diverse temporal information, seamlessly embedding the proposed temporal cues into the network. Extensive experiments on nuScenes demonstrate that GDFusion significantly outperforms established baselines, achieving 2.2%–4.7% mIoU improvement and reducing memory consumption by 30%–72%. Codes are available at https: //github.com/cdb342/GDFusion.
Dubing Chen, Xingping Dong, Xianfei Li, Wenlong Liao, Jianbing Shen
CVPR1
2025 ALOcc: Adaptive Lifting-Based 3D Semantic Occupancy and Cost Volume-Based Flow Predictions
Dubing Chen, Wencheng Han, Xinjing Cheng, Junbo Yin, Chenzhong Xu, Fahad Shahbaz Khan, Jianbing Shen
ICCV1
2025 Semantic Causality-Aware Vision-Based 3D Occupancy Prediction
abstract
Vision-based 3D semantic occupancy prediction is a critical task in 3D vision that integrates volumetric 3D reconstruction with semantic understanding. Existing methods, however, often rely on modular pipelines. These modules are typically optimized independently or use pre-configured inputs, leading to cascading errors. In this paper, we address this limitation by designing a novel causal loss that enables holistic, end-to-end supervision of the modular 2D-to-3D transformation pipeline. Grounded in the principle of 2D-to-3D semantic causality, this loss regulates the gradient flow from 3D voxel representations back to the 2D features. Consequently, it renders the entire pipeline differentiable, unifying the learning process and making previously non-trainable components fully learnable. Building on this principle, we propose the Semantic Causality-Aware 2D-to-3D Transformation, which comprises three components guided by our causal loss: Channel-Grouped Lifting for adaptive semantic mapping, Learnable Camera Offsets for enhanced robustness against camera perturbations, and Normalized Convolution for effective feature propagation. Extensive experiments demonstrate that our method achieves state-of-the-art performance on the Occ3D benchmark, demonstrating significant robustness to camera perturbations and improved 2D-to-3D semantic consistency.
Dubing Chen, Yucheng Zhou 0001, Xianfei Li, Wenlong Liao, Jianbing Shen
ICCV1
2024 Evolutionary Generalized Zero-Shot Learning
Dubing Chen, Chenyi Jiang, Haofeng Zhang 0001
IJCAI1
2024 Estimation of Near-Instance-Level Attribute Bottleneck for Zero-Shot Learning
Chenyi Jiang, Yuming Shen, Dubing Chen, Haofeng Zhang 0001, Ling Shao 0001, Philip Torr 0001
Int. J. Comput. Vis.3
2023 Deconstructed Generation-Based Zero-Shot Model
abstract
Recent research on Generalized Zero-Shot Learning (GZSL) has focused primarily on generation-based methods. However, current literature has overlooked the fundamental principles of these methods and has made limited progress in a complex manner. In this paper, we aim to deconstruct the generator-classifier framework and provide guidance for its improvement and extension. We begin by breaking down the generator-learned unseen class distribution into class-level and instance-level distributions. Through our analysis of the role of these two types of distributions in solving the GZSL problem, we generalize the focus of the generation-based approach, emphasizing the importance of (i) attribute generalization in generator learning and (ii) independent classifier learning with partially biased data. We present a simple method based on this analysis that outperforms SotAs on four public GZSL datasets, demonstrating the validity of our deconstruction. Furthermore, our proposed method remains effective even without a generative model, representing a step towards simplifying the generator-classifier structure. Our code is available at https://github.com/cdb342/DGZ.
Dubing Chen, Yuming Shen, Haofeng Zhang 0001, Philip Torr 0001
AAAI1
2023 Combining Pixel-Level and Structure-Level Adaptation for Semantic Segmentation
Xiwen Bi, Dubing Chen, Haofeng Zhang 0001
Neural Process. Lett.2
2022 Weighted Contrastive Hashing
Jiaguo Yu, Huming Qiu, Dubing Chen, Haofeng Zhang 0001
ACCV (5)3
2022 Zero-Shot Logit Adjustment
abstract
Semantic-descriptor-based Generalized Zero-Shot Learning (GZSL) poses challenges in recognizing novel classes in the test phase. The development of generative models enables current GZSL techniques to probe further into the semantic-visual link, culminating in a two-stage form that includes a generator and a classifier. However, existing generation-based methods focus on enhancing the generator's effect while neglecting the improvement of the classifier. In this paper, we first analyze of two properties of the generated pseudo unseen samples: bias and homogeneity. Then, we perform variational Bayesian inference to back-derive the evaluation metrics, which reflects the balance of the seen and unseen classes. As a consequence of our derivation, the aforementioned two properties are incorporated into the classifier training as seen-unseen priors via logit adjustment. The Zero-Shot Logit Adjustment further puts semantic-based classifiers into effect in generation-based GZSL. Our experiments demonstrate that the proposed technique achieves state-of-the-art when combined with the basic generator, and it can improve various generative Zero-Shot Learning frameworks. Our codes are available on https://github.com/cdb342/IJCAI-2022-ZLA.
Dubing Chen, Yuming Shen, Haofeng Zhang 0001, Philip Torr 0001
IJCAI1