EDBT 2026 Demo / reviewers in the wild / expert
Guangcong Zheng
dblp:279/5938
· DBLP profile ↗
14ranked-venue papers
3as first author
13since 2021 · last 2026
0000-0002-9943-1577ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 10 since 2021Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CamI2V-Epipolar: Epipolar-Constrained Block Sparse Attention for Camera-Controlled Image-to-Video Diffusion ModelabstractRecently, camera pose has emerged as a physics-informed condition for video diffusion models. Existing methods that directly adopt 3D full attention for cross-frame feature interaction achieve moderate camera controllability with extensive training, yet camera controllability and geometric consistency remain a challenge in complex scenarios such as large camera rotation or movement. This underscores the value of integrating physical priors into attention mechanisms. Instead of pixel-level attention masking, we propose a key innovation to infuse epipolar geometric priors into the block selection of video block sparse attention. This design resolves memory and speed bottlenecks at high resolutions like 768P, enables compatibility with FlashAttention-3, and establishes a principled framework for embedding physical priors into sparse attention mechanisms. Experiments on static RealEstate10 K and dynamic RealCam-Vid datasets demonstrate that our method outperforms the state-of-the-art RealCam-I2V, effectively enhancing camera controllability while preserving generation quality and generalization. Guangcong Zheng, Xi Li 0001 |
IEEE Signal Process. Lett. | 1 |
| 2025 | Energy-Guided Optimization for Personalized Image Editing with Pretrained Text-to-Image Diffusion ModelsabstractThe rapid advancement of pretrained text-driven diffusion models has significantly enriched applications in image generation and editing. However, as the demand for personalized content editing increases, new challenges emerge especially when dealing with arbitrary objects and complex scenes. Existing methods usually mistakes mask as the object shape prior, which struggle to achieve a seamless integration result. The mostly used inversion noise initialization also hinders the identity consistency towards the target object. To address these challenges, we propose a novel training-free framework that formulates personalized content editing as the optimization of edited images in the latent space, using diffusion models as the energy function guidance conditioned by reference text-image pairs. A coarse-to-fine strategy is proposed that employs text energy guidance at the early stage to achieve a natural transition toward the target class and uses point-to-point feature-level image energy guidance to perform fine-grained appearance alignment with the target object. Additionally, we introduce the latent space content composition to enhance overall identity consistency with the target. Extensive experiments demonstrate that our method excels in object replacement even with a large domain gap, highlighting its potential for high-quality, personalized image editing. Xinghe Fu, Guangcong Zheng, Taiping Yao |
AAAI | 3 |
| 2025 | CustomCrafter: Customized Video Generation with Preserving Motion and Concept Composition AbilitiesabstractCustomized video generation aims to generate high-quality videos guided by text prompts and subject's reference images. However, since it is only trained on static images, the fine-tuning process of subject learning disrupts abilities of video diffusion models (VDMs) to combine concepts and generate motions. To restore these abilities, some methods use additional video similar to the prompt to fine-tune or guide the model. This requires frequent changes of guiding videos and even re-tuning of the model when generating different motions, which is very inconvenient for users. In this paper, we propose CustomCrafter, a novel framework that preserves the model's motion generation and conceptual combination abilities without additional video and fine-tuning to recovery. For preserving conceptual combination ability, we design a plug-and-play module to update few parameters in VDMs, enhancing the model's ability to capture the appearance details and the ability of concept combinations for new subjects. For motion generation, we observed that VDMs tend to restore the motion of video in the early stage of denoising, while focusing on the recovery of subject details in the later stage. Therefore, we propose Dynamic Weighted Video Sampling Strategy. Using the pluggability of our subject learning modules, we reduce the impact of this module on motion generation in the early stage of denoising, preserving the ability to generate motion of VDMs. In the later stage of denoising, we restore this module to repair the appearance details of the specified subject, thereby ensuring the fidelity of the subject's appearance. Experimental results show that our method has a significant improvement compared to previous methods. Yong Zhang 0034, Xintao Wang 0002, Xianpan Zhou, Guangcong Zheng, Zhongang Qi, Ying Shan, Xi Li 0001 |
AAAI | 5 |
| 2025 | RealCam-I2V: Real-World Image-to-Video Generation with Interactive Complex Camera ControlabstractRecent advancements in camera-trajectory-guided image-to-video generation offer higher precision and better support for complex camera control compared to text-based approaches. However, they also introduce significant usability challenges, as users often struggle to provide precise camera parameters when working with arbitrary real-world images without knowledge of their depth nor scene scale. To address these real-world application issues, we propose RealCam-I2V, a novel diffusion-based video generation framework that integrates monocular metric depth estimation to establish 3D scene reconstruction in a preprocessing step. During training, the reconstructed 3D scene enables scaling camera parameters from relative to metric scales, ensuring compatibility and scale consistency across diverse real-world images. In inference, RealCam-I2V offers an intuitive interface where users can precisely draw camera trajectories by dragging within the 3D scene. To further enhance precise camera control and scene consistency, we propose scene-constrained noise shaping, which shapes high-level noise and also allows the framework to maintain dynamic and coherent video generation in lower noise stages. RealCam-I2V achieves significant improvements in controllability and video quality on the RealEstate10K and out-of-domain images. We further enables applications like camera-controlled looping video generation and generative frame interpolation. Project page: https://zgctroy.github.io/RealCam-I2V. Guangcong Zheng, Shuigen Zhan, Yehao Lu, Yining Lin, Chuanyun Deng, Yepan Xiong |
ICCV | 2 |
| 2025 | Dynamic-DINO: Fine-Grained Mixture of Experts Tuning for Real-Time Open-Vocabulary Object DetectionabstractThe Mixture of Experts (MoE) architecture has excelled in Large Vision-Language Models (LVLMs), yet its potential in real-time open-vocabulary object detectors, which also leverage large-scale vision-language datasets but smaller models, remains unexplored. This work investigates this domain, revealing intriguing insights. In the shallow layers, experts tend to cooperate with diverse peers to expand the search space. While in the deeper layers, fixed collaborative structures emerge, where each expert maintains 2-3 fixed partners and distinct expert combinations are specialized in processing specific patterns. Concretely, we propose Dynamic-DINO, which extends Grounding DINO 1.5 Edge from a dense model to a dynamic inference framework via an efficient MoE-Tuning strategy. Additionally, we design a granularity decomposition mechanism to decompose the Feed-Forward Network (FFN) of base model into multiple smaller expert networks, expanding the subnet search space. To prevent performance degradation at the start of fine-tuning, we further propose a pre-trained weight allocation strategy for the experts, coupled with a specific router initialization. During inference, only the input-relevant experts are activated to form a compact subnet. Experiments show that, pretrained with merely 1.56M open-source data, Dynamic-DINO outperforms Grounding DINO 1.5 Edge, pretrained on the private Grounding20M dataset. Yehao Lu, Minghe Weng, Zekang Xiao, Guangcong Zheng, Ping Luo 0002 |
ICCV | 6 |
| 2025 | Relationship-Incremental Scene Graph Generation by a Divide-and-Conquer Pipeline With Feature AdapterabstractAs a challenging computer vision task, Scene Graph Generation (SGG) finds the latent semantic relationships among objects from a given image, which may be limited by the datasets and real-world scenarios. In this paper, we consider a novel incremental learning task called Relationship-Incremental Scene Graph Generation (RISGG) that learns the semantic relationships among objects in an incremental way. Compared with classic Class-Incremental Learning (CIL) problem, RISGG suffers from its special issues: 1) Old class shift - the relationship-labeled object pair may have different labels during different learning sessions; 2) Background shift - the relationship-unlabeled object pair may not be a real unlabeled one. In this work, we address the above issues from the following aspects. First, we present a Divide-and-Conquer (DaC) pipeline to deal with the old class shift via decoupling the recognition of relationship classes and recognizing relationships individually. In this way, label confusion and interaction among different relationships are eliminated during training. Second, we propose a Feature Adapter (FA) to bridge the feature space gap between the current session and the previous one and use our extra supervision to mine old relationship information in the current session. Our proposed network combined DaC and FA, abbreviated DaCFA-Net, for RISGG. Experimental results on the benchmark dataset demonstrate the significant performance gain of DaCFA-Net in RISGG. It gains about 20% improvement against the SGG baselines on the popular VG dataset. Xuewei Li 0003, Guangcong Zheng, Yunlong Yu 0001, Naye Ji, Xi Li 0001 |
IEEE Trans. Image Process. | 2 |
| 2025 | Decoupling Discriminative Attributes for Few-Shot Fine-Grained RecognitionabstractFew-shot fine-tuning of pre-trained vision-language models (VLMs) for downstream tasks has gained widespread attention for reducing data annotation efforts while maintaining high performance. However, we observe that VLMs excel in excluding most incorrect classes in fine-grained recognition tasks, but struggles with a small set of confusing categories, which are typically highly similar subspecies. Existing few-shot fine-tuning methods attempt to directly recognize the correct category among all predefined classes, limiting their ability to capture discriminative features for those confusing categories. This raises an intriguing question: Can we specifically extract useful information from confusing classes to enhance fine-grained recognition performance? Based on this insight, we propose a hierarchical few-shot fine-tuning framework to address the severe confusion problem while ensuring the interpretability, namely Attribute-Decoupled Discriminator (AttrDD). Instead of thinking once among all classes, AttrDD employs a two-stage recognition, "think through" then "think smart". Specifically, in the first phase, a representative VLM, CLIP, is fine-tuned to select the Top-K confusing classes. In the second phase, we leverage the knowledge of large language models (LLMs) to generate fixed format descriptions of attribute differences between these confusing classes via in-context learning. Attribute-decoupled classifications are then conducted to capture fine-grained discriminative features. To achieve parameter-efficient fine-tuning, we introduce a lightweight attention adapter for each phase to align image features with task-specific textual features and LLM-generated textual features. Extensive experiments on 9 fine-grained recognition benchmarks demonstrate that AttrDD consistently outperforms existing baselines by wide margins. Yehao Lu, Chaoxiang Cai, Wei Su 0009, Guangcong Zheng, Xuewei Li 0003, Xi Li 0001 |
IEEE Trans. Image Process. | 4 |
| 2024 | BEVSpread: Spread Voxel Pooling for Bird's-Eye-View Representation in Vision-Based Roadside 3D Object DetectionabstractVision-based roadside 3D object detection has attracted rising attention in autonomous driving domain, since it en-compasses inherent advantages in reducing blind spots and expanding perception range. While previous work mainly focuses on accurately estimating depth or height for 2D-to-3D mapping, ignoring the position approximation error in the voxel pooling process. Inspired by this insight, we propose a novel voxel pooling strategy to reduce such error, dubbed BEVSpread. Specifically, instead of bringing the image features contained in a frustum point to a single BEV grid, BEVSpread considers each frustum point as a source and spreads the image features to the surrounding BEV grids with adaptive weights. To achieve superior prop- agation performance, a specific weight function is designed to dynamically control the decay speed of the weights according to distance and depth. Aided by customized CUDA parallel acceleration, BEVSpread achieves comparable inference time as the original voxel pooling. Extensive experiments on two large-scale roadside benchmarks demonstrate that, as a plug-in, BEVSpread can significantly improve the performance of existing frustum-based BEV methods by a large margin of (1.12, 5.26, 3.01) AP in vehicle, pedestrian and cyclist. The source code will be made publicly available at BEVSpread. Yehao Lu, Guangcong Zheng, Shuigen Zhan, Xiaoqing Ye, Zichang Tan, Jingdong Wang 0001, Gaoang Wang, Xi Li 0001 |
CVPR | 3 |
| 2024 | A Survey of Multimodal Controllable Diffusion Models
Guangcong Zheng, Tian-Rui Yang, Jingdong Wang 0001, Xi Li 0001 |
J. Comput. Sci. Technol. | 2 |
| 2023 | LayoutDiffusion: Controllable Diffusion Model for Layout-to-Image GenerationabstractRecently, diffusion models have achieved great success in image synthesis. However, when it comes to the layout-to-image generation where an image often has a complex scene of multiple objects, how to make strong control over both the global layout map and each detailed object remains a challenging task. In this paper, we propose a diffusion model named LayoutDiffusion that can obtain higher generation quality and greater controllability than the previous works. To overcome the difficult multimodal fusion of image and layout, we propose to construct a structural image patch with region information and transform the patched image into a special layout to fuse with the normal layout in a unified form. Moreover, Layout Fusion Module (LFM) and Object-aware Cross Attention (OaCA) are proposed to model the relationship among multiple objects and designed to be object-aware and position-sensitive, allowing for precisely controlling the spatial related information. Extensive experiments show that our LayoutDiffusion out-performs the previous SOTA methods on FID, CAS by relatively 46.35%,26.70% on COCO-stuff and 44.29%,41.82% on VG. Code is available at https://github.com/ZGCTroy/LayoutDiffusion. Guangcong Zheng, Xianpan Zhou, Xuewei Li 0003, Zhongang Qi, Ying Shan, Xi Li 0001 |
CVPR | 1 |
| 2023 | Uncertainty-Aware Scene Graph Generation
Xuewei Li 0003, Guangcong Zheng, Yunlong Yu 0001, Xi Li 0001 |
Pattern Recognit. Lett. | 3 |
| 2022 | Entropy-Driven Sampling and Training Scheme for Conditional Diffusion Generation
Guangcong Zheng, Shengming Li, Hui Wang 0107, Taiping Yao, Shouhong Ding, Xi Li 0001 |
ECCV (22) | 1 |
| 2021 | BEKT: Deep Knowledge Tracing with Bidirectional Encoder Representations from Transformers
Zejie Tian, Guangcong Zheng, Brendan Flanagan |
ICCE | 2 |
| 2020 | Modeling on virtual network embedding using reinforcement learningabstractSummary It is well known that virtual network (VN) embedding (VNE) aims to solve how to efficiently allocate physical resources to a VN. However, this issue has been proved to be an NP‐hard problem. Besides, as most of the existing approaches are based on heuristic algorithms, which is easy to fall into local optimal. To address the challenge, we formalize the problem as a mixed integer programming problem and propose a novel VNE method based on reinforcement learning in this article. And to solve the problem, we introduce a pointer network to generate virtual node mapping strategies through an attention mechanism, and design a reward function related to link resource consumption to build the connection between node mapping and link mapping stages of VNE. In addition, we present a policy gradient optimization mechanism to leverage the reward information obtained from the sampled solutions, and design an active search based process to automatically update the parameters of the neural network and to obtain near‐optimal embedding solution. The experimental results show that the proposed method can improve the performance in average physical node utilization and long‐term revenue to cost ratio comparing than that of the existing models. Cong Wang 0009, Fanghui Zheng, Guangcong Zheng, Sancheng Peng, Zejie Tian, Yujia Guo, Guorui Li, Ying Yuan 0001 |
Concurr. Comput. Pract. Exp. | 3 |