EDBT 2026 Demo / reviewers in the wild / expert
Dian Zheng
dblp:321/3587
· DBLP profile ↗
16ranked-venue papers
4as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 3 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EchoGen: Cycle-Consistent Learning for Unified Layout-Image Generation and UnderstandingabstractIn this work, we present EchoGen, a unified framework for layout-to-image generation and image grounding, capable of generating images with both accurate layout and high fidelity to the text description.(e.g., spatial relationship), and grounding the image robustly at the same time. We believe that image grounding possesses strong text and layout understanding abilities, which can compensate for the corresponding limitations in layout-to-image generation. At the same time, images generated from layouts exhibit high diversity in content, thereby enhancing the robustness of image grounding. Jointly training both tasks within a unified model can promote performance improvements for each. However, we identify that this joint training paradigm encounters several optimization challenges and results in restricted performance. To address these issues, we propose progressive training strategies. First, the Parallel Multi-Task Pre-training (PMTP) stage equips the model with basic abilities for both tasks, leveraging shared tokens to accelerate training. Next, the Dual Joint Optimization (DJO) stage exploits task duality to sequentially integrate the two tasks, enabling unified optimization. Finally, the Cycle RL stage eliminates reliance on visual supervision by using consistency constraints as rewards, significantly enhancing the model’s unified capabilities via the GRPO strategy. Extensive experiments demonstrate state-of-the-art results on both layout-to-image generation and image grounding benchmarks, and reveal clear synergistic gains from optimizing the two tasks together. Dian Zheng, Jianxiong Gao, Bin Liu 0016 |
AAAI | 3 |
| 2026 | Uni-MMMU: A Massive Multi-discipline Multimodal Unified BenchmarkabstractKai Zou, Ziqi Huang, Yuhao Dong, Shulin Tian, Dian Zheng, Hongbo Liu, Jingwen He, Bin Liu, Yu Qiao, Ziwei Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yuhao Dong, Shulin Tian, Dian Zheng, Jingwen He, Bin Liu 0016, Yu Qiao 0001, Ziwei Liu 0002 |
ACL (1) | 5 |
| 2025 | SpatialDreamer: Self-supervised Stereo Video Synthesis from Monocular InputabstractStereo video synthesis from monocular input is challenging in spatial computing and virtual reality due to the lack of high-quality stereo video pairs for training and the difficulty of maintaining spatio-temporal consistency between frames. Existing methods primarily address these issues by directly applying novel view synthesis (NVS) techniques to video, while facing limitations such as the inability to effectively represent dynamic scenes and the requirement for extensive training data. In this paper, we introduce a novel self-supervised stereo video synthesis paradigm via a video diffusion model, termed SpatialDreamer, which meets the challenges head-on. Firstly, to address the stereo video data insufficiency, we propose a Depth based Video Generation module DVG, which employs a forward-backward rendering mechanism to generate paired videos with geometric and temporal priors. Leveraging data generated by DVG, we propose RefinerNet along with a self-supervised synthetic framework designed to facilitate efficient and dedicated training. More importantly, we devise a consistency control module, which consists of a metric of stereo deviation strength and a Temporal Interaction Learning module TIL for geometric and temporal consistency ensurance respectively. We evaluated the proposed method against various benchmark methods, with the results showcasing its superior performance. Our project website is at: https://spatialdreamer.github.io. Yangqi Long, Congzhentao Huang, Cao Li, Chengfei Lv, Dian Zheng |
CVPR | 7 |
| 2025 | Panorama Generation From NFoV Image Done RightabstractGenerating 360-degree panoramas from narrow field of view (NFoV) image is a promising computer vision task for Virtual Reality (VR) applications. Existing methods mostly assess the generated panoramas with InceptionNet or CLIP based metrics, which tend to perceive the image quality and is not suitable for evaluating the distortion. In this work, we first propose a distortion-specific CLIP, named Distort-CLIP to accurately evaluate the panorama distortion and discover the "visual cheating" phenomenon in previous works (i.e., tending to improve the visual results by sacrificing distortion accuracy). This phenomenon arises because prior methods employ a single network to learn the distinct panorama distortion and content completion at once, which leads the model to prioritize optimizing the latter. To address the phenomenon, we propose PanoDecouple, a decoupled diffusion model framework, which decouples the panorama generation into distortion guidance and content completion, aiming to generate panoramas with both accurate distortion and visual appeal. Specifically, we design a DistortNet for distortion guidance by imposing panorama-specific distortion prior and a modified condition registration mechanism; and a ContentNet for content completion by imposing perspective image information. Additionally, a distortion correction loss function with Distort-CLIP is introduced to constrain the distortion explicitly. The extensive experiments validate that PanoDecouple surpasses existing methods both in distortion and visual metrics. Dian Zheng, Xiao-Ming Wu 0002, Cao Li, Chengfei Lv, Jianfang Hu, Wei-Shi Zheng 0001 |
CVPR | 1 |
| 2025 | Decoupled Distillation to Erase: A General Unlearning Method for Any Class-centric TasksabstractIn this work, we present DEcoupLEd Distillation To Erase (DELETE), a general and strong unlearning method for any class-centric tasks. To derive this, we first propose a theoretical framework to analyze the general form of unlearning loss and decompose it into forgetting and retention terms. Through the theoretical framework, we point out that a class of previous methods could be mainly formulated as a loss that implicitly optimizes the forgetting term while lacking supervision for the retention term, disturbing the distribution of pre-trained model and struggling to adequately preserve knowledge of the remaining classes. To address it, we refine the retention term using "dark knowledge" and propose a mask distillation unlearning method. By applying a mask to separate forgetting logits from retention logits, our approach optimizes both the forgetting and refined retention components simultaneously, retaining knowledge of the remaining classes while ensuring thorough forgetting of the target class. Without access to the remaining data or intervention (i.e., used in some works), we achieve state-of-the-art performance across various benchmarks. What’s more, DELETE is a general solution that can be applied to various downstream tasks, including face recognition, backdoor defense, and semantic segmentation with great performance. Dian Zheng, Qijie Mo, Renjie Lu 0002, Kun-Yu Lin, Wei-Shi Zheng 0001 |
CVPR | 2 |
| 2025 | ShotBench: Expert-Level Cinematic Understanding in Vision-Language ModelsabstractRecent Vision-Language Models (VLMs) have shown strong performance in general-purpose visual understanding and reasoning, but their ability to comprehend the visual grammar of movie shots remains underexplored and insufficiently evaluated. To bridge this gap, we present \textbf{ShotBench}, a dedicated benchmark for assessing VLMs’ understanding of cinematic language. ShotBench includes 3,049 still images and 500 video clips drawn from more than 200 films, with each sample annotated by trained annotators or curated from professional cinematography resources, resulting in 3,608 high-quality question-answer pairs. We conduct a comprehensive evaluation of over 20 state-of-the-art VLMs across eight core cinematography dimensions. Our analysis reveals clear limitations in fine-grained perception and cinematic reasoning of current VLMs. To improve VLMs capability in cinematography understanding, we construct a large-scale multimodal dataset, named ShotQA, which contains about 70k Question-Answer pairs derived from movie shots.
Besides, we propose ShotVL and train this VLM model with a two-stage training strategy, integrating both supervised fine-tuning and Group Relative Policy Optimization (GRPO). Experimental results demonstrate that our model achieves substantial improvements, surpassing all existing strongest open-source and proprietary models evaluated on ShotBench, establishing a new state-of-the-art performance. Jingwen He, Dian Zheng, Yuhao Dong, Fan Zhang 0045, Yinan He, Weichao Chen 0001, Yu Qiao 0001, Wanli Ouyang, Shengjie Zhao 0001, Ziwei Liu 0002 |
NeurIPS | 4 |
| 2025 | Deep Concept Forgetting in Text-to-Image Diffusion Models
Dian Zheng, Xiao-Ming Wu 0002, Wei-Shi Zheng 0001 |
PRCV (2) | 2 |
| 2025 | DiffuVolume: Diffusion Model for Volume based Stereo Matching
Dian Zheng, Xiao-Ming Wu 0002, Zuhao Liu 0002, Jingke Meng, Wei-Shi Zheng 0001 |
Int. J. Comput. Vis. | 1 |
| 2024 | Dexterous Grasp TransformerabstractIn this work, we propose a novel discriminative frame-work for dexterous grasp generation, named Dexterous Grasp TRansformer (DGTR), capable of predicting a di-verse set of feasible grasp poses by processing the object point cloud with only one forward pass. We formulate dex-terous grasp generation as a set prediction task and design a transformer-based grasping model for it. However, we identify that this set prediction paradigm encounters sev-eral optimization challenges in the field of dexterous grasping and results in restricted performance. To address these issues, we propose progressive strategies for both the training and testing phases. First, the dynamic-static matching training (DSMT) strategy is presented to enhance the opti-mization stability during the training phase. Second, we in-troduce the adversarial-balanced test-time adaptation (AB-TTA) with a pair of adversarial losses to improve grasping quality during the testing phase. Experimental results on the DexGraspNet dataset demonstrate the capability of DGTR to predict dexterous grasp poses with both high quality and diversity. Notably, while keeping high qual-ity, the diversity of grasp poses predicted by DGTR sig-nificantly outperforms previous works in multiple metrics without any data pre-processing. Codes are available at https://github.com/iSEE-Laboratory/DGTR. Guo-Hao Xu, Yi-Lin Wei, Dian Zheng, Xiao-Ming Wu 0002, Wei-Shi Zheng 0001 |
CVPR | 3 |
| 2024 | Selective Hourglass Mapping for Universal Image Restoration Based on Diffusion ModelabstractUniversal image restoration is a practical and poten-tial computer vision task for real-world applications. The main challenge of this task is handling the different degra-dation distributions at once. Existing methods mainly utilize task-specific conditions (e.g., prompt) to guide the model to learn different distributions separately, named multi-partite mapping. However, it is not suitable for universal model learning as it ignores the shared information between different tasks. In this work, we propose an advanced selective hourglass mapping strategy based on diffusion model, termed DiffUIR. Two novel considerations make our Dif-fUIR non-trivial. Firstly, we equip the model with strong condition guidance to obtain accurate generation direction of diffusion model (selective). More importantly, DiffUIR integrates a flexible shared distribution term (SDT) into the diffusion algorithm elegantly and naturally, which gradually maps different distributions into a shared one. In the reverse process, combined with SDT and strong condition guidance, DiffUIR iteratively guides the shared distribution to the task-specific distribution with high image quality (hourglass). Without bells and whistles, by only modifying the mapping strategy, we achieve state-of-the-art performance on five image restoration tasks, 22 benchmarks in the universal setting and zero-shot generalization setting. Surprisingly, by only using a lightweight model (only 0.89M), we could achieve outstanding performance. The source code and pre-trained models are available at https://github.com/iSEE-Laboratory/DiffUIR. Dian Zheng, Xiao-Ming Wu 0002, Shuzhou Yang, Jianfang Hu, Wei-Shi Zheng 0001 |
CVPR | 1 |
| 2024 | An Economic Framework for 6-DoF Grasp Detection
Xiao-Ming Wu 0002, Jia-Feng Cai, Jian-Jian Jiang, Dian Zheng, Yi-Lin Wei, Wei-Shi Zheng 0001 |
ECCV (27) | 4 |
| 2023 | Generating Anomalies for Video Anomaly Detection with Prompt-based Feature MappingabstractAnomaly detection in surveillance videos is a challenging computer vision task where only normal videos are available during training. Recent work released the first virtual anomaly detection dataset to assist real-world detection. However, an anomaly gap exists because the anomalies are bounded in the virtual dataset but unbounded in the real world, so it reduces the generalization ability of the virtual dataset. There also exists a scene gap between virtual and real scenarios, including scene-specific anomalies (events that are abnormal in one scene but normal in another) and scene-specific attributes, such as the viewpoint of the surveillance camera. In this paper, we aim to solve the problem of the anomaly gap and scene gap by proposing a prompt-based feature mapping framework (PFMF). The PFMF contains a mapping network guided by an anomaly prompt to generate unseen anomalies with unbounded types in the real scenario, and a mapping adaptation branch to narrow the scene gap by applying domain classifier and anomaly classifier. The proposed framework outperforms the state-of-the-art on three benchmark datasets. Extensive ablation experiments also show the effectiveness of our framework design. Zuhao Liu 0002, Xiao-Ming Wu 0002, Dian Zheng, Kun-Yu Lin, Wei-Shi Zheng 0001 |
CVPR | 3 |
| 2023 | Estimator Meets Equilibrium Perspective: A Rectified Straight Through Estimator for Binary Neural Networks TrainingabstractBinarization of neural networks is a dominant paradigm in neural networks compression. The pioneering work BinaryConnect uses Straight Through Estimator (STE) to mimic the gradients of the sign function, but it also causes the crucial inconsistency problem. Most of the previous methods design different estimators instead of STE to mitigate it. However, they ignore the fact that when reducing the estimating error, the gradient stability will decrease concomitantly. These highly divergent gradients will harm the model training and increase the risk of gradient vanishing and gradient exploding. To fully take the gradient stability into consideration, we present a new perspective to the BNNs training, regarding it as the equilibrium between the estimating error and the gradient stability. In this view, we firstly design two indicators to quantitatively demonstrate the equilibrium phenomenon. In addition, in order to balance the estimating error and the gradient stability well, we revise the original straight through estimator and propose a power function based estimator, Rectified Straight Through Estimator (ReSTE for short). Comparing to other estimators, ReSTE is rational and capable of flexibly balancing the estimating error with the gradient stability. Extensive experiments on CIFAR-10 and ImageNet datasets show that ReSTE has excellent performance and surpasses the state-of-the-art methods without any auxiliary modules or losses. Xiao-Ming Wu 0002, Dian Zheng, Zuhao Liu 0002, Wei-Shi Zheng 0001 |
ICCV | 2 |
| 2023 | MarsNet: Automated Rock Segmentation With Transformers for Tianwen-1 MissionabstractThe Mars exploration mission of China named Tianwen-1 is being carried out as scheduled. The Navigation and Terrain Cameras (NaTeCam) equipped on the Zhurong Rover play an essential role in obstacle recognition. The main obstacles on the Martian surface are rocks of different sizes, which influence the path planning of Zhurong Rover in scientific exploration. Most existing semantic segmentation methods are based on the U-Net architecture with ResNet or other backbones, and features extracted by these methods lack long-range dependencies. To fully exploit the context information, we propose the MarsNet framework for the Mars image, which combines transformers with the convolutional neural network (CNN) as the backbone, and hybrid dilated convolution (HDC) is also employed to the decoder path to help detect the huge rocks. Besides, since there are few open-source datasets for rock segmentation for Mars, we establish a segmentation dataset from the Martian surface image, named TWMARS, captured by NaTeCam. Extensive experiments are conducted on the TWMARS dataset, and the experimental results demonstrate that MarsNet achieves accurate rock segmentation and outperforms state-of-the-art methods. The source code is available athttps://github.com/BUPT-ANT-1007/MarsNet. Weikun Lv, Linhui Wei, Dian Zheng, Yu Liu 0001, Yumei Wang |
IEEE Geosci. Remote. Sens. Lett. | 3 |
| 2022 | Underwater Stereo Matching Via Unsupervised Appearance And Feature Adaptation NetworksabstractStereo matching has been widely used to estimate depth maps in terrestrial environments. However, it is difficult to achieve appealing performance in underwater environments, since adequate underwater stereo data with groundtruth depth information is not easily available for training an underwater depth estimation model. In addition, the domain gap also leads to the failure of directly applying existing models of terrestrial scenes to underwater scenes. Therefore, this paper proposes a novel underwater depth estimation network which can infer depth maps from real underwater stereo images in an unsupervised adaptation manner. The proposed learning pipeline contains style adaptation (SA) in appearance space and feature adaptation (FA) in semantic space to progressively adapt the depth estimation models to underwater domain. Experimental results show that by integrating the proposed adaptation modules into the off-the-shelf stereo matching backbones, our method achieves a superior performance of underwater depth estimation compared to other state-of-the-art methods. Yazhi Yuan, Xinchen Ye, Dian Zheng, Rui Xu 0002 |
ICASSP | 4 |
| 2022 | MetaMars: 3DoF+ Roaming With Panoramic Stitching for Tianwen-1 MissionabstractIn China’s Tianwen-1 mission for Mars exploration, Zhurong rover carries the Navigation and Terrain Camera (NaTeCam), and collects a lot of terrains and topographic data. For the scientific popularization of Mars, a virtual reality roaming based on panoramic stitching presents the Martian surface in detail and provides an immersive experience for users. However, the current Mars roaming systems focus on global information instead of details, and the traditional panoramic stitching methods are not suitable for Mars images. This letter proposes a panoramic stitching method and develops a 3DoF+ Mars roaming system. The proposed Stitching-Combines-Features-and-Projection (SCFP) jointly considers the extracted features of images and projection based on intrinsic camera parameters, which improves the matching effect and execution efficiency. Experimental results indicate that SCFP reduces the computational time and improves stitching results than state-of-the-art methods. Further, we design and deploy the 3DoF+ roaming system based on panoramic stitching with datasets obtained from the Tianwen-1 mission. Dian Zheng, Linhui Wei, Yu Liu 0001, Yumei Wang |
IEEE Geosci. Remote. Sens. Lett. | 1 |