EDBT 2026 Demo / reviewers in the wild / expert
Haoyu Chen 0003
dblp:146/8170-3
· DBLP profile ↗
22ranked-venue papers
5as first author
21since 2021 · last 2026
0000-0001-7618-9733ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 18 · 5 first-author · 17 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 4 first-author · 14 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Standing on the Giants: Informative Messenger Prompts With Self-Adapter for Image RestorationabstractDespite the recent advances in the application of diffusion models to the realm of image restoration, their inherent stochastic nature can often lead to inaccuracies in reconstructing spatial structures and fine details. This paper introduces an innovative paradigm that harnesses the rich knowledge encapsulated in existing high-level pre-trained models to guide the diffusion process, offering a flexible and potent approach to restoration tasks. However, there are significant challenges in leveraging a pre-trained model for image restoration, including a data gap between natural clean images and degradation images, and paradigm differences that cause insufficient intermediate knowledge. To tackle these issues, we introduce informative Messenger prompts and a Self-adapter for the pre-trained model, whose appropriate information acts as explicit constraints for diffusion, enabling reliable result generation. Specifically, ourMeSa-IRsuccessfully adapts to feature exploration for degraded samples, via disseminating distinctive information from degraded instances. Furthermore, it bolsters knowledge representation through the bidirectional interchange of hierarchical information facilitated by the innovative use of messenger prompts. Experimental results demonstrate the state-of-the-art performance of our framework on five tasks in terms of perceptual and distortion metrics. We will release codes at https://github.com/Ephemeral182/MeSa-IR. Sixiang Chen, Tian Ye 0001, Yulun Zhang 0001, Haoyu Chen 0003, Zhaohu Xing, Fugee Tsung, Lei Zhu 0003 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | PromptHaze: Prompting Real-world Dehazing via Depth Anything ModelabstractReal-world image dehazing remains a challenging task due to the diverse nature of haze degradation and the lack of large-scale paired datasets. Existing methods based on hand-crafted priors or generative priors struggle to recover accurate backgrounds and fine details from dense haze regions. In this work, we propose a novel paradigm, PromptHaze, for real-world image dehazing via the depth prompt from the Depth Anything model. By employing a prompt-by-prompt strategy, our method iteratively updates the depth prompt and progressively restores the background through a dehazing network with controllable dehazing strength. Extensive experiments on widely-used real-world dehazing benchmarks demonstrate the superiority of PromptHaze in recovering authentic backgrounds and fine details from various haze scenes, outperforming state-of-the-art methods across multiple quality metrics. Tian Ye 0001, Sixiang Chen, Haoyu Chen 0003, Wenhao Chai, Zhaohu Xing, Wenxue Li 0003, Lei Zhu 0003 |
AAAI | 3 |
| 2025 | POSTA: A Go-to Framework for Customized Artistic Poster GenerationabstractPoster design is a critical medium for visual communication. Prior work has explored automatic poster design using deep learning techniques, but these approaches lack text accuracy, user customization, and aesthetic appeal, limiting their applicability in artistic domains such as movies and exhibitions, where both clear content delivery and visual impact are essential. To address these limitations, we present POSTA: a modular framework powered by diffusion models and multimodal large language models (MLLMs) for customized artistic poster generation. The framework consists of three modules. Background Diffusion creates a themed background based on user input. Design MLLM then generates layout and typography elements that align with and complement the background style. Finally, to enhance the poster’s aesthetic appeal, ArtText Diffusion applies additional stylization to key text elements. The final result is a visually cohesive and appealing poster, with a fully modular process that allows for complete customization. To train our models, we develop the PosterArt dataset, comprising high-quality artistic posters annotated with layout, typography, and pixel-level stylized text segmentation. Our comprehensive experimental analysis demonstrates POSTA’s exceptional controllability and design diversity, outperforming existing models in both text accuracy and aesthetic quality. Haoyu Chen 0003, Wenbo Li 0002, Tian Ye 0001, Songhua Liu, Ying-Cong Chen, Lei Zhu 0003, Xinchao Wang |
CVPR | 1 |
| 2025 | JarvisIR: Elevating Autonomous Driving Perception with Intelligent Image RestorationabstractVision-centric perception systems struggle with unpredictable and coupled weather degradations in the wild. Current solutions are often limited, as they either depend on specific degradation priors or suffer from significant domain gaps. To enable robust and autonomous operation in real-world conditions, we propose JarvisIR, a VLM-powered agent that leverages the VLM as a controller to manage multiple expert restoration models. To further enhance system robustness, reduce hallucinations, and improve generalizability in real-world adverse weather, JarvisIR employs a novel two-stage framework consisting of supervised fine-tuning and human feedback alignment. Specifically, to address the lack of paired data in real-world scenarios, the human feedback alignment enables the VLM to be fine-tuned effectively on large-scale real-world data in an unsupervised manner. To support the training and evaluation of JarvisIR, we introduce CleanBench, a comprehensive dataset consisting of high-quality and large-scale instruction-responses pairs, including 150K synthetic entries and 80K real entries. Extensive experiments demonstrate that JarvisIR exhibits superior decision-making and restoration capabilities. Compared with existing methods, it achieves a 50% improvement in the average of all perception metrics on CleanBench-Real. Yunlong Lin, Zixu Lin, Haoyu Chen 0003, Panwang Pan, Chenxin Li, Sixiang Chen, Kairun Wen, Yeying Jin, Wenbo Li 0002, Xinghao Ding |
CVPR | 3 |
| 2025 | Genhaze: Pioneering Controllable One-Step Realistic Haze Generation for Real-World DehazingabstractReal-world image dehazing is crucial for enhancing visual quality in computer vision applications. However, existing physics-based haze generation paradigms struggle to model the complexities of real-world haze and lack controllability, limiting the performance of existing baselines on real-world images. In this paper, we introduce GenHaze, a pioneering haze generation framework that enables the one-step generation of high-quality, reference-controllable hazy images. GenHaze leverages the pre-trained latent diffusion model (LDM) with a carefully designed clean-to-haze generation protocol to produce realistic hazy images. Additionally, by leveraging its fast, controllable generation of paired highquality hazy images, we illustrate that existing dehazing baselines can be unleashed in a simple and efficient manner. Extensive experiments indicate that GenHaze achieves visually convincing and quantitatively superior hazy images. It also significantly improves multiple existing dehazing models across 7 non-reference metrics with minimal fine-tuning epochs. Our work demonstrates that LDM possesses the potential to generate realistic degradations, providing an effective alternative to prior generation pipelines. Sixiang Chen, Tian Ye 0001, Yunlong Lin, Yeying Jin, Haoyu Chen 0003, Jianyu Lai, Song Fei, Zhaohu Xing, Fugee Tsung, Lei Zhu 0003 |
ICCV | 6 |
| 2025 | Turbo2K: Towards Ultra-Efficient and High-Quality 2K Video Synthesis
Wenbo Li 0002, Zhongdao Wang, Haoze Sun, Bangzhen Liu, Haoyu Chen 0003, Aoxue Li, Lei Zhu 0003 |
ICCV | 6 |
| 2025 | JarvisArt: Liberating Human Artistic Creativity via an Intelligent Photo Retouching AgentabstractPhoto retouching has become integral to contemporary visual storytelling, enabling users to capture aesthetics and express creativity. While professional tools such as Adobe Lightroom offer powerful capabilities, they demand substantial expertise and manual effort. In contrast, existing AI-based solutions provide automation but often suffer from limited adjustability and poor generalization, failing to meet diverse and personalized editing needs. To bridge this gap, we introduce JarvisArt, a multi-modal large language model (MLLM)-driven agent that understands user intent, mimics the reasoning process of professional artists, and intelligently coordinates over 200 retouching tools within Lightroom. JarvisArt undergoes a two-stage training process: an initial Chain-of-Thought supervised fine-tuning to establish basic reasoning and tool-use skills, followed by Group Relative Policy Optimization for Retouching (GRPO-R) to further enhance its decision-making and tool proficiency. We also propose the Agent-to-Lightroom Protocol to facilitate seamless integration with Lightroom. To evaluate performance, we develop MMArt-Bench, a novel benchmark constructed from real-world user edits. JarvisArt demonstrates user-friendly interaction, superior generalization, and fine-grained control over both global and local adjustments, paving a new avenue for intelligent photo retouching. Notably, it outperforms GPT-4o with a 60\% improvement in average pixel-level metrics on MMArt-Bench for content fidelity, while maintaining comparable instruction-following capabilities. Yunlong Lin, Zixu Lin, Kunjie Lin, Jinbin Bai, Panwang Pan, Chenxin Li, Haoyu Chen 0003, Zhongdao Wang, Xinghao Ding, Wenbo Li 0002, Shuicheng Yan |
NeurIPS | 7 |
| 2025 | CamEdit: Continuous Camera Parameter Control for Photorealistic Image EditingabstractRecent advances in diffusion models have substantially improved text-driven image editing. However, existing frameworks based on discrete textual tokens struggle to support continuous control over camera parameters and smooth transitions in visual effects. These limitations hinder their applications to realistic, camera-aware, and fine-grained editing tasks. In this paper, we present CamEdit, a diffusion-based framework for photorealistic image editing that enables continuous and semantically meaningful manipulation of common camera parameters such as aperture and shutter speed. CamEdit incorporates a continuous parameter prompting mechanism and a parameter-aware modulation module that guides the model in smoothly adjusting focal plane, aperture, and shutter speed, reflecting the effects of varying camera settings within the diffusion process. To support supervised learning in this setting, we introduce CamEdit50K, a dataset specifically designed for photorealistic image editing with continuous camera parameter settings. It contains over 50k image pairs combining real and synthetic data with dense camera parameter variations across diverse scenes. Extensive experiments demonstrate that CamEdit enables flexible, consistent, and high-fidelity image editing, achieving state-of-the-art performance in camera-aware visual manipulation and fine-grained photographic control. Xinran Qin, Zhixin Wang, Haoyu Chen 0003, Renjing Pei, Wenbo Li 0002, Xiaochun Cao |
NeurIPS | 4 |
| 2025 | PocketSR: The Super-Resolution Expert in Your Pocket MobilesabstractReal-world image super-resolution (RealSR) aims to enhance the visual quality of in-the-wild images, such as those captured by mobile phones. While existing methods leveraging large generative models demonstrate impressive results, the high computational cost and latency make them impractical for edge deployment. In this paper, we introduce PocketSR, an ultra-lightweight, single-step model that brings generative modeling capabilities to RealSR while maintaining high fidelity. To achieve this, we design LiteED, a highly efficient alternative to the original computationally intensive VAE in SD, reducing parameters by 97.5\% while preserving high-quality encoding and decoding. Additionally, we propose online annealing pruning for the U-Net, which progressively shifts generative priors from heavy modules to lightweight counterparts, ensuring effective knowledge transfer and further optimizing efficiency. To mitigate the loss of prior knowledge during pruning, we incorporate a multi-layer feature distillation loss. Through an in-depth analysis of each design component, we provide valuable insights for future research. PocketSR, with a model size of 146M parameters, processes 4K images in just 0.8 seconds, achieving a remarkable speedup over previous methods. Notably, it delivers performance on par with state-of-the-art single-step and even multi-step RealSR models, making it a highly practical solution for edge-device applications. Haoze Sun, Linfeng Jiang, Renjing Pei, Zhixin Wang, Haoyu Chen 0003, Fenglong Song, Yujiu Yang 0001, Wenbo Li 0002 |
NeurIPS | 8 |
| 2025 | Triplane-Smoothed Video Dehazing with CLIP-Enhanced Generalization
Haoyu Chen 0003, Tian Ye 0001, Lei Zhu 0003 |
Int. J. Comput. Vis. | 2 |
| 2024 | Low-Res Leads the Way: Improving Generalization for Super-Resolution by Self-Supervised LearningabstractFor image super-resolution (SR), bridging the gap between the performance on synthetic datasets and real-world degradation scenarios remains a challenge. This work introduces a novel “Low-Res Leads the Way” (LWay) training framework, merging Supervised Pre-training with Self-supervised Learning to enhance the adaptability of SR models to real-world images. Our approach utilizes a low-resolution (LR) reconstruction network to extract degradation embeddings from LR images, merging them with super-resolved outputs for LR reconstruction. Leveraging unseen LR images for self-supervised learning guides the model to adapt its modeling space to the target domain, facili-tating fine-tuning of SR models without requiring paired high-resolution (HR) images. The integration of Discrete Wavelet Transform (DWT)further refines the focus on high-frequency details. Extensive evaluations show that our method significantly improves the generalization and de-tail restoration capabilities of SR models on unseen real-world datasets, outperforming existing methods. Our training regime is universally compatible, requiring no network architecture modifications, making it a practical solution for real-world SR applications. Haoyu Chen 0003, Wenbo Li 0002, Jinjin Gu, Haoze Sun, Xueyi Zou, Zhensong Zhang, Youliang Yan, Lei Zhu 0003 |
CVPR | 1 |
| 2024 | CoSeR: Bridging Image and Language for Cognitive Super-ResolutionabstractExisting super-resolution (SR) models primarily focus on restoring local texture details, often neglecting the global semantic information within the scene. This oversight can lead to the omission of crucial semantic details or the intro-duction of inaccurate textures during the recovery process. In our work, we introduce the Cognitive Super-Resolution (CoSeR) framework, empowering SR models with the ca-pacity to comprehend low-resolution images. We achieve this by marrying image appearance and language under-standing to generate a cognitive embedding, which not only activates prior information from large text-to-image diffusion models but also facilitates the generation of high-quality reference images to optimize the SR process. To fur-ther improve image fidelity, we propose a novel condition injection scheme called “Ali-in-Attention ”, consolidating all conditional information into a single module. Conse-quently, our method successfully restores semantically cor-rect and photorealistic details, demonstrating state-of-the-art performance across multiple benchmarks. Project page: https://coser-main.github.io/ Haoze Sun, Wenbo Li 0002, Jianzhuang Liu, Haoyu Chen 0003, Renjing Pei, Xueyi Zou, Youliang Yan, Yujiu Yang 0001 |
CVPR | 4 |
| 2024 | Semi-supervised Video Desnowing Network via Temporal Decoupling Experts and Distribution-Driven Contrastive Regularization
Angelica I. Avilés-Rivero, Sixiang Chen, Haoyu Chen 0003, Lei Zhu 0003 |
ECCV (10) | 6 |
| 2024 | RestoreAgent: Autonomous Image Restoration Agent via Multimodal Large Language ModelsabstractNatural images captured by mobile devices often suffer from multiple types of degradation, such as noise, blur, and low light. Traditional image restoration methods require manual selection of specific tasks, algorithms, and execution sequences, which is time-consuming and may yield suboptimal results. All-in-one models, though capable of handling multiple tasks, typically support only a limited range and often produce overly smooth, low-fidelity outcomes due to their broad data distribution fitting. To address these challenges, we first define a new pipeline for restoring images with multiple degradations, and then introduce RestoreAgent, an intelligent image restoration system leveraging multimodal large language models. RestoreAgent autonomously assesses the type and extent of degradation in input images and performs restoration through (1) determining the appropriate restoration tasks, (2) optimizing the task sequence, (3) selecting the most suitable models, and (4) executing the restoration. Experimental results demonstrate the superior performance of RestoreAgent in handling complex degradation, surpassing human experts. Furthermore, the system’s modular design facilitates the fast integration of new tasks and models. Haoyu Chen 0003, Wenbo Li 0002, Jinjin Gu, Sixiang Chen, Tian Ye 0001, Renjing Pei, Kaiwen Zhou 0001, Fenglong Song, Lei Zhu 0003 |
NeurIPS | 1 |
| 2024 | UltraPixel: Advancing Ultra High-Resolution Image Synthesis to New PeaksabstractUltra-high-resolution image generation poses great challenges, such as increased semantic planning complexity and detail synthesis difficulties, alongside substantial training resource demands. We present UltraPixel, a novel architecture utilizing cascade diffusion models to generate high-quality images at multiple resolutions (\textit{e.g.}, 1K, 2K, and 4K) within a single model, while maintaining computational efficiency. UltraPixel leverages semantics-rich representations of lower-resolution images in a later denoising stage to guide the whole generation of highly detailed high-resolution images, significantly reducing complexity. Specifically, we introduce implicit neural representations for continuous upsampling and scale-aware normalization layers adaptable to various resolutions. Notably, both low- and high-resolution processes are performed in the most compact space, sharing the majority of parameters with less than 3$\%$ additional parameters for high-resolution outputs, largely enhancing training and inference efficiency. Our model achieves fast training with reduced data requirements, producing photo-realistic high-resolution images and demonstrating state-of-the-art performance in extensive experiments. Wenbo Li 0002, Haoyu Chen 0003, Renjing Pei, Long Peng 0003, Fenglong Song, Lei Zhu 0003 |
NeurIPS | 3 |
| 2023 | Masked Image Training for Generalizable Deep Image DenoisingabstractWhen capturing and storing images, devices inevitably introduce noise. Reducing this noise is a critical task called image denoising. Deep learning has become the de facto method for image denoising, especially with the emergence of Transformer-based models that have achieved notable state-of-the-art results on various image tasks. However, deep learning-based methods often suffer from a lack of generalization ability. For example, deep models trained on Gaussian noise may perform poorly when tested on other noise distributions. To address this issue, we present a novel approach to enhance the generalization performance of denoising networks, known as masked training. Our method involves masking random pixels of the input image and reconstructing the missing information during training. We also mask out the features in the self-attention layers to avoid the impact of training-testing inconsistency. Our approach exhibits better generalization ability than other deep learning models and is directly applicable to real-world scenarios. Additionally, our interpretability analysis demonstrates the superiority of our method. Haoyu Chen 0003, Jinjin Gu, Yihao Liu 0001, Salma Abdel Magid, Chao Dong 0005, Qiong Wang 0001, Hanspeter Pfister, Lei Zhu 0003 |
CVPR | 1 |
| 2023 | Snow Removal in Video: A New Dataset and A Novel MethodabstractSnowfall is a common weather phenomenon that can severely affect computer vision tasks by obscuring objects and scenes. However, existing deep learning-based snow removal methods are designed for single images only. In this paper, we target a more complex task - video snow removal, which aims to restore the clear video from the snowy video. To facilitate this task, we propose the first high-quality video dataset, which simulates realistic physical characteristics of snow and haze using a rendering engine and augmentation techniques. We also develop a deep learning framework for video snow removal. Specifically, we propose a snow-query temporal aggregation module and a snow-aware contrastive learning loss function. The module aggregates features between video frames and removes snow effectively, while the loss function helps identify and eliminate snow features. We conduct extensive experiments and demonstrate that our proposed dataset is more realistic than previous datasets, and the models trained on it achieve better performance in real-world snowing images. Our proposed method outperforms state-of-the-art video and image-based methods on both synthetic and real snowy videos. Haoyu Chen 0003, Jinjin Gu, Xuequan Lu, Haoming Cai, Lei Zhu 0003 |
ICCV | 1 |
| 2023 | Crafting Training Degradation Distribution for the Accuracy-Generalization Trade-off in Real-World Super-ResolutionabstractSuper-resolution (SR) techniques designed for real-world applications commonly encounter two primary challenges: generalization performance and restoration accuracy. We demonstrate that when methods are trained using complex, large-range degradations to enhance generalization, a decline in accuracy is inevitable. However, since the degradation in a certain real-world applications typically exhibits a limited variation range, it becomes feasible to strike a trade-off between generalization performance and testing accuracy within this scope. In this work, we introduce a novel approach to craft training degradation distributions using a small set of reference images. Our strategy is founded upon the binned representation of the degradation space and the Frechet distance between degradation distributions. Our results indicate that the proposed technique significantly improves the performance of test images while preserving generalization capabilities in real-world applications. Ruofan Zhang, Jinjin Gu, Haoyu Chen 0003, Chao Dong 0005, Yulun Zhang 0001, Wenming Yang |
ICML | 3 |
| 2023 | CPLFormer: Cross-scale Prototype Learning Transformer for Image Snow RemovalabstractRemoving snow from a single image poses a significant challenge within the image restoration domain, as snowfall's effects are in various scales and forms. Existing methods have tried to tackle this issue by using multi-scale approaches, but their reliance on targeted design for handling each single-scale feature has resulted in unsatisfactory performance. This is primarily due to a lack of cross-scale knowledge, making it difficult to effectively handle degradations. To this end, we propose a novel approach, CPLFormer, which uses snow prototypes to own comprehensive clean scene understanding through learning from cross-scale features, outperforming convolutional network and vanilla transformer-based solutions. CPLFormer has several advantages: firstly, learnable snow prototypes learn global context information from multiple scales to uncover hidden clean cues; secondly, prototypes can propagate cross-scale information to each patch through cross-attention to assist with clean patch reconstruction; thirdly, CPLFormer surpasses advanced state-of-the-art desnowing networks and the prevalent universal image restoration transformers on six synthetic and real-world benchmark tests. Sixiang Chen, Tian Ye 0001, Yun Liu 0002, Jinbin Bai, Haoyu Chen 0003, Yunlong Lin, Erkang Chen |
ACM Multimedia | 5 |
| 2023 | Uncertainty-Driven Dynamic Degradation Perceiving and Background Modeling for Efficient Single Image DesnowingabstractSingle-image snow removal aims to restore clean images from heterogeneous and irregular snow degradations. Recent methods utilize neural networks to remove various degradations directly. However, these approaches suffer from the limited ability to flexibly perceive complicated snow degradation patterns and insufficient representation of background structure information. To further improve the performance and generalization ability of snow removal, this paper aims to develop a novel and efficient paradigm from the perspective of degradation perceiving and background modeling. Sixiang Chen, Tian Ye 0001, Chenghao Xue, Haoyu Chen 0003, Yun Liu 0002, Erkang Chen, Lei Zhu 0003 |
ACM Multimedia | 4 |
| 2023 | Mask-Guided Progressive Network for Joint Raindrop and Rain Streak Removal in VideosabstractVideos captured in rainy weather are unavoidably corrupted by both rain streaks and raindrops in driving scenarios, and it is desirable and challenging to recover background details obscured by rain streaks and raindrops. However, existing video rain removal methods often address either video rain streak removal or video raindrop removal, thereby suffer from degraded performance when deal with both simultaneously. The bottleneck is a lack of a video dataset, where each video frame contains both rain streaks and raindrops. To address this issue, we in this work generate a synthesized dataset, namely VRDS, with 102 rainy videos from diverse scenarios, and each video frame has the corresponding rain streak map, raindrop mask, and the underlying rain-free clean image (ground truth). Moreover, we devise a mask-guided progressive video deraining network (ViMP-Net) to remove both rain streaks and raindrops of each video frame. Specifically, we develop an intensity-guided alignment block to predict the rain streak intensity map and remove the rain streaks of the input rainy video at the first stage. Then, we predict a raindrop mask and pass it into a devised mask-guided dual transformer block to learn inter-frame and intra-frame transformer features, which are then fed into a decoder for further eliminating raindrops. Experimental results demonstrate that our ViMP-Net outperforms state-of-the-art methods on our synthetic dataset and real-world rainy videos. Our code is available at https://github.com/TonyHongtaoWu/ViMP-Net. Haoyu Chen 0003, Lei Zhu 0003 |
ACM Multimedia | 3 |
| 2020 | PIPAL: A Large-Scale Image Quality Assessment Dataset for Perceptual Image Restoration
Jinjin Gu, Haoming Cai, Haoyu Chen 0003, Xiaoxing Ye, Jimmy S. J. Ren, Chao Dong 0005 |
ECCV (11) | 3 |