Pengyang Ling

dblp:349/1362 · DBLP profile ↗
← Back
17ranked-venue papers
5as first author
17since 2021 · last 2026
0009-0001-1672-7242ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 3 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 2 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Efficient Haze Removal via Scene Depth Ordering for Robust Traffic Monitoring
abstract
The reliability of vision-based systems, such as traffic monitoring and intelligent driving, is typically compromised in hazy weather owing to diminished visibility. In this paper, we propose a novel efficient image dehazing framework guided by depth order, leveraging the consistency of depth perception to establish strong global constraints for enhanced haze removal. The consistent depth perception ensures that the regions that look farther or closer in hazy images also appear farther or closer in the corresponding dehazing results, substantially avoiding potential visual degradation. To this end, the depth order in hazy images is approximated by the reverse order of color difference between pixel values and global atmospheric light, offering an effective and efficient alternative for depth perception modeling. Subsequently, we have developed a depth order embedded transformation model to estimate the transmission maps jointly constrained by depth order and haze imaging model, ensuring that the depth order remains unchanged in corresponding dehazing results. This model harnesses the extracted depth order as a powerful global constraint for the dehazing process, facilitating the efficient use of global information and thus achieving superior image restoration. Extensive experiments demonstrate that the proposed method can better recover potential structure and vivid color with higher computational efficiency, offering an efficient solution for robust traffic monitoring against hazy weather.
Pengyang Ling, Huaian Chen, Haoxuan Wang 0004, Yuxuan Gu 0001, Yi Jin 0002, Jinjin Zheng, Enhong Chen
IEEE Trans. Intell. Transp. Syst.1
2025 ByTheWay: Boost Your Text-to-Video Generation Model to Higher Quality in a Training-free Way
abstract
The text-to-video (T2V) generation models, offering convenient visual creation, have recently garnered increasing attention. Despite their substantial potential, the generated videos may present artifacts, including structural implausibility, temporal inconsistency, and a lack of motion, often resulting in near-static video. In this work, we have identified a correlation between the disparity of temporal attention maps across different blocks and the occurrence of temporal inconsistencies. Additionally, we have observed that the energy contained within the temporal attention maps is directly related to the magnitude of motion amplitude in the generated videos. Based on these observations, we present ByTheWay, a training-free method to improve the quality of text-to-video generation without introducing additional parameters, augmenting memory or sampling time. Specifically, ByTheWay is composed of two principal components: 1) Temporal Self-Guidance improves the structural plausibility and temporal consistency of generated videos by reducing the disparity between the temporal attention maps across various decoder blocks. 2) Fourier-based Motion Enhancement enhances the magnitude and richness of motion by amplifying the energy of the map. Extensive experiments demonstrate that ByTheWay significantly improves the quality of text-to-video generation with negligible additional cost. Our code is available at: https://github.com/Bujiazi/ByTheWay.
Jiazi Bu, Pengyang Ling, Pan Zhang 0001, Xiaoyi Dong, Yuhang Zang, Yuhang Cao, Dahua Lin, Jiaqi Wang 0003
CVPR2
2025 Improving Visual and Downstream Performance of Low-Light Enhancer with Vision Foundation Models Collaboration
abstract
In this paper, we observe that the collaboration of various foundation models can perceive semantic and degraded information within images, thereby guiding the low-light enhancement process. Specifically, we propose a self-supervised low-light enhancement framework based on the multiple foundation models collaboration (dubbed FoCo), aimed at improving both the visual quality of enhanced images and the performance in high-level applications. At the feature level, FoCo leverages the rich features from various foundation models to enhance the model’s semantic perception during training, thereby reducing the gap between enhanced results and high-quality images from a high-level perspective. At the task level, we exploit the robustness-gap between strong foundation models and weak models, applying high-level task guidance to the low-light enhancement training process. Through the collaboration of multiple foundation models, the proposed framework shows better enhancement performance and adapts better to high-level tasks. Extensive experiments across various enhancement and application benchmarks demonstrate the qualitative and quantitative superiority of the proposed method over numerous state-of-the-art techniques.
Yuxuan Gu 0001, Haoxuan Wang 0004, Pengyang Ling, Zhixiang Wei, Huaian Chen, Yi Jin 0002, Enhong Chen
CVPR3
2025 Light-a-Video: Training-Free Video Relighting via Progressive Light Fusion
abstract
Recent advancements in image relighting models, driven by large-scale datasets and pre-trained diffusion models, have enabled the imposition of consistent lighting. However, video relighting still lags, primarily due to the excessive training costs and the scarcity of diverse, high-quality video relighting datasets. A simple application of image relighting models on a frame-by-frame basis leads to several issues: lighting source inconsistency and relighted appearance inconsistency, resulting in flickers in the generated videos. In this work, we propose Light-A-Video, a training-free approach to achieve temporally smooth video relighting. Adapted from image relighting models, Light-A-Video introduces two key techniques to enhance lighting consistency. First, we design a Consistent Light Attention (CLA) module, which enhances cross-frame interactions within the self-attention layers of the image relight model to stabilize the generation of the background lighting source. Second, leveraging the physical principle of light transport independence, we apply linear blending between the source video's appearance and the relighted appearance, using a Progressive Light Fusion (PLF) strategy to ensure smooth temporal transitions in illumination. Experiments show that Light-A-Video improves the temporal consistency of relighted video while maintaining the relighted image quality, ensuring coherent lighting transitions across frames. Project page: https://bujiazi.github.io/light-a-video.github.io/.
Jiazi Bu, Pengyang Ling, Pan Zhang 0001, Qidong Huang, Jinsong Li 0001, Xiaoyi Dong, Yuhang Zang, Yuhang Cao, Anyi Rao, Jiaqi Wang 0003, Li Niu 0002
ICCV3
2025 MotionClone: Training-Free Motion Cloning for Controllable Video Generation
abstract
Motion-based controllable video generation offers the potential for creating captivating visual content. Existing methods typically necessitate model training to encode particular motion cues or incorporate fine-tuning to inject certain motion patterns, resulting in limited flexibility and generalization. In this work, we propose MotionClone, a training-free framework that enables motion cloning from reference videos to versatile motion-controlled video generation, including text-to-video and image-to-video. Based on the observation that the dominant components in temporal-attention maps drive motion synthesis, while the rest mainly capture noisy or very subtle motions, MotionClone utilizes sparse temporal attention weights as motion representations for motion guidance, facilitating diverse motion transfer across varying scenarios. Meanwhile, MotionClone allows for the direct extraction of motion representation through a single denoising step, bypassing the cumbersome inversion processes and thus promoting both efficiency and flexibility. Extensive experiments demonstrate that MotionClone exhibits proficiency in both global camera motion and local object motion, with notable superiority in terms of motion fidelity, textual alignment, and temporal consistency.
Pengyang Ling, Jiazi Bu, Pan Zhang 0001, Xiaoyi Dong, Yuhang Zang, Huaian Chen, Jiaqi Wang 0003, Yi Jin 0002
ICLR1
2025 HiFlow: Training-free High-Resolution Image Generation with Flow-Aligned Guidance
abstract
Text-to-image (T2I) diffusion/flow models have drawn considerable attention recently due to their remarkable ability to deliver flexible visual creations. Still, high-resolution image synthesis presents formidable challenges due to the scarcity and complexity of high-resolution content. Recent approaches have investigated training-free strategies to enable high-resolution image synthesis with pre-trained models. However, these techniques often struggle with generating high-quality visuals and tend to exhibit artifacts or low-fidelity details, as they typically rely solely on the endpoint of the low-resolution sampling trajectory while neglecting intermediate states that are critical for preserving structure and synthesizing finer detail. To this end, we present HiFlow, a training-free and model-agnostic framework to unlock the resolution potential of pre-trained flow models. Specifically, HiFlow establishes a virtual reference flow within the high-resolution space that effectively captures the characteristics of low-resolution flow information, offering guidance for high-resolution generation through three key aspects: initialization alignment for low-frequency consistency, direction alignment for structure preservation, and acceleration alignment for detail fidelity. By leveraging such flow-aligned guidance, HiFlow substantially elevates the quality of high-resolution image synthesis of T2I models and demonstrates versatility across their personalized variants. Extensive experiments validate HiFlow's capability in achieving superior high-resolution image quality over state-of-the-art methods.
Jiazi Bu, Pengyang Ling, Pan Zhang 0001, Xiaoyi Dong, Yuhang Zang, Yuhang Cao, Dahua Lin, Jiaqi Wang 0003
NeurIPS2
2025 Towards Better & Faster Autoregressive Image Generation: From the Perspective of Entropy
abstract
In this work, we first revisit the sampling issues in current autoregressive (AR) image generation models and identify that image tokens, unlike text tokens, exhibit lower information density and non-uniform spatial distribution. Accordingly, we present an entropy-informed decoding strategy that facilitates higher autoregressive generation quality with faster synthesis speed. Specifically, the proposed method introduces two main innovations: 1) dynamic temperature control guided by spatial entropy of token distributions, enhancing the balance between content diversity, alignment accuracy, and structural coherence in both mask-based and scale-wise models, without extra computational overhead, and 2) entropy-aware acceptance rules in speculative decoding, achieving near-lossless generation at about 85% of the inference cost of conventional acceleration methods. Extensive experiments across multiple benchmarks using diverse AR image generation models demonstrate the effectiveness and generalizability of our approach in enhancing both generation quality and sampling speed.
Feng Zhao 0004, Pengyang Ling, Haibo Qiu, Zhixiang Wei, Hu Yu 0001, Jie Huang 0017, Zhixiong Zeng, Lin Ma 0002
NeurIPS3
2025 Prior-assisted unpaired image dehazing framework for enhanced visibility in real-world hazy scenarios
Pengyang Ling, Haoxuan Wang 0004, Huaian Chen, Yuxuan Gu 0001, Yi Jin 0002, Jinjin Zheng
Expert Syst. Appl.1
2025 Seed Optimization With Frozen Generator for Superior Zero-Shot Low-Light Image Enhancement
abstract
In this work, we observe that the generators, which are pre-trained on massive natural images, inherently hold the promising potential for superior low-light image enhancement against varying scenarios. Specifically, for the low-light image enhancement process of a single image, we introduce the pre-trained generators to restore the details and colors degraded by low-light conditions, thereby improving the visual effect. Taking one step further, we introduce a novel optimization strategy, which backpropagates the gradients to the input seeds rather than the parameters of the low-light image enhancement model, thus intactly retaining the generative knowledge learned from natural images and achieving faster convergence speed. Benefiting from the pre-trained knowledge and seed-optimization strategy, the low-light image enhancement model can significantly regularize the visibility and fidelity of the enhanced result, thus rapidly generating high-quality images without training on any low-light dataset. Extensive experiments on various benchmarks demonstrate the effectiveness of the proposed method, showing its potential advantages over numerous state-of-the-art methods both qualitatively and quantitatively.
Yuxuan Gu 0001, Yi Jin 0002, Ben Wang 0005, Zhixiang Wei, Xiaoxiao Ma 0006, Haoxuan Wang 0004, Pengyang Ling, Huaian Chen, Enhong Chen
IEEE Trans. Circuits Syst. Video Technol.7
2025 Masked Video Pretraining Advances Real-World Video Denoising
abstract
Learning-based video denoisers have attained state-of-the-art (SOTA) performances on public evaluation benchmarks. Nevertheless, they typically encounter significant performance drops when applied to unseen real-world data, owing to inherent data discrepancies. To address this problem, this work delves into the model pretraining techniques and proposes masked central frame modeling (MCFM), a new video pretraining approach that significantly improves the generalization ability of the denoiser. This proposal stems from a key observation: pretraining denoiser by reconstructing intact videos from the corrupted sequences, where the central frames are masked at a suitable probability, contributes to achieving superior performance on real-world data. Building upon MCFM, we introduce a robust video denoiser, named MVDenoiser, which is firstly pretrained on massive available ordinary videos for general video modeling, and then finetuned on costful real-world noisy/clean video pairs for noisy-to-clean mapping. Additionally, beyond the denoising model, we further establish a new paired real-world noisy video dataset (RNVD) to facilitate cross-dataset evaluation of generalization ability. Extensive experiments conducted across different datasets demonstrate that the proposed method achieves superior performance compared to existing methods. Code and dataset are available athttps://github.com/mxxx99/MVDenoiser.
Yi Jin 0002, Xiaoxiao Ma 0006, Rui Zhang 0120, Huaian Chen, Yuxuan Gu 0001, Pengyang Ling, Enhong Chen
IEEE Trans. Multim.6
2024 FreeDrag: Feature Dragging for Reliable Point-Based Image Editing
abstract
To serve the intricate and varied demands of image editing, precise and flexible manipulation in image content is indispensable. Recently, Drag-based editing methods have gained impressive performance. However, these methods predominantly center on point dragging, resulting in two noteworthy drawbacks, namely “miss tracking ”, where dif-ficulties arise in accurately tracking the predetermined han-dle points, and “ambiguous tracking”, where tracked points are potentially positioned in wrong regions that closely re-semble the handle points. To address the above issues, we propose FreeDrag, a feature dragging methodology designed to free the burden on point tracking. The Free-Drag incorporates two key designs, i.e., template feature via adaptive updating and line search with backtracking, the former improves the stability against drastic content change by elaborately controlling the feature updating scale after each dragging, while the latter alleviates the misguidance from similar points by actively restricting the search area in a line. These two technologies together contribute to a more stable semantic dragging with higher efficiency. Comprehensive experimental results substantiate that our approach significantly outperforms pre-existing methodologies, offering reliable point-based editing even in various complex scenarios.
Pengyang Ling, Lin Chen 0026, Pan Zhang 0001, Huaian Chen, Yi Jin 0002, Jinjin Zheng
CVPR1
2024 Stronger, Fewer, & Superior: Harnessing Vision Foundation Models for Domain Generalized Semantic Segmentation
abstract
In this paper, we first assess and harness various Vision Foundation Models (VFMs) in the context of Domain Generalized Semantic Segmentation (DGSS). Driven by the motivation that Leveraging Stronger pre-trained models and Fewer trainable parameters for Superior generalizability, we introduce a robust fine-tuning approach, namely “Rein”, to parameter-efficiently harness VFMs for DGSS. Built upon a set of trainable tokens, each linked to distinct instances, Rein precisely refines and forwards the feature maps from each layer to the next layer within the backbone. This process produces diverse refinements for different categories within a single image. With fewer trainable parameters, Rein efficiently fine-tunes VFMs for DGSS tasks, surprisingly surpassing full parameter fine-tuning. Extensive experiments across various settings demonstrate that Rein significantly outperforms state-of-the-art methods. Remarkably, with just an extra 1% of trainable parameters within the frozen backbone, Rein achieves a mIoU of 78.4% on the Cityscapes, without accessing any real urban-scene datasets. Code is available at https://github.com/w1oves/Rein.git.
Zhixiang Wei, Lin Chen 0026, Yi Jin 0002, Xiaoxiao Ma 0006, Pengyang Ling, Ben Wang 0005, Huaian Chen, Jinjin Zheng
CVPR6
2024 Masked Pre-training Enables Universal Zero-shot Denoiser
abstract
In this work, we observe that model trained on vast general images via masking strategy, has been naturally embedded with their distribution knowledge, thus spontaneously attains the underlying potential for strong image denoising. Based on this observation, we propose a novel zero-shot denoising paradigm, i.e., $\textbf{M}$asked $\textbf{P}$re-train then $\textbf{I}$terative fill ($\textbf{MPI}$). MPI first trains model via masking and then employs pre-trained weight for high-quality zero-shot image denoising on a single noisy image. Concretely, MPI comprises two key procedures: $\textbf{1) Masked Pre-training}$ involves training model to reconstruct massive natural images with random masking for generalizable representations, gathering the potential for valid zero-shot denoising on images with varying noise degradation and even in distinct image types. $\textbf{2) Iterative filling}$ exploits pre-trained knowledge for effective zero-shot denoising. It iteratively optimizes the image by leveraging pre-trained weights, focusing on alternate reconstruction of different image parts, and gradually assembles fully denoised image within limited number of iterations. Comprehensive experiments across various noisy scenarios underscore the notable advances of MPI over previous approaches with a marked reduction in inference time.
Xiaoxiao Ma 0006, Zhixiang Wei, Yi Jin 0002, Pengyang Ling, Ben Wang 0005, Junkang Dai, Huaian Chen
NeurIPS4
2024 All-in-One Hardware-Oriented Model Compression for Efficient Multi-Hardware Deployment
abstract
Structured pruning is an efficient compression technique that significantly reduces the inference latency and energy consumption of convolutional neural networks (CNNs) by eliminating redundant filters. However, existing works suffer from expensive algorithm costs in multi-hardware deployment scenarios involving several budgets across multiple hardware devices. To tackle this challenge, we propose a novel all-in-one hardware-oriented compression framework (AHC), which integrates structured pruning and data pruning to rapidly generate vast hardware-efficient models with ultra-low pruning and fine-tuning costs. Specifically, AHC develops a unified hardware-aware pruning (UHP), which rapidly generates numerous hardware-efficient models for several budgets across multiple hardware devices in once pruning process, thereby reducing pruning costs in multi-hardware deployment scenarios. Moreover, AHC proposes a progressive data pruning (PDP), which gradually removes samples that have a negligible impact on enhancing the predictive ability of pruned models, thereby accelerating the fine-tuning process with negligible performance loss. Extensive experiments demonstrate the superiority of the AHC over state-of-the-art (SOTA) structured pruning methods in terms of algorithm costs, latency, and accuracy. In particular, compared with SOTA hardware-oriented pruning method, AHC achieves comparable performances while reducing$5.3\times $pruning costs and$2.7\times $fine-tuning costs in multi-hardware deployment scenarios. Code is available athttps://github.com/HXuan-Wang/AHC.
Haoxuan Wang 0004, Pengyang Ling, Xin Fan 0005, Tao Tu 0006, Jinjin Zheng, Huaian Chen, Yi Jin 0002, Enhong Chen
IEEE Trans. Circuits Syst. Video Technol.2
2024 Collaborative Filter Pruning for Efficient Automatic Surface Defect Detection
abstract
Surface defect detection is a critical task in industrial production, and numerous methods have been proposed to achieve high detection accuracy. Although deep-learning-based approaches have achieved state-of-the-art (SOTA) performances, their vast computational cost and high memory footprint prevent their deployment in resource-constrained environments. To address this problem, we propose a collaborative filter pruning method for the defect detection model, which significantly reduces the number of required calculations and parameters while maintaining high performance, even in cases with tasks suffering from the class imbalance problem. Our method aims to obtain lightweight pruned models by removing unimportant filters according to their importance evaluated by both structural similarity and detail richness of corresponding feature maps. Moreover, to improve the performance of pruned models, we propose a knowledge-fused fine-tuning approach that fuses the knowledge derived from two teacher networks to look after both representation learning and classifier learning, alleviating the class imbalance problem. Experimental results on four public datasets demonstrate that the proposed approach performs favorably relative to the SOTA methods. In particular, the proposed method achieves 39× and 59× parameter compression for VGG-16 and ResNet-50, respectively, on the NEU-CLS dataset, with a very small detection accuracy loss (<0.2%).
Haoxuan Wang 0004, Xin Fan 0005, Pengyang Ling, Ben Wang 0005, Huaian Chen, Yi Jin 0002
IEEE Trans. Ind. Informatics3
2023 Disentangle then Parse: Night-time Semantic Segmentation with Illumination Disentanglement
abstract
Most prior semantic segmentation methods have been developed for day-time scenes, while typically underperforming in night-time scenes due to insufficient and complicated lighting conditions. In this work, we tackle this challenge by proposing a novel night-time semantic segmentation paradigm, i.e., disentangle then parse (DTP). DTP explicitly disentangles night-time images into light-invariant reflectance and light-specific illumination components and then recognizes semantics based on their adaptive fusion. Concretely, the proposed DTP comprises two key components: 1) Instead of processing lighting-entangled features as in prior works, our Semantic-Oriented Disentanglement (SOD) framework enables the extraction of reflectance component without being impeded by lighting, allowing the network to consistently recognize the semantics under cover of varying and complicated lighting conditions. 2) Based on the observation that the illumination component can serve as a cue for some semantically confused regions, we further introduce an Illumination-Aware Parser (IAParser) to explicitly learn the correlation between semantics and lighting, and aggregate the illumination features to yield more precise predictions. Extensive experiments on the night-time segmentation task with various settings demonstrate that DTP significantly outperforms state-of-the-art methods. Furthermore, with negligible additional parameters, DTP can be directly used to benefit existing day-time methods for night-time segmentation. Code and dataset are available at https://github.com/w1oves/DTP.git.
Zhixiang Wei, Lin Chen 0026, Tao Tu 0006, Pengyang Ling, Huaian Chen, Yi Jin 0002
ICCV4
2023 Single Image Dehazing Using Saturation Line Prior
abstract
Saturation information in hazy images is conducive to effective haze removal, However, existing saturation-based dehazing methods just focus on the saturation value of each pixel itself, while the higher-level distribution characteristic between pixels regarding saturation remains to be harnessed. In this paper, we observe that the pixels, which share the same surface reflectance coefficient in the local patches of haze-free images, exhibit a linear relationship between their saturation component and the reciprocal of their brightness component in the corresponding hazy images normalized by atmospheric light. Furthermore, the intercept of the line described by this linear relationship on the saturation axis is exactly the saturation value of these pixels in the haze-free images. Using this characteristic of saturation, termed saturation line prior (SLP), the transmission estimation is translated into the construction of saturation lines. Accordingly, a new dehazing framework using SLP is proposed, which employs the intrinsic relevance between pixels to achieve a reliable saturation line construction for transmission estimation. This approach can recover the fine details and attain realistic colors from hazy scenes, resulting in a remarkable visibility improvement. Extensive experiments in real-world and synthetic hazy images show that the proposed method performs favorably against state-of-the-art dehazing methods. Code is available on https://github.com/LPengYang/Saturation-Line-Prior.
Pengyang Ling, Huaian Chen, Xiao Tan 0004, Yi Jin 0002, Enhong Chen
IEEE Trans. Image Process.1