Yi Jin 0002

dblp:38/4674-2 · DBLP profile ↗
← Back
34ranked-venue papers
1as first author
30since 2021 · last 2026
0000-0001-8232-3863ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 21 · 1 first-author · 17 since 2021Artificial intelligence and machine learning · 15 · 14 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Learning Deep ISP for High-Speed Cameras: Achieving DSLR-Quality Imaging Under High Frame Rates
abstract
High-speed imaging, which captures the fleeting dynamics of moving objects at extreme frame rates, has become an indispensable tool across a wide range of scientific disciplines. Yet, the pursuit of high temporal resolution often comes at the cost of significant image degradation, due to the inherent limitations of imaging sensors and the extreme conditions of ultra-short exposure and massive data throughput. As a result, high-speed cameras often produce images marred by strong noise and severe color distortions. In this work, we propose a deep image signal processing (ISP) paradigm that enables high-speed cameras to maintain extremely high frame rates while achieving image quality comparable to that of digital single-lens reflex (DSLR) cameras. To this end, we make two key contributions: 1) constructing RHID, the first large-scale real-world high-speed imaging ISP dataset, comprising 282,912 RAW images captured by high-speed cameras and corresponding sRGB images captured by DSLRs, featuring complex degradations intrinsic to high-speed acquisition; and 2) proposing a misalignment-robust ISP learning framework (MisISP), equipped with a prior mapper-guided image alignment module (PMIA) and a spectrum-guided weakly-aligned image supervisory loss, which effectively addresses inherent pixel misalignments caused by heterogeneous sensor characteristics. Extensive experiments demonstrate that our paradigm substantially advances the performance of existing deep ISP models for high-speed imaging, achieving remarkable improvements in noise suppression, brightness enhancement, and color preservation.
Huaian Chen, Tao Tu 0006, Wenjun Wei, Yi Jin 0002, Enhong Chen
IEEE Trans. Pattern Anal. Mach. Intell.5
2026 Rein++: Efficient Generalization and Adaptation for Semantic Segmentation With Vision Foundation Models
abstract
Vision Foundation Models(VFMs) have achieved remarkable success in various computer vision tasks.However, their application to semantic segmentation is hindered by two significant challenges: (1) the disparity in data scale, as segmentation datasets are typically much smaller than those used for VFM pre-training, and (2) domain distribution shifts, where real-world segmentation scenarios are diverse and often underrepresented during pre-training. To overcome these limitations, we present Rein++, an efficient VFM-based segmentation framework that demonstrates superior generalization from limited data and enables effective adaptation to diverse unlabeled scenarios. Specifically, Rein++ comprises a domain generalization solution Rein-G and a domain adaptation solution Rein-A. Rein-G introduces a set of trainable, instance-aware tokens that effectively refine the VFM's features for the segmentation task. This parameter-efficient approach fine-tunes less than 1% of the backbone's parameters, enabling robust generalization. Building on the Rein-G, Rein-A performs unsupervised domain adaptation at both the instance and logit levels to mitigate domain shifts. In addition, it incorporates a semantic transfer module that leverages the class-agnostic capabilities of the segment anything model to enhance boundary details in the target domain. The integrated Rein++ pipeline first learns a generalizable model on a source domain (e.g., daytime scenes) and subsequently adapts it to diverse target domains (e.g., nighttime scenes) without any target labels. Comprehensive experiments demonstrate that Rein++ significantly outperforms state-of-the-art methods with efficient training, underscoring its roles an efficient, generalizable, and adaptive segmentation solution for VFMs, even for large models with billions of parameters.
Zhixiang Wei, Xiaoxiao Ma 0006, Ruishen Yan, Tao Tu 0006, Huaian Chen, Jinjin Zheng, Yi Jin 0002, Enhong Chen
IEEE Trans. Pattern Anal. Mach. Intell.7
2026 Fundus Image Enhancement With Pyramid Conditional Flow
abstract
Deep learning-based approaches, which learn pixel-to-pixel mapping from input to output images, have demonstrated exceptional performance in enhancing low-quality fundus images. However, due to the ambiguous definition of the ground-truth high-quality image, the pixel-to-pixel mapping encounters an ill-posed problem arising from the complex one-to-many relationship between low-quality fundus images and their corresponding high-quality versions. To address this problem, this work proposes a PCFlow, the first normalizing flow method that learns the complex distributions of high-quality fundus images rather than a pixel-to-pixel mapping. Unlike the existing image natural enhancement methods that aim to restore images with comfortable visual quality, PCFlow enhances fundus images by prioritizing clinically significant information. To this end, we design a condition module that utilizes retinal structure as a conditioning factor to constrain the optimization of PCFlow, and then build an invertible coupling layer that employs a pyramid structure for identifying each frequency component of retinal features. With the cooperation and interactions of these key components, the proposed PCFlow preserves the retinal structures and pathological characteristics essential for clinical applications. Extensive experiments on the real and synthetic fundus datasets demonstrate that our method achieves better performance.
Wenjun Wei, Huaian Chen, Yi Jin 0002
IEEE J. Biomed. Health Informatics5
2026 Efficient Haze Removal via Scene Depth Ordering for Robust Traffic Monitoring
abstract
The reliability of vision-based systems, such as traffic monitoring and intelligent driving, is typically compromised in hazy weather owing to diminished visibility. In this paper, we propose a novel efficient image dehazing framework guided by depth order, leveraging the consistency of depth perception to establish strong global constraints for enhanced haze removal. The consistent depth perception ensures that the regions that look farther or closer in hazy images also appear farther or closer in the corresponding dehazing results, substantially avoiding potential visual degradation. To this end, the depth order in hazy images is approximated by the reverse order of color difference between pixel values and global atmospheric light, offering an effective and efficient alternative for depth perception modeling. Subsequently, we have developed a depth order embedded transformation model to estimate the transmission maps jointly constrained by depth order and haze imaging model, ensuring that the depth order remains unchanged in corresponding dehazing results. This model harnesses the extracted depth order as a powerful global constraint for the dehazing process, facilitating the efficient use of global information and thus achieving superior image restoration. Extensive experiments demonstrate that the proposed method can better recover potential structure and vivid color with higher computational efficiency, offering an efficient solution for robust traffic monitoring against hazy weather.
Pengyang Ling, Huaian Chen, Haoxuan Wang 0004, Yuxuan Gu 0001, Yi Jin 0002, Jinjin Zheng, Enhong Chen
IEEE Trans. Intell. Transp. Syst.5
2025 Improving Visual and Downstream Performance of Low-Light Enhancer with Vision Foundation Models Collaboration
abstract
In this paper, we observe that the collaboration of various foundation models can perceive semantic and degraded information within images, thereby guiding the low-light enhancement process. Specifically, we propose a self-supervised low-light enhancement framework based on the multiple foundation models collaboration (dubbed FoCo), aimed at improving both the visual quality of enhanced images and the performance in high-level applications. At the feature level, FoCo leverages the rich features from various foundation models to enhance the model’s semantic perception during training, thereby reducing the gap between enhanced results and high-quality images from a high-level perspective. At the task level, we exploit the robustness-gap between strong foundation models and weak models, applying high-level task guidance to the low-light enhancement training process. Through the collaboration of multiple foundation models, the proposed framework shows better enhancement performance and adapts better to high-level tasks. Extensive experiments across various enhancement and application benchmarks demonstrate the qualitative and quantitative superiority of the proposed method over numerous state-of-the-art techniques.
Yuxuan Gu 0001, Haoxuan Wang 0004, Pengyang Ling, Zhixiang Wei, Huaian Chen, Yi Jin 0002, Enhong Chen
CVPR6
2025 Integral Fast Fourier Color Constancy
abstract
Traditional auto white balance (AWB) algorithms typically assume a single global illuminant source, which leads to color distortions in multi-illuminant scenes. While recent neural network-based methods have shown excellent accuracy in such scenarios, their high parameter count and computational demands limit their practicality for real-time video applications. The Fast Fourier Color Constancy (FFCC) algorithm was proposed for single-illuminant-source scenes, predicting a global illuminant source with high efficiency. However, it cannot be directly applied to multi-illuminant scenarios unless specifically modified. To address this, we propose Integral Fast Fourier Color Constancy (IFFCC), an extension of FFCC tailored for multi-illuminant scenes. IFFCC leverages the proposed integral UV histogram to accelerate histogram computations across all possible regions in Cartesian space and parallelizes Fourier-based convolution operations, resulting in a spatially-smooth illumination map. This approach enables high-accuracy, real-time AWB in multi-illuminant scenes. Extensive experiments show that IFFCC achieves accuracy that is on par with or surpasses that of pixel-level neural networks, while reducing the parameter count by over 400× and processing speed by 20 − 100× faster than network-based approaches.
Wenjun Wei, Yanlin Qian, Huaian Chen, Junkang Dai, Yi Jin 0002
CVPR5
2025 HQ-CLIP: Leveraging Large Vision-Language Models to Create High-Quality Image-Text Datasets and CLIP Models
Zhixiang Wei, Guangting Wang, Xiaoxiao Ma 0006, Ke Mei, Huaian Chen, Yi Jin 0002, Fengyun Rao
ICCV6
2025 MotionClone: Training-Free Motion Cloning for Controllable Video Generation
abstract
Motion-based controllable video generation offers the potential for creating captivating visual content. Existing methods typically necessitate model training to encode particular motion cues or incorporate fine-tuning to inject certain motion patterns, resulting in limited flexibility and generalization. In this work, we propose MotionClone, a training-free framework that enables motion cloning from reference videos to versatile motion-controlled video generation, including text-to-video and image-to-video. Based on the observation that the dominant components in temporal-attention maps drive motion synthesis, while the rest mainly capture noisy or very subtle motions, MotionClone utilizes sparse temporal attention weights as motion representations for motion guidance, facilitating diverse motion transfer across varying scenarios. Meanwhile, MotionClone allows for the direct extraction of motion representation through a single denoising step, bypassing the cumbersome inversion processes and thus promoting both efficiency and flexibility. Extensive experiments demonstrate that MotionClone exhibits proficiency in both global camera motion and local object motion, with notable superiority in terms of motion fidelity, textual alignment, and temporal consistency.
Pengyang Ling, Jiazi Bu, Pan Zhang 0001, Xiaoyi Dong, Yuhang Zang, Huaian Chen, Jiaqi Wang 0003, Yi Jin 0002
ICLR9
2025 Prior-assisted unpaired image dehazing framework for enhanced visibility in real-world hazy scenarios
Pengyang Ling, Haoxuan Wang 0004, Huaian Chen, Yuxuan Gu 0001, Yi Jin 0002, Jinjin Zheng
Expert Syst. Appl.5
2025 Seed Optimization With Frozen Generator for Superior Zero-Shot Low-Light Image Enhancement
abstract
In this work, we observe that the generators, which are pre-trained on massive natural images, inherently hold the promising potential for superior low-light image enhancement against varying scenarios. Specifically, for the low-light image enhancement process of a single image, we introduce the pre-trained generators to restore the details and colors degraded by low-light conditions, thereby improving the visual effect. Taking one step further, we introduce a novel optimization strategy, which backpropagates the gradients to the input seeds rather than the parameters of the low-light image enhancement model, thus intactly retaining the generative knowledge learned from natural images and achieving faster convergence speed. Benefiting from the pre-trained knowledge and seed-optimization strategy, the low-light image enhancement model can significantly regularize the visibility and fidelity of the enhanced result, thus rapidly generating high-quality images without training on any low-light dataset. Extensive experiments on various benchmarks demonstrate the effectiveness of the proposed method, showing its potential advantages over numerous state-of-the-art methods both qualitatively and quantitatively.
Yuxuan Gu 0001, Yi Jin 0002, Ben Wang 0005, Zhixiang Wei, Xiaoxiao Ma 0006, Haoxuan Wang 0004, Pengyang Ling, Huaian Chen, Enhong Chen
IEEE Trans. Circuits Syst. Video Technol.2
2025 Masked Video Pretraining Advances Real-World Video Denoising
abstract
Learning-based video denoisers have attained state-of-the-art (SOTA) performances on public evaluation benchmarks. Nevertheless, they typically encounter significant performance drops when applied to unseen real-world data, owing to inherent data discrepancies. To address this problem, this work delves into the model pretraining techniques and proposes masked central frame modeling (MCFM), a new video pretraining approach that significantly improves the generalization ability of the denoiser. This proposal stems from a key observation: pretraining denoiser by reconstructing intact videos from the corrupted sequences, where the central frames are masked at a suitable probability, contributes to achieving superior performance on real-world data. Building upon MCFM, we introduce a robust video denoiser, named MVDenoiser, which is firstly pretrained on massive available ordinary videos for general video modeling, and then finetuned on costful real-world noisy/clean video pairs for noisy-to-clean mapping. Additionally, beyond the denoising model, we further establish a new paired real-world noisy video dataset (RNVD) to facilitate cross-dataset evaluation of generalization ability. Extensive experiments conducted across different datasets demonstrate that the proposed method achieves superior performance compared to existing methods. Code and dataset are available athttps://github.com/mxxx99/MVDenoiser.
Yi Jin 0002, Xiaoxiao Ma 0006, Rui Zhang 0120, Huaian Chen, Yuxuan Gu 0001, Pengyang Ling, Enhong Chen
IEEE Trans. Multim.1
2025 Data and Prior-Driven Low-Light Enhancement Boosting the Visibility of Imaging Systems
Huaian Chen, Ben Wang 0005, Zhixiang Wei, Yi Jin 0002, Enhong Chen
IEEE Trans. Syst. Man Cybern. Syst.5
2024 FreeDrag: Feature Dragging for Reliable Point-Based Image Editing
abstract
To serve the intricate and varied demands of image editing, precise and flexible manipulation in image content is indispensable. Recently, Drag-based editing methods have gained impressive performance. However, these methods predominantly center on point dragging, resulting in two noteworthy drawbacks, namely “miss tracking ”, where dif-ficulties arise in accurately tracking the predetermined han-dle points, and “ambiguous tracking”, where tracked points are potentially positioned in wrong regions that closely re-semble the handle points. To address the above issues, we propose FreeDrag, a feature dragging methodology designed to free the burden on point tracking. The Free-Drag incorporates two key designs, i.e., template feature via adaptive updating and line search with backtracking, the former improves the stability against drastic content change by elaborately controlling the feature updating scale after each dragging, while the latter alleviates the misguidance from similar points by actively restricting the search area in a line. These two technologies together contribute to a more stable semantic dragging with higher efficiency. Comprehensive experimental results substantiate that our approach significantly outperforms pre-existing methodologies, offering reliable point-based editing even in various complex scenarios.
Pengyang Ling, Lin Chen 0026, Pan Zhang 0001, Huaian Chen, Yi Jin 0002, Jinjin Zheng
CVPR5
2024 Stronger, Fewer, & Superior: Harnessing Vision Foundation Models for Domain Generalized Semantic Segmentation
abstract
In this paper, we first assess and harness various Vision Foundation Models (VFMs) in the context of Domain Generalized Semantic Segmentation (DGSS). Driven by the motivation that Leveraging Stronger pre-trained models and Fewer trainable parameters for Superior generalizability, we introduce a robust fine-tuning approach, namely “Rein”, to parameter-efficiently harness VFMs for DGSS. Built upon a set of trainable tokens, each linked to distinct instances, Rein precisely refines and forwards the feature maps from each layer to the next layer within the backbone. This process produces diverse refinements for different categories within a single image. With fewer trainable parameters, Rein efficiently fine-tunes VFMs for DGSS tasks, surprisingly surpassing full parameter fine-tuning. Extensive experiments across various settings demonstrate that Rein significantly outperforms state-of-the-art methods. Remarkably, with just an extra 1% of trainable parameters within the frozen backbone, Rein achieves a mIoU of 78.4% on the Cityscapes, without accessing any real urban-scene datasets. Code is available at https://github.com/w1oves/Rein.git.
Zhixiang Wei, Lin Chen 0026, Yi Jin 0002, Xiaoxiao Ma 0006, Pengyang Ling, Ben Wang 0005, Huaian Chen, Jinjin Zheng
CVPR3
2024 Masked Pre-training Enables Universal Zero-shot Denoiser
abstract
In this work, we observe that model trained on vast general images via masking strategy, has been naturally embedded with their distribution knowledge, thus spontaneously attains the underlying potential for strong image denoising. Based on this observation, we propose a novel zero-shot denoising paradigm, i.e., $\textbf{M}$asked $\textbf{P}$re-train then $\textbf{I}$terative fill ($\textbf{MPI}$). MPI first trains model via masking and then employs pre-trained weight for high-quality zero-shot image denoising on a single noisy image. Concretely, MPI comprises two key procedures: $\textbf{1) Masked Pre-training}$ involves training model to reconstruct massive natural images with random masking for generalizable representations, gathering the potential for valid zero-shot denoising on images with varying noise degradation and even in distinct image types. $\textbf{2) Iterative filling}$ exploits pre-trained knowledge for effective zero-shot denoising. It iteratively optimizes the image by leveraging pre-trained weights, focusing on alternate reconstruction of different image parts, and gradually assembles fully denoised image within limited number of iterations. Comprehensive experiments across various noisy scenarios underscore the notable advances of MPI over previous approaches with a marked reduction in inference time.
Xiaoxiao Ma 0006, Zhixiang Wei, Yi Jin 0002, Pengyang Ling, Ben Wang 0005, Junkang Dai, Huaian Chen
NeurIPS3
2024 All-in-One Hardware-Oriented Model Compression for Efficient Multi-Hardware Deployment
abstract
Structured pruning is an efficient compression technique that significantly reduces the inference latency and energy consumption of convolutional neural networks (CNNs) by eliminating redundant filters. However, existing works suffer from expensive algorithm costs in multi-hardware deployment scenarios involving several budgets across multiple hardware devices. To tackle this challenge, we propose a novel all-in-one hardware-oriented compression framework (AHC), which integrates structured pruning and data pruning to rapidly generate vast hardware-efficient models with ultra-low pruning and fine-tuning costs. Specifically, AHC develops a unified hardware-aware pruning (UHP), which rapidly generates numerous hardware-efficient models for several budgets across multiple hardware devices in once pruning process, thereby reducing pruning costs in multi-hardware deployment scenarios. Moreover, AHC proposes a progressive data pruning (PDP), which gradually removes samples that have a negligible impact on enhancing the predictive ability of pruned models, thereby accelerating the fine-tuning process with negligible performance loss. Extensive experiments demonstrate the superiority of the AHC over state-of-the-art (SOTA) structured pruning methods in terms of algorithm costs, latency, and accuracy. In particular, compared with SOTA hardware-oriented pruning method, AHC achieves comparable performances while reducing$5.3\times $pruning costs and$2.7\times $fine-tuning costs in multi-hardware deployment scenarios. Code is available athttps://github.com/HXuan-Wang/AHC.
Haoxuan Wang 0004, Pengyang Ling, Xin Fan 0005, Tao Tu 0006, Jinjin Zheng, Huaian Chen, Yi Jin 0002, Enhong Chen
IEEE Trans. Circuits Syst. Video Technol.7
2024 Collaborative Filter Pruning for Efficient Automatic Surface Defect Detection
abstract
Surface defect detection is a critical task in industrial production, and numerous methods have been proposed to achieve high detection accuracy. Although deep-learning-based approaches have achieved state-of-the-art (SOTA) performances, their vast computational cost and high memory footprint prevent their deployment in resource-constrained environments. To address this problem, we propose a collaborative filter pruning method for the defect detection model, which significantly reduces the number of required calculations and parameters while maintaining high performance, even in cases with tasks suffering from the class imbalance problem. Our method aims to obtain lightweight pruned models by removing unimportant filters according to their importance evaluated by both structural similarity and detail richness of corresponding feature maps. Moreover, to improve the performance of pruned models, we propose a knowledge-fused fine-tuning approach that fuses the knowledge derived from two teacher networks to look after both representation learning and classifier learning, alleviating the class imbalance problem. Experimental results on four public datasets demonstrate that the proposed approach performs favorably relative to the SOTA methods. In particular, the proposed method achieves 39× and 59× parameter compression for VGG-16 and ResNet-50, respectively, on the NEU-CLS dataset, with a very small detection accuracy loss (<0.2%).
Haoxuan Wang 0004, Xin Fan 0005, Pengyang Ling, Ben Wang 0005, Huaian Chen, Yi Jin 0002
IEEE Trans. Ind. Informatics6
2023 Disentangle then Parse: Night-time Semantic Segmentation with Illumination Disentanglement
abstract
Most prior semantic segmentation methods have been developed for day-time scenes, while typically underperforming in night-time scenes due to insufficient and complicated lighting conditions. In this work, we tackle this challenge by proposing a novel night-time semantic segmentation paradigm, i.e., disentangle then parse (DTP). DTP explicitly disentangles night-time images into light-invariant reflectance and light-specific illumination components and then recognizes semantics based on their adaptive fusion. Concretely, the proposed DTP comprises two key components: 1) Instead of processing lighting-entangled features as in prior works, our Semantic-Oriented Disentanglement (SOD) framework enables the extraction of reflectance component without being impeded by lighting, allowing the network to consistently recognize the semantics under cover of varying and complicated lighting conditions. 2) Based on the observation that the illumination component can serve as a cue for some semantically confused regions, we further introduce an Illumination-Aware Parser (IAParser) to explicitly learn the correlation between semantics and lighting, and aggregate the illumination features to yield more precise predictions. Extensive experiments on the night-time segmentation task with various settings demonstrate that DTP significantly outperforms state-of-the-art methods. Furthermore, with negligible additional parameters, DTP can be directly used to benefit existing day-time methods for night-time segmentation. Code and dataset are available at https://github.com/w1oves/DTP.git.
Zhixiang Wei, Lin Chen 0026, Tao Tu 0006, Pengyang Ling, Huaian Chen, Yi Jin 0002
ICCV6
2023 BRAS: Bidirectional Reflectance Adjustment Strategy for 3-D Reconstruction of Mirror-Like Surface
abstract
A mirror-like surface (MLS) reflects highlight, aggravating the image saturation in structured light 3-D reconstruction systems and precluding defect detection based on reconstructed 3-D profiles. Previous studies have focused on a strategy of limiting the luminous flux entering the camera. However, the high intensity of the reflected highlight forces the limitation to be strengthened, which heavily reduces the modulation in the images for the 3-D reconstruction. Therefore, we propose a new strategy to adjust the source of the reflected highlight, i.e., the bidirectional reflectance (BR) of the MLS, which fundamentally suppresses the highlight and removes the limitation on the luminous flux. To execute the proposed strategy, an unfixed view structured light system (UVSLS) is established. The UVSLS converts the reflection viewer from the real camera to the virtual camera, realizing the flexible adjustment of the MLS BR. Finally, an MLS 3-D reconstruction framework is constructed to obtain the 3-D profile of the MLS. Experiments demonstrate that the proposed framework reduces the saturated pixels by 73.72% compared with the conventional method. Compared with the previous methods, the saturated pixels are reduced by an average of 47.62%.
Ben Wang 0005, Yabing Zheng, Minghui Duan, Xin Fan 0005, Yi Jin 0002, Jinjin Zheng
IEEE Trans. Ind. Informatics6
2023 Single Image Dehazing Using Saturation Line Prior
abstract
Saturation information in hazy images is conducive to effective haze removal, However, existing saturation-based dehazing methods just focus on the saturation value of each pixel itself, while the higher-level distribution characteristic between pixels regarding saturation remains to be harnessed. In this paper, we observe that the pixels, which share the same surface reflectance coefficient in the local patches of haze-free images, exhibit a linear relationship between their saturation component and the reciprocal of their brightness component in the corresponding hazy images normalized by atmospheric light. Furthermore, the intercept of the line described by this linear relationship on the saturation axis is exactly the saturation value of these pixels in the haze-free images. Using this characteristic of saturation, termed saturation line prior (SLP), the transmission estimation is translated into the construction of saturation lines. Accordingly, a new dehazing framework using SLP is proposed, which employs the intrinsic relevance between pixels to achieve a reliable saturation line construction for transmission estimation. This approach can recover the fine details and attain realistic colors from hazy scenes, resulting in a remarkable visibility improvement. Extensive experiments in real-world and synthetic hazy images show that the proposed method performs favorably against state-of-the-art dehazing methods. Code is available on https://github.com/LPengYang/Saturation-Line-Prior.
Pengyang Ling, Huaian Chen, Xiao Tan 0004, Yi Jin 0002, Enhong Chen
IEEE Trans. Image Process.4
2023 Deep Multi-Exposure Image Fusion for Dynamic Scenes
abstract
Recently, learning-based multi-exposure fusion (MEF) methods have made significant improvements. However, these methods mainly focus on static scenes and are prone to generate ghosting artifacts when tackling a more common scenario, i.e., the input images include motion, due to the lack of a benchmark dataset and solution for dynamic scenes. In this paper, we fill this gap by creating an MEF dataset of dynamic scenes, which contains multi-exposure image sequences and their corresponding high-quality reference images. To construct such a dataset, we propose a 'static-for-dynamic' strategy to obtain multi-exposure sequences with motions and their corresponding reference images. To the best of our knowledge, this is the first MEF dataset of dynamic scenes. Correspondingly, we propose a deep dynamic MEF (DDMEF) framework to reconstruct a ghost-free high-quality image from only two differently exposed images of a dynamic scene. DDMEF is achieved through two steps: pre-enhancement-based alignment and privilege-information-guided fusion. The former pre-enhances the input images before alignment, which helps to address the misalignments caused by the significant exposure difference. The latter introduces a privilege distillation scheme with an information attention transfer loss, which effectively improves the deghosting ability of the fusion network. Extensive qualitative and quantitative experimental results show that the proposed method outperforms state-of-the-art dynamic MEF methods. The source code and dataset are released at https://github.com/Tx000/Deep_dynamicMEF.
Xiao Tan 0004, Huaian Chen, Rui Zhang 0120, Yan Kan, Jinjin Zheng, Yi Jin 0002, Enhong Chen
IEEE Trans. Image Process.7
2023 Video Denoising for Scenes With Challenging Motion: A Comprehensive Analysis and a New Framework
abstract
Challenging motion, which tends to cause artifacts, is a key problem in the video denoising task. Recent video denoising methods have attempted to address this problem. However, they usually provide general performance evaluation on the overall dataset and cannot provide a comprehensive analysis for the influence of different motion levels. Thus, we questioned whether these methods can effectively deal with different scene motions. To this end, we synthesize a dataset containing videos with different motion levels and capture a new dataset that consists of videos involving large-scale motion. Then, we provide a comprehensive analysis on the elaborately collected datasets and find that, as the motion level increases, the performance of the denoising models based on implicit motion estimation (IME) declines sharply, while explicit motion estimation (EME) contributes to a more robust denoising quality. Therefore, in this work, we present an EME-embedded progressive denoising framework that fully considers the relationship between the noise removal and motion estimation. Specifically, we decouple video denoising into spatial denoising, EME-based frame reconstruction, and temporal refining processes. Spatial denoising improves the accuracy of EME process in the case of videos suffering from heavy noise, while the temporal refining process refines the denoised frame by utilizing temporal redundancy of the reconstructed motion-free frames. Extensive experiments demonstrate that the proposed method outperforms existing state-of-the-art methods, especially for videos containing large-scale motion.
Huaian Chen, Minghui Duan, Yi Jin 0002, Yan Kan, Changan Zhu
IEEE Trans. Multim.4
2023 Deep SR-HDR: Joint Learning of Super-Resolution and High Dynamic Range Imaging for Dynamic Scenes
abstract
The visual quality of a single image captured by a digital camera usually suffers from limited spatial resolution and low dynamic range (LDR) due to sensor constraints. To address these problems, recent works have independently applied convolutional neural networks (CNNs) to super-resolution (SR) and high dynamic range (HDR) imaging and made significant improvements in visual quality. However, directly connecting SR and HDR networks is an inefficient way to enhance image quality, because these two tasks share most of the same processing steps. To this end, we propose a deep neural network for the joint task of SR and HDR imaging, termed Deep SR-HDR, which reconstructs a high-resolution (HR) HDR image from a set of differently exposed low-resolution (LR) LDR images of a dynamic scene. Specifically, we merge the shared processing steps, including feature extraction and alignment of these two tasks. In particular, to handle large-scale complex motions, we design a multi-scale deformable module (MSDM) that estimates the sampling location offsets in a coarse-to-fine manner and then flexibly integrates useful information to compensate for the missing content in the motion regions. Then, we divide the fusion stage into two branches for HDR generation and high-frequency information extraction. With the cooperation and interactions of these modules, the proposed network reconstructs high-quality HR HDR images. Extensive qualitative and quantitative experimental results demonstrate the superiority and high efficiency of the proposed network.
Xiao Tan 0004, Huaian Chen, Yi Jin 0002, Changan Zhu
IEEE Trans. Multim.4
2022 Reusing the Task-specific Classifier as a Discriminator: Discriminator-free Adversarial Domain Adaptation
abstract
Adversarial learning has achieved remarkable performances for unsupervised domain adaptation (UDA). Existing adversarial UDA methods typically adopt an additional discriminator to play the min-max game with a feature extractor. However, most of these methods failed to effectively leverage the predicted discriminative information, and thus cause mode collapse for generator. In this work, we address this problem from a different perspective and design a simple yet effective adversarial paradigm in the form of a discriminator-free adversarial learning network (DALN), wherein the category classifier is reused as a discriminator, which achieves explicit domain alignment and category distinguishment through a unified objective, enabling the DALN to leverage the predicted discriminative information for sufficient feature alignment. Basically, we introduce a Nuclear-norm Wasserstein discrepancy (NWD) that has definite guidance meaning for performing discrimination. Such NWD can be coupled with the classifier to serve as a discriminator satisfying the K-Lipschitz constraint without the requirements of additional weight clipping or gradient penalty strategy. Without bells and whistles, DALN compares favorably against the existing state-of-the-art (SOTA) methods on a variety of public datasets. Moreover, as a plug-and-play technique, NWD can be directly used as a generic regularizer to benefit existing UDA algorithms. Code is available at https://github.com/xiaoachen98/DALN.
Lin Chen 0026, Huaian Chen, Zhixiang Wei, Xin Jin 0014, Xiao Tan 0004, Yi Jin 0002, Enhong Chen
CVPR6
2022 Deliberated Domain Bridging for Domain Adaptive Semantic Segmentation
abstract
In unsupervised domain adaptation (UDA), directly adapting from the source to the target domain usually suffers significant discrepancies and leads to insufficient alignment. Thus, many UDA works attempt to vanish the domain gap gradually and softly via various intermediate spaces, dubbed domain bridging (DB). However, for dense prediction tasks such as domain adaptive semantic segmentation (DASS), existing solutions have mostly relied on rough style transfer and how to elegantly bridge domains is still under-explored. In this work, we resort to data mixing to establish a deliberated domain bridging (DDB) for DASS, through which the joint distributions of source and target domains are aligned and interacted with each in the intermediate space. At the heart of DDB lies a dual-path domain bridging step for generating two intermediate domains using the coarse-wise and the fine-wise data mixing techniques, alongside a cross-path knowledge distillation step for taking two complementary models trained on generated intermediate samples as ‘teachers’ to develop a superior ‘student’ in a multi-teacher distillation manner. These two optimization steps work in an alternating way and reinforce each other to give rise to DDB with strong adaptation power. Extensive experiments on adaptive segmentation tasks with different settings demonstrate that our DDB significantly outperforms state-of-the-art methods.
Lin Chen 0026, Zhixiang Wei, Xin Jin 0014, Huaian Chen, Miao Zheng, Kai Chen 0026, Yi Jin 0002
NeurIPS7
2022 Structure-Texture Aware Network for Low-Light Image Enhancement
abstract
Global structure and local detailed texture have different effects on image enhancement tasks. However, most existing works treated these two components in the same way, without fully considering the characteristics of the global structure and local detailed texture. In this work, we propose a structure-texture aware network (STANet) that successfully exploits structure and texture features of low-light images to improve perceptual quality. To construct STANet, a fine-scale contour map guided filter is introduced to decompose the image into a structure component and a texture component. Then, structure-attention and texture-attention subnetworks are designed to fully exploit the characteristics of these two components. Finally, a fusion subnetwork with attention mechanisms is utilized to explore the internal correlations among the global and local features. Furthermore, to optimize the proposed STANet model, we propose a hybrid loss function; specifically, a color loss function is introduced to alleviate color distortion in the enhanced image. Extensive experiments demonstrate that the proposed method improves the visual quality of images; moreover, STANet outperforms most other state-of-the-art approaches.
Huaian Chen, Yi Jin 0002, Changan Zhu
IEEE Trans. Circuits Syst. Video Technol.4
2022 Multiframe-to-Multiframe Network for Video Denoising
abstract
Most existing studies performed video denoising by using multiple adjacent noisy frames to recover one clean frame; however, despite achieving relatively good quality for each individual frame, these approaches may result in visual flickering when the denoised frames are considered in sequence. In this paper, instead of separately restoring each clean frame, we propose a multiframe-to-multiframe (MM) denoising scheme that simultaneously recovers multiple clean frames from consecutive noisy frames. The proposed MM denoising scheme uses a training strategy that optimizes the denoised video from both the spatial and temporal dimensions, enabling better temporal consistency in the denoised video. Furthermore, we present an MM network (MMNet), which adopts a spatiotemporal convolutional architecture that considers both the interframe similarity and single-frame characteristics. Benefiting from the underlying parallel mechanism of the MM denoising scheme, MMNet achieves a highly competitive denoising efficiency. Extensive analyses and experiments demonstrate that MMNet outperforms the state-of-the-art video denoising methods, yielding temporal consistency improvements of at least 13.3$\%$and running more than 2 times faster than the other methods.
Huaian Chen, Yi Jin 0002, Changan Zhu
IEEE Trans. Multim.2
2022 Semisupervised Semantic Segmentation by Improving Prediction Confidence
abstract
Most of the recent image segmentation methods have tried to achieve the utmost segmentation results using large-scale pixel-level annotated data sets. However, obtaining these pixel-level annotated training data is usually tedious and expensive. In this work, we address the task of semisupervised semantic segmentation, which reduces the need for large numbers of pixel-level annotated images. We propose a method for semisupervised semantic segmentation by improving the confidence of the predicted class probability map via two parts. First, we build an adversarial framework that regards the segmentation network as the generator and uses a fully convolutional network as the discriminator. The adversarial learning makes the prediction class probability closer to 1. Second, the information entropy of the predicted class probability map is computed to represent the unpredictability of the segmentation prediction. Then, we infer the label-error map of the segmentation prediction and minimize the uncertainty on misclassified regions for unlabeled images. In contrast to existing semisupervised and weakly supervised semantic segmentation methods, the proposed method results in more confident predictions by focusing on the misclassified regions, especially the boundary regions. Our experimental results on the PASCAL VOC 2012 and PASCAL-CONTEXT data sets show that the proposed method achieves competitive segmentation performance.
Huaian Chen, Yi Jin 0002, Guoqiang Jin, Changan Zhu, Enhong Chen
IEEE Trans. Neural Networks Learn. Syst.2
2021 Single Image Specular Highlight Removal on Natural Scenes
Huaian Chen, Chenggang Hou, Minghui Duan, Xiao Tan 0004, Yi Jin 0002, Panlang Lv, Shaoqian Qin
PRCV (3)5
2021 DOF: A Demand-Oriented Framework for Image Denoising
abstract
Most existing image denoising methods focus on improving denoising quality. However, when applying denoising methods to practical tasks, in addition to the denoising quality, the number of parameters, and the computational complexity should be fully considered. In this article, we propose a demand-oriented framework (DOF) for image denoising, which can give preference to the number of parameters, the computational complexity, and the denoising quality or balance these three performance metrics. To perform the demand-oriented denoising, we first design a scale encoder to help the denoising model extract fewer but more representative features. Then, the split-flow module is introduced to fully exploit the input features by sharing the information of one network branch with other network branches. Finally, the scale decoder is utilized to reconstruct the final noise map without using any parameters. Through extensive experiments, we demonstrate that the proposed framework can be applied to several existing methods to help them achieve a more competitive denoising performance in terms of the number of parameters, and computational complexity.
Huaian Chen, Yi Jin 0002, Minghui Duan, Changan Zhu, Enhong Chen
IEEE Trans. Ind. Informatics2
2020 A multiscale dilated residual network for image denoising
Huaian Chen, Guoqiang Jin, Yi Jin 0002, Changan Zhu, Enhong Chen
Multim. Tools Appl.4
2020 Facial expression recognition with convolutional neural networks via a new face cropping and rotation strategy
Kuan Li, Yi Jin 0002, Muhammad Waqar Akram, Rui Han 0001, Jiongwei Chen
Vis. Comput.2
2019 Novel multi-convolutional neural network fusion approach for smile recognition
Jiongwei Chen, Yi Jin 0002, Muhammad Waqar Akram, Kuan Li, Enhong Chen
Multim. Tools Appl.2
2010 Efficient control system development using real-time virtual hardware-in-the-loop simulation
abstract
In this paper, an efficient control system development method based on real-time virtual hardware-in-the-loop simulation is presented. Instead of developing and testing control system after physical prototype is produced, it can be conducted immediately after concept of controlled system is designed by developing a pure software real-time simulator. In this simulator, all the hardware is virtual. The appearance and operation of these virtual hardware is programmed to imitate corresponding physical prototypes so as to implement real-time virtual hardware-in-the-loop simulation. So The control software and the control logic of control system can be developed quickly. And the source code is compatible with both the simulation platform and physical platform. Moreover, It's much easier and faster to debug and test the control logic in such a way. An example of a large and high-groove density ruling engine is provided to explain detail implementation of the proposed method.
Yi Jin 0002, Guofu Lian, Guoliang Ding, Changan Zhu
ICARCV2