VLDB 2026 Research / reviewers in the wild / expert
Wenhan Yang
dblp:156/2359
· DBLP profile ↗
206ranked-venue papers
29as first author
137since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 148 · 20 first-author · 87 since 2021Artificial intelligence and machine learning · 92 · 14 first-author · 77 since 2021Systems, architecture and hardware · 9 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Computer networks · 3 · 3 since 2021Security and privacy · 3 · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Event-Guided Super-Resolving Blurry Image via Asymmetric Integral Driven ConsistencyabstractSuper-Resolution from a Blurry low-resolution image (SRB) constitutes a severely ill-posed inverse problem. Current learning-based SRB approaches primarily rely on synthetic, well-labeled paired datasets to regularize solution spaces, yet they exhibit limited generalizability in practical applications due to significant domain discrepancies between simulated degradations and real-world imaging conditions. To bridge this synthetic-to-real gap, we propose a novel Self-supervised Event-based SRB (SE-SRB) framework that leverages neuromorphic event streams as physical priors and adopts a lightweight neural architecture tailored for effective domain adaptation. Specifically, the proposed SE-SRB introduces a self-supervised learning paradigm based on asymmetric integral driven consistency, which enforces temporal coherence between predictions derived from RGB and asynchronous event streams at different time points. Extensive experiments validate that SE-SRB consistently outperforms state-of-the-art methods on both synthetic and real-world datasets. Built upon a lightweight parallel two-stream architecture, SE-SRB achieves high computational efficiency, featuring reduced parameter count, lower FLOPs, and real-time inference capability (40 FPS). Chi Zhang 0027, Xiang Zhang 0022, Lei Yu 0006, Gui-Song Xia, Yuming Fang 0001, Wenhan Yang |
AAAI | 6 |
| 2026 | Malice Hides in Equivalence: Attacking Text-Image Alignment Assessment with Subtle Text Variations
Kang Xiao, Xuelin Shen, Baoliang Chen, Wenhan Yang, Meng Wang 0001 |
ISCAS | 4 |
| 2026 | Contrastive Mean Teacher for Robust Low-Light Image Enhancement
Zhangkai Ni, Menglin Han, Wenhan Yang, Hanli Wang, Lin Ma 0002, Sam Kwong |
Int. J. Comput. Vis. | 3 |
| 2026 | Towards Generalized Image Coding for Machine Through Meta Adversarial Adaptation
Xuelin Shen, Kangsheng Yin, Wenhan Yang |
Int. J. Comput. Vis. | 3 |
| 2026 | QuadPrior++: Multi-Dimension Augmented Physical Prior for Zero-Reference Illumination EnhancementabstractExisting low-light enhancement methods typically rely on fitting data mappings (pixel-wise mappings through fully supervised methods or distribution-wise mappings through weakly supervised or self-supervised methods). However, their performance is heavily dependent on specific scenes and fails to adequately model the intrinsic prior of natural images, resulting in poor generalization. To tackle this challenge, we leverage the strengths of powerful generative diffusion models, conditioned on a thoughtfully designed prior, and propose a novel zero-reference low-light enhancement framework that gets rid of dependence on the distribution of low-light images. In detail, we address the most fundamental core by proposing an illumination-invariant prior derived from the theory of physical light transfer, bridging the gap between normal and low-light domains, and enabling zero-shot enhancement without the need for low-light-specific training. A prior-to-image restoration framework is built upon generative diffusion models, pre-trained on normal-light data. During inference, the framework extracts the illumination-invariant prior from low-light inputs and maps them back to high-quality images, naturally for low-light enhancement. Additionally, such intrinsic properties of illumination-invariant prior open up opportunities for distilling diffusion models into compact CNN-based networks. We propose a novel prior-injected distillation paradigm incorporating intensity, frequency, and gradient domain-augmented regularization comprehensively. This distillation framework not only reduces computational costs but also maintains high fidelity and perceptual quality in enhanced outputs, making it more efficient and practical for real-world applications. The approach further extends seamlessly to handle over-exposure scenarios, demonstrating its versatility in addressing complex lighting conditions. Extensive experiments demonstrate the superiority of our framework in various scenarios, as well as its strong interpretability, robustness, and efficiency. Haofeng Huang, Wenjing Wang 0001, Wenhan Yang, Ling-Yu Duan, Jiaying Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | CTNet: Color transformation network for low-light image enhancement
Lidong Xie, Runmin Cong, Ju Dai, Wenhan Yang, JunJun Pan |
Pattern Recognit. | 4 |
| 2026 | Syntax-Driven Multi-Realism Image Compression With Consistency Guided Diffusion ModelabstractGiven the challenge of balancing high fidelity with perceptual quality, multi-realism image compression is developed to adapt flexibly to varying requirements. It allows images with different levels of realism to be decoded from the same bit stream. Diffusion models are known for generating images with high perceptual quality. However, their inherent process of adding noise and denoising is often difficult to control and will bring more distortion. This limits their direct application in image compression, especially in multi-realism image compression which requires precise control to adapt to different requirements. To address this issue, we propose aConsistency Guided Diffusion Modelas a post-processing network for multi-realism image compression, aiming to control the addition of detail representations, thereby adjusting the trade-off between subjective quality and fidelity. In detail, our proposed novel method is crafted to introduce an additional consistency guided feature branch into the diffusion model to constrain the deviation caused by randomness in the diffusion process to ensure fidelity. Furthermore, a syntax-driven feature fusion module is constructed to guide the information adaptive fusion of two branches with an input extra ultra-low stream, which contains the context information and trade-off control information. In addition, we design a warm-up based training strategy and adopt a continuous online optimization method to improve coding efficiency and trade-off control precision. Extensive experiments validate the superiority of our method over existing compression techniques, as well as the effectiveness of each component. Haowei Kuang, Wenhan Yang, Zongming Guo, Jiaying Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | Low-Light Image Enhancement via Diffusion Models With Semantic Priors of Any Region
Lingyu Zhu 0006, Wenhan Yang, Howard Leung, Shiqi Wang 0001, Sam Kwong |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | Prompting Rain Off: Evolving Compact Dual Prompts for Continual De-RainingabstractIn recent years, there has been notable progress in single-image rain removal, particularly focusing on static data distributions in these approaches. When dealing with data that constantly changes, the challenge of catastrophic forgetting arises, which is quite common and critical in real-world scenarios. To address this, we propose Evolving COmpact Dual Prompt Learning (EcoDPL), an efficient rehearsal-free continual learning deraining framework designed specifically for low-level vision tasks. Specifically, we design two prompt pools at both image and feature levels and insert these prompts into images and embedding tokens, for better knowledge transfer across tasks. Our adaptive weight generation module, P-Fuser, attaches an attention map to each prompt, to adaptively pay attention to different inputs, and get different weights to fuse prompts, making the inserted prompts more flexible with various inputs. Also, we introduce Grad-Tuner, a dictionary learning strategy, to compress knowledge into fewer prompts. This makes the knowledge more compact and provides more space for new prompts to learn new tasks. Our method stands out by leveraging small, learnable prompts for efficient knowledge retention across tasks, not increasing training time or parameters. Furthermore, we present an augmented method that upgrades the distance function $\gamma $ from simple cosine distance to a more advanced weight generation network. We also employ a fine-tuned dictionary learning technique, compressing knowledge into a more compact form, and enhancing the ability of prompts to learn new tasks. With our new designs, the model becomes more flexible with various inputs and it compresses knowledge into fewer prompts to free up spaces to learn new tasks. Through extensive experiments on various rain removal datasets, our EcoDPL method consistently outperforms previous continual learning techniques. Notably, although EcoDPL is designed for continual learning with changing data, it also performs well with stationary data, proving its robustness and versatility. Our website is available at: https://starymoon.github.io/Prompting-Rain-Off. Minghao Liu 0019, Wenhan Yang, Jiaying Liu 0001 |
IEEE Trans. Image Process. | 2 |
| 2026 | Open-Set Anomaly Segmentation in Complex ScenariosabstractPrecise segmentation of out-of-distribution (OoD) objects, herein referred to as anomalies, is crucial for the reliable deployment of semantic segmentation models in open-set, safety-critical applications, such as autonomous driving. Current anomalous segmentation benchmarks predominantly focus on favorable weather conditions, resulting in untrustworthy evaluations that overlook the risks posed by diverse meteorological conditions in open-set environments, such as low illumination, dense fog, and heavy rain. To bridge this gap, this paper introduces the ComSAmy, a Complex Scenarios Anomaly segmentation benchmark. ComSAmy encompasses a wide spectrum of adverse weather conditions, dynamic driving environments, and diverse anomaly types to comprehensively evaluate the model performance in realistic open-world scenarios. Our extensive evaluation of several state-of-the-art anomalous segmentation models reveals that existing methods demonstrate significant deficiencies in such challenging scenarios, highlighting their serious safety risks for real-world deployment. To solve that, we propose a novel energy-entropy learning (EEL) strategy that integrates the complementary information from energy and entropy to bolster the robustness of anomaly segmentation under complex open-world environments. Additionally, a diffusion-based anomalous training data synthesizer is proposed to generate diverse and high-quality anomalous images to enhance the existing copy-paste training data synthesizer. Extensive experimental results on both public and ComSAmy benchmarks demonstrate that our proposed diffusion-based synthesizer with energy and entropy learning (DiffEEL) framework serves as an effective and generalizable plug-and-play method to enhance existing models, yielding an average improvement of around 4.96% in AUPRC and 9.87% in $\rm {FPR}_{95}$ . Song Xia, Yi Yu 0011, Henghui Ding, Wenhan Yang, Shifei Liu, Alex Chichung Kot, Xudong Jiang 0001 |
IEEE Trans. Image Process. | 4 |
| 2026 | Seeing in the Dark with Ambient GuidanceabstractA low-light image taken in a dark scene usually suffers from severe distortions, which does not accurately characterize the ambient lighting. Long exposure is an accustomed way to capture more supplementary light and alleviate the degradation, but sometimes it induces other distortions, e.g., blurriness. To address this issue, we propose a new paradigm that introduces additional captured ambient guidance, i.e., a long-exposure image to steer the low-light enhancement. In practice, this long-exposure image can be obtained conveniently, but usually suffers from blurriness and misalignment. To effectively extract and fuse information from degraded and misaligned low-light and guidance image pairs, we propose a Long Exposure Compensation Network (LECNet). Adaptive Band Regression is introduced to disentangle the image into multi-scale representations and coarse-to-fine aggregate them with an attention mechanism. For stable image-guidance registration and artifact suppression, we propose a Bounded Cross-domain Deformable Alignment to warp the guidance based on extracted feature pyramids step by step. To integrate knowledge about the degradation into our LECNet for better fidelity, a dual learned back projection is enforced between the predicted result and the paired inputs in illumination and texture detail consistency, serving the model training for both offline training and online sample-adaptive finetuning. For training and evaluation of this new paradigm, we build a dataset with both synthetic and real-captured image triplets of long/short exposure pairs and extra blurry guidance. The experimental evaluation demonstrates the significance of our new paradigm, as well as the superiority of our LECNet and its usability in the real world. Haofeng Huang, Wenhan Yang, Mengnan Wang, Ling-Yu Duan, Jiaying Liu 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2026 | Enhanced industrial anomaly detection via CutMask data augmentation: a self-supervised approach
Dingjian Yao, Wenhan Yang, Zhengsong Wang |
Vis. Comput. | 3 |
| 2025 | Backdoor Attacks Against No-Reference Image Quality Assessment Models via a Scalable TriggerabstractNo-Reference Image Quality Assessment (NR-IQA), responsible for assessing the quality of a single input image without using any reference, plays a critical role in evaluating and optimizing computer vision systems, e.g., low-light enhancement. Recent research indicates that NR-IQA models are susceptible to adversarial attacks, which can significantly alter predicted scores with visually imperceptible perturbations. Despite revealing vulnerabilities, these attack methods have limitations, including high computational demands, untargeted manipulation, limited practical utility in white-box scenarios, and reduced effectiveness in black-box scenarios. To address these challenges, we shift our focus to another significant threat and present a novel poisoning-based backdoor attack against NR-IQA (BAIQA), allowing the attacker to manipulate the IQA model's output to any desired target value by simply adjusting a scaling coefficient alpha for the trigger. We propose to inject the trigger in the discrete cosine transform (DCT) domain to improve the local invariance of the trigger for countering trigger diminishment in NR-IQA models due to widely adopted data augmentations. Furthermore, the universal adversarial perturbations (UAP) in the DCT space are designed as the trigger, to increase IQA model susceptibility to manipulation and improve attack effectiveness. In addition to the heuristic method for poison-label BAIQA (P-BAIQA), we explore the design of clean-label BAIQA (C-BAIQA), focusing on alpha sampling and image data refinement, driven by theoretical insights we reveal. Extensive experiments on diverse datasets and various NR-IQA models demonstrate the effectiveness of our attacks. Yi Yu 0011, Song Xia, Xun Lin, Wenhan Yang, Shijian Lu, Yap-Peng Tan, Alex Chichung Kot |
AAAI | 4 |
| 2025 | UP-Restorer: When Unrolling Meets Prompts for Unified Image RestorationabstractAll-in-one restoration needs to implicitly distinguish between different degradation conditions and apply specific prior constraints accordingly. To fulfill this goal, our work makes the first effort to create an all-in-one restoration via unrolling from the typical maximum a-posterior optimization function. This unrolling framework naturally leads to the construction of progressively solving models, which are equivalent to a diffusion enhancer taking as input dynamically generated prompts. Under a score-based diffusion model, the prompts are integrated for propogating and updating several context-related variables, i.e. transmission map, atmospheric light map and noise or rain map progressively. Such learned prompt generation process, which simulates the nonlinear operations in the unrolled solution, is combined with linear operations owning clear physics implications to make the diffusion models well reguarlized and more effective in learning degradation-related visual priors. Experimental results demonstrate that our method achieves significant performance improvements across various image restoration tasks, realizing true all-in-one image restoration. Minghao Liu 0019, Wenhan Yang, Jinyi Luo, Jiaying Liu 0001 |
AAAI | 2 |
| 2025 | Fast Omni-Directional Image Super-Resolution: Adapting the Implicit Image Function with Pixel and Semantic-Wise Spherical Geometric PriorsabstractIn the context of Omni-Directional Image (ODI) Super-Resolution (SR), the unique challenge arises from the non-uniform oversampling characteristics caused by EquiRectangular Projection (ERP). Considerable efforts in designing complex spherical convolutions or polyhedron reprojection offer significant performance improvements but at the expense of cumbersome processing procedures and slower inference speeds. Under these circumstances, this paper proposes a new ODI-SR model characterized by its capacity to perform Fast and Arbitrary-scale ODI-SR processes, denoted as FAOR. The key innovation lies in adapting the implicit image function from the planar image domain to the ERP image domain by incorporating spherical geometric priors at both the latent representation and image reconstruction stages, in a low-overhead manner. Specifically, at the latent representation stage, we adopt a pair of pixel-wise and semantic-wise sphere-to-planar distortion maps to perform affine transformations on the latent representation, thereby incorporating it with spherical properties. Moreover, during the image reconstruction stage, we introduce a geodesic-based resampling strategy, aligning the implicit image function with spherical geometrics without introducing additional parameters. As a result, the proposed FAOR outperforms the state-of-the-art ODI-SR models with a much faster inference speed. Extensive experimental results and ablation studies have demonstrated the effectiveness of our design. Xuelin Shen, Silin Zheng, Kang Xiao, Wenhan Yang |
AAAI | 5 |
| 2025 | Unified Coding for Both Human Perception and Generalized Machine Analytics with CLIP SupervisionabstractThe image compression model has long struggled with adaptability and generalization, as the decoded bitstream typically serves only human or machine needs and fails to preserve information for unseen visual tasks. Therefore, this paper innovatively introduces supervision obtained from multimodal pre-training models and incorporates adaptive multi-objective optimization tailored to support both human visual perception and machine vision simultaneously with a single bitstream, denoted as Unified and Generalized Image Coding for Machine (UG-ICM). Specifically, to get rid of the reliance between compression models with downstream task supervision, we introduce Contrastive Language-Image Pre-training (CLIP) models into the training constraint for improved generalization. Global-to-instance-wise CLIP supervision is applied to help obtain hierarchical semantics that make models more generalizable for the tasks relying on the information of different granularity. Furthermore, for supporting both human and machine visions with only a unifying bitstream, we incorporate a conditional decoding strategy that takes as conditions human or machine preferences, enabling the bitstream to be decoded into different versions for corresponding preferences. As such, our proposed UG-ICM is fully trained in a self-supervised manner, i.e., without awareness of any specific downstream models and tasks. The extensive experiments have shown that the proposed UG-ICM is capable of achieving remarkable improvements in various unseen machine analytics tasks, while simultaneously providing perceptually satisfying images. Kangsheng Yin, Xuelin Shen, Yu-Lin He, Wenhan Yang, Shiqi Wang 0001 |
AAAI | 5 |
| 2025 | Theoretical Insights in Model Inversion Robustness and Conditional Entropy Maximization for Collaborative Inference SystemsabstractBy locally encoding raw data into intermediate features, collaborative inference enables end users to leverage powerful deep learning models without exposure of sensitive raw data to cloud servers. However, recent studies have revealed that these intermediate features may not sufficiently preserve privacy, as information can be leaked and raw data can be reconstructed via model inversion attacks (MIAs). Obfuscation-based methods, such as noise corruption, adversarial representation learning, and information filters, enhance the inversion robustness by obfuscating the task-irrelevant redundancy empirically. However, methods for quantifying such redundancy remain elusive, and the explicit mathematical relation between this redundancy minimization and inversion robustness enhancement has not yet been established. To address that, this work first theoretically proves that the conditional entropy of inputs given intermediate features provides a guaranteed lower bound on the reconstruction mean square error (MSE) under any MIA. Then, we derive a differentiable and solvable measure for bounding this conditional entropy based on the Gaussian mixture estimation and propose a conditional entropy maximization (CEM) algorithm to enhance the inversion robustness. Experimental results on four datasets demonstrate the effectiveness and adaptability of our proposed CEM; without compromising feature utility and computing efficiency, plugging the proposed CEM into obfuscation-based defense mechanisms consistently boosts their inversion robustness, achieving average gains ranging from 12.9% to 48.2%. Code is available at https://github.com/xiasong0501/CEM. Song Xia, Yi Yu 0011, Wenhan Yang, Meiwen Ding, Zhuo Chen 0006, Ling-Yu Duan, Alex Chichung Kot, Xudong Jiang 0001 |
CVPR | 3 |
| 2025 | Adaptive Dual Uncertainty Optimization: Boosting Monocular 3D Object Detection under Test-Time ShiftsabstractAccurate monocular 3D object detection (M3OD) is pivotal for safety-critical applications like autonomous driving, yet its reliability deteriorates significantly under real-world domain shifts caused by environmental or sensor variations. To address these shifts, Test-Time Adaptation (TTA) methods have emerged, enabling models to adapt to target distributions during inference. While prior TTA approaches recognize the positive correlation between low uncertainty and high generalization ability, they fail to address the dual uncertainty inherent to M3OD: semantic uncertainty (ambiguous class predictions) and geometric uncertainty (unstable spatial localization). To bridge this gap, we propose Dual Uncertainty Optimization (DUO), the first TTA framework designed to jointly minimize both uncertainties for robust M3OD. Through a convex optimization lens, we introduce an innovative convex structure of the focal loss and further derive a novel unsupervised version, enabling label-agnostic uncertainty weighting and balanced learning for high-uncertainty objects. In parallel, we design a semantic-aware normal field constraint that preserves geometric coherence in regions with clear semantic cues, reducing uncertainty from the unstable 3D representation. This dual-branch mechanism forms a complementary loop: enhanced spatial perception improves semantic classification, and robust semantic predictions further refine spatial understanding. Extensive experiments demonstrate the superiority of DUO over existing methods across various datasets and domain shift types. Xinzhu Ma, Shixiang Tang, Wenhan Yang, Ling-Yu Duan |
ICCV | 6 |
| 2025 | Cross-Granularity Online Optimization with Masked Compensated Information for Learned Image Compression
Haowei Kuang, Wenhan Yang, Zongming Guo, Jiaying Liu 0001 |
ICCV | 2 |
| 2025 | AFUNet: Cross-Iterative Alignment-Fusion Synergy for HDR Reconstruction via Deep Unfolding ParadigmabstractExisting learning-based methods effectively reconstruct HDR images from multi-exposure LDR inputs with extended dynamic range and improved detail, but they rely more on empirical design rather than theoretical foundation, which can impact their reliability. To address these limitations, we propose the cross-iterative Alignment and Fusion deep Unfolding Network (AFUNet), where HDR reconstruction is systematically decoupled into two interleaved subtasks -- alignment and fusion -- optimized through alternating refinement, achieving synergy between the two subtasks to enhance the overall performance. Our method formulates multi-exposure HDR reconstruction from a Maximum A Posteriori (MAP) estimation perspective, explicitly incorporating spatial correspondence priors across LDR images and naturally bridging the alignment and fusion subproblems through joint constraints. Building on the mathematical foundation, we reimagine traditional iterative optimization through unfolding -- transforming the conventional solution process into an end-to-end trainable AFUNet with carefully designed modules that work progressively. Specifically, each iteration of AFUNet incorporates an Alignment-Fusion Module (AFM) that alternates between a Spatial Alignment Module (SAM) for alignment and a Channel Fusion Module (CFM) for adaptive feature fusion, progressively bridging misaligned content and exposure discrepancies. Extensive qualitative and quantitative evaluations demonstrate AFUNet's superior performance, consistently surpassing state-of-the-art methods. Our code is available at: https://github.com/eezkni/AFUNet Zhangkai Ni, Wenhan Yang |
ICCV | 3 |
| 2025 | PRE-Mamba: A 4D State Space Model for Ultra-High-Frequent Event Camera DerainingabstractEvent cameras excel in high temporal resolution and dynamic range but suffer from dense noise in rainy conditions. Existing event deraining methods face trade-offs between temporal precision, deraining effectiveness, and computational efficiency. In this paper, we propose PRE-Mamba, a novel point-based event camera deraining framework that fully exploits the spatiotemporal characteristics of raw event and rain. Our framework introduces a 4D event cloud representation that integrates dual temporal scales to preserve high temporal precision, a Spatio-Temporal Decoupling and Fusion module (STDF) that enhances deraining capability by enabling shallow decoupling and interaction of temporal and spatial information, and a Multi-Scale State Space Model (MS3M) that captures deeper rain dynamics across dual-temporal and multi-spatial scales with linear computational complexity. Enhanced by frequency-domain regularization, PRE-Mamba achieves superior performance (0.95 SR, 0.91 NR, and 0.4s/M events) with only 0.26M parameters on EventRain-27K, a comprehensive dataset with labeled synthetic and real-world sequences. Moreover, our method generalizes well across varying rain intensities, viewpoints, and even snowy conditions. Ciyu Ruan, Ruishan Guo, Zihang Gong, Jingao Xu, Wenhan Yang, Xinlei Chen |
ICCV | 5 |
| 2025 | Which Tasks Should Be Compressed Together? A Causal Discovery Approach for Efficient Multi-Task Representation CompressionabstractConventional image compression methods are inadequate for intelligent analysis, as they overemphasize pixel-level precision while neglecting semantic significance and the interaction among multiple tasks. This paper introduces a Taskonomy-Aware Multi-Task Compression framework comprising (1) inter-coherent task grouping, which organizes synergistic tasks into shared representations to improve multi-task accuracy and reduce encoding volume, and (2) a conditional entropy-based directed acyclic graph (DAG) that captures causal dependencies among grouped representations. By leveraging parent representations as contextual priors for child representations, the framework effectively utilizes cross-task information to improve entropy model accuracy. Experiments on diverse vision tasks, including Keypoint 2D, Depth Z-buffer, Semantic Segmentation, Surface Normal, Edge Texture, and Autoencoder, demonstrate significant bitrate-performance gains, validating the method’s capability to reduce system entropy uncertainty. These findings underscore the potential of leveraging representation disentanglement, synergy, and causal modeling to learn compact representations, which enable efficient multi-task compression in intelligent systems. Sha Guo, Zhuo Chen 0006, Wenhan Yang, Ling-Yu Duan |
ICLR | 5 |
| 2025 | Mini-batch Coresets for Memory-efficient Language Model Training on Data MixturesabstractTraining with larger mini-batches improves the convergence rate and can yield superior performance. However, training with large mini-batches becomes prohibitive for Large Language Models (LLMs), due to the large GPU memory requirement. To address this problem, an effective approach is finding small mini-batch coresets that closely match the gradient of larger mini-batches. However, this approach becomes infeasible and ineffective for LLMs, due to the highly imbalanced mixture of sources in language data, use of the Adam optimizer, and the very large gradient dimensionality of LLMs. In this work, we address the above challenges by proposing *Coresets for Training LLMs* (CoLM). First, we show that mini-batch coresets found by gradient matching do not contain representative examples of the small sources w.h.p., and thus including all examples of the small sources in the mini-batch coresets is crucial for optimal performance. Second, we normalize the gradients by their historical exponential to find mini-batch coresets for training with Adam. Finally, we leverage zeroth-order methods to find smooth gradient of the last *V*-projection matrix and sparsify it to keep the dimensions with the largest normalized gradient magnitude. We apply CoLM to fine-tuning Phi-2, Phi-3, Zephyr, and Llama-3 models with LoRA on MathInstruct and SuperGLUE benchmark. Remarkably, CoLM reduces the memory requirement of fine-tuning by 2x and even outperforms training with 4x larger mini-batches. Moreover, CoLM seamlessly integrates with existing memory-efficient training methods like LoRA, further reducing the memory requirements of training LLMs. Wenhan Yang, Rathul Anand, Yu Yang 0007, Baharan Mirzasoleiman |
ICLR | 2 |
| 2025 | Adaptive Gradient Quantization with Bit Allocation for Distributed Deep LearningabstractGradient compression plays a crucial role in mitigating communication overhead in distributed deep learning. Existing gradient compression methods usually employ fix-bit quantization across all layers, neglecting the varying sensitivities of different layers to compression, resulting in suboptimal performance. In this paper, we introduce a layer-wise bit allocation mechanism for gradient quantization that minimizes overall quantization error within a specified bit budget. To address the heavy computational load of conventional greedy search approach for bit allocation, we develop two acceleration techniques to reduce computational overhead, thereby making the proposed bit allocation method feasible for real-time deep learning training. Specifically, by observing the bit allocation statistics, we propose Bit Searching Range Optimization to narrow the available bit options, while the Bit Pre-Assignment selectively bypasses certain searching processes. Experimental results across various neural network models and datasets demonstrate the effectiveness of our proposed bit allocation methods for gradient quantization. The combination of proposed acceleration techniques offers an advantageous trade-off among quantization error, model training performance and time consumption. Moreover, our proposed bit allocation methods can be seamlessly integrated with existing gradient compression approaches, improving overall performance. Fei Gao 0019, Wenhan Yang, Ling-Yu Duan, Zhuo Chen 0006 |
ICME | 4 |
| 2025 | MTL-UE: Learning to Learn Nothing for Multi-Task LearningabstractMost existing unlearnable strategies focus on preventing unauthorized users from training single-task learning (STL) models with personal data. Nevertheless, the paradigm has recently shifted towards multi-task data and multi-task learning (MTL), targeting generalist and foundation models that can handle multiple tasks simultaneously. Despite their growing importance, MTL data and models have been largely neglected while pursuing unlearnable strategies. This paper presents MTL-UE, the first unified framework for generating unlearnable examples for multi-task data and MTL models. Instead of optimizing perturbations for each sample, we design a generator-based structure that introduces label priors and class-wise feature embeddings which leads to much better attacking performance. In addition, MTL-UE incorporates intra-task and inter-task embedding regularization to increase inter-class separation and suppress intra-class variance which enhances the attack robustness greatly. Furthermore, MTL-UE is versatile with good supports for dense prediction tasks in MTL. It is also plug-and-play allowing integrating existing surrogate-dependent unlearnable methods with little adaptation. Extensive experiments show that MTL-UE achieves superior attacking performance consistently across 4 MTL datasets, 3 base UE methods, 5 model backbones, and 5 MTL task-weighting strategies. Code is available at https://github.com/yuyi-sd/MTL-UE. Yi Yu 0011, Song Xia, Siyuan Yang 0001, Chenqi Kong, Wenhan Yang, Shijian Lu, Yap-Peng Tan, Alex Chichung Kot |
ICML | 5 |
| 2025 | Privacy-Shielded Image Compression: Defending Against Exploitation from Vision-Language Pretrained ModelsabstractThe improved semantic understanding of vision-language pretrained (VLP) models has made it increasingly difficult to protect publicly posted images from being exploited by search engines and other similar tools. In this context, this paper seeks to protect users' privacy by implementing defenses at the image compression stage to prevent exploitation. Specifically, we propose a flexible coding method, termed Privacy-Shielded Image Compression (PSIC), that can produce bitstreams with multiple decoding options. By default, the bitstream is decoded to preserve satisfactory perceptual quality while preventing interpretation by VLP models. Our method also retains the original image compression functionality. With a customizable input condition, the proposed scheme can reconstruct the image that preserves its full semantic information. A Conditional Latent Trigger Generation (CLTG) module is proposed to produce bias information based on customizable conditions to guide the decoding process into different reconstructed versions, and an Uncertainty-Aware Encryption-Oriented (UAEO) optimization function is designed to leverage the soft labels inferred from the target VLP model's uncertainty on the training data. This paper further incorporates an adaptive multi-objective optimization strategy to obtain improved encrypting performance and perceptual quality simultaneously within a unified training process. The proposed scheme is plug-and-play and can be seamlessly integrated into most existing Learned Image Compression (LIC) models. Extensive experiments across multiple downstream tasks have demonstrated the effectiveness of our design. Xuelin Shen, Jiayin Xu, Kangsheng Yin, Wenhan Yang |
ICML | 4 |
| 2025 | End-to-End Low-Light Enhancement for Object Detection with Learned Metadata from RAWsabstractAlthough RAW images offer advantages over sRGB by avoiding ISP-induced distortion and preserving more information in low-light conditions, their widespread use is limited due to high storage costs, transmission burdens, and the need for significant architectural changes for downstream tasks. To address the issues, this paper explores a new raw-based machine vision paradigm, termed Compact RAW Metadata-guided Image Refinement (CRM-IR). In particular, we propose a Machine Vision-oriented Image Refinement (MV-IR) module that refines sRGB images to better suit machine vision preferences, guided by learned raw metadata. Such a design allows the CRM-IR to focus on extracting the most essential metadata from raw images to support downstream machine vision tasks, while remaining plug-and-play and fully compatible with existing imaging pipelines, without any changes to model architectures or ISP modules. We implement our CRM-IR scheme on various object detection networks, and extensive experiments under low-light conditions demonstrate that it can significantly improve performance with an additional bitrate cost of less than $10^{-3}$ bits per pixel. Xuelin Shen, Haifeng Jiao, Yu-Lin He, Wenhan Yang |
NeurIPS | 5 |
| 2025 | Prompt-Guided Alignment with Information Bottleneck Makes Image Compression Also a RestorerabstractLearned Image Compression (LIC) models face critical challenges in real-world scenarios due to various environmental degradations, such as fog and rain. Due to the distribution mismatch between degraded inputs and clean training data, well-trained LIC models suffer from reduced compression efficiency, while retraining dedicated models for diverse degradation types is costly and impractical. Our method addresses the above issue by leveraging prompt learning under the information bottleneck principle, enabling compact extraction of shared components between degraded and clean images for improved latent alignment and compression efficiency. In detail, we propose an Information Bottleneck-constrained Latent Representation Unifying (IB-LRU) scheme, in which a Probabilistic Prompt Generator (PPG) is deployed to simultaneously capture the distribution of different degradations. Such a design dynamically guides the latent-representation process at the encoder through a gated modulation process. Moreover, to promote the degradation distribution capture process, the probabilistic prompt learning is guided by the Information Bottleneck (IB) principle. That is,IB constrains the information encoded in the prompt to focus solely on degradation characteristics while avoiding the inclusion of redundant image contextual information. We apply our IB-LRU method to a variety of state-of-the-art LIC backbones, and extensive experiments under various degradation scenarios demonstrate the effectiveness of our design. Our code will be publicly available. Xuelin Shen, Jiayin Xu, Wenhan Yang |
NeurIPS | 4 |
| 2025 | Towards Data-Centric Face Anti-spoofing: Improving Cross-Domain Generalization via Physics-Based Data Synthesis
Rizhao Cai, Cecelia Soh, Zitong Yu, Haoliang Li, Wenhan Yang, Alex Chichung Kot |
Int. J. Comput. Vis. | 5 |
| 2025 | Semantic Masking with Curriculum Learning for Robust HDR Image Reconstruction
Zhangkai Ni, Kerui Ren, Wenhan Yang, Hanli Wang, Sam Kwong |
Int. J. Comput. Vis. | 4 |
| 2025 | Large Models for Aerial Edges: An Edge-Cloud Model Evolution and Communication ParadigmabstractThe future sixth-generation (6G) of wireless networks is expected to surpass its predecessors by offering ubiquitous coverage through integrated air-ground deployments in both communication and computing domains. In such networks, aerial platforms, such as unmanned aerial vehicles (UAVs), conduct artificial intelligence (AI) computations based on multi-modal data to support diverse applications including surveillance and environment construction. However, these multi-domain inference and content generation tasks require large AI models, demanding powerful computing capabilities and finely tuned inference models trained on rich datasets, thus posing significant challenges for UAVs. To tackle this problem, we propose an integrated air-ground edge-cloud model framework, in which UAVs serve as edge nodes for data collection and small model computation. Through wireless channels, UAVs collaborate with ground cloud servers providing large model computation and model updating for edge UAVs. With limited wireless communication bandwidth, the proposed framework faces the challenge of information exchange scheduling between the edge UAVs and the cloud server. To tackle this, we present joint task allocation, transmission resource allocation, transmission data quantization design, and edge model update design to enhance the inference accuracy of the integrated air-ground edge-cloud model evolution framework by mean average precision (mAP) maximization. A closed-form lower bound on the mAP of the proposed framework is derived based on the mAP of the edge model and mAP of the cloud model, and the solution to the mAP maximization problem is optimized accordingly. Simulations, based on results from vision-based classification experiments, consistently demonstrate that the mAP of the proposed integrated air-ground edge-cloud model evolution framework outperforms both a centralized cloud model framework and a distributed edge model framework across various communication bandwidths and data sizes. Shuhang Zhang, Ke Chen 0004, Boya Di, Hongliang Zhang 0001, Wenhan Yang, Dusit Niyato, Zhu Han 0001, H. Vincent Poor |
IEEE J. Sel. Areas Commun. | 6 |
| 2025 | CLMS: Bridging domain gaps in medical imaging segmentation with source-free continual learning for robust knowledge transfer and adaptation
Weilu Li, Hao Zhou 0030, Wenhan Yang, Zhi Xie |
Medical Image Anal. | 4 |
| 2025 | Interpretable Optimization-Inspired Unfolding Network for Low-Light Image EnhancementabstractRetinex model-based methods have shown to be effective in layer-wise manipulation with well-designed priors for low-light image enhancement (LLIE). However, the hand-crafted priors and conventional optimization algorithm adopted to solve the layer decomposition problem result in the lack of adaptivity and efficiency. To this end, this paper proposes a Retinex-based deep unfolding network (URetinex-Net++), which unfolds an optimization problem into a learnable network to decompose a low-light image into reflectance and illumination layers. By formulating the decomposition problem as an implicit priors regularized model, three learning-based modules are carefully designed, responsible for data-dependent initialization, high-efficient unfolding optimization, and fairly-flexible component adjustment, respectively. Particularly, the proposed unfolding optimization module, introducing two networks to adaptively fit implicit priors in the data-driven manner, can realize noise suppression and details preservation for decomposed components. URetinex-Net++ is a further augmented version of URetinex-Net, which introduces a cross-stage fusion block to alleviate the color defect in URetinex-Net. Therefore, boosted performance on LLIE can be obtained in both visual quality and quantitative metrics, where only a few parameters are introduced and little time is cost. Extensive experiments on real-world low-light images qualitatively and quantitatively demonstrate the effectiveness and superiority of the proposed URetinex-Net++ over state-of-the-art methods. Wenhui Wu 0001, Jian Weng 0009, Xu Wang 0006, Wenhan Yang, Jianmin Jiang |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2025 | Robust and Transferable Backdoor Attacks Against Deep Image Compression With Selective Frequency PriorabstractRecent advancements in deep learning-based compression techniques have demonstrated remarkable performance surpassing traditional methods. Nevertheless, deep neural networks have been observed to be vulnerable to backdoor attacks, where an added pre-defined trigger pattern can induce the malicious behavior of the models. In this paper, we propose a novel approach to launch a backdoor attack with multiple triggers against learned image compression models. Drawing inspiration from the widely used discrete cosine transform (DCT) in existing compression codecs and standards, we propose a frequency-based trigger injection model that adds triggers in the DCT domain. In particular, we design several attack objectives that are adapted for a series of diverse scenarios, including: 1) attacking compression quality in terms of bit-rate and reconstruction quality; 2) attacking task-driven measures, such as face recognition and semantic segmentation in downstream applications. To facilitate more efficient training, we develop a dynamic loss function that dynamically balances the impact of different loss terms with fewer hyper-parameters, which also results in more effective optimization of the attack objectives with improved performance. Furthermore, we consider several advanced scenarios. We evaluate the resistance of the proposed backdoor attack to the defensive pre-processing methods and then propose a two-stage training schedule along with the design of robust frequency selection to further improve resistance. To strengthen both the cross-model and cross-domain transferability on attacking downstream CV tasks, we propose to shift the classification boundary in the attack loss during training. Extensive experiments also demonstrate that by employing our trained trigger injection models and making slight modifications to the encoder parameters of the compression model, our proposed attack can successfully inject multiple backdoors accompanied by their corresponding triggers into a single image compression model. Yi Yu 0011, Yufei Wang 0006, Wenhan Yang, Lanqing Guo, Shijian Lu, Ling-Yu Duan, Yap-Peng Tan, Alex Chichung Kot |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | Any Fashion Attribute Editing: Dataset and Pretrained ModelsabstractFashion attribute editing is essential for combining the expertise of fashion designers with the potential of generative artificial intelligence. In this work, we focus on 'any' fashion attribute editing: 1) the ability to edit 78 fine-grained design attributes commonly observed in daily life; 2) the capability to modify desired attributes while keeping the rest components still; and 3) the flexibility to continuously edit on the edited image. To this end, we present the Any Fashion Attribute Editing (AFED) dataset, which includes 830 K high-quality fashion images from sketch and product domains, filling the gap for a large-scale, openly accessible fine-grained dataset. We also propose Twin-Net, a twin encoder-decoder GAN inversion method that offers diverse and precise information for high-fidelity image reconstruction. This inversion model, trained on the new dataset, serves as a robust foundation for attribute editing. Additionally, we introduce PairsPCA to identify semantic directions in latent space, enabling accurate editing without manual supervision. Comprehensive experiments, including comparisons with ten state-of-the-art image inversion methods and four editing algorithms, demonstrate the effectiveness of our Twin-Net and editing algorithm. All data and models are available at https://github.com/ArtmeScienceLab/AnyFashionAttributeEditing. Shumin Zhu, Xingxing Zou, Wenhan Yang, Wai Keung Wong |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | Facial Image Compression via Neural Image Manifold CompressionabstractAlthough the recent learning-based image and video coding techniques achieve rapid development, the signal fidelity-driven target in these methods leads to the divergence to a highly effective and efficient coding framework for both human and machine. In this paper, we aim to address the issue by making use of the power of generative models to bridge the gap between full fidelity (for human vision) and high discrimination (for machine vision). Therefore, relying on existing pretrained generative adversarial networks (GAN), we build a GAN inversion framework that projects the image into a low-dimensional natural image manifold. In this manifold, the feature is highly discriminative and also encodes the appearance information of the image, named aslatent code. Taking a variational bit-rate constraint with a hyperprior model to model/suppress the entropy of image manifold code, our method is capable of fulfilling the needs of both machine and human visions at very low bit-rates. To improve the visual quality of image reconstruction, we further proposemultiple latent codesandscalable inversion. The former gets several latent codes in the inversion, while the latter additionally compresses and transmits a shallow compact feature to support visual reconstruction. Experimental results demonstrate the superiority of our method in both human vision tasks,i.e. image reconstruction, and machine vision tasks, including semantic parsing and attribute prediction. Wenhan Yang, Haofeng Huang, Jiaying Liu 0001, Alex Chichung Kot |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2025 | Toward Model Resistant to Transferable Adversarial Examples via Trigger ActivationabstractAdversarial examples, characterized by imperceptible perturbations, pose significant threats to deep neural networks by misleading their predictions. A critical aspect of these examples is their transferability, allowing them to deceive unseen models in closed-box scenarios. Despite the widespread exploration of defense methods, including those on transferability, they show limitations: inefficient deployment, ineffective defense, and degraded performance on clean images. In this work, we introduce a novel training paradigm aimed at enhancing robustness against transferable adversarial examples (TAEs) in a more efficient and effective way. We propose a model that exhibits random guessing behavior when presented with clean data$\boldsymbol {x}$as input, and generates accurate predictions when with triggered data$\boldsymbol {x}+\boldsymbol {\tau }$. Importantly, the trigger$\boldsymbol {\tau }$remains constant for all data instances. We refer to these models as models with trigger activation. We are surprised to find that these models exhibit certain robustness against TAEs. Through the consideration of first-order gradients, we provide a theoretical analysis of this robustness. Moreover, through the joint optimization of the learnable trigger and the model, we achieve improved robustness to transferable attacks. Extensive experiments conducted across diverse datasets, evaluating a variety of attacking methods, underscore the effectiveness and superiority of our approach. Yi Yu 0011, Song Xia, Xun Lin, Chenqi Kong, Wenhan Yang, Shijian Lu, Yap-Peng Tan, Alex Chichung Kot |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2025 | Rethinking Artifact Mitigation in HDR Reconstruction: From Detection to OptimizationabstractArtifact remains a long-standing challenge in High Dynamic Range (HDR) reconstruction. Existing methods focus on model designs for artifact mitigation but ignore explicit detection and suppression strategies. Because artifact lacks clear boundaries, distinct shapes, and semantic consistency, and there is no existing dedicated dataset for HDR artifact, progress in direct artifact detection and recovery is impeded. To bridge the gap, we propose a unified HDR reconstruction framework that integrates artifact detection and model optimization. Firstly, we build the first HDR artifact dataset (HADataset), comprising 1,213 diverse multi-exposure Low Dynamic Range (LDR) image sets and 1,765 HDR image pairs with per-pixel artifact annotations. Secondly, we develop an effective HDR artifact detector (HADetector), a robust artifact detection model capable of accurately localizing HDR reconstruction artifact. HADetector plays two pivotal roles: (1) enhancing existing HDR reconstruction models through fine-tuning, and (2) serving as a non-reference image quality assessment (NR-IQA) metric, the Artifact Score (AS), which aligns closely with human visual perception for reliable quality evaluation. Extensive experiments validate the effectiveness and generalizability of our framework, including the HADataset, HADetector, fine-tuning paradigm, and AS metric. The code and datasets are available at: https://github.com/xinyueliii/hdr-artifact-detect-optimize. Zhangkai Ni, Wenhan Yang, Hanli Wang, Lianghua He, Sam Kwong |
IEEE Trans. Image Process. | 4 |
| 2025 | Structural Similarity-Inspired Unfolding for Lightweight Image Super-ResolutionabstractMajor efforts in data-driven image super-resolution (SR) primarily focus on expanding the receptive field of the model to better capture contextual information. However, these methods are typically implemented by stacking deeper networks or leveraging transformer-based attention mechanisms, which consequently increases model complexity. In contrast, model-driven methods based on the unfolding paradigm show promise in improving performance while effectively maintaining model compactness through sophisticated module design. Based on these insights, we propose a Structural Similarity-Inspired Unfolding (SSIU) method for efficient image SR. This method is designed through unfolding an SR optimization function constrained by structural similarity, aiming to combine the strengths of both data-driven and model-driven approaches. Our model operates progressively following the unfolding paradigm. Each iteration consists of multiple Mixed-Scale Gating Modules (MSGM) and an Efficient Sparse Attention Module (ESAM). The former implements comprehensive constraints on features, including a structural similarity constraint, while the latter aims to achieve sparse activation. In addition, we design a Mixture-of-Experts-based Feature Selector (MoE-FS) that fully utilizes multi-level feature information by combining features from different steps. Extensive experiments validate the efficacy and efficiency of our unfolding-inspired network. Our model outperforms current state-of-the-art models, boasting lower parameter counts and reduced memory consumption. Our code will be available at: https://github.com/eezkni/SSIU. Zhangkai Ni, Wenhan Yang, Hanli Wang, Shiqi Wang 0001, Sam Kwong |
IEEE Trans. Image Process. | 3 |
| 2025 | Breaking Boundaries: Unifying Imaging and Compression for HDR Image CompressionabstractHigh Dynamic Range (HDR) images present unique challenges for Learned Image Compression (LIC) due to their complex domain distribution compared to Low Dynamic Range (LDR) images. In coding practice, HDR-oriented LIC typically adopts preprocessing steps (e.g., perceptual quantization and tone mapping operation) to align the distributions between LDR and HDR images, which inevitably comes at the expense of perceptual quality. To address this challenge, we rethink the HDR imaging process which involves fusing multiple exposure LDR images to create an HDR image and propose a novel HDR image compression paradigm, Unifying Imaging and Compression (HDR-UIC). The key innovation lies in establishing a seamless pipeline from image capture to delivery and enabling end-to-end training and optimization. Specifically, a Mixture-ATtention (MAT)-based compression backbone merges LDR features while simultaneously generating a compact representation. Meanwhile, the Reference-guided Misalignment-aware feature Enhancement (RME) module mitigates ghosting artifacts caused by misalignment in the LDR branches, maintaining fidelity without introducing additional information. Furthermore, we introduce an Appearance Redundancy Removal (ARR) module to optimize coding resource allocation among LDR features, thereby enhancing the final HDR compression performance. Extensive experimental results demonstrate the efficacy of our approach, showing significant improvements over existing state-of-the-art HDR compression schemes. Our code is available at: https://github.com/plf1999/HDR-UIC. Xuelin Shen, Linfeng Pan, Zhangkai Ni, Yu-Lin He, Wenhan Yang, Shiqi Wang 0001, Sam Kwong |
IEEE Trans. Image Process. | 5 |
| 2025 | Shell-Guided Compression of Voxel Radiance FieldsabstractIn this paper, we address the challenge of significant memory consumption and redundant components in large-scale voxel-based model, which are commonly encountered in real-world 3D reconstruction scenarios. We propose a novel method called Shell-guided compression of Voxel Radiance Fields (SVRF), aimed at optimizing voxel-based model into a shell-like structure to reduce storage costs while maintaining rendering accuracy. Specifically, we first introduce a Shell-like Constraint, operating in two main aspects: 1) enhancing the influence of voxels neighboring the surface in determining the rendering outcomes, and 2) expediting the elimination of redundant voxels both inside and outside the surface. Additionally, we introduce an Adaptive Thresholds to ensure appropriate pruning criteria for different scenes. To prevent the erroneous removal of essential object parts, we further employ a Dynamic Pruning Strategy to conduct smooth and precise model pruning during training. The compression method we propose does not necessitate the use of additional labels. It merely requires the guidance of self-supervised learning based on predicted depth. Furthermore, it can be seamlessly integrated into any voxel-grid-based method. Extensive experimental results demonstrate that our method achieves comparable rendering quality while compressing the original number of voxel grids by more than 70%. Our code will be available at: https://github.com/eezkni/SVRF. Peiqi Yang, Zhangkai Ni, Hanli Wang, Wenhan Yang, Shiqi Wang 0001, Sam Kwong |
IEEE Trans. Image Process. | 4 |
| 2025 | M2Trans: Multi-Modal Regularized Coarse-to-Fine Transformer for Ultrasound Image Super-ResolutionabstractUltrasound image super-resolution (SR) aims to transform low-resolution images into high-resolution ones, thereby restoring intricate details crucial for improved diagnostic accuracy. However, prevailing methods relying solely on image modality guidance and pixel-wise loss functions struggle to capture the distinct characteristics of medical images, such as unique texture patterns and specific colors harboring critical diagnostic information. To overcome these challenges, this paper introduces the Multi-Modal Regularized Coarse-to-fine Transformer (M2Trans) for Ultrasound Image SR. By integrating the text modality, we establish joint image-text guidance during training, leveraging the medical CLIP model to incorporate richer priors from text descriptions into the SR optimization process, enhancing detail, structure, and semantic recovery. Furthermore, we propose a novel coarse-to-fine transformer comprising multiple branches infused with self-attention and frequency transforms to efficiently capture signal dependencies across different scales. Extensive experimental results demonstrate significant improvements over state-of-the-art methods on benchmark datasets, including CCA-US, US-CASE, and our newly created dataset MMUS1K, with a minimum improvement of 0.17dB, 0.30dB, and 0.28dB in terms of PSNR. Zhangkai Ni, Runyu Xiao, Wenhan Yang, Hanli Wang, Zhihua Wang 0002, Lihua Xiang |
IEEE J. Biomed. Health Informatics | 3 |
| 2025 | HFGlobalFormer: When High-Frequency Recovery Meets Global Context Modeling for Compressed Image DeraindropabstractWhen transmission medium and compression degradation are intertwined, new challenges emerge. This study addresses the problem of raindrop removal from compressed images, where raindrops obscure large areas of the background and compression leads to the loss of high-frequency (HF) information. The restoration of the former requires global contextual information, while the latter necessitates guidance for high-frequency details, resulting in a conflict in utilizing these two types of information when designing existing methods. To address this issue, we propose a novel transformer architecture that leverages the advantages of attention mechanism and HF-friendly design to effectively restore the compressed raindrop images at the framework, component, and module levels. Specifically, at the framework level, we integrate relative position multi-head self-attention and convolutional layers into the proposed low-high-frequency transformer (LHFT), where the former captures global contextual information and the latter focuses on high-frequency information. Their combination effectively resolves the issue of mixed degradation. At the component level, we utilize high-frequency depth-wise convolution (HFDC) with zero-mean kernels to improve the capability to extract high-frequency features, drawing inspiration from typical high-frequency filters like Prewitt and Sobel operators. Finally, at the module level, we introduce a low-high-attention module (LHAM) to adaptively allocate the importance of low and high frequencies along channels for effective fusion. We establish the JPEG-compressed raindrop image dataset and conduct extensive experiments on different compression rates. Experimental results demonstrate that the proposed method outperforms state-of-the-art methods without increasing computational costs. Rongqun Lin, Wenhan Yang, Baoliang Chen, Shiqi Wang 0001, Sam Kwong |
IEEE Trans. Multim. | 2 |
| 2025 | Transferring From Distortion to Perception-Oriented Optimization: Just-Noticeable-Distortion-Based Domain AdaptationabstractTheperception-distortion- tradeoffreveals the limitation of current low-level deep learning paradigms,i.e., minimizing reconstruction distortion does not guarantee improved perceptual quality. Acknowledging the lack of a reliableperception-oriented optimization function, we are motivated to explore a flexible approach for enhancing perceptual quality by steering thetradeoffto prioritizeperception. To this end, we reconsider theperception-distortionfunction by incorporating the Just-Noticeable-Distortion (JND) mechanism. We mathematically demonstrate that in the common image restoration process, altering the optimization target from natural images to distorted images—where the distortion intensity is constrained by the JND threshold and the distortion type aligns with that arising from the restorer itself—effectively obtained improvedperceptionindices without any changes to the restorer or optimization function. Accordingly, to facilitate various low-level learning models, we are motivated to construct the first large-scale CNN-oriented JND image dataset. Our dataset comprises 500 natural images and 4,500 degraded versions generated by a series of autoencoders, as well as the actual JND judgment results collected through rigorous subjective testing from twenty volunteers. Finally, a learning-based JND inference model is established on the proposed dataset and employed in the proposed JND-based adaptation scheme, where the inferred JND images serve as pseudo-ground truth for the training or fine-tuning processes of low-level vision models. Extensive experiments on image super-resolution and end-to-end image compression across multiple models have shown encouraging improvements in perceptual quality, demonstrating the effectiveness of the proposed scheme. Our dataset is available at:https://github.com/ohq17/CNN-Oriented-JND-Dataset. Xuelin Shen, Haoqiao Ou, Zhangkai Ni, Wenhan Yang, Shiqi Wang 0001, Sam Kwong |
IEEE Trans. Multim. | 4 |
| 2024 | Omnipotent Distillation with LLMs for Weakly-Supervised Natural Language Video Localization: When Divergence Meets ConsistencyabstractNatural language video localization plays a pivotal role in video understanding, and leveraging weakly-labeled data is considered a promising approach to circumvent the laborintensive process of manual annotations. However, this approach encounters two significant challenges: 1) limited input distribution, namely that the limited writing styles of the language query, annotated by human annotators, hinder the model’s generalization to real-world scenarios with diverse vocabularies and sentence structures; 2) the incomplete ground truth, whose supervision guidance is insufficient. To overcome these challenges, we propose an omnipotent distillation algorithm with large language models (LLM). The distribution of the input sample is enriched to obtain diverse multi-view versions while a consistency then comes to regularize the consistency of their results for distillation. Specifically, we first train our teacher model with the proposed intra-model agreement, where multiple sub-models are supervised by each other. Then, we leverage the LLM to paraphrase the language query and distill the teacher model to a lightweight student model by enforcing the consistency between the localization results of the paraphrased sentence and the original one. In addition, to assess the generalization of the model across different dimensions of language variation, we create extensive datasets by building upon existing datasets. Our experiments demonstrate substantial performance improvements adaptively to diverse kinds of language queries. Peijun Bao, Zihao Shao, Wenhan Yang, Boon Poh Ng, Meng Hwa Er, Alex Chichung Kot |
AAAI | 3 |
| 2024 | Local-Global Multi-Modal Distillation for Weakly-Supervised Temporal Video GroundingabstractThis paper for the first time leverages multi-modal videos for weakly-supervised temporal video grounding. As labeling the video moment is labor-intensive and subjective, the weakly-supervised approaches have gained increasing attention in recent years. However, these approaches could inherently compromise performance due to inadequate supervision. Therefore, to tackle this challenge, we for the first time pay attention to exploiting complementary information extracted from multi-modal videos (e.g., RGB frames, optical flows), where richer supervision is naturally introduced in the weaklysupervised context. Our motivation is that by integrating different modalities of the videos, the model is learned from synergic supervision and thereby can attain superior generalization capability. However, addressing multiple modalities† would also inevitably introduce additional computational overhead, and might become inapplicable if a particular modality is inaccessible. To solve this issue, we adopt a novel route: building a multi-modal distillation algorithm to capitalize on the multi-modal knowledge as supervision for model training, while still being able to work with only the single modal input during inference. As such, we can utilize the benefits brought by the supplementary nature of multiple modalities, without compromising the applicability in practical scenarios. Specifically, we first propose a cross-modal mutual learning framework and train a sophisticated teacher model to learn collaboratively from the multi-modal videos. Then we identify two sorts of knowledge from the teacher model, i.e., temporal boundaries and semantic activation map. And we devise a local-global distillation algorithm to transfer this knowledge to a student model of single-modal input at both local and global levels. Extensive experiments on large-scale datasets demonstrate that our method achieves state-of-the-art performance with/without multi-modal inputs. Peijun Bao, Wenhan Yang, Boon Poh Ng, Meng Hwa Er, Alex Chichung Kot |
AAAI | 3 |
| 2024 | Seeing Dark Videos via Self-Learned Bottleneck Neural RepresentationabstractEnhancing low-light videos in a supervised style presents a set of challenges, including limited data diversity, misalignment, and the domain gap introduced through the dataset construction pipeline. Our paper tackles these challenges by constructing a self-learned enhancement approach that gets rid of the reliance on any external training data. The challenge of self-supervised learning lies in fitting high-quality signal representations solely from input signals. Our work designs a bottleneck neural representation mechanism that extracts those signals. More in detail, we encode the frame-wise representation with a compact deep embedding and utilize a neural network to parameterize the video-level manifold consistently. Then, an entropy constraint is applied to the enhanced results based on the adjacent spatial-temporal context to filter out the degraded visual signals, e.g. noise and frame inconsistency. Last, a novel Chromatic Retinex decomposition is proposed to effectively align the reflectance distribution temporally. It benefits the entropy control on different components of each frame and facilitates noise-to-noise training, successfully suppressing the temporal flicker. Extensive experiments demonstrate the robustness and superior effectiveness of our proposed method. Our project is publicly available at: https://huangerbai.github.io/SLBNR/. Haofeng Huang, Wenhan Yang, Ling-Yu Duan, Jiaying Liu 0001 |
AAAI | 2 |
| 2024 | DeS3: Adaptive Attention-Driven Self and Soft Shadow Removal Using ViT SimilarityabstractRemoving soft and self shadows that lack clear boundaries from a single image is still challenging. Self shadows are shadows that are cast on the object itself. Most existing methods rely on binary shadow masks, without considering the ambiguous boundaries of soft and self shadows. In this paper, we present DeS3, a method that removes hard, soft and self shadows based on adaptive attention and ViT similarity. Our novel ViT similarity loss utilizes features extracted from a pre-trained Vision Transformer. This loss helps guide the reverse sampling towards recovering scene structures. Our adaptive attention is able to differentiate shadow regions from the underlying objects, as well as shadow regions from the object casting the shadow. This capability enables DeS3 to better recover the structures of objects even when they are partially occluded by shadows. Different from existing methods that rely on constraints during the training phase, we incorporate the ViT similarity during the sampling stage. Our method outperforms state-of-the-art methods on the SRD, AISTD, LRSS, USR and UIUC datasets, removing hard, soft, and self shadows robustly. Specifically, our method outperforms the SOTA method by 16% of the RMSE of the whole image on the LRSS dataset. Yeying Jin, Wei Ye 0005, Wenhan Yang, Yuan Yuan 0039, Robby T. Tan |
AAAI | 3 |
| 2024 | ColNeRF: Collaboration for Generalizable Sparse Input Neural Radiance FieldabstractNeural Radiance Fields (NeRF) have demonstrated impressive potential in synthesizing novel views from dense input, however, their effectiveness is challenged when dealing with sparse input. Existing approaches that incorporate additional depth or semantic supervision can alleviate this issue to an extent. However, the process of supervision collection is not only costly but also potentially inaccurate. In our work, we introduce a novel model: the Collaborative Neural Radiance Fields (ColNeRF) designed to work with sparse input. The collaboration in ColNeRF includes the cooperation among sparse input source images and the cooperation among the output of the NeRF. Through this, we construct a novel collaborative module that aligns information from various views and meanwhile imposes self-supervised constraints to ensure multi-view consistency in both geometry and appearance. A Collaborative Cross-View Volume Integration module (CCVI) is proposed to capture complex occlusions and implicitly infer the spatial location of objects. Moreover, we introduce self-supervision of target rays projected in multiple directions to ensure geometric and color consistency in adjacent regions. Benefiting from the collaboration at the input and output ends, ColNeRF is capable of capturing richer and more generalized scene representation, thereby facilitating higher-quality results of the novel view synthesis. Our extensive experimental results demonstrate that ColNeRF outperforms state-of-the-art sparse input generalizable NeRF methods. Furthermore, our approach exhibits superiority in fine-tuning towards adapting to new scenes, achieving competitive performance compared to per-scene optimized NeRF-based methods while significantly reducing computational costs. Our code is available at: https://github.com/eezkni/ColNeRF. Zhangkai Ni, Peiqi Yang, Wenhan Yang, Hanli Wang, Lin Ma 0002, Sam Kwong |
AAAI | 3 |
| 2024 | Misalignment-Robust Frequency Distribution Loss for Image TransformationabstractThis paper aims to address a common challenge in deep learning-based image transformation methods, such as im-age enhancement and super-resolution, which heavily rely on precisely aligned paired datasets with pixel-level align-ments. However, creating precisely aligned paired images presents significant challenges and hinders the advance-ment of methods trained on such data. To overcome this challenge, this paper introduces a novel and simple frequency Distribution Loss (FDL) for computing distribution distance within the frequency domain. Specifically, we transform image features into the frequency domain using Discrete Fourier Transformation (DFT). Subsequently, frequency components (amplitude and phase) are processed separately to form the FDL loss function. Our method is empirically proven effective as a training constraint due to the thoughtful utilization of global information in the frequency domain. Extensive experimental evaluations, fo-cusing on image enhancement and super-resolution tasks, demonstrate that FDL outperforms existing misalignment-robust loss functions. Furthermore, we explore the poten-tial of our FDL for image style transfer that relies solely on completely misaligned data. Our code is available at: https://github.com/eezkni/FDL Zhangkai Ni, Juncheng Wu, Wenhan Yang, Hanli Wang, Lin Ma 0002 |
CVPR | 4 |
| 2024 | SinSR: Diffusion-Based Image Super-Resolution in a Single StepabstractWhile super-resolution (SR) methods based on diffusion models exhibit promising results, their practical application is hindered by the substantial number of required inference steps. Recent methods utilize the degraded images in the initial state, thereby shortening the Markov chain. Nevertheless, these solutions either rely on a precise formulation of the degradation process or still necessitate a relatively lengthy generation path (e.g., 15 iterations). To enhance inference speed, we propose a simple yet effective method for achieving single-step SR generation, named SinSR. Specifically, we first derive a deterministic sampling process from the most recent state-of-the-art (SOTA) method for accelerating diffusion-based SR. This allows the mapping between the input random noise and the generated high-resolution image to be obtained in a reduced and acceptable number of inference steps during training. We show that this deterministic mapping can be distilled into a student model that performs SR within only one inference step. Additionally, we propose a novel consistency-preserving loss to simultaneously leverage the ground-truth image during the distillation process, ensuring that the performance of the student model is not solely bound by the feature manifold of the teacher model, resulting in further performance improvement. Extensive experiments conducted on synthetic and real-world datasets demonstrate that the proposed method can achieve comparable or even superior performance compared to both previous SOTA methods and the teacher model, in just one sampling step, resulting in a remarkable up to × 10 speedup for inference. Our code will be released at https://github.com/wyf0912/SinSR/. Yufei Wang 0006, Wenhan Yang, Yaohui Wang 0001, Lanqing Guo, Lap-Pui Chau, Ziwei Liu 0002, Yu Qiao 0001, Alex Chichung Kot, Bihan Wen |
CVPR | 2 |
| 2024 | E3M: Zero-Shot Spatio-Temporal Video Grounding with Expectation-Maximization Multimodal Modulation
Peijun Bao, Zihao Shao, Wenhan Yang, Boon Poh Ng, Alex Chichung Kot |
ECCV (83) | 3 |
| 2024 | A Unified Image Compression Method for Human Perception and Multiple Vision Tasks
Sha Guo, Lin Sui, Chen-Lin Zhang, Zhuo Chen 0006, Wenhan Yang, Ling-Yu Duan |
ECCV (71) | 5 |
| 2024 | Region-Adaptive Transform with Segmentation Prior for Image Compression
Yuxi Liu 0020, Wenhan Yang, Huihui Bai 0001, Yunchao Wei, Yao Zhao 0001 |
ECCV (46) | 2 |
| 2024 | Unrolled Decomposed Unpaired Learning for Controllable Low-Light Video Enhancement
Lingyu Zhu 0006, Wenhan Yang, Baoliang Chen, Hanwei Zhu, Zhangkai Ni, Qi Mao 0002, Shiqi Wang 0001 |
ECCV (23) | 2 |
| 2024 | Unlearnable Examples Detection via Iterative Filtering
Yi Yu 0011, Qichen Zheng, Siyuan Yang 0001, Wenhan Yang, Jun Liu 0036, Shijian Lu, Yap-Peng Tan, Kwok-Yan Lam, Alex Chichung Kot |
ICANN (10) | 4 |
| 2024 | Image Coding for Analytics via Adversarially Augmented AdaptationabstractImage Coding for Machine (ICM) aims to compress an image so that the reconstructed one can meet the requirements of both human vision and machine vision. Existing methods apply the constraint from the downstream models to improve machine analytics performance while compromising the visual quality. This paper proposes a novel adversarially augmented adaptation route that achieves a better trade-off between the utility of the human and machine perspectives by making slight changes to the image manifold. In detail, a targeted adversarial attack is employed to generate subtle image perturbations that are nearly imperceptible to humans but significantly improve machine analytic performance. These perturbed images would be subsequently employed as ground truth to guide training/fine-tuning of an end-to-end image compression network. Note that, our method is a plug-and-play framework that does not rely on any change in existing architecture or loss functions. Extensive experimental results demonstrate the superiority of the proposed scheme over conventional ICM frameworks and the effectiveness of our design. Xuelin Shen, Kangsheng Yin, Xu Wang 0006, Yu-Lin He, Shiqi Wang 0001, Wenhan Yang |
ICASSP | 6 |
| 2024 | Learned Image Compression for Both Humans and Machines via Dynamic AdaptationabstractRecent advancements in neural image compression have shown great potential in outperforming conventional standard codecs in terms of both rate-distortion and rate-analysis performance. However, there is an issue of divergent preferences in information preservation or reconstruction in the process of compression for humans and machines, respectively. Compression for humans tends to retain the signal fidelity or perceptual quality of visual appearance while compression for machines requires preserving critical semantic information, resulting in the limitation of the bitstream supporting only a single requirement during the compression. To bridge this gap, we propose a dynamic adaptation approach that generates a single bitstream serving both humans and machines. This approach aims to mitigate the domain gap among tasks, which facilitates maintaining the performance of out-of-scope tasks. Specifically, the proposed method concentrates on learning a dynamic adaptation process, i.e., optimizing the latent representation in the compressed domain in an end-to-end manner while adhering to the rate-performance constraint. Extensive results reveal that our paradigm significantly reduces the domain gap, surpassing existing codecs. Lingyu Zhu 0006, Binzhe Li, Riyu Lu, Peilin Chen 0001, Qi Mao 0002, Zhao Wang 0004, Wenhan Yang, Shiqi Wang 0001 |
ICIP | 7 |
| 2024 | Image Coding For Machine Via Analytics-Driven Appearance Redundancy ReductionabstractAmong various technical approaches in machine vision coding, Image Coding for Machine (ICM) stands out for its capability to simultaneously fulfill both human perception and machine vision needs. However, it is often criticized for its lack of efficiency regarding rate-analytics performance. In this paper, we propose an Appearance Redundancy Reduction (ARR) module, designed to function as a plug-in for existing ICM frameworks, aiming to further enhance the coding efficiency regarding rate analytics without any changes to the ICM itself. To be specific, our work pays additional attention to the intrinsic correlation between the low-level image structure and high-level vision analytics, and subsequently proposes a novel colour quantization mechanism to squeeze out the analytics-free redundant appearance information. Moreover, a differentiable soften quantization operation is derived to enable end-to-end training within the ICM framework. Extensive experimental results have shown that integrating the proposed ARR module yields substantial improvements regarding rate-analytic performance, even surpassing the performance of the feature coding paradigm, while maintaining the generalizability across different tasks and acceptable perceptual representation. Xuelin Shen, Haoqiao Ou, Wenhan Yang |
ICIP | 3 |
| 2024 | Solving Diffusion ODEs with Optimal Boundary Conditions for Better Image Super-ResolutionabstractDiffusion models, as a kind of powerful generative model, have given impressive results on image super-resolution (SR) tasks. However, due to the randomness introduced in the reverse process of diffusion models, the performances of diffusion-based SR models are fluctuating at every time of sampling, especially for samplers with few resampled steps. This inherent randomness of diffusion models results in ineffectiveness and instability, making it challenging for users to guarantee the quality of SR results. However, our work takes this randomness as an opportunity: fully analyzing and leveraging it leads to the construction of an effective plug-and-play sampling method that owns the potential to benefit a series of diffusion-based SR methods. More in detail, we propose to steadily sample high-quality SR images from pre-trained diffusion-based SR models by solving diffusion ordinary differential equations (diffusion ODEs) with optimal boundary conditions (BCs) and analyze the characteristics between the choices of BCs and their corresponding SR results. Our analysis shows the route to obtain an approximately optimal BC via an efficient exploration in the whole space. The quality of SR results sampled by the proposed method with fewer steps outperforms the quality of results sampled by current methods with randomness from the same pre-trained diffusion-based SR model, which means that our sampling method ''boosts'' current diffusion-based SR models without any additional training. Yiyang Ma, Huan Yang 0005, Wenhan Yang, Jianlong Fu, Jiaying Liu 0001 |
ICLR | 3 |
| 2024 | Correcting Diffusion-Based Perceptual Image Compression with Privileged End-to-End DecoderabstractThe images produced by diffusion models can attain excellent perceptual quality. However, it is challenging for diffusion models to guarantee distortion, hence the integration of diffusion models and image compression models still needs more comprehensive explorations. This paper presents a diffusion-based image compression method that employs a privileged end-to-end decoder model as correction, which achieves better perceptual quality while guaranteeing the distortion to an extent. We build a diffusion model and design a novel paradigm that combines the diffusion model and an end-to-end decoder, and the latter is responsible for transmitting the privileged information extracted at the encoder side. Specifically, we theoretically analyze the reconstruction process of the diffusion models at the encoder side with the original images being visible. Based on the analysis, we introduce an end-to-end convolutional decoder to provide a better approximation of the score function $\nabla_{\mathbf{x}_t}\log p(\mathbf{x}_t)$ at the encoder side and effectively transmit the combination. Experiments demonstrate the superiority of our method in both distortion and perception compared with previous perceptual compression methods. Yiyang Ma, Wenhan Yang, Jiaying Liu 0001 |
ICML | 2 |
| 2024 | Better Safe than Sorry: Pre-training CLIP against Targeted Data Poisoning and Backdoor AttacksabstractContrastive Language-Image Pre-training (CLIP) on large image-caption datasets has achieved remarkable success in zero-shot classification and enabled transferability to new domains. However, CLIP is extremely more vulnerable to targeted data poisoning and backdoor attacks compared to supervised learning. Perhaps surprisingly, poisoning 0.0001% of CLIP pre-training data is enough to make targeted data poisoning attacks successful. This is four orders of magnitude smaller than what is required to poison supervised models. Despite this vulnerability, existing methods are very limited in defending CLIP models during pre-training. In this work, we propose a strong defense, SAFECLIP, to safely pre-train CLIP against targeted data poisoning and backdoor attacks. SAFECLIP warms up the model by applying unimodal contrastive learning (CL) on image and text modalities separately. Then, it divides the data into safe and risky sets by applying a Gaussian Mixture Model to the cosine similarity of image-caption pair representations. SAFECLIP pre-trains the model by applying the CLIP loss to the safe set and applying unimodal CL to image and text modalities of the risky set separately. By gradually increasing the size of the safe set during pre-training, SAFECLIP effectively breaks targeted data poisoning and backdoor attacks without harming the CLIP performance. Our extensive experiments on CC3M, Visual Genome, and MSCOCO demonstrate that SAFECLIP significantly reduces the success rate of targeted data poisoning attacks from 93.75% to 0% and that of various backdoor attacks from up to 100% to 0%, without harming CLIP’s performance. Wenhan Yang, Jingdong Gao, Baharan Mirzasoleiman |
ICML | 1 |
| 2024 | Purify Unlearnable Examples via Rate-Constrained Variational AutoencodersabstractUnlearnable examples (UEs) seek to maximize testing error by making subtle modifications to training examples that are correctly labeled. Defenses against these poisoning attacks can be categorized based on whether specific interventions are adopted during training. The first approach is training-time defense, such as adversarial training, which can mitigate poisoning effects but is computationally intensive. The other approach is pre-training purification, e.g., image short squeezing, which consists of several simple compressions but often encounters challenges in dealing with various UEs. Our work provides a novel disentanglement mechanism to build an efficient pre-training purification method. Firstly, we uncover rate-constrained variational autoencoders (VAEs), demonstrating a clear tendency to suppress the perturbations in UEs. We subsequently conduct a theoretical analysis for this phenomenon. Building upon these insights, we introduce a disentangle variational autoencoder (D-VAE), capable of disentangling the perturbations with learnable class-wise embeddings. Based on this network, a two-stage purification approach is naturally developed. The first stage focuses on roughly eliminating perturbations, while the second stage produces refined, poison-free results, ensuring effectiveness and robustness across various scenarios. Extensive experiments demonstrate the remarkable performance of our method across CIFAR-10, CIFAR-100, and a 100-class ImageNet-subset. Code is available at https://github.com/yuyi-sd/D-VAE. Yi Yu 0011, Yufei Wang 0006, Song Xia, Wenhan Yang, Shijian Lu, Yap-Peng Tan, Alex Chichung Kot |
ICML | 4 |
| 2024 | Consistency Guided Diffusion Model with Neural Syntax for Perceptual Image CompressionabstractDiffusion models show impressive performances in image generation with excellent perceptual quality. However, its tendency to introduce additional distortion prevents its direct application in image compression. To address the issue, this paper introduces a Consistency Guided Diffusion Model (CGDM) tailored for perceptual image compression, which integrates an end-to-end image compression model with a diffusion-based post-processing network, aiming to learn richer detail representations with less fidelity loss. In detail, the compression and post-processing networks are cascaded and a branch of consistency guided features is added to constrain the deviation in the diffusion process for better reconstruction quality. Furthermore, a Syntax driven Feature Fusion (SFF) module is constructed to take an extra ultra-low bitstream from the encoding end as input, guiding the adaptive fusion of information from the two branches. In addition, we design a globally uniform boundary control strategy with overlapped patches and adopt a continuous online optimization mode to improve both coding efficiency and global consistency. Extensive experiments validate the superiority of our method to existing perceptual compression techniques. Our project is publicly available at: https://ellisonkuang.github.io/CGDM.github.io/. Haowei Kuang, Yiyang Ma, Wenhan Yang, Zongming Guo, Jiaying Liu 0001 |
ACM Multimedia | 3 |
| 2024 | ContextGS : Compact 3D Gaussian Splatting with Anchor Level Context ModelabstractRecently, 3D Gaussian Splatting (3DGS) has become a promising framework for novel view synthesis, offering fast rendering speeds and high fidelity. However, the large number of Gaussians and their associated attributes require effective compression techniques.
Existing methods primarily compress neural Gaussians individually and independently, i.e., coding all the neural Gaussians at the same time, with little design for their interactions and spatial dependence. Inspired by the effectiveness of the context model in image compression, we propose the first autoregressive model at the anchor level for 3DGS compression in this work. We divide anchors into different levels and the anchors that are not coded yet can be predicted based on the already coded ones in all the coarser levels, leading to more accurate modeling and higher coding efficiency. To further improve the efficiency of entropy coding, e.g., to code the coarsest level with no already coded anchors, we propose to introduce a low-dimensional quantized feature as the hyperprior for each anchor, which can be effectively compressed. Our work pioneers the context model in the anchor level for 3DGS representation, yielding an impressive size reduction of over 100 times compared to vanilla 3DGS and 15 times compared to the most recent state-of-the-art work Scaffold-GS, while achieving comparable or even higher rendering quality. Yufei Wang 0006, Lanqing Guo, Wenhan Yang, Alex Chichung Kot, Bihan Wen |
NeurIPS | 4 |
| 2024 | DDR: Exploiting Deep Degradation Response as Flexible Image DescriptorabstractImage deep features extracted by pre-trained networks are known to contain rich and informative representations. In this paper, we present Deep Degradation Response (DDR), a method to quantify changes in image deep features under varying degradation conditions. Specifically, our approach facilitates flexible and adaptive degradation, enabling the controlled synthesis of image degradation through text-driven prompts. Extensive evaluations demonstrate the versatility of DDR as an image descriptor, with strong correlations observed with key image attributes such as complexity, colorfulness, sharpness, and overall quality. Moreover, we demonstrate the efficacy of DDR across a spectrum of applications. It excels as a blind image quality assessment metric, outperforming existing methodologies across multiple datasets. Additionally, DDR serves as an effective unsupervised learning objective in image restoration tasks, yielding notable advancements in image deblurring and single-image super-resolution. Our code is available at: https://github.com/eezkni/DDR. Juncheng Wu, Zhangkai Ni, Hanli Wang, Wenhan Yang, Yuyin Zhou, Shiqi Wang 0001 |
NeurIPS | 4 |
| 2024 | Transferable Adversarial Attacks on SAM and Its Downstream ModelsabstractThe utilization of large foundational models has a dilemma: while fine-tuning downstream tasks from them holds promise for making use of the well-generalized knowledge in practical applications, their open accessibility also poses threats of adverse usage.
This paper, for the first time, explores the feasibility of adversarial attacking various downstream models fine-tuned from the segment anything model (SAM), by solely utilizing the information from the open-sourced SAM.
In contrast to prevailing transfer-based adversarial attacks, we demonstrate the existence of adversarial dangers even without accessing the downstream task and dataset to train a similar surrogate model.
To enhance the effectiveness of the adversarial attack towards models fine-tuned on unknown datasets, we propose a universal meta-initialization (UMI) algorithm to extract the intrinsic vulnerability inherent in the foundation model, which is then utilized as the prior knowledge to guide the generation of adversarial perturbations.
Moreover, by formulating the gradient difference in the attacking process between the open-sourced SAM and its fine-tuned downstream models, we theoretically demonstrate that a deviation occurs in the adversarial update direction by directly maximizing the distance of encoded feature embeddings in the open-sourced SAM.
Consequently, we propose a gradient robust loss that simulates the associated uncertainty with gradient-based noise augmentation to enhance the robustness of generated adversarial examples (AEs) towards this deviation, thus improving the transferability.
Extensive experiments demonstrate the effectiveness of the proposed universal meta-initialized and gradient robust adversarial attack (UMI-GRAT) toward SAMs and their downstream models.
Code is available at https://github.com/xiasong0501/GRAT. Song Xia, Wenhan Yang, Yi Yu 0011, Xun Lin, Henghui Ding, Ling-Yu Duan, Xudong Jiang 0001 |
NeurIPS | 2 |
| 2024 | Graph Contrastive Learning under Heterophily via Graph FiltersabstractGraph contrastive learning (CL) methods learn node representations in a self-supervised manner by maximizing the similarity between the augmented node representations obtained via a GNN-based encoder. However, CL methods perform poorly on graphs with heterophily, where connected nodes tend to belong to different classes. In this work, we address this problem by proposing an effective graph CL method, namely HLCL, for learning graph representations under heterophily. HLCL first identifies a homophilic and a heterophilic subgraph based on the cosine similarity of node features. It then uses a low-pass and a high-pass graph filter to aggregate representations of nodes connected in the homophilic subgraph and differentiate representations of nodes in the heterophilic subgraph. The final node representations are learned by contrasting both the augmented high-pass filtered views and the augmented low-pass filtered node views. Our extensive experiments show that HLCL outperforms state-of-the-art graph CL methods on benchmark datasets with heterophily, as well as large-scale real-world graphs, by up to 7%, and outperforms graph supervised learning methods on datasets with heterophily by up to 10%. Wenhan Yang, Baharan Mirzasoleiman |
UAI | 1 |
| 2024 | SkatingVerse: A large-scale benchmark for comprehensive evaluation on human action understandingabstractAbstract Human action understanding (HAU) is a broad topic that involves specific tasks, such as action localisation, recognition, and assessment. However, most popular HAU datasets are bound to one task based on particular actions. Combining different but relevant HAU tasks to establish a unified action understanding system is challenging due to the disparate actions across datasets. A large‐scale and comprehensive benchmark, namely SkatingVerse is constructed for action recognition, segmentation, proposal, and assessment. SkatingVerse focus on fine‐grained sport action, hence figure skating is chosen as the task object, which eliminates the biases of the object, scene, and space that exist in most previous datasets. In addition, skating actions have inherent complexity and similarity, which is an enormous challenge for current algorithms. A total of 1687 official figure skating competition videos was collected with a total of 184.4 h, exceeding four times over other datasets with a similar topic. SkatingVerse enables to formulate a unified task to output fine‐grained human action classification and assessment results from a raw figure skating competition video. In addition, SkatingVerse can facilitate the study of HAU foundation model due to its large scale and abundant categories. Moreover, image modality is incorporated for human pose estimation task into SkatingVerse . Extensive experimental results show that (1) SkatingVerse significantly helps the training and evaluation of HAU methods, (2) the performance of existing HAU methods has much room to improve, and SkatingVerse helps to reduce such gaps, and (3) unifying relevant tasks in HAU through a uniform dataset can facilitate more practical applications. SkatingVerse will be publicly available to facilitate further studies on relevant problems. Ziliang Gan, Lei Jin 0003, Yu Cheng 0009, Yinglei Teng, Zun Li 0001, Yawen Li 0001, Wenhan Yang, Junliang Xing, Jian Zhao 0006 |
IET Comput. Vis. | 8 |
| 2024 | Beyond Learned Metadata-Based Raw Image Reconstruction
Yufei Wang 0006, Yi Yu 0011, Wenhan Yang, Lanqing Guo, Lap-Pui Chau, Alex Chichung Kot, Bihan Wen |
Int. J. Comput. Vis. | 3 |
| 2024 | Temporally Consistent Enhancement of Low-Light Videos via Spatial-Temporal Compatible LearningabstractAbstract Temporal inconsistency is the annoying artifact that has been commonly introduced in low-light video enhancement, but current methods tend to overlook the significance of utilizing both data-centric clues and model-centric design to tackle this problem. In this context, our work makes a comprehensive exploration from the following three aspects. First, to enrich the scene diversity and motion flexibility, we construct a synthetic diverse low/normal-light paired video dataset with a carefully designed low-light simulation strategy, which can effectively complement existing real captured datasets. Second, for better temporal dependency utilization, we develop a Temporally Consistent Enhancer Network (TCE-Net) that consists of stacked 3D convolutions and 2D convolutions to exploit spatial-temporal clues in videos. Last, the temporal dynamic feature dependencies are exploited to obtain consistency constraints for different frame indexes. All these efforts are powered by a Spatial-Temporal Compatible Learning (STCL) optimization technique, which dynamically constructs specific training loss functions adaptively on different datasets. As such, multiple-frame information can be effectively utilized and different levels of information from the network can be feasibly integrated, thus expanding the synergies on different kinds of data and offering visually better results in terms of illumination distribution, color consistency, texture details, and temporal coherence. Extensive experimental results on various real-world low-light video datasets clearly demonstrate the proposed method achieves superior performance to state-of-the-art methods. Our code and synthesized low-light video database will be publicly available at https://github.com/lingyzhu0101/low-light-video-enhancement.git . Lingyu Zhu 0006, Wenhan Yang, Baoliang Chen, Hanwei Zhu, Xiandong Meng, Shiqi Wang 0001 |
Int. J. Comput. Vis. | 2 |
| 2024 | Towards 360$^{\circ }$ image compression for machines via modulating pixel significance
Silin Zheng, Xuelin Shen, Qiudan Zhang, Zhuo Chen 0006, Wenhan Yang, Xu Wang 0006 |
Multim. Tools Appl. | 5 |
| 2024 | Unsupervised Illumination Adaptation for Low-Light VisionabstractInsufficient lighting poses challenges to both human and machine visual analytics. While existing low-light enhancement methods prioritize human visual perception, they often neglect machine vision and high-level semantics. In this paper, we make pioneering efforts to build an illumination enhancement model for high-level vision. Drawing inspiration from camera response functions, our model could enhance images from the machine vision perspective despite being lightweight in architecture and simple in formulation. We also introduce two approaches that leverage knowledge from base enhancement curves and self-supervised pretext tasks to train for different downstream normal-to-low-light adaptation scenarios. Our proposed framework overcomes the limitations of existing algorithms without requiring access to labeled data in low-light conditions. It facilitates more effective illumination restoration and feature alignment, significantly improving the performance of downstream tasks in a plug-and-play manner. This research advances the field of low-light machine analytics and broadly applies to various high-level vision tasks, including classification, face detection, optical flow estimation, and video action recognition. Wenjing Wang 0001, Rundong Luo, Wenhan Yang, Jiaying Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Video Coding for Machines: Compact Visual Representation Compression for Intelligent Collaborative AnalyticsabstractAs an emerging research practice leveraging recent advanced AI techniques, e.g. deep models based prediction and generation, Video Coding for Machines (VCM) is committed to bridging to an extent separate research tracks of video/image compression and feature compression, and attempts to optimize compactness and efficiency jointly from a unified perspective of high accuracy machine vision and full fidelity human vision. With the rapid advances of deep feature representation and visual data compression in mind, in this paper, we summarize VCM methodology and philosophy based on existing academia and industrial efforts. The development of VCM follows a general rate-distortion optimization, and the categorization of key modules or techniques is established including feature-assisted coding, scalable coding, intermediate feature compression/optimization, and machine vision targeted codec, from broader perspectives of vision tasks, analytics resources, etc. From previous works, it is demonstrated that, although existing works attempt to reveal the nature of scalable representation in bits when dealing with machine and human vision tasks, there remains a rare study in the generality of low bit rate representation, and accordingly how to support a variety of visual analytic tasks. Therefore, we investigate a novel visual information compression for the analytics taxonomy problem to strengthen the capability of compact visual representations extracted from multiple tasks for visual analytics. A new perspective of task relationships versus compression is revisited. By keeping in mind the transferability among different machine vision tasks (e.g. high-level semantic and mid-level geometry-related), we aim to support multiple tasks jointly at low bit rates. In particular, to narrow the dimensionality gap between neural network generated features extracted from pixels and a variety of machine vision features/labels (e.g. scene class, segmentation labels), a codebook hyperprior is designed to compress the neural network-generated features. As demonstrated in our experiments, this new hyperprior model is expected to improve feature compression efficiency by estimating the signal entropy more accurately, which enables further investigation of the granularity of abstracting compact features among different tasks. Wenhan Yang, Haofeng Huang, Yueyu Hu, Ling-Yu Duan, Jiaying Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | Learning to Remove Rain in Video With Self-SupervisionabstractIn heavy rain video, rain streak and rain accumulation are the most common causes of degradation. They occlude background information and can significantly impair the visibility. Most existing methods rely heavily on the synthetic training data, and thus raise the domain gap problem that prevents the trained models from performing adequately in real testing cases. Unlike these methods, we introduce a self-learning method to remove both rain streaks and rain accumulation without using any ground-truth clean images in training our model, which consequently can alleviate the domain gap issue. The main idea is based on the assumptions that (1) adjacent clean frames can be aligned or warped from one frame to another frame, (2) rain streaks are distributed randomly in the temporal domain, (3) the rain streak/accumulation related variables/priors can be inferred reliably from the information within the images/sequences. Based on these assumptions, we construct an augmented Self-Learned Deraining Network (SLDNet+) to remove both rain streaks and rain accumulation by utilizing temporal correlation, consistency, and rain-related priors. For the temporal correlation, our SLDNet+ takes rain degraded adjacent frames as its input, aligns them, and learns to predict the clean version of the current frame. For the temporal consistency, a new loss is designed to build a robust mapping between the predicted clean frame and non-rain regions from the adjacent rain frames. For the rain-streak-related prior, the rain streak removal network is optimized jointly with motion estimation and rain region detection; while for the rain-accumulation-related prior, a novel non-local video rain accumulation removal method is developed to estimate the accumulation-lines from the whole input video and to offer better color constancy and temporal smoothness. Extensive experiments show the effectiveness of our approach, which provides superior results compared with the existing state of the art methods both quantitatively and qualitatively. The source code will be made publicly available at: https://github.com/flyywh/CVPR-2020-Self-Rain-Removal-Journal. Wenhan Yang, Robby T. Tan, Shiqi Wang 0001, Alex Chichung Kot, Jiaying Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | Cross-Image Disentanglement for Low-Light Enhancement in Real WorldabstractImages captured in the low-light condition suffer from low visibility and various imaging artifacts, e.g., real noise. Existing supervised algorithms for low-light image enhancement require a large set of pixel-aligned training image pairs, which are hard to prepare in practice. Though some recent unsupervised methods can alleviate such data challenges, many real world artifacts inevitably get falsely amplified in the enhanced results due to the lack of corresponding supervision. In this paper, instead of using perfectly aligned images for training, we creatively employ the misaligned real world images as the guidance, which are considerably easier to collect. Specifically, we propose a Cross-Image Disentanglement Network (CIDN) with weakly supervised learning, to separately extract cross-image brightness and image-specific content features from low/normal-light images. Based on that, CIDN can simultaneously correct the brightness and suppress image artifacts in the feature domain, which largely increases the robustness of the pixel shifts between training pairs. By considering real world corruptions, we propose a new training dataset with misaligned and noisy image pairs and its corresponding evaluation dataset. Experimental results show that our model achieves state-of-the-art performances on both the newly proposed dataset and other popular low-light datasets. The code implementation is publicly available at:https://github.com/GuoLanqing/CIDN. Lanqing Guo, Renjie Wan, Wenhan Yang, Alex Chichung Kot, Bihan Wen |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | Benchmarking Joint Face Spoofing and Forgery Detection With Visual and Physiological CuesabstractFace anti-spoofing (FAS) and face forgery detection play vital roles in securing face biometric systems from presentation attacks (PAs) and vicious digital manipulation (e.g., deepfakes). Despite satisfactory performance upon large-scale data and powerful deep models, recent advances in face spoofing and forgery detection approaches usually focus on 1) unimodal visual appearance or physiological (i.e., remote photoplethysmography (rPPG)) cues; and 2) separated feature representation for FAS or face forgery detection. On one side, unimodal appearance and rPPG features are respectively vulnerable to high-fidelity face 3D mask and video replay attacks, inspiring us to design reliable multi-modal fusion mechanisms for generalized FAS. On the other side, there are rich common features across FAS and face forgery detection tasks (e.g., periodic rPPG rhythms and vanilla appearance for bonafides), providing solid evidence to design a joint FAS and face forgery detection system in a multi-task learning fashion. In this paper, we establish the first joint face spoofing and forgery detection benchmark using both visual appearance and physiological rPPG cues. To enhance the rPPG periodicity discrimination, we design a two-branch physiological network using both facial spatio-temporal rPPG signal map and its continuous wavelet transformed counterpart as inputs. To mitigate the modality bias and improve the fusion efficacy, we conduct a weighted batch and layer normalization for both appearance and rPPG features before multi-modal fusion. We also investigate prevalent deep models, feature fusion strategies and multi-task learning configurations for joint face spoofing and forgery detection. We find that the generalization capacities of both unimodal (appearance or rPPG) and multi-modal (appearance+rPPG) models can be obviously improved via joint training on these two tasks. We hope this new benchmark will facilitate the future research of both FAS and deepfake detection communities. The codes will be released athttps://github.com/ZitongYu/Benchmarking. Zitong Yu, Rizhao Cai, Zhi Li 0054, Wenhan Yang, Jingang Shi, Alex Chichung Kot |
IEEE Trans. Dependable Secur. Comput. | 4 |
| 2024 | Diffusion Enhancement for Cloud Removal in Ultra-Resolution Remote Sensing ImageryabstractThe presence of cloud layers severely compromises the quality and effectiveness of optical remote sensing (RS) images. However, existing deep-learning (DL)-based cloud removal (CR) techniques, which usually take the fidelity-driven losses as constraints, e.g.,$L_{1}$or$L_{2}$losses, tend to generate smooth results, often failing to reconstruct visually pleasing results and cause semantic loss. To tackle this challenge, this work proposes to encompass enhancements at the data and methodology fronts. On the data side, an ultra-resolution benchmark named CUHK cloud removal (CUHK-CR) of 0.5 m spatial resolution is established. This benchmark incorporates rich detailed textures and diverse cloud coverage, serving as a robust foundation for designing and assessing CR models. From the methodology perspective, a novel diffusion-based framework for CR named diffusion enhancement (DE) is introduced. This framework aims to gradually recover texture details, leveraging a reference visual prior providing foundational structure of the images to enhance inference accuracy. Additionally, a weight allocation (WA) network is developed to dynamically adjust the weights for feature fusion, thereby further improving performance, particularly in the context of ultra-resolution image generation. Furthermore, a coarse-to-fine training strategy is applied to effectively expedite training convergence while reducing the computational complexity required to handle ultra-resolution images. Extensive experiments on the newly established CUHK-CR and existing datasets such as RICE confirm that the proposed DE framework outperforms existing DL-based methods in terms of both perceptual quality and signal fidelity. Jialu Sui, Yiyang Ma, Wenhan Yang, Man-On Pun, Jiaying Liu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Toward Real-World Super Resolution With Adaptive Self-Similarity MiningabstractDespite efforts to construct super-resolution (SR) training datasets with a wide range of degradation scenarios, existing supervised methods based on these datasets still struggle to consistently offer promising results due to the diversity of real-world degradation scenarios and the inherent complexity of model learning. Our work explores a new route: integrating the sample-adaptive property learned through image intrinsic self-similarity and the universal knowledge acquired from large-scale data. We achieve this by uniting internal learning and external learning by an unrolled optimization process. With the merits of both, the tuned fully-supervised SR models can be augmented to broadly handle the real-world degradation in a plug-and-play style. Furthermore, to promote the efficiency of combining internal/external learning, we apply an attention-based weight-updating method to guide the mining of self-similarity, and various data augmentations are adopted while applying the exponential moving average strategy. We conduct extensive experiments on real-world degraded images and our approach outperforms other methods in both qualitative and quantitative comparisons. Our project is available at: https://github.com/ZahraFan/AdaSSR/. Zejia Fan, Wenhan Yang, Zongming Guo, Jiaying Liu 0001 |
IEEE Trans. Image Process. | 2 |
| 2024 | UCL-Dehaze: Toward Real-World Image Dehazing via Unsupervised Contrastive LearningabstractWhile the wisdom of training an image dehazing model on synthetic hazy data can alleviate the difficulty of collecting real-world hazy/clean image pairs, it brings the well-known domain shift problem. From a different yet new perspective, this paper explores contrastive learning with an adversarial training effort to leverage unpaired real-world hazy and clean images, thus alleviating the domain shift problem and enhancing the network's generalization ability in real-world scenarios. We propose an effective unsupervised contrastive learning paradigm for image dehazing, dubbed UCL-Dehaze. Unpaired real-world clean and hazy images are easily captured, and will serve as the important positive and negative samples respectively when training our UCL-Dehaze network. To train the network more effectively, we formulate a new self-contrastive perceptual loss function, which encourages the restored images to approach the positive samples and keep away from the negative samples in the embedding space. Besides the overall network architecture of UCL-Dehaze, adversarial training is utilized to align the distributions between the positive samples and the dehazed images. Compared with recent image dehazing works, UCL-Dehaze does not require paired data during training and utilizes unpaired positive/negative data to better enhance the dehazing performance. We conduct comprehensive experiments to evaluate our UCL-Dehaze and demonstrate its superiority over the state-of-the-arts, even only 1,800 unpaired real-world images are used to train our network. Source code is publicly available at https://github.com/yz-wang/UCL-Dehaze. Yongzhen Wang 0001, Xuefeng Yan 0001, Fu Lee Wang, Haoran Xie 0001, Wenhan Yang, Xiao-Ping Zhang 0002, Harry Qin, Mingqiang Wei |
IEEE Trans. Image Process. | 5 |
| 2024 | Gap-Closing Matters: Perceptual Quality Evaluation and Optimization of Low-Light Image EnhancementabstractThere is a growing consensus in the research community that the optimization of low-light image enhancement approaches should be guided by the visual quality perceived by end users. Despite the substantial efforts invested in the design of low-light enhancement algorithms, there has been comparatively limited focus on assessing subjective and objective quality systematically. To mitigate this gap and provide a clear path towards optimizing low-light image enhancement for better visual quality, we propose a gap-closing framework. In particular, our gap-closing framework starts with the creation of a large-scale dataset for Subjective QUality Assessment of REconstructed LOw-Light Images (SQUARE-LOL). This database serves as the foundation for studying the quality of enhanced images and conducting a comprehensive subjective user study. Subsequently, we propose an objective quality assessment measure that plays a critical role in bridging the gap between visual quality and enhancement. Finally, we demonstrate that our proposed objective quality measure can be incorporated into the process of optimizing the learning of the enhancement model toward perceptual optimality. We validate the effectiveness of our proposed framework through both the accuracy of quality prediction and the perceptual quality of image enhancement. Baoliang Chen, Lingyu Zhu 0006, Hanwei Zhu, Wenhan Yang, Linqi Song, Shiqi Wang 0001 |
IEEE Trans. Multim. | 4 |
| 2024 | Opinion-Unaware Blind Image Quality Assessment Using Multi-Scale Deep Feature StatisticsabstractDeep learning-based methods have significantly influenced the blind image quality assessment (BIQA) field, however, these methods often require training using large amounts of human rating data. In contrast, traditional knowledge-based methods are cost-effective for training but face challenges in effectively extracting features aligned with human visual perception. To bridge these gaps, we propose integrating deep features from pre-trained visual models with a statistical analysis model into a Multi-scale Deep Feature Statistics (MDFS) model for achieving opinion-unaware BIQA (OU-BIQA), thereby eliminating the reliance on human rating data and significantly improving training efficiency. Specifically, we extract patch-wise multi-scale features from pre-trained vision models, which are subsequently fitted into a multivariate Gaussian (MVG) model. The final quality score is determined by quantifying the distance between the MVG model derived from the test image and the benchmark MVG model derived from the high-quality image set. A comprehensive series of experiments conducted on various datasets show that our proposed model exhibits superior consistency with human visual perception compared to state-of-the-art BIQA models. Furthermore, it shows improved generalizability across diverse target-specific BIQA tasks. Our code is available at:https://github.com/eezkni/MDFS Zhangkai Ni, Keyan Ding, Wenhan Yang, Hanli Wang, Shiqi Wang 0001 |
IEEE Trans. Multim. | 4 |
| 2024 | Glow in the Dark: Low-Light Image Enhancement With External MemoryabstractDeep learning-based methods have achieved remarkable success with powerful modeling capabilities. However, the weights of these models are learned over the entire training dataset, which inevitably leads to the ignorance of sample specific properties in the learned enhancement mapping. This situation causes ineffective enhancement in the testing phase for the samples that differ significantly from the training distribution. In this paper, we introduce external memory to form an external memory-augmented network (EMNet) for low-light image enhancement. The external memory aims to capture the sample specific properties of the training dataset to guide the enhancement in the testing phase. Benefiting from the learned memory, more complex distributions of reference images in the entire dataset can be “remembered” to facilitate the adjustment of the testing samples more adaptively. To further augment the capacity of the model, we take the transformer as our baseline network, which specializes in capturing long-range spatial redundancy. Experimental results demonstrate that our proposed method has a promising performance and outperforms state-of-the-art methods. It is noted that, the proposed external memory is a plug-and-play mechanism that can be integrated with any existing method to further improve the enhancement quality. More practices of integrating external memory with other image enhancement methods are qualitatively and quantitatively analyzed. The results further confirm that the effectiveness of our proposed memory mechanism when combing with existing enhancement methods. Dongjie Ye, Zhangkai Ni, Wenhan Yang, Hanli Wang, Shiqi Wang 0001, Sam Kwong |
IEEE Trans. Multim. | 3 |
| 2023 | Cross-Modal Label Contrastive Learning for Unsupervised Audio-Visual Event LocalizationabstractThis paper for the first time explores audio-visual event localization in an unsupervised manner. Previous methods tackle this problem in a supervised setting and require segment-level or video-level event category ground-truth to train the model. However, building large-scale multi-modality datasets with category annotations is human-intensive and thus not scalable to real-world applications. To this end, we propose cross-modal label contrastive learning to exploit multi-modal information among unlabeled audio and visual streams as self-supervision signals. At the feature representation level, multi-modal representations are collaboratively learned from audio and visual components by using self-supervised representation learning. At the label level, we propose a novel self-supervised pretext task i.e. label contrasting to self-annotate videos with pseudo-labels for localization model training. Note that irrelevant background would hinder the acquisition of high-quality pseudo-labels and thus lead to an inferior localization model. To address this issue, we then propose an expectation-maximization algorithm that optimizes the pseudo-label acquisition and localization model in a coarse-to-fine manner. Extensive experiments demonstrate that our unsupervised approach performs reasonably well compared to the state-of-the-art supervised methods. Peijun Bao, Wenhan Yang, Boon Poh Ng, Meng Hwa Er, Alex Chichung Kot |
AAAI | 2 |
| 2023 | Estimating Reflectance Layer from a Single Image: Integrating Reflectance Guidance and Shadow/Specular Aware LearningabstractEstimating the reflectance layer from a single image is a challenging task. It becomes more challenging when the input image contains shadows or specular highlights, which often render an inaccurate estimate of the reflectance layer. Therefore, we propose a two-stage learning method, including reflectance guidance and a Shadow/Specular-Aware (S-Aware) network to tackle the problem. In the first stage, an initial reflectance layer free from shadows and specularities is obtained with the constraint of novel losses that are guided by prior-based shadow-free and specular-free images. To further enforce the reflectance layer to be independent of shadows and specularities in the second-stage refinement, we introduce an S-Aware network that distinguishes the reflectance image from the input image. Our network employs a classifier to categorize shadow/shadow-free, specular/specular-free classes, enabling the activation features to function as attention maps that focus on shadow/specular regions. Our quantitative and qualitative evaluations show that our method outperforms the state-of-the-art methods in the reflectance layer estimation that is free from shadows and specularities. Yeying Jin, Ruoteng Li, Wenhan Yang, Robby T. Tan |
AAAI | 3 |
| 2023 | ShadowDiffusion: When Degradation Prior Meets Diffusion Model for Shadow RemovalabstractRecent deep learning methods have achieved promising results in image shadow removal. However, their restored images still suffer from unsatisfactory boundary artifacts, due to the lack of degradation prior embedding and the deficiency in modeling capacity. Our work addresses these issues by proposing a unified diffusion framework that integrates both the image and degradation priors for highly effective shadow removal. In detail, we first propose a shadow degradation model, which inspires us to build a novel unrolling diffusion model, dubbed ShandowDiffusion. It remarkably improves the model's capacity in shadow removal via progressively refining the desired output with both degradation prior and diffusive generative prior, which by nature can serve as a new strong baseline for image restoration. Furthermore, ShadowDiffusion progressively refines the estimated shadow mask as an auxiliary task of the diffusion generator, which leads to more accurate and robust shadow-free image generation. We conduct extensive experiments on three popular public datasets, including ISTD, ISTD+, and SRD, to validate our method's effectiveness. Compared to the state-of-the-art methods, our model achieves a significant improvement in terms of PSNR, increasing from 31.69dB to 34. 73dB over SRD dataset.11https://github.com/GuoLanqing/ShadowDiffusion Lanqing Guo, Chong Wang 0011, Wenhan Yang, Siyu Huang, Yufei Wang 0006, Hanspeter Pfister, Bihan Wen |
CVPR | 3 |
| 2023 | Raw Image Reconstruction with Learned Compact MetadataabstractWhile raw images exhibit advantages over sRGB images (e.g., linearity and fine-grained quantization level), they are not widely used by common users due to the large storage requirements. Very recent works propose to compress raw images by designing the sampling masks in the raw image pixel space, leading to suboptimal image representations and redundant metadata. In this paper, we propose a novel framework to learn a compact representation in the latent space serving as the metadata in an end-to-end manner. Furthermore, we propose a novel sRGB-guided context model with the improved entropy estimation strategies, which leads to better reconstruction quality, smaller size of metadata, and faster speed. We illustrate how the proposed raw image compression scheme can adaptively allocate more bits to image regions that are important from a global perspective. The experimental results show that the proposed method can achieve superior raw image reconstruction results using a smaller size of the metadata on both uncompressed sRGB images and JPEG images. The code will be released at https://github.com/wyf0912/R2LCM Yufei Wang 0006, Yi Yu 0011, Wenhan Yang, Lanqing Guo, Lap-Pui Chau, Alex Chichung Kot, Bihan Wen |
CVPR | 3 |
| 2023 | Backdoor Attacks Against Deep Image Compression via Adaptive Frequency TriggerabstractRecent deep-learning-based compression methods have achieved superior performance compared with traditional approaches. However, deep learning models have proven to be vulnerable to backdoor attacks, where some specific trigger patterns added to the input can lead to malicious behavior of the models. In this paper, we present a novel backdoor attack with multiple triggers against learned image compression models. Motivated by the widely used discrete cosine transform (DCT) in existing compression systems and standards, we propose a frequency-based trigger injection model that adds triggers in the DCT domain. In particular, we design several attack objectives for various attacking scenarios, including: 1) attacking compression quality in terms of bit-rate and reconstruction quality; 2) attacking task-driven measures, such as downstream face recognition and semantic segmentation. Moreover, a novel simple dynamic loss is designed to balance the influence of different loss terms adaptively, which helps achieve more efficient training. Extensive experiments show that with our trained trigger injection models and simple modification of encoder parameters (of the compression model), the proposed attack can successfully inject several backdoors with corresponding triggers in a single image compression model. Yi Yu 0011, Yufei Wang 0006, Wenhan Yang, Shijian Lu, Yap-Peng Tan, Alex Chichung Kot |
CVPR | 3 |
| 2023 | Boundary-Aware Divide and Conquer: A Diffusion-based Solution for Unsupervised Shadow RemovalabstractRecent deep learning methods have achieved superior results in shadow removal. However, most of these supervised methods rely on training over a huge amount of shadow and shadow-free image pairs, which require laborious annotations and may end up with poor model generalization. Shadows, in fact, only form partial degradation in images, while their non-shadow regions provide rich structural information potentially for unsupervised learning. In this paper, we propose a novel diffusion-based solution for unsupervised shadow removal, which separately modeling the shadow, non-shadow, and their boundary regions. We employ a pretrained unconditional diffusion model fused with non-corrupted information to generate the natural shadow-free image. While the diffusion model can restore the clear structure in the boundary region by utilizing its adjacent non-corrupted contextual information, it fails to address the inner shadow area due to the isolation of the non-corrupted contexts. Thus we further propose a Shadow-Invariant Intrinsic Decomposition module to exploit the underlying reflectance in the shadow region to maintain structural consistency during the diffusive sampling. Extensive experiments on the publicly available shadow removal datasets show that the proposed method achieves a significant improvement compared to existing unsupervised methods, and even is comparable with some existing supervised methods. Lanqing Guo, Chong Wang 0011, Wenhan Yang, Yufei Wang 0006, Bihan Wen |
ICCV | 3 |
| 2023 | Similarity Min-Max: Zero-Shot Day-Night Domain AdaptationabstractLow-light conditions not only hamper human visual experience but also degrade the model’s performance on downstream vision tasks. While existing works make remarkable progress on day-night domain adaptation, they rely heavily on domain knowledge derived from the task-specific nighttime dataset. This paper challenges a more complicated scenario with border applicability, i.e., zero-shot day-night domain adaptation, which eliminates reliance on any nighttime data. Unlike prior zero-shot adaptation approaches emphasizing either image-level translation or model-level adaptation, we propose a similarity min-max paradigm that considers them under a unified framework. On the image level, we darken images towards minimum feature similarity to enlarge the domain gap. Then on the model level, we maximize the feature similarity between the darkened images and their normal-light counterparts for better model adaptation. To the best of our knowledge, this work represents the pioneering effort in jointly optimizing both levels, resulting in a significant improvement of model generalizability. Extensive experiments demonstrate our method’s effectiveness and broad applicability on various nighttime vision tasks, including classification, semantic segmentation, visual place recognition, and video action recognition. Our project page is available at https://red-fairy.github.io/ZeroShotDayNightDA-Webpage/ Rundong Luo, Wenjing Wang 0001, Wenhan Yang, Jiaying Liu 0001 |
ICCV | 3 |
| 2023 | ExposureDiffusion: Learning to Expose for Low-light Image EnhancementabstractPrevious raw image-based low-light image enhancement methods predominantly relied on feed-forward neural networks to learn deterministic mappings from low-light to normally-exposed images. However, they failed to capture critical distribution information, leading to visually undesirable results. This work addresses the issue by seamlessly integrating a diffusion model with a physics-based exposure model. Different from a vanilla diffusion model that has to perform Gaussian denoising, with the injected physics-based exposure model, our restoration process can directly start from a noisy image instead of pure noise. As such, our method obtains significantly improved performance and reduced inference time compared with vanilla diffusion models. To make full use of the advantages of different intermediate steps, we further propose an adaptive residual layer that effectively screens out the side-effect in the iterative refinement when the intermediate results have been already well-exposed. The proposed framework can work with both real-paired datasets, SOTA noise models, and different backbone networks. We evaluate the proposed method on various public benchmarks, achieving promising results with consistent improvements using different exposure models and backbones. Besides, the proposed method achieves better generalization capacity for unseen amplifying ratios and better performance than a larger feedforward neural model when few parameters are adopted. The code is released at https://github.com/wyf0912/ExposureDiffusion. Yufei Wang 0006, Yi Yu 0011, Wenhan Yang, Lanqing Guo, Lap-Pui Chau, Alex Chichung Kot, Bihan Wen |
ICCV | 3 |
| 2023 | Flash Compensated Low-Light Enhancement Via Hierarchical Network PredictionabstractPhotography in low-light conditions suffers from dense noise and insufficient light. Flash photography, introducing extra light sources, performs better at suppressing noise and revealing details, while being interrupted by unnatural ambient illumination. This paper offers an analysis of the pros and cons to utilize low-light and flash images for enhancement, which inspires us to design a unified sample-adaptive CNN to capture diverse focuses from different inputs in a complementary way. Specifically, a Flash Compensated Dynamic Filtering Network is proposed to utilize the revealed details of flash images to compensate for fine structure reconstruction in low-light enhancement. To adaptively fuse information from misaligned low-light and flash image pairs, our network is designed with three distinctive features. Firstly, we adopt a layer-wise regression strategy, where results are predicted from the single input first and then fused to sufficiently leverage complementary information. Secondly, we employ a sample-adaptive mechanism, where each pixel is estimated with its distinctive parameters augmented by weighted residual connections. Finally, we utilize a coarse-to-fine architecture, where features are extracted by diversified receptive fields to utilize hierarchical contextual information. Experimental results demonstrate that the three design principles lead to the significant superiority of the proposed method over state-of-the-art methods. Haowei Kuang, Haofeng Huang, Wenhan Yang, Jiaying Liu 0001 |
ICIP | 3 |
| 2023 | Enhancing Low-Light Images Using Infrared Encoded ImagesabstractLow-light image enhancement task is essential yet challenging as it is ill-posed intrinsically. Previous arts mainly focus on the low-light images captured in the visible spectrum using pixel-wise loss, which limits the capacity of recovering the brightness, contrast, and texture details due to the small number of income photons. In this work, we propose a novel approach to increase the visibility of images captured under low-light environments by removing the in-camera infrared (IR) cut-off filter, which allows for the capture of more photons and results in improved signal-to-noise ratio due to the inclusion of information from the IR spectrum. To verify the proposed strategy, we collect a paired dataset of low-light images captured without the IR cut-off filter, with corresponding long-exposure reference images with an external filter. The experimental results on the proposed dataset demonstrate the effectiveness of the proposed method, showing better performance quantitatively and qualitatively. The dataset and code are publicly available at https://wyf0912.github.io/ELIEI/ Shulin Tian, Yufei Wang 0006, Renjie Wan, Wenhan Yang, Alex Chichung Kot, Bihan Wen |
ICIP | 4 |
| 2023 | Modality Meets Long-Term Tracker: A Siamese Dual Fusion Framework for Tracking UAVabstractTracking an Unmanned Aerial Vehicle (UAV) to obtain its locations and trajectory is a crucial task to avoid the unlawful use of UAVs. However, most existing UAV tracking methods fail when facing cluster environments, out-of-view, and occlusions because of their insufficient representation of global context information capacity. To mitigate these issues, we propose a new tracker, namely SiamFusion, to innovate a dual fusion procedure that leverages the advantages in both the feature and decision levels. In particular, we propose a novel feature fusion module named Modality-Fusion to utilize multi-modal information, enhancing the perception of the target. From the decision level, we further develop a local-global converter based on a multi-modal fusion decision-making mechanism to reduce the accumulation during tracking, which significantly increases the robustness of the tracking process. Extensive experiments demonstrate the superiority of the proposed SiamFusion, which achieves the best performance on Anti-UAV in terms of accuracy and speed. In particular, we exceed the state-of-the-art tracking algorithm in the tracking accuracy by 4.2% at a similar frame rate. Our source codes, pre-trained models, and online demos will be released upon acceptance. Lei Jin 0003, Shengjie Li 0003, Jianqiang Xia, Jun Wang 0041, Zun Li 0001, Wenhan Yang, Pengfei Zhang 0016, Jian Zhao 0006, Bo Zhang 0007 |
ICIP | 8 |
| 2023 | Collaborative Spatial-Temporal Distillation for Efficient Video DerainingabstractIn this paper, we propose a novel knowledge distillation framework to improve the efficiency of deep networks for video deraining. The knowledge is transferred from a large-scale powerful teacher network to a compact efficient student network via the proposed collaborative spatial-temporal distillation framework. The framework is equipped with three collaboration schemes of different granularities that make use of spatial-temporal redundancy in a complementary way for better distillation performance. First, the spatial alignment module applies distillation constraints at different spatial scales to achieve better scale invariance in transferred knowledge. Second, the temporal alignment module traces both temporal status between teacher and student separately and collaboratively, to comprehensively utilize inter-frame information. Third, these two alignment modules interact through a spatial-temporal adaptor, where spatial-temporal knowledge is transferred in a unified framework. Extensive experiments demonstrate the superiority of our distillation framework as well as the effectiveness of each module. Our code is available at: https://github.com/HuYuzhang/Knowledge-Distillation. Yuzhang Hu, Minghao Liu 0019, Wenhan Yang, Jiaying Liu 0001, Zongming Guo |
ICME | 3 |
| 2023 | Removing Image Artifacts From Scratched Lens ProtectorsabstractA protector is placed in front of the camera lens for mobile devices to avoid damage, while the protector itself can be easily scratched accidentally, especially for plastic ones. The artifacts appear in a wide variety of patterns, making it difficult to see through them clearly. Removing image artifacts from the scratched lens protector is inherently challenging due to the occasional flare artifacts and the co-occurring interference within mixed artifacts. Though different methods have been proposed for some specific distortions, they seldom consider such inherent challenges. In our work, we consider the inherent challenges in a unified framework with two cooperative modules, which facilitate the performance boost of each other. We also collect a new dataset from the real world to facilitate training and evaluation purposes. The experimental results demonstrate that our method outperforms the baselines qualitatively and quantitatively. The code and datasets will be released at https://github.com/wyf0912/flare-removal Yufei Wang 0006, Renjie Wan, Wenhan Yang, Bihan Wen, Lap-Pui Chau, Alex Chichung Kot |
ISCAS | 3 |
| 2023 | Robust Contrastive Language-Image Pretraining against Data Poisoning and Backdoor AttacksabstractContrastive vision-language representation learning has achieved state-of-the-art performance for zero-shot classification, by learning from millions of image-caption pairs crawled from the internet. However, the massive data that powers large multimodal models such as CLIP, makes them extremely vulnerable to various types of targeted data poisoning and backdoor attacks. Despite this vulnerability, robust contrastive vision-language pre-training against such attacks has remained unaddressed. In this work, we propose RoCLIP, the first effective method for robust pre-training multimodal vision-language models against targeted data poisoning and backdoor attacks. RoCLIP effectively breaks the association between poisoned image-caption pairs by considering a relatively large and varying pool of random captions, and matching every image with the text that is most similar to it in the pool instead of its own caption, every few epochs.It also leverages image and text augmentations to further strengthen the defense and improve the performance of the model. Our extensive experiments show that RoCLIP renders state-of-the-art targeted data poisoning and backdoor attacks ineffective during pre-training CLIP models. In particular, RoCLIP decreases the success rate for targeted data poisoning attacks from 93.75% to 12.5% and that of backdoor attacks down to 0%, while improving the model's linear probe performance by 10% and maintains a similar zero shot performance compared to CLIP. By increasing the frequency of matching, RoCLIP is able to defend strong attacks, which add up to 1% poisoned examples to the data, and successfully maintain a low attack success rate of 12.5%, while trading off the performance on some tasks. Wenhan Yang, Jingdong Gao, Baharan Mirzasoleiman |
NeurIPS | 1 |
| 2023 | Unsupervised Face Detection in the DarkabstractLow-light face detection is challenging but critical for real-world applications, such as nighttime autonomous driving and city surveillance. Current face detection models rely on extensive annotations and lack generality and flexibility. In this paper, we explore how to learn face detectors without low-light annotations. Fully exploiting existing normal light data, we propose adapting face detectors from normal light to low light. This task is difficult because the gap between brightness and darkness is too large and complicated at the object level and pixel level. Accordingly, the performance of current low-light enhancement or adaptation methods is unsatisfactory. To solve this problem, we propose a joint High-Low Adaptation (HLA) framework. We design bidirectional low-level adaptation and multitask high-level adaptation. For low-level, we enhance the dark images and degrade the normal-light images, making both domains move toward each other. For high-level, we combine context-based and contrastive learning to comprehensively close the features on different domains. Experiments show that our HLA-Face v2 model obtains superior low-light face detection performance even without the use of low-light annotations. Moreover, our adaptation scheme can be extended to a wide range of applications, such as improving supervised learning and generic object detection. Project publicly available at: https://daooshee.github.io/HLA-Face-v2-Website/. Wenjing Wang 0001, Wenhan Yang, Jiaying Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2023 | Automatic Model-Based Dataset Generation for High-Level Vision Tasks of Autonomous Driving in Haze WeatherabstractImproving the performance of high-level computer vision tasks in adverse weather (e.g., haze) is highly critical for autonomous driving safety. However, collecting and annotating training sets for various high-level tasks in haze weather are expensive and time-consuming. To address this issue, we propose a novel haze generation model called HazeGEN by coupling the variational autoencoder and the generative adversarial network to automatically generate annotated datasets. The proposed HazeGEN leverages a shared latent space assumption based on an optimized encoder–decoder architecture, which guarantees high fidelity in the cross-domain image translations. To ensure that the generated image can truly facilitate high-level vision task performance, a semisupervised learning strategy is developed for HazeGEN to efficiently learn the useful knowledge from both the real-world images (with unsupervised losses) and the synthetic images generated following the atmosphere scattering model (with supervised losses). Extensive experiments and ablation studies demonstrate that training the model with our generated haze dataset greatly improves accuracy in high-level tasks such as semantic segmentation and object detection. Furthermore, one important but under-exploited issue is investigated to find out whether the developed dataset can be a good substitute for the real ones. Results show that the generated dataset has the most similar performance to the real-world collected haze dataset on multiple challenging industrial scenarios compared with prior works. Tianqi Su, Siyi Chen 0004, Wenhan Yang, Jiaying Liu 0001, Zhongfeng Wang 0001 |
IEEE Trans. Ind. Informatics | 4 |
| 2023 | Purifying Low-Light Images via Near-Infrared Enlightened ImageabstractCameras usually produce low-quality images under low-light conditions. Though many methods have been proposed to enhance the visibility of low-light images, they are mainly designed for illumination correction and less capable of sup-pressing the artifacts. In this paper, we propose to enhance the visibility and suppress artifacts by purifying low-light images under the guidance of the NIR enlightened image captured by using the near-infrared light as compensation. Specifically, we introduce a disentanglement framework to disentangle the structure and color components from the NIR enlightened and RGB images, respectively. Correspondingly, we introduce a new dataset with the RGB and NIR enlightened images for training and evaluation purposes. The experimental results show that our proposed method achieves promising results. Renjie Wan, Boxin Shi, Wenhan Yang, Bihan Wen, Ling-Yu Duan, Alex Chichung Kot |
IEEE Trans. Multim. | 3 |
| 2023 | Deep Inter Prediction with Error-Corrected Auto-Regressive Network for Video CodingabstractModern codecs remove temporal redundancy of a video via inter prediction, i.e., searching previously coded frames for similar blocks and storing motion vectors to save bit-rates. However, existing codecs adopt block-level motion estimation, where a block is regressed by reference blocks linearly and is doomed to fail to deal with non-linear motions. In this article, we generate virtual reference frames (VRFs) with previously reconstructed frames via deep networks to offer an additional candidate, which is not constrained to linear motion structure and further significantly improves coding efficiency. More specifically, we propose a novel deep Auto-Regressive Moving-Average (ARMA) model, Error-Corrected Auto-Regressive Network (ECAR-Net), equipped with the powers of the conventional statistic ARMA models and deep networks jointly for reference frame prediction. Similar to conventional ARMA models, the ECAR-Net consists of two stages: Auto-Regression (AR) stage and Error-Correction (EC) stage, where the first part predicts the signal at the current time-step based on previously reconstructed frames, while the second one compensates for the output of the AR stage to obtain finer details. Different from the statistic AR models only focusing on short-term temporal dependency, the AR model of our ECAR-Net is further injected with the long-term dynamics mechanism, where long temporal information is utilized to help predict motions more accurately. Furthermore, ECAR-Net works in a configuration-adaptive way, i.e., using different dynamics and error definitions for the Low Delay B and Random Access configurations, which helps improve the adaptivity and generality in diverse coding scenarios. With the well-designed network, our method surpasses HEVC on average 5.0% and 6.6% BD-rate saving for the luma component under the Low Delay B and Random Access configurations and also obtains on average 1.54% BD-rate saving over VVC. Furthermore, ECAR-Net works in a configuration-adaptive way, i.e., using different dynamics and error definitions for the Low Delay B and Random Access configurations, which helps improve the adaptivity and generality in diverse coding scenarios. Yuzhang Hu, Wenhan Yang, Jiaying Liu 0001, Zongming Guo |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2022 | Semantically Contrastive Learning for Low-Light Image EnhancementabstractLow-light image enhancement (LLE) remains challenging due to the unfavorable prevailing low-contrast and weak-visibility problems of single RGB images. In this paper, we respond to the intriguing learning-related question -- if leveraging both accessible unpaired over/underexposed images and high-level semantic guidance, can improve the performance of cutting-edge LLE models? Here, we propose an effective semantically contrastive learning paradigm for LLE (namely SCL-LLE). Beyond the existing LLE wisdom, it casts the image enhancement task as multi-task joint learning, where LLE is converted into three constraints of contrastive learning, semantic brightness consistency, and feature preservation for simultaneously ensuring the exposure, texture, and color consistency. SCL-LLE allows the LLE model to learn from unpaired positives (normal-light)/negatives (over/underexposed), and enables it to interact with the scene semantics to regularize the image enhancement network, yet the interaction of high-level semantic knowledge and the low-level signal prior is seldom investigated in previous methods. Training on readily available open data, extensive experiments demonstrate that our method surpasses the state-of-the-arts LLE models over six independent cross-scenes datasets. Moreover, SCL-LLE's potential to benefit the downstream semantic segmentation under extremely dark conditions is discussed. Source Code: https://github.com/LingLIx/SCL-LLE. Dong Liang 0008, Ling Li 0010, Mingqiang Wei, Wenhan Yang, Huiyu Zhou 0001 |
AAAI | 6 |
| 2022 | Low-Light Image Enhancement with Normalizing FlowabstractTo enhance low-light images to normally-exposed ones is highly ill-posed, namely that the mapping relationship between them is one-to-many. Previous works based on the pixel-wise reconstruction losses and deterministic processes fail to capture the complex conditional distribution of normally exposed images, which results in improper brightness, residual noise, and artifacts. In this paper, we investigate to model this one-to-many relationship via a proposed normalizing flow model. An invertible network that takes the low-light images/features as the condition and learns to map the distribution of normally exposed images into a Gaussian distribution. In this way, the conditional distribution of the normally exposed images can be well modeled, and the enhancement process, i.e., the other inference direction of the invertible network, is equivalent to being constrained by a loss function that better describes the manifold structure of natural images during the training. The experimental results on the existing benchmark datasets show our method achieves better quantitative and qualitative results, obtaining better-exposed illumination, less noise and artifact, and richer colors. Yufei Wang 0006, Renjie Wan, Wenhan Yang, Haoliang Li, Lap-Pui Chau, Alex Chichung Kot |
AAAI | 3 |
| 2022 | Structure Representation Network and Uncertainty Feedback Learning for Dense Non-uniform Fog Removal
Yeying Jin, Wending Yan, Wenhan Yang, Robby T. Tan |
ACCV (3) | 3 |
| 2022 | MSDN: Mutually Semantic Distillation Network for Zero-Shot LearningabstractThe key challenge of zero-shot learning (ZSL) is how to infer the latent semantic knowledge between visual and attribute features on seen classes, and thus achieving a desirable knowledge transfer to unseen classes. Prior works either simply align the global features of an image with its associated class semantic vector or utilize unidirectional attention to learn the limited latent semantic representations, which could not effectively discover the intrinsic semantic knowledge (e.g., attribute semantics) between visual and attribute features. To solve the above dilemma, we propose a Mutually Semantic Distillation Network (MSDN), which progressively distills the intrinsic semantic representations between visual and attribute features for ZSL. MSDN incorporates an attribute→visual attention sub-net that learns attribute-based visual features, and a visual→attribute attention sub-net that learns visual-based attribute features. By further introducing a semantic distillation loss, the two mutual attention sub-nets are capable of learning collaboratively and teaching each other throughout the training process. The proposed MSDN yields significant improvements over the strong baselines, leading to new state-of-the-art performances on three popular challenging benchmarks. Our codes have been available at: https://github.com/shiming-chen/MSDN. Shiming Chen 0002, Ziming Hong, Guosen Xie, Wenhan Yang, Qinmu Peng, Kai Wang 0036, Jian Zhao 0006, Xinge You |
CVPR | 4 |
| 2022 | Neural Data-Dependent Transform for Learned Image CompressionabstractLearned image compression has achieved great success due to its excellent modeling capacity, but seldom further considers the Rate-Distortion Optimization (RDO) of each input image. To explore this potential in the learned codec, we make the first attempt to build a neural data-dependent transform and introduce a continuous online mode decision mechanism to jointly optimize the coding efficiency for each individual image. Specifically, apart from the image content stream, we employ an additional model stream to generate the transform parameters at the decoder side. The pres-ence of a model stream enables our model to learn more abstract neural-syntax, which helps cluster the latent repre-sentations of images more compactly. Beyond the transform stage, we also adopt neural-syntax based post-processing for the scenarios that require higher quality reconstructions regardless of extra decoding overhead. Moreover, the in-volvement of the model stream further makes it possible to optimize both the representation and the decoder in an on-line way, i. e. RDO at the testing time. It is equivalent to a continuous online mode decision, like coding modes in the traditional codecs, to improve the coding efficiency based on the individual input image. The experimental results show the effectiveness of the proposed neural-syntax de-sign and the continuous online mode decision mechanism, demonstrating the superiority of our method in coding effi-ciency. Our project is available at: https://dezhao-wang.github.io/Neural-Syntax-Website/. Dezhao Wang, Wenhan Yang, Yueyu Hu, Jiaying Liu 0001 |
CVPR | 2 |
| 2022 | URetinex-Net: Retinex-based Deep Unfolding Network for Low-light Image EnhancementabstractRetinex model-based methods have shown to be effective in layer-wise manipulation with well-designed priors for low-light image enhancement. However, the commonly used handcrafted priors and optimization-driven solutions lead to the absence of adaptivity and efficiency. To address these issues, in this paper, we propose a Retinex-based deep unfolding network (URetinex-Net), which unfolds an optimization problem into a learnable network to decompose a low-light image into reflectance and illumination layers. By formulating the decomposition problem as an implicit priors regularized model, three learning-based modules are carefully designed, responsible for data-dependent initialization, high-efficient unfolding optimization, and user-specified illumination enhancement, respectively. Particularly, the proposed unfolding optimization module, introducing two networks to adaptively fit implicit priors in data-driven manner, can realize noise suppression and details preservation for the final decomposition results. Extensive experiments on real-world low-light images qualitatively and quantitatively demonstrate the effectiveness and superiority of the proposed method over state-of-the-art methods. The code is available at https://github.com/AndersonYong/URetinex-Net. Wenhui Wu 0001, Jian Weng 0009, Xu Wang 0006, Wenhan Yang, Jianmin Jiang |
CVPR | 5 |
| 2022 | Towards Robust Rain Removal Against Adversarial Attacks: A Comprehensive Benchmark Analysis and BeyondabstractRain removal aims to remove rain streaks from images/videos and reduce the disruptive effects caused by rain. It not only enhances image/video visibility but also allows many computer vision algorithms to function properly. This paper makes the first attempt to conduct a comprehensive study on the robustness of deep learning-based rain removal methods against adversarial attacks. Our study shows that, when the image/video is highly degraded, rain removal methods are more vulnerable to the adversarial attacks as small distortions/perturbations become less noticeable or detectable. In this paper, we first present a comprehensive empirical evaluation of various methods at different levels of attacks and with various losses/targets to generate the perturbations from the perspective of human perception and machine analysis tasks. A systematic evaluation of key modules in existing methods is performed in terms of their robustness against adversarial attacks. From the insights of our analysis, we construct a more robust deraining method by integrating these effective modules. Finally, we examine various types of adversarial attacks that are specific to deraining problems and their effects on both human and machine vision tasks, including 1) rain region attacks, adding perturbations only in the rain regions to make the perturbations in the attacked rain images less visible; 2) object-sensitive attacks, adding perturbations only in regions near the given objects. Code is available at https://github.com/yuyi-sd/Robust_Rain_Removal. Yi Yu 0011, Wenhan Yang, Yap-Peng Tan, Alex Chichung Kot |
CVPR | 2 |
| 2022 | Unsupervised Night Image Enhancement: When Layer Decomposition Meets Light-Effects Suppression
Yeying Jin, Wenhan Yang, Robby T. Tan |
ECCV (37) | 2 |
| 2022 | Self-Learned Video Super-Resolution with Augmented Spatial and Temporal ContextabstractVideo super-resolution methods typically rely on paired training data, in which the low-resolution frames are usually synthetically generated under predetermined degradation conditions (e.g., Bicubic downsampling). However, in real applications, it is labor-consuming and expensive to obtain this kind of training data, which limits the practical performance of these methods. To address the issue and get rid of the synthetic paired data, in this paper, we make exploration in utilizing the internal self-similarity redundancy within the video to build a Self-Learned Video Super-Resolution (SLVSR) method, which only needs to be trained on the input testing video itself. We employ a series of data augmentation strategies to make full use of the spatial and temporal context of the target video clips. The idea is applied to two branches of mainstream SR methods: frame fusion and frame recurrence methods. Since the former takes advantage of the short-term temporal consistency and the latter of the long-term one, our method can satisfy different practical situations. The experimental results show the superiority of our proposed method, especially in addressing the video super-resolution problems in real applications. Zejia Fan, Jiaying Liu 0001, Wenhan Yang, Zongming Guo |
ICASSP | 3 |
| 2022 | Rain-Prior Injected Knowledge Distillation for Single Image DerainingabstractThis paper makes efforts in improving the efficiency of deep networks for single image deraining with a newly proposed knowledge distillation framework. Specifically, we propose a rain-prior injected distillation scheme to transfer the knowledge from a large-scale teacher network to a more compact student network. Previous works directly calculate the distillation loss between the features extracted from the student and teacher networks. Differently, our distillation scheme adaptively removes the noisy background patterns by calculating the distillation loss based on the residual feature, which is inferred from the features extracted from the rain and ground truth images. This residual operation makes the student network focus on transferring only the knowledge on the rain streaks instead of the background, which facilitates more effective distillation results. Furthermore, our method can be applied to reduce both the network size and the deraining recurrence stage, which makes it a plug-and-play module that can be integrated into diverse existing deraining methods. Experimental results prove the efficiency of our method to build an efficient deraining network and the superiority over existing distillation methods. Yuzhang Hu, Wenhan Yang, Jiaying Liu 0001, Zongming Guo |
ICIP | 2 |
| 2022 | Semantic Compression Embedding for Generative Zero-Shot LearningabstractGenerative methods have been successfully applied in zero-shot learning (ZSL) by learning an implicit mapping to alleviate the visual-semantic domain gaps and synthesizing unseen samples to handle the data imbalance between seen and unseen classes. However, existing generative methods simply use visual features extracted by the pre-trained CNN backbone. These visual features lack attribute-level semantic information. Consequently, seen classes are indistinguishable, and the knowledge transfer from seen to unseen classes is limited. To tackle this issue, we propose a novel Semantic Compression Embedding Guided Generation (SC-EGG) model, which cascades a semantic compression embedding network (SCEN) and an embedding guided generative network (EGGN). The SCEN extracts a group of attribute-level local features for each sample and further compresses them into the new low-dimension visual feature. Thus, a dense-semantic visual space is obtained. The EGGN learns a mapping from the class-level semantic space to the dense-semantic visual space, thus improving the discriminability of the synthesized dense-semantic unseen visual features. Extensive experiments on three benchmark datasets, i.e., CUB, SUN and AWA2, demonstrate the significant performance gains of SC-EGG over current state-of-the-art methods and its baselines. Ziming Hong, Shiming Chen 0002, Guosen Xie, Wenhan Yang, Jian Zhao 0006, Yuanjie Shao, Qinmu Peng, Xinge You |
IJCAI | 4 |
| 2022 | Collaborative Scalable Visual Compression for Human-Centered VideosabstractMachine intelligence systems have been increasingly widely deployed in real-world circumstances, while the conventional human-vision oriented video coding schemes are inefficient to be embedded in large-scale systems and further support a wide range of applications. There have been urgent demands for a new generation of compression framework to efficiently encodes visual data, where the compression and analytics for machine vision and human perception can be jointly optimized. To this end, we propose a novel visual compression framework to provide visual contents with different granularity for both human and machine vision tasks collaboratively. The proposed scalable compression framework maintains the critical semantic information in a basic layer, so that it is capable of supporting the accurate machine vision analysis under a tight bit-rate constraint. It is scalable to provide visual representations of different granularity to support various kinds of tasks, including video reconstruction that serves human vision examination. Experimental results on the human-centered videos have demonstrated the promising functionality of scalable visual coding with improved efficiency for high-performance machine analysis and human perception. Haofeng Huang, Wenhan Yang, Jiaying Liu 0001, Ling-Yu Duan |
ISCAS | 2 |
| 2022 | Meta-Interpolation: Time-Arbitrary Frame Interpolation via Dual Meta-LearningabstractExisting video frame interpolation methods can only interpolate the frame at a given intermediate time-step, e.g. 1/2. In this paper, we aim to explore a more generalized kind of video frame interpolation, that at an arbitrary time-step. To this end, we consider processing different time-steps with adaptively generated convolutional kernels in a unified way with the help of meta-learning. Specifically, we develop a dual meta-learned frame interpolation framework to synthesize intermediate frames with the guidance of context information and optical flow as well as taking the time-step as side information. First, a content-aware meta-learned flow refinement module is built to improve the accuracy of the optical flow estimation based on the down-sampled version of the input frames. Second, with the refined optical flow and the time-step as the input, a motion-aware meta-learned frame interpolation module generates the convolutional kernels for every pixel used in the convolution operations on the feature map of the coarse warped version of the input frames to generate the predicted frame. Extensive qualitative and quantitative evaluations, as well as ablation studies, demonstrate that, via introducing meta-learning in our framework in such a well-designed way, our method not only achieves superior performance to state-of-the-art frame interpolation approaches but also owns an extended capacity to support the interpolation at an arbitrary time-step. Shixing Yu, Yiyang Ma, Wenhan Yang, Jiaying Liu 0001 |
ISCAS | 3 |
| 2022 | Cycle-Interactive Generative Adversarial Network for Robust Unsupervised Low-Light EnhancementabstractGetting rid of the fundamental limitations in fitting to the paired training data, recent unsupervised low-light enhancement methods excel in adjusting illumination and contrast of images. However, for unsupervised low light enhancement, the remaining noise suppression issue due to the lacking of supervision of detailed signal largely impedes the wide deployment of these methods in real-world applications. Herein, we propose a novel Cycle-Interactive Generative Adversarial Network (CIGAN) for unsupervised low-light image enhancement, which is capable of not only better transferring illumination distributions between low/normal-light images but also manipulating detailed signals between two domains, e.g., suppressing/synthesizing realistic noise in the cyclic enhancement/degradation process. In particular, the proposed low-light guided transformation feed-forwards the features of low-light images from the generator of enhancement GAN (eGAN) into the generator of degradation GAN (dGAN). With the learned information of real low-light images, dGAN can synthesize more realistic diverse illumination and contrast in low-light images. Moreover, the feature randomized perturbation module in dGAN learns to increase the feature randomness to produce diverse feature distributions, persuading the synthesized low-light images to contain realistic noise. Extensive experiments demonstrate both the superiority of the proposed method and the effectiveness of each module in CIGAN. Zhangkai Ni, Wenhan Yang, Hanli Wang, Shiqi Wang 0001, Lin Ma 0002, Sam Kwong |
ACM Multimedia | 2 |
| 2022 | Image Inpainting Detection via Enriched Attentive Pattern with Near Original Image AugmentationabstractAs deep learning-based inpainting methods have achieved increasingly better results, its malicious use, e.g. removing objects to report fake news or to provide fake evidence, is becoming threatening. Previous works have provided rich discussions on network architectures, e.g. even performing Neural Architecture Search to obtain the optimal model architecture. However, there are rooms in other aspects. In our work, we provide comprehensive efforts from data and feature aspects. From the data aspect, as harder samples in the training data usually lead to stronger detection models, we propose near original image augmentation that pushes the inpainted images closer to the original ones (without distortion and inpainting) as the input images, which is proved to improve the detection accuracy. From the feature aspect, we propose to extract the attentive pattern. With the designed attentive pattern, the knowledge of different inpainting methods can be better exploited during the training phase. Finally, extensive experiments are conducted. In our evaluation, we consider the scenarios where the inpainting masks, which are used to generate the testing set, have a distribution gap from those masks used to produce the training set. Thus, the comparisons are conducted on a newly proposed dataset, where testing masks are inconsistent with the training ones. The experimental results show the superiority of the proposed method and the effectiveness of each component. All our codes and data will be online available. Wenhan Yang, Rizhao Cai, Alex Chichung Kot |
ACM Multimedia | 1 |
| 2022 | Spatial-temporal interaction learning based two-stream network for action recognition
Yujun Ma, Wenhan Yang, Wanting Ji, Ruili Wang 0001 |
Inf. Sci. | 3 |
| 2022 | Learning End-to-End Lossy Image Compression: A BenchmarkabstractImage compression is one of the most fundamental techniques and commonly used applications in the image and video processing field. Earlier methods built a well-designed pipeline, and efforts were made to improve all modules of the pipeline by handcrafted tuning. Later, tremendous contributions were made, especially when data-driven methods revitalized the domain with their excellent modeling capacities and flexibility in incorporating newly designed modules and constraints. Despite great progress, a systematic benchmark and comprehensive analysis of end-to-end learned image compression methods are lacking. In this paper, we first conduct a comprehensive literature survey of learned image compression methods. The literature is organized based on several aspects to jointly optimize the rate-distortion performance with a neural network, i.e., network architecture, entropy model and rate control. We describe milestones in cutting-edge learned image-compression methods, review a broad range of existing works, and provide insights into their historical development routes. With this survey, the main challenges of image compression methods are revealed, along with opportunities to address the related issues with recent advanced learning methods. This analysis provides an opportunity to take a further step towards higher-efficiency image compression. By introducing a coarse-to-fine hyperprior model for entropy estimation and signal reconstruction, we achieve improved rate-distortion performance, especially on high-resolution images. Extensive benchmark experiments demonstrate the superiority of our model in rate-distortion performance and time complexity on multi-core CPUs and GPUs. Yueyu Hu, Wenhan Yang, Zhan Ma 0001, Jiaying Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Recurrent Multi-Frame Deraining: Combining Physics Guidance and Adversarial LearningabstractExisting video rain removal methods mainly focus on rain streak removal and are solely trained based on the synthetic data, which neglect more complex degradation factors, e.g., rain accumulation, and the prior knowledge in real rain data. Thus, in this paper, we build a more comprehensive rain model with several degradation factors and construct a novel two-stage video rain removal method that combines the power of synthetic videos and real data. Specifically, a novel two-stage progressive network is proposed: recovery guided by a physics model, and further restoration by adversarial learning. The first stage performs an inverse recovery process guided by our proposed rain model. An initially estimated background frame is obtained based on the input rain frame. The second stage employs adversarial learning to refine the result, i.e., recovering the overall color and illumination distributions of the frame, the background details that are failed to be recovered in the first stage, and removing the artifacts generated in the first stage. Furthermore, we also introduce a more comprehensive rain model that includes degradation factors, e.g., occlusion and rain accumulation, which appear in real scenes yet ignored by existing methods. This model, which generates more realistic rain images, will train and evaluate our models better. Extensive evaluations on synthetic and real videos show the effectiveness of our method in comparisons to the state-of-the-art methods. Our datasets, results and code are available at: https://github.com/flyywh/Recurrent-Multi-Frame-Deraining. Wenhan Yang, Robby T. Tan, Jiashi Feng, Shiqi Wang 0001, Bin Cheng 0001, Jiaying Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Appearance Matters, So Does Audio: Revealing the Hidden Face via Cross-Modality TransferabstractRecently, there has been an exponential increase in the security concerns raised by faking face (e.g., deepfake), which automatically changes the identity with a specifically learned deep generative model. With numerous approaches proposed to identify the fake content, much less work has been dedicated to automatically revealing the authentic one that is originally acquired. Here, we propose a new paradigm that seeks to reveal the authentic face hidden behind the fake one by leveraging the joint information of face and audio. More specifically, given the fake face as well as the audio segment, the cross-modality transferable capability is exploited by learning to generate the feature of the authentic face, based on the underlying clues from the audio as well as the fake face appearance. The effectiveness of the proposed scheme is validated through a series of evaluations, and experimental results show that the proposed model achieves promising face reconstruction performance in revealing the hidden faces, in terms of reconstruction quality, as well as identity and face attribute inference accuracy. Chenqi Kong, Baoliang Chen, Wenhan Yang, Haoliang Li, Peilin Chen 0001, Shiqi Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | Enlightening Low-Light Images With Dynamic Guidance for Context EnrichmentabstractImages acquired in low-light conditions suffer from a series of visual quality degradations,e.g., low visibility, degraded contrast, and intensive noise. These complicated degradations based on various contexts (e.g., noise in smooth regions, over-exposure in well-exposed regions and low contrast around edges) cast major challenges to the low-light image enhancement. Herein, we propose a new methodology by imposing a learnable guidance map from the signal and deep priors, making the deep neural network adaptively enhance low-light images in a region-dependent manner. The enhancement capability of the learnable guidance map is further exploited with the multi-scale dilated context collaboration, leading to contextually enriched feature representations extracted by the model with various receptive fields. Through assimilating the intrinsic perceptual information from the learned guidance map, richer and more realistic textures are generated. Extensive experiments on real low-light images demonstrate the effectiveness of our method, which delivers superior results quantitatively and qualitatively. The code is available athttps://github.com/lingyzhu0101/GEMSCto facilitate future research. Lingyu Zhu 0006, Wenhan Yang, Baoliang Chen, Fangbo Lu, Shiqi Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Towards Low Light Enhancement With RAW ImagesabstractIn this paper, we make the first benchmark effort to elaborate on the superiority of using RAW images in the low light enhancement and develop a novel alternative route to utilize RAW images in a more flexible and practical way. Inspired by a full consideration on the typical image processing pipeline, we are inspired to develop a new evaluation framework, Factorized Enhancement Model (FEM), which decomposes the properties of RAW images into measurable factors and provides a tool for exploring how properties of RAW images affect the enhancement performance empirically. The empirical benchmark results show that the Linearity of data and Exposure Time recorded in meta-data play the most critical role, which brings distinct performance gains in various measures over the approaches taking the sRGB images as input. With the insights obtained from the benchmark results in mind, a RAW-guiding Exposure Enhancement Network (REENet) is developed, which makes trade-offs between the advantages and inaccessibility of RAW images in real applications in a way of using RAW images only in the training phase. REENet projects sRGB images into linear RAW domains to apply constraints with corresponding RAW images to reduce the difficulty of modeling training. After that, in the testing phase, our REENet does not rely on RAW images. Experimental results demonstrate not only the superiority of REENet to state-of-the-art sRGB-based methods and but also the effectiveness of the RAW guidance and all components. Haofeng Huang, Wenhan Yang, Yueyu Hu, Jiaying Liu 0001, Ling-Yu Duan |
IEEE Trans. Image Process. | 2 |
| 2022 | Feature-Aligned Video Raindrop Removal With Temporal ConstraintsabstractExisting adherent raindrop removal methods focus on the detection of the raindrop locations, and then use inpainting techniques or generative networks to recover the background behind raindrops. Yet, as adherent raindrops are diverse in sizes and appearances, the detection is challenging for both single image and video. Moreover, unlike rain streaks, adherent raindrops tend to cover the same area in several frames. Addressing these problems, our method employs a two-stage video-based raindrop removal method. The first stage is the single image module, which generates initial clean results. The second stage is the multiple frame module, which further refines the initial results using temporal constraints, namely, by utilizing multiple input frames in our process and applying temporal consistency between adjacent output frames. Our single image module employs a raindrop removal network to generate initial raindrop removal results, and create a mask representing the differences between the input and initial output. Once the masks and initial results for consecutive frames are obtained, our multiple-frame module aligns the frames in both the image and feature levels and then obtains the clean background. Our method initially employs optical flow to align the frames, and then utilizes deformable convolution layers further to achieve feature-level frame alignment. To remove small raindrops and recover correct backgrounds, a target frame is predicted from adjacent frames. A series of unsupervised losses are proposed so that our second stage, which is the video raindrop removal module, can self-learn from video data without ground truths. Experimental results on real videos demonstrate the state-of-art performance of our method both quantitatively and qualitatively. Wending Yan, Wenhan Yang, Robby T. Tan |
IEEE Trans. Image Process. | 3 |
| 2022 | Towards Analysis-Friendly Face Representation With Scalable Feature and Texture CompressionabstractCompactly representing visual information plays a fundamental role in optimizing the ultimate utility of myriad visual data-centered applications. Numerous approaches have been proposed to efficiently compress the texture and visual features for human visual perception and machine intelligence, respectively; however, much less work has been dedicated to studying the interactions between them. Here, we investigate the integration of feature and texture compression and show that a universal and collaborative visual information representation can be achieved in a hierarchical way. In particular, we study feature and texture compression in a scalable coding framework, where the base layer serves as the deep learning feature and the enhancement layer targets to perfectly reconstruct the texture. Based on the strong generative capability of deep neural networks, the gap between the base feature layer and enhancement layer is further filled with feature-level texture reconstruction, with the goal of further constructing texture representations from features. As such, the residuals between the original and reconstructed texture could be further conveyed in the enhancement layer. To improve the efficiency of the proposed framework, the base layer neural network is trained in a multitask manner such that the learned features enjoy both high-quality reconstruction and high-accuracy analysis. The framework and optimization strategies are further applied in face image compression, and promising coding performance has been achieved in terms of both rate-fidelity and rate-accuracy evaluations. Shurun Wang, Shiqi Wang 0001, Wenhan Yang, Xinfeng Zhang 0001, Shanshe Wang, Siwei Ma 0001, Wen Gao 0001 |
IEEE Trans. Multim. | 3 |
| 2021 | HLA-Face: Joint High-Low Adaptation for Low Light Face DetectionabstractFace detection in low light scenarios is challenging but vital to many practical applications, e.g., surveillance video, autonomous driving at night. Most existing face detectors heavily rely on extensive annotations, while collecting data is time-consuming and laborious. To reduce the burden of building new datasets for low light conditions, we make full use of existing normal light data and explore how to adapt face detectors from normal light to low light. The challenge of this task is that the gap between normal and low light is too huge and complex for both pixel-level and object-level. Therefore, most existing low-light enhancement and adaptation methods do not achieve desirable performance. To address the issue, we propose a joint High-Low Adaptation (HLA) framework. Through a bidirectional low-level adaptation and multi-task high-level adaptation scheme, our HLA-Face outperforms state-of-the-art methods even without using dark face labels for training. Our project is publicly available at: https://daooshee.github.io/HLA-Face-Website/. Wenjing Wang 0001, Wenhan Yang, Jiaying Liu 0001 |
CVPR | 2 |
| 2021 | Self-Aligned Video Deraining With Transmission-Depth ConsistencyabstractIn this paper, we address the problem of rain streaks and rain accumulation removal in video, by developing a self-alignment network with transmission-depth consistency. Existing video based deraining methods focus only on rain streak removal, and commonly use optical flow to align the rain video frames. However, besides rain streaks, rain accummulation can considerably degrade visibility; and, optical flow estimation in a rain video is still erroneous, making the deraining performance tend to be inaccurate. Our method employs deformable convolution layers in our encoder to achieve feature-level frame alignment, and hence avoids using optical flow. For rain streaks, our method predicts the current frame from its adjacent frames, such that rain streaks that appear randomly in the temporal domain can be removed. For rain accumulation, our method employs a transmission-depth consistency loss to resolve the ambiguity between the depth and water-droplet density. Our network estimates the depth from consecutive rain-accumulation-removal outputs, and calculates the transmission map using a commonly used physics model. To ensure photometric-temporal and depth-temporal consistencies, our method estimates the camera poses, so that it can warp one frame to its adjacent frames. Experimental results show that our method is effective in removing both rain streaks and rain accumulation, outperforming those of state-of-the-art methods quantitatively and qualitatively. Wending Yan, Robby T. Tan, Wenhan Yang, Dengxin Dai |
CVPR | 3 |
| 2021 | Teacher-Student Learning With Multi-Granularity Constraint Towards Compact Facial Feature RepresentationabstractIn this paper, we propose a novel end-to-end feature compression scheme by leveraging the representation and learning capability of deep neural networks, towards intelligent front-end equipped analysis with promising accuracy and efficiency. In particular, the extracted features are compactly coded in an end-to-end manner by optimizing the rate- distortion cost to achieve feature-in-feature representation. The multi-granularity constraint is further imposed, serving as the optimization objective to make the feature compression more "healthier" from the perspective of ultimate utility. More specifically, the analysis accuracy is considered in the coarse granularity level constraint, ensuring the capability of facial analysis with the reconstructed feature. Furthermore, at the fine granularity level the feature fidelity is involved to preserve the original feature quality. Moreover, a latent code level teacher-student enhancement model is proposed to efficiently transfer the low bit-rate representation into a high bit- rate one. Such a strategy further allows us to adaptively shift the representation cost to decoding computations, leading to more flexible feature compression with enhanced decoding capability. We verify the effectiveness of the proposed model with the facial feature, and experimental results reveal better compression performance in terms of rate-accuracy compared with existing models. Shurun Wang, Shiqi Wang 0001, Wenhan Yang, Xinfeng Zhang 0001, Shanshe Wang, Siwei Ma 0001 |
ICASSP | 3 |
| 2021 | Benchmarking Low-Light Image Enhancement and Beyond
Jiaying Liu 0001, Dejia Xu, Wenhan Yang, Minhao Fan, Haofeng Huang |
Int. J. Comput. Vis. | 3 |
| 2021 | Bridging the Gap Between Computational Photography and Visual RecognitionabstractWhat is the current state-of-the-art for image restoration and enhancement applied to degraded images acquired under less than ideal circumstances? Can the application of such algorithms as a pre-processing step improve image interpretability for manual analysis or automatic visual recognition to classify scene content? While there have been important advances in the area of computational photography to restore or enhance the visual quality of an image, the capabilities of such techniques have not always translated in a useful way to visual recognition tasks. Consequently, there is a pressing need for the development of algorithms that are designed for the joint problem of improving visual appearance and recognition, which will be an enabling factor for the deployment of visual recognition tools in many real-world scenarios. To address this, we introduce the UG$^2$dataset as a large-scale benchmark composed of video imagery captured under challenging conditions, and two enhancement tasks designed to test algorithmic impact on visual quality and automatic object recognition. Furthermore, we propose a set of metrics to evaluate the joint improvement of such tasks as well as individual algorithmic advances, including a novel psychophysics-based evaluation regime for human assessment and a realistic set of quantitative measures for object recognition performance. We introduce six new algorithms for image restoration or enhancement, which were created as part of the IARPA sponsored UG$^2$Challenge workshop held at CVPR 2018. Under the proposed evaluation regime, we present an in-depth analysis of these algorithms and a host of deep learning-based and classic baseline approaches. From the observed results, it is evident that we are in the early days of building a bridge between computational photography and visual recognition, leaving many opportunities for innovation in this area. Rosaura G. VidalMata, Sreya Banerjee, Brandon RichardWebster, Michael Albright, Pedro Davalos, Scott McCloskey, Ben Miller, Asong Tambo, Sushobhan Ghosh, Sudarshan Nagesh, Ye Yuan 0012, Yueyu Hu, Wenhan Yang, Xiaoshuai Zhang, Jiaying Liu 0001, Zhangyang Wang, Hwann-Tzong Chen, Tzu-Wei Huang, Wen-Chi Chin, Yi-Chun Li, Mahmoud Lababidi, Charles Otto, Walter J. Scheirer |
IEEE Trans. Pattern Anal. Mach. Intell. | 14 |
| 2021 | Single Image Deraining: From Model-Based to Data-Driven and BeyondabstractThe goal of single-image deraining is to restore the rain-free background scenes of an image degraded by rain streaks and rain accumulation. The early single-image deraining methods employ a cost function, where various priors are developed to represent the properties of rain and background layers. Since 2017, single-image deraining methods step into a deep-learning era, and exploit various types of networks, i.e., convolutional neural networks, recurrent neural networks, generative adversarial networks, etc., demonstrating impressive performance. Given the current rapid development, in this paper, we provide a comprehensive survey of deraining methods over the last decade. We summarize the rain appearance models, and discuss two categories of deraining approaches: model-based and data-driven approaches. For the former, we organize the literature based on their basic models and priors. For the latter, we discuss the developed ideas related to architectures, constraints, loss functions, and training datasets. We present milestones of single-image deraining methods, review a broad selection of previous works in different categories, and provide insights on the historical development route from the model-based to data-driven methods. We also summarize performance comparisons quantitatively and qualitatively. Beyond discussing the technicality of deraining methods, we also discuss the future possible directions. Wenhan Yang, Robby T. Tan, Shiqi Wang 0001, Yuming Fang 0001, Jiaying Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | Camera Invariant Feature Learning for Generalized Face Anti-SpoofingabstractThere has been an increasing consensus in learning based face anti-spoofing that the divergence in terms of camera models is causing a large domain gap in real application scenarios. We describe a framework that eliminates the influence of inherent variance from acquisition cameras at the feature level, leading to the generalized face spoofing detection model that could be highly adaptive to different acquisition devices. In particular, the framework is composed of two branches. The first branch aims to learn the camera invariant spoofing features via feature level decomposition in the high frequency domain. Motivated by the fact that the spoofing features exist not only in the high frequency domain, in the second branch the discrimination capability of extracted spoofing features is further boosted from the enhanced image based on the recomposition of the high-frequency and low-frequency information. Finally, the classification results of the two branches are fused together by a weighting strategy. Experiments show that the proposed method can achieve better performance in both intra-dataset and cross-dataset settings, demonstrating the high generalization capability in various application scenarios. Baoliang Chen, Wenhan Yang, Haoliang Li, Shiqi Wang 0001, Sam Kwong |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2021 | Compressed Domain Deep Video Super-ResolutionabstractReal-world video processing algorithms are often faced with the great challenges of processing the compressed videos instead of pristine videos. Despite the tremendous successes achieved in deep-learning based video super-resolution (SR), much less work has been dedicated to the SR of compressed videos. Herein, we propose a novel approach for compressed domain deep video SR by jointly leveraging the coding priors and deep priors. By exploiting the diverse and ready-made spatial and temporal coding priors (e.g., partition maps and motion vectors) extracted directly from the video bitstream in an effortless way, the video SR in the compressed domain allows us to accurately reconstruct the high resolution video with high flexibility and substantially economized computational complexity. More specifically, to incorporate the spatial coding prior, the Guided Spatial Feature Transform (GSFT) layer is proposed to modulate features of the prior with the guidance of the video information, making the prior features more fine-grained and content-adaptive. To incorporate the temporal coding prior, a guided soft alignment scheme is designed to generate local attention off-sets to compensate for decoded motion vectors. Our soft alignment scheme combines the merits of explicit and implicit motion modeling methods, rendering the alignment of features more effective for SR in terms of the computational complexity and robustness to inaccurate motion fields. Furthermore, to fully make use of the deep priors, the multi-scale fused features are generated from a scale-wise convolution reconstruction network for final SR video reconstruction. To promote the compressed domain video SR research, we build an extensive Compressed Videos with Coding Prior (CVCP) dataset, including compressed videos of diverse content and various coding priors extracted from the bitstream. Extensive experimental results show the effectiveness of coding priors in compressed domain video SR. Peilin Chen 0001, Wenhan Yang, Meng Wang 0017, Kangkang Hu, Shiqi Wang 0001 |
IEEE Trans. Image Process. | 2 |
| 2021 | Just Noticeable Distortion Profile Inference: A Patch-Level Structural Visibility Learning ApproachabstractIn this paper, we propose an effective approach to infer the just noticeable distortion (JND) profile based on patch-level structural visibility learning. Instead of pixel-level JND profile estimation, the image patch, which is regarded as the basic processing unit to better correlate with the human perception, can be further decomposed into three conceptually independent components for visibility estimation. In particular, to incorporate the structural degradation into the patch-level JND model, a deep learning-based structural degradation estimation model is trained to approximate the masking of structural visibility. In order to facilitate the learning process, a JND dataset is further established, including 202 pristine images and 7878 distorted images generated by advanced compression algorithms based on the upcoming Versatile Video Coding (VVC) standard. Extensive experimental results further show the superiority of the proposed approach over the state-of-the-art. Our dataset is available at: https://github.com/ShenXuelin-CityU/PWJNDInfer. Xuelin Shen, Zhangkai Ni, Wenhan Yang, Xinfeng Zhang 0001, Shiqi Wang 0001, Sam Kwong |
IEEE Trans. Image Process. | 3 |
| 2021 | Combining Progressive Rethinking and Collaborative Learning: A Deep Framework for In-Loop FilteringabstractIn this paper, we aim to address issues of (1) joint spatial-temporal modeling and (2) side information injection for deep-learning based in-loop filter. For (1), we design a deep network with both progressive rethinking and collaborative learning mechanisms to improve quality of the reconstructed intra-frames and inter-frames, respectively. For intra coding, a Progressive Rethinking Network (PRN) is designed to simulate the human decision mechanism for effective spatial modeling. Our designed block introduces an additional inter-block connection to bypass a high-dimensional informative feature before the bottleneck module across blocks to review the complete past memorized experiences and rethinks progressively. For inter coding, the current reconstructed frame interacts with reference frames (peak quality frame and the nearest adjacent frame) collaboratively at the feature level. For (2), we extract both intra-frame and inter-frame side information for better context modeling. A coarse-to-fine partition map based on HEVC partition trees is built as the intra-frame side information. Furthermore, the warped features of the reference frames are offered as the inter-frame side information. Our PRN with intra-frame side information provides 9.0% BD-rate reduction on average compared to HEVC baseline under All-intra (AI) configuration. While under Low-Delay B (LDB), Low-Delay P (LDP) and Random Access (RA) configuration, our PRN with inter-frame side information provides 9.0%, 10.6% and 8.0% BD-rate reduction on average respectively. Our project webpage is https://dezhao-wang.github.io/PRN-v2/. Dezhao Wang, Sifeng Xia, Wenhan Yang, Jiaying Liu 0001 |
IEEE Trans. Image Process. | 3 |
| 2021 | Band Representation-Based Semi-Supervised Low-Light Image Enhancement: Bridging the Gap Between Signal Fidelity and Perceptual QualityabstractIt has been widely acknowledged that under-exposure causes a variety of visual quality degradation because of intensive noise, decreased visibility, biased color, etc. To alleviate these issues, a novel semi-supervised learning approach is proposed in this paper for low-light image enhancement. More specifically, we propose a deep recursive band network (DRBN) to recover a linear band representation of an enhanced normal-light image based on the guidance of the paired low/normal-light images. Such design philosophy enables the principled network to generate a quality improved one by reconstructing the given bands based upon another learnable linear transformation which is perceptually driven by an image quality assessment neural network. On one hand, the proposed network is delicately developed to obtain a variety of coarse-to-fine band representations, of which the estimations benefit each other in a recursive process mutually. On the other hand, the extracted band representation of the enhanced image in the recursive band learning stage of DRBN is capable of bridging the gap between the restoration knowledge of paired data and the perceptual quality preference to high-quality images. Subsequently, the band recomposition learns to recompose the band representation towards fitting perceptual regularization of high-quality images with the perceptual guidance. The proposed architecture can be flexibly trained with both paired and unpaired data. Extensive experiments demonstrate that our method produces better enhanced results with visually pleasing contrast and color distributions, as well as well-restored structural details. Wenhan Yang, Shiqi Wang 0001, Yuming Fang 0001, Yue Wang 0032, Jiaying Liu 0001 |
IEEE Trans. Image Process. | 1 |
| 2021 | Sparse Gradient Regularized Deep Retinex Network for Robust Low-Light Image EnhancementabstractDue to the absence of a desirable objective for low-light image enhancement, previous data-driven methods may provide undesirable enhanced results including amplified noise, degraded contrast and biased colors. In this work, inspired by Retinex theory, we design an end-to-end signal prior-guided layer separation and data-driven mapping network with layer-specified constraints for single-image low-light enhancement. A Sparse Gradient Minimization sub-Network (SGM-Net) is constructed to remove the low-amplitude structures and preserve major edge information, which facilitates extracting paired illumination maps of low/normal-light images. After the learned decomposition, two sub-networks (Enhance-Net and Restore-Net) are utilized to predict the enhanced illumination and reflectance maps, respectively, which helps stretch the contrast of the illumination map and remove intensive noise in the reflectance map. The effects of all these configured constraints, including the signal structure regularization and losses, combine together reciprocally, which leads to good reconstruction results in overall visual quality. The evaluation on both synthetic and real images, particularly on those containing intensive noise, compression artifacts and their interleaved artifacts, shows the effectiveness of our novel models, which significantly outperforms the state-of-the-art methods. Wenhan Yang, Wenjing Wang 0001, Haofeng Huang, Shiqi Wang 0001, Jiaying Liu 0001 |
IEEE Trans. Image Process. | 1 |
| 2021 | Towards Coding for Human and Machine Vision: Scalable Face Image CodingabstractThe past decades have witnessed the rapid development of image and video coding techniques in the era of big data. However, the signal fidelity-driven coding pipeline design limits the capability of the existing image/video coding frameworks to fulfill the needs of both machine and human vision. In this paper, we come up with a novel face image coding framework by leveraging both the compressive and the generative models, to support machine vision and human perception tasks jointly. Given an input image, the feature analysis is first applied, and then the generative model is employed to reconstruct image with compact structure and color features, where sparse edges are extracted to connect both kinds of vision and a key reference pixel selection method is proposed to determine the priorities of the reference color pixels for scalable coding. The compact edge map serves as the basic layer for machine vision tasks, and the reference pixels act as an enhanced layer to guarantee signal fidelity for human vision. By introducing advanced generative models, we train a decoding network to reconstruct images from compact structure and color representations, which is flexible to accept inputs in a scalable way and to control the imagery effect of the outputs between signal fidelity and visual realism. Experimental results and comprehensive performance analysis over the face image dataset demonstrate the superiority of our framework in both human vision tasks and machine vision tasks, which provide useful evidence on the emerging standardization efforts on MPEG VCM (Video Coding for Machine). Shuai Yang 0001, Yueyu Hu, Wenhan Yang, Ling-Yu Duan, Jiaying Liu 0001 |
IEEE Trans. Multim. | 3 |
| 2020 | Coarse-to-Fine Hyper-Prior Modeling for Learned Image CompressionabstractApproaches to image compression with machine learning now achieve superior performance on the compression rate compared to existing hybrid codecs. The conventional learning-based methods for image compression exploits hyper-prior and spatial context model to facilitate probability estimations. Such models have limitations in modeling long-term dependency and do not fully squeeze out the spatial redundancy in images. In this paper, we propose a coarse-to-fine framework with hierarchical layers of hyper-priors to conduct comprehensive analysis of the image and more effectively reduce spatial redundancy, which improves the rate-distortion performance of image compression significantly. Signal Preserving Hyper Transforms are designed to achieve an in-depth analysis of the latent representation and the Information Aggregation Reconstruction sub-network is proposed to maximally utilize side-information for reconstruction. Experimental results show the effectiveness of the proposed network to efficiently reduce the redundancies in images and improve the rate-distortion performance, especially for high-resolution images. Our project is publicly available at https://huzi96.github.io/coarse-to-fine-compression.html. Yueyu Hu, Wenhan Yang, Jiaying Liu 0001 |
AAAI | 2 |
| 2020 | Towards Scale-Free Rain Streak Removal via Self-Supervised Fractal Band LearningabstractData-driven rain streak removal methods, which most of rely on synthesized paired data, usually come across the generalization problem when being applied in real cases. In this paper, we propose a novel deep-learning based rain streak removal method injected with self-supervision to improve the ability to remove rain streaks in various scales. To realize this goal, we made efforts in two aspects. First, considering that rain streak removal is highly correlated with texture characteristics, we create a fractal band learning (FBL) network based on frequency band recovery. It integrates commonly seen band feature operations with neural modules and effectively improves the capacity to capture discriminative features for deraining. Second, to further improve the generalization ability of FBL for rain streaks in various scales, we add cross-scale self-supervision to regularize the network training. The constraint forces the extracted features of inputs in different scales to be equivalent after rescaling. Therefore, FBL can offer similar responses based on solely image content without the interleave of scale and is capable to remove rain streaks in various scales. Extensive experiments in quantitative and qualitative evaluations demonstrate the superiority of our FBL for rain streak removal, especially for the real cases where very large rain streaks exist, and prove the effectiveness of its each component. Our code will be public available at: https://github.com/flyywh/AAAI-2020-FBL-SS. Wenhan Yang, Shiqi Wang 0001, Dejia Xu, Jiaying Liu 0001 |
AAAI | 1 |
| 2020 | Raw-Guided Enhancing Reprocess of Low-Light Image via Deep Exposure Adjustment
Haofeng Huang, Wenhan Yang, Yueyu Hu, Jiaying Liu 0001 |
ACCV (2) | 2 |
| 2020 | From Fidelity to Perceptual Quality: A Semi-Supervised Approach for Low-Light Image EnhancementabstractUnder-exposure introduces a series of visual degradation, i.e. decreased visibility, intensive noise, and biased color, etc. To address these problems, we propose a novel semi-supervised learning approach for low-light image enhancement. A deep recursive band network (DRBN) is proposed to recover a linear band representation of an enhanced normal-light image with paired low/normal-light images, and then obtain an improved one by recomposing the given bands via another learnable linear transformation based on a perceptual quality-driven adversarial learning with unpaired data. The architecture is powerful and flexible to have the merit of training with both paired and unpaired data. On one hand, the proposed network is well designed to extract a series of coarse-to-fine band representations, whose estimations are mutually beneficial in a recursive process. On the other hand, the extracted band representation of the enhanced image in the first stage of DRBN (recursive band learning) bridges the gap between the restoration knowledge of paired data and the perceptual quality preference to real high-quality images. Its second stage (band recomposition) learns to recompose the band representation towards fitting perceptual properties of high-quality images via adversarial learning. With the help of this two-stage design, our approach generates enhanced results with well-reconstructed details and visually promising contrast and color distributions. Qualitative and quantitative evaluations demonstrate the superiority of our DRBN. Wenhan Yang, Shiqi Wang 0001, Yuming Fang 0001, Yue Wang 0032, Jiaying Liu 0001 |
CVPR | 1 |
| 2020 | Self-Learning Video Rain Streak Removal: When Cyclic Consistency Meets Temporal CorrespondenceabstractIn this paper, we address the problem of rain streaks removal in video by developing a self-learned rain streak removal method, which does not require any clean groundtruth images in the training process. The method is inspired by fact that the adjacent frames are highly correlated and can be regarded as different versions of identical scene, and rain streaks are randomly distributed along the temporal dimension. With this in mind, we construct a two-stage Self-Learned Deraining Network (SLDNet) to remove rain streaks based on both temporal correlation and consistency. In the first stage, SLDNet utilizes the temporal correlations and learns to predict the clean version of the current frame based on its adjacent rain video frames. In the second stage, SLDNet enforces the temporal consistency among different frames. It takes both the current rain frame and adjacent rain video frames to recover structural details. The first stage is responsible for reconstructing main structures, and the second stage is responsible for extracting structural details. We build our network architecture with two sub-tasks, i.e. motion estimation, and rain region detection, and optimize them jointly. Our extensive experiments demonstrate the effectiveness of our method, offering better results both quantitatively and qualitatively. Wenhan Yang, Robby T. Tan, Shiqi Wang 0001, Jiaying Liu 0001 |
CVPR | 1 |
| 2020 | Towards Coding For Human And Machine Vision: A Scalable Image Coding ApproachabstractThe past decades have witnessed the rapid development of image and video coding techniques in the era of big data. However, the signal fidelity-driven coding pipeline design limits the capability of the existing image/video coding frameworks to fulfill the needs of both machine and human vision. In this paper, we come up with a novel image coding framework by leveraging both the compressive and the generative models, to support machine vision and human perception tasks jointly. Given an input image, the feature analysis is first applied, and then the generative model is employed to perform image reconstruction with features and additional reference pixels, in which compact edge maps are extracted in this work to connect both kinds of vision in a scalable way. The compact edge map serves as the basic layer for machine vision tasks, and the reference pixels act as a sort of enhanced layer to guarantee signal fidelity for human vision. By introducing advanced generative models, we train a flexible network to reconstruct images from compact feature representations and the reference pixels. Experimental results demonstrate the superiority of our framework in both human visual quality and facial landmark detection, which provide useful evidence on the emerging standardization efforts on MPEG VCM (Video Coding for Machine). Our project website is available at https://williamyang1991.github.io/projects/VCM-Face/. Yueyu Hu, Shuai Yang 0001, Wenhan Yang, Ling-Yu Duan, Jiaying Liu 0001 |
ICME | 3 |
| 2020 | An Emerging Coding Paradigm Vcm: A Scalable Coding Approach Beyond Feature And SignalabstractIn this paper, we study a new problem arising from the emerging MPEG standardization effort Video Coding for Machine (VCM)1, which aims to bridge the gap between visual feature compression and classical video coding. VCM is committed to address the requirement of compact signal representation for both machine and human vision in a more or less scalable way. To this end, we make endeavors in leveraging the strength of predictive and generative models to support advanced compression techniques for both machine and human vision tasks simultaneously, in which visual features serve as a bridge to connect signal-level and task-level compact representations in a scalable manner. Specifically, we employ a conditional deep generation network to reconstruct video frames with the guidance of learned motion pattern. By learning to extract sparse motion pattern via a predictive model, the network elegantly leverages the feature representation to generate the appearance of to-be-coded frames via a generative model, relying on the appearance of the coded key frames. Meanwhile, the sparse motion pattern is compact and highly effective for high-level vision tasks, e.g. action recognition. Experimental results demonstrate that our method yields much better reconstruction quality compared with the traditional video codecs (0.0063 gain in SSIM), as well as state-of-the-art action recognition performance over highly compressed videos (9.4% gain in recognition accuracy), which showcases a promising paradigm of coding signal for both human and machine vision. Sifeng Xia, Kunchangtai Liang, Wenhan Yang, Ling-Yu Duan, Jiaying Liu 0001 |
ICME | 3 |
| 2020 | Memory-Augmented Auto-Regressive Network for Frame Recurrent Inter PredictionabstractInter prediction is quite important for the modern codecs to remove temporal redundancy. In this paper, we make endeavors in generating artificial reference frames with previous reconstructed frames for inter prediction, to offer a better choice when the traditional block-wise motion estimation fails to find a good reference block. Long-term temporal dynamics are tracked during the whole coding process to generate more accurate and realistic artificial reference frames. Specifically, we propose a Memory-Augmented Auto-Regressive Network (MAAR-Net) for frame prediction in video coding. MAAR-Net regresses the current frame with two nearest frames via an auto-regressive (AR) model to better capture the main spatial and temporal structures. The AR regression coefficients are generated based on adjacent frame information as well as the long-term motion dynamics accumulated and propagated by a convolutional Long Short-Term Memory (LSTM). To generate the target frame with higher quality, a quality attention mechanism is introduced for the temporal regularization between different reconstructed frames. With the well-designed network, our method surpasses HEVC on average 4.0% BD-rate saving and up to 10.6% BD-rate saving for the luma component under the low-delay configuration. Yuzhang Hu, Sifeng Xia, Wenhan Yang, Jiaying Liu 0001 |
ISCAS | 3 |
| 2020 | When Bitstream Prior Meets Deep Prior: Compressed Video Super-resolution with Learning from DecodingabstractThe standard paradigm of video super-resolution (SR) is to generate the spatial-temporal coherent high-resolution (HR) sequence from the corresponding low-resolution (LR) version which has already been decoded from the bitstream. However, a highly practical while relatively under-studied way is enabling the built-in SR functionality in the decoder, in the sense that almost all videos are compactly represented. In this paper, we systematically investigate the SR of compressed LR videos by leveraging the interactivity between decoding prior and deep prior. By fully exploiting the compact video stream information, the proposed bitstream prior embedded SR framework achieves compressed video SR and quality enhancement simultaneously in a single feed-forward process. More specifically, we propose a motion vector guided multi-scale local attention module that explicitly exploits the temporal dependency and suppresses coding artifacts with substantially economized computational complexity. Moreover, a scale-wise deep residual-in-residual network is learned to reconstruct the SR frames from the multi-scale fused features. To facilitate the research of compressed video SR, we also build a large-scale dataset with compressed videos of diverse content, including ready-made diversified kinds of side information extracted from the bitstream. Both quantitative and qualitative evaluations show that our model achieves superior performance for compressed video SR, and offers competitive performance compared to the sequential combinations of the state-of-the-art methods for compressed video artifacts removal and SR. Peilin Chen 0001, Wenhan Yang, Shiqi Wang 0001 |
ACM Multimedia | 2 |
| 2020 | Integrating Semantic Segmentation and Retinex Model for Low-Light Image EnhancementabstractRetinex model is widely adopted in various low-light image enhancement tasks. The basic idea of the Retinex theory is to decompose images into reflectance and illumination. The ill-posed decomposition is usually handled by hand-crafted constraints and priors. With the recently emerging deep-learning based approaches as tools, in this paper, we integrate the idea of Retinex decomposition and semantic information awareness. Based on the observation that various objects and backgrounds have different material, reflection and perspective attributes, regions of a single low-light image may require different adjustment and enhancement regarding contrast, illumination and noise. We propose an enhancement pipeline with three parts that effectively utilize the semantic layer information. Specifically, we extract the segmentation, reflectance as well as illumination layers, and concurrently enhance every separate region, i.e. sky, ground and objects for outdoor scenes. Extensive experiments on both synthetic data and real world images demonstrate the superiority of our method over current state-of-the-art low-light enhancement algorithms. Minhao Fan, Wenjing Wang 0001, Wenhan Yang, Jiaying Liu 0001 |
ACM Multimedia | 3 |
| 2020 | MS2L: Multi-Task Self-Supervised Learning for Skeleton Based Action RecognitionabstractIn this paper, we address self-supervised representation learning from human skeletons for action recognition. Previous methods, which usually learn feature presentations from a single reconstruction task, may come across the overfitting problem, and the features are not generalizable for action recognition. Instead, we propose to integrate multiple tasks to learn more general representations in a self-supervised manner. To realize this goal, we integrate motion prediction, jigsaw puzzle recognition, and contrastive learning to learn skeleton features from different aspects. Skeleton dynamics can be modeled through motion prediction by predicting the future sequence. And temporal patterns, which are critical for action recognition, are learned through solving jigsaw puzzles. We further regularize the feature space by contrastive learning. Besides, we explore different training strategies to utilize the knowledge from self-supervised tasks for action recognition. We evaluate our multi-task self-supervised learning approach with action classifiers trained under different configurations, including unsupervised, semi-supervised and fully-supervised settings. Our experiments on the NW-UCLA, NTU RGB+D, and PKUMMD datasets show remarkable performance for action recognition, demonstrating the superiority of our method in learning more discriminative and general features. Our project website is available at https://langlandslin.github.io/projects/MSL/. Lilang Lin, Sijie Song, Wenhan Yang, Jiaying Liu 0001 |
ACM Multimedia | 3 |
| 2020 | Unpaired Image Enhancement with Quality-Attention Generative Adversarial NetworkabstractIn this work, we aim to learn an unpaired image enhancement model, which can enrich low-quality images with the characteristics of high-quality images provided by users. We propose a quality attention generative adversarial network (QAGAN) trained on unpaired data based on the bidirectional Generative Adversarial Network (GAN) embedded with a quality attention module (QAM). The key novelty of the proposed QAGAN lies in the injected QAM for the generator such that it learns domain-relevant quality attention directly from the two domains. More specifically, the proposed QAM allows the generator to effectively select semantic-related characteristics from the spatial-wise and adaptively incorporate style-related attributes from the channel-wise, respectively. Therefore, in our proposed QAGAN, not only discriminators but also the generator can directly access both domains which significantly facilitate the generator to learn the mapping function. Extensive experimental results show that, compared with the state-of-the-art methods based on unpaired learning, our proposed method achieves better performance in both objective and subjective evaluations. Zhangkai Ni, Wenhan Yang, Shiqi Wang 0001, Lin Ma 0002, Sam Kwong |
ACM Multimedia | 2 |
| 2020 | Sensitivity-Aware Bit Allocation for Intermediate Deep Feature CompressionabstractIn this paper, we focus on compressing and transmitting deep intermediate features to support the prosperous applications at the cloud side efficiently, and propose a sensitivity-aware bit allocation algorithm for the deep intermediate feature compression. Considering that different channels' contributions to the final inference result of the deep learning model might differ a lot, we design a channel-wise bit allocation mechanism to maintain the accuracy while trying to reduce the bit-rate cost. The algorithm consists of two passes. In the first pass, only one channel is exposed to compression degradation while other channels are kept as the original ones in order to test this channel's sensitivity to the compression degradation. This process will be repeated until all channels' sensitivity is obtained. Then, in the second pass, bits allocated to each channel will be automatically decided according to the sensitivity obtained in the first pass to make sure that the channel with higher sensitivity can be allocated with more bits to maintain accuracy as much as possible. With the well-designed algorithm, our method surpasses state-of-the-art compression tools with on average 6.4% BD-rate saving. Yuzhang Hu, Sifeng Xia, Wenhan Yang, Jiaying Liu 0001 |
VCIP | 3 |
| 2020 | Joint Rain Detection and Removal from a Single Image with Contextualized Deep NetworksabstractRain streaks, particularly in heavy rain, not only degrade visibility but also make many computer vision algorithms fail to function properly. In this paper, we address this visibility problem by focusing on single-image rain removal, even in the presence of dense rain streaks and rain-streak accumulation, which is visually similar to mist or fog. To achieve this, we introduce a new rain model and a deep learning architecture. Our rain model incorporates a binary rain map indicating rain-streak regions, and accommodates various shapes, directions, and sizes of overlapping rain streaks, as well as rain accumulation, to model heavy rain. Based on this model, we construct a multi-task deep network, which jointly learns three targets: the binary rain-streak map, rain streak layers, and clean background, which is our ultimate output. To generate features that can be invariant to rain steaks, we introduce a contextual dilated network, which is able to exploit regional contextual information. To handle various shapes and directions of overlapping rain streaks, our strategy is to utilize a recurrent process that progressively removes rain streaks. Our binary map provides a constraint and thus additional information to train our network. Extensive evaluation on real images, particularly in heavy rain, shows the effectiveness of our model and architecture. Wenhan Yang, Robby T. Tan, Jiashi Feng, Zongming Guo, Shuicheng Yan, Jiaying Liu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2020 | Video Coding for Machines: A Paradigm of Collaborative Compression and Intelligent AnalyticsabstractVideo coding, which targets to compress and reconstruct the whole frame, and feature compression, which only preserves and transmits the most critical information, stand at two ends of the scale. That is, one is with compactness and efficiency to serve for machine vision, and the other is with full fidelity, bowing to human perception. The recent endeavors in imminent trends of video compression, e.g. deep learning based coding tools and end-to-end image/video coding, and MPEG-7 compact feature descriptor standards, i.e. Compact Descriptors for Visual Search and Compact Descriptors for Video Analysis, promote the sustainable and fast development in their own directions, respectively. In this paper, thanks to booming AI technology, e.g. prediction and generation models, we carry out exploration in the new area, Video Coding for Machines (VCM), arising from the emerging MPEG standardization efforts1. Towards collaborative compression and intelligent analytics, VCM attempts to bridge the gap between feature coding for machine vision and video coding for human vision. Aligning with the rising Analyze then Compress instance Digital Retina, the definition, formulation, and paradigm of VCM are given first. Meanwhile, we systematically review state-of-the-art techniques in video compression and feature compression from the unique perspective of MPEG standardization, which provides the academic and industrial evidence to realize the collaborative compression of video and feature streams in a broad range of AI applications. Finally, we come up with potential VCM solutions, and the preliminary results have demonstrated the performance and efficiency gains. Further direction is discussed as well. Ling-Yu Duan, Jiaying Liu 0001, Wenhan Yang, Tiejun Huang 0001, Wen Gao 0001 |
IEEE Trans. Image Process. | 3 |
| 2020 | A Comprehensive Benchmark for Single Image Compression Artifact ReductionabstractWe present a comprehensive study and evaluation of existing single image compression artifact removal algorithms using a new 4K resolution benchmark. This benchmark is called the Large-Scale Ideal Ultra high-definition 4K (LIU4K), and it includes including diversified foreground objects and background scenes with rich structures. Compression artifact removal, as a common post-processing technique, aims at alleviating undesirable artifacts, such as blockiness, ringing, and banding caused by quantization and approximation in the compression process. In this work, a systematic listing of the reviewed methods is presented based on their basic models (handcrafted models and deep networks). The main contributions and novelties of these methods are highlighted, and the main development directions are summarized, including architectures, multi-domain sources, signal structures, and new targeted units. Furthermore, based on a unified deep learning configuration (i.e.same training data, loss function, optimization algorithm,etc.), we evaluate recent deep learning-based methods based on diversified evaluation measures. The experimental results show state-of-the-art performance comparisons of existing methods based on both full-reference, non-reference, and task-driven metrics. Our survey gives a comprehensive reference source for future research on single image compression artifact removal and inspires new directions in related fields. Jiaying Liu 0001, Dong Liu 0002, Wenhan Yang, Sifeng Xia, Xiaoshuai Zhang, Yuanying Dai |
IEEE Trans. Image Process. | 3 |
| 2020 | Towards Unsupervised Deep Image Enhancement With Generative Adversarial NetworkabstractImproving the aesthetic quality of images is challenging and eager for the public. To address this problem, most existing algorithms are based on supervised learning methods to learn an automatic photo enhancer for paired data, which consists of low-quality photos and corresponding expert-retouched versions. However, the style and characteristics of photos retouched by experts may not meet the needs or preferences of general users. In this paper, we present an unsupervised image enhancement generative adversarial network (UEGAN), which learns the corresponding image-to-image mapping from a set of images with desired characteristics in an unsupervised manner, rather than learning on a large number of paired images. The proposed model is based on single deep GAN which embeds the modulation and attention mechanisms to capture richer global and local features. Based on the proposed model, we introduce two losses to deal with the unsupervised image enhancement: (1) fidelity loss, which is defined as a l2 regularization in the feature domain of a pre-trained VGG network to ensure the content between the enhanced image and the input image is the same, and (2) quality loss that is formulated as a relativistic hinge adversarial loss to endow the input image the desired characteristics. Both quantitative and qualitative results show that the proposed model effectively improves the aesthetic quality of images. Zhangkai Ni, Wenhan Yang, Shiqi Wang 0001, Lin Ma 0002, Sam Kwong |
IEEE Trans. Image Process. | 2 |
| 2020 | LR3M: Robust Low-Light Enhancement via Low-Rank Regularized Retinex ModelabstractNoise causes unpleasant visual effects in low-light image/video enhancement. In this paper, we aim to make the enhancement model and method aware of noise in the whole process. To deal with heavy noise which is not handled in previous methods, we introduce a robust low-light enhancement approach, aiming at well enhancing low-light images/videos and suppressing intensive noise jointly. Our method is based on the proposed Low-Rank Regularized Retinex Model (LR3M), which is the first to inject low-rank prior into a Retinex decomposition process to suppress noise in the reflectance map. Our method estimates a piece-wise smoothed illumination and a noise-suppressed reflectance sequentially, avoiding remaining noise in the illumination and reflectance maps which are usually presented in alternative decomposition methods. After getting the estimated illumination and reflectance, we adjust the illumination layer and generate our enhancement result. Furthermore, we apply our LR3M to video low-light enhancement. We consider inter-frame coherence of illumination maps and find similar patches through reflectance maps of successive frames to form the low-rank prior to make use of temporal correspondence. Our method performs well for a wide variety of images and videos, and achieves better quality both in enhancing and denoising, compared with the state-of-the-art methods. Xutong Ren, Wenhan Yang, Wen-Huang Cheng, Jiaying Liu 0001 |
IEEE Trans. Image Process. | 2 |
| 2020 | Removing Arbitrary-Scale Rain Streaks via Fractal Band Learning With Self-SupervisionabstractData-driven rain streak removal methods, most of which rely on synthesized paired data, usually come across the generalization problem when being applied in real scenarios. In this paper, we propose a novel deep-learning based rain streak removal method injected with self-supervision to obtain the capacity of removing more varied-scale rain streaks in practical applications. To this end, in this work, efforts are made from two perspectives. First, considering that rain streak removal is highly correlated with texture characteristics, we create a fractal band learning (FBL) network based on frequency band recovery. It integrates commonly seen band feature operations as neural forms and effectively improves the capacity to capture discriminative features for deraining. Second, to further improve the generalization ability of FBL to remove rain streaks of varied scales, we incorporate scale-robust self-supervision to regularize the network training. The constraint forces the extracted features of an input rain image at different scales to be equivalent after rescaling operations. Therefore, our method can offer similar responses based on solely image content without the interference of scale change and is capable to remove varied-scale rain streaks. Extensive experiments in quantitative and qualitative evaluations demonstrate the superiority of our method for rain streak removal, especially for the real cases where very large rain streaks exist, and prove the effectiveness of each component. Wenhan Yang, Shiqi Wang 0001, Jiaying Liu 0001 |
IEEE Trans. Image Process. | 1 |
| 2020 | Advancing Image Understanding in Poor Visibility Environments: A Collective Benchmark StudyabstractExisting enhancement methods are empirically expected to help the high-level end computer vision task: however, that is observed to not always be the case in practice. We focus on object or face detection in poor visibility enhancements caused by bad weathers (haze, rain) and low light conditions. To provide a more thorough examination and fair comparison, we introduce three benchmark sets collected in real-world hazy, rainy, and low-light conditions, respectively, with annotated objects/faces. We launched the UG2+ challenge Track 2 competition in IEEE CVPR 2019, aiming to evoke a comprehensive discussion and exploration about whether and how low-level vision techniques can benefit the high-level automatic visual recognition in various scenarios. To our best knowledge, this is the first and currently largest effort of its kind. Baseline results by cascading existing enhancement and detection models are reported, indicating the highly challenging nature of our new data as well as the large room for further technical innovations. Thanks to a large participation from the research community, we are able to analyze representative team solutions, striving to better identify the strengths and limitations of existing mindsets as well as the future directions. Wenhan Yang, Ye Yuan 0012, Wenqi Ren, Jiaying Liu 0001, Walter J. Scheirer, Zhangyang Wang, Taiheng Zhang, Qiaoyong Zhong, Di Xie, Shiliang Pu, Yuqiang Zheng, Yanyun Qu, Yuhong Xie, Hao Jiang 0014, Siyuan Yang 0001, Yan Liu 0041, Xiaochao Qu, Pengfei Wan 0001, Shuai Zheng 0005, Minhui Zhong, Taiyi Su, Lingzhi He, Yandong Guo, Yao Zhao 0001, Zhenfeng Zhu, Jinxiu Liang, Jingwen Wang 0003, Yuhui Quan, Yong Xu 0007, Bo Liu 0112, Xin Liu 0012, Tingyu Lin 0003, Xiaochuan Li 0001, Feng Lu 0005, Lin Gu 0003, Shengdi Zhou, Cong Cao 0005, Cheng Chi 0003, Chubin Zhuang, Zhen Lei 0001, Stan Z. Li, Shizheng Wang, Ruizhe Liu, Dong Yi, Zheming Zuo, Jianning Chi, Huan Wang 0014, Kai Wang 0036, Yixiu Liu, Xingyu Gao 0001, Zhenyu Chen 0003, Yongzhou Li, Huicai Zhong, Jing Huang 0017, Heng Guo 0003, Jianfei Yang 0001, Wenjuan Liao, Jiangang Yang, Liguo Zhou, Mingyue Feng, Likun Qin |
IEEE Trans. Image Process. | 1 |
| 2020 | Deep Reference Generation With Multi-Domain Hierarchical Constraints for Inter PredictionabstractInter prediction is an important module in video coding for temporal redundancy removal, where similar reference blocks are searched from previously coded frames and employed to predict the block to be coded. Although existing video codecs can estimate and compensate for block-level motions, their inter prediction performance is still heavily affected by the remaining inconsistent pixel-wise displacement caused by irregular rotation and deformation. In this paper, we address the problem by proposing a deep frame interpolation network to generate additional reference frames in coding scenarios. First, we summarize the previous adaptive convolutions used for frame interpolation and propose a factorized kernel convolutional network to improve the modeling capacity and simultaneously keep its compact form. Second, to better train this network, multi-domain hierarchical constraints are introduced to regularize the training of our factorized kernel convolutional network. For spatial domain, we use a gradually down-sampled and up-sampled auto-encoder to generate the factorized kernels for frame interpolation at different scales. For quality domain, considering the inconsistent quality of the input frames, the factorized kernel convolution is modulated with quality-related features to learn to exploit more information from high quality frames. For frequency domain, a sum of absolute transformed difference loss that performs frequency transformation is utilized to facilitate network optimization from the view of coding performance. With the well-designed frame interpolation network regularized by multi-domain hierarchical constraints, our method surpasses HEVC on average 3.8% BD-rate saving for the luma component under the random access configuration and also obtains on average 0.83% BD-rate saving over the upcoming VVC. Jiaying Liu 0001, Sifeng Xia, Wenhan Yang |
IEEE Trans. Multim. | 3 |
| 2019 | Frame-Consistent Recurrent Video Deraining With Dual-Level FlowabstractIn this paper, we address the problem of rain removal from videos by proposing a more comprehensive framework that considers the additional degradation factors in real scenes neglected in previous works. The proposed framework is built upon a two-stage recurrent network with dual-level flow regularizations to perform the inverse recovery process of the rain synthesis model for video deraining. The rain-free frame is estimated from the single rain frame at the first stage. It is then taken as guidance along with previously recovered clean frames to help obtain a more accurate clean frame at the second stage. This two-step architecture is capable of extracting more reliable motion information from the initially estimated rain-free frame at the first stage for better frame alignment and motion modeling at the second stage. Furthermore, to keep the motion consistency between frames that facilitates a frame-consistent deraining model at the second stage, a dual-level flow based regularization is proposed at both coarse flow and fine pixel levels. To better train and evaluate the proposed video deraining network, a novel rain synthesis model is developed to produce more visually authentic paired training and evaluation videos. Extensive experiments on a series of synthetic and real videos verify not only the superiority of the proposed method over state-of-the-art but also the effectiveness of network design and its each component. Wenhan Yang, Jiaying Liu 0001, Jiashi Feng |
CVPR | 1 |
| 2019 | Partition Tree Guided Progressive Rethinking Network for in-Loop Filtering of HEVCabstractIn-Loop filter is a key part in High Efficiency Video Coding (HEVC) which effectively removes the compression artifacts. Recently, many newly proposed methods combine residual learning and dense connection to construct a deeper network for better in-loop filtering performance. However, the long-term dependency between blocks is neglected, and information usually passes between blocks only after dimension compression. To address these issues, we propose the Progressive Rethinking Block (PRB) to deliver long-term memory between the neighboring blocks and allow information to flow without compression, which is similar to human decision mechanism - usually reviewing the complete past memorized experiences to decide in the present, not just based on simple principles summarized before. PRBs further establish the Progressive Rethinking Network (PRN). In addition, we calculate the Multi-scale Mean value of Coding Units (MM-CU) to generate the side information maps which guide the training of the network by novelly telling the network architecture of the entire coding partition tree. Experimental results show that our proposed partition tree guided PRN provides 10.1% BD-rate reduction on average compared to the HEVC baseline. Dezhao Wang, Sifeng Xia, Wenhan Yang, Yueyu Hu, Jiaying Liu 0001 |
ICIP | 3 |
| 2019 | Deep Inter Prediction Via Pixel-Wise Motion Oriented Reference GenerationabstractInter prediction is an important module in video coding for temporal redundancy removal, where the reference blocks are searched from the previously coded frames and employed to predict the block to be coded. However, apart from regular block-wise shift motion, there usually exists inconsistent pixel-wise motion such as rotation and deformation between blocks, which will largely degrade the prediction performance. In this paper, we propose a Multiscale Adaptive Separable Convolutional Neural Network (MASCNN) to generate pixel-wise closer reference frames for inter prediction. A multiscale network is built to interpolate the target frame from coarse to fine. Reconstruction losses are enforced on each scale to make the network infer the main structure at small scales, which improves the interpolation accuracy of the network. Furthermore, a sum of absolute transformed difference (SATD) loss function is proposed to regularize the network training, which further improves the coding performance. Compared with HEVC, our method can obtain on average 5.7% BD-rate saving and up to 9.9% BD-rate saving for the luma component under the random access configuration. Sifeng Xia, Wenhan Yang, Yueyu Hu, Jiaying Liu 0001 |
ICIP | 2 |
| 2019 | Deep Pyramid Variation Learning for Image InterpolationabstractPrevious learning-based interpolation methods do not consider multi-scale structural information, which is generally effective for image modeling. In this paper, we design a deep network based on a novel pyramid variation learning approach with multi-scale structure modeling. An image is represented as multi-dimensional features. Besides two spatial dimensions, the features include a neighboring variation dimension where every pixel is encoded as the variation to its nearest low-resolution pixel, and a scale dimension along which the feature maps generated by a gradual down-sampling process are stacked. Thus, these multi-dimensional features are constructed to model local dependency and multi-scale similarity jointly. Inspired by this feature design, we build an end-to-end trainable Recurrent Multi-Path Aggregation Network (RMPAN) for image interpolation, where the scale dimension is unfolded to form a multi-path aggregation network to apply joint filters at different scales recurrently. Location aware sampling layers are used in RMPAN to transform feature maps into different scales with only location changes in each convolution path, which aggregate the context information without resolution loss. Comprehensive experiments demonstrate that our method leads to a superior performance and offers new state-of-the-art benchmark. Wenhan Yang, Jiaying Liu 0001 |
ICME | 2 |
| 2019 | Switch Mode Based Deep Fractional Interpolation in Video CodingabstractFractional interpolation is a significant technology in motion compensation of video coding. It generates sub-pixel level reference samples in inter prediction to facilitate temporal redundancy removal between video frames. Recently, some methods explore to introduce the deep learning technique for fractional interpolation and have obtained better compression results. However, existing deep learning based methods still treat fractional interpolation as a traditional interpolation problem but fail to adjust it to the motion compensation scenario. In this paper, we design a switch mode based deep fractional interpolation method to introduce integer pixels of different positions to the interpolation of sub-pixel position samples. By switching between integer pixels of different positions, our method can infer the sub-pixels with smaller variations and achieve better fractional interpolation results. Consequently the motion compensation performance can be further improved. Experimental results have also verified the efficiency of the switch mode based deep fractional interpolation. Compared with High Efficiency Video Coding, our method achieves 2.8% bit saving on average and up to 6.2% bit saving under low-delay P configuration. Sifeng Xia, Wenhan Yang, Yueyu Hu, Wen-Huang Cheng, Jiaying Liu 0001 |
ISCAS | 2 |
| 2019 | Reference-Guided Deep Super-Resolution via Manifold Localized External CompensationabstractThe rapid development of social network and online multimedia technology makes it possible to address traditional image and video enhancement problems, with the aid of online similar reference data. In this paper, we tackle the problem of super-resolution (SR) in this way, specifically aiming to handle the “one-to-many” problem between the image patches of low resolution (LR) and high resolution (HR). We propose a manifold localized deep external compensation (MALDEC) network to additionally utilize reference images, i. e., retrieved similar images in cloud database and reference HR frame in a video, to provide an accurate localization and mapping to the HR manifold, and compensate the lost high-frequency details. The proposed network employs a three-step recovery: 1) internal structure inference, which uses the LR image itself and the internally inferred high frequency information to preserve main structure of the HR image; 2) manifold localization, which localizes the HR manifold and constructs the correspondence between the internal inferred image and the external images; and 3) external compensation, which introduces the external references of retrieved similar patches based on manifold localization information to reconstruct the high-frequency details. The learnable components of MALDEC, internal structure inference, and external compensation, are trained jointly to make a good tradeoff between these two terms for an optimal SR result. Finally, the proposed method is examined under three tasks: cloud-based image SR, multi-pose face reconstruction, and reference frame-guided video SR. Extensive experiments demonstrate the superiority of our method than the state-of-the-art SR methods in both objective and subjective evaluations, and our method offers new state-of-the-art performance. Wenhan Yang, Sifeng Xia, Jiaying Liu 0001, Zongming Guo |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2019 | One-for-All: Grouped Variation Network-Based Fractional Interpolation in Video CodingabstractFractional interpolation is used to provide sub-pixel level references for motion compensation in the interprediction of video coding, which attempts to remove temporal redundancy in video sequences. Traditional handcrafted fractional interpolation filters face the challenge of modeling discontinuous regions in videos, while existing deep learning-based methods are either designed for a single quantization parameter (QP), only generating half-pixel samples, or need to train a model for each sub-pixel position. In this paper, we present a one-for-all fractional interpolation method based on a grouped variation convolutional neural network (GVCNN). Our method can deal with video frames coded using different QPs and is capable of generating all sub-pixel positions at one sub-pixel level. Also, by predicting variations between integer-position pixels and sub-pixels, our network offers more expressive power. Moreover, we perform specific measurements in training data generation to simulate practical situations in video coding, including blurring the down-sampled sub-pixel samples to avoid aliasing effects and coding integer pixels to simulate reconstruction errors. In addition, we analyze the impact of the size of blur kernels theoretically. Experimental results verify the efficiency of GVCNN. Compared with HEVC, our method achieves 2.2% in bit saving on average and up to 5.2% under low-delay P configuration. Jiaying Liu 0001, Sifeng Xia, Wenhan Yang, Mading Li, Dong Liu 0002 |
IEEE Trans. Image Process. | 3 |
| 2019 | D3R-Net: Dynamic Routing Residue Recurrent Network for Video Rain RemovalabstractIn this paper, we address the problem of video rain removal by considering rain occlusion regions, i.e., very low light transmittance for rain streaks. Different from additive rain streaks, in such occlusion regions, the details of backgrounds are completely lost. Therefore, we propose a hybrid rain model to depict both rain streaks and occlusions. Integrating the hybrid model and useful motion segmentation context information, we present a Dynamic Routing Residue Recurrent Network (D3R-Net). D3R-Net first extracts the spatial features by a residual network. Then, the spatial features are aggregated by recurrent units along the temporal axis. In the temporal fusion, the context information is embedded into the network in a "dynamic routing" way. A heap of recurrent units takes responsibility for handling the temporal fusion in given contexts, e.g., rain or non-rain regions. In the certain forward and backward processes, one of these recurrent units is mainly activated. Then, a context selection gate is employed to detect the context and select one of these temporally fused features generated by these recurrent units as the final fused feature. Finally, this last feature plays a role of "residual feature." It is combined with the spatial feature and then used to reconstruct the negative rain streaks. In such a D3R-Net, we incorporate motion segmentation, which denotes whether a pixel belongs to fast moving edges or not, and rain type indicator, indicating whether a pixel belongs to rain streaks, rain occlusions, and non-rain regions, as the context variables. Extensive experiments on a series of synthetic and real videos with rain streaks verify not only the superiority of the proposed method over state of the art but also the effectiveness of our network design and its each component. Jiaying Liu 0001, Wenhan Yang, Shuai Yang 0001, Zongming Guo |
IEEE Trans. Image Process. | 2 |
| 2019 | Context-Aware Text-Based Binary Image Stylization and SynthesisabstractIn this work, we present a new framework for the stylization of text-based binary images. First, our method stylizes the stroke-based geometric shape like text, symbols and icons in the target binary image based on an input style image. Second, the composition of the stylized geometric shape and a background image is explored. To accomplish the task, we propose legibilitypreserving structure and texture transfer algorithms, which progressively narrow the visual differences between the binary image and the style image. The stylization is then followed by a contextaware layout design algorithm, where cues for both seamlessness and aesthetics are employed to determine the optimal layout of the shape in the background. Given the layout, the binary image is seamlessly embedded into the background by texture synthesis under a context-aware boundary constraint. According to the contents of binary images, our method can be applied to many fields.We show that the proposed method is capable of addressing the unsupervised text stylization problem and is superior to stateof- the-art style transfer methods in automatic artistic typography creation. Besides, extensive experiments on various tasks, such as visual-textual presentation synthesis, icon/symbol rendering and structure-guided image inpainting, demonstrate the effectiveness of the proposed method. Shuai Yang 0001, Jiaying Liu 0001, Wenhan Yang, Zongming Guo |
IEEE Trans. Image Process. | 3 |
| 2019 | Scale-Free Single Image Deraining Via Visibility-Enhanced Recurrent Wavelet LearningabstractIn this paper, we address a rain removal problem from a single image, even in the presence of large rain streaks and rain streak accumulation (where individual streaks cannot be seen, and thus visually similar to mist or fog). For rain streak removal, the mismatch problem between different streak sizes in training and testing phases leads to a poor performance, especially when there are large streaks. To mitigate this problem, we embed a hierarchical representation of wavelet transform into a recurrent rain removal process: 1) rain removal on the low-frequency component; 2) recurrent detail recovery on highfrequency components under the guidance of the recovered lowfrequency component. Benefiting from the recurrent multi-scale modeling of wavelet transform-like design, the proposed network trained on streaks with one size can adapt to those with larger sizes, which significantly favors real rain streak removal. The dilated residual dense network is used as the basic model of the recurrent recovery process. The network includes multiple paths with different receptive fields, thus can make full use of multi-scale redundancy and utilize context information in large regions. Furthermore, to handle heavy rain cases where rain streak accumulation is presented, we construct a detail appearing rain accumulation removal to not only improve the visibility but also enhance the details in dark regions. The evaluation on both synthetic and real images, particularly on those containing large rain streaks and heavy accumulation, shows the effectiveness of our novel models, which significantly outperforms the state-ofthe- art methods. Wenhan Yang, Jiaying Liu 0001, Shuai Yang 0001, Zongming Guo |
IEEE Trans. Image Process. | 1 |
| 2019 | Progressive Spatial Recurrent Neural Network for Intra PredictionabstractIntra prediction is an important component of modern video codecs, which is able to efficiently squeeze out the spatial redundancy in video frames. With preceding pixels as the context, traditional intra prediction schemes generate linear predictions based on several predefined directions (i.e., modes) for blocks to be encoded. However, these modes are relatively simple and their predictions may fail when facing blocks with complex textures, which leads to additional bits encoding the residue. In this paper, we design a progressive spatial recurrent neural network (PS-RNN) that learns to conduct intra prediction. Specifically, our PS-RNN consists of three spatial recurrent units and progressively generates predictions by passing information along from preceding contents to blocks to be encoded. To make our network generate predictions considering both distortion and bit rate, we propose using sum of absolute transformed difference (SATD) as the loss function to train PS-RNN since SATD is able to measure rate-distortion cost of encoding a residue block. Moreover, our method supports variable-block-size for intra prediction, which is more practical in real coding conditions. The proposed intra prediction scheme achieves on average 2.5% bit-rate reduction on variable-block-size settings under the same reconstruction quality compared with HEVC. Yueyu Hu, Wenhan Yang, Mading Li, Jiaying Liu 0001 |
IEEE Trans. Multim. | 2 |
| 2018 | Deep Retinex Decomposition for Low-Light Enhancement
Chen Wei 0005, Wenjing Wang 0001, Wenhan Yang, Jiaying Liu 0001 |
BMVC | 3 |
| 2018 | Erase or Fill? Deep Joint Recurrent Rain Removal and Reconstruction in VideosabstractIn this paper, we address the problem of video rain removal by constructing deep recurrent convolutional networks. We visit the rain removal case by considering rain occlusion regions, i.e. the light transmittance of rain streaks is low. Different from additive rain streaks, in such rain occlusion regions, the details of background images are completely lost. Therefore, we propose a hybrid rain model to depict both rain streaks and occlusions. With the wealth of temporal redundancy, we build a Joint Recurrent Rain Removal and Reconstruction Network (J4R-Net) that seamlessly integrates rain degradation classification, spatial texture appearances based rain removal and temporal coherence based background details reconstruction. The rain degradation classification provides a binary map that reveals whether a location is degraded by linear additive streaks or occlusions. With this side information, the gate of the recurrent unit learns to make a trade-off between rain streak removal and background details reconstruction. Extensive experiments on a series of synthetic and real videos with rain streaks verify the superiority of the proposed method over previous state-of-the-art methods. Jiaying Liu 0001, Wenhan Yang, Shuai Yang 0001, Zongming Guo |
CVPR | 2 |
| 2018 | Attentive Generative Adversarial Network for Raindrop Removal From a Single ImageabstractRaindrops adhered to a glass window or camera lens can severely hamper the visibility of a background scene and degrade an image considerably. In this paper, we address the problem by visually removing raindrops, and thus transforming a raindrop degraded image into a clean one. The problem is intractable, since first the regions occluded by raindrops are not given. Second, the information about the background scene of the occluded regions is completely lost for most part. To resolve the problem, we apply an attentive generative network using adversarial training. Our main idea is to inject visual attention into both the generative and discriminative networks. During the training, our visual attention learns about raindrop regions and their surroundings. Hence, by injecting this information, the generative network will pay more attention to the raindrop regions and the surrounding structures, and the discriminative network will be able to assess the local consistency of the restored regions. This injection of visual attention to both generative and discriminative networks is the main contribution of this paper. Our experiments show the effectiveness of our approach, which outperforms the state of the art methods quantitatively and qualitatively. Rui Qian 0003, Robby T. Tan, Wenhan Yang, Jiajun Su, Jiaying Liu 0001 |
CVPR | 3 |
| 2018 | Enhanced Intra Prediction with Recurrent Neural Network in Video CodingabstractIntra prediction is one of the important parts in video/image codec. With intra prediction mechanism, spatial redundancy can be largely removed for further bit saving. However, current state-of-the-art intra prediction method does not produce satisfactory prediction result due to its limits in reference samples and modeling ability. To enhance the intra prediction in HEVC, in this paper, a deep neural network featuring spatial RNN, which models the spatial dependency of pixels as sequential dynamics, is proposed to generate better prediction signals. Experimental results show improvement in BD-Rate for the proposed method compared with the original HEVC prediction scheme. Yueyu Hu, Wenhan Yang, Sifeng Xia, Wen-Huang Cheng, Jiaying Liu 0001 |
DCC | 2 |
| 2018 | A Group Variational Transformation Neural Network for Fractional Interpolation of Video CodingabstractMotion compensation is an important technology in video coding to remove the temporal redundancy between coded video frames. In motion compensation, fractional interpolation is used to obtain more reference blocks at sub-pixel level. Existing video coding standards commonly use fixed interpolation filters for fractional interpolation, which are not efficient enough to handle diverse video signals well. In this paper, we design a group variational transformation convolutional neural network (GVTCNN) to improve the fractional interpolation performance of the luma component in motion compensation. GVTCNN infers samples at different sub-pixel positions from the input integer-position sample. It first extracts a shared feature map from the integer-position sample to infer various sub-pixel position samples. Then a group variational transformation technique is used to transform a group of copied shared feature maps to samples at different sub-pixel positions. Experimental results have identified the interpolation efficiency of our GVTCNN. Compared with the interpolation method of High Efficiency Video Coding, our method achieves 1.9% bit saving on average and up to 5.6% bit saving under low-delay P configuration. Sifeng Xia, Wenhan Yang, Yueyu Hu, Siwei Ma 0001, Jiaying Liu 0001 |
DCC | 2 |
| 2018 | GLADNet: Low-Light Enhancement Network with Global AwarenessabstractIn this paper, we address the problem of lowlight enhancement. Our key idea is to first calculate a global illumination estimation for the low-light input, then adjust the illumination under the guidance of the estimation and supplement the details using a concatenation with the original input. Considering that, we propose a GLobal illumination Aware and Detail-preserving Network (GLADNet). The input image is rescaled to a certain size and then put into an encoder-decoder network to generate global priori knowledge of the illumination. Based on the global prior and the original input image, a convolutional network is employed for detail reconstruction. For training GLADNet, we use a synthetic dataset generated from RAW images. Extensive experiments demonstrate the superiority of our method over other compared methods on the real low-light images captured in various conditions. Wenjing Wang 0001, Chen Wei 0005, Wenhan Yang, Jiaying Liu 0001 |
FG | 3 |
| 2018 | Dmcnn: Dual-Domain Multi-Scale Convolutional Neural Network for Compression Artifacts RemovalabstractJPEG is one of the most commonly used standards among lossy image compression methods. However, JPEG compression inevitably introduces various kinds of artifacts, especially at high compression rates, which could greatly affect the Quality of Experience (QoE). Recently, convolutional neural network (CNN) based methods have shown excellent performance for removing the JPEG artifacts. Lots of efforts have been made to deepen the CNN s and extract deeper features, while relatively few works pay attention to the receptive field of the network. In this paper, we illustrate that the quality of output images can be significantly improved by enlarging the receptive fields in many cases. One step further, we propose a Dual-domain Multi-scale CNN (DMCNN) to take full advantage of redundancies on both the pixel and DCT domains. Experiments show that DMCNN sets a new state-of-the-art for the task of JPEG artifact removal. Xiaoshuai Zhang, Wenhan Yang, Yueyu Hu, Jiaying Liu 0001 |
ICIP | 2 |
| 2018 | Dual Recovery Network with Online Compensation for Image Super-ResolutionabstractImage super-resolution (SR) methods essentially lead to a loss of some high-frequency (HF) information when predicting high-resolution (HR) images from low-resolution (LR) images without using external references. To address this issue, we additionally utilize online retrieved data to facilitate image SR in a unified deep framework. A novel dual high-frequency recovery network (DHN) is proposed to predict an HR image with three parts: an LR image, an internal inferred HF (IHF) map (HF missing part inferred solely from the LR image) and an external extracted HF (EHF) map. In particular, we infer the HF information based on both the LR image and similar HR references which are retrieved online. For the EHF map, we align the references with affine transformation and then in the aligned references, part of HF signals are extracted by the proposed DHN to compensate for the HF loss. Extensive experimental results demonstrate that our DHN achieves notably better performance than state-of-the-art SR methods. Sifeng Xia, Wenhan Yang, Jiaying Liu 0001, Zongming Guo |
ISCAS | 2 |
| 2018 | Context-Aware Unsupervised Text StylizationabstractIn this work, we present a novel algorithm to stylize the text without supervision, which provides a flexible and convenient way to invoke fantastic text expressions. Rather than employing the fixed pair of target text and source style images, our unsupervised framework establishes an implicit mapping for them by using an abstract imagery of the style image as bridges. Based on the mapping, we progressively narrow the visual discrepancy between text and style images by the proposed legibility-preserving structure transfer and texture transfer algorithms, which effectively balance the text legibility and style consistency. Furthermore, we explore a seamless composition of the stylized text and a background image, in which the optimal text layout is determined by a context-aware layout design algorithm utilizing cues for both seamlessness and aesthetics. Given the layout, the text can be seamlessly embedded into the background by texture synthesis under a context-aware boundary constraint. Experimental results demonstrate the effectiveness of the proposed method in automatic artistic typography creation and visual-textual presentation synthesis. Shuai Yang 0001, Jiaying Liu 0001, Wenhan Yang, Zongming Guo |
ACM Multimedia | 3 |
| 2018 | Optimized Spatial Recurrent Network for Intra Prediction in Video CodingabstractIntra prediction in modern video codecs is able to efficiently reduce spatial redundancy in video frames. With preceding pixels as context, traditional intra prediction schemes generate linear predictions based on several predefined directions (i.e. modes) for the current prediction unit (PU). However, these modes are relatively simple and are not able to handle complex textures, which leads to additional bits encoding the residue. In this paper, we design a convolutional neural network (CNN) guided spatial recurrent neural network (RNN) to improve the intra prediction in High-Efficiency Video Coding (HEVC). By exploring the correlations between pixels, the network learns to generate prediction signal in a progressive manner. The progressive model solves the problem of asymmetry in intra prediction naturally. As the model is designed for global context modeling, no flags for intra prediction modes selection need to be encoded. Our proposed intra prediction scheme achieves on average 1.2% bit-rate saving compared with HEVC. Yueyu Hu, Wenhan Yang, Sifeng Xia, Jiaying Liu 0001 |
VCIP | 2 |
| 2018 | Video super-resolution based on spatial-temporal recurrent residual networks
Wenhan Yang, Jiashi Feng, Guosen Xie, Jiaying Liu 0001, Zongming Guo, Shuicheng Yan |
Comput. Vis. Image Underst. | 1 |
| 2018 | Blind visual quality assessment for image super-resolution by convolutional neural network
Yuming Fang 0001, Chi Zhang 0027, Wenhan Yang, Jiaying Liu 0001, Zongming Guo |
Multim. Tools Appl. | 3 |
| 2018 | Automatic portrait oil painter: joint domain stylization for portrait images
Saboya Yang, Shuai Yang 0001, Wenhan Yang, Jiaying Liu 0001 |
Multim. Tools Appl. | 3 |
| 2018 | Isophote-Constrained Autoregressive Model With Adaptive Window Extension for Image InterpolationabstractThe autoregressive (AR) model is widely used in image interpolations. Traditional AR models consider utilizing the dependence between pixels to model the image signal. However, they ignore the valuable patch-level information for image modeling. In this paper, we propose to integrate both the pixel-level and patch-level information to depict the relationship between high-resolution and low-resolution pixels and obtain better image interpolation results. In particular, we propose an isophote-constrained AR (ICAR) model to perform AR-flavored interpolation within an identified joint stable region and further develop an AR interpolation with an adaptive window extension. Considering the smoothness along the isophote curve, the ICAR model searches only several successive similar patches along the isophote curve over a large region to construct an adaptive window. These overlapped patches, representing the patch-level structure similarity, are used to construct a joint AR model. To better characterize the piecewise stationarity and determine whether a pixel is suitable for AR estimation, we further propose pixel-level and patch-level similarity metrics and embed them into the ICAR model, introducing a weighted ICAR model. Comprehensive experiments demonstrate that our method can effectively reconstruct the edge structures and suppress jaggy or ringing artifacts. In the objective quality evaluation, our method achieves the best results in terms of both peak signal-to-noise ratio and structural similarity for both simple size doubling (two times) and for arbitrary scale enlargements. Wenhan Yang, Jiaying Liu 0001, Mading Li, Zongming Guo |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2018 | Structure-Revealing Low-Light Image Enhancement Via Robust Retinex ModelabstractLow-light image enhancement methods based on classic Retinex model attempt to manipulate the estimated illumination and to project it back to the corresponding reflectance. However, the model does not consider the noise, which inevitably exists in images captured in low-light conditions. In this paper, we propose the robust Retinex model, which additionally considers a noise map compared with the conventional Retinex model, to improve the performance of enhancing low-light images accompanied by intensive noise. Based on the robust Retinex model, we present an optimization function that includes novel regularization terms for the illumination and reflectance. Specifically, we use norm to constrain the piece-wise smoothness of the illumination, adopt a fidelity term for gradients of the reflectance to reveal the structure details in low-light images, and make the first attempt to estimate a noise map out of the robust Retinex model. To effectively solve the optimization problem, we provide an augmented Lagrange multiplier based alternating direction minimization algorithm without logarithmic transformation. Experimental results demonstrate the effectiveness of the proposed method in low-light image enhancement. In addition, the proposed method can be generalized to handle a series of similar problems, such as the image enhancement for underwater or remote sensing and in hazy or dusty conditions. Mading Li, Jiaying Liu 0001, Wenhan Yang, Xiaoyan Sun 0001, Zongming Guo |
IEEE Trans. Image Process. | 3 |
| 2018 | Robust LSTM-Autoencoders for Face De-Occlusion in the WildabstractFace recognition techniques have been developed significantly in recent years. However, recognizing faces with partial occlusion is still challenging for existing face recognizers, which is heavily desired in real-world applications concerning surveillance and security. Although much research effort has been devoted to developing face de-occlusion methods, most of them can only work well under constrained conditions, such as all of faces are from a pre-defined closed set of subjects. In this paper, we propose a robust LSTM-Autoencoders (RLA) model to effectively restore partially occluded faces even in the wild. The RLA model consists of two LSTM components, which aims at occlusion-robust face encoding and recurrent occlusion removal respectively. The first one, named multi-scale spatial LSTM encoder, reads facial patches of various scales sequentially to output a latent representation, and occlusion-robustness is achieved owing to the fact that the influence of occlusion is only upon some of the patches. Receiving the representation learned by the encoder, the LSTM decoder with a dual channel architecture reconstructs the overall face and detects occlusion simultaneously, and by feat of LSTM, the decoder breaks down the task of face de-occlusion into restoring the occluded part step by step. Moreover, to minimize identify information loss and guarantee face recognition accuracy over recovered faces, we introduce an identity-preserving adversarial training scheme to further improve RLA. Extensive experiments on both synthetic and real data sets of faces with occlusion clearly demonstrate the effectiveness of our proposed RLA in removing different types of facial occlusion at various locations. The proposed method also provides significantly larger performance gain than other de-occlusion methods in promoting recognition performance over partially-occluded faces. Fang Zhao 0006, Jiashi Feng, Jian Zhao 0006, Wenhan Yang, Shuicheng Yan |
IEEE Trans. Image Process. | 4 |
| 2018 | Photo Stylistic Brush: Robust Style Transfer via Superpixel-Based Bipartite GraphabstractWith the rapid development of social network and multimedia technology, customized image and video stylization have been widely used for various social-media applications. In this paper, we explore the problem of exemplar-based photo style transfer, which provides a flexible and convenient way to invoke fantastic visual impression. Rather than investigating some fixed artistic patterns to represent certain styles as was done in some previous works, our work emphasizes styles related to a series of visual effects in the photograph (e.g., color, tone, and contrast). We propose a photo stylistic brush, an automatic robust style transfer approach based on Super pixel-based BIpartite Graph (SuperBIG). A two-step bipartite graph algorithm with different granularity levels is employed to aggregate pixels into superpixels and find their correspondences. In the first step, with the extracted hierarchical features, a bipartite graph is constructed to describe the content similarity for pixel partition to produce superpixels. In the second step, superpixels in the input/reference image are rematched to form a new superpixel-based bipartite graph, and superpixel-level correspondences are generated by bipartite matching. Finally, the refined correspondence guides SuperBIG to perform the transformation in a decorrelated color space. Extensive experimental results demonstrate the effectiveness and robustness of the proposed method for transferring various styles of exemplar images, even for some challenging cases, such as night images. Jiaying Liu 0001, Wenhan Yang, Xiaoyan Sun 0001, Wenjun Zeng 0001 |
IEEE Trans. Multim. | 2 |
| 2017 | Deep Joint Rain Detection and Removal from a Single ImageabstractIn this paper, we address a rain removal problem from a single image, even in the presence of heavy rain and rain streak accumulation. Our core ideas lie in our new rain image model and new deep learning architecture. We add a binary map that provides rain streak locations to an existing model, which comprises a rain streak layer and a background layer. We create a model consisting of a component representing rain streak accumulation (where individual streaks cannot be seen, and thus visually similar to mist or fog), and another component representing various shapes and directions of overlapping rain streaks, which usually happen in heavy rain. Based on the model, we develop a multi-task deep learning architecture that learns the binary rain streak map, the appearance of rain streaks, and the clean background, which is our ultimate output. The additional binary map is critically beneficial, since its loss function can provide additional strong information to the network. To handle rain streak accumulation (again, a phenomenon visually similar to mist or fog) and various shapes and directions of overlapping rain streaks, we propose a recurrent rain detection and removal network that removes rain streaks and clears up the rain accumulation iteratively and progressively. In each recurrence of our method, a new contextualized dilated network is developed to exploit regional contextual information and to produce better representations for rain detection. The evaluation on real images, particularly on heavy rain, shows the effectiveness of our models and architecture. Wenhan Yang, Robby T. Tan, Jiashi Feng, Jiaying Liu 0001, Zongming Guo, Shuicheng Yan |
CVPR | 1 |
| 2017 | General scale interpolation via context-aware autoregressive model and multiplanar constraintabstractIn this paper, we propose a novel image interpolation algorithm suitable for general scale enlargement. Different from previous AR-based interpolation algorithms which employ predetermined reference configuration to predict pixel values, we consider the context information when building AR models. Optimal references are selected by incorporating nonlocal-based correlation coefficient and the indicator for local edge direction. Furthermore, the multiplanar constraint among similar patches is applied to enhance the correlation within the estimation window and serves as a kind of supplement to data fidelity term in AR model. The experimental results show that our method is effective in several enlargement scales and successfully alleviate the artifacts nearby edges and preserve their sharpness. The comparison experiments demonstrate that the proposed method can obtain desirable performance in terms of both objective and subjective results. Shihong Deng, Jiaying Liu 0001, Mading Li, Wenhan Yang, Zongming Guo |
ICASSP | 4 |
| 2017 | Variation learning guided convolutional network for image interpolationabstractIn this paper, we propose a variational learning model that effectively exploits the structural similarities for image representation, and construct a deep network based on this model for image interpolation. Based on the local dependency, our learning model represents an image as the three-dimensional features. Besides two coordinate dimensions, an additional neighboring variation dimension is added to encode every pixel as the variation to its nearest low-resolution pixel by the local similarity. This added dimension lowers the risk of over-fitting for learning approaches and constructs abundant structural correspondences for inferring the missing information lost in image degradation. Then, this three-dimensional features are naturally modeled, extracted and refined by an end-to-end trainable recurrent convolutional network for image interpolation. Comprehensive experiments demonstrate that our method leads to a surprisingly superior performance and offers new state-of-the-art benchmark. Wenhan Yang, Jiaying Liu 0001, Sifeng Xia, Zongming Guo |
ICIP | 1 |
| 2017 | Soft segmentation-guided bipartite graph image stylizationabstractIn this paper, we propose a photo stylistic brush, an automatic robust style transfer approach based on soft segmentation-guided bipartite graph. A two-step bipartite graph algorithm with different granularity levels is employed to aggregate pixels into superpixel and find their correspondences. In the first step, with the extracted hierarchical features, a bipartite graph is constructed to describe the content similarity for pixel partition to produce superpixels. In the second step, superpixels in the input/reference image are rematched to form a new soft segmentation-guided bipartite graph, and superpixel-level correspondences are generated by a bipartite matching. Finally, the refined correspondence guides our approach to perform the transfer in a decorrelated color space. Extensive experimental results demonstrate the effectiveness and robustness of the proposed method for transferring various styles of exemplar images, even for some challenging cases, such as night images. Saboya Yang, Jiaying Liu 0001, Wenhan Yang, Shuai Yang 0001, Chunpeng Li |
ICIP | 3 |
| 2017 | Joint-domain unsupervised stylization for portraitsabstractPeople wish to own a portrait painting of themselves by Da Vinci. Unfortunately, it is impossible to make this dream come true; nevertheless, it may give us an opportunity by transferring some artistic features from one single reference painting. To address this issue, we propose a joint-domain image stylization approach, particularly for portrait oil paintings. From the view of artistic appreciation, we analyze an amount of oil painting artworks and summarize three critical factors to depict the figure, i.e. color, structure and texture. First, the tone of the input image is recolored based on semantic regions corresponding to the reference. Those semantic regions are segmented automatically via the color swatch, by considering the constraints of colors and positions. Then, we exploit sparse representation to reconstruct the layout by acquiring the structure from the reference. The paired training set for sparse dictionary learning is built with the guidance of edge features. Third, considering that texture is usually locally stochastic but regularly repetitive in global, a coarse-to-fine texture synthesis is used to enhance the detail pattern. Subjective results demonstrate the proposed method achieves desirable results compared with state-of-art methods while keeping consistent with artist's style. Saboya Yang, Jiaying Liu 0001, Shuai Yang 0001, Wenhan Yang, Zongming Guo |
ISCAS | 4 |
| 2017 | Real-Time Deep Video SpaTial Resolution UpConversion SysTem (STRUCT++ Demo)abstractImage and video super-resolution (SR) has been explored for several decades. However, few works are integrated into practical systems for real-time image and video SR. In this work, we present a real-time deep video SpaTial Resolution UpConversion SysTem (STRUCT++). Our demo system achieves real-time performance (50 fps on CPU for CIF sequences and 45 fps on GPU for HDTV videos) and provides several functions: 1) batch processing; 2) full resolution comparison; 3) local region zooming in. These functions are convenient for super-resolution of a batch of videos (at most 10 videos in parallel), comparisons with other approaches and observations of local details of the SR results. The system is built on a Global context aggregation and Local queue jumping Network (GLNet). It has a thinner and deeper network structure to aggregate global context with an additional local queue jumping path to better model local structures of the signal. GLNet achieves state-of-the-art performance for real-time video SR. Wenhan Yang, Shihong Deng, Yueyu Hu, Junliang Xing, Jiaying Liu 0001 |
ACM Multimedia | 1 |
| 2017 | Real-time deep image super-resolution via global context aggregation and local queue jumpingabstractDeep learning-based image super-resolution has provided very impressive reconstruction quality. However, their running time still sets barriers for real-time applications. In this paper, we propose a Global context aggregation and Local queue jumping Network (GLNet) which provides the more effective image SR given a certain number of model parameters. In our GLNet, we reconsider the model design of the real-time image SR paradigm. Then, we construct a deep network with fewer channels but a deeper structure to effectively aggregate the global context. The dilated convolutions are used as parts of basic units of our GLNet, which further enlarges the receptive field. Besides, an additional local queue jumping path is employed to connect the first-layer feature map and the last-layer feature map to better model the local signal structure. Extensive experiments demonstrate the superiority of our GLNet which offers new state-of-the-art performance considering both reconstruction quality and time consumption. Yueyu Hu, Jiaying Liu 0001, Wenhan Yang, Shihong Deng, Luyao Zhang 0007, Zongming Guo |
VCIP | 3 |
| 2017 | LG-CNN: From local parts to global discrimination for fine-grained recognition
Guosen Xie, Xu-Yao Zhang, Wenhan Yang, Mingliang Xu 0001, Shuicheng Yan, Cheng-Lin Liu 0001 |
Pattern Recognit. | 3 |
| 2017 | Deep Edge Guided Recurrent Residual Learning for Image Super-ResolutionabstractIn this paper, we consider the image super-resolution (SR) problem. The main challenge of image SR is to recover high-frequency details of a low-resolution (LR) image that are important for human perception. To address this essentially ill-posed problem, we introduce a Deep Edge Guided REcurrent rEsidual (DEGREE) network to progressively recover the high-frequency details. Different from most of the existing methods that aim at predicting high-resolution (HR) images directly, the DEGREE investigates an alternative route to recover the difference between a pair of LR and HR images by recurrent residual learning. DEGREE further augments the SR process with edge-preserving capability, namely the LR image and its edge map can jointly infer the sharp edge details of the HR image during the recurrent recovery process. To speed up its training convergence rate, by-pass connections across the multiple layers of DEGREE are constructed. In addition, we offer an understanding on DEGREE from the view-point of sub-band frequency decomposition on image signal and experimentally demonstrate how the DEGREE can recover different frequency bands separately. Extensive experiments on three benchmark data sets clearly demonstrate the superiority of DEGREE over the well-established baselines and DEGREE also provides new state-of-the-arts on these data sets. We also present addition experiments for JPEG artifacts reduction to demonstrate the good generality and flexibility of our proposed DEGREE network to handle other image processing tasks. Wenhan Yang, Jiashi Feng, Jianchao Yang, Fang Zhao 0006, Jiaying Liu 0001, Zongming Guo, Shuicheng Yan |
IEEE Trans. Image Process. | 1 |
| 2017 | Retrieval Compensated Group Structured Sparsity for Image Super-ResolutionabstractSparse representation-based image super-resolution is a well-studied topic; however, a general sparse framework that can utilize both internal and external dependencies remains unexplored. In this paper, we propose a group-structured sparse representation approach to make full use of both internal and external dependencies to facilitate image super-resolution. External compensated correlated information is introduced by a two-stage retrieval and refinement. First, in the global stage, the content-based features are exploited to select correlated external images. Then, in the local stage, the patch similarity, measured by the combination of content and high-frequency patch features, is utilized to refine the selected external data. To better learn priors from the compensated external data based on the distribution of the internal data and further complement their advantages, nonlocal redundancy is incorporated into the sparse representation model to form a group sparsity framework based on an adaptive structured dictionary. Our proposed adaptive structured dictionary consists of two parts: one trained on internal data and the other trained on compensated external data. Both are organized in a cluster-based form. To provide the desired over-completeness property, when sparsely coding a given LR patch, the proposed structured dictionary is generated dynamically by combining several of the nearest internal and external orthogonal subdictionaries to the patch instead of selecting only the nearest one as in previous methods. Extensive experiments on image super-resolution validate the effectiveness and state-of-the-art performance of the proposed method. Additional experiments on contaminated and uncorrelated external data also demonstrate its superior robustness. Jiaying Liu 0001, Wenhan Yang, Xinfeng Zhang 0001, Zongming Guo |
IEEE Trans. Multim. | 2 |
| 2016 | Robust and automatic video colorization via multiframe reordering refinementabstractIn this paper, we propose a robust video colorization method automatically through limited color references in a video sequence. The proposed method first estimates motion vectors between a monochrome frame and colored reference frames for initial matching by optical flow. Then it transfers color information to matched points in the monochrome frame and further propagates color information of matched points to other parts of the monochrome frame. Furthermore, we design a multiframe reordering refinement to colorize video sequences robustly. Experimental results demonstrate that the proposed method achieves much better performance in video colorization than state-of-the-art methods. Sifeng Xia, Jiaying Liu 0001, Yuming Fang 0001, Wenhan Yang, Zongming Guo |
ICIP | 4 |
| 2016 | Human activity recognition based on weighted limb featuresabstractHuman activity recognition plays an important role in personal assistive robot, being able to recognize human activity and perform corresponding assistive action is a great challenges for personal assistive robot. Human body is an articulated system of rigid segments that can be divided into five parts, but many existing methods always identify actions based on the motion trajectories of whole body. In this paper, taking into account the fact that most actions can be performed by a few limbs and the other limbs should not impact on the action recognition, we proposed an activity recognition method based on limb weights. The weight of each limb is composed of consistency weight and uniqueness weight, which are learned according to the similarity degree among different sequences for each specific action. The covariance descriptor, which is the concatenation of eigenvalues extracted from covariance matrices, is adopted to represent the motion trajectory of each limb. In order to distinguish action instances from each other in the feature sequences, a simple annotation method is used. Experimental results on the Cornell activity dataset and the Lab dataset show that the proposed method not only can outperform the state-of-the-art algorithms, but also is appropriate to recognize the actions whose non-core limbs' trajectories are different from each other. Liang Zhang 0010, Wenhan Yang, Guangming Zhu 0001, Peiyi Shen, Juan Song |
IROS | 2 |
| 2016 | Autoregressive image interpolation via context modeling and multiplanar constraintabstractIn this paper, we propose a novel image interpolation algorithm by context-aware autoregressive (AR) model and multiplanar constraint. Different from existing AR based methods which employ predetermined reference configuration to predict pixel values, the proposed method considers the anisotropic pixel dependencies in natural images and adaptively chooses the optimal prediction context by utilizing the nonlocal redundancy to interpolate pixels. Furthermore, the multiplanar constraint is applied to enhance the correlations within the estimation window by exploiting the self-similarity property of natural images. Similar patches are collected by the combination of patch-wise pixel values and the gradient information. And the inter-patch dependencies are adopted to improve the interpolation. The experimental results show that our method is effective in image interpolation and successfully decreases the artifacts nearby the sharp edges. The comparison experiments demonstrate that the proposed method can obtain better performance than other related ones in terms of both objective and subjective results. Shihong Deng, Jiaying Liu 0001, Mading Li, Wenhan Yang, Zongming Guo |
VCIP | 4 |
| 2015 | Neighborhood regression for edge-preserving image super-resolutionabstractThere have been many proposed works on image super-resolution via employing different priors or external databases to enhance HR results. However, most of them do not work well on the reconstruction of high-frequency details of images, which are more sensitive for human vision system. Rather than reconstructing the whole components in the image directly, we propose a novel edge-preserving super-resolution algorithm, which reconstructs low- and high-frequency components separately. In this paper, a Neighborhood Regression method is proposed to reconstruct high-frequency details on edge maps, and low-frequency part is reconstructed by the traditional bicubic method. Then, we perform an iterative combination method to obtain the estimated high resolution result, based on an energy minimization function which contains both low-frequency consistency and high-frequency adaptation. Extensive experiments evaluate the effectiveness and performance of our algorithm. It shows that our method is competitive or even better than the state-of-art methods. Yanghao Li, Jiaying Liu 0001, Wenhan Yang, Zongming Guo |
ICASSP | 3 |
| 2015 | Novel autoregressive model based on adaptive window-extension and patch-geodesic distance for image interpolationabstractIn this paper, we propose a novel autoregressive (AR) model based on the adaptive window and the patch-geodesic distance for the image interpolation. The model combines the information of inner/inter-patch correlation. To model the inner-patch correlation, we introduce a patch-geodesic distance similarity metric. The proposed metric shows the desirable capacity to depict the piecewise-stationarity of natural images. For the inter-patch correlation, we introduce the inter-patch structure variation and propose an adaptive window-extension AR model. The model extends the interpolation window according to the local structural variation, increasing the adaptation without violating the consistency. Comprehensive experiments demonstrate that the proposed method is better than or competitive with state-of-the-art interpolation methods in both objective and subjective quality evaluations. Wenhan Yang, Jiaying Liu 0001, Shuai Yang 0001, Zongming Guo |
ICASSP | 1 |
| 2015 | Multi-pose face hallucination via neighbor embedding for facial componentsabstractIn this paper, we propose a novel multi-pose face hallucination method based on Neighbor Embedding for Facial Components (NEFC) to magnify face images with various poses and expressions. To represent the structure of a face, a facial component decomposition is employed on each face image. Then, a neighbor embedding reconstruction method with locality-constraint is performed for each facial component. For the video scenario, we utilize optical flow to locate the position of each patch among the neighboring frames and make use of the Intra and Inter Nonlocal Means method to preserve consistency between neighboring frames. Experimental results evaluate the effectiveness and adaptability of our algorithm. It shows that our method achieves better performance than the state-of-the-art methods, especially on the face images with various poses and expressions. Yanghao Li, Jiaying Liu 0001, Wenhan Yang, Zongming Guo |
ICIP | 3 |
| 2015 | Adaptive autoregressive model with window extension via explicit geometry for image interpolationabstractIn this paper, we propose a novel adaptive autoregressive (AR) model constructed with an explicit geometry based extended window for image interpolation. Geometric features are chosen as criterions to include more useful pixels. These features are estimated explicitly and guide the interpolation window to extend adaptively. To characterize the piecewise stationary of images, the patch-geodesic distance based similarity is proposed and modulated into the adaptive AR model. For increasing the precision of the parameter estimation, a weighted ridge regression based estimation is employed. With the estimation, the multicollinearity between parameters, which occurs in piecewise stationarity conditions, is eliminated. Experimental results demonstrate that the proposed method is better than or competitive with state-of-the-art interpolation methods in both objective and subjective quality evaluations. Qingyun Wang 0007, Jiaying Liu 0001, Wenhan Yang, Zongming Guo |
ICIP | 3 |
| 2015 | Image super-resolution via nonlocal similarity and group structured sparse representationabstractSparse prior provides an effective tool for the image reconstruction. However, the sparse coding for independent patches leads to the unstable sparse decomposition. In this paper, we propose a group structured sparse representation model by considering the nonlocal similarity. The nonlocal similar patches are collected and classified into groups. Patches in the same group are reconstructed based the same basis of dictionaries. The dictionary is organized as the combination of many orthogonal sub-dictionaries. To provide the redundancy, the dictionary used for the sparse coding is generated online with several sub-dictionaries, thus it is over-complete. We apply the proposed model into a gradual SR framework. The framework enlarges LR to HR by a patch enhancement and an alternative sparse reconstruction on the patch and group. Objective quality evaluation shows that our proposed SR method achieves highest PSNR results comparing with the state-of-the-art methods. And subjective results demonstrate the proposed method reduces artifacts and preserves more details. Wenhan Yang, Jiaying Liu 0001, Saboya Yang, Zongming Quo |
VCIP | 1 |
| 2015 | Image Super-Resolution Based on Structure-Modulated Sparse RepresentationabstractSparse representation has recently attracted enormous interests in the field of image restoration. The conventional sparsity-based methods enforce sparse coding on small image patches with certain constraints. However, they neglected the characteristics of image structures both within the same scale and across the different scales for the image sparse representation. This drawback limits the modeling capability of sparsity-based super-resolution methods, especially for the recovery of the observed low-resolution images. In this paper, we propose a joint super-resolution framework of structure-modulated sparse representations to improve the performance of sparsity-based image super-resolution. The proposed algorithm formulates the constrained optimization problem for high-resolution image recovery. The multistep magnification scheme with the ridge regression is first used to exploit the multiscale redundancy for the initial estimation of the high-resolution image. Then, the gradient histogram preservation is incorporated as a regularization term in sparse modeling of the image super-resolution problem. Finally, the numerical solution is provided to solve the super-resolution problem of model parameter estimation and sparse representation. Extensive experiments on image super-resolution are carried out to validate the generality, effectiveness, and robustness of the proposed algorithm. Experimental results demonstrate that our proposed algorithm, which can recover more fine structures and details from an input low-resolution image, outperforms the state-of-the-art methods both subjectively and objectively in most cases. Yongqin Zhang, Jiaying Liu 0001, Wenhan Yang, Zongming Guo |
IEEE Trans. Image Process. | 3 |
| 2014 | General scale interpolation based on fine-grained isophote model with consistency constraintabstractIn this paper, we propose a fine-grained isophote model with consistency constraint to characterize the piecewise-stationarity of image signals. According to this model, we present a novel interpolation algorithm. In this model, the displacement coefficient is used to model the isophote. Then fine-grained pixel intensity information is introduced to correct the displacement calculation and make the isophote estimation more robust. In order to handle the piecewise-stationarity, we force the isophote direction consistent in the local window when an interpolated line is piecewise-stationary. The proposed algorithm can accommodate the general scale enlargement. Experimental results demonstrate that the proposed approach achieves better performances in both objective and subjective quality assessment. Wenhan Yang, Jiaying Liu 0001, Mading Li, Zongming Guo |
ICIP | 1 |