Huiyuan Fu

dblp:39/10696 · DBLP profile ↗
← Back
52ranked-venue papers
18as first author
25since 2021 · last 2026
0000-0002-5276-4366ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 43 · 15 first-author · 24 since 2021Artificial intelligence and machine learning · 17 · 6 first-author · 8 since 2021Computer networks · 2 · 2 first-author · 1 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 HDRMovieformer: A Transformer Framework and Benchmark for Cinematic SDR-to-HDR Conversion
abstract
With the growing prevalence of HDR-capable cinema venues such as Cinity LED theaters, there is an increasing demand to convert existing Standard Dynamic Range (SDR) films into High Dynamic Range (HDR) formats for theatrical presentation. However, existing SDR-to-HDR conversion methods are primarily tailored for consumer-grade content such as television and therefore fall short of the stringent requirements of professional cinematic material. To bridge this gap, we present HDRMovie7K, the first large-scale, lossless dataset of cinematic SDR-HDR frame pairs sourced from professional Digital Cinema Distribution Master (DCDM) workflows. Based on this foundation, we introduce HDRMovieformer, a transformer-based framework featuring a Luminance Estimator module for luminance guidance, a Luminance-Guided Multi-Head Self-Attention to focus on critical fine-detail recovery, and a Chroma Refiner for color accuracy, optimized with a novel Wide Color Gamut Loss. To further evaluate our model in online streaming media scenarios, we introduce HDRMovie1K, a dataset curated from publicly available HDR film clips. Extensive experiments on both HDRMovie7K and HDRMovie1K demonstrate that our method achieves state-of-the-art performance.
Huiyuan Fu, Chuanming Wang, Huadong Ma
AAAI2
2026 Selective Diffusion Distillation for Real-World High-Scale Image Super-Resolution
abstract
High-scale image super-resolution (SR) has become increasingly important with the rapid growth of mobile devices and high-resolution displays. However, current SR methods primarily focus on lower scales and generalize poorly to high-scale scenarios due to severe information loss and complex real-world degradations. In this paper, we propose a novel Selective Diffusion Distillation (SDD) framework for real-world high-scale SR, which distills reliable knowledge from a low-scale diffusion teacher to a high-scale student. Specifically, considering severe information loss in high-scale inputs, directly distilling from low-scale models may result in feature misalignment. To address this, we introduce a Degradation-aware Metric Learning (DML) approach to align feature distributions across different degradation levels. In addition, since the diffusion-based teacher may hallucinate artifacts in ambiguous regions, blindly imitating these unreliable outputs can degrade the student’s fidelity. To tackle this, we propose a Region-aware Selective Distillation (RSD) strategy to filter out uncertain predictions and adaptively supervise only on reliable areas. To evaluate the effectiveness of our method, we introduce Real-UltraSR, a new real-world benchmark that contains diverse high-scale LR-HR pairs, including x8, x10, x12, and x14. Extensive experiments demonstrate that our SDD framework achieves state-of-the-art performance across multiple benchmarks.
Wenli Zheng, Huiyuan Fu, Xin Wang 0001, Huadong Ma
AAAI2
2026 Learning Continuous Degradation for Real-World Arbitrary-Scale Video Super-Resolution
abstract
Arbitrary-scale video super-resolution (VSR) aims to enhance video resolution at continuous scales and has attracted increasing attention in recent years. However, existing methods typically rely on fixed degradation modes, such as bicubic downsampling, which often fail to handle the complex degradations of real-world videos. Current real-world datasets only cover limited scales (e.g., ×2, ×4) and are insufficient to capture the diverse degradations required for arbitrary-scale VSR. To address this, we present RealArbVSR, the first real-world VSR dataset with both integer and decimal scale factors, providing a wider range of degradation levels. Moreover, to generate continuous degradations beyond the collected scales, we propose the Continuous Degradation Generation Network (CDGN), which synthesizes realistic LR videos with arbitrary degradations. Specifically, we design a Scale-aware Degradation Module (SDM) to adaptively learn scale-specific degradations and an Implicit Filter Module (IFM) that represents spatial-temporal features as a continuous feature domain for arbitrary-scale LR frame generation. Extensive experiments demonstrate that our CDGN trained on RealArbVSR produces high-fidelity LR videos with arbitrary degradations and significantly enhances the performance of VSR models in real-world scenarios. The RealArbVSR dataset and source code will be publicly released for further research.
Wenli Zheng, Huiyuan Fu, Chuanming Wang, Enyuan Zhang, Hengming Mao, Heng Zhang 0042, Huadong Ma
IEEE Trans. Circuits Syst. Video Technol.2
2025 MUSE: Multi-Subject Unified Synthesis Via Explicit Layout Semantic Expansion
abstract
Existing text-to-image diffusion models have demonstrated remarkable capabilities in generating high-quality images guided by textual prompts. However, achieving multi-subject compositional synthesis with precise spatial control remains a significant challenge. In this work, we address the task of layout-controllable multi-subject synthesis (LMS), which requires both faithful reconstruction of reference subjects and their accurate placement in specified regions within a unified image. While recent advancements have separately improved layout control and subject synthesis, existing approaches struggle to simultaneously satisfy the dual requirements of spatial precision and identity preservation in this composite task. To bridge this gap, we propose MUSE, a unified synthesis framework that employs concatenated cross-attention (CCA) to seamlessly integrate layout specifications with textual guidance through explicit semantic space expansion. The proposed CCA mechanism enables bidirectional modality alignment between spatial constraints and textual descriptions without interference. Furthermore, we design a progressive two-stage training strategy that decomposes the LMS task into learnable sub-objectives for effective optimization. Extensive experiments demonstrate that MUSE achieves zero-shot end-to-end generation with superior spatial accuracy and identity consistency compared to existing solutions, advancing the frontier of controllable image synthesis. Our code and model are available at https://github.com/pf0607/MUSE.
Fei Peng 0003, Junqiang Wu, Yan Li 0043, Tingting Gao, Di Zhang 0026, Huiyuan Fu
ICCV6
2025 From Abyssal Darkness to Blinding Glare: a Benchmark on Extreme Exposure Correction in Real World
Bo Wang 0108, Huiyuan Fu, Zhiye Huang, Siru Zhang 0002, Xin Wang 0001, Huadong Ma
ICCV2
2025 Severe Light, Textureless Sight: A Benchmark for Extreme Exposure Correction
abstract
Exposure correction aims to restore underexposed and overexposed images to normal exposed images in a single network. However, conventional methods primarily focus on correcting non-extreme exposure cases and struggle to accurately restore lightness and structure information in extreme exposure scenarios. Through a thorough investigation, we observe that the extreme exposure correction task is limited by the lack of high-quality benchmark datasets. To address the above challenges, in this paper, we construct the first Extreme Exposure Dataset named EED by manually collecting a large number of diverse scenes. By introducing probabilistic blur kernel, EED not only ensures the rich diversity and brightness distribution of scenes but also approaches the degradation of the real world. To achieve exposure correction in extreme conditions, we propose a novel Extreme Exposure Correction Network by leveraging the mask-aware Fourier transform prior, which decouples lightness and structure components precisely. To restore severe abnormal lightness and lost structure information in extreme exposure scenes, we introduce a well-exposed referenced image to guide the coarse restoration and employ a Timestep-guided Frequency Diffusion Module for further refinement. Extensive experiments demonstrate the superiority of our dataset and method. The dataset will be available at https://github.com/juvenoia/EED.
Bo Wang 0108, Jin Liu 0024, Huiyuan Fu, Xin Wang 0001, Heng Zhang 0042, Huadong Ma
ACM Multimedia3
2025 EvRAW: Event-guided Structural and Color Modeling for RAW-to-sRGB Image Reconstruction
abstract
Event-based image reconstruction has achieved remarkable progress, benefiting from the high temporal resolution and high dynamic range of event cameras. However, most event-based methods focus on enhancing sRGB image quality, neglecting the potential of leveraging event data for RAW-to-sRGB conversion. Due to the limitations of camera sensors, images processed through standard ISP pipelines often suffer from motion blur and color distortion in dynamic scenes. In contrast, RAW images preserve uncompressed scene information, integrating event signals at this stage enables finer texture recovery and more accurate color correction. To tackle these challenges, we propose EvRAW, a novel event-assisted RAW-to-sRGB image reconstruction network that integrates event signals to promote high-fidelity sRGB image reconstruction. Specifically, we introduce a Motion-guided Structural Enhancement (MSE) module that extracts motion patterns from event streams and aggregates dynamic features to restore fine textures. Additionally, we propose an Adaptive Color Correction (ACC) module that performs region-wise gamma correction and channel-wise color decoding to enhance color fidelity under complex lighting conditions. To evaluate performance in challenging real-world scenarios, we collect a pixel-aligned RAW-Event dataset specifically for this task. Extensive experiments demonstrate that EvRAW achieves state-of-the-art performance in RAW-to-sRGB reconstruction on both synthetic and real-world datasets.
Wenli Zheng, Huiyuan Fu, Xicong Wang, Hao Kang, Chuanming Wang, Jin Liu 0024, Heng Zhang 0042, Huadong Ma
ACM Multimedia2
2025 RegPalm: Toward Large-Scale Open-Set Palmprint Recognition by Reducing Pattern Variance
abstract
Despite the recent significant progress in palmprint recognition, there are still challenges in scaling up this technology for real-world scenarios. One major challenge in developing practical, highly accurate recognition models is the shortage of comprehensive public datasets that can be used to evaluate performance at extremely low false accept rates (FAR). Furthermore, obtaining high-precision recognition models is greatly hindered by pattern variance, a notable challenge with the palmprint modality given the current technology pipeline. To address the above problems, we first collect a palmprint dataset, WebPalm, that contains the largest number of identities as well as images that have been disclosed so far. To reduce pattern variance, we propose RegPalm, a novel framework that unifies palmprint orientations (UPO) and learns pairwise spatial registration of palmprints (PPR) in an end-to-end manner. UPO harmonizes the pattern variance between left and right orientations, hence enhancing the network’s perceptual capabilities. PPR decreases both inter-class and intra-class pattern variance to improve the model’s ability to recognize hard examples. RegPalm reinforces the model by discriminating subtle palmprint features, thereby improving its performance under extremely low FAR. RegPalm not only surpasses the current state-of-the-art by 9.3 percentage points (pp) and 12.2 pp in TAR@FAR=1e-6 under the 1:1 and 1:3 open-set protocols, respectively, but also consistently achieves a 16 pp improvement in TAR@FAR=1e-9 on the WebPalm benchmark. The experimental results fully reveal the practicability and superiority of RegPalm in the real world.
Yaoyao Zhong, Weilong Chai, Huiyuan Fu, Huadong Ma
IEEE Trans. Inf. Forensics Secur.5
2025 Rethinking the Low-Light Video Enhancement: Benchmark Datasets and Methods
abstract
Low-light video enhancement is a critical task in computer vision with a wide range of applications. However, there is a lack of high-quality benchmark datasets in this field. To address this issue, we collect a high-quality low-light video dataset using a well-designed camera system. The videos in our dataset feature apparent camera motion and strict spatial alignment. In order to achieve general low-light video enhancement, we propose a Retinex-based method called Light Adjustable Network (LAN). LAN iteratively adjusts the brightness and adapts to different lighting conditions in various real-world scenarios, producing visually appealing results. We further develop a new dataset capture method and low-light video enhancement method to address the limitation of our previous dataset in capturing dynamic scenes and previous method. The new camera setup and capture method enable the recording of real continuous videos and generate the new dataset. Our new low-light video enhancement method, LAN++, leverages a new inter-frame relationship, difference images. It utilizes the texture information contained in the difference images of dynamic scenes to supplement the high-frequency details of the original features, which produce sharper and more realistic output images. The extensive experiments demonstrate the superiority of our low-light video dataset and enhancement method. Our dataset can be downloaded at https://pan.baidu.com/s/1d3EljvVduVM0wUOvzjWaqA?pwd=p45g.
Huiyuan Fu, Wenkai Zheng, Xicong Wang, Xin Wang 0001, Heng Zhang 0042, Huadong Ma
IEEE Trans. Image Process.2
2025 Part-Level Relationship Learning for Fine-Grained Few-Shot Image Classification
abstract
Recently, an increasing number of few-shot image classification methods have been proposed, and they aim at seeking a learning paradigm to train a high-performance classification model with limited labeled samples. However, the neglect of part-level relationships causes few-shot methods to struggle to distinguish between closely similar subcategories, which makes it difficult for them to solve the fine-grained image classification problem. To tackle this challenging task, this paper proposes a fine-grained few-shot image classification method that exploits both intra-part and inter-part relationships among different samples. To establish comprehensive relationships, we first extract multiple discriminative descriptors from the input image, representing its different parts. Then, we propose to define the metric spaces by interpolating intra-part relationships, which can help the model adaptively find clear boundaries for these confusing classes. Finally, since the unlabeled image has high similarities to all classes, we project these similarities into a high-dimension space according to the inter-part relationship and interpolate a parameterized classifier to discover the subtle differences among these similar classes. To evaluate our proposed method, we conduct extensive experiments on various fine-grained datasets. Without any pre-train/fine-tuning process, our approach clearly outperforms previous few-shot learning methods, which demonstrates the effectiveness of our approach.
Chuanming Wang, Huiyuan Fu, Peiye Liu, Huadong Ma
IEEE Trans. Multim.2
2024 Region-Aware Exposure Consistency Network for Mixed Exposure Correction
abstract
Exposure correction aims to enhance images suffering from improper exposure to achieve satisfactory visual effects. Despite recent progress, existing methods generally mitigate either overexposure or underexposure in input images, and they still struggle to handle images with mixed exposure, i.e., one image incorporates both overexposed and underexposed regions. The mixed exposure distribution is non-uniform and leads to varying representation, which makes it challenging to address in a unified process. In this paper, we introduce an effective Region-aware Exposure Correction Network (RECNet) that can handle mixed exposure by adaptively learning and bridging different regional exposure representations. Specifically, to address the challenge posed by mixed exposure disparities, we develop a region-aware de-exposure module that effectively translates regional features of mixed exposure scenarios into an exposure-invariant feature space. Simultaneously, as de-exposure operation inevitably reduces discriminative information, we introduce a mixed-scale restoration unit that integrates exposure-invariant features and unprocessed features to recover local information. To further achieve a uniform exposure distribution in the global image, we propose an exposure contrastive regularization strategy under the constraints of intra-regional exposure consistency and inter-regional exposure continuity. Extensive experiments are conducted on various datasets, and the experimental results demonstrate the superiority and generalization of our proposed method. The code is released at: https://github.com/kravrolens/RECNet.
Jin Liu 0024, Huiyuan Fu, Chuanming Wang, Huadong Ma
AAAI2
2024 Continuous Optical Zooming: A Benchmark for Arbitrary-Scale Image Super-Resolution in Real World
abstract
Most current arbitrary-scale image super-resolution (SR) methods has commonly relied on simulated data generated by simple synthetic degradation models (e.g., bicubic down-sampling) at continuous various scales, thereby falling short in capturing the complex degradation of real-world images. This limitation hinders the visual quality of these methods when applied to real-world images. To address this issue, we propose the Continuous Optical Zooming dataset (COZ), by constructing an automatic imaging system to collect images at fine-grained various focal lengths within a specific range and providing strict image pair alignment. The COZ dataset serves as a benchmark to provide real-world data for training and testing arbitrary-scale SR models. To enhance the model's robustness against real-world image degradation, we propose a Local Mix Implicit network (LMI) based on the MLP-mixer architecture and meta-learning, which directly learns the local texture information by simultaneously mixing features and coordinates of multiple independent points. The extensive experiments demonstrate the superior performance of the arbitrary-scale SR models trained on the COZ dataset compared to models trained on simulated data. Our LMI model exhibits the superior effectiveness compared to other models. This study is of great significance in developing more efficient algorithms and improving the performance of arbitrary-scale image SR methods in practical applications. Our dataset and codes are available at https://github.com/pf0607/COZ.
Huiyuan Fu, Fei Peng 0003, Yejun Li, Xin Wang 0001, Huadong Ma
CVPR1
2024 Learning Exposure Correction in Dynamic Scenes
abstract
Exposure correction aims to enhance visual data suffering from improper exposures, which can greatly improve satisfactory visual effects. However, previous methods mainly focus on the image modality, and the video counterpart is less explored in the literature. Directly applying prior image-based methods to videos results in temporal incoherence with low visual quality. Through thorough investigation, we find that the development of relevant communities is limited by the absence of a benchmark dataset. Therefore, in this paper, we construct the first real-world paired video dataset, including both underexposure and overexposure dynamic scenes. To achieve spatial alignment, we utilize two DSLR cameras and a beam splitter to simultaneously capture improper and normal exposure videos. Additionally, we propose an end-to-end video exposure correction network, in which a dual-stream module is designed to deal with both underexposure and overexposure factors, enhancing the illumination based on Retinex theory. The extensive experiments based on various metrics and user studies demonstrate the significance of our dataset and the effectiveness of our method. The code and dataset are available at https://github.com/kravrolens/VECNet.
Jin Liu 0024, Bo Wang 0108, Chuanming Wang, Huiyuan Fu, Huadong Ma
ACM Multimedia4
2024 Exploring in Extremely Dark: Low-Light Video Enhancement with Real Events
abstract
Due to the limitations of sensor, traditional cameras struggle to capture details within extremely dark areas of videos. The absence of such details can significantly impact the effectiveness of low-light video enhancement. In contrast, event cameras offer a visual representation with higher dynamic range, facilitating the capture of motion information even in exceptionally dark conditions. Motivated by this advantage, we propose the Real-Event Embedded Network for low-light video enhancement. To better utilize events for enhancing extremely dark regions, we propose an Event-Image Fusion module, which can identify these dark regions and enhance them significantly. To ensure temporal stability of the video and restore details within extremely dark areas, we design unsupervised temporal consistency loss and detail contrast loss. Alongside the supervised loss, these loss functions collectively contribute to the semi-supervised training of the network on unpaired real data. Experimental results on synthetic and real data demonstrate the superiority of the proposed method compared to the state-of-the-art methods.
Xicong Wang, Huiyuan Fu, Xin Wang 0001, Heng Zhang 0042, Huadong Ma
ACM Multimedia2
2024 SwinIT: Hierarchical Image-to-Image Translation Framework Without Cycle Consistency
abstract
Image-to-image (I2I) translation often requires establishing cycle consistency between the source and the translated images across different domains. However, cycle consistency requires redundant reconstruction, and is too restrictive to satisfy the bijection assumption between the two domains. In this paper, we propose SwinIT, a hierarchical Swin-transformer I2I Translation framework without using cycle consistency. Specifically, we carefully design symmetrical encoders for content and style flows, then explore newly proposed adaptive denormalization and normalization strategies. This framework can effectively capture and fuse content and style representations in a coarse-to-fine manner, ensuring our method achieves high performance without cycle consistency. Guided by element-wise feature adaptive denormalization, our model focuses on preserving semantic structure information. Due to the semantic mismatch between unpaired source and exemplar images, we introduce cross-attention adaptive instance normalization to help achieve better alignment. However, because the original optimization objective lacks direct supervision to preserve high-frequency information, rich edge details are lost during the translation. We propose a wavelet transformation matching loss to recover the details by converting the image into multi-frequency parts. We validate our proposed method in various I2I translation tasks, including arbitrary style transfer, multi-modal image synthesis, and semantic image synthesis, demonstrating its effectiveness in both qualitative and quantitative evaluations.
Jin Liu 0024, Huiyuan Fu, Xin Wang 0001, Huadong Ma
IEEE Trans. Circuits Syst. Video Technol.2
2024 Mutual Distillation Learning for Person Re-Identification
abstract
With the rapid advancements in deep learning technologies, person re-identification (ReID) has witnessed remarkable performance improvements. However, the majority of prior works have traditionally focused on solving the problem via extracting features solely from a single perspective, such as uniform partitioning, attention mechanisms, or semantic masks. While these approaches have demonstrated efficacy within specific contexts, they fall short in diverse situations. In this paper, we propose a novel approach, Mutual Distillation Learning For Person Re-identification (termed as MDPR), which addresses the challenging problem from multiple perspectives within a single unified model, leveraging the power of mutual distillation to enhance the feature representations collectively. Specifically, our approach encompasses two branches: a hard content branch to extract local features via a uniform horizontal partitioning strategy and a soft content branch to dynamically distinguish between foreground and background and facilitate the extraction of multi-granularity features via a carefully designed attention mechanism. To facilitate knowledge exchange between these two branches, a mutual distillation and fusion process is employed, promoting the capability of the outputs of each branch. Extensive experiments are conducted on widely used person ReID datasets to validate the effectiveness and superiority of our approach. Notably, our method achieves an impressive 88.7%/94.4% in mAP/Rank-1 on the DukeMTMC-reID dataset, surpassing the current state-of-the-art results. Our source code is available athttps://github.com/KuilongCui/MDPR.
Huiyuan Fu, Kuilong Cui, Chuanming Wang, Mengshi Qi, Huadong Ma
IEEE Trans. Multim.1
2024 Learning Mutually Exclusive Part Representations for Fine-Grained Image Classification
abstract
Fine-grained image classification (FGIC) aims to separate different subcategories from one general superclass, which requires the classification model to extract distinctive representations from subtle yet discriminative regions of the objects. Learning multiple part representations can give a detailed description of the object from different perspectives, boosting the classification performance. However, it still remains a challenging problem to effectively locate diverse parts and extract their features without the assistance of part annotations. In this article, we present a novel method to achieve accurate fine-grained image classification by learning a set of diverse and discriminative part representations without requiring additional supervision. Firstly, our method utilizes a simple attention interaction module to lead learned spatial attentions to focus on different parts, resulting in mutually exclusive part representations. Then, to reduce the impairment of channel coupling among part representations, a part-wise channel weighting module is designed to adjust the amplitudes of different representations adaptively, making them to be diverse along the channel dimension. Moreover, to ensure comprehensive and sufficient part representations, our method introduces multi-granularity feature learning. It enables the extraction of part representations from different semantic and content levels, capturing fine-grained details effectively. To evaluate our method, extensive experiments are conducted on various benchmark fine-grained image datasets, and the results show that our method can achieve outstanding performance for FGIC, demonstrating its effectiveness.
Chuanming Wang, Huiyuan Fu, Huadong Ma
IEEE Trans. Multim.2
2024 Multi-Domain Image-to-Image Translation with Cross-Granularity Contrastive Learning
abstract
The objective of multi-domain image-to-image translation is to learn the mapping from a source domain to a target domain in multiple image domains while preserving the content representation of the source domain. Despite the importance and recent efforts, most previous studies disregard the large style discrepancy between images and instances in various domains, or fail to capture instance details and boundaries properly, resulting in poor translation results for rich scenes. To address these problems, we present an effective architecture for multi-domain image-to-image translation that only requires one generator. Specifically, we provide detailed procedures for capturing the features of instances throughout the learning process, as well as learning the relationship between the style of the global image and that of a local instance in the image by enforcing the cross-granularity consistency. In order to capture local details within the content space, we employ a dual contrastive learning strategy that operates at both the instance and patch levels. Extensive studies on different multi-domain image-to-image translation datasets reveal that our proposed method outperforms state-of-the-art approaches.
Huiyuan Fu, Jin Liu 0024, Xin Wang 0001, Huadong Ma
ACM Trans. Multim. Comput. Commun. Appl.1
2023 You Do Not Need Additional Priors or Regularizers in Retinex-Based Low-Light Image Enhancement
abstract
Images captured in low-light conditions often suffer from significant quality degradation. Recent works have built a large variety of deep Retinex-based networks to enhance low-light images. The Retinex-based methods require decomposing the image into reflectance and illumination components, which is a highly ill-posed problem and there is no available ground truth. Previous works addressed this problem by imposing some additional priors or regularizers. However, finding an effective prior or regularizer that can be applied in various scenes is challenging, and the performance of the model suffers from too many additional constraints. We propose a contrastive learning method and a self-knowledge distillation method for Retinex decomposition that allow training our Retinex-based model without elaborate hand-crafted regularization functions. Rather than estimating reflectance and illuminance images and representing the final images as their element-wise products as in previous works, our regularizer-free Retinex decomposition and synthesis network (RFR) extracts reflectance and illuminance features and synthesizes them end-to-end. In addition, we propose a loss function for contrastive learning and a progressive learning strategy for self-knowledge distillation. Extensive experimental results demonstrate that our proposed methods can achieve superior performance compared with state-of-the-art approaches.
Huiyuan Fu, Wenkai Zheng, Xin Wang 0001, Chuanming Wang, Huadong Ma
CVPR1
2023 Dancing in the Dark: A Benchmark towards General Low-light Video Enhancement
abstract
Low-light video enhancement is a challenging task with broad applications. However, current research in this area is limited by the lack of high-quality benchmark datasets. To address this issue, we design a camera system and collect a high-quality low-light video dataset with multiple exposures and cameras. Our dataset provides dynamic video pairs with pronounced camera motion and strict spatial alignment. To achieve general low-light video enhancement, we also propose a novel Retinex-based method named Light Adjustable Network (LAN). LAN iteratively refines the illumination and adaptively adjusts it under varying lighting conditions, leading to visually appealing results even in diverse real-world scenarios. The extensive experiments demonstrate the superiority of our low-light video dataset and enhancement method. Our dataset is available at https://github.com/ciki000/DID.
Huiyuan Fu, Wenkai Zheng, Xicong Wang, Heng Zhang 0042, Huadong Ma
ICCV1
2023 Multi-Part Token Transformer with Dual Contrastive Learning for Fine-grained Image Classification
abstract
Fine-grained image classification focuses on distinguishing objects from different similar subcategories, which requires the classification model to extract subtle yet discriminative descriptors. Recent Vision Transformer (ViT) has shown an enormous potential for this challenging task, but previous ViT-based methods have primarily focused on improving the relationship between image patches, neglecting the limited expressive capability caused by the single class token.To address this limitation, we propose to learn a Multi-part Token Transformer (MpT-Trans), which extends the class token to multiple tokens presenting various parts, enhancing the model's capability of extracting discriminative information. Specifically, our MpT-Trans model interpolates the vision transformer framework with two modules: (i) the Part-wise Shift Learning (PwSL) module is proposed to extend the single class token to a set of part tokens with differentiable shifts, enabling the model to extract informative representations from different perspectives; (ii) the Dual Contrastive Learning (DuCL) module is introduced to exploit the inter-class and inter-part relationships to regularize the learning of part tokens, enhancing their diversity and discrimination for accurate classification. Extensive experiments and ablation study demonstrate that the proposed MpT-Trans achieves state-of-the-art performance on various fine-grained image benchmark datasets, demonstrating the effectiveness of our proposed method.
Chuanming Wang, Huiyuan Fu, Huadong Ma
ACM Multimedia2
2022 PaCL: Part-level Contrastive Learning for Fine-grained Few-shot Image Classification
abstract
Recently, it is gaining increasingly attention to incorporate self-supervised technologies into few-shot learning. Previous methods have exclusively focused on image-level self-supervision, but they ignore that capturing subtle part features plays an important role in distinguishing fine-grained images. In this paper, we propose an approach named PaCL that embeds part-level contrastive learning into fine-grained few-shot image classification, strengthening the models' capability to extract discriminative features from indistinguishable images. PaCL treats parts as the inputs of contrastive learning, and it uses a transformation module to involve image-specific information into pre-defined meta parts, generating multiple features from each meta part depending on different images. To alleviate the impact of changes in views or occlusions, we propose to adopt part prototypes in contrastive learning. Part prototypes are generated by aggregating the features of each certain type of part, which are more reliable than directly using part features. A few-shot classifier is adopted to predict query images, which calculates the classification loss to optimize the transformation module and meta parts in conjunction with the loss calculated in contrastive learning. The optimization process will enforce the model to learn to extract discriminative and diverse features from different parts of the objects, even for the samples of unseen classes. Extensive studies show that our proposed method improves the performance of fine-grained few-shot image classification across several backbones, datasets, and tasks, achieving superior results compared with state-of-the-art methods.
Chuanming Wang, Huiyuan Fu, Huadong Ma
ACM Multimedia2
2021 Stacked Semantically-Guided Learning for Image De-distortion
abstract
Image de-distortion is very important because distortions will degrade the image quality significantly. It can benefit many computational visual media applications that are primarily designed for high-quality images. In order to address this challenging issue, we propose a stacked semantically-guided network, which is the first try on this task. It can capture and restore the distortions around the humans and the adjacent background effectively with the stacked network architecture and the semantically-guided scheme. In addition, a discriminative restoration loss function is proposed to recover different distorted regions in the images discriminatively. As another important effort, we construct a large-scale dataset for image de-distortion. Extensive qualitative and quantitative experiments show that our proposed method achieves a superior performance compared with the state-of-the-art approaches.
Huiyuan Fu, Changhao Tian, Xin Wang 0001, Huadong Ma
ACM Multimedia1
2021 See clearly on rainy days: Hybrid multiscale loss guided multi-feature fusion network for single image rain removal
abstract
The quality of photos is highly susceptible to severe weather such as heavy rain; it can also degrade the performance of various visual tasks like object detection. Rain removal is a challenging problem because rain streaks have different appearances even in one image. Regions where rain accumulates appear foggy or misty, while rain streaks can be clearly seen in areas where rain is less heavy. We propose removing various rain effects in pictures using a hybrid multiscale loss guided multiple feature fusion de-raining network (MSGMFFNet). Specially, to deal with rain streaks, our method generates a rain streak attention map, while preprocessing uses gamma correction and contrast enhancement to enhanced images to address the problem of rain accumulation. Using these tools, the model can restore a result with abundant details. Furthermore, a hybrid multiscale loss combining L 1 loss and edge loss is used to guide the training process to pay attention to edge and content information. Comprehensive experiments conducted on both synthetic and real-world datasets demonstrate the effectiveness of our method.
Huiyuan Fu, Yu Zhang 0133, Huadong Ma
Comput. Vis. Media1
2021 An end-to-end convolutional network for joint detecting and denoising adversarial perturbations in vehicle classification
abstract
Deep convolutional neural networks (DCNNs) have been widely deployed in real-world scenarios. However, DCNNs are easily tricked by adversarial examples, which present challenges for critical applications, such as vehicle classification. To address this problem, we propose a novel end-to-end convolutional network for joint detection and removal of adversarial perturbations by denoising (DDAP). It gets rid of adversarial perturbations using the DDAP denoiser based on adversarial examples discovered by the DDAP detector. The proposed method can be regarded as a pre-processing step—it does not require modifying the structure of the vehicle classification model and hardly affects the classification results on clean images. We consider four kinds of adversarial attack (FGSM, BIM, DeepFool, PGD) to verify DDAP's capabilities when trained on BIT-Vehicle and other public datasets. It provides better defense than other state-of-the-art defensive methods.
Huiyuan Fu, Huadong Ma
Comput. Vis. Media2
2020 Region-Based Global Reasoning Networks
Chuanming Wang, Huiyuan Fu, Charles Ling 0001, Peilun Du, Huadong Ma
AAAI2
2020 Global Structure Graph Guided Fine-Grained Vehicle Recognition
abstract
Fine-grained vehicle recognition is a challenging problem due to the subtle intra-category appearance variation, which requires the recognition model can capture discriminative features from distinguishing regions. The structure is an important characteristic of vehicles which can help to find substantial parts and learn distinguishing representations. In this paper, we propose an approach that introduces the structure graph into consideration to learn distinguishing representations for vehicle recognition. Our proposed method first constructs a global structure graph from the features generated by the convolutional network and then it applies the graph as the guidance to produce effective representations of vehicles. The results of extensive experiments demonstrate that our proposed method can produce more promising results than other state-of-the-art methods. The results of the visualization illustrate that our approach can construct a suitable structure graph and the global structure information facilitates learning discriminative representations at crucial parts of vehicles.
Chuanming Wang, Huiyuan Fu, Huadong Ma
ICASSP2
2020 Cross-Granularity Learning for Multi-Domain Image-to-Image Translation
abstract
Image translation across diverse domains has attracted more and more attention. Existing multi-domain image-to-image translation algorithms only learn the features of the complete image without considering specific features of local instances. To ensure the important instance to be more realistically translated, we propose a cross-granularity learning model for multi-domain image-to-image translation. We provide detailed procedures to capture the features of instances during the learning process, and specifically learn the relationship between style of the global image and the style of an instance on the image through the enforcing of the cross-granularity consistency. In our design, we only need one generator to perform the instance-aware multi-domain image translation. Our extensive experiments on several multi-domain image-to-image translation datasets show that our proposed method can achieve superior performance compared with the state-of-the-art approaches.
Huiyuan Fu, Xin Wang 0001, Huadong Ma
ACM Multimedia1
2020 Enhancing Anomaly Detection in Surveillance Videos with Transfer Learning from Action Recognition
abstract
Anomaly detection in surveillance videos, as a special case of video-based action recognition, has been of increasing interest in multimedia community and public security. Action recognition in videos faces some challenges, such as cluttered background, illumination conditions. Besides these above difficulties, detecting anomaly in surveillance videos has several unique problems to be solved. For example, the lack of sufficient training samples is one of the main challenges for detecting anomalies in surveillance videos. In this paper, we propose to utilize transfer learning to leverage the good results from action recognition for anomaly detection in surveillance videos. More specially, we explore some techniques based on action recognition models from the following aspects: training samples, temporal modules for action recognition, network backbones. We draw some conclusions. First, more training samples from surveillance videos lead to higher classification accuracy. Second, stronger temporal modules designed for recognizing action and deeper networks do not achieve better results. This conclusion is reasonable since deeper networks tend to over-fitting, especially for the small-scale training set. Besides, to distinguish the hard examples from normal activities, we separately train a neural network to classify the hard category and normal events. Then we fuse the binary network and previous network to generate the final prediction for general anomaly detection. On the benchmarks of CitySCENE, our framework achieves promising performance and obtains the first prize for general anomaly detection and the second prize for specific anomaly detection.
Kun Liu 0016, Minzhi Zhu, Huiyuan Fu, Huadong Ma, Tat-Seng Chua
ACM Multimedia3
2020 MCFF-CNN: Multiscale comprehensive feature fusion convolutional neural network for vehicle color recognition based on residual learning
Huiyuan Fu, Huadong Ma, Gaoya Wang, Xiaomou Zhang
Neurocomputing1
2020 FOM: Fourth-order moment based causal direction identification on the heteroscedastic data
Ruichu Cai, Jincheng Ye, Jie Qiao, Huiyuan Fu
Neural Networks4
2019 Esnet: Edge-Based Segmentation Network for Real-Time Semantic Segmentation in Traffic Scenes
abstract
Semantic segmentation is widely used in the industry recently, especially in the field of scene understanding, surveillance and autonomous driving. However, majority of current state-of-the-art algorithms run accompany with high consumption of computation resources. Thus, our work focuses on real-time semantic segmentation which could reduce a large proportion of computation. Traditional methods to speed up segmentation process tend to down sample image. However, down sampling would cause the loss of information. Hence, we propose a real-time edge-based segmentation network (ESNet) that incorporate high-resolution global edge information with low-resolution classification-level semantic information. Our network performs real-time inference on single GPU card on high-resolution Cityscapes dataset.
Haoran Lyu, Huiyuan Fu, Xiaojun Hu, Liang Liu 0001
ICIP2
2019 Hippocampus Segmentation in MRI Using Side U-Net Model
Wenbin Yao, Huiyuan Fu
ICONIP (3)3
2019 A Convolutional Neural Network Pruning Method Based On Attention Mechanism
abstract
Pruning effectively reduces the size of neural networks, which facilitates deployment of neural networks in production environment, especially in embedded systems with limited computing resources.In this paper, we propose a convolutional neural network pruning method based on attention mechanism.We add a attention module to model to generate scaling factors for channels.The scaling factors are considered as channels' importance score, thus filters and convolution kernels corresponding to channels with lower importance score are removed.Our method has the ability to learn importance of channels during training, instead of considering only the direct impact of parameters like existing methods.Moreover, it does not depend on any dedicated libraries, so could be combined with other compression methods for better performance.In experiments, we prune about 90% parameters in VGGNet with 0.67% accuracy drop and prune about 50% parameters in ResNet-56 with 1.02% accuracy drop.
Wenbin Yao, Huiyuan Fu
SEKE3
2018 An End-to-End Neural Network for Multi-line License Plate Recognition
abstract
Currently, license plate recognition plays an important role in numerous applications and a number of technologies have been proposed. However, most of them can only work with single-line license plates. In the practical application scenarios, there are also existing many multi-line license plates. The traditional approaches need to segment the original input images for double-line license plates. This is a very difficult problem in the complex scenes. In order to solve this problem, we propose an end-to-end neural network for both single-line and double-line license plate recognition. It is segmentation-free for the original input license plate images. We view each of these whole images as a unit on feature maps after deep convolution neural network directly. A large number of experiments show that our method is effective. It is better than the state-of-the-art algorithms in SYSU-ITS license plate library data.
Huiyuan Fu, Huadong Ma
ICPR2
2017 Multi-attribute Based Fire Detection in Diverse Surveillance Videos
Shuangqun Li, Wu Liu 0005, Huadong Ma, Huiyuan Fu
MMM (1)4
2016 Siamese neural network based gait recognition for human identification
abstract
As the remarkable characteristics of remote accessed, robust and security, gait recognition has gained significant attention in the biometrics based human identification task. However, the existed methods mainly employ the handcrafted gait features, which cannot well handle the indistinctive inter-class differences and large intra-class variations of human gait in real-world situation. In this paper, we have developed a Siamese neural network based gait recognition framework to automatically extract robust and discriminative gait features for human identification. Different from conventional deep neural network, the Siamese network can employ distance metric learning to drive the similarity metric to be small for pairs of gait from the same person, and large for pairs from different persons. In particular, to further learn effective model with limited training data, we composite the gait energy images instead of raw sequence of gaits. Consequently, the experiments on the world's largest gait database show our framework impressively outperforms state-of-the-arts.
Cheng Zhang 0014, Wu Liu 0005, Huadong Ma, Huiyuan Fu
ICASSP4
2016 Large-scale vehicle re-identification in urban surveillance videos
abstract
Vehicle, as a significant object class in urban surveillance, attracts massive focuses in computer vision field, such as detection, tracking, and classification. Among them, vehicle re-identification (Re-Id) is an important yet frontier topic, which not only faces the challenges of enormous intra-class and subtle inter-class differences of vehicles in multicameras, but also suffers from the complicated environments in urban surveillance scenarios. Besides, the existing vehicle related datasets all neglect the requirements of vehicle Re-Id: 1) massive vehicles captured in real-world traffic environment; and 2) applicable recurrence rate to give cross-camera vehicle search for vehicle Re-Id. To facilitate vehicle Re-Id research, we propose a large-scale benchmark dataset for vehicle Re-Id in the real-world urban surveillance scenario, named “VeRi”. It contains over 40,000 bounding boxes of 619 vehicles captured by 20 cameras in unconstrained traffic scene. Moreover, each vehicle is captured by 2~18 cameras in different viewpoints, illuminations, and resolutions to provide high recurrence rate for vehicle Re-Id. Finally, we evaluate six competitive vehicle Re-Id methods on VeRi and propose a baseline which combines the color, texture, and highlevel semantic information extracted by deep neural network.
Xinchen Liu, Wu Liu 0005, Huadong Ma, Huiyuan Fu
ICME4
2016 A vehicle classification system based on hierarchical multi-SVMs in crowded traffic scenes
Huiyuan Fu, Huadong Ma, Yinxin Liu
Neurocomputing1
2016 Scene-free multi-class weather classification on single images
Huadong Ma, Huiyuan Fu, Cheng Zhang 0014
Neurocomputing3
2015 Automatically Stereoscopic Camera Control for 3D Animation Production
abstract
This paper proposes a novel approach for automatically controlling stereoscopic camera parameters that specifically addresses challenges in stereo 3D animation production process.Our proposed camera control method produces stereo contents with preferable depth perception and guarantees visual comfort by optimization of camera parameters. We introduce an attention tracking method to calculate convergence plane, avoiding window violation and minimizing visual conflict. Moreover, we derive an smoothing function on convergence plane that reduces depth jump over time. Then, we calculate the inter-axial separation using a perceived depth mapping. We describe how to implement our method on the Maya plug-in and test the stereo effect using professional stereo 3D animation scenes. The experimental results, including a user study, show that our method enhances the stereo effect. Our controller provides automatic camera control that can be helpful in creating comfortable and faster stereo 3D animations.
Huadong Ma, Liang Liu 0001, Huiyuan Fu
ACM Multimedia5
2015 Patch-Based Disparity Remapping for Stereoscopic Images
Huadong Ma, Liang Liu 0001, Huiyuan Fu
MMM (1)4
2015 Outdoor Air Quality Inference from Single Image
Huadong Ma, Huiyuan Fu, Xinpeng Wang 0007
MMM (2)3
2015 Hotspot-entropy based data forwarding in opportunistic social networks
Peiyan Yuan, Huadong Ma, Huiyuan Fu
Pervasive Mob. Comput.3
2014 Real-time crowd detection based on gradient magnitude entropy model
abstract
Reliable and real-time crowd detection is one of the most important tasks in intelligent video surveillance system. Previous works focus on counting the number of pedestrians in the crowd directly or use holistic features of crowd scenes for crowd detection. However, the former methods will be invalid in complex crowded scenes, and the latter methods will be confused for feature selection. In this paper, we propose a simple but effective model - Gradient Magnitude Entropy (GME) model for crowd detection. Our model is based on a key observation - the value of GME in a region which will increase as the number of pedestrians grows. Thus, we can estimate the degree of crowd when the value of GME is larger than some threshold, without counting the number of pedestrians. Extensive experiments show that our GME model outperforms state-of-the-art techniques on several challenging datasets. Furthermore, our method can process in real time for practical surveillance applications.
Huiyuan Fu, Huadong Ma
ACM Multimedia1
2014 Crowd Counting via Head Detection and Motion Flow Estimation
abstract
Crowd counting with heavy occlusions is highly desired for public security. However, few works have been studied towards this goal. Most previous systems only count passing people robustly without heavy occlusions, else they have to estimate the crowds to a certain extent. To solve this difficult problem, this paper proposes an effective algorithm by combining head detection and motion flow estimation together. First, we detect each head in the crowd by using our proposed scene-adaptive scheme on depth data. Then, we estimate the motion flow based on the interest points in each head region on color data. We can ultimately achieve multi-direction crowd counting results with the loop of above steps. Based on this approach, we have built a practical system for reliable crowd counting. Extensive experimental results show that our system is effective.
Huiyuan Fu, Huadong Ma, Hongtian Xiao
ACM Multimedia1
2014 Scene-adaptive accurate and fast vertical crowd counting via joint using depth and color information
Huiyuan Fu, Huadong Ma, Hongtian Xiao
Multim. Tools Appl.1
2012 From video to text: Semantic driving scene understanding using a coarse-to-fine method
abstract
Semantic understanding from video is one of the most challenging tasks in video analysis. However, it has not been taken enough attention. In this paper, we focus on understanding the semantics of video in the driving scene. We present a coarse-to-fine method to parse the driving scene, and obtain the high-level semantic information of the scene. In the coarse phase, we divide the captured frame into four separate parts based on edge density entropy and scene context. In the fine phase, we join multi-class object segmentation and detection algorithms together in a unified Conditional Random Filed (CRF) model for each part understanding. Moreover, the object probabilistic location prior knowledge based on training and previous edge density entropy result is also integrated into our approach for better object localization. Experimental results show that our proposed method is effective comparing to current state-of-the-art approaches.
Huiyuan Fu, Huadong Ma
ICASSP1
2012 Real-time accurate crowd counting based on RGB-D information
abstract
Real-time accurate crowd counting is one of important tasks in intelligent visual surveillance systems. Most previous works can only count passing people robustly without heavy occlusions which are very common in the practical surveillance scenes. To solve this difficult problem, we propose a new method for crowd counting for RGB-D (RGB plus depth) data using a commodity depth camera. In our method, we first detect each head-shoulder of the passing or still person in the surveillance region with fast template matching based on depth information including pedestrian filling with convex hull segmentation. Then, we track and count each detected head-shoulder based on RGB information bidirectionally. By using this approach, we have built a practical system for robust and fast crowd counting. Extensive experimental results show that our method achieves significant improvement comparing to states-of-the-art approach, and the built system is not only robust to heavy occlusions, but also can be deployed in the real time crowd counting application scenes.
Huiyuan Fu, Huadong Ma, Hongtian Xiao
ICIP1
2012 Night Removal by Color Estimation and sparse representation
Huiyuan Fu, Huadong Ma, Shixin Wu
ICPR1
2011 EGMM: An enhanced Gaussian mixture model for detecting moving objects with intermittent stops
abstract
Moving object detection is one of the most important tasks in intelligent visual surveillance systems. Gaussian Mixture Model (GMM) has been most widely used for moving object detection, because of its robustness to variable scenes. However, to the best of our knowledge, existing GMM based methods can not detect moving objects which gradually stop and keep still state for a while. In this paper, we present an Enhanced Gaussian MixtureModel, called EGMM, to handle this problem. We integrate an Initial Gaussian Background Model (IGBM) and an extended Kalman filter based tracker with GMM, to enhance its performance. Experimental results show that our EGMM based method has a lower miss rate at the same false positives per image comparing to GMM based method for moving pedestrian detection, and it also has a higher detection rate for abandoned object detection comparing to GMM based method.
Huiyuan Fu, Huadong Ma, Anlong Ming
ICME1
2011 Robust Human Detection with Low Energy Consumption in Visual Sensor Network
abstract
In this paper, we try to address the difficult problem of detecting humans robustly with low energy consumption in the visual sensor network. The proposed method contains two parts: one is an ESOBS (Enhanced Self-Organizing Background Subtraction) based foreground segmentation module to obtain active areas in the observed area from the visual sensor; the other is a HOG (Histograms of Oriented Gradients) based detection module to detect the appearance shape from the foreground areas. Moreover, we create a large pedestrian dataset according to the specific scene in visual sensor networks. Numerous experiments are conducted. The experimental results show the effectiveness of our method.
Huiyuan Fu, Huadong Ma, Liang Liu 0001
MSN1