EDBT 2026 Demo / reviewers in the wild / expert
Ying Fu 0001
dblp:89/1229-1
· DBLP profile ↗
124ranked-venue papers
21as first author
91since 2021 · last 2026
0000-0002-6677-694XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 95 · 15 first-author · 70 since 2021Graphics, computer vision, multimedia, augmented reality and games · 75 · 10 first-author · 52 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Real Noise Decoupling for Hyperspectral Image DenoisingabstractHyperspectral image (HSI) denoising is a crucial step in enhancing the quality of HSIs. Noise modeling methods can fit noise distributions to generate synthetic HSIs to train denoising networks. However, the noise in captured HSIs is usually complex and difficult to model accurately, which significantly limits the effectiveness of these approaches. In this paper, we propose a multi-stage noise-decoupling framework that decomposes complex noise into explicitly modeled and implicitly modeled components. This decoupling reduces the complexity of noise and enhances the learnability of HSI denoising methods when applied to real paired data. Specifically, for explicitly modeled noise, we utilize an existing noise model to generate paired data for pre-training a denoising network, equipping it with prior knowledge to handle the explicitly modeled noise effectively. For implicitly modeled noise, we introduce a high-frequency wavelet guided network. Leveraging the prior knowledge from the pre-trained module, this network adaptively extracts high-frequency features to target and remove the implicitly modeled noise from real paired HSIs. Furthermore, to effectively eliminate all noise components and mitigate error accumulation across stages, a multi-stage learning strategy, comprising separate pre-training and joint fine-tuning, is employed to optimize the entire framework. Extensive experiments on public and our captured datasets demonstrate that our proposed framework outperforms state-of-the-art methods, effectively handling complex real-world noise and significantly enhancing HSI quality. Yingkai Zhang, Tao Zhang 0042, Ying Fu 0001 |
AAAI | 4 |
| 2026 | Atlantis++: Enabling Underwater Depth Estimation with Stable Diffusion and Beyond
Fan Zhang 0123, Shaodi You, Yu Li 0003, Ying Fu 0001 |
Int. J. Comput. Vis. | 4 |
| 2026 | MSFA Image Denoising Using Physics-Based Noise Model and Noise-Decoupled NetworkabstractMultispectral filter array (MSFA) camera is increasingly used due to its compact size and fast capturing speed. However, because of its narrow-band property, it often suffers from the light-deficient problem, and images captured are easily overwhelmed by noise. As a type of commonly used denoising method, neural networks have shown their power to achieve satisfactory denoising results. However, their performance highly depends on high-quality noisy-clean image pairs. For the task of MSFA image denoising, there is currently neither a paired real dataset nor an accurate noise model capable of generating realistic noisy images. To this end, we present a physics-based noise model that is capable to match the real noise distribution and synthesize realistic noisy images. In our noise model, those different types of noise can be divided into SimpleDist component and ComplexDist component. The former contains all the types of noise that can be described using a simple probability distribution like Gaussian or Poisson distribution, and the latter contains the complicated color bias noise that cannot be modeled using a simple probability distribution. Besides, we design a noise-decoupled network consisting of a SimpleDist noise removal network (SNRNet) and a ComplexDist noise removal network (CNRNet) to sequentially remove each component. Moreover, according to the non-uniformity of color bias noise in our noise model, we introduce a learnable position embedding in CNRNet to indicate the position information. To verify the effectiveness of our physics-based noise model and noise-decoupled network, we collect a real MSFA denoising dataset with paired long-exposure clean images and short-exposure noisy images. Experiments are conducted to prove that the network trained using synthetic data generated by our noise model performs as well as trained using paired real data, and our noise-decoupled network outperforms other state-of-the-art denoising methods. Ying Fu 0001, Qiankun Liu 0001, Jun Zhang 0007 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Frequency Dynamic Convolution for Dense Image PredictionabstractWhile Dynamic Convolution (DY-Conv) has shown promising performance by enabling adaptive weight selection through multiple parallel weights combined with an attention mechanism, the frequency response of these weights tends to exhibit high similarity, resulting in high parameter costs but limited adaptability. In this work, we introduce Frequency Dynamic Convolution (FDConv), a novel approach that mitigates these limitations by learning a fixed parameter budget in the Fourier domain. FDConv divides this budget into frequency-based groups with disjoint Fourier indices, enabling the construction of frequency-diverse weights without increasing the parameter cost. To further enhance adaptability, we propose Kernel Spatial Modulation (KSM) and Frequency Band Modulation (FBM). KSM dynamically adjusts the frequency response of each filter at the spatial level, while FBM decomposes weights into distinct frequency bands in the frequency domain and modulates them dynamically based on local content. Extensive experiments on object detection, segmentation, and classification validate the effectiveness of FD-Conv. We demonstrate that when applied to ResNet-50, FDConv achieves superior performance with a modest increase of +3.6M parameters, outperforming previous methods that require substantial increases in parameter budgets (e.g., CondConv +90M, KW +76.5M). Moreover, FD-Conv seamlessly integrates into a variety of architectures, including ConvNeXt, Swin-Transformer, offering a flexible and efficient solution for modern vision tasks. The code is made publicly available at https://github.com/Linwei-Chen/FDConv. Lin Gu 0003, Liang Li 0003, Chenggang Yan 0001, Ying Fu 0001 |
CVPR | 5 |
| 2025 | Multi-Granularity Class Prototype Topology Distillation for Class-Incremental Source-Free Unsupervised Domain AdaptationabstractThis paper explores the Class-Incremental Source-Free Unsupervised Domain Adaptation (CI-SFUDA) problem, where the unlabeled target data come incrementally without access to labeled source instances. This problem poses two challenges, the interference of similar source-class knowledge in target-class representation learning and the shocks of new target knowledge to old ones. To address them, we propose the Multi-Granularity Class Prototype Topology Distillation (GROTO) algorithm, which effectively transfers the source knowledge to the class-incremental target domain. Concretely, we design the multi-granularity class prototype self-organization module and the prototype topology distillation module. First, we mine the positive classes by modeling accumulation distributions. Next, we introduce multi-granularity class prototypes to generate reliable pseudo-labels, and exploit them to promote the positive-class target feature self-organization. Second, the positive-class prototypes are leveraged to construct the topological structures of source and target feature spaces. Then, we perform the topology distillation to continually mitigate the shocks of new target knowledge to old ones. Extensive experiments demonstrate that our proposed method achieves state-of-the-art performance on three public datasets. Peihua Deng, Xichun Sheng, Chenggang Yan 0001, Yaoqi Sun, Ying Fu 0001, Liang Li 0003 |
CVPR | 6 |
| 2025 | Noise Calibration and Spatial-Frequency Interactive Network for STEM Image EnhancementabstractScanning Transmission Electron Microscopy (STEM) enables the observation of atomic arrangements at sub-angstrom resolution, allowing for atomically resolved analysis of the physical and chemical properties of materials. However, due to the effects of noise, electron beam damage, sample thickness, etc, obtaining satisfactory atomic-level images is often challenging. Enhancing STEM images can reveal clearer structural details of materials. Nonetheless, existing STEM image enhancement methods usually overlook unique features in the frequency domain, and existing datasets lack realism and generality. To resolve these issues, in this paper, we develop noise calibration, data synthesis, and enhancement methods for STEM images. We first present a STEM noise calibration method, which is used to synthesize more realistic STEM images. The parameters of background noise, scan noise, and pointwise noise are obtained by statistical analysis and fitting of real STEM images containing atoms. Then we use these parameters to develop a more general dataset that considers both regular and random atomic arrangements and includes both HAADF and BF mode images. Finally, we design a spatial-frequency interactive network for STEM image enhancement, which can explore the information in the frequency domain formed by the periodicity of atomic arrangement. Experimental results show that our data is closer to real STEM images and achieves better enhancement performances together with our network. Code will be available at https://github.com/HeasonLee/SFIN. Hesong Li, Ziqi Wu, Ruiwen Shao, Tao Zhang 0042, Ying Fu 0001 |
CVPR | 5 |
| 2025 | Frequency-Dynamic Attention Modulation for Dense PredictionabstractVision Transformers (ViTs) have significantly advanced computer vision, demonstrating strong performance across various tasks. However, the attention mechanism in ViTs makes each layer function as a low-pass filter, and the stacked-layer architecture in existing transformers suffers from frequency vanishing. This leads to the loss of critical details and textures. We propose a novel, circuit-theory-inspired strategy called Frequency-Dynamic Attention Modulation (FDAM), which can be easily plugged into ViTs. FDAM directly modulates the overall frequency response of ViTs and consists of two techniques: Attention Inversion (AttInv) and Frequency Dynamic Scaling (FreqScale). Since circuit theory uses low-pass filters as fundamental elements, we introduce AttInv, a method that generates complementary high-pass filtering by inverting the low-pass filter in the attention matrix, and dynamically combining the two. We further design FreqScale to weight different frequency components for fine-grained adjustments to the target response function. Through feature similarity analysis and effective rank evaluation, we demonstrate that our approach avoids representation collapse, leading to consistent performance improvements across various models, including SegFormer, DeiT, and MaskDINO. These improvements are evident in tasks such as semantic segmentation, object detection, and instance segmentation. Additionally, we apply our method to remote sensing detection, achieving state-of-the-art results in single-scale settings. The code is available at https://github.com/Linwei-Chen/FDAM. Lin Gu 0003, Ying Fu 0001 |
ICCV | 3 |
| 2025 | Learning Dense Feature Matching via Lifting Single 2D Image to 3D SpaceabstractFeature matching plays a fundamental role in many computer vision tasks, yet existing methods heavily rely on scarce and clean multi-view image collections, which constrains their generalization to diverse and challenging scenarios. Moreover, conventional feature encoders are typically trained on single-view 2D images, limiting their capacity to capture 3D-aware correspondences. In this paper, we propose a novel two-stage framework that lifts 2D images to 3D space, named as \textbf{Lift to Match (L2M)}, taking full advantage of large-scale and diverse single-view images. To be specific, in the first stage, we learn a 3D-aware feature encoder using a combination of multi-view image synthesis and 3D feature Gaussian representation, which injects 3D geometry knowledge into the encoder. In the second stage, a novel-view rendering strategy, combined with large-scale synthetic data generation from single-view images, is employed to learn a feature decoder for robust feature matching, thus achieving generalization across diverse domains. Extensive experiments demonstrate that our method achieves superior generalization across zero-shot evaluation benchmarks, highlighting the effectiveness of the proposed framework for robust feature matching. Yingping Liang, Yutao Hu 0002, Wenqi Shao, Ying Fu 0001 |
ICCV | 4 |
| 2025 | RobuSTereo: Robust Zero-Shot Stereo Matching under Adverse WeatherabstractLearning-based stereo matching models struggle in adverse weather conditions due to the scarcity of corresponding training data and the challenges in extracting discriminative features from degraded images. These limitations significantly hinder zero-shot generalization to out-of-distribution weather conditions. In this paper, we propose \textbf{RobuSTereo}, a novel framework that enhances the zero-shot generalization of stereo matching models under adverse weather by addressing both data scarcity and feature extraction challenges. First, we introduce a diffusion-based simulation pipeline with a stereo consistency module, which generates high-quality stereo data tailored for adverse conditions. By training stereo matching models on our synthetic datasets, we reduce the domain gap between clean and degraded images, significantly improving the models' robustness to unseen weather conditions. The stereo consistency module ensures structural alignment across synthesized image pairs, preserving geometric integrity and enhancing depth estimation accuracy. Second, we design a robust feature encoder that combines a specialized ConvNet with a denoising transformer to extract stable and reliable features from degraded images. The ConvNet captures fine-grained local structures, while the denoising transformer refines global representations, effectively mitigating the impact of noise, low visibility, and weather-induced distortions. This enables more accurate disparity estimation even under challenging visual conditions. Extensive experiments demonstrate that \textbf{RobuSTereo} significantly improves the robustness and generalization of stereo matching models across diverse adverse weather scenarios. Yingping Liang, Yutao Hu 0002, Ying Fu 0001 |
ICCV | 4 |
| 2025 | Boosting Zero-shot Stereo Matching Using Large-Scale Mixed Images Sources in the Real WorldabstractStereo matching methods rely on dense pixel-wise ground truth labels, which are laborious to obtain, especially for real-world datasets. The scarcity of labeled data and domain gaps between synthetic and real-world images also pose notable challenges. In this paper, we propose a novel framework, BooSTer, that leverages both vision foundation models and large-scale mixed image sources, including synthetic, real, and single-view images. First, to fully unleash the potential of large-scale single-view images, we design a data generation strategy combining monocular depth estimation and diffusion models to generate dense stereo matching data from single-view images. Second, to tackle sparse labels in real-world datasets, we transfer knowledge from monocular depth estimation models, using pseudo-mono depth labels and a dynamic scale- and shift-invariant loss for additional supervision. Furthermore, we incorporate vision foundation model as an encoder to extract robust and transferable features, boosting accuracy and generalization. Extensive experiments on benchmark datasets demonstrate the effectiveness of our approach, achieving significant improvements in accuracy over existing methods, particularly in scenarios with limited labeled data and domain shifts. Yingping Liang, Ying Fu 0001 |
IJCAI | 3 |
| 2025 | FCDFusion: A Fast, Low Color Deviation Method for Fusing Visible and Infrared Image PairsabstractVisible and infrared image fusion (VIF) aims to combine information from visible and infrared images into a single fused image. Previous VIF methods usually employ a color space transformation to keep the hue and saturation from the original visible image. However, for fast VIF methods, this operation accounts for the majority of the calculation and is the bottleneck preventing faster processing. In this paper, we propose a fast fusion method, FCDFusion, with little color deviation. It preserves color information without color space transformations, by directly operating in RGB color space. It incorporates gamma correction at little extra cost, allowing color and contrast to be rapidly improved. We regard the fusion process as a scaling operation on 3D color vectors, greatly simplifying the calculations. A theoretical analysis and experiments show that our method can achieve satisfactory results in only 7 FLOPs per pixel. Compared to state-of-the-art fast, color-preserving methods using HSV color space, our method provides higher contrast at only half of the computational cost. We further propose a new metric, color deviation, to measure the ability of a VIF method to preserve color. It is specifically designed for VIF tasks with color visible-light images, and overcomes deficiencies of existing VIF metrics used for this purpose. Our code is available at https://github.com/HeasonLee/FCDFusion. Hesong Li, Ying Fu 0001 |
Comput. Vis. Media | 2 |
| 2025 | LucIE: Language-Guided Local Image Editing for Fashion ImagesabstractLanguage-guided fashion image editing is challenging, as fashion image editing is local and requires high precision, while natural language cannot provide precise visual information for guidance. In this paper, we propose LucIE, a novel unsupervised language-guided local image editing method for fashion images. LucIE adopts and modifies recent text-to-image synthesis network, DF-GAN, as its backbone. However, the synthesis backbone often changes the global structure of the input image, making local image editing impractical. To increase structural consistency between input and edited images, we propose Content-Preserving Fusion Module (CPFM). Different from existing fusion modules, CPFM prevents iterative refinement on visual feature maps and accumulates additive modifications on RGB maps. LucIE achieves local image editing explicitly with language-guided image segmentation and mask-guided image blending while only using image and text pairs. Results on the DeepFashion dataset shows that LucIE achieves state-of-the-art results. Compared with previous methods, images generated by LucIE also exhibit fewer artifacts. We provide visualizations and perform ablation studies to validate LucIE and the CPFM. We also demonstrate and analyze limitations of LucIE, to provide a better understanding of LucIE. Huanglu Wen, Shaodi You, Ying Fu 0001 |
Comput. Vis. Media | 3 |
| 2025 | Relation-Guided Adversarial Learning for Data-Free Knowledge Transfer
Yingping Liang, Ying Fu 0001 |
Int. J. Comput. Vis. | 2 |
| 2025 | Unaligned RGB Guided Hyperspectral Image Super-Resolution with Spatial-Spectral Concordance
Yingkai Zhang, Zeqiang Lai, Ying Fu 0001, Chenghu Zhou |
Int. J. Comput. Vis. | 4 |
| 2025 | Spatial Frequency Modulation for Semantic SegmentationabstractHigh spatial frequency information, including fine details like textures, significantly contributes to the accuracy of semantic segmentation. However, according to the Nyquist-Shannon Sampling Theorem, high-frequency components are vulnerable to aliasing or distortion when propagating through downsampling layers such as strided-convolution. Here, we propose a novel Spatial Frequency Modulation (SFM) that modulates high-frequency features to a lower frequency before downsampling and then demodulates them back during upsampling. Specifically, we implement modulation through adaptive resampling (ARS) and design a lightweight add-on that can densely sample the high-frequency areas to scale up the signal, thereby lowering its frequency in accordance with the Frequency Scaling Property. We also propose Multi-Scale Adaptive Upsampling (MSAU) to demodulate the modulated feature and recover high-frequency information through non-uniform upsampling This module further improves segmentation by explicitly exploiting information interaction between densely and sparsely resampled areas at multiple scales. Both modules can seamlessly integrate with various architectures, extending from convolutional neural networks to transformers. Feature visualization and analysis demonstrate that our method effectively alleviates aliasing while successfully retaining details after demodulation. As a result, the proposed approach considerably enhances existing state-of-the-art segmentation models (e.g., Mask2Former-Swin-T +1.5 mIoU, InternImage-T +1.4 mIoU on ADE20 K). Furthermore, ARS also enhances the performance of powerful Deformable Convolution (+0.8 mIoU on Cityscapes) by maintaining relative positional order during non-uniform sampling. Finally, we validate the broad applicability and effectiveness of SFM by extending it to image classification, adversarial robustness, instance segmentation, and panoptic segmentation tasks. Ying Fu 0001, Lin Gu 0003, Dezhi Zheng, Jifeng Dai |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Latent Diffusion Enhanced Rectangle Transformer for Hyperspectral Image RestorationabstractThe restoration of hyperspectral image (HSI) plays a pivotal role in subsequent hyperspectral image applications. Despite the remarkable capabilities of deep learning, current HSI restoration methods face challenges in effectively exploring the spatial non-local self-similarity and spectral low-rank property inherently embedded with HSIs. This paper addresses these challenges by introducing a latent diffusion enhanced rectangle Transformer for HSI restoration, tackling the non-local spatial similarity and HSI-specific latent diffusion low-rank property. In order to effectively capture non-local spatial similarity, we propose the multi-shape spatial rectangle self-attention module in both horizontal and vertical directions, enabling the model to utilize informative spatial regions for HSI restoration. Meanwhile, we propose a spectral latent diffusion enhancement module that generates the image-specific latent dictionary based on the content of HSI for low-rank vector extraction and representation. This module utilizes a diffusion model to generatively obtain representations of global low-rank vectors, thereby aligning more closely with the desired HSI. A series of comprehensive experiments were carried out on four common hyperspectral image restoration tasks, including HSI denoising, HSI super-resolution, HSI reconstruction, and HSI inpainting. The results of these experiments highlight the effectiveness of our proposed method, as demonstrated by improvements in both objective metrics and subjective visual quality. Miaoyu Li, Ying Fu 0001, Tao Zhang 0042, Ji Liu 0003, Dejing Dou, Chenggang Yan 0001, Yulun Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Flow-Anything: Learning Real-World Optical Flow Estimation From Large-Scale Single-View ImagesabstractOptical flow estimation is a crucial subfield of computer vision, serving as a foundation for video tasks. However, the real-world robustness is limited by animated synthetic datasets for training. This introduces domain gaps when applied to real-world applications and limits the benefits of scaling up datasets. To address these challenges, we propose Flow-Anything, a large-scale data generation framework designed to learn optical flow estimation from any single-view images in the real world. We employ two effective steps to make data scaling-up promising. First, we convert a single-view image into a 3D representation using advanced monocular depth estimation networks. This allows us to render optical flow and novel view images under a virtual camera. Second, we develop an Object-Independent Volume Rendering module and a Depth-Aware Inpainting module to model the dynamic objects in the 3D representation. These two steps allow us to generate realistic datasets for training from large-scale single-view images, namely FA-Flow Dataset. For the first time, we demonstrate the benefits of generating optical flow training data from large-scale real-world images, outperforming the most advanced unsupervised methods and supervised methods on synthetic datasets. Moreover, our models serve as a foundation model and enhance the performance of various downstream video tasks. Yingping Liang, Ying Fu 0001, Yutao Hu 0002, Wenqi Shao, Debing Zhang |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Learning Rain Location Prior for Nighttime Deraining and BeyondabstractMost deraining methods work on day scenes while leaving nighttime deraining underexplored, where darkness and non-uniform illuminations pose additional challenges. Consequently, night rain has a quite different appearance varying by location and cannot be effectively handled. To accommodate this issue, we propose a Rain Location Prior (RLP) by implicitly learning it from rainy images to reflect rain location information and boost the performance of deraining models by prior injection. Then, we introduce a Rain Prior Injection Module (RPIM) with a multi-scale scheme to modulate it by attention and emphasize the features of rain streak areas for better injection efficiency. Finally, to alleviate the data scarcity issue and facilitate the research on nighttime deraining, we propose the GTAV-NightRain dataset by considering the interaction between rain streaks and non-uniform illuminations, and provide detailed instructions on data collection pipeline which is highly replicable and flexible to integrate challenging factors of rainy night in the future. Our method outperforms state-of-the-art backbone by 1.3 dB in PSNR and generalizes better on real data such as heavy rain and the presence of glow and glaring lights. Ablation studies are conducted to validate the effectiveness of each component and we visualize RLP to show good interpretability. Moreover, we apply our method to daytime deraining and desnow to show good generalizability on other location-dependent degradations. Our method is a step forward in nighttime deraining and the GTAV-NightRain dataset may become a good complement to previous datasets. Fan Zhang 0123, Shaodi You, Yu Li 0003, Ying Fu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | EventHDR: From Event to High-Speed HDR Videos and BeyondabstractEvent cameras are innovative neuromorphic sensors that asynchronously capture the scene dynamics. Due to the event-triggering mechanism, such cameras record event streams with much shorter response latency and higher intensity sensitivity compared to conventional cameras. On the basis of these features, previous works have attempted to reconstruct high dynamic range (HDR) videos from events, but have either suffered from unrealistic artifacts or failed to provide sufficiently high frame rates. In this paper, we present a recurrent convolutional neural network that reconstruct high-speed HDR videos from event sequences, with a key frame guidance to prevent potential error accumulation caused by the sparse event data. Additionally, to address the problem of severely limited real dataset, we develop a new optical system to collect a real-world dataset with paired high-speed HDR videos and event streams, facilitating future research in this field. Our dataset provides the first real paired dataset for event-to-HDR reconstruction, avoiding potential inaccuracies from simulation strategies. Experimental results demonstrate that our method can generate high-quality, high-speed HDR videos. We further explore the potential of our work in cross-camera reconstruction and downstream computer vision tasks, including object detection, panoramic segmentation, optical flow estimation, and monocular depth estimation under HDR scenarios. Yunhao Zou, Ying Fu 0001, Tsuyoshi Takatani, Yinqiang Zheng |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Calibration-Free Raw Image Denoising via Fine-Grained Noise EstimationabstractImage denoising has progressed significantly due to the development of effective deep denoisers. To improve the performance in real-world scenarios, recent trends prefer to formulate superior noise models to generate realistic training data, or estimate noise levels to steer non-blind denoisers. In this paper, we bridge both strategies by presenting an innovative noise estimation and realistic noise synthesis pipeline. Specifically, we integrates a fine-grained statistical noise model and contrastive learning strategy, with a unique data augmentation to enhance learning ability. Then, we use this model to estimate noise parameters on evaluation dataset, which are subsequently used to craft camera-specific noise distribution and synthesize realistic noise. One distinguishing feature of our methodology is its adaptability: our pre-trained model can directly estimate unknown cameras, making it possible to unfamiliar sensor noise modeling using only testing images, without calibration frames or paired training data. Another highlight is our attempt in estimating parameters for fine-grained noise models, which extends the applicability to even more challenging low-light conditions. Through empirical testing, our calibration-free pipeline demonstrates effectiveness in both normal and low-light scenarios, further solidifying its utility in real-world noise synthesis and denoising tasks. Yunhao Zou, Ying Fu 0001, Yulun Zhang 0001, Tao Zhang 0042, Chenggang Yan 0001, Radu Timofte |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Restoration of Images Taken Through a Dirty Window Using Optics-Guided TransformerabstractTaking photographs through windows is an inevitable scenario in the real world, but glass windows are not ideally clean in most cases. Although there exists various raindrop removal methods, the occlusion of dirt, as another dirty window case, has not been well valued. The vital reasons include i) the limitation of the optical imaging model proposed in previous methods, and ii) the shortage of a practical dataset for sufficient types of dirty glass windows. To fill this research gap, in this paper, we first propose a general optical imaging model that fits widely used dirty window cases. Following this, training and testing synthetic datasets are generated, and real-world dirty window data are collected to evaluate the effectiveness of our imaging model and synthetic data. For the methodology part, we propose an optics-guided Transformer network to solve this special image restoration problem, i.e., the dirt removal for images taken through a dirty window. Experimental results demonstrate that our imaging model is effective and robust. Our proposed network leads to higher performance than existing methods on both synthetic and real-world dirty window images. Code and data are available at https://github.com/Zongliang-Wu/ReDNet. Zongliang Wu, Juzheng Zhang, Ying Fu 0001, Yulun Zhang 0001, Xin Yuan 0002 |
IEEE Trans. Image Process. | 3 |
| 2025 | Hyperspectral Image Super Resolution With Real Unaligned RGB GuidanceabstractFusion-based hyperspectral image (HSI) super-resolution has become increasingly prevalent for its capability to integrate high-frequency spatial information from the paired high-resolution (HR) RGB reference (Ref-RGB) image. However, most of the existing methods either heavily rely on the accurate alignment between low-resolution (LR) HSIs and RGB images or can only deal with simulated unaligned RGB images generated by rigid geometric transformations, which weakens their effectiveness for real scenes. In this article, we explore the fusion-based HSI super-resolution with real Ref-RGB images that have both rigid and nonrigid misalignments. To properly address the limitations of existing methods for unaligned reference images, we propose an HSI fusion network (HSIFN) with heterogeneous feature extractions, multistage feature alignments, and attentive feature fusion. Specifically, our network first transforms the input HSI and RGB images into two sets of multiscale features with an HSI encoder and an RGB encoder, respectively. The features of Ref-RGB images are then processed by a multistage alignment module to explicitly align the features of Ref-RGB with the LR HSI. Finally, the aligned features of Ref-RGB are further adjusted by an adaptive attention module to focus more on discriminative regions before sending them to the fusion decoder to generate the reconstructed HR HSI. Additionally, we collect a real-world HSI fusion dataset, consisting of paired HSI and unaligned Ref-RGB, to support the evaluation of the proposed model for real scenes. Extensive experiments are conducted on both simulated and our real-world datasets, and it shows that our method obtains a clear improvement over existing single-image and fusion-based super-resolution methods on quantitative assessment as well as visual comparison. The code and dataset are publicly available at https://zeqiang-lai.github.io/HSI-RefSR/. Zeqiang Lai, Ying Fu 0001, Jun Zhang 0007 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | Supervise-Assisted Self-Supervised Deep-Learning Method for Hyperspectral Image RestorationabstractHyperspectral image (HSI) restoration is a challenging research area, covering a variety of inverse problems. Previous works have shown the great success of deep learning in HSI restoration. However, facing the problem of distribution gaps between training HSIs and target HSI, those data-driven methods falter in delivering satisfactory outcomes for the target HSIs. In addition, the degradation process of HSIs is usually disturbed by noise, which is not well taken into account in existing restoration methods. The existence of noise further exacerbates the dissimilarities within the data, rendering it challenging to attain desirable results without an appropriate learning approach. To track these issues, in this article, we propose a supervise-assisted self-supervised deep-learning method to restore noisy degraded HSIs. Initially, we facilitate the restoration network to acquire a generalized prior through supervised learning from extensive training datasets. Then, the self-supervised learning stage is employed and utilizes the specific prior of the target HSI. Particularly, to restore clean HSIs during the self-supervised learning stage from noisy degraded HSIs, we introduce a noise-adaptive loss function that leverages inner statistics of noisy degraded HSIs for restoration. The proposed noise-adaptive loss consists of Stein's unbiased risk estimator (SURE) and total variation (TV) regularizer and fine-tunes the network with the presence of noise. We demonstrate through experiments on different HSI tasks, including denoising, compressive sensing, super-resolution, and inpainting, that our method outperforms state-of-the-art methods on benchmarks under quantitative metrics and visual quality. The code is available at https://github.com/ying-fu/SSDL-HSI. Miaoyu Li, Ying Fu 0001, Tao Zhang 0042, Guanghui Wen |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2025 | ProbIBR: Fast Image-Based Rendering With Learned Probability-Guided SamplingabstractWe present a general, fast, and practical solution for interpolating novel views of diverse real-world scenes given a sparse set of nearby views. Existing generic novel view synthesis methods rely on time-consuming scene geometry pre-computation or redundant sampling of the entire space for neural volumetric rendering, limiting the overall efficiency. Instead, we incorporate learned MVS priors into the neural volume rendering pipeline while improving the rendering efficiency by reducing sampling points under the guidance of depth probability distributions. Specifically, fewer but important points are sampled under the guidance of depth probability distributions extracted from the learned MVS architecture. Based on the learned probability-guided sampling, we develop a sophisticated neural volume rendering module that effectively integrates source view information with the learned scene structures. We further propose confidence-aware refinement to improve the rendering results in uncertain, occluded, and unreferenced regions. Moreover, we build a four-view camera system for holographic display and provide a real-time version of our framework for free-viewpoint experience, where novel view images of a spatial resolution of 512×512 can be rendered at around 20 fps on a single GTX 3090 GPU. Experiments show that our method achieves 15 to 40 times faster rendering compared to state-of-the-art baselines, with strong generalization capacity and comparable high-quality novel view synthesis performance. Yuemei Zhou, Tao Yu 0007, Zerong Zheng, Gaochang Wu, Guihua Zhao, Ying Fu 0001, Yebin Liu |
IEEE Trans. Vis. Comput. Graph. | 7 |
| 2024 | Frequency-Adaptive Dilated Convolution for Semantic SegmentationabstractDilated convolution, which expands the receptive field by inserting gaps between its consecutive elements, is widely employed in computer vision. In this study, we propose three strategies to improve individual phases of dilated convolution from the perspective of spectrum analysis. Departing from the conventional practice of fixing a global dilation rate as a hyperparameter, we introduce Frequency-Adaptive Dilated Convolution (FADC), which dynamically adjusts dilation rates spatially based on local frequency components. Subsequently, we design two plug-in modules to directly enhance effective bandwidth and receptive field size. The Adaptive Kernel (AdaKern) module decomposes convolution weights into low-frequency and high-frequency components, dynamically adjusting the ratio between these components on a per-channel basis. By increasing the high-frequency part of convolution weights, AdaKern captures more high-frequency components, thereby improving effective bandwidth. The Frequency Selection (FreqSelect) module optimally balances high- and low-frequency components in feature representations through spatially variant reweighting. It suppresses high frequencies in the background to encourage FADC to learn a larger dilation, thereby increasing the receptive field for an expanded scope. Extensive experiments on segmentation and object detection consistently validate the efficacy of our approach. The code is made publicly available at https://github.com/ying-fu/FADC. Lin Gu 0003, Dezhi Zheng, Ying Fu 0001 |
CVPR | 4 |
| 2024 | Suppress and Rebalance: Towards Generalized Multi-Modal Face Anti-SpoofingabstractFace Anti-Spoofing (FAS) is crucial for securing face recognition systems against presentation attacks. With ad-vancements in sensor manufacture and multi-modal learning techniques, many multi-modal FAS approaches have emerged. However, they face challenges in generalizing to unseen attacks and deployment conditions. These chal-lenges arise from (1) modality unreliability, where some modality sensors like depth and infrared undergo signifi-cant domain shifts in varying environments, leading to the spread of unreliable information during cross-modal feature fusion, and (2) modality imbalance, where training overly relies on a dominant modality hinders the conver-gence of others, reducing effectiveness against attack types that are indistinguishable by sorely using the dominant modality. To address modality unreliability, we propose the Uncertainty-Guided Cross-Adapter (U-Adapter) to recognize unreliably detected regions within each modality and suppress the impact of unreliable regions on other modal-ities. For modality imbalance, we propose a Rebalanced Modality Gradient Modulation (ReGrad) strategy to rebal-ance the convergence speed of all modalities by adaptively adjusting their gradients. Besides, we provide the first large-scale benchmark for evaluating multi-modal FAS per-formance under domain generalization scenarios. Exten-sive experiments demonstrate that our method outperforms state-of-the-art methods. Source codes and protocols are released on https://github.com/OMGGGGG/mmdg. Xun Lin, Shuai Wang 0049, Rizhao Cai, Yizhong Liu, Ying Fu 0001, Wenzhong Tang, Zitong Yu, Alex Chichung Kot |
CVPR | 5 |
| 2024 | Infrared Small Target Detection with Scale and Location SensitivityabstractRecently, infrared small target detection (IRSTD) has been dominated by deep-learning-based methods. However, these methods mainly focus on the design of complex model structures to extract discriminative features, leaving the loss functions for IRSTD under-explored. For ex-ample, the widely used Intersection over Union (IoU) and Dice losses lack sensitivity to the scales and locations of targets, limiting the detection performance of detectors. In this paper, we focus on boosting detection performance with a more effective loss but a simpler model structure. Specifically, we first propose a novel Scale and Location Sensitive (SLS) loss to handle the limitations of existing losses: 1) for scale sensitivity, we compute a weight for the IoU loss based on target scales to help the detector distinguish targets with different scales: 2) for location sensitivity, we introduce a penalty term based on the center points of targets to help the detector localize targets more precisely. Then, we design a simple Multi-Scale Head to the plain U-Net (MSHNet). By applying SLS loss to each scale of the predictions, our MSHNet outperforms existing state-of-the-art methods by a large margin. In addition, the detection performance of existing detectors can be further improved when trained with our SLS loss, demonstrating the effectiveness and generalization of our SLS loss. The code is available at https://github.com/ying-fu/MSHNet. Qiankun Liu 0001, Rui Liu 0039, Bolun Zheng, Hongkui Wang, Ying Fu 0001 |
CVPR | 5 |
| 2024 | Learning Visual Prompt for Gait RecognitionabstractGait, a prevalent and complex form of human motion, plays a significant role in the field of long-range pedestrian retrieval due to the unique characteristics inherent in individual motion patterns. However, gait recognition in real-world scenarios is challenging due to the limitations of capturing comprehensive cross-viewing and crossclothing data. Additionally, distractors such as occlusions, directional changes, and lingering movements further complicate the problem. The widespread application of deep learning techniques has led to the development of various potential gait recognition methods. However, these methods utilize convolutional networks to extract shared information across different views and attire conditions. Once trained, the parameters and non-linear function become constrained to fixed patterns, limiting their adaptability to various distractors in real-world scenarios. In this paper, we present a unified gait recognition framework to extract global motion patterns and develop a novel dynamic transformer to generate representative gait features. Specifically, we develop a trainable part-based prompt pool with numerous key-value pairs that can dynamically select prompt templates to incorporate into the gait sequence, thereby providing task-relevant shared knowledge information. Furthermore, we specifically design dynamic attention to extract robust motion patterns and address the length generalization issue. Extensive experiments on four widely recognized gait datasets, i.e., Gait3D, GREW, OUMVLP, and CASIA-B, reveal that the proposed method yields substantial improvements compared to current state-of-the-art approaches. Ying Fu 0001, Chunshui Cao, Saihui Hou, Yongzhen Huang, Dezhi Zheng |
CVPR | 2 |
| 2024 | Multi-Object Tracking in the DarkabstractLow-light scenes are prevalent in real-world applications (e.g. autonomous driving and surveillance at night). Recently, multi-object tracking in various practical use cases have received much attention, but multi-object tracking in dark scenes is rarely considered. In this paper, we focus on multi-object tracking in dark scenes. To address the lack of datasets, we first build a Low-light Multi-Object Tracking (LMOT) dataset. LMOT provides well-aligned low-light video pairs captured by our dual-camera system, and high-quality multi-object tracking annotations for all videos. Then, we propose a low-light multi-object tracking method, termed as LTrack. We introduce the adaptive low-pass downsample module to enhance low-frequency components of images outside the sensor noises. The degradation suppression learning strategy enables the model to learn invariant information under noise disturbance and image quality degradation. These components improve the robustness of multi-object tracking in dark scenes. We conducted a comprehensive analysis of our LMOT dataset and proposed LTrack. Experimental results demonstrate the superiority of the proposed method and its competitiveness in real night low-light scenes. Dataset and Code: https:/github.com/ying-fu/LMOT Qiankun Liu 0001, Yunhao Zou, Ying Fu 0001 |
CVPR | 5 |
| 2024 | Atlantis: Enabling Underwater Depth Estimation with Stable DiffusionabstractMonocular depth estimation has experienced significant progress on terrestrial images in recent years thanks to deep learning advancements. But it remains inadequate for underwater scenes primarily due to data scarcity. Given the inherent challenges of light attenuation and backscat-ter in water, acquiring clear underwater images or precise depth is notably difficult and costly. To mitigate this issue, learning-based approaches often rely on synthetic data or turn to self- or unsupervised manners. Nonetheless, their performance is often hindered by domain gap and looser constraints. In this paper, we propose a novel pipeline for generating photorealistic underwater images using accurate terrestrial depth. This approach facilitates the supervised training of models for underwater depth estimation, effectively reducing the performance disparity between ter-restrial and underwater environments. Contrary to previous synthetic datasets that merely apply style transfer to terres-trial images without scene content change, our approach uniquely creates vivid non-existent underwater scenes by leveraging terrestrial depth data through the innovative Stable Diffusion model. Specifically, we introduce a specialized Depth2Underwater ControlNet, trained on prepared {Underwater, Depth, Text} data triplets, for this generation task. Our newly developed dataset, Atlantis, enables terres-trial depth estimation models to achieve considerable improvements on unseen underwater scenes, surpassing their terrestrial pretrained counterparts both quantitatively and qualitatively. Moreover, we further show its practical utility by applying the improved depth in underwater image enhancement, and its smaller domain gap from the LLVM perspective. Code and dataset are publicly available at https://github.com/zkawfanx/Atlantis. Fan Zhang 0123, Shaodi You, Yu Li 0003, Ying Fu 0001 |
CVPR | 4 |
| 2024 | Binarized Low-Light Raw Video EnhancementabstractRecently, deep neural networks have achieved excellent performance on low-light raw video enhancement. How-ever, they often come with high computational complexity and large memory costs, which hinder their applications on resource-limited devices. In this paper, we explore the feasibility of applying the extremely compact binary neural network (BNN) to low-light raw video enhancement. Nev-ertheless, there are two main issues with binarizing video enhancement models. One is how to fuse the temporal in-formation to improve low-light denoising without complex modules. The other is how to narrow the performance gap between binary convolutions with the full precision ones. To address the first issue, we introduce a spatial-temporal shift operation, which is easy-to-binarize and effective. The temporal shift efficiently aggregates the features of neigh-bor frames and the spatial shift handles the misalignment caused by the large motion in videos. For the second issue, we present a distribution-aware binary convolution, which captures the distribution characteristics of real-valued in-put and incorporates them into plain binary convolutions to alleviate the degradation in performance. Extensive quantitative and qualitative experiments have shown our high-efficiency binarized low-light raw video enhancement method can attain a promising performance. The code is available at https://github.com/ying-fuIBRVE. Gengchen Zhang, Yulun Zhang 0001, Xin Yuan 0002, Ying Fu 0001 |
CVPR | 4 |
| 2024 | Object-Aware NIR-to-Visible Translation
Yunyi Gao, Lin Gu 0003, Qiankun Liu 0001, Ying Fu 0001 |
ECCV (23) | 4 |
| 2024 | Latent Diffusion Prior Enhanced Deep Unfolding for Snapshot Spectral Compressive Imaging
Zongliang Wu, Ruiying Lu, Ying Fu 0001, Xin Yuan 0002 |
ECCV (33) | 3 |
| 2024 | A Cross-modal Fusion Method for Multispectral Small Ship DetectionabstractThe fusion module of RGB and infrared (IR) remote sensing images is the key of multispectral ship detection. Existing works have shown that the cross-attention-based feature fusion can achieve good performance by extracting the complementary information of RGB and IR modalities. However, the existing commonly used cross-attention mechanisms introduce lots of redundancy parameters and mainly focus on global feature interaction of multispectral images, ignoring local detail information that is also important for small ship detection. In this paper, we propose a novel multispectral ship detection approach named LoGFusion. In LoGFusion, we design the cross stage partial module with partial convolution (CSPMPC) to reduce feature redundancy and utilize the local cross-modal fusion module (LoCFM) and global cross-modal fusion module (GCFM) to capture both local and global cross-modal features. Furthermore, we introduce a Multispectral Small Ship Dataset (MSSD) containing over 5k ship targets for small target detection. Experiments on MSSD validate the effectiveness of our method in terms of small ship detection in multispectral images. Yang Liu 0119, Yu Liu 0005, Xueqian Wang 0002, Linping Zhang, Zhizhuo Jiang, Yaowen Li, Chenggang Yan 0001, Ying Fu 0001, Tao Zhang 0042 |
FUSION | 8 |
| 2024 | When Semantic Segmentation Meets Frequency AliasingabstractDespite recent advancements in semantic segmentation, where and what pixels are hard to segment remains largely unexplored.
Existing research only separates an image into easy and hard regions and empirically observes the latter are associated with object boundaries.
In this paper, we conduct a comprehensive analysis of hard pixel errors, categorizing them into three types: false responses, merging mistakes, and displacements.
Our findings reveal a quantitative association between hard pixels and aliasing,
which is distortion caused by the overlapping of frequency components in the Fourier domain during downsampling.
To identify the frequencies responsible for aliasing, we propose using the equivalent sampling rate to calculate the Nyquist frequency, which marks the threshold for aliasing.
Then, we introduce the aliasing score as a metric to quantify the extent of aliasing.
While positively correlated with the proposed aliasing score, three types of hard pixels exhibit different patterns.
Here, we propose two novel de-aliasing filter (DAF) and frequency mixing (FreqMix) modules to alleviate aliasing degradation by accurately removing or adjusting frequencies higher than the Nyquist frequency.
The DAF precisely removes the frequencies responsible for aliasing before downsampling,
while the FreqMix dynamically selects high-frequency components within the encoder block.
Experimental results demonstrate consistent improvements in semantic segmentation and low-light instance segmentation tasks.
The code is at: \url{https://github.com/Linwei-Chen/Seg-Aliasing}. Lin Gu 0003, Ying Fu 0001 |
ICLR | 3 |
| 2024 | End-to-End Video Text Spotting with Transformer
Weijia Wu 0001, Yuanqiang Cai, Chunhua Shen, Debing Zhang, Ying Fu 0001, Ping Luo 0002 |
Int. J. Comput. Vis. | 5 |
| 2024 | Frequency-Aware Feature Fusion for Dense Image PredictionabstractDense image prediction tasks demand features with strong category information and precise spatial boundary details at high resolution. To achieve this, modern hierarchical models often utilize feature fusion, directly adding upsampled coarse features from deep layers and high-resolution features from lower levels. In this paper, we observe rapid variations in fused feature values within objects, resulting in intra-category inconsistency due to disturbed high-frequency features. Additionally, blurred boundaries in fused features lack accurate high frequency, leading to boundary displacement. Building upon these observations, we propose Frequency-Aware Feature Fusion (FreqFusion), integrating an Adaptive Low-Pass Filter (ALPF) generator, an offset generator, and an Adaptive High-Pass Filter (AHPF) generator. The ALPF generator predicts spatially-variant low-pass filters to attenuate high-frequency components within objects, reducing intra-class inconsistency during upsampling. The offset generator refines large inconsistent features and thin boundaries by replacing inconsistent features with more consistent ones through resampling, while the AHPF generator enhances high-frequency detailed boundary information lost during downsampling. Comprehensive visualization and quantitative analysis demonstrate that FreqFusion effectively improves feature consistency and sharpens object boundaries. Extensive experiments across various dense prediction tasks confirm its effectiveness. Ying Fu 0001, Lin Gu 0003, Chenggang Yan 0001, Tatsuya Harada, Gao Huang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2024 | Transformer Based Pluralistic Image Completion With Reduced Information LossabstractTransformer based methods have achieved great success in image inpainting recently. However, we find that these solutions regard each pixel as a token, thus suffering from an information loss issue from two aspects: 1) They downsample the input image into much lower resolutions for efficiency consideration. 2) They quantize 2563RGB values to a small number (such as 512) of quantized color values. The indices of quantized pixels are used as tokens for the inputs and prediction targets of the transformer. To mitigate these issues, we propose a new transformer based framework called “PUT”. Specifically, to avoid input downsampling while maintaining computation efficiency, we design a patch-based auto-encoder P-VQVAE. The encoder converts the masked image into non-overlapped patch tokens and the decoder recovers the masked regions from the inpainted tokens while keeping the unmasked regions unchanged. To eliminate the information loss caused by input quantization, an Un-quantized Transformer is applied. It directly takes features from the P-VQVAE encoder as input without any quantization and only regards the quantized tokens as prediction targets.Furthermore, to make the inpainting process more controllable, we introduce semantic and structural conditions as extra guidance. Extensive experiments show that our method greatly outperforms existing transformer based methods on image fidelity and achieves much higher diversity and better fidelity than state-of-the-art pluralistic inpainting methods on complex large-scale datasets (e.g., ImageNet). Codes are available athttps://github.com/liuqk3/PUT. Qiankun Liu 0001, Zhentao Tan, Dongdong Chen 0001, Ying Fu 0001, Qi Chu 0001, Gang Hua 0001, Nenghai Yu |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | Raw Image Based Over-Exposure Correction Using Channel-Guidance StrategyabstractMost existing methods for over-exposure in image correction are developed based on sRGB images, which can result in complex and non-linear degradation due to the image signal processing pipeline. By contrast, data-driven approaches based on RAW image data offer natural advantages for image processing tasks. RAW images, characterized by their near-linear correlation with scene radiance and enriched information content due to higher bit depth, demonstrate superior performance compared to sRGB-based techniques. Further, the spectral sensitivity characteristics intrinsic to digital camera sensors indicate that the blue and red channels in a Bayer pattern RAW image typically encompass more contextual information than the green channels. This property renders them less susceptible to over-exposure, thereby making them more effective for data extraction in high dynamic range scenes. In this paper, we introduce a Channel-Guidance Network (CGNet) that leverages the benefits of RAW images for over-exposure correction. The CGNet estimates the properly-exposed sRGB image directly from the over-exposed RAW image in an end-to-end manner. Specifically, we introduce a RAW-based channel-guidance branch to the U-net-based backbone, which exploits the color channel intensity prior of RAW images to achieve superior over-exposure correction performance. To further facilitate research in over-exposure correction, we present synthetic and real-world over-exposure correction benchmark datasets. These datasets comprise a large set of paired RAW and sRGB images across a variety of scenarios. Experiments on our RAW-sRGB datasets validate the advantages of our RAW-based channel guidance strategy and proposed CGNet over state-of-the-art sRGB-based methods on over-exposure correction. Our code and dataset are publicly available athttps://github.com/whiteknight-WJN/CGNet. Ying Fu 0001, Yunhao Zou, Qiankun Liu 0001, Chenggang Yan 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | Category-Level Band Learning-Based Feature Extraction for Hyperspectral Image ClassificationabstractHyperspectral image (HSI) classification is a classical task in remote sensing image analysis. With the development of deep learning, schemes based on deep learning have gradually become the mainstream of HSI classification. However, existing HSI classification schemes either lack the exploration of category-specific information in the spectral bands and the intrinsic value of information contained in features at different scales, or are unable to extract multiscale spatial information and global spectral properties simultaneously. To solve these problems, in this article, we propose a novel HSI classification framework named CL-MGNet, which can fully exploit the category-specific properties in spectral bands and obtain features with multiscale spatial information and global spectral properties. Specifically, we first propose a spectral weight learning (SWL) module with a category consistency loss to achieve the enhancement of information in important bands and the mining of category-specific properties. Then, a multiscale backbone is proposed to extract the spatial information at different scales and the cross-channel attention via multiscale convolution and a grouping attention module. Finally, we employ an attention multilayer perceptron (attention-MLP) block to exploit the global spectral properties of HSI, which is helpful for the final fully connected layer to obtain the classification result. The experimental results on five representative hyperspectral remote sensing datasets demonstrate the superiority of our method. Ying Fu 0001, Hongrong Liu, Yunhao Zou, Shuai Wang 0049, Zhongxiang Li, Dezhi Zheng |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2024 | Siamese-DETR for Generic Multi-Object TrackingabstractThe ability to detect and track the dynamic objects in different scenes is fundamental to real-world applications, e.g., autonomous driving and robot navigation. However, traditional Multi-Object Tracking (MOT) is limited to track objects belonging to the pre-defined closed-set categories. Recently, Generic MOT (GMOT) is proposed to track interested objects beyond pre-defined categories and it can be divided into Open-Vocabulary MOT (OVMOT) and Template-Image-based MOT (TIMOT). Taking the consideration that the expensive well pre-trained (vision-)language model and fine-grained category annotations are required to train OVMOT models, in this paper, we focus on TIMOT and propose a simple but effective method, Siamese-DETR. Only the commonly used detection datasets (e.g., COCO) are required for training. Different from existing TIMOT methods, which train a Single Object Tracking (SOT) based detector to detect interested objects and then apply a data association based MOT tracker to get the trajectories, we leverage the inherent object queries in DETR variants. Specifically: 1) The multi-scale object queries are designed based on the given template image, which are effective for detecting different scales of objects with the same category as the template image; 2) A dynamic matching training strategy is introduced to train Siamese-DETR on commonly used detection datasets, which takes full advantage of provided annotations; 3) The online tracking pipeline is simplified through a tracking-by-query manner by incorporating the tracked boxes in the previous frame as additional query boxes. The complex data association is replaced with the much simpler Non-Maximum Suppression (NMS). Extensive experimental results show that Siamese-DETR surpasses existing MOT methods on GMOT-40 dataset by a large margin. Qiankun Liu 0001, Ying Fu 0001 |
IEEE Trans. Image Process. | 4 |
| 2024 | SASAN: Spectrum-Axial Spatial Approach Networks for Medical Image SegmentationabstractOphthalmic diseases such as central serous chorioretinopathy (CSC) significantly impair the vision of millions of people globally. Precise segmentation of choroid and macular edema is critical for diagnosing and treating these conditions. However, existing 3D medical image segmentation methods often fall short due to the heterogeneous nature and blurry features of these conditions, compounded by medical image clarity issues and noise interference arising from equipment and environmental limitations. To address these challenges, we propose the Spectrum Analysis Synergy Axial-Spatial Network (SASAN), an approach that innovatively integrates spectrum features using the Fast Fourier Transform (FFT). SASAN incorporates two key modules: the Frequency Integrated Neural Enhancer (FINE), which mitigates noise interference, and the Axial-Spatial Elementum Multiplier (ASEM), which enhances feature extraction. Additionally, we introduce the Self-Adaptive Multi-Aspect Loss (LSM), which balances image regions, distribution, and boundaries, adaptively updating weights during training. We compiled and meticulously annotated the Choroid and Macular Edema OCT Mega Dataset (CMED-18k), currently the world’s largest dataset of its kind. Comparative analysis against 13 baselines shows our method surpasses these benchmarks, achieving the highest Dice scores and lowest HD95 in the CMED and OIMHS datasets. Our code is publicly available at https://github.com/IMOP-lab/SASAN-Pytorch. Xingru Huang, Jian Huang 0015, Tianyun Zhang, Changpeng Yue, Xuanbin Chen, Qianni Zhang, Ying Fu 0001, Yangyundou Wang |
IEEE Trans. Medical Imaging | 11 |
| 2024 | AnimeDiff: Customized Image Generation of Anime Characters Using Diffusion ModelabstractDue to the unprecedented power of text-to-image diffusion models, customizing these models to generate new concepts has gained increasing attention. Existing works have achieved some success on real-world concepts, but fail on the concepts of anime characters. We empirically find that such low quality comes from the newly introduced identifier text tokens, which are optimized to identify different characters. In this paper, we proposeAnimeDiffwhich focuses on customized image generation of anime characters. Our AnimeDiff directly binds anime characters with their names and keeps the embeddings of text tokens unchanged. Furthermore, when composing multiple characters in a single image, the model tends to confuse the properties of those characters. To address this issue, our AnimeDiff incorporates aCut-and-Pastedata augmentation strategy that produces multi-character images for training by cutting and pasting multiple characters onto background images. Experiments are conducted to prove the superiority of AnimeDiff over other methods. Qiankun Liu 0001, Dongdong Chen 0001, Lu Yuan 0001, Ying Fu 0001 |
IEEE Trans. Multim. | 5 |
| 2023 | Spatial-Spectral Transformer for Hyperspectral Image DenoisingabstractHyperspectral image (HSI) denoising is a crucial preprocessing procedure for the subsequent HSI applications. Unfortunately, though witnessing the development of deep learning in HSI denoising area, existing convolution-based methods face the trade-off between computational efficiency and capability to model non-local characteristics of HSI. In this paper, we propose a Spatial-Spectral Transformer (SST) to alleviate this problem. To fully explore intrinsic similarity characteristics in both spatial dimension and spectral dimension, we conduct non-local spatial self-attention and global spectral self-attention with Transformer architecture. The window-based spatial self-attention focuses on the spatial similarity beyond the neighboring region. While, the spectral self-attention exploits the long-range dependencies between highly correlative bands. Experimental results show that our proposed method outperforms the state-of-the-art HSI denoising methods in quantitative quality and visual results. The code is released at https://github.com/MyuLi/SST. Miaoyu Li, Ying Fu 0001, Yulun Zhang 0001 |
AAAI | 2 |
| 2023 | Spectral Enhanced Rectangle Transformer for Hyperspectral Image DenoisingabstractDenoising is a crucial step for hyperspectral image (HSI) applications. Though witnessing the great power of deep learning, existing HSI denoising methods suffer from limitations in capturing the nonlocal self-similarity. Trans-formers have shown potential in capturing longrange de-pendencies, but few attempts have been made with specifically designed Transformer to model the spatial and spec-tral correlation in HSIs. In this paper, we address these issues by proposing a spectral enhanced rectangle Trans-former, driving it to explore the nonlocal spatial similarity and global spectral low-rank property of HSIs. For the former, we exploit the rectangle self-attention horizontally and vertically to capture the nonlocal similarity in the spatial domain. For the latter, we design a spectral enhancement module that is capable of extracting global underlying low-rank property of spatial-spectral cubes to suppress noise, while enabling the interactions among non-overlapping spatial rectangles. Extensive experiments have been conducted on both synthetic noisy HSIs and real noisy HSIs, showing the effectiveness of our proposed method in terms of both objective metric and subjective visual quality. The code is available at https://github.com/MyuLi/SERT. Miaoyu Li, Ji Liu 0003, Ying Fu 0001, Yulun Zhang 0001, Dejing Dou |
CVPR | 3 |
| 2023 | Dynamic Aggregated Network for Gait RecognitionabstractGait recognition is beneficial for a variety of applications, including video surveillance, crime scene investigation, and social security, to mention a few. However, gait recognition often suffers from multiple exterior factors in real scenes, such as carrying conditions, wearing overcoats, and diverse viewing angles. Recently, various deep learning-based gait recognition methods have achieved promising results, but they tend to extract one of the salient features using fixed-weighted convolutional networks, do not well consider the relationship within gait features in key regions, and ignore the aggregation of complete motion patterns. In this paper, we propose a new perspective that actual gait features include global motion patterns in multiple key regions, and each global motion pattern is composed of a series of local motion patterns. To this end, we propose a Dynamic Aggregation Network (DANet) to learn more discriminative gait features. Specifically, we create a dynamic attention mechanism between the features of neighboring pixels that not only adaptively focuses on key regions but also generates more expressive local motion patterns. In addition, we develop a selfattention mechanism to select representative local motion patterns and further learn robust global motion patterns. Extensive experiments on three popular public gait datasets, i.e., CASIA-B, OUMVLP, and Gait3D, demonstrate that the proposed method can provide substantial improvements over the current state-of-the-art methods.1 Ying Fu 0001, Dezhi Zheng, Chunshui Cao, Xuecai Hu, Yongzhen Huang |
CVPR | 2 |
| 2023 | LG-BPN: Local and Global Blind-Patch Network for Self-Supervised Real-World DenoisingabstractDespite the significant results on synthetic noise under simplified assumptions, most self-supervised denoising methods fail under real noise due to the strong spatial noise correlation, including the advanced self-supervised blindspot networks (BSNs). For recent methods targeting real-world denoising, they either suffer from ignoring this spatial correlation, or are limited by the destruction of fine textures for under-considering the correlation. In this paper, we present a novel method called LG-BPN for self-supervised real-world denoising, which takes the spatial correlation statistic into our network design for local detail restoration, and also brings the long-range dependencies modeling ability to previously CNN-based BSN methods. First, based on the correlation statistic, we propose a densely-sampled patch-masked convolution module. By taking more neighbor pixels with low noise correlation into account, we enable a denser local receptive field, preserving more useful information for enhanced fine structure recovery. Second, we propose a dilated Transformer block to allow distant context exploitation in BSN. This global perception addresses the intrinsic deficiency of BSN, whose receptive field is constrained by the blind spot requirement, which can not be fully resolved by the previous CNN-based BSNs. These two designs enable LG-BPN to fully exploit both the detailed structure and the global interaction in a blind manner. Extensive results on real-world datasets demonstrate the superior performance of our method. https://github.com/Wang-XIaoDingdd/LGBPN Ying Fu 0001, Ji Liu 0003, Yulun Zhang 0001 |
CVPR | 2 |
| 2023 | Hybrid Spectral Denoising Transformer with Guided AttentionabstractIn this paper, we present a Hybrid Spectral Denoising Transformer (HSDT) for hyperspectral image denoising. Challenges in adapting transformer for HSI arise from the capabilities to tackle existing limitations of CNN-based methods in capturing the global and local spatial-spectral correlations while maintaining efficiency and flexibility. To address these issues, we introduce a hybrid approach that combines the advantages of both models with a Spatial-Spectral Separable Convolution (S3Conv), Guided Spectral Self-Attention (GSSA), and Self-Modulated Feed-Forward Network (SM-FFN). Our S3Conv works as a lightweight alternative to 3D convolution, which extracts more spatial-spectral correlated features while keeping the flexibility to tackle HSIs with an arbitrary number of bands. These features are then adaptively processed by GSSA which performs 3D self-attention across the spectral bands, guided by a set of learnable queries that encode the spectral signatures. This not only enriches our model with powerful capabilities for identifying global spectral correlations but also maintains linear complexity. Moreover, our SM-FFN proposes the self-modulation that intensifies the activations of more informative regions, which further strengthens the aggregated features. Extensive experiments are conducted on various datasets under both simulated and real-world noise, and it shows that our HSDT significantly outperforms the existing state-of-the-art methods while maintaining low computational overhead. Code is at https://github.com/Zeqiang-Lai/HSDT. Zeqiang Lai, Chenggang Yan 0001, Ying Fu 0001 |
ICCV | 3 |
| 2023 | Pixel Adaptive Deep Unfolding Transformer for Hyperspectral Image ReconstructionabstractHyperspectral Image (HSI) reconstruction has made gratifying progress with the deep unfolding framework by formulating the problem into a data module and a prior module. Nevertheless, existing methods still face the problem of insufficient matching with HSI data. The issues lie in three aspects: 1) fixed gradient descent step in the data module while the degradation of HSI is agnostic in the pixel-level. 2) inadequate prior module for 3D HSI cube. 3) stage interaction ignoring the differences in features at different stages. To address these issues, in this work, we propose a Pixel Adaptive Deep Unfolding Transformer (PADUT) for HSI reconstruction. In the data module, a pixel adaptive descent step is employed to focus on pixel-level agnostic degradation. In the prior module, we introduce the Non-local Spectral Transformer (NST) to emphasize the 3D characteristics of HSI for recovering. Moreover, inspired by the diverse expression of features in different stages and depths, the stage interaction is improved by the Fast Fourier Transform (FFT). Experimental results on both simulated and real scenes exhibit the superior performance of our method compared to state-of-the-art HSI reconstruction methods. The code is released at: https://github.com/MyuLi/PADUT Miaoyu Li, Ying Fu 0001, Ji Liu 0003, Yulun Zhang 0001 |
ICCV | 2 |
| 2023 | MPI-Flow: Learning Realistic Optical Flow with Multiplane ImagesabstractThe accuracy of learning-based optical flow estimation models heavily relies on the realism of the training datasets. Current approaches for generating such datasets either employ synthetic data or generate images with limited realism. However, the domain gap of these data with real-world scenes constrains the generalization of the trained model to real-world applications. To address this issue, we investigate generating realistic optical flow datasets from real-world images. Firstly, to generate highly realistic new images, we construct a layered depth representation, known as multiplane images (MPI), from single-view images. This allows us to generate novel view images that are highly realistic. To generate optical flow maps that correspond accurately to the new image, we calculate the optical flows of each plane using the camera matrix and plane depths. We then project these layered optical flows into the output optical flow map with volume rendering. Secondly, to ensure the realism of motion, we present an independent object motion module that can separate the camera and dynamic object motion in MPI. This module addresses the deficiency in MPI-based single-view methods, where optical flow is generated only by camera motion and does not account for any object movement. We additionally devise a depth-aware inpainting module to merge new images with dynamic objects and address unnatural motion occlusions. We show the superior performance of our method through extensive experiments on real-world datasets. Moreover, our approach achieves state-of-the-art performance in both unsupervised and supervised training of learning-based models. The code will be made publicly available at: https://github.com/Sharpiless/MPI-Flow. Yingping Liang, Debing Zhang, Ying Fu 0001 |
ICCV | 4 |
| 2023 | Fine-grained Unsupervised Domain Adaptation for Gait RecognitionabstractGait recognition has emerged as a promising technique for the long-range retrieval of pedestrians, providing numerous advantages such as accurate identification in challenging conditions and non-intrusiveness, making it highly desirable for improving public safety and security. However, the high cost of labeling datasets, which is a prerequisite for most existing fully supervised approaches, poses a significant obstacle to the development of gait recognition. Recently, some unsupervised methods for gait recognition have shown promising results. However, these methods mainly rely on a fine-tuning approach that does not sufficiently consider the relationship between source and target domains, leading to the catastrophic forgetting of source domain knowledge. This paper presents a novel perspective that adjacent-view sequences exhibit overlapping views, which can be leveraged by the network to gradually attain cross-view and cross-dressing capabilities without pre-training on the labeled source domain. Specifically, we propose a fine-grained Unsupervised Domain Adaptation (UDA) framework that iteratively alternates between two stages. The initial stage involves offline clustering, which transfers knowledge from the labeled source domain to the unlabeled target domain and adaptively generates pseudo-labels according to the expressiveness of each part. Subsequently, the second stage encompasses online training, which further achieves cross-dressing capabilities by continuously learning to distinguish numerous features of source and target domains. The effectiveness of the proposed method is demonstrated through extensive experiments conducted on widely-used public gait datasets. Ying Fu 0001, Dezhi Zheng, Yunjie Peng, Chunshui Cao, Yongzhen Huang |
ICCV | 2 |
| 2023 | Learning Rain Location Prior for Nighttime DerainingabstractRain can significantly degrade image quality and visibility, making deraining a critical area of research in computer vision. Despite recent progress in learning-based deraining methods, there is a lack of focus on nighttime deraining due to the unique challenges posed by non-uniform local illuminations from artificial light sources. Rain streaks in these scenes have diverse appearances that are tightly related to their relative positions to light sources, making it difficult for existing deraining methods to effectively handle them. In this paper, we highlight the importance of rain streak location information in nighttime deraining. Specifically, we propose a Rain Location Prior (RLP) that is learned implicitly from rainy images using a recurrent residual model. This learned prior contains location information of rain streaks and, when injected into deraining models, can significantly improve their performance. To further improve the effectiveness of the learned prior, we also propose a Rain Prior Injection Module (RPIM) to modulate the prior before injection, increasing the importance of features within rain streak areas. Experimental results demonstrate that our approach outperforms existing state-of-the-art methods by about 1dB and effectively improves the performance of deraining models. We also evaluate our method on real night rainy images to show the capability to handle real scenes with fully synthetic data for training. Our method represents a significant step forward in the area of nighttime deraining and highlights the importance of location information in this challenging problem. The code is publicly available at https://github.com/zkawfanx/RLP. Fan Zhang 0123, Shaodi You, Yu Li 0003, Ying Fu 0001 |
ICCV | 4 |
| 2023 | RawHDR: High Dynamic Range Image Reconstruction from a Single Raw ImageabstractHigh dynamic range (HDR) images capture much more intensity levels than standard ones. Current methods predominantly generate HDR images from 8-bit low dynamic range (LDR) sRGB images that have been degraded by the camera processing pipeline. However, it becomes a formidable task to retrieve extremely high dynamic range scenes from such limited bit-depth data. Unlike existing methods, the core idea of this work is to incorporate more informative Raw sensor data to generate HDR images, aiming to recover scene information in hard regions (the darkest and brightest areas of an HDR scene). To this end, we propose a model tailor-made for Raw images, harnessing the unique features of Raw data to facilitate the Raw-to-HDR mapping. Specifically, we learn exposure masks to separate the hard and easy regions of a high dynamic scene. Then, we introduce two important guidances, dual intensity guidance, which guides less informative channels with more informative ones, and global spatial guidance, which extrapolates scene specifics over an extended spatial domain. To verify our Raw-to-HDR approach, we collect a large Raw/HDR paired dataset for both training and testing. Our empirical evaluations validate the superiority of the proposed Raw-to-HDR reconstruction model, as well as our newly captured dataset in the experiments. Yunhao Zou, Chenggang Yan 0001, Ying Fu 0001 |
ICCV | 3 |
| 2023 | Iterative Denoiser and Noise Estimator for Self-Supervised Image DenoisingabstractWith the emergence of powerful deep learning tools, more and more effective deep denoisers have advanced the field of image denoising. However, the huge progress made by these learning-based methods severely relies on large-scale and high-quality noisy/clean training pairs, which limits the practicality in real-world scenarios. To overcome this, researchers have been exploring self-supervised approaches that can denoise without paired data. However, the unavailable noise prior and inefficient feature extraction take these methods away from high practicality and precision. In this paper, we propose a Denoise-Corrupt-Denoise pipeline (DCD-Net) for self-supervised image denoising. Specifically, we design an iterative training strategy, which iteratively optimizes the denoiser and noise estimator, and gradually approaches high denoising performances using only single noisy images without any noise prior. The proposed self-supervised image denoising framework provides very competitive results compared with state-of-the-art methods on widely used synthetic and real-world image denoising benchmarks. Yunhao Zou, Chenggang Yan 0001, Ying Fu 0001 |
ICCV | 3 |
| 2023 | Recurrent Self-Supervised Video Denoising with Denser Receptive FieldabstractSelf-supervised video denoising has seen decent progress through the use of blind spot networks. However, under their blind spot constraints, previous self-supervised video denoising methods suffer from significant information loss and texture destruction in either the whole reference frame or neighbor frames, due to their inadequate consideration of the receptive field. Moreover, the limited number of available neighbor frames in previous methods leads to the discarding of distant temporal information. Nonetheless, simply adopting existing recurrent frameworks does not work, since they easily break the constraints on the receptive field imposed by self-supervision. In this paper, we propose RDRF for selfsupervised video denoising, which not only fully exploits both the reference and neighbor frames with a denser receptive field, but also better leverages the temporal information from both local and distant neighbor features. First, towards a comprehensive utilization of information from both reference and neighbor frames, RDRF realizes a denser receptive field by taking more neighbor pixels along the spatial and temporal dimensions. Second, it features a self-supervised recurrent video denoising framework, which concurrently integrates distant and near-neighbor temporal features. This enables long-term bidirectional information aggregation, while mitigating error accumulation in the plain recurrent framework. Our method exhibits superior performance on both synthetic and real video denoising datasets. Codes will be available at https://github.com/Wang-XIaoDingdd/RDRF. Yulun Zhang 0001, Debing Zhang, Ying Fu 0001 |
ACM Multimedia | 4 |
| 2023 | MGT: Modality-Guided Transformer for Infrared and Visible Image Fusion
Taoying Zhang, Hesong Li, Qiankun Liu 0001, Xiaoyong Wang, Ying Fu 0001 |
PRCV (1) | 5 |
| 2023 | Instance Segmentation in the Dark
Ying Fu 0001, Kaixuan Wei, Dezhi Zheng, Felix Heide |
Int. J. Comput. Vis. | 2 |
| 2023 | Deep external and internal learning for noisy compressive sensing
Tao Zhang 0042, Ying Fu 0001, Debing Zhang, Chun Hu |
Neurocomputing | 2 |
| 2023 | Image manipulation detection by multiple tampering traces and edge artifact enhancement
Xun Lin, Shuai Wang 0049, Jiahao Deng, Ying Fu 0001, Xiao Bai 0001, Xinlei Chen, Xiaolei Qu, Wenzhong Tang |
Pattern Recognit. | 4 |
| 2023 | Level-Aware Consistent Multilevel Map Translation From Satellite ImageryabstractWith the rapid development of remote sensing technology, the quality of satellite imagery (SI) is getting higher, which contains rich cartographic information that can be translated into maps. However, existing methods either only focus on generating single-level map or do not fully consider the challenges of multilevel translation from satellite imageries, i.e., the large domain gap, level-dependent content differences, and main content consistency. In this article, we propose a novel level-aware fusion network for the SI-based multilevel map generation (MLMG) task. It aims to tackle these three challenges. To deal with the large domain gap, we propose to generate maps in a coarse-to-fine way. To well-handle the level-dependent content differences, we design a level classifier to explore different levels of the map. Besides, we use a map element extractor to extract the major geographic element features from satellite imageries, which is helpful to keep the main content consistency. Next, we design a multilevel fusion generator to generate a consistent multilevel map from the multilevel preliminary map, which further ensures the main content consistency. In addition, we collect a high-quality multilevel dataset for SI-based MLMG. Experimental results show that the proposed method can provide substantial improvements over the state-of-the-art alternatives in terms of both objective metric and visual quality. Ying Fu 0001, Tao Song 0004 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2023 | Blind Super-Resolution of Single Remotely Sensed Hyperspectral ImageabstractHyperspectral image (HSI) super-resolution has recently advanced with significant progress by utilizing the powerful representation capabilities of deep neural networks. These approaches, however, inevitably rely on a sizable amount of training data which can be difficult to acquire for remotely sensed HSIs. In many cases, these methods are designed and tailored for only one or a few specific super-resolution scenarios, making them inflexible for handling images with different unknown degradations. In this paper, we introduce a two-step framework for blind remotely sensed HSI super-resolution, where the degradation is unknown. Specifically, in the first step, we propose to leverage the abundant remotely sensed color images to address the data insufficiency for remotely sensed HSI super-resolution. It is achieved by exploring the spatial knowledge from remotely sensed color images with a super-resolution network for a predefined degradation, which is then transferred to HSIs via band-by-band super-resolution. Direct use of the results from the transferred super-resolution network is suboptimal as it neglects the spectral correlations of different bands and the gap between predefined degradation and the real one. To make further refinements, we present an unsupervised scheme that simultaneously refines the super-resolved HSI and the unknown degradation by a non-negative matrix factorization network and a learnable degradation prior. To validate the effectiveness of our method, we conducted extensive experiments on a variety of remotely sensed HSI datasets. The results demonstrate that our method could generalize on various unknown degradations with superior performance against the state-of-the-art methods. Zhiyuan Liang, Shuai Wang 0049, Tao Zhang 0042, Ying Fu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Low-Light Raw Video Denoising With a High-Quality Realistic Motion DatasetabstractRecently, supervised deep-learning methods have shown their effectiveness on raw video denoising in low-light. However, existing training datasets have specific drawbacks, e.g., inaccurate noise modeling in synthetic datasets, simple motion created by hand or fixed motion, and limited-quality ground truth caused by the beam splitter in real captured datasets. These defects significantly decline the performance of network when tackling real low-light video sequences, where noise distribution and motion patterns are extremely complex. In this paper, we collect a raw video denoising dataset in low-light with complex motion and high-quality ground truth, overcoming the drawbacks of previous datasets. Specifically, we capture 210 paired videos, each containing short/long exposure pairs of real video frames with dynamic objects and diverse scenes displayed on a high-end monitor. Besides, since spatial self-similarity has been extensively utilized in image tasks, harnessing this property for network design is more crucial for video denoising as temporal redundancy. To effectively exploit the intrinsic temporal-spatial self-similarity of complex motion in real videos, we propose a new Transformer-based network, which can effectively combine the locality of convolution with the long-range modeling ability of 3D temporal-spatial self-attention. Extensive experiments verify the value of our dataset and the effectiveness of our method on various metrics. Ying Fu 0001, Tao Zhang 0042, Jun Zhang 0007 |
IEEE Trans. Multim. | 1 |
| 2023 | ∇-Prox: Differentiable Proximal Algorithm Modeling for Large-Scale OptimizationabstractTasks across diverse application domains can be posed as large-scale optimization problems, these include graphics, vision, machine learning, imaging, health, scheduling, planning, and energy system forecasting. Independently of the application domain, proximal algorithms have emerged as a formal optimization method that successfully solves a wide array of existing problems, often exploiting problem-specific structures in the optimization. Although model-based formal optimization provides a principled approach to problem modeling with convergence guarantees, at first glance, this seems to be at odds with black-box deep learning methods. A recent line of work shows that, when combined with learning-based ingredients, model-based optimization methods are effective, interpretable, and allow for generalization to a wide spectrum of applications with little or no extra training data. However, experimenting with such hybrid approaches for different tasks by hand requires domain expertise in both proximal optimization and deep learning, which is often error-prone and time-consuming. Moreover, naively unrolling these iterative methods produces lengthy compute graphs, which when differentiated via autograd techniques results in exploding memory consumption, making batch-based training challenging. In this work, we introduce ∇-Prox, a domain-specific modeling language and compiler for large-scale optimization problems using differentiable proximal algorithms. ∇-Prox allows users to specify optimization objective functions of unknowns concisely at a high level, and intelligently compiles the problem into compute and memory-efficient differentiable solvers. One of the core features of ∇-Prox is its full differentiability, which supports hybrid model- and learning-based solvers integrating proximal optimization with neural network pipelines. Example applications of this methodology include learning-based priors and/or sample-dependent inner-loop optimization schedulers, learned with deep equilibrium learning or deep reinforcement learning. With a few lines of code, we show ∇-Prox can generate performant solvers for a range of image optimization problems, including end-to-end computational optics, image deraining, and compressive magnetic resonance imaging. We also demonstrate ∇-Prox can be used in a completely orthogonal application domain of energy system planning, an essential task in the energy crisis and the clean energy transition, where it outperforms state-of-the-art CVXPY and commercial Gurobi solvers. Zeqiang Lai, Kaixuan Wei, Ying Fu 0001, Philipp Härtel, Felix Heide |
ACM Trans. Graph. | 3 |
| 2022 | Deep Spatial Adaptive Network for Real Image DemosaicingabstractDemosaicing is the crucial step in the image processing pipeline and is a highly ill-posed inverse problem. Recently, various deep learning based demosaicing methods have achieved promising performance, but they often design the same nonlinear mapping function for different spatial location and are not well consider the difference of mosaic pattern for each color. In this paper, we propose a deep spatial adaptive network (SANet) for real image demosaicing, which can adaptively learn the nonlinear mapping function for different locations. The weights of spatial adaptive convolution layer are generated by the pattern information in the receptive filed. Besides, we collect a paired real demosaicing dataset to train and evaluate the deep network, which can make the learned demosaicing network more practical in the real world. The experimental results show that our SANet outperforms the state-of-the-art methods under both comprehensive quantitative metrics and perceptive quality in both noiseless and noisy cases. Tao Zhang 0042, Ying Fu 0001, Cheng Li 0009 |
AAAI | 2 |
| 2022 | Mutual Contrastive Low-rank Learning to Disentangle Whole Slide Image Representations for Glioma Grading
Lipei Zhang, Yiran Wei 0002, Ying Fu 0001, Stephen J. Price, Carola-Bibiane Schönlieb, Chao Li 0031 |
BMVC | 3 |
| 2022 | Estimating Fine-Grained Noise Model via Contrastive LearningabstractImage denoising has achieved unprecedented progress as great efforts have been made to exploit effective deep denoisers. To improve the denoising performance in real-world, two typical solutions are used in recent trends: devising better noise models for the synthesis of more realistic training data, and estimating noise level function to guide non-blind denoisers. In this work, we combine both noise modeling and estimation, and propose an innovative noise model estimation and noise synthesis pipeline for realistic noisy image generation. Specifically, our model learns a noise estimation model with fine-grained statistical noise model in a contrastive manner. Then, we use the estimated noise parameters to model camera-specific noise distribution, and synthesize realistic noisy training data. The most striking thing for our work is that by calibrating noise models of several sensors, our model can be extended to predict other cameras. In other words, we can estimate camera-specific noise models for unknown sensors with only testing images, without laborious calibration frames or paired noisy/clean data. The proposed pipeline endows deep denoisers with competitive performances with state-of-the-art real noise modeling methods. Yunhao Zou, Ying Fu 0001 |
CVPR | 2 |
| 2022 | Event-guided Video Clip Generation from Blurry ImagesabstractDynamic and active pixel vision sensors (DAVIS) can simultaneously produce streams of asynchronous events captured by the dynamic vision sensor (DVS) and intensity frames from the active pixel sensor (APS). Event sequences show high temporal resolution and high dynamic range, while intensity images easily suffer from motion blur due to the low frame rate of APS. In this paper, we present an end-to-end convolutional neural network based method under the local and global constraints of events to restore clear, sharp intensity frames through collaborative learning from a blurry image and its associated event streams. Specifically, we first learn a function of the relationship between the sharp intensity frame and the corresponding blurry image with its event data. Then we propose a generation module to realize it with a supervision module to constrain the restoration in the motion process. We also capture the first realistic dataset with paired blurry frame/events and sharp frames by synchronizing a DAVIS camera and a high-speed camera. Experimental results show that our method can reconstruct high-quality sharp video clips, and outperform the state-of-the-art on both simulated and real-world data. Tsuyoshi Takatani, Zhongyuan Wang 0001, Ying Fu 0001, Yinqiang Zheng |
ACM Multimedia | 4 |
| 2022 | Plug-and-Play algorithm for under-sampling Fourier single-pixel imaging
Ying Fu 0001, Jun Zhang 0007 |
Sci. China Inf. Sci. | 2 |
| 2022 | Guided Hyperspectral Image Denoising with Realistic Data
Tao Zhang 0042, Ying Fu 0001, Jun Zhang 0007 |
Int. J. Comput. Vis. | 2 |
| 2022 | Hybrid supervised instance segmentation by learning label noise suppression
Ying Fu 0001, Shaodi You, Hongzhe Liu 0001 |
Neurocomputing | 2 |
| 2022 | Deep plug-and-play prior for hyperspectral image restoration
Zeqiang Lai, Kaixuan Wei, Ying Fu 0001 |
Neurocomputing | 3 |
| 2022 | Dynamic proximal unrolling network for compressive imaging
Yixiao Yang, Ran Tao 0003, Kaixuan Wei, Ying Fu 0001 |
Neurocomputing | 4 |
| 2022 | TFPnP: Tuning-free Plug-and-Play Proximal Algorithms with Applications to Inverse Imaging ProblemsabstractPlug-and-Play (PnP) is a non-convex optimization framework that combines proximal algorithms, for example, the alternating direction method of multipliers (ADMM), with advanced denoising priors. Over the past few years, great empirical success has been obtained by PnP algorithms, especially for the ones that integrate deep learning-based denoisers. However, a key problem of PnP approaches is the need for manual parameter tweaking which is essential to obtain high-quality results across the high discrepancy in imaging conditions and varying scene content. In this work, we present a class of tuning-free PnP proximal algorithms that can determine parameters such as denoising strength, termination time, and other optimization-specific parameters automatically. A core part of our approach is a policy network for automated parameter search which can be effectively learned via a mixture of model-free and model-based deep reinforcement learning strategies. We demonstrate, through rigorous numerical and visual experiments, that the learned policy can customize parameters to different settings, and is often more efficient and effective than existing handcrafted criteria. Moreover, we discuss several practical considerations of PnP denoisers, which together with our learned policy yield state-of-the-art results. This advanced performance is prevalent on both linear and nonlinear exemplar inverse imaging problems, and in particular shows promising results on compressed sensing MRI, sparse-view CT, single-photon imaging, and phase retrieval. Kaixuan Wei, Angelica I. Avilés-Rivero, Jingwei Liang, Ying Fu 0001, Hua Huang 0001, Carola-Bibiane Schönlieb |
J. Mach. Learn. Res. | 4 |
| 2022 | LE-GAN: Unsupervised low-light image enhancement network using attention module and identity invariant loss
Ying Fu 0001, Shaodi You |
Knowl. Based Syst. | 1 |
| 2022 | A large-scale hyperspectral dataset for flower classification
Yongrong Zheng, Tao Zhang 0042, Ying Fu 0001 |
Knowl. Based Syst. | 3 |
| 2022 | Coded Hyperspectral Image Reconstruction Using Deep External and Internal LearningabstractTo solve the low spatial and/or temporal resolution problem which the conventional hyperspectral cameras often suffer from, coded hyperspectral imaging systems have attracted more attention recently. Recovering a hyperspectral image (HSI) from its corresponding coded image is an ill-posed inverse problem, and learning accurate prior of HSI is essential to solve this inverse problem. In this paper, we present an effective convolutional neural network (CNN) based method for coded HSI reconstruction, which learns the deep prior from the external dataset as well as the internal information of input coded image with spatial-spectral constraint. Specifically, we first develop a CNN-based channel attention reconstruction network to effectively exploit the spatial-spectral correlation of the HSI. Then, the reconstruction network is learned by leveraging an arbitrary external hyperspectral dataset to exploit the general spatial-spectral correlation under adversarial loss. Finally, we customize the network by internal learning with spatial-spectral constraint and total variation regularization for each coded image, which can make use of the internal imaging model to learn specific prior for current desirable image and effectively avoids overfitting. Experimental results using both synthetic data and real images show that our method outperforms the state-of-the-art methods on several popular coded hyperspectral imaging systems under both comprehensive quantitative metrics and perceptive quality. Ying Fu 0001, Tao Zhang 0042, Lizhi Wang 0001, Hua Huang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Joint Camera Spectral Response Selection and Hyperspectral Image RecoveryabstractHyperspectral image (HSI) recovery from a single RGB image has attracted much attention, whose performance has recently been shown to be sensitive to the camera spectral response (CSR). In this paper, we present an efficient convolutional neural network (CNN) based method, which can jointly select the optimal CSR from a candidate dataset and learn a mapping to recover HSI from a single RGB image captured with this algorithmically selected camera under multi-chip or single-chip setups. Given a specific CSR, we first present a HSI recovery network, which accounts for the underlying characteristics of the HSI, including spectral nonlinear mapping and spatial similarity. Later, we append a CSR selection layer onto the recovery network, and the optimal CSR under both multi-chip and single-chip setups can thus be automatically determined from the network weights under the nonnegative sparse constraint. Experimental results on three hyperspectral datasets and two camera spectral response datasets demonstrate that our HSI recovery network outperforms state-of-the-art methods in terms of both quantitative metrics and perceptive quality, and the selection layer always returns a CSR consistent to the best one determined by exhaustive search. Finally, we show that our method can also perform well in the real capture system, and collect a hyperspectral flower dataset to evaluate the effect from HSI recovery on classification problem. Ying Fu 0001, Tao Zhang 0042, Yinqiang Zheng, Debing Zhang, Hua Huang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Physics-Based Noise Modeling for Extreme Low-Light PhotographyabstractEnhancing the visibility in extreme low-light environments is a challenging task. Under nearly lightless condition, existing image denoising methods could easily break down due to significantly low SNR. In this paper, we systematically study the noise statistics in the imaging pipeline of CMOS photosensors, and formulate a comprehensive noise model that can accurately characterize the real noise structures. Our novel model considers the noise sources caused by digital camera electronics which are largely overlooked by existing methods yet have significant influence on raw measurement in the dark. It provides a way to decouple the intricate noise structure into different statistical distributions with physical interpretations. Moreover, our noise model can be used to synthesize realistic training data for learning-based low-light denoising algorithms. In this regard, although promising results have been shown recently with deep convolutional neural networks, the success heavily depends on abundant noisy-clean image pairs for training, which are tremendously difficult to obtain in practice. Generalizing their trained models to images from new devices is also problematic. Extensive experiments on multiple low-light denoising datasets - including a newly collected one in this work covering various devices - show that a deep neural network trained with our proposed noise formation model can reach surprisingly-high accuracy. The results are on par with or sometimes even outperform training with paired real data, opening a new door to real-world extreme low-light photography. Kaixuan Wei, Ying Fu 0001, Yinqiang Zheng, Jiaolong Yang |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | Translation of Aerial Image Into Digital Map via Discriminative Segmentation and Creative GenerationabstractAutomatic translation of aerial images into digital maps is an important and challenging task which is widely used in practical applications. Most of the existing works view it either as a creative image-to-image translation problem or a discriminative semantic segmentation problem. However, we notice that human annotators need to extract and understand the information in aerial images first and then translate them to online maps in a creative way, which helps them draw accurate and visually appealing online maps. In this article, we propose an end-to-end online map generation method that combines a discriminative module with a creative module based on this observation to mimic human behavior. Specifically, we first utilize a semantic segmentation module to obtain a rough aerial map, in which each region is labeled with its category, and then further improve its quality with a creative module. To train a robust network that generalizes well to unfamiliar regions, we also collect a large aerial image dataset for online map generation (AIDOMG). AIDOMG consists of 40 087 pairs of aerial images and corresponding online maps collected from nine regions of six continents. We conduct extensive experiments to verify the superiority of the new design that combines discrimination and creativity and experimental results show that the performance of the proposed method significantly outperforms baseline methods. Ying Fu 0001, Shuaizhe Liang, Dongdong Chen 0001, Zhanlong Chen |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2021 | Crafting Object Detection in Very Low Light
Kaixuan Wei, Ying Fu 0001 |
BMVC | 4 |
| 2021 | Learning Temporal Consistency for Low Light Video Enhancement From Single ImagesabstractSingle image low light enhancement is an important task and it has many practical applications. Most existing methods adopt a single image approach. Although their performance is satisfying on a static single image, we found, however, they suffer serious temporal instability when handling low light videos. We notice the problem is because existing data-driven methods are trained from single image pairs where no temporal information is available. Unfortunately, training from real temporally consistent data is also problematic because it is impossible to collect pixel-wisely paired low and normal light videos under controlled environments in large scale and diversities with noise of identical statistics. In this paper, we propose a novel method to enforce the temporal stability in low light video enhancement with only static images. The key idea is to learn and infer motion field (optical flow) from a single image and synthesize short range video sequences. Our strategy is general and can extend to large scale datasets directly. Based on this idea, we propose our method which can infer motion prior for single image low light video enhancement and enforce temporal consistency. Rigorous experiments and user study demonstrate the state-of-the-art performance of our proposed method. Our code and model will be publicly available at https://github.com/zkawfanx/StableLLVE. Fan Zhang 0123, Yu Li 0003, Shaodi You, Ying Fu 0001 |
CVPR | 4 |
| 2021 | Cross-MPI: Cross-Scale Stereo for Image Super-Resolution Using Multiplane ImagesabstractVarious combinations of cameras enrich computational photography, among which reference-based super-resolution (RefSR) plays a critical role in multiscale imaging systems. However, existing RefSR approaches fail to accomplish high-fidelity super-resolution under a large resolution gap, e.g., 8× upscaling, due to the lower consideration of the underlying scene structure. In this paper, we aim to solve the RefSR problem in actual multiscale camera systems inspired by multiplane image (MPI) representation. Specifically, we propose Cross-MPI, an end-to-end RefSR network composed of a novel plane-aware attention-based MPI mechanism, a multiscale guided upsampling module as well as a super-resolution (SR) synthesis and fusion module. Instead of using a direct and exhaustive matching between the cross-scale stereo, the proposed plane-aware attention mechanism fully utilizes the concealed scene structure for efficient attention-based correspondence searching. Further combined with a gentle coarse-to-fine guided upsampling strategy, the proposed Cross-MPI can achieve a robust and accurate detail transmission. Experimental results on both digitally synthesized and optical zoom cross-scale data show that the Cross-MPI framework can achieve superior performance against the existing RefSR methods and is a real fit for actual multiscale camera systems even with large-scale differences. Yuemei Zhou, Gaochang Wu, Ying Fu 0001, Kun Li 0001, Yebin Liu |
CVPR | 3 |
| 2021 | Learning To Reconstruct High Speed and High Dynamic Range Videos From EventsabstractEvent cameras are novel sensors that capture the dynamics of a scene asynchronously. Such cameras record event streams with much shorter response latency than images captured by conventional cameras, and are also highly sensitive to intensity change, which is brought by the triggering mechanism of events. On the basis of these two features, previous works attempt to reconstruct high speed and high dynamic range (HDR) videos from events. However, these works either suffer from unrealistic artifacts, or cannot provide sufficiently high frame rate. In this paper, we present a convolutional recurrent neural network which takes a sequence of neighboring events to reconstruct high speed HDR videos, and temporal consistency is well considered to facilitate the training process. In addition, we setup a prototype optical system to collect a real-world dataset with paired high speed HDR videos and event streams, which will be made publicly accessible for future researches in this field. Experimental results on both simulated and real scenes verify that our method can generate high speed HDR videos with high quality, and outperform the state-of-the-art reconstruction methods. Yunhao Zou, Yinqiang Zheng, Tsuyoshi Takatani, Ying Fu 0001 |
CVPR | 4 |
| 2021 | LocalTrans: A Multiscale Local Transformer Network for Cross-Resolution Homography EstimationabstractCross-resolution image alignment is a key problem in multiscale gigapixel photography, which requires to estimate homography matrix using images with large resolution gap. Existing deep homography methods concatenate the input images or features, neglecting the explicit formulation of correspondences between them, which leads to degraded accuracy in cross-resolution challenges. In this paper, we consider the cross-resolution homography estimation as a multimodal problem, and propose a local transformer network embedded within a multiscale structure to explicitly learn correspondences between the multimodal inputs, namely, input images with different resolutions. The proposed local transformer adopts a local attention map specifically for each position in the feature. By combining the local transformer with the multiscale structure, the network is able to capture long-short range correspondences efficiently and accurately. Experiments on both the MS-COCO dataset and the real-captured cross-resolution dataset show that the proposed network outperforms existing state-of-the-art feature-based and deep-learning-based homography estimation methods, and is able to accurately align images under 10× resolution gap. Ruizhi Shao, Gaochang Wu, Yuemei Zhou, Ying Fu 0001, Lu Fang 0001, Yebin Liu |
ICCV | 4 |
| 2021 | Hyperspectral Image Denoising with Realistic DataabstractThe hyperspectral image (HSI) denoising has been widely utilized to improve HSI qualities. Recently, learning-based HSI denoising methods have shown their effectiveness, but most of them are based on synthetic dataset and lack the generalization capability on real testing HSI. Moreover, there is still no public paired real HSI denoising dataset to learn HSI denoising network and quantitatively evaluate HSI methods. In this paper, we mainly focus on how to produce realistic dataset for learning and evaluating HSI denoising network. On the one hand, we collect a paired real HSI denoising dataset, which consists of short-exposure noisy HSIs and the corresponding long-exposure clean HSIs. On the other hand, we propose an accurate HSI noise model which matches the distribution of real data well and can be employed to synthesize realistic dataset. On the basis of the noise model, we present an approach to calibrate the noise parameters of the given hyperspectral camera. The extensive experimental results show that a network learned with only synthetic data generated by our noise model performs as well as it is learned with paired real data. Our code and data are available at: https://github.com/ColinTaoZhang/HSIDwRD. Tao Zhang 0042, Ying Fu 0001, Cheng Li 0009 |
ICCV | 2 |
| 2021 | Disentangled Face Attribute Editing via Instance-Aware Latent Space SearchabstractRecent works have shown that a rich set of semantic directions exist in the latent space of Generative Adversarial Networks (GANs), which enables various facial attribute editing applications. However, existing methods may suffer poor attribute variation disentanglement, leading to unwanted change of other attributes when altering the desired one. The semantic directions used by existing methods are at attribute level, which are difficult to model complex attribute correlations, especially in the presence of attribute distribution bias in GAN's training set. In this paper, we propose a novel framework (IALS) that performs Instance-Aware Latent-Space Search to find semantic directions for disentangled attribute editing. The instance information is injected by leveraging the supervision from a set of attribute classifiers evaluated on the input images. We further propose a Disentanglement-Transformation (DT) metric to quantify the attribute transformation and disentanglement efficacy and find the optimal control factor between attribute-level and instance-specific directions based on it. Experimental results on both GAN-generated and real-world images collectively show that our method outperforms state-of-the-art methods proposed recently by a wide margin. Code is available at https://github.com/yxuhan/IALS. Jiaolong Yang, Ying Fu 0001 |
IJCAI | 3 |
| 2021 | $\mathrm 3D^2Unet$: 3D Deformable Unet for Low-Light Video Enhancement
Yuhang Zeng, Yunhao Zou, Ying Fu 0001 |
PRCV (3) | 3 |
| 2021 | Residual scale attention network for arbitrary scale image super-resolution
Ying Fu 0001, Tao Zhang 0042, Yonggang Lin |
Neurocomputing | 1 |
| 2021 | Cross-modal dynamic convolution for multi-modal emotion recognition
Huanglu Wen, Shaodi You, Ying Fu 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2021 | Cross-modal context-gated convolution for multi-modal sentiment analysisabstractWhen inferring sentiments, using verbal clues only is problematic because of the ambiguity. Adding related vocal and visual contexts as complements for verbal clues can be helpful. To infer sentiments from multi-modal temporal sequences, we need to identify both sentiment-related clues and their cross-modal interactions. However, sentiment-related behaviors of different modalities may not occur at the same time. These behaviors and their interactions are also sparse in time, making it hard to infer the correct sentiments. Besides, unaligned sequences from sensors also have varying sampling rates , which amplify the misalignment and sparsity mentioned above. While most previous multi-modal sentiment analysis works only focus on word-aligned sequences, we propose cross-modal context-gated convolution for unaligned sequences. Cross-modal context-gated convolution captures the local cross-modal interactions, dealing with the misalignment while reducing the effect of unrelated information. Cross-modal context-gated convolution introduces the concept of cross-modal context gate, enabling itself to catch useful cross-modal interactions more effectively. Cross-modal context-gated convolution also brings more possibilities to the layer design for multi-modal sequential modeling. Experiments on multi-modal sentiment analysis datasets under both word-aligned and unaligned conditions show the validity of our approach. Huanglu Wen, Shaodi You, Ying Fu 0001 |
Pattern Recognit. Lett. | 3 |
| 2021 | 3-D Quasi-Recurrent Neural Network for Hyperspectral Image DenoisingabstractIn this article, we propose an alternating directional 3-D quasi-recurrent neural network for hyperspectral image (HSI) denoising, which can effectively embed the domain knowledge-structural spatiospectral correlation and global correlation along spectrum (GCS). Specifically, 3-D convolution is utilized to extract structural spatiospectral correlation in an HSI, while a quasi-recurrent pooling function is employed to capture the GCS. Moreover, the alternating directional structure is introduced to eliminate the causal dependence with no additional computation cost. The proposed model is capable of modeling spatiospectral dependence while preserving the flexibility toward HSIs with an arbitrary number of bands. Extensive experiments on HSI denoising demonstrate significant improvement over the state-of-the-art under various noise settings, in terms of both restoration accuracy and computation time. Our code is available at https://github.com/Vandermode/QRNN3D. Kaixuan Wei, Ying Fu 0001, Hua Huang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2020 | DNU: Deep Non-Local Unrolling for Computational Spectral ImagingabstractComputational spectral imaging has been striving to capture the spectral information of the dynamic world in the last few decades. In this paper, we propose an interpretable neural network for computational spectral imaging. First, we introduce a novel data-driven prior that can adaptively exploit both the local and non-local correlations among the spectral image. Our data-driven prior is integrated as a regularizer into the reconstruction problem. Then, we propose to unroll the reconstruction problem into an optimization-inspired deep neural network. The architecture of the network has high interpretability by explicitly characterizing the image correlation and the system imaging model. Finally, we learn the complete parameters in the network through end-to-end training, enabling robust performance with high spatial-spectral fidelity. Extensive simulation and hardware experiments validate the superior performance of our method over state-of-the-art methods. Lizhi Wang 0001, Maoqing Zhang, Ying Fu 0001, Hua Huang 0001 |
CVPR | 4 |
| 2020 | A Physics-Based Noise Formation Model for Extreme Low-Light Raw DenoisingabstractLacking rich and realistic data, learned single image denoising algorithms generalize poorly in real raw images that not resemble the data used for training. Although the problem can be alleviated by the heteroscedastic Gaussian noise model, the noise sources caused by digital camera electronics are still largely overlooked, despite their significant effect on raw measurement, especially under extremely low-light condition. To address this issue, we present a highly accurate noise formation model based on the characteristics of CMOS photosensors, thereby enabling us to synthesize realistic samples that better match the physics of image formation process. Given the proposed noise model, we additionally propose a method to calibrate the noise parameters for available modern digital cameras, which is simple and reproducible for any new device. We systematically study the generalizability of a neural network trained with existing schemes, by introducing a new low-light denoising dataset that covers many modern digital cameras from diverse brands. Extensive empirical results collectively show that by utilizing our proposed noise formation model, a network can reach the capability as if it had been trained with rich real data, which demonstrates the effectiveness of our noise formation model. Kaixuan Wei, Ying Fu 0001, Jiaolong Yang, Hua Huang 0001 |
CVPR | 2 |
| 2020 | Tuning-free Plug-and-Play Proximal Algorithm for Inverse Imaging ProblemsabstractPlug-and-play (PnP) is a non-convex framework that combines ADMM or other proximal algorithms with advanced denoiser priors. Recently, PnP has achieved great empirical success, especially with the integration of deep learning-based denoisers. However, a key problem of PnP based approaches is that they require manual parameter tweaking. It is necessary to obtain high-quality results across the high discrepancy in terms of imaging conditions and varying scene content. In this work, we present a tuning-free PnP proximal algorithm, which can automatically determine the internal parameters including the penalty parameter, the denoising strength and the terminal time. A key part of our approach is to develop a policy network for automatic search of parameters, which can be effectively learned via mixed model-free and model-based deep reinforcement learning. We demonstrate, through numerical and visual experiments, that the learned policy can customize different parameters for different states, and often more efficient and effective than existing handcrafted criteria. Moreover, we discuss the practical considerations of the plugged denoisers, which together with our learned policy yield state-of-the-art results. This is prevalent on both linear and nonlinear exemplary inverse imaging problems, and in particular, we show promising results on Compressed Sensing MRI and phase retrieval. Kaixuan Wei, Angelica I. Avilés-Rivero, Jingwei Liang, Ying Fu 0001, Carola-Bibiane Schönlieb, Hua Huang 0001 |
ICML | 4 |
| 2020 | GPS-Net: Graph-based Photometric Stereo NetworkabstractLearning-based photometric stereo methods predict the surface normal either in a per-pixel or an all-pixel manner. Per-pixel methods explore the inter-image intensity variation of each pixel but ignore features from the intra-image spatial domain. All-pixel methods explore the intra-image intensity variation of each input image but pay less attention to the inter-image lighting variation. In this paper, we present a Graph-based Photometric Stereo Network, which unifies per-pixel and all-pixel processings to explore both inter-image and intra-image information. For per-pixel operation, we propose the Unstructured Feature Extraction Layer to connect an arbitrary number of input image-light pairs into graph structures, and introduce Structure-aware Graph Convolution filters to balance the input data by appropriately weighting shadows and specular highlights. For all-pixel operation, we propose the Normal Regression Network to make efficient use of the intra-image spatial information for predicting a surface normal map with rich details. Experimental results on the real-world benchmark show that our method achieves excellent performance under both sparse and dense lighting distributions. Zhuokun Yao, Kun Li 0001, Ying Fu 0001, Haofeng Hu, Boxin Shi |
NeurIPS | 3 |
| 2020 | Simultaneous hyperspectral image super-resolution and geometric alignment with a hybrid camera system
Ying Fu 0001, Yongrong Zheng, Yinqiang Zheng, Hua Huang 0001 |
Neurocomputing | 1 |
| 2020 | Global Topology Constraint Network for Fine-Grained Vehicle RecognitionabstractFine-grained vehicle recognition is challenging due to the large intra-class variation in vehicle pose and viewpoint. Many existing methods, especially convolution neural network (CNN)-based methods, solve this problem via detecting and aligning parts individually, and do not consider the interaction between parts, which is very important for effective part detection and vehicle recognition. In this paper, we propose a global topology constraint network for fine-grained vehicle recognition, which adopts the constraint of global topology relationship to depict the interaction between parts and integrates it into CNN in an efficient way. Different CNNs for image classification truncated at intermediate layer can be used for part detection. The global topology relationship between parts is encoded into kernel of depthwise convolution layer, and can be learned from training images. The response under global topology constraint reflects the probability of topology relationship between parts existing in input image. All responses build the description for classification with translation invariance. Through training the whole network, the back-propagation of gradient information of kernel for global topology relationship will guide former layers to better detect useful parts, and thereby improve vehicle recognition. Our proposed method does not require additional annotation, such as bounding box or part annotation, and can be trained in an end-to-end way. We conduct comparison experiments on public Stanford Cars and CompCars datasets, which both show that our method achieves the state-of-the-art performance. Ye Xiang, Ying Fu 0001, Hua Huang 0001 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2019 | Hyperspectral Image Super-Resolution With Optimized RGB GuidanceabstractTo overcome the limitations of existing hyperspectral cameras on spatial/temporal resolution, fusing a low resolution hyperspectral image (HSI) with a high resolution RGB (or multispectral) image into a high resolution HSI has been prevalent. Previous methods for this fusion task usually employ hand-crafted priors to model the underlying structure of the latent high resolution HSI, and the effect of the camera spectral response (CSR) of the RGB camera on super-resolution accuracy has rarely been investigated. In this paper, we first present a simple and efficient convolutional neural network (CNN) based method for HSI super-resolution in an unsupervised way, without any prior training. Later, we append a CSR optimization layer onto the HSI super-resolution network, either to automatically select the best CSR in a given CSR dataset, or to design the optimal CSR under some physical restrictions. Experimental results show our method outperforms the state-of-the-arts, and the CSR optimization can further boost the accuracy of HSI super-resolution. Ying Fu 0001, Tao Zhang 0042, Yinqiang Zheng, Debing Zhang, Hua Huang 0001 |
CVPR | 1 |
| 2019 | Hyperspectral Image Reconstruction Using a Deep Spatial-Spectral PriorabstractRegularization is a fundamental technique to solve an ill-posed optimization problem robustly and is essential to reconstruct compressive hyperspectral images. Various hand-crafted priors have been employed as a regularizer but are often insufficient to handle the wide variety of spectra of natural hyperspectral images, resulting in poor reconstruction quality. Moreover, the prior-regularized optimization requires manual tweaking of its weight parameters to achieve a balance between the spatial and spectral fidelity of result images. In this paper, we present a novel hyperspectral image reconstruction algorithm that substitutes the traditional hand-crafted prior with a data-driven prior, based on an optimization-inspired network. Our method consists of two main parts: First, we learn a novel data-driven prior that regularizes the optimization problem with a goal to boost the spatial-spectral fidelity. Our data-driven prior learns both local coherence and dynamic characteristics of natural hyperspectral images. Second, we combine our regularizer with an optimization-inspired network to overcome the heavy computation problem in the traditional iterative optimization methods. We learn the complete parameters in the network through end-to-end training, enabling robust performance with high accuracy. Extensive simulation and hardware experiments validate the superior performance of our method over the state-of-the-art methods. Lizhi Wang 0001, Ying Fu 0001, Min H. Kim 0001, Hua Huang 0001 |
CVPR | 3 |
| 2019 | Single Image Reflection Removal Exploiting Misaligned Training Data and Network EnhancementsabstractRemoving undesirable reflections from a single image captured through a glass window is of practical importance to visual computing systems. Although state-of-the-art methods can obtain decent results in certain situations, performance declines significantly when tackling more general real-world cases. These failures stem from the intrinsic difficulty of single image reflection removal -- the fundamental ill-posedness of the problem, and the insufficiency of densely-labeled training data needed for resolving this ambiguity within learning-based neural network pipelines. In this paper, we address these issues by exploiting targeted network enhancements and the novel use of misaligned data. For the former, we augment a baseline network architecture by embedding context encoding modules that are capable of leveraging high-level contextual clues to reduce indeterminacy within areas containing strong reflections. For the latter, we introduce an alignment-invariant loss function that facilitates exploiting misaligned real-world training data that is much easier to collect. Experimental results collectively show that our method outperforms the state-of-the-art with aligned data, and that significant improvements are possible when using additional misaligned data. Kaixuan Wei, Jiaolong Yang, Ying Fu 0001, David P. Wipf, Hua Huang 0001 |
CVPR | 3 |
| 2019 | Incremental Learning Using Conditional Adversarial NetworksabstractIncremental learning using Deep Neural Networks (DNNs) suffers from catastrophic forgetting. Existing methods mitigate it by either storing old image examples or only updating a few fully connected layers of DNNs, which, however, requires large memory footprints or hurts the plasticity of models. In this paper, we propose a new incremental learning strategy based on conditional adversarial networks. Our new strategy allows us to use memory-efficient statistical information to store old knowledge, and fine-tune both convolutional layers and fully connected layers to consolidate new knowledge. Specifically, we propose a model consisting of three parts, i.e., a base sub-net, a generator, and a discriminator. The base sub-net works as a feature extractor which can be pre-trained on large scale datasets and shared across multiple image recognition tasks. The generator conditioned on labeled embeddings aims to construct pseudo-examples with the same distribution as the old data. The discriminator combines real-examples from new data and pseudo-examples generated from the old data distribution to learn representation for both old and new classes. Through adversarial training of the discriminator and generator, we accomplish the multiple continuous incremental learning. Comparison with the state-of-the-arts on public CIFAR-100 and CUB-200 datasets shows that our method achieves the best accuracies on both old and new classes while requiring relatively less memory storage. Ye Xiang, Ying Fu 0001, Pan Ji, Hua Huang 0001 |
ICCV | 2 |
| 2019 | Hyperspectral Image Reconstruction Using Deep External and Internal LearningabstractTo solve the low spatial and/or temporal resolution problem which the conventional hypelrspectral cameras often suffer from, coded snapshot hyperspectral imaging systems have attracted more attention recently. Recovering a hyperspectral image (HSI) from its corresponding coded image is an ill-posed inverse problem, and learning accurate prior of HSI is essential to solve this inverse problem. In this paper, we present an effective convolutional neural network (CNN) based method for coded HSI reconstruction, which learns the deep prior from the external dataset as well as the internal information of input coded image with spatial-spectral constraint. Our method can effectively exploit spatial-spectral correlation and sufficiently represent the variety nature of HSIs. Experimental results show our method outperforms the state-of-the-art methods under both comprehensive quantitative metrics and perceptive quality. Tao Zhang 0042, Ying Fu 0001, Lizhi Wang 0001, Hua Huang 0001 |
ICCV | 2 |
| 2019 | Computational Hyperspectral Imaging Based on Dimension-Discriminative Low-Rank Tensor RecoveryabstractExploiting the prior information is fundamental for the image reconstruction in computational hyperspectral imaging. Existing methods usually unfold the 3D signal as a 1D vector and treat the prior information within different dimensions in an indiscriminative manner, which ignores the high-dimensionality nature of hyperspectral image (HSI) and thus results in poor quality reconstruction. In this paper, we propose to make full use of the high-dimensionality structure of the desired HSI to boost the reconstruction quality. We first build a high-order tensor by exploiting the nonlocal similarity in HSI. Then, we propose a dimension-discriminative low-rank tensor recovery (DLTR) model to characterize the structure prior adaptively in each dimension. By integrating the structure prior in DLTR with the system imaging process, we develop an optimization framework for HSI reconstruction, which is finally solved via the alternating minimization algorithm. Extensive experiments implemented with both synthetic and real data demonstrate that our method outperforms state-of-the-art methods. Lizhi Wang 0001, Ying Fu 0001, Xiaoming Zhong, Hua Huang 0001 |
ICCV | 3 |
| 2019 | An Effective Network with ConvLSTM for Low-Light Image Enhancement
Yixi Xiang, Ying Fu 0001, Lei Zhang 0021, Hua Huang 0001 |
PRCV (2) | 2 |
| 2019 | Fast HSI super resolution using linear regressionabstractHyperspectral imaging has great achievements in agriculture, astronomy, surveillance, and so on. However, the inherent low spatial resolution of hyperspectral imaging, unfortunately, limits its more widespread applications. Recently, hyperspectral image (HSI) super resolution addresses this problem by fusing a low spatial resolution HSI (LR‐HSI) with a high spatial resolution multispectral image (HR‐MSI), but most of these methods did not consider real‐time restoration of high spatial resolution HSI. In this study, the authors propose a fast HSI super‐resolution method which fills this blank. Specifically, they model the hyperspectral super resolution as a linear regression problem according to the fact that the imaging process is a linear transform and the inverse of this transform can be approximately estimated, as the spectra of a typical scene lie in a very low‐dimensional space. To further exploit the low‐dimensional nature of the spectra, they divide the HR‐MSI and LR‐HSI into several patches and learn the inverse transform patch‐by‐patch. Experiments on several public datasets show that their method approximates state‐of‐the‐art methods in accuracy, but is several orders of magnitude faster than all of them. Furthermore, they provide an efficient C language implementation of their methods, which can meet the real‐time request. Lingfei Song, Ying Fu 0001, Hua Huang 0001 |
IET Image Process. | 2 |
| 2019 | Image restoration from patch-based compressed sensing measurement
Hua Huang 0001, Guangtao Nie, Yinqiang Zheng, Ying Fu 0001 |
Neurocomputing | 4 |
| 2019 | Low-rank Bayesian tensor factorization for hyperspectral image denoising
Kaixuan Wei, Ying Fu 0001 |
Neurocomputing | 2 |
| 2019 | Global relative position space based pooling for fine-grained vehicle recognition
Ye Xiang, Ying Fu 0001, Hua Huang 0001 |
Neurocomputing | 2 |
| 2019 | Fast Parallel Implementation of Dual-Camera Compressive Hyperspectral Imaging SystemabstractCoded aperture snapshot spectral imager (CASSI) provides a potential solution to recover the 3D hyperspectral image (HSI) from a single 2D measurement. The latest proposed design of the dual-camera compressive hyperspectral imager (DCCHI) can collect more information simultaneously with the CASSI to improve the reconstruction quality. The main bottleneck now lies in the high computation complexity of the reconstruction methods, which hinders the practical application. In this paper, we propose a fast parallel implementation based on DCCHI to reach a stable and efficient HSI reconstruction. Specifically, we develop a new optimization method for the reconstruction problem, which integrates the alternative direction multiplier method with the total variation-based regularization to boost the convergence rate. Then, to improve the time efficiency, a novel parallel implementation based on GPU is proposed. The performance of the proposed method is validated on both synthetic and real data. The experimental results demonstrate that our method has a significant advantage in time efficiency, while maintaining a comparable reconstruction fidelity. Hua Huang 0001, Ying Fu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2019 | HyperReconNet: Joint Coded Aperture Optimization and Image Reconstruction for Compressive Hyperspectral ImagingabstractCoded aperture snapshot spectral imaging (CASSI) system encodes the 3D hyperspectral image (HSI) within a single 2D compressive image and then reconstructs the underlying HSI by employing an inverse optimization algorithm, which equips with the distinct advantage of snapshot but usually results in low reconstruction accuracy. To improve the accuracy, existing methods attempt to design either alternative coded apertures or advanced reconstruction methods, but cannot connect these two aspects via a unified framework, which limits the accuracy improvement. In this paper, we propose a convolution neural network (CNN) based endto- end method to boost the accuracy by jointly optimizing the coded aperture and the reconstruction method. On the one hand, based on the nature of CASSI forward model, we design a repeated pattern for the coded aperture, whose entities are learned by acting as the network weights. On the other hand, we conduct the reconstruction through simultaneously exploiting intrinsic properties within HSI - the extensive correlations across the spatial and the spectral dimensions. By leveraging the power of deep learning, the coded aperture design and the image reconstruction are connected and optimized via a unified framework. Experimental results show that our method outperforms the state-of-the-art methods under both comprehensive quantitative metrics and perceptive quality. Lizhi Wang 0001, Tao Zhang 0042, Ying Fu 0001, Hua Huang 0001 |
IEEE Trans. Image Process. | 3 |
| 2018 | Joint Camera Spectral Sensitivity Selection and Hyperspectral Image Recovery
Ying Fu 0001, Tao Zhang 0042, Yinqiang Zheng, Debing Zhang, Hua Huang 0001 |
ECCV (3) | 1 |
| 2018 | Hyperspectral image super-resolution under misaligned hybrid camera systemabstractHyperspectral imaging has been widely used for agriculture, astronomy, surveillance, and so on. However, hyperspectral imaging usually suffers from low‐spatial resolution, due to the limited photons in individual bands. Recently, more hyperspectral image super‐resolution methods have been developed by fusing the low‐resolution hyperspectral image and high‐resolution RGB image, but most of them did not consider the misalignment between two input images. In this study, the authors present an effective method to restore a high‐resolution hyperspectral image from the misaligned low‐resolution hyperspectral image and high‐resolution RGB image, which exploits spectral and spatial correlation in hyperspectral and RGB images. Specifically, they employ the spectral sparsity to restore the high‐resolution hyperspectral image on the misaligned part, and then simultaneously employ spectral and spatial structure correlation to restore the high‐resolution hyperspectral image on the aligned area, which can be fused to obtain the high‐quality hyperspectral image restoration under a misaligned hybrid camera system. Experimental results show that the proposed method outperforms the state‐of‐the‐art hyperspectral image super‐resolution methods under a misaligned hybrid camera system in terms of both objective metric and subjective visual quality. Yonggang Lin, Yongrong Zheng, Ying Fu 0001, Hua Huang 0001 |
IET Image Process. | 3 |
| 2018 | Hyperspectral Image Super-Resolution With a Mosaic RGB ImageabstractRecently, many hyperspectral (HS) image superresolution methods that merge a low spatial resolution HS image and a high spatial resolution three-channel RGB image have been proposed in spectral imaging. A largely ignored fact is that most existing commercial RGB cameras capture high resolution images by a single CCD/CMOS sensor equipped with a color filter array (CFA). In this paper, we account for the common imaging mechanism of commercial RGB cameras, and propose to use a mosaic RGB image for HS image super-resolution, which prevents demosaicing error and thus its propagation into the HS image super-resolution results. We design a proper nonlocal low-rank regularization to exploit the intrinsic properties - rich self-repeating patterns and high correlation across spectra - within HS images of natural scenes, and formulate the HS image super-resolution task into a variational optimization problem, which can be efficiently solved via the alternating direction method of multipliers (ADMM). The effectiveness of the proposed method has been evaluated on two benchmark datasets, demonstrating that the proposed method can provide substantial improvement over the current state-of-the-art HS image superresolution methods without considering the mosaicing effect. Finally, we show that our method can also perform well in the real capture system. Ying Fu 0001, Yinqiang Zheng, Hua Huang 0001, Imari Sato, Yoichi Sato 0001 |
IEEE Trans. Image Process. | 1 |
| 2017 | Camera spectral sensitivity, illumination and spectral reflectance estimation for a hybrid hyperspectral image capture systemabstractA variety of methods have been proposed to restore high resolution hyperspectral image (HSI) from a hybrid camera system, which captures high spatial resolution RGB images and low spatial resolution HSI. They focused unanimously on HSI super-resolution via fusion, yet did not explore the potential of this kind of system for camera spectral sensitivity (CSS), illumination spectrum, and high spatial resolution spectral reflectance recovery. In this paper, we present a sparse representation based method to estimate the CSS of the RGB camera under unknown illumination for the hybrid camera system. Furthermore, the illumination and high spatial resolution spectral reflectance are simultaneously recovered. Experimental results show the effectiveness of the proposed methods on camera spectral sensitivity, illumination spectrum and spectral reflectance recovery. Ying Fu 0001, Yinqiang Zheng, Hua Huang 0001 |
ICIP | 2 |
| 2017 | Adaptive Spatial-Spectral Dictionary Learning for Hyperspectral Image Restoration
Ying Fu 0001, Antony Lam, Imari Sato, Yoichi Sato 0001 |
Int. J. Comput. Vis. | 1 |
| 2016 | Direct and Global Component Separation from a Single Image Using Basis Representation
Art Subpa-Asa, Ying Fu 0001, Yinqiang Zheng, Toshiyuki Amano, Imari Sato |
ACCV (3) | 2 |
| 2016 | Exploiting Spectral-Spatial Correlation for Coded Hyperspectral Image RestorationabstractConventional scanning and multiplexing techniques for hyperspectral imaging suffer from limited temporal and/or spatial resolution. To resolve this issue, coding techniques are becoming increasingly popular in developing snapshot systems for high-resolution hyperspectral imaging. For such systems, it is a critical task to accurately restore the 3D hyperspectral image from its corresponding coded 2D image. In this paper, we propose an effective method for coded hyperspectral image restoration, which exploits extensive structure sparsity in the hyperspectral image. Specifically, we simultaneously explore spectral and spatial correlation via low-rank regularizations, and formulate the restoration problem into a variational optimization model, which can be solved via an iterative numerical algorithm. Experimental results using both synthetic data and real images show that the proposed method can significantly outperform the state-of-the-art methods on several popular coding-based hyperspectral imaging systems. Ying Fu 0001, Yinqiang Zheng, Imari Sato, Yoichi Sato 0001 |
CVPR | 1 |
| 2016 | Separating Reflective and Fluorescent Components Using High Frequency Illumination in the Spectral DomainabstractHyperspectral imaging is beneficial to many applications but most traditional methods do not consider fluorescent effects which are present in everyday items ranging from paper to even our food. Furthermore, everyday fluorescent items exhibit a mix of reflection and fluorescence so proper separation of these components is necessary for analyzing them. In recent years, effective imaging methods have been proposed but most require capturing the scene under multiple illuminants. In this paper, we demonstrate efficient separation and recovery of reflectance and fluorescence emission spectra through the use of two high frequency illuminations in the spectral domain. With the obtained fluorescence emission spectra from our high frequency illuminants, we then describe how to estimate the fluorescence absorption spectrum of a material given its emission spectrum. In addition, we provide an in depth analysis of our method and also show that filters can be used in conjunction with standard light sources to generate the required high frequency illuminants. We also test our method under ambient light and demonstrate an application of our method to synthetic relighting of real scenes. Ying Fu 0001, Antony Lam, Imari Sato, Takahiro Okabe, Yoichi Sato 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2016 | Reflectance and Fluorescence Spectral Recovery via Actively Lit RGB ImagesabstractIn recent years, fluorescence analysis of scenes has received attention in computer vision. Fluorescence can provide additional information about scenes, and has been used in applications such as camera spectral sensitivity estimation, 3D reconstruction, and color relighting. In particular, hyperspectral images of reflective-fluorescent scenes provide a rich amount of data. However, due to the complex nature of fluorescence, hyperspectral imaging methods rely on specialized equipment such as hyperspectral cameras and specialized illuminants. In this paper, we propose a more practical approach to hyperspectral imaging of reflective-fluorescent scenes using only a conventional RGB camera and varied colored illuminants. The key idea of our approach is to exploit a unique property of fluorescence: the chromaticity of fluorescent emissions are invariant under different illuminants. This allows us to robustly estimate spectral reflectance and fluorescent emission chromaticity. We then show that given the spectral reflectance and fluorescent chromaticity, the fluorescence absorption and emission spectra can also be estimated. We demonstrate in results that all scene spectra can be accurately estimated from RGB images. Finally, we show that our method can be used to accurately relight scenes under novel lighting. Ying Fu 0001, Antony Lam, Imari Sato, Takahiro Okabe, Yoichi Sato 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2015 | Adaptive Spatial-Spectral Dictionary Learning for Hyperspectral Image DenoisingabstractHyperspectral imaging is beneficial in a diverse range of applications from diagnostic medicine, to agriculture, to surveillance to name a few. However, hyperspectral images often times suffer from degradation due to the limited light, which introduces noise into the imaging process. In this paper, we propose an effective model for hyperspectral image (HSI) denoising that considers underlying characteristics of HSIs: sparsity across the spatial-spectral domain, high correlation across spectra, and non-local self-similarity over space. We first exploit high correlation across spectra and non-local self-similarity over space in the noisy HSI to learn an adaptive spatial-spectral dictionary. Then, we employ the local and non-local sparsity of the HSI under the learned spatial-spectral dictionary to design an HSI denoising model, which can be effectively solved by an iterative numerical algorithm with parameters that are adaptively adjusted for different clusters and different noise levels. Experimental results on HSI denoising show that the proposed method can provide substantial improvements over the current state-of-the-art HSI denoising methods in terms of both objective metric and subjective visual quality. Ying Fu 0001, Antony Lam, Imari Sato, Yoichi Sato 0001 |
ICCV | 1 |
| 2015 | Separating Fluorescent and Reflective Components by Using a Single Hyperspectral ImageabstractThis paper introduces a novel method to separate fluorescent and reflective components in the spectral domain. In contrast to existing methods, which require to capture two or more images under varying illuminations, we aim to achieve this separation task by using a single hyperspectral image. After identifying the critical hurdle in single-image component separation, we mathematically design the optimal illumination spectrum, which is shown to contain substantial high-frequency components in the frequency domain. This observation, in turn, leads us to recognize a key difference between reflectance and fluorescence in response to the frequency modulation effect of illumination, which fundamentally explains the feasibility of our method. On the practical side, we successfully find an off-the-shelf lamp as the light source, which is strong in irradiance intensity and cheap in cost. A fast linear separation algorithm is developed as well. Experiments using both synthetic data and real images have confirmed the validity of the selected illuminant and the accuracy of our separation algorithm. Yinqiang Zheng, Ying Fu 0001, Antony Lam, Imari Sato, Yoichi Sato 0001 |
ICCV | 2 |
| 2014 | Reflectance and Fluorescent Spectra Recovery Based on Fluorescent Chromaticity Invariance under Varying IlluminationabstractIn recent years, fluorescence analysis of scenes has received attention. Fluorescence can provide additional information about scenes, and has been used in applications such as camera spectral sensitivity estimation, 3D reconstruction, and color relighting. In particular, hyperspectral images of reflective-fluorescent scenes provide a rich amount of data. However, due to the complex nature of fluorescence, hyperspectral imaging methods rely on specialized equipment such as hyperspectral cameras and specialized illuminants. In this paper, we propose a more practical approach to hyperspectral imaging of reflective-fluorescent scenes using only a conventional RGB camera and varied colored illuminants. The key idea of our approach is to exploit a unique property of fluorescence: the chromaticity of fluorescence emissions are invariant under different illuminants. This allows us to robustly estimate spectral reflectance and fluorescence emission chromaticity. We then show that given the spectral reflectance and fluorescent chromaticity, the fluorescence absorption and emission spectra can also be estimated. We demonstrate in results that all scene spectra can be accurately estimated from RGB images. Finally, we show that our method can be used to accurately relight scenes under novel lighting. Ying Fu 0001, Antony Lam, Yasuyuki Kobashi, Imari Sato, Takahiro Okabe, Yoichi Sato 0001 |
CVPR | 1 |
| 2014 | Interreflection Removal Using Fluorescence
Ying Fu 0001, Antony Lam, Yasuyuki Matsushita, Imari Sato, Yoichi Sato 0001 |
ECCV (5) | 1 |
| 2013 | Separating Reflective and Fluorescent Components Using High Frequency Illumination in the Spectral DomainabstractHyper spectral imaging is beneficial to many applications but current methods do not consider fluorescent effects which are present in everyday items ranging from paper, to clothing, to even our food. Furthermore, everyday fluorescent items exhibit a mix of reflectance and fluorescence. So proper separation of these components is necessary for analyzing them. In this paper, we demonstrate efficient separation and recovery of reflective and fluorescent emission spectra through the use of high frequency illumination in the spectral domain. With the obtained fluorescent emission spectra from our high frequency illuminants, we then present to our knowledge, the first method for estimating the fluorescent absorption spectrum of a material given its emission spectrum. Conventional bispectral measurement of absorption and emission spectra needs to examine all combinations of incident and observed light wavelengths. In contrast, our method requires only two hyper spectral images. The effectiveness of our proposed methods are then evaluated through a combination of simulation and real experiments. We also demonstrate an application of our method to synthetic relighting of real scenes. Ying Fu 0001, Antony Lam, Imari Sato, Takahiro Okabe, Yoichi Sato 0001 |
ICCV | 1 |