EDBT 2026 Demo / reviewers in the wild / expert
Gang Fu 0003
dblp:42/5813-3
· DBLP profile ↗
33ranked-venue papers
8as first author
26since 2021 · last 2026
0000-0002-7432-3174ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 25 · 7 first-author · 18 since 2021Artificial intelligence and machine learning · 11 · 4 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RIR-Agent: An interactive framework for effective and adaptive restoration of remote sensing imagery
Junyu Liu, Tianyu Li 0003, Lanyue Liang, Gang Fu 0003, Guoqing Wang 0001, Quan Rui, Xiongxin Tang, Shuyuan Zhu, Yang Yang 0002 |
Expert Syst. Appl. | 4 |
| 2025 | PHR-DIFF: Portrait Highlights Removal via Patch-aware Diffusion ModelabstractPortraits often suffer from specular highlights due to factors like skin oiliness, lighting conditions, and shooting angles, which degrade aesthetics and affect downstream tasks. Thus, portrait highlight removal is imperative. Previous methods struggle to remove highlights and achieve high-fidelity restoration of disturbed regions simultaneously. In this work, we propose a novel patch-based diffusion model for this task, named PHR-DIFF. Specifically, in the training, we present a patchify training strategy that divides the portrait into equal-sized patches and performs diffusion on these patches individually. This patchify can extract more compact facial features and reduce training costs. Besides, to learn the global coherence of the face, we propose a patch-residual approach. It encodes the full-resolution highlight-free portrait into latent features, which are further used as residual terms to constrain the forward training. In the sampling, we remove portrait highlights in a patch-wise manner and propose a Patch-Aware Highlight Removal (PAHR) mechanism. PAHR leverages features from non-highlight regions to effectively guide the patch-wise removal of highlight components. Experimental results on multiple public datasets demonstrate that PHR-DIFF removes highlights more cleanly and avoids artifacts. Hongsheng Zheng, Zhongyun Bao, Gang Fu 0003, Xuze Jiao, Chunxia Xiao |
AAAI | 3 |
| 2025 | Hierarchical Adaptive Filtering Network for Text Image Specular Highlight RemovalabstractDespite significant advances in the field of specular highlight removal in recent years, existing methods predominantly focus on natural images, where highlights typically appear on raised or edged surfaces of objects. These highlights are often small and sparsely distributed. However, for text images such as cards and posters, the flat surfaces reflect light uniformly, resulting in large areas of highlights. Current methods struggle with these large-area highlights in text images, often producing severe visual artifacts or noticeable discrepancies between filled pixels and the original image in the central high-intensity highlight areas. To address these challenges, we propose the Hierarchical Adaptive Filtering Network (HAFNet). Our approach performs filtering at both the downsampled deep feature layer and the upsampled image reconstruction layer. By designing and applying the Adaptive Comprehensive Filtering Module (ACFM) and Adaptive Dilated Filtering Module (ADFM) at different layers, our method effectively restores semantic information in large-area specular highlight regions and recovers detail loss at various scales. The required filtering kernels are pre-generated by a prediction network, allowing them to adaptively adjust according to different images and their semantic content, enabling robust performance across diverse scenarios. Additionally, we utilize Unity3D to construct a comprehensive large-area highlight dataset featuring images with rich texts and complex textures. Experimental results on various datasets demonstrate that our method outperforms state-of-the-art approaches. Jingbo Hu, Ling Zhang 0017, Gang Fu 0003, Chunxia Xiao |
CVPR | 4 |
| 2025 | Neural Solver of Dichromatic Reflection Model for Specular Highlight Removal
Gang Fu 0003 |
ICCV | 1 |
| 2025 | Multi-Attention Guided Knowledge Distillation For High-Performance Object DetectionabstractKnowledge distillation is beneficial for improving the performance of object detection models. However, the existing methods utilizing attention maps for feature weighting embrace limited flexibility, and may result in the loss of crucial channel and location information. To fully utilize these critical information, this paper introduce a novel Multi-Attention Guided Distillation framework that aims to enrich the channel of detail attention representations by enhancing local feature expressions, while patching feature maps concurrently. Meanwhile, a global spatial attention map is utilized to supplement global feature information. To alleviate the discrepancy in the feature attention map between the teacher and student, we use the original student’s features to mimic the weighted teacher’s features. Extensive ablation experiments have proven the effectiveness of our method. Compared with other distillation methods, the SOTA results demostrate the advanced nature of our model, and some student models even outperform the teacher model. Zhihao Kong, Qifeng Lin, Qishen Shen, Jiayi Qiu, Gang Fu 0003, Yuanlong Yu 0001 |
ICME | 5 |
| 2025 | I 2HDiffuser: Image Illumination Harmonization Meets the Diffusion ModelabstractRecently, since diffusion models show great potential in image generation, many pretrained diffusion models based image composition methods have been proposed for image illumination harmonization. However, they mainly face two key challenges: 1) the effective preservation of foreground appearance (i.e., content structure and texture details, etc); 2) Reasonable generation of the foreground casting shadow. To this end, we propose a novel Image Illumination Harmonization Diffusion model called I 2 HDiffuser to achieve image illumination harmonization with high-fidelity foreground appearance and reasonable cast shadows. I 2 HDiffuser mainly consists of frequency domain feature enhancement branch (FDFEB) and illumination-shadow consistency generation branch (ISCGB). Specifically, FDFEB first introduces the Wavelet Transform Module (WTM) for decomposing composite image features into low-frequency (i.e., illumination features, etc) and high-frequency (i.e., texture and content structure features, etc) components using the Haar wavelet transform. Then the Multi-Condition Guidance Mechanism (M-CGM) is proposed to interact these components as prior conditions, which are further injected into the ISCGB with a noise-to-denoise process for guiding high-fidelity content and background illumination-aware foreground regeneration. Meanwhile, a shadow mask step-wise iterative optimization strategy is introduced to the ISCGB to explicitly provide a reasonable shadow generation space for foreground objects. Extensive experiments on public image harmonization datasets DESOBAv2 and iHarmony4 and real illumination harmonization dataset IH-SG show that the I 2HDiffuser achieves the superiority. Zhongyun Bao, Gang Fu 0003, Jianchi Sun, Chunxia Xiao |
ACM Multimedia | 2 |
| 2025 | Attention-Based Mean-Max Balance Assignment for Oriented Object Detection in Optical Remote Sensing ImagesabstractFor objects with arbitrary angles in optical remote sensing (RS) images, the oriented bounding box regression task often faces the problem of ambiguous boundaries between positive and negative samples. The statistical analysis of existing label assignment strategies reveals that anchors with low Intersection over Union (IoU) between ground truth (GT) may also accurately surround the GT after decoding. Therefore, this article proposes an attention-based mean-max balance assignment (AMMBA) strategy, which consists of two parts: mean-max balance assignment (MMBA) strategy and balance feature pyramid with attention (BFPA). MMBA employs the mean-max assignment (MMA) and balance assignment (BA) to dynamically calculate a positive threshold and adaptively match better positive samples for each GT for training. Meanwhile, to meet the need of MMBA for more accurate feature maps, we construct a BFPA module that integrates spatial and scale attention mechanisms to promote global information propagation. Combined with S2ANet, our AMMBA method can effectively achieve state-of-the-art performance, with a precision of 80.91% on the DOTA dataset in a simple plug-and-play fashion. Extensive experiments on three challenging optical RS image datasets (DOTA-v1.0, HRSC, and DIOR-R) further demonstrate the balance between precision and speed in single-stage object detectors. Our AMMBA has enough potential to assist all existing RS models in a simple way to achieve better detection performance. The code is available athttps://github.com/promisekoloer/AMMBA. Qifeng Lin, Daoye Zhu, Gang Fu 0003, Chuanxi Chen, Yuanlong Yu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | Multiple Region Proposal Experts Network for Wide-Scale Remote Sensing Object DetectionabstractFaced with the wide-scale characteristics of objects in optical remote sensing images, the current object detection models are always unable to provide satisfactory detection capabilities for remote sensing tasks. To achieve better wide-scale coverage for various remote sensing regions of interest, this article introduces a multiprediction mechanism to build a novel region generation model, namely, a multiple region proposal experts network (MRPENet). Meanwhile, to achieve both region proposal coverage and receptive field coverage of wide-scale objects, we constructed a prior design of an anchor (PDA) module and an adaptive features compensation (AFC) module to achieve the coverage of wide-scale remote sensing objects. To better utilize the multiexpert characteristics of our model, we customized a new training sample allocation strategy, dynamic scale-assigned expert learning (DSAEL), to cultivate the ability of experts to deal with objects at various scales. To the best of our knowledge, this is the first time that a multiple region proposal network (RPN) mechanism has been used in the object detection of optical remote sensing images. Extensive experiments have shown the generality and effectiveness of our MRPENet. Without bells and whistles, MRPENet achieves a new state-of-the-art (SOTA) on standard benchmarks, i.e., DOTA-v1.0 [82.02% mean average precision (mAP)], HRSC2016 (98.16% mAP), and FAIR1M-v1.0 (48.80% mAP). Qifeng Lin, Daoye Zhu, Gang Fu 0003, Yuanlong Yu 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Two-Stage Video Shadow Detection via Temporal-Spatial Adaption
Xin Duan, Yu Cao 0019, Lei Zhu 0003, Gang Fu 0003, Xin Wang 0118, Ping Li 0016 |
ECCV (48) | 4 |
| 2024 | CFDiffusion: Controllable Foreground Relighting in Image Compositing via Diffusion ModelabstractInserting foreground objects into specific background scenes and eliminating the illumination inconsistency (eg., color, brightness) between them is an important and challenging task. It typically involves multiple processing tasks, such as image harmonization and shadow generation. In these two domains, there are already many mature solutions, but they often only focus on one of the tasks. Recently, some image composition methods have utilized diffusion models to address both of these issues simultaneously, but they cannot guarantee complete reconstruction of the foreground content. In this work, we propose CFDiffusion, which can simultaneously handle image harmonization and shadow generation. We first employ a shadow mask predictor to estimate the shadow mask of the foreground object. Next, we design a harmonization-shadow generator based on a diffusion model to harmonize the foreground and generate shadows concurrently. Additionally, we propose a foreground content enhancement module to ensure the complete preservation of foreground content at the insertion location, and we also develop an adaptive encoder to guide the harmonization process in the foreground area. The experimental results on the iHarmony4 dataset and the IH-SG dataset demonstrate the superiority of our CFDiffusion approach. Zhongyun Bao, Gang Fu 0003, Weilei He, Chao Liang 0001, Chunxia Xiao |
ACM Multimedia | 4 |
| 2024 | HighlightRemover: Spatially Valid Pixel Learning for Image Specular Highlight Removal
Ling Zhang 0017, Yidong Ma, Weilei He, Zhongyun Bao, Gang Fu 0003, Wenju Xu, Chunxia Xiao |
ACM Multimedia | 6 |
| 2024 | Foreground Harmonization and Shadow Generation for Composite Image
Zhongyun Bao, Gang Fu 0003, Weilei He, Chao Liang 0001, Chunxia Xiao |
ACM Multimedia | 4 |
| 2024 | Illuminator: Image-based illumination editing for indoor scene harmonizationabstractIllumination harmonization is an important but challenging task that aims to achieve illumination compatibility between the foreground and background under different illumination conditions. Most current studies mainly focus on achieving seamless integration between the appearance (illumination or visual style) of the foreground object itself and the background scene or producing the foreground shadow. They rarely considered global illumination consistency (i.e., the illumination and shadow of the foreground object). In our work, we introduce “Illuminator”, an image-based illumination editing technique. This method aims to achieve more realistic global illumination harmonization, ensuring consistent illumination and plausible shadows in complex indoor environments. The Illuminator contains a shadow residual generation branch and an object illumination transfer branch. The shadow residual generation branch introduces a novel attention-aware graph convolutional mechanism to achieve reasonable foreground shadow generation. The object illumination transfer branch primarily transfers background illumination to the foreground region. In addition, we construct a real-world indoor illumination harmonization dataset called RIH, which consists of various foreground objects and background scenes captured under diverse illumination conditions for training and evaluating our Illuminator. Our comprehensive experiments, conducted on the RIH dataset and a collection of real-world everyday life photos, validate the effectiveness of our method. Zhongyun Bao, Gang Fu 0003, Zipei Chen, Chunxia Xiao |
Comput. Vis. Media | 2 |
| 2024 | Towards High-Resolution Specular Highlight Detection
Gang Fu 0003, Qing Zhang 0006, Lei Zhu 0003, Qifeng Lin, Siyuan Fan, Chunxia Xiao |
Int. J. Comput. Vis. | 1 |
| 2024 | Eyeglass Reflection Removal With Joint Learning of Reflection Elimination and Content InpaintingabstractEyeglass reflection removal is of great importance to the portrait image processing. However, it remains a challenge to eliminate the reflections on the glass and restore the textual contents of eyes without introducing visual artifacts. Addressing this problem, in this paper, we propose an Eyeglass Reflection Removal Network (ER2Net) by learning reflection elimination and content inpainting jointly. The reflection elimination branch is effective in weak reflection regions, and the content inpainting branch is dedicated to content reasoning in strong reflection regions. We then propose a result fusion module (RFM), which adaptively fuses the elimination result and the inpainting result according to the reflection intensity of each pixel, to produce high-quality result. We also design a memory module for improving the content inpainting result, and propose an eye-symmetry loss to avoid visual artifacts. Additionally, we construct the first Real-world eyeglass Reflection (ReyeR) dataset for eyeglass reflection removal. Extensive quantitative and qualitative experiments demonstrate the superiority of the ER2Net over state-of-the-art methods for eyeglass reflection removal. Wentao Zou, Xiao Lu 0002, Zhilv Yi, Ling Zhang 0017, Gang Fu 0003, Ping Li 0016, Chunxia Xiao |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | Towards High-Quality Photorealistic Image Style TransferabstractPreserving important textures of the content image and achieving prominent style transfer results remains a challenge in the field of image style transfer. This challenge arises from the entanglement between color and texture during the style transfer process. To address this challenge, we propose an end-to-end network that incorporates adaptive weighted least squares (AWLS) filter, iterative least squares (ILS) filter, and channel separation. Given a content image ($\mathcal {C}$) and a reference style image ($\mathcal {S}$), we begin by separating the RGB channels and utilizing ILS filter to decompose them into structure and texture layers. We then perform style transfer on the structural layers using WCT$^{2}$(incorporating wavelet pooling and unpooling techniques for whitening and coloring transforms) in the R, G, and B channels, respectively. We address the texture distortion caused by WCT$^{2}$with a texture enhancing (TE) module in the structural layer. Furthermore, we propose an estimating and compensating for the structure loss (ECSL) module. In the ECSL module, with the AWLS filter and the ILS filter, we estimate the texture loss caused by TE, convert the loss of the structural layer to the loss of the texture layer, and compensate for the loss in the texture layer. The final structural layer and the texture layer are merged into the channel style transfer results in the separated R, G, and B channels into the final style transfer result. Thereby, this enables a more complete texture preservation and a significant style transfer process. To evaluate our method, we utilize quantitative experiments using various metrics, including NIQE, AG, SSIM, PSNR, and a user study. The experimental results demonstrate the superiority of our approach over the previous state-of-the-art methods. Haimin Zhang 0001, Gang Fu 0003, Caoqing Jiang, Fei Luo 0004, Chunxia Xiao, Min Xu 0001 |
IEEE Trans. Multim. | 3 |
| 2023 | Towards High-Quality Specular Highlight Removal by Leveraging Large-Scale Synthetic DataabstractThis paper aims to remove specular highlights from a single object-level image. Although previous methods have made some progresses, their performance remains somewhat limited, particularly for real images with complex specular highlights. To this end, we propose a three-stage network to address them. Specifically, given an input image, we first decompose it into the albedo, shading, and specular residue components to estimate a coarse specular-free image. Then, we further refine the coarse result to alleviate its visual artifacts such as color distortion. Finally, we adjust the tone of the refined result to match the tone of the input as closely as possible. In addition, to facilitate network training and quantitative evaluation, we present a large-scale synthetic dataset of object-level images, covering diverse objects and illumination conditions. Extensive experiments illustrate that our network is able to generalize well to unseen real object-level images, and even produce good results for scene-level images with multiple background objects and complex lighting. Gang Fu 0003, Qing Zhang 0006, Lei Zhu 0003, Chunxia Xiao, Ping Li 0016 |
ICCV | 1 |
| 2023 | A Multiple Prediction Mechanisms Ensemble for Complex Remote Sensing ScenesabstractFacing complex remote sensing scenes, detection models with single detection mechanisms cannot always provide satisfactory detection capabilities. In order to obtain better detection performance in various remote sensing scenes, this paper constructs a novel ensemble model, namely: the multiple prediction mechanisms ensemble (MPME). In order to improve the feature representation ability and region recognition ability of the ensemble model, we build the ensemble of feature pyramids (EFP) and the ensemble of detection heads (EDH) respectively. In order to further improve the detection accuracy of the ensemble model, we propose a training strategy (k-Nearest Loss Learning), so that each sub-detector does not need to learn a trade-off among all training samples, and also reduces the possibility of model over-fitting. The experimental results show that our MPME is a more efficient and effective ensemble model. Compared with other ensemble models, our MPME has a faster detection speed and better detection accuracy. Compared with other state-of-the-art detectors, our detector also achieves superior detection performance. Qifeng Lin, Luojun Lin, Yuanlong Yu 0001, Gang Fu 0003 |
ACM Multimedia | 4 |
| 2022 | Deep Image-based Illumination HarmonizationabstractIntegrating a foreground object into a background scene with illumination harmonization is an important but challenging task in computer vision and augmented reality community. Existing methods mainly focus on foreground and background appearance consistency or the foreground object shadow generation, which rarely consider global appearance and illumination harmonization. In this paper, we formulate seamless illumination harmonization as an illumination exchange and aggregation problem. Specifically, we firstly apply a physically-based rendering method to construct a large-scale, high-quality dataset (named IH) for our task, which contains various types of foreground objects and background scenes with different lighting conditions. Then, we propose a deep image-based illumination harmonization GAN framework named DIH-GAN, which makes full use of a multi-scale attention mechanism and illumination exchange strategy to directly infer mapping relationship between the inserted foreground object and the corresponding background scene. Meanwhile, we also use adversarial learning strategy to further refine the illumination harmonization result. Our method can not only achieve harmonious appearance and illumination for the foreground object but also can generate compelling shadow cast by the foreground object. Comprehensive experiments on both our IH dataset and real-world images show that our proposed DIH-GAN provides a practical and effective solution for image-based object illumination harmonization editing, and validate the superiority of our method against state-of-the-art methods. Our IH dataset is available at https://github.com/zhongyunbao/Dataset. Zhongyun Bao, Chengjiang Long, Gang Fu 0003, Daquan Liu, Yuanzhen Li, Chunxia Xiao |
CVPR | 3 |
| 2022 | Photorealistic Style Transfer via Adaptive Filtering and Channel SeperationabstractThe problem of color and texture distortion remains unsolved in the photorealistic style transfer task. It is mainly caused by the interference between color and texture during transferring. To address this problem, we propose a end-to-end network via adaptive filtering and channel separation. Given a pair of content image and reference image, we firstly decompose them into two structure layers through adaptive weighted least squares filter (AWLSF), which could better perceive the color structure and illumination. Then, we carry out RGB transfer in a channel separation way on the two generated structure layers. To deal with texture in a relatively independent manner, we use a module and a subtraction operation to get more complete and clear content features. Finally, we merge the color structure and texture detail into the ultimate result. We conduct solid quantitative experiments on four metrics NIQE, AG, SSIM, and PSNR, and make a user study. The experimental results demonstrate that our method is able to produce better results than previous state-of-the-art methods, and validate the effectiveness and superiority of our method. Fei Luo 0004, Caoqing Jiang, Gang Fu 0003, Zipei Chen, Shenghong Hu, Chunxia Xiao |
ACM Multimedia | 4 |
| 2022 | Deep attentive style transfer for images with wavelet decomposition
Gang Fu 0003, Qinan Yan, Caoqing Jiang, Tuo Cao, Shenghong Hu, Chunxia Xiao |
Inf. Sci. | 2 |
| 2022 | DDBN: Dual detection branch network for semantic diversity predictions
Qifeng Lin, Chengjiang Long, Jianhui Zhao 0001, Gang Fu 0003 |
Pattern Recognit. | 4 |
| 2022 | MEDNet: Multiexpert Detection Network With Unsupervised Clustering of Training SamplesabstractFor various remote sensing objects, the current detection framework based on a single detection pipeline fails to provide satisfactory detection accuracy. In order to further improve the object recognition ability of the detection model, this article introduces the effective “multiexpert” mechanism into the field of remote sensing object detection and then constructs a multiexpert detection network (MEDNet). In this model, we first construct multiple feature pyramids (MFPs) to replace the traditional single feature pyramid to enrich the semantic representation ability of the model. Then, we equip multiple detection experts (MDEs) to leverage multiple kinds of features from MFP to perform different semantic predictions. As the first CNN-based multiexpert detection model for remote sensing images, we tailor a loss distance-based k-experts clustering (LD-kEC) strategy to assign training samples to different detection experts in an unsupervised fashion. By this strategy, we can directly use the existing remote sensing dataset without expert labels for end-to-end training of our multiexpert model. The experimental results prove that the proposed multiexpert-based detector can indeed significantly improve the object detection performance for remote sensing images. Qifeng Lin, Jianhui Zhao 0001, Bo Du 0001, Gang Fu 0003 |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2022 | CRPN-SFNet: A High-Performance Object Detector on Large-Scale Remote Sensing ImagesabstractLimited by the GPU memory, the current mainstream detectors fail to directly apply to large-scale remote sensing images for object detection. Moreover, the scale range of objects in remote sensing images is much wider than that of general images, which also greatly hinders the existing methods to effectively detect geospatial objects of various scales. For achieving high-performance object detection on large-scale remote sensing images, this article proposes a much faster and more accurate detecting framework, called cropping region proposal network-based scale folding network (CRPN-SFNet). In our framework, the CRPN includes a weak semantic RPN for quickly locating interesting regions and a strategy of generating cropping regions to effectively filter out meaningless regions, which can greatly reduce the computation and storage burden. Meanwhile, the proposed SFNet leverages the scale folding-based training and testing methods to extend the valid detection range of existing detectors, which is beneficial for detecting remote sensing objects of various scales, including very small and very large geospatial objects. Extensive experiments on the public Dataset for Object deTection in Aerial images data set indicate that our CRPN can help our detector deal the larger image faster with the limited GPU memory; meanwhile, the SFNet is beneficial to achieve more accurate detection of geospatial objects with wide-scale range. For large-scale remote sensing images, the proposed detection framework outperforms the state-of-the-art object detection methods in terms of accuracy and speed. Qifeng Lin, Jianhui Zhao 0001, Gang Fu 0003 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | Interactive lighting editing system for single indoor low-light scene images with corresponding depth mapsabstractWe propose a novel interactive lighting editing system for lighting a single indoor RGB image based on spherical harmonic lighting. It allows users to intuitively edit illumination and relight the complicated low-light indoor scene. Our method not only achieves plausible global relighting but also enhances the local details of the complicated scene according to the spatially-varying spherical harmonic lighting, which only requires a single RGB image along with a corresponding depth map. To this end, we first present a joint optimization algorithm, which is based on the geometric optimization of the depth map and intrinsic image decomposition avoiding texture-copy, for refining the depth map and obtaining the shading map. Then we propose a lighting estimation method based on spherical harmonic lighting, which not only achieves the global illumination estimation of the scene, but also further enhances local details of the complicated scene. Finally, we use a simple and intuitive interactive method to edit the environment lighting map to adjust lighting and relight the scene. Through extensive experimental results, we demonstrate that our proposed approach is simple and intuitive for relighting the low-light indoor scene, and achieve state-of-the-art results. Zhongyun Bao, Gang Fu 0003, Chunxia Xiao |
Vis. Informatics | 2 |
| 2021 | A Multi-Task Network for Joint Specular Highlight Detection and RemovalabstractSpecular highlight detection and removal are fundamental and challenging tasks. Although recent methods have achieved promising results on the two tasks by training on synthetic training data in a supervised manner, they are typically solely designed for highlight detection or removal, and their performance usually deteriorates significantly on real-world images. In this paper, we present a novel network that aims to detect and remove highlights from natural images. To remove the domain gap between synthetic training samples and real test images, and support the investigation of learning-based approaches, we first introduce a dataset with about 16K real images, each of which has the corresponding ground truths of highlight detection and removal. Using the presented dataset, we develop a multi-task network for joint highlight detection and removal, based on a new specular highlight image formation model. Experiments on the benchmark datasets and our new dataset show that our approach clearly outperforms state-of-the-art methods for both highlight detection and removal. Gang Fu 0003, Qing Zhang 0006, Lei Zhu 0003, Ping Li 0016, Chunxia Xiao |
CVPR | 1 |
| 2020 | Learning to Detect Specular Highlights from Real-world ImagesabstractSpecular highlight detection is a challenging problem, and has many applications such as shiny object detection and light source estimation. Although various highlight detection methods have been proposed, they fail to disambiguate bright material surfaces from highlights, and cannot handle non-white-balanced images. Moreover, at present, there is still no benchmark dataset for highlight detection. In this paper, we present a large-scale real-world highlight dataset containing a rich variety of material categories, with diverse highlight shapes and appearances, in which each image is with an annotated ground-truth mask. Based on the dataset, we develop a deep learning-based specular highlight detection network (SHDNet) leveraging multi-scale context contrasted features to accurately detect specular highlights of varying scales. In addition, we design a binary cross-entropy (BCE) loss and an intersection-over-union edge (IoUE) loss for our network. Compared with existing highlight detection methods, our method can accurately detect highlights of different sizes, while effectively excluding the non-highlight regions, such as bright materials, non-specular as well as colored lighting, and even light sources. Gang Fu 0003, Qing Zhang 0006, Qifeng Lin, Lei Zhu 0003, Chunxia Xiao |
ACM Multimedia | 1 |
| 2020 | Shading-aware shadow detection and removal from a single image
Xinyun Fan, Ling Zhang 0017, Qingan Yan, Gang Fu 0003, Zipei Chen, Chengjiang Long, Chunxia Xiao |
Vis. Comput. | 5 |
| 2019 | A Hybrid L2 -LP Variational Model For Single Low-Light Image Enhancement With Bright Channel PriorabstractIn this paper, we consider and study the norm variable and propose a hybrid L2-Lpvariational model with bright channel prior based on Retinex to decompose an observed image into a reflectance layer and an illumination layer. Different from the existing methods, our proposed model can preserve the reflectance layer with more fine details while enforcing the illumination layer to be texture-less, avoiding the texture-copy problem. Moreover, for solving our non-linear optimization, we adopt an alternating minimization scheme to find the optimal. Finally, we test our algorithm on a large number of images and the experimental results illustrate that the proposed method has achieved the better result than other state-of-the-art methods both qualitatively and quantitatively. Gang Fu 0003, Chunxia Xiao |
ICIP | 1 |
| 2019 | Towards High-Quality Intrinsic Images in the WildabstractWe address the intrinsic image decomposition problem for separating an image into its intrinsic images, i.e, a reflectance layer and a shading layer. Although this problem has been studied for decades, it remains a significant challenge, particularly for real-world images. In this paper, we present a novel method for estimating high-quality intrinsic images for real-world images. Our method is built upon two observations on real-world images: (i) reflectance is generally sparse and there are limited number of reflectance values in an image; (ii) shading usually has locally smooth transition. Based on the two observations, we formulate the decomposition problem into an optimization framework, where we encourage the reflectance sparseness by globally confining the number of reflectance discontinuities among neighboring pixels using an L_0 norm, and utilize a total variation for maintaining locally smooth shading. We employ two benchmark datasets and perform various experiments to evaluate our method. Experimental results show that our method outperforms state-of-the-art methods, both qualitatively and quantitatively. Gang Fu 0003, Qing Zhang 0006, Chunxia Xiao |
ICME | 1 |
| 2019 | Cropping Region Proposal Network Based Framework for Efficient Object Detection on Large Scale Remote Sensing ImagesabstractIt is very difficult to directly detect objects on the entire large scale remote sensing image, due to the limited GPU memory. Moreover, there are no objects of interest in most areas of such a huge image, thus a lot of computational costs is wasted in dealing with these vain areas. Therefore, this paper proposes a Cropping Region Proposal Network (CRPN), which includes a weak semantic RPN for quickly locating interesting regions, and a dual-scale strategy for generating effective cropping regions. Cropping regions consist of small and large cropping scales for detecting various-scale objects including very small and very large objects, which is hard for existing methods. CRPN helps to detect effective regions of remote sensing image. Meanwhile, it is also modularized and can be easily connected with mainstream detectors to form an end-to-end detecting framework. Experiments on public DOTA dataset show that our CRPN is effective for filtering invalid regions to greatly reduce the computation burden, and helps to achieve more accurate object detection on large scale remote sensing images. Qifeng Lin, Jianhui Zhao 0001, Qianqian Tong 0001, Guian Zhang, Gang Fu 0003 |
ICME | 6 |
| 2019 | Wavelet Flow: Optical Flow Guided Wavelet Facial Image FusionabstractAbstract Estimating the correspondence between the images using optical flow is the key component for image fusion, however, computing optical flow between a pair of facial images including backgrounds is challenging due to large differences in illumination, texture, color and background in the images. To improve optical flow results for image fusion, we propose a novel flow estimation method, wavelet flow, which can handle both the face and background in the input images. The key idea is that instead of computing flow directly between the input image pair, we estimate the image flow by incorporating multi‐scale image transfer and optical flow guided wavelet fusion. Multi‐scale image transfer helps to preserve the background and lighting detail of input, while optical flow guided wavelet fusion produces a series of intermediate images for further fusion quality optimizing. Our approach can significantly improve the performance of the optical flow algorithm and provide more natural fusion results for both faces and backgrounds in the images. We evaluate our method on a variety of datasets to show its high outperformance. Qingan Yan, Gang Fu 0003, Chunxia Xiao |
Comput. Graph. Forum | 3 |
| 2019 | Specular Highlight Removal for Real-world ImagesabstractAbstract Removing specular highlight in an image is a fundamental research problem in computer vision and computer graphics. While various methods have been proposed, they typically do not work well for real‐world images due to the presence of rich textures, complex materials, hard shadows, occlusions and color illumination, etc. In this paper, we present a novel specular highlight removal method for real‐world images. Our approach is based on two observations of the real‐world images: (i) the specular highlight is often small in size and sparse in distribution; (ii) the remaining diffuse image can be represented by linear combination of a small number of basis colors with the sparse encoding coefficients. Based on the two observations, we design an optimization framework for simultaneously estimating the diffuse and specular highlight images from a single image. Specifically, we recover the diffuse components of those regions with specular highlight by encouraging the encoding coefficients sparseness using L0 norm. Moreover, the encoding coefficients and specular highlight are also subject to the non‐negativity according to the additive color mixing theory and the illumination definition, respectively. Extensive experiments have been performed on a variety of images to validate the effectiveness of the proposed method and its superiority over the previous methods. Gang Fu 0003, Qing Zhang 0006, Chengfang Song, Qifeng Lin, Chunxia Xiao |
Comput. Graph. Forum | 1 |