Fengying Xie

dblp:121/9085 · DBLP profile ↗
← Back
56ranked-venue papers
3as first author
32since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 34 · 2 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 22 · 9 since 2021Artificial intelligence and machine learning · 9 · 1 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Rectification Reimagined: A Unified Mamba Model for Image Correction and Rectangling with Prompts
abstract
Image correction and rectangling are valuable tasks in practical photography systems such as smartphones. Recent remarkable advancements in deep learning have undeniably brought about substantial performance improvements in these fields. Nevertheless, existing methods mainly rely on task-specific architectures. This significantly restricts their generalization ability and effective application across a wide range of different tasks. In this paper, we introduce the Unified Rectification Framework (UniRect), a comprehensive approach that addresses these practical tasks from a consistent distortion rectification perspective. Our approach incorporates various task-specific inverse problems into a general distortion model by simulating different types of lenses. To handle diverse distortions, UniRect adopts one task-agnostic rectification framework with a dual-component structure: a Deformation Module, which utilizes a novel Residual Progressive Thin-Plate Spline (RP-TPS) model to address complex geometric deformations, and a subsequent Restoration Module, which employs Residual Mamba Blocks (RMBs) to counteract the degradation caused by the deformation process and enhance the fidelity of the output image. Moreover, a Sparse Mixture-of-Experts (SMoEs) structure is designed to circumvent heavy task competition in multi-task learning due to varying distortions. Extensive experiments demonstrate that our models have achieved state-of-the-art performance compared with other up-to-date methods.
Linwei Qiu, Gongzhe Li, Xiaozhe Zhang, Qi Sun 0001, Fengying Xie
AAAI5
2025 Beyond Spatial Domain: Cross-domain Promoted Fourier Convolution Helps Single Image Dehazing
abstract
Vanilla convolution and window-based self-attention have shown significant success in image dehazing. However, they are constrained by limited receptive fields and ignore frequency gaps between dehazed and clear images. The former hampers the modeling of global dependencies, while the latter impedes the learning of high-frequency features, leading to suboptimal performance. In this paper, we propose the Joint Spatial and Fourier Convolutional Network (JSFC-Net), which leverages Fourier transformation to simultaneously address the two aforementioned problems with low computational overhead. We introduce the Frequency-Spatial Promoted and Physical Learning Block, which extracts high-level features from the spatial domain and frequency domain in parallel. We design a simple yet effective solution that uses spatial features to promote and modulate frequency features in a multi-scale manner, achieving refinement of frequency features and addressing robustness issue caused by global sensitivity. Additionally, we present the Receptive Field Selection Module to facilitate improved fusion of spatial and frequency domain features. Finally, we introduce frequency loss to further narrow frequency gaps. Comprehensive experiments on multiple datasets demonstrate that JSFC-Net is significantly superior to SOTA dehazing methods.
Xiaozhe Zhang, Haidong Ding, Fengying Xie, Linpeng Pan, Yue Zi, Haopeng Zhang 0001
AAAI3
2025 Real-Time Scene-Adaptive Tone Mapping for High-Dynamic Range Object Detection
abstract
High dynamic range (HDR) images, with their rich tone and detail reproduction, hold significant potential to enhance computer vision systems, particularly in autonomous driving. However, most neural networks for embedded vision are trained on low dynamic range (LDR) inputs and suffer substantial performance degradation when handling high-bit-depth HDR images due to the challenges posed by extreme dynamic ranges. In this paper, we propose a novel tone mapping method that not only bridges the gap between HDR RAW inputs and the LDR sRGB requirements of detection networks but also achieves end-to-end optimization with the downstream tasks. Instead of relying on traditional image signal processing (ISP) pipeline, we introduce neural photometric calibration to regularize dynamic ranges and a scaling-invariant local tone mapping module to preserve image details. In addition, our architecture also supports performance transfer finetuning, enabling efficient adaptation from the LDR model to the HDR RAW model with minimal cost. The proposed method outperforms traditional tone mapping algorithms and advanced AI-ISP methods in challenging automotive HDR scenes. Moreover, our pipeline achieves real-time processing of 4K high-bit-depth HDR inputs on the Nvidia Jetson platform.
Gongzhe Li, Linwei Qiu, Peibei Cao, Fengying Xie, Xiangyang Ji, Qilin Sun 0001
NeurIPS4
2025 Infrared Small Target Detection Based on Prior Guided Dense Nested Network
abstract
Infrared small target detection (IRSTD) has been widely applied and developed in military and civilian fields, playing a vital role. Despite the extensive research foundation of traditional manual feature-based methods, they are still constrained by the inherent problem of infrared small targets lacking prior features. In recent years, the advancement of deep learning methods has enriched the research landscape in this field, yet they are still constrained by the imbalance of positive and negative samples between the target and the background. To address these issues, we propose a novel prior guided dense nested network (PGDN-Net), which ingeniously integrates traditional manual features with a deep learning network model. First, three prior features are extracted, including the high-order Riesz transform feature, the compactness and heterogeneity feature (CH), and the corner feature of the structure tensor (ST). Then, these features are input into a dense nested network for guidance, supported by a two-orientation attention aggregation module and a channel and spatial attention module. Different features play their respective guiding roles in different depths of the network. Through multiple attention mechanisms and feature fusion operations on the interested target area, the extraction and preservation of target features can be improved, while easily removing irrelevant backgrounds. Experiments on public datasets demonstrate the effectiveness and progressiveness of our PGDN-Net. Compared with other state-of-the-art methods, it achieves better performance in background suppression, target enhancement, probability of detection, and false alarm rate. In addition, the PGDN-Net model can effectively maintain and restore the original shape of the target while performing robust detection, which is beneficial for subsequent fine-grained recognition tasks.
Chang Liu 0090, Xuedong Song, Dianyu Yu, Linwei Qiu, Fengying Xie, Yue Zi, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.5
2025 M3-CR: Multiscale Multibranch Mamba for SAR-Assisted Optical Image Thick Cloud Removal
abstract
SAR-assisted thick cloud removal from optical remote sensing images has long been a challenging task. Current mainstream methods face challenges in achieving an effective global receptive field, fully utilizing multi-scale features, and deeply integrating features from both modalities. To overcome these limitations, we propose the Multi-scale Multi-branch Mamba model(M3-CR) for SAR-assisted thick cloud removal. Specifically, we integrate the Mamba model into the task of SAR-assisted cloud removal, effectively modeling global dependencies within the images. Concurrently, a Multi-scale Multi-branch structure is introduced to extract and integrate multi-scale information, and in combination with a convolutional branch to fully exploit the global and local geographic proximities inherent in remote sensing images. Furthermore, we present a novel feature fusion module leveraging the Modal-Traversing 2D Selective Scan(MTSS2D) to enable deep interaction and integration of features from optical and SAR images. The experimental results on two benchmark databases show that the M3-CR achieves superior performance compared to state-of-theart cloud removal approaches, while requiring fewer parameters and reduced FLOPs. The code for M3-CR will be made publicly available at https://github.com/LinpengPan/M3CR.
Linpeng Pan, Xuedong Song, Fengying Xie, Xiaozhe Zhang, Haolin Ji, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.3
2025 Radiation-Tolerant Unsupervised Deep Image Stitching for Remote Sensing
Linwei Qiu, Fengying Xie, Chang Liu 0090, Xiaoling Che, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.2
2025 HUNTNet: Homomorphic Unified Nexus Topology for Camouflaged Object Detection
abstract
Camouflaged object detection (COD) is challenging for both human and computer vision, as targets often blend into the background by sharing similar color, texture, or shape. While many feature enhancement techniques exist, single-view methods tend to overemphasize certain Recognizing that camouflaged objects exhibit different concealment strategies under varying observational perspectives, we propose HUNTNet, a network that establishes a dynamic detection mechanism to decouple target features from RGB images and perform topological decamouflage across multiple homomorphic feature spaces through a unified feature focusing architecture. We adopt PVTv2 as the backbone to extract multi-perspective spatial features. Detail representation is enhanced via a feature module that integrates Dual-Channel Recursive (DCR), Wavelet-Gabor Transform (WGT), and Anisotropic Gradient Responding (AGR), which together improve boundary discrimination and edge contour detection. To further boost performance, the Simplicial Feature Integration (SFI) module recursively fuses multi-layer features, enabling high-resolution focus on target regions. Experiments show that HUNTNet surpasses state-of-the-art methods in both accuracy and generalization, offering a robust solution for COD and improving segmentation in complex scenes. Our code is available at https://github.com/HaolinJi817/HUNTNet.
Haolin Ji, Fengying Xie, Linpeng Pan, Yushan Zheng, Zhenwei Shi 0001
IEEE Trans. Image Process.2
2025 Self-Supervised Representation Distribution Learning for Reliable Data Augmentation in Histopathology WSI Classification
abstract
Multiple instance learning (MIL) based whole slide image (WSI) classification is often carried out on the representations of patches extracted from WSI with a pre-trained patch encoder. The performance of classification relies on both patch-level representation learning and MIL classifier training. Most MIL methods utilize a frozen model pre-trained on ImageNet or a model trained with self-supervised learning on histopathology image dataset to extract patch image representations and then fix these representations in the training of the MIL classifiers for efficiency consideration. However, the invariance of representations cannot meet the diversity requirement for training a robust MIL classifier, which has significantly limited the performance of the WSI classification. In this paper, we propose a Self-Supervised Representation Distribution Learning framework (SSRDL) for patch-level representation learning with an online representation sampling strategy (ORS) for both patch feature extraction and WSI-level data augmentation. The proposed method was evaluated on three datasets under three MIL frameworks. The experimental results have demonstrated that the proposed method achieves the best performance in histopathology image representation learning and data augmentation and outperforms state-of-the-art methods under different WSI classification frameworks. The code is available at https://github.com/lazytkm/SSRDL.
Kunming Tang, Zhiguo Jiang 0001, Kun Wu 0010, Jun Shi 0006, Fengying Xie, Wei Wang 0380, Yushan Zheng
IEEE Trans. Medical Imaging5
2025 Pan-Cancer Histopathology WSI Pre-Training With Position-Aware Masked Autoencoder
abstract
Large-scale pre-training models have promoted the development of histopathology image analysis. However, existing self-supervised methods for histopathology images primarily focus on learning patch features, while there is a notable gap in the availability of pre-training models specifically designed for WSI-level feature learning. In this paper, we propose a novel self-supervised learning framework for pan-cancer WSI-level representation pre-training with the designed position-aware masked autoencoder (PAMA). Meanwhile, we propose the position-aware cross-attention (PACA) module with a kernel reorientation (KRO) strategy and an anchor dropout (AD) mechanism. The KRO strategy can capture the complete semantic structure and eliminate ambiguity in WSIs, and the AD contributes to enhancing the robustness and generalization of the model. We evaluated our method on 7 large-scale datasets from multiple organs for pan-cancer classification tasks. The results have demonstrated the effectiveness and generalization of PAMA in discriminative WSI representation learning and pan-cancer WSI pre-training. The proposed method was also compared with 8 WSI analysis methods. The experimental results have indicated that our proposed PAMA is superior to the state-of-the-art methods. The code and checkpoints are available at https://github.com/WkEEn/PAMA.
Kun Wu 0010, Zhiguo Jiang 0001, Kunming Tang, Jun Shi 0006, Fengying Xie, Wei Wang 0380, Yushan Zheng
IEEE Trans. Medical Imaging5
2024 Prototypical Information Bottlenecking and Disentangling for Multimodal Cancer Survival Prediction
abstract
Multimodal learning significantly benefits cancer survival prediction, especially the integration of pathological images and genomic data. Despite advantages of multimodal learning for cancer survival prediction, massive redundancy in multimodal data prevents it from extracting discriminative and compact information: (1) An extensive amount of intra-modal task-unrelated information blurs discriminability, especially for gigapixel whole slide images (WSIs) with many patches in pathology and thousands of pathways in genomic data, leading to an "intra-modal redundancy" issue. (2) Duplicated information among modalities dominates the representation of multimodal data, which makes modality-specific information prone to being ignored, resulting in an "inter-modal redundancy" issue. To address these, we propose a new framework, Prototypical Information Bottlenecking and Disentangling (PIBD), consisting of Prototypical Information Bottleneck (PIB) module for intra-modal redundancy and Prototypical Information Disentanglement (PID) module for inter-modal redundancy. Specifically, a variant of information bottleneck, PIB, is proposed to model prototypes approximating a bunch of instances for different risk levels, which can be used for selection of discriminative instances within modality. PID module decouples entangled multimodal data into compact distinct components: modality-common and modality-specific knowledge, under the guidance of the joint prototypical distribution. Extensive experiments on five cancer benchmark datasets demonstrated our superiority over other methods. The code is released.
Yilan Zhang, Yingxue Xu, Jianqi Chen, Fengying Xie
ICLR4
2024 Thick Cloud Removal in Multitemporal Remote Sensing Images Using a Coarse-to-Fine Framework
abstract
Abstract—Remote sensing (RS) images are widely used for Earth observation. However, cloud contamination greatly degrades the quality of RS images and limits their applications. In this letter, we propose a coarse-to-fine thick cloud removal method for a single pair of multitemporal RS images. First, we perform a global color transformation on a cloud-free reference image using linear regression coefficients between the pixels in the cloudy target image and the reference image in the same cloud-free regions, and obtain a coarse result. Then, a convolutional neural network (CNN) based on internal constraint is used to refine the coarse result, which does not require any construction of additional external training dataset in advance. We further design a multiscale feature extraction and fusion module and an auxiliary loss involving cloud regions to improve the performance of the CNN. Finally, Poisson image fusion is employed to generate a seamless cloud-free result. On a simulated test set containing 500 pairs of multitemporal RS images, the proposed method achieves satisfactory results with 25.1277 dB in peak signal-to-noise ratio (PSNR), 0.9077 in structural similarity (SSIM), and 0.9342 in correlation coefficient (CC). Qualitative and quantitative comparisons of our proposed against several state-of-the-art methods on the simulated and real cloudy images demonstrate the superiority of the proposed method.
Yue Zi, Xuedong Song, Fengying Xie, Zhiguo Jiang 0001
IEEE Geosci. Remote. Sens. Lett.3
2024 Histopathology language-image representation learning for fine-grained digital pathology cross-modal retrieval
Dingyi Hu, Zhiguo Jiang 0001, Jun Shi 0006, Fengying Xie, Kun Wu 0010, Kunming Tang, Jianguo Huai, Yushan Zheng
Medical Image Anal.4
2024 Robust Haze and Thin Cloud Removal via Conditional Variational Autoencoders
abstract
Existing methods for remote sensing image dehazing and thin cloud removal treat this image restoration task as a clear pixel estimation problem, yielding a single prediction result through a deterministic pipeline. However, image restoration is a highly ill-posed problem, as the sharp pixel value corresponding to the input cannot be uniquely determined solely from the degraded image. In this paper, we present a novel algorithm for haze and thin cloud removal using Conditional Variational Autoencoders (CVAE) to generate multiple realistic restored images for each input. By sampling from the latent space to capture the pixel diversity, the proposed method mitigates the limitations arising from inaccuracies in a single estimation. In this uncertainty pipeline, we can generate a more accurate restored image based on these multiple predictions. Furthermore, we have developed a Dynamic Fusion Network (DFN) for combining multiple plausible outcomes to obtain a more accurate result. DFN dynamically predicts the kernels used for restored result generation conditioned on inputs, improving haze and thin cloud thanks to its adaptive nature. Quantitative and qualitative experiments demonstrate that the proposed method outperforms existing state-of-the-art techniques by a significant margin on dehazing and thin cloud removal benchmarks.
Haidong Ding, Fengying Xie, Linwei Qiu, Xiaozhe Zhang, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.2
2024 Infrared Small Target Detection Based on Monogenic Signal Decomposition
abstract
Robust detection of infrared small target under complex background is of great significance for infrared search and tracking applications. However, the inherent problem of limited prior features for infrared small target has always made its detection task a challenging research topic. In order to solve the problem, we propose a novel infrared small target detection method based on monogenic signal decomposition and feature expansion, which can effectively enrich and extract the potential features of the target. First, a series of local information of the original image is obtained through the monogenic signal constructed by Riesz transform. Then, various features of the small target are extracted from different local signals, including the direction feature, edge feature, and local saliency feature. Finally, the fusion of target features is completed through signal reconstruction, thereby achieving target detection. This method not only pays attention to the local salient characteristic of the target, but also supplements the consideration of other characteristics of the target, providing a new idea for small target detection. The experimental results on real infrared images show that the proposed method framework is reasonable and effective, and possesses better detection performance and good generalization compared to other state-of-the-art methods.
Chang Liu 0090, Fengying Xie, Linwei Qiu, Haolin Ji, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.2
2024 Remote Sensing Image Rectangling With Iterative Warping Kernel Self-Correction Transformer
abstract
Stitched remote sensing images often exhibit irregular boundaries, which can be frustrating for general users and detrimental to downstream tasks such as object detection and segmentation. However, this issue has received insufficient attention and remains unexplored within the remote sensing domain. In this study, we investigate mesh-based rectangling techniques for remote sensing images, aiming to produce rectangular outputs while preserving the original field-of-view (FoV) and avoiding the introduction of unreliable content. Observing that prior rectangling algorithms tend to generate unsatisfactory boundaries or discernible distortions, that is, under-rectangling or over-rectangling, we propose the concept of a warping kernel associated with mesh deformations to account for these phenomena. Consequently, we introduce the iterative warping kernel self-correction transformer (IWKFormer), designed to enhance warping kernel estimation and generate superior rectangular outcomes. It primarily comprises two components: a mesh feature extractor built upon the partial swin transformer block (PSTB) and a corrector module using the swin transformer block (STB). These modules collaborate to derive warping kernels implicitly. The extractor extracts latent features pertinent to mesh deformation, whereas the corrector iteratively refines the warping kernel estimation to improve the ultimate prediction. Furthermore, to bolster further research, we have constructed an aerial imagery stitching rectangling dataset (AIRD), featuring a wide array of stitching scenes. Extensive experimentation on the AIRD demonstrates that our method yields visually appealing and naturally rectangled images, achieving state-of-the-art performance. The code and data will be available athttps://github.com/yyywxk/IWKFormer.
Linwei Qiu, Fengying Xie, Chang Liu 0090, Xuedong Song, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.2
2024 Proxy and Cross-Stripes Integration Transformer for Remote Sensing Image Dehazing
abstract
Existing Transformer-based dehazing methods for remote sensing (RS) images, to avoid quadratic computation complexity with respect to the feature map size, either perform self-attention mechanisms within local windows or capture long-range dependencies in the channel dimension rather than spatial. Each of these methods has its drawbacks. To address these limitations, we propose the Proxy and Cross-Stripes Integration Transformer (PCSformer) for RS image dehazing. PCSformer introduces two innovative Transformer blocks, i.e., sliding cross-stripes Transformer block and local proxy-based global Transformer block. The former allows us to directly model long-range dependencies and capture rich contextual information for large-scale objects in RS images. The latter seeks valuable information for thick haze regions within the whole feature map, generating more consistent and realistic scene details for such regions. Both achieve a large receptive field with cost-effective computational complexity within a single Transformer block. Furthermore, we introduce a shallow deep model with a small receptive field to conduct local refinement, which can mitigate artifacts associated with a large receptive field. Finally, to facilitate the better application of dehazing models to downstream visual tasks, we contribute two large-scale datasets for RS image dehazing. Experiments indicate that the dehazing models trained on our datasets can better assist downstream visual tasks under hazy atmospheric conditions compared to the dehazing models trained on existing datasets. Quantitative and qualitative experiments demonstrate that the proposed PCSformer significantly outperforms existing state-of-the-art techniques on dehazing benchmarks, particularly excelling in the restoration of thick haze scenes. The code and datasets are available athttps://github.com/SmileShaun/PCSformer.
Xiaozhe Zhang, Fengying Xie, Haidong Ding, Shaocheng Yan, Zhenwei Shi 0001
IEEE Trans. Geosci. Remote. Sens.2
2023 Position-Aware Masked Autoencoder for Histopathology WSI Representation Learning
Kun Wu 0010, Yushan Zheng, Jun Shi 0006, Fengying Xie, Zhiguo Jiang 0001
MICCAI (6)4
2023 ECL: Class-Enhancement Contrastive Learning for Long-Tailed Skin Lesion Classification
Yilan Zhang, Jianqi Chen, Fengying Xie
MICCAI (2)4
2023 Feedback Network for Compact Thin Cloud Removal
abstract
The thin cloud removal (CR) technique has great practical value for the application of remote sensing images. Existing deep-learning-based methods have attained remarkable achievements. However, most of them neglect the inherent feature correlations in deeper layers due to learning in a successive manner. In this letter, we propose a compact thin cloud removal network based on the feedback (FB) mechanism, called CRFB-Net, which leverages the high-level features as feedback information to modulate shallow representations. CRFB-Net employs the recurrent architecture to achieve such a feedback scheme. Specifically, the restoration process does not terminate after obtaining an output. In this case, the output of intermediate iterations will flow into the next iteration as feedback. For better utilization of feedback, a multiscale feature fusion block (MFFB) is designed to refine the low-level representations from three scales. Furthermore, we introduce a curriculum learning strategy to train the CRFB-Net by gradually increasing the complexity of restoration, through which a sharper result is produced step by step. Extensive experiments demonstrate the superiority of our CRFB-Net, outperforming state-of-the-art.
Haidong Ding, Fengying Xie, Yue Zi, Xuedong Song
IEEE Geosci. Remote. Sens. Lett.2
2023 Kernel Attention Transformer for Histopathology Whole Slide Image Analysis and Assistant Cancer Diagnosis
abstract
Transformer has been widely used in histopathology whole slide image analysis. However, the design of token-wise self-attention and positional embedding strategy in the common Transformer limits its effectiveness and efficiency when applied to gigapixel histopathology images. In this paper, we propose a novel kernel attention Transformer (KAT) for histopathology WSI analysis and assistant cancer diagnosis. The information transmission in KAT is achieved by cross-attention between the patch features and a set of kernels related to the spatial relationship of the patches on the whole slide images. Compared to the common Transformer structure, KAT can extract the hierarchical context information of the local regions of the WSI and provide diversified diagnosis information. Meanwhile, the kernel-based cross-attention paradigm significantly reduces the computational amount. The proposed method was evaluated on three large-scale datasets and was compared with 8 state-of-the-art methods. The experimental results have demonstrated the proposed KAT is effective and efficient in the task of histopathology WSI analysis and is superior to the state-of-the-art methods.
Yushan Zheng, Jun Li 0106, Jun Shi 0006, Fengying Xie, Jianguo Huai, Zhiguo Jiang 0001
IEEE Trans. Medical Imaging4
2022 Uncertainty-Based Thin Cloud Removal Network via Conditional Variational Autoencoders
Haidong Ding, Yue Zi, Fengying Xie
ACCV (3)3
2022 Histopathology Cross-Modal Retrieval based on Dual-Transformer Network
abstract
Computer-aided cancer diagnosis (CAD) methods based on the histopathological images have achieved great development. The content-based whole slide image (WSI) retrieval is one of the important application that can search for the informative data to assist clinical diagnosis. It is notable that the current retrieval system are mainly developed based on the image content and image labels. The diagnosis report for the WSIs given by the pathologists are also valuable data, but have not yet been adequately considered in modeling. In this paper, we propose a cross-modal retrieval framework based on histopathology WSIs and diagnosis report, which can simultaneously achieve four retrieval tasks for histopathology database across WSIs and diagnosis reports. The compact binary features from both WSIs and diagnosis reports are first extracted, and then built in a common vision-language semantic feature space by the constraint of the designed cross hashing loss function. The method was verified on a gastric histopathology dataset that contains 932 gastric cases with 4 lesion categories. Experimental results have demonstrated the effectiveness of the proposed method in the cross-modal retrieval tasks for digital pathology system.
Dingyi Hu, Fengying Xie, Zhiguo Jiang 0001, Yushan Zheng, Jun Shi 0006
BIBE2
2022 Multi-Frame Super-Resolution With Raw Images Via Modified Deformable Convolution
abstract
In this paper we propose a novel model towards multi-frame super-resolution, which leverages multiple RAW images and yields a super-resolved RGB image. To facilitate the pixel misalignment in burst photography, we apply a refined Pyramid Cascading and Deformable Convolution (PCD) feature alignment module. A new 3D deformable convolution fusion module is proposed subsequently to merge the information from all frames adaptively. In addition, we employ an encoder-decoder network to restore color and details in sRGB space after super-resolving images in linear space. Extensive experiments demonstrate the superiority of our architecture and the strength of multi-frame super-resolution with RAW images.
Gongzhe Li, Linwei Qiu, Haopeng Zhang 0001, Fengying Xie, Zhiguo Jiang 0001
ICASSP4
2022 Lesion-Aware Contrastive Representation Learning for Histopathology Whole Slide Images Analysis
Jun Li 0106, Yushan Zheng, Kun Wu 0010, Jun Shi 0006, Fengying Xie, Zhiguo Jiang 0001
MICCAI (2)5
2022 Kernel Attention Transformer (KAT) for Histopathology Whole Slide Image Classification
Yushan Zheng, Jun Li 0106, Jun Shi 0006, Fengying Xie, Zhiguo Jiang 0001
MICCAI (2)4
2022 Semantic Segmentation of Remote Sensing Image Based on Regional Self-Attention Mechanism
abstract
In remote sensing images (RSIs), accurate semantic segmentation faces more challenges because of small targets, unbalanced categories, and complex scenes. Restricted by local receptive field of convolution layers, the traditional semantic segmentation models cannot use global information of RSIs. According to the characteristics of RSIs, we propose an RSANet based on regional self-attention mechanism. Our model is no longer limited by the locality of convolution, but transfers the information flow in the whole image. It can mine out the relationship between pixels in the surrounding areas, which is more logical for understanding images content. Moreover, compared with the traditional self-attention mechanism, RSANet can effectively reduce the noise of feature maps and the interference of redundant features. Our model can get better semantic segmentation results than other current models on the DroneDeploy data set and the Chreos semantic segmentation data set. The experiments show that our RSANet achieves 2% higher mean intersection over union (mIoU) than the baseline model, especially in terms of fineness, edge integrity, and classification accuracy.
Danpei Zhao, Chenxu Wang 0017, Yue Gao 0008, Zhenwei Shi 0001, Fengying Xie
IEEE Geosci. Remote. Sens. Lett.5
2022 Thin Cloud Removal for Remote Sensing Images Using a Physical-Model-Based CycleGAN With Unpaired Data
abstract
Thin cloud removal from remote sensing (RS) images is challenging. Recently, deep-learning-based methods have achieved excellent results using supervised training on paired image data. However, in practice, real paired image data are unavailable. Therefore, in this letter, we propose a novel thin cloud removal method, a physical-model-based CycleGAN (PM-CycleGAN), which can be trained using only unpaired data. The PM-CycleGAN training process comprises forward and backward loops. The forward loop first decomposes a cloudy image into a cloud-free image, thin cloud thickness map, and thickness coefficient using three generators. Then, it combines these three components using a physical model to reconstruct the original cloudy image to obtain the cycle consistency constraint. The backward loop first uses the physical model to synthesize a cloud-free image, thin cloud thickness map, and thickness coefficient into a cloudy image, which are then decomposed into the original three components using the three generators. Visual and quantitative comparisons against several state-of-the-art (SOTA) methods on a cloudy image dataset demonstrated the superiority of PM-CycleGAN.
Yue Zi, Fengying Xie, Xuedong Song, Zhiguo Jiang 0001, Haopeng Zhang 0001
IEEE Geosci. Remote. Sens. Lett.2
2022 Dermoscopic image retrieval based on rotation-invariance deep hashing
Yilan Zhang, Fengying Xie, Xuedong Song, Yushan Zheng, Jie Liu 0085
Medical Image Anal.2
2022 Encoding histopathology whole slide images with location-aware graphs for diagnostically relevant regions retrieval
Yushan Zheng, Zhiguo Jiang 0001, Jun Shi 0006, Fengying Xie, Haopeng Zhang 0001, Dingyi Hu, Shujiao Sun, Zhongmin Jiang, Chenghai Xue
Medical Image Anal.4
2021 Single-shot weakly-supervised object detection guided by empirical saliency model
Danpei Zhao, Zhichao Yuan, Zhenwei Shi 0001, Fengying Xie
Neurocomputing4
2021 Stain Standardization Capsule for Application-Driven Histopathological Image Normalization
abstract
Color consistency is crucial to developing robust deep learning methods for histopathological image analysis. With the increasing application of digital histopathological slides, the deep learning methods are probably developed based on the data from multiple medical centers. This requirement makes it a challenging task to normalize the color variance of histopathological images from different medical centers. In this paper, we propose a novel color standardization module named stain standardization capsule based on the capsule network and the corresponding dynamic routing algorithm. The proposed module can learn and generate uniform stain separation outputs for histopathological images in various color appearance without the reference to manually selected template images. The proposed module is light and can be jointly trained with the application-driven CNN model. The proposed method was validated on three histopathology datasets and a cytology dataset, and was compared with state-of-the-art methods. The experimental results have demonstrated that the SSC module is effective in improving the performance of histopathological image analysis and has achieved the best performance in the compared methods.
Yushan Zheng, Zhiguo Jiang 0001, Haopeng Zhang 0001, Fengying Xie, Dingyi Hu, Shujiao Sun, Jun Shi 0006, Chenghai Xue
IEEE J. Biomed. Health Informatics4
2021 Diagnostic Regions Attention Network (DRA-Net) for Histopathology WSI Recommendation and Retrieval
abstract
The development of whole slide imaging techniques and online digital pathology platforms have accelerated the popularization of telepathology for remote tumor diagnoses. During a diagnosis, the behavior information of the pathologist can be recorded by the platform and then archived with the digital case. The browsing path of the pathologist on the WSI is one of the valuable information in the digital database because the image content within the path is expected to be highly correlated with the diagnosis report of the pathologist. In this article, we proposed a novel approach for computer-assisted cancer diagnosis named session-based histopathology image recommendation (SHIR) based on the browsing paths on WSIs. To achieve the SHIR, we developed a novel diagnostic regions attention network (DRA-Net) to learn the pathology knowledge from the image content associated with the browsing paths. The DRA-Net does not rely on the pixel-level or region-level annotations of pathologists. All the data for training can be automatically collected by the digital pathology platform without interrupting the pathologists' diagnoses. The proposed approaches were evaluated on a gastric dataset containing 983 cases within 5 categories of gastric lesions. The quantitative and qualitative assessments on the dataset have demonstrated the proposed SHIR framework with the novel DRA-Net is effective in recommending diagnostically relevant cases for auxiliary diagnosis. The MRR and MAP for the recommendation are respectively 0.816 and 0.836 on the gastric dataset. The source code of the DRA-Net is available at https://github.com/zhengyushan/dpathnet.
Yushan Zheng, Zhiguo Jiang 0001, Fengying Xie, Jun Shi 0006, Haopeng Zhang 0001, Jianguo Huai, Xiaomiao Yang
IEEE Trans. Medical Imaging3
2020 Investigating the Incorporation of Machine Learning Concepts in Data Structure Education
abstract
This Research to Practice Work-In-Progress paper discussed the incorporation of machine learning (ML) concepts in data structure education. The thriving of the ML especially deep learning techniques has led to an increased demand for trained professionals with ML skills to solve challenging engineering problems in many fields. Getting students familiar with ML as early as from CS2 (the data structure course) could benefit them in many aspects, but this direction has not been explored yet. In this paper, we discussed possible ways to integrate the ML concepts into data structure (DS) course. First, after introducing the concept of tensor in DS classroom teaching, we propose a practical experiment to implement the forward propagation of a pretrained convolutional neural network (CNN) aiming at classifying handwritten digits. Second, an experiment of decision tree based classification is set to give students an illuminating context via practicing the usage of tree structure. Finally, we design the experiment of computing graph to help the understanding of Directed Acyclic Graph (DAG), in which the students are required to implement the calculation of a multiple-variable function and its gradient based on DAG. Practicing DS knowledge in interesting ML-related problem contexts would intrigue the study enthusiasm of students and give them a general understanding of the application of DS knowledge in frontier technology, which could benefit the education of both DS and ML-related courses.
Bo Liu 0027, Fengying Xie
FIE2
2020 Tracing Diagnosis Paths on Histopathology WSIs for Diagnostically Relevant Case Recommendation
Yushan Zheng, Zhiguo Jiang 0001, Haopeng Zhang 0001, Fengying Xie, Jun Shi 0006
MICCAI (5)4
2020 Finding Arbitrary-Oriented Ships From Remote Sensing Images Using Corner Detection
abstract
Ship detection in remote sensing images is a challenging task. In this letter, a novel anchor-free framework is proposed for detecting arbitrary-oriented ships in remote sensing images. First, an end-to-end fully convolutional network is designed to detect the three key points, including the bow, stern, and center of the ship, as well as its angle. Second, the key points of the bow and stern are combined to generate possible rotated bounding boxes. Third, the predicted center and angle information of the ship are used to confirm the bounding box. In the designed network, feature fusion and feature enhancement modules are introduced to improve the performance in complex scenes. The proposed method avoids complicated anchor design compared with anchor-based methods. The experimental results show that with good robustness to haze occlusion, scale variation, and adjacent ship disturbances, our method outperforms other state-of-the-art methods.
Fengying Xie, Yuanyao Lu, Zhiguo Jiang 0001
IEEE Geosci. Remote. Sens. Lett.2
2020 A Local Flatness Based Variational Approach to Retinex
abstract
A topic of continued interest in Retinex over the years has been finding ways to implement it with computational models of improved accuracy and efficiency. We have devised a new approach to digitally implementing the Retinex using a local deviation based variational model. The new model leads to improvements in the computed image quality with respect to illumination correction and image enhancement. Several contributions are made: 1) a new prior constraint, which we call local flatness, is proposed, and a new measure of Local Deviation (LD) is developed to quantify the degree of local illumination flatness; 2) a variational problem is defined and the solution is found by a logical sequence of steps; 3) discrete implementation of the variational solution is shown to effectively estimate and remove uneven illumination, yielding an accurate recovered image. Unlike other physical prior based variational Retinex models, which use the L2 norm of the illumination gradient to enforce smoothness of illumination, our LD prior selectively imposes local flatness on illumination by calculating the deviation between the estimated illumination surface to a reference plane. In the experiments, pseudo ground truth images are created by superimposing uneven illumination on real scenes, providing an effective way to objectively assess algorithm performance. The experimental results show that our method can reconstruct more accurate recovered images than other state-of-the-art methods, while maintaining good contrast.
Fengying Xie, Rui Zhang 0069, Zhiguo Jiang 0001, Alan C. Bovik
IEEE Trans. Image Process.2
2019 A Comparative Study of CNN and FCN for Histopathology Whole Slide Image Analysis
Shujiao Sun, Bonan Jiang, Yushan Zheng, Fengying Xie
ICIG (2)4
2019 Encoding Histopathological WSIs Using GNN for Scalable Diagnostically Relevant Regions Retrieval
Yushan Zheng, Bonan Jiang, Jun Shi 0006, Haopeng Zhang 0001, Fengying Xie
MICCAI (1)5
2018 Size-Scalable Content-Based Histopathological Image Retrieval From Database That Consists of WSIs
abstract
Content-based image retrieval (CBIR) has been widely researched for histopathological images. It is challenging to retrieve contently similar regions from histopathological whole slide images (WSIs) for regions of interest (ROIs) in different size. In this paper, we propose a novel CBIR framework for database that consists of WSIs and size-scalable query ROIs. Each WSI in the database is encoded into a matrix of binary codes. When retrieving, a group of region proposals that have similar size with the query ROI are firstly located in the database through an efficient table-lookup approach. Then, these regions are ranked by a designed multi-binary-code-based similarity measurement. Finally, the top relevant regions and their locations in the WSIs as well as the corresponding diagnostic information are returned to assist pathologists. The effectiveness of the proposed framework is evaluated on a fine-annotated WSI database of epithelial breast tumors. The experimental results have proved that the proposed framework is effective for retrieval from database that consists of WSIs. Specifically, for query ROIs of 4096 4096 pixels, the retrieval precision of the top 20 return has reached 96% and the retrieval time is less than 1.5 s.
Yushan Zheng, Zhiguo Jiang 0001, Haopeng Zhang 0001, Fengying Xie, Yibing Ma, Huaqiang Shi, Yu Zhao 0029
IEEE J. Biomed. Health Informatics4
2018 Histopathological Whole Slide Image Analysis Using Context-Based CBIR
abstract
Histopathological image classification (HIC) and content-based histopathological image retrieval (CBHIR) are two promising applications for the histopathological whole slide image (WSI) analysis. HIC can efficiently predict the type of lesion involved in a histopathological image. In general, HIC can aid pathologists in locating high-risk cancer regions from a WSI by providing a cancerous probability map for the WSI. In contrast, CBHIR was developed to allow searches for regions with similar content for a region of interest (ROI) from a database consisting of historical cases. Sets of cases with similar content are accessible to pathologists, which can provide more valuable references for diagnosis. A drawback of the recent CBHIR framework is that a query ROI needs to be manually selected from a WSI. An automatic CBHIR approach for a WSI-wise analysis needs to be developed. In this paper, we propose a novel aided-diagnosis framework of breast cancer using whole slide images, which shares the advantages of both HIC and CBHIR. In our framework, CBHIR is automatically processed throughout the WSI, based on which a probability map regarding the malignancy of breast tumors is calculated. Through the probability map, the malignant regions in WSIs can be easily recognized. Furthermore, the retrieval results corresponding to each sub-region of the WSIs are recorded during the automatic analysis and are available to pathologists during their diagnosis. Our method was validated on fully annotated WSI data sets of breast tumors. The experimental results certify the effectiveness of the proposed method.
Yushan Zheng, Zhiguo Jiang 0001, Haopeng Zhang 0001, Fengying Xie, Yibing Ma, Huaqiang Shi, Yu Zhao 0029
IEEE Trans. Medical Imaging4
2017 No Reference Assessment of Image Visibility for Dehazing
Manjun Qin, Fengying Xie, Zhiguo Jiang 0001
ICIG (1)2
2017 Segmentation of dermoscopy images based on fully convolutional neural network
abstract
Lesion segmentation is one of the crucial steps for computerized dermoscopy image analysis. To accurately extract lesion borders from dermoscopy images, a novel segmentation method based on fully convolutional neural network is proposed in this paper. The designed network contains a low-level trunk followed by two brunches (global brunch and local brunch). The low-level trunk is fine-tuned from VGG16 net. Two brunches with different receptive fields extract global and local features respectively. After the combination of the global and local features, the final segmentation results are obtained through pixel-wise softmax classification. Experiments are conducted on the challenge dataset ISBI 2016. The results demonstrate that our designed network is more adaptive to dermoscopy images, which obtain more accurate lesion borders with good robust than other state-of-the-art methods.
Zilin Deng, Haidi Fan, Fengying Xie
ICIP3
2017 Kernel estimation for motion blur removal using deep convolutional neural network
abstract
Blind deblurring can restore the sharp image from the blur version when the blur kernel is unknown, which is a challenging task. Kernel estimation is crucial for blind deblurring. In this paper, a novel blur kernel estimation method based on regression model is proposed for motion blur. The motion blur features are firstly mined through convolutional neural network (CNN), and then mapped to motion length and orientation by support vector regression (SVR). Experiments show that the proposed model, namely CNNSVR, can give more accurate kernel estimation and generate better deblurring result compared with other state-of-the-art algorithms.
Yanan Lu, Fengying Xie, Zhiguo Jiang 0001
ICIP2
2017 Feature extraction from histopathological images based on nucleus-guided convolutional neural network for breast lesion classification
Yushan Zheng, Zhiguo Jiang 0001, Fengying Xie, Haopeng Zhang 0001, Yibing Ma, Huaqiang Shi, Yu Zhao 0029
Pattern Recognit.3
2017 Breast Histopathological Image Retrieval Based on Latent Dirichlet Allocation
abstract
In the field of pathology, whole slide image (WSI) has become the major carrier of visual and diagnostic information. Content-based image retrieval among WSIs can aid the diagnosis of an unknown pathological image by finding its similar regions in WSIs with diagnostic information. However, the huge size and complex content of WSI pose several challenges for retrieval. In this paper, we propose an unsupervised, accurate, and fast retrieval method for a breast histopathological image. Specifically, the method presents a local statistical feature of nuclei for morphology and distribution of nuclei, and employs the Gabor feature to describe the texture information. The latent Dirichlet allocation model is utilized for high-level semantic mining. Locality-sensitive hashing is used to speed up the search. Experiments on a WSI database with more than 8000 images from 15 types of breast histopathology demonstrate that our method achieves about 0.9 retrieval precision as well as promising efficiency. Based on the proposed framework, we are developing a search engine for an online digital slide browsing and retrieval platform, which can be applied in computer-aided diagnosis, pathology education, and WSI archiving and management.
Yibing Ma, Zhiguo Jiang 0001, Haopeng Zhang 0001, Fengying Xie, Yushan Zheng, Huaqiang Shi, Yu Zhao 0029
IEEE J. Biomed. Health Informatics4
2017 Melanoma Classification on Dermoscopy Images Using a Neural Network Ensemble Model
abstract
We develop a novel method for classifying melanocytic tumors as benign or malignant by the analysis of digital dermoscopy images. The algorithm follows three steps: first, lesions are extracted using a self-generating neural network (SGNN); second, features descriptive of tumor color, texture and border are extracted; and third, lesion objects are classified using a classifier based on a neural network ensemble model. In clinical situations, lesions occur that are too large to be entirely contained within the dermoscopy image. To deal with this difficult presentation, new border features are proposed, which are able to effectively characterize border irregularities on both complete lesions and incomplete lesions. In our model, a network ensemble classifier is designed that combines back propagation (BP) neural networks with fuzzy neural networks to achieve improved performance. Experiments are carried out on two diverse dermoscopy databases that include images of both the xanthous and caucasian races. The results show that classification accuracy is greatly enhanced by the use of the new border features and the proposed classifier model.
Fengying Xie, Haidi Fan, Zhiguo Jiang 0001, Rusong Meng, Alan C. Bovik
IEEE Trans. Medical Imaging1
2016 Cloud detection of remote sensing images by deep learning
abstract
Cloud detection plays a major role for remote sensing image processing. Most of the existed cloud detection methods use the low-level feature of the cloud, which often cause error result especially for thin cloud and complex scene. In this paper, a novel cloud detection method based on deep learning framework is proposed. The designed deep Convolutional Neural Networks (CNNs) consists of four convolutional layers and two fully-connected layers, which can mine the deep features of cloud. The image is firstly clustered into superpixels as sub-region through simple linear iterative cluster (SLIC) method. Through the designed network model, the probability of each superpixel that belongs to cloud region is predicted, so that the cloud probability map of the image is generated. Lastly, the cloud region is obtained according to the gradient of the cloud map. Through the proposed method, both thin cloud and thick cloud can be detected well, and the result is insensitive to complex scene. Experimental results indicate that the proposed method is more robust and effective than compared methods.
Mengyun Shi, Fengying Xie, Yue Zi, Jihao Yin
IGARSS2
2016 No-Reference Assessment on Haze for Remote-Sensing Images
abstract
Assessment on haze can filter out images with dense haze to improve the reliability of remote-sensing image interpretation. In this letter, a novel no-reference haze assessment method based on haze distribution is proposed for remote-sensing images. First, range channel of an image is defined and the haze distribution map (HDM) is extracted from the hazy image. Then, the haze assessment metric HDM-based haze assessment (HDMHA) is designed according to the HDM. Finally, the degree of haze in remote-sensing images is predicted using the proposed metric. In order to objectively verify the effectiveness of the proposed metric HDMHA, a method of simulating hazy remote-sensing images based on the haze imaging model is proposed in this letter, and the simulated hazy images are greatly similar to real ones in vision. A series of experiments are done on both real images and simulated images, and the results show that the proposed metric achieves good consistency when compared with subjective experiments and outperforms typical blind image quality assessment methods.
Xiaoxi Pan, Fengying Xie, Zhiguo Jiang 0001, Zhenwei Shi 0001, Xiaoyan Luo
IEEE Geosci. Remote. Sens. Lett.2
2015 Pattern Classification for Dermoscopic Images Based on Structure Textons and Bag-of-Features Model
Fengying Xie, Zhiguo Jiang 0001, Rusong Meng
ICIG (3)2
2015 Adaptive segmentation based on multi-classification model for dermoscopy images
Fengying Xie, Yefen Wu, Zhiguo Jiang 0001, Rusong Meng
Frontiers Comput. Sci.1
2015 No Reference Quality Assessment for Multiply-Distorted Images Based on an Improved Bag-of-Words Model
abstract
Multiple distortion assessment is a big challenge in image quality assessment (IQA). In this letter, a no reference IQA model for multiply-distorted images is proposed. The features, which are sensitive to each distortion type even in the presence of other distortions, are first selected from three kinds of NSS features. An improved Bag-of-Words (BoW) model is then applied to encode the selected features. Lastly, a simple yet effective linear combination is used to map the image features to the quality score. The combination weights are obtained through lasso regression. A series of experiments show that the feature selection strategy and the improved BoW model are effective in improving the accuracy of quality prediction for multiple distortion IQA. Compared with other algorithms, the proposed method delivers the best result for multiple distortion IQA.
Yanan Lu, Fengying Xie, Tongliang Liu, Zhiguo Jiang 0001, Dacheng Tao
IEEE Signal Process. Lett.2
2015 No Reference Uneven Illumination Assessment for Dermoscopy Images
abstract
For the dermoscopy image, uneven illumination will influence segmentation accuracy and lead to wrong aided diagnosis result. In this paper, a no reference uneven illumination assessment metric is proposed for dermoscopy images. Firstly, the distorted image is decomposed to illumination and reflectance components through variational framework for Retinex (VFR). Then, the illumination component is extracted by basis function fitting. Lastly, average gradient of the illumination component (AGIC) is calculated as the uneven illumination metric. A series of experiments show that, the proposed illumination extraction method is insensitive to the image content, and the proposed metric delivers an accurate illumination assessment result.
Yanan Lu, Fengying Xie, Yefen Wu, Zhiguo Jiang 0001, Rusong Meng
IEEE Signal Process. Lett.2
2015 Haze Removal for a Single Remote Sensing Image Based on Deformed Haze Imaging Model
abstract
The contrast of remote sensing images captured in haze condition is poor, which influences their interpretation. In this letter, a novel dehazing algorithm based on the deformed haze imaging model is proposed. First, the model is deformed by introducing a translation term. Second, the atmospheric light and transmission are estimated according to the new model combined with dark channel prior. Lastly, the haze is successfully removed from remote sensing images using the proposed estimation algorithm. The estimated transmission is insensitive to the texture of ground objects, and the dehazing effect for nonuniform haze is more satisfactory than the compared method. Moreover, our approach can be used for general haze removal through adjusting the translation term. Experimental results reveal that the proposed method can recover the real scene clearly from haze remote sensing images along with the advantage of good color consistency.
Xiaoxi Pan, Fengying Xie, Zhiguo Jiang 0001, Jihao Yin
IEEE Signal Process. Lett.2
2013 Automatic Skin Lesion Segmentation Based on Supervised Learning
abstract
The accuracy of automatic skin lesion detection is important in the computer-aided diagnosis (CAD) of skin cancers. In this paper, a novel method of automatic skin lesion segmentation to get the accurate border is proposed. The initial lesion is extracted by the Otsu's threshold firstly. Secondly, the outer peripheral region around the initial lesion is obtained with the affinity propagation clustering method (AP). The outer periphery is divided into small homogeneous sub-regions using simple linear iterative clustering (SLIC). Finally, the homogeneous sub-regions are classified into the background skin and lesion by supervised learning and the accuracy border is obtained. A series of experiments done on the proposed method and the other four state-of-the-art automatic methods show that the proposed method delivers better accuracy and robust segmentation results.
Yefen Wu, Fengying Xie, Zhiguo Jiang 0001, Rusong Meng
ICIG2
2013 Automatic segmentation of dermoscopy images using self-generating neural networks seeded by genetic algorithm
Fengying Xie, Alan C. Bovik
Pattern Recognit.1
2012 Automatic Skin Lesion Segmentation Based on Texture Analysis and Supervised Learning
Yingding He, Fengying Xie
ACCV (2)2