Yingying Fang

dblp:190/5759 · DBLP profile ↗
← Back
27ranked-venue papers
4as first author
23since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 2 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 3 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 1 first-author · 6 since 2021Systems, architecture and hardware · 1Security and privacy · 1
YearPublicationVenuePosition
2026 Physics-Aware Accelerated Unrolling Model for Sparse-View CT Reconstruction
abstract
Deep unrolling models (DUMs) have shown great poten-tial in sparse-view CT reconstruction by combining itera-tive optimization and deep learning. However, most DUMsinsufficiently account for physical degradation from sparse-view imaging, leading to slow convergence and persistentartifacts. To address this, we propose PAUM, a Physics-Aware Accelerated Unrolling Model explicitly incorporatingCT imaging physics into the iterative reconstruction. PAUMfirst introduces a Dual-Domain Physics-Aware Extrapolation(DDPE) module. By modeling dual-domain degradations, itperforms row-wise extrapolation in the sinogram domain toimprove missing view recovery, and pixel-wise extrapolationin the image domain to address spatially variant degradationfrom incomplete backprojection. This physics-aware extrap-olation aligns optimization dynamics with underlying physi-cal imaging degradation, significantly enhances structural up-dates, thereby accelerating convergence. Subsequently, wedevelop a lightweight Block-Attention Deformable Regu-larization Network (BDRN), leveraging deformable convo-lutions and block-wise attention to model spatially variantand structured artifact physical characteristics. This enablesspatially adaptive regularization on extrapolated results, ef-fectively improving reconstruction quality. Extensive exper-iments demonstrate PAUM achieves over 1dB improvementcompared to SOTA methods, while reducing iteration countby 50%.
Shaojie Guo, Yingying Fang, Junkang Zhang, Yan Wang 0033
AAAI2
2026 GEMA-Score: Granular Explainable Multi-Agent Scoring Framework for Radiology Report Evaluation
abstract
Automatic medical report generation has the potential to support clinical diagnosis, reduce the workload of radiologists, and demonstrate potential for enhancing diagnostic consistency. However, current evaluation metrics often fail to reflect the clinical reliability of generated reports. Overlap-based methods overlook fine-grained details (e.g., location, severity), diagnostic metrics are constrained by fixed vocabularies. Some diagnostic metrics are limited by fixed vocabularies or templates, reducing their ability to capture diverse clinical expressions. LLM-based metrics lack interpretable reasoning, limiting trust in clinical settings. Therefore, we propose a Granular Explainable Multi-Agent Score (GEMA-Score) in this paper, which conducts both objective quantification and subjective evaluation through a large language model-based multi-agent workflow. Our GEMA-Score parses structured reports and employs stable calculations through interactive exchanges of information among agents to assess disease diagnosis, location, severity, and uncertainty. Additionally, an LLM-based scoring agent evaluates completeness, readability, and clinical terminology while providing explanatory feedback. Extensive experiments show that GEMA-Score achieves the highest correlation with human experts on public datasets (Kendall = 0.69 on ReXVal; 0.45 on RadEvalX), demonstrating improved clinical scoring reliability.
Zhenxuan Zhang, Kinhei Lee, Peiyuan Jing, Weihang Deng, Huichi Zhou, Zihao Jin, Zhifan Gao, Dominic C. Marshall, Yingying Fang, Guang Yang 0006
AAAI10
2026 RPE-PAD: Relative Pose Estimation for Pose-agnostic Anomaly Detection
abstract
Pose-agnostic Anomaly Detection (PAD) aims to detect anomalies when the poses of query images are unknown and differ from those in the training set. Therefore, accurately estimating the camera poses for the query images in the test set is critical for this task. Existing query-specific framework methods require re-optimizing a new set of parameters for each query image, limiting their generalization and increasing computational burden. To overcome these limitations, we propose a novel method, Relative Pose Estimation for Pose-agnostic Anomaly Detection (RPE-PAD), which enhances both generalization and efficiency with a query-independent framework. Specifically, we propose a Random View Synthesis Scheme (RVSS) that generates new poses by adding Gaussian perturbations to the original poses, then renders the corresponding views to augment the dataset. To estimate the relative camera pose between two input images, we introduce an Iterative Relative Pose Refinement Network (IRPRN), which incorporates a hierarchical coarse-to-fine refinement strategy. Furthermore, we employ a Multi-Pair Training Strategy (MPTS) to train the proposed IRPRN, leveraging multiple image pairs to expand the relative pose transformation space during training. Extensive experiments demonstrate that our method achieves robust anomaly detection performance while significantly improving inference efficiency.
Mengzan Qi, Rongkang Ma, Yingying Fang, Guixu Zhang, Tieyong Zeng, Zhi Li 0080
AAAI4
2026 Reason like a radiologist: Chain-of-thought and reinforcement learning for verifiable report generation
abstract
Radiology report generation is critical for efficiency, but current models often lack the structured reasoning of experts and the ability to explicitly ground findings in anatomical evidence, which limits clinical trust and explainability. This paper introduces BoxMed-RL, a unified training framework to generate spatially verifiable and explainable chest X-ray reports. BoxMed-RL advances chest X-ray report generation through two integrated phases: (1) Pretraining Phase. BoxMed-RL learns radiologist-like reasoning through medical concept learning and enforces spatial grounding with reinforcement learning. (2) Downstream Adapter Phase. Pretrained weights are frozen while a lightweight adapter ensures fluency and clinical credibility. Experiments on two widely used public benchmarks (MIMIC-CXR and IU X-Ray) demonstrate that BoxMed-RL achieves an average 7 % improvement in both METEOR and ROUGE-L metrics compared to state-of-the-art methods. An average 5 % improvement in large language model-based metrics further underscores BoxMed-RL's robustness in generating high-quality reports. Related code and training templates are publicly available at https://github.com/ayanglab/BoxMed-RL.
Peiyuan Jing, Kinhei Lee, Zhenxuan Zhang, Huichi Zhou, Zhengqing Yuan, Zhifan Gao, Lei Zhu 0003, Giorgos Papanastasiou, Yingying Fang, Guang Yang 0006
Medical Image Anal.9
2026 Dynamical multi-order responses and global semantic-infused adversarial learning: A robust airway segmentation method
abstract
Automated airway segmentation in computerized tomography (CT) images is crucial for the accurate diagnosis of lung diseases. However, the scarcity of manual annotations hinders the efficacy of supervised learning, while unconstrained intensities and sample imbalance lead to discontinuity and false-negative issues. To address these challenges, we propose a novel airway segmentation model named Dynamical Multi-order responses and Global Semantic-infused Adversarial network (DMGSA), integrating the unsupervised and supervised learning in parallel to alleviate the label scarcity of airway. In the unsupervised branch, (1) we propose several novel strategies of Dynamic Mask-Ratio (DMR) to empower the model to perceive context information of varying sizes, mimicking the laws of human learning vividly; (2) we present a novel target of Multi-Order Normalized Responses (MONR), exploiting the distinct order exponential operation of raw images and oriented gradients to enhance the textural representations of bronchioles; (3) we introduce the Adversarial Learning (AL) on the top of MONR module to discern nuances between real and fake images, focusing on capturing the textural features of terminal bronchioles. For the supervised branch, we propose an innovative Generalized Mean pooling based Global Semantic-infused (GMGS) module to ulteriorly improve the robustness. Ultimately, we have verified the method performance and robustness by training on normal lung disease datasets, while testing on lung cancer, COVID-19 and Lung fibrosis datasets. All experimental results have proved that our method exceeds state-of-the-art methods significantly.
Sheng Zhang 0024, Yang Nan 0002, Yingying Fang, Yongkai Liu, Giorgos Papanastasiou, Zhifan Gao, Shuo Li 0001, Simon Walsh, Guang Yang 0006
Medical Image Anal.3
2025 Cyclic Vision-Language Manipulator: Towards Reliable and Fine-Grained Image Interpretation for Automated Report Generation
abstract
Despite significant advancements in automated report generation, the opaqueness of text interpretability continues to cast doubt on the reliability of the content produced. This paper introduces a novel approach to identify specific image features in X-ray images that influence the outputs of report generation models. Specifically, we propose Cyclic Vision-Language Manipulator (CVLM), a module to generate a manipulated X-ray from an original X-ray and its report from a designated report generator. The essence of CVLM is that cycling manipulated X-rays to the report generator produces altered reports aligned with the alterations pre-injected into the reports for X-ray generation, achieving the term ``cyclic manipulation''. This process allows direct comparison between original and manipulated X-rays, clarifying the critical image features driving changes in reports and enabling model users to assess the reliability of the generated texts. Empirical evaluations demonstrate that CVLM can identify more precise and reliable features compared to existing explanation methods, significantly enhancing the transparency and applicability of AI-generated reports.
Yingying Fang, Zihao Jin, Shaojie Guo, Jinda Liu, Zhiling Yue, Yijian Gao, Junzhi Ning, Simon Walsh, Guang Yang 0006
IJCAI1
2025 A Parallel Network for LRCT Segmentation and Uncertainty Mitigation with Fuzzy Sets
abstract
Accurate segmentation of airways in Low-Resolution CT (LRCT) scans is vital for diagnostics in scenarios such as reduced radiation exposure, emergency response, or limited resources. Yet manual annotation is labor-intensive and prone to variability, while existing automated methods often fail to capture small airway branches in lower-resolution 3D data. To address this, we introduce \textbf{FuzzySR}, a parallel framework that merges super-resolution (SR) and segmentation. By concurrently producing high-resolution reconstructions and precise airway masks, it enhances anatomic fidelity and captures delicate bronchi. FuzzySR employs a deep fuzzy set mechanism, leveraging learnable $t$-distribution and triangular membership functions via cross-attention. Through parameters $\mu$, $\sigma$, and $d_f$, it preserves uncertain features and mitigates boundary noise. Extensive evaluations on lung cancer, COVID-19, and pulmonary fibrosis datasets confirm FuzzySR’s superior segmentation accuracy on LRCT, surpassing even high-resolution baselines. By uniting fuzzy-logic-driven uncertainty handling with SR-based resolution enhancement, FuzzySR effectively bridges the gap for robust airway delineation from LRCT data.
Yang Nan 0002, Xiaodan Xing, Yingying Fang, Simon Walsh, Guang Yang 0006
UAI4
2025 A lung structure and function information-guided residual diffusion model for predicting idiopathic pulmonary fibrosis progression
Caiwen Jiang, Xiaodan Xing, Yang Nan 0002, Yingying Fang, Sheng Zhang 0024, Simon Walsh, Guang Yang 0006, Dinggang Shen
Medical Image Anal.4
2025 Unpaired translation of chest X-ray images for lung opacity diagnosis via adaptive activation masks and cross-domain alignment
abstract
Chest X-ray radiographs (CXRs) play a pivotal role in diagnosing and monitoring cardiopulmonary diseases. However, lung opacities in CXRs frequently obscure anatomical structures, impeding clear identification of lung borders and complicating localisation of pathology. This challenge significantly hampers segmentation accuracy and precise lesion identification, crucial for diagnosis. To tackle these issues, our study proposes an unpaired CXR translation framework that converts CXRs with lung opacities into counterparts without lung opacities while preserving semantic features. Central to our approach is the use of adaptive activation masks to selectively modify opacity regions in lung CXRs. Cross-domain alignment ensures translated CXRs without opacity issues align with feature maps and prediction labels from a pre-trained CXR lesion classifier, facilitating the interpretability of the translation process. We validate our method using RSNA, MIMIC-CXR-JPG and JSRT datasets, demonstrating superior translation quality through lower Fréchet Inception Distance (FID) and Kernel Inception Distance (KID) scores compared to existing methods (FID: 67.18 vs. 210.4, KID: 0.01604 vs. 0.225). Evaluation on RSNA opacity, MIMIC acute respiratory distress syndrome (ARDS) patient CXRs and JSRT CXRs shows our method enhances segmentation accuracy of lung borders and improves lesion classification, further underscoring its potential in clinical settings (RSNA: mIoU: 76.58% vs. 62.58%, Sensitivity: 85.58% vs. 77.03%; MIMIC ARDS: mIoU: 86.20% vs. 72.07%, Sensitivity: 92.68% vs. 86.85%; JSRT: mIoU: 91.08% vs. 85.6%, Sensitivity: 97.62% vs. 95.04%). Our approach advances CXR imaging analysis, especially in investigating segmentation impacts through image translation techniques. • Unpaired translation removes lung opacities yet keeps key features in X-rays. • Adaptive masks highlight and constrain opacity changes for better interpretability. • Cross-domain alignment reduces artefacts and preserves real diagnostic features. • Experiments show improved image fidelity, segmentation, and lesion classification.
Junzhi Ning, Dominic C. Marshall, Yijian Gao, Xiaodan Xing, Yang Nan 0002, Yingying Fang, Sheng Zhang 0024, Matthieu Komorowski, Guang Yang 0006
Pattern Recognit. Lett.6
2025 Affective Evaluation Based on Gamified Experience and Cascaded Fuzzy Reasoning for Wrist Rehabilitation
abstract
In order to establish the link between rehabilitation experience and patients' emotional feedback to assist decision-making in wrist rehabilitation program, this paper integrates gamified experience and proposes affective evaluation method based on cascaded fuzzy reasoning. Primarily, gamified interactive tasks are developed for wrist rehabilitation training, and the prototype of rehabilitation interactive system is built to collect multi-dimensional physiological signals (MPS). Subsequently, with MPS as input and arousal valence (AV) as output, fuzzy reasoning rule I is defined to establish MPS-AV model and calculate the value of AV. Besides, with AV as input and emotion label (EL) as output, fuzzy reasoning rule II is defined to establish AV-EL model and visually analyze the emotional state of patients. Lastly, the task completion rate of rehabilitation interaction behavior are calculated based on behavior coding, and the effectiveness of the proposed affective evaluation method is verified by integrating subjective evaluation. The gamified rehabilitation experience and the affective evaluation method based on cascaded fuzzy reasoning provide methods and experiences for wrist rehabilitation related products and services, and this protocol can also be generalized in the field of rehabilitation digital therapy.
Weihua Lu, Yingying Fang, Wenxin He, Yicha Zhang
IEEE Trans. Fuzzy Syst.2
2024 DiffExplainer: Unveiling Black Box Models Via Counterfactual Generation
Yingying Fang, Shuang Wu 0002, Zihao Jin, Caiwen Xu, Simon Walsh, Guang Yang 0006
MICCAI (10)1
2024 Diff3Dformer: Leveraging Slice Sequence Diffusion for Enhanced 3D CT Classification with Transformer Networks
Zihao Jin, Yingying Fang, Caiwen Xu, Simon Walsh, Guang Yang 0006
MICCAI (1)2
2024 Fuzzy Attention-Based Border Rendering Network for Lung Organ Segmentation
Sheng Zhang 0024, Yang Nan 0002, Yingying Fang, Xiaodan Xing, Zhifan Gao, Guang Yang 0006
MICCAI (9)3
2024 Purified Distillation: Bridging Domain Shift and Category Gap in Incremental Object Detection
abstract
Incremental Object Detection (IOD) simulates the dynamic data flow in real-world applications, which require detectors to learn new classes or adapt to new domains while retaining knowledge from previous tasks. Most existing IOD methods focus only on class incremental learning, assuming all data comes from the same domain. However, this is hardly achievable in practical applications, as images collected under different conditions often exhibit completely different characteristics, such as lighting, weather, style, etc. Class IOD methods suffer from performance degradation in these scenarios with domain shifts. To bridge domain shifts and category gaps in IOD, we propose Purified Distillation (PD), where we use a set of trainable queries to transfer the teacher's attention on old tasks to the student and adopt the gradient reversal layer to guide the student to learn the teacher's feature space structure from a micro perspective, which has not been extensively studied in previous works. Meanwhile, PD combines classification confidence with localization confidence to purify the most meaningful output nodes, so that the student model inherits a more comprehensive teacher knowledge. Extensive experiments across various IOD settings on six widely used datasets show that PD significantly outperforms state-of-the-art methods. Even after five steps of incremental learning, our method can preserve 60.6% mAP on the first task, while compared methods can only maintain up to 55.9%.
Shilong Jia, Tingting Wu 0001, Yingying Fang, Tieyong Zeng, Guixu Zhang, Zhi Li 0080
ACM Multimedia3
2024 Dynamic Multimodal Information Bottleneck for Multimodality Classification
abstract
Effectively leveraging multimodal data such as various images, laboratory tests and clinical information is becoming increasingly attractive in a variety of AI-based medical diagnosis and prognosis tasks. Most existing multi-modal techniques only focus on enhancing their performance by leveraging the differences or shared features from various modalities and fusing feature across different modalities. These approaches are generally not optimal for clinical settings, which pose the additional challenges of limited training data, as well as being rife with redundant data or noisy modality channels, leading to subpar performance. To address this gap, we study the robustness of existing methods to data redundancy and noise and propose a generalized dynamic multimodal information bottleneck framework for attaining a robust fused feature representation. Specifically, our information bottleneck module serves to filter out the task-irrelevant information and noises in the fused feature, and we further introduce a sufficiency loss to prevent dropping of task-relevant information, thus explicitly preserving the sufficiency of prediction information in the distilled feature. We validate our model on an in-house and a public COVID19 dataset for mortality prediction as well as two public biomedical datasets for diagnostic tasks. Extensive experiments show that our method surpasses the state-of-the-art and is significantly more robust, being the only method to remain performance when large-scale noisy channels exist. Our code is publicly available at https://github.com/ayanglab/DMIB.
Yingying Fang, Shuang Wu 0002, Sheng Zhang 0024, Chaoyan Huang, Tieyong Zeng, Xiaodan Xing, Simon Walsh, Guang Yang 0006
WACV1
2024 Probing perfection: The relentless art of meddling for pulmonary airway segmentation from HRCT via a human-AI collaboration based active learning method
abstract
In the realm of pulmonary tracheal segmentation, the scarcity of annotated data stands as a prevalent pain point in most medical segmentation endeavors. Concurrently, most Deep Learning (DL) methodologies employed in this domain invariably grapple with other dual challenges: the inherent opacity of 'black box' models and the ongoing pursuit of performance enhancement. In response to these intertwined challenges, the core concept of our Human-Computer Interaction (HCI) based learning models (RS_UNet, LC_UNet, UUNet and WD_UNet) hinge on the versatile combination of diverse query strategies and an array of deep learning models. We train four HCI models based on the initial training dataset and sequentially repeat the following steps 1-4: (1) Query Strategy: Our proposed HCI models selects those samples which contribute the most additional representative information when labeled in each iteration of the query strategy (showing the names and sequence numbers of the samples to be annotated). Additionally, in this phase, the model selects the unlabeled samples with the greatest predictive disparity by calculating the Wasserstein Distance, Least Confidence, Entropy Sampling, and Random Sampling. (2) Central line correction: The selected samples in previous stage are then used for domain expert correction of the system-generated tracheal central lines in each training round. (3) Update training dataset: When domain experts are involved in each epoch of the DL model's training iterations, they update the training dataset with greater precision after each epoch, thereby enhancing the trustworthiness of the 'black box' DL model and improving the performance of models. (4) Model training: Proposed HCI model is trained using the updated training dataset and an enhanced version of existing UNet. Experimental results validate the effectiveness of this Human-Computer Interaction-based approaches, demonstrating that our proposed WD-UNet, LC-UNet, UUNet, RS-UNet achieve comparable or even superior performance than the state-of-the-art DL models, such as WD-UNet with only 15 %-35 % of the training data, leading to substantial reductions (65 %-85 % reduction of annotation effort) in physician annotation time.
Yang Nan 0002, Sheng Zhang 0024, Federico Felder, Xiaodan Xing, Yingying Fang, Javier Del Ser, Simon Walsh, Guang Yang 0006
Artif. Intell. Medicine6
2024 Guest Editorial: Special Issue on the British Machine Vision Conference 2022
Guang Yang 0006, Angelica I. Avilés-Rivero, Yingying Fang, Zhenhua Feng 0001, Gianluigi Ciocca, Yulia Hicks, Constantino Carlos Reyes-Aldasoro
Int. J. Comput. Vis.3
2024 WeaFU: Weather-Informed Image Blind Restoration via Multi-Weather Distribution Diffusion
abstract
The extraction of distribution from images with diverse weather conditions is crucial for enhancing the robustness of visual algorithms. When addressing image degradation caused by different weather, accurately perceiving the data distribution of weather-informed degradation becomes a fundamental challenge. However, given the highly stochastic nature, modelling weather distribution poses a formidable task. In this paper, we propose a novel multi-Weather distribution difFUsion blind restoration model, named WeaFU. Firstly, the model employs representation learning to map image distribution into a latent space. Subsequently, WeaFU utilizes a diffusion-based approach, with the assistance of Diffusion Distribution Generator (DDG), to perceive and extract corresponding weather distribution. This strategy ingeniously injects data distribution into the recovery process, significantly enhancing the robustness of the model in diverse weather scenarios. Finally, a Conditional Distribution-Aware Transformer (CDAT) is constructed to align the distribution information with pixels, thereby obtaining clear images. Extensive experiments on real and synthetic datasets demonstrate that WeaFU achieves superior performance.
Bodong Cheng, Juncheng Li 0003, Jun Shi 0004, Yingying Fang, Guixu Zhang, Tieyong Zeng, Zhi Li 0080
IEEE Trans. Circuits Syst. Video Technol.4
2024 Fuzzy Attention-Based Border Rendering Orthogonal Network for Lung Organ Segmentation
abstract
Automatic lung organ segmentation on computerized tomography images is crucial for lung disease diagnosis. However, the unlimited voxel values and class imbalance of lung organs can lead to false-negative/positive and leakage issues in numerous state-of-the-art methods. In addition, some lung organs are easily lost during therecycleddown/up-sample procedure, e.g., bronchioles and arterioles, which can cause severe discontinuity issue. Inspired by these, this article introduces an effective lung organ segmentation method called fuzzy attention-based border rendering feature orthogonal network, which 1) integrates an efficient transformer-like fuzzy-attention module into deep networks to cope with the uncertainty in feature representations; 2) decouples and depicts the lung organ regions as cube-trees by focusing only onrecycle-sampling border vulnerable points, rendering the severely discontinuous, false-negative/positive organ regions with two novel global-local cube-tree fusion and sparse patched feature orthogonal modules; 3) develops a multiscale self-knowledge guidance module to improve model performance and robustness. We have demonstrated the efficacy of proposed method on five challenging datasets of lung organ segmentation, i.e., airway and artery. All experimental results demonstrate that our method can achieve the favorable performance significantly.
Sheng Zhang 0024, Yingying Fang, Yang Nan 0002, Weiping Ding 0001, Yew-Soon Ong, Alejandro F. Frangi, Witold Pedrycz, Simon Walsh, Guang Yang 0006
IEEE Trans. Fuzzy Syst.2
2024 Fuzzy Attention Neural Network to Tackle Discontinuity in Airway Segmentation
abstract
Airway segmentation is crucial for the examination, diagnosis, and prognosis of lung diseases, while its manual delineation is unduly burdensome. To alleviate this time-consuming and potentially subjective manual procedure, researchers have proposed methods to automatically segment airways from computerized tomography (CT) images. However, some small-sized airway branches (e.g., bronchus and terminal bronchioles) significantly aggravate the difficulty of automatic segmentation by machine learning models. In particular, the variance of voxel values and the severe data imbalance in airway branches make the computational module prone to discontinuous and false-negative predictions, especially for cohorts with different lung diseases. The attention mechanism has shown the capacity to segment complex structures, while fuzzy logic can reduce the uncertainty in feature representations. Therefore, the integration of deep attention networks and fuzzy theory, given by the fuzzy attention layer, should be an escalated solution for better generalization and robustness. This article presents an efficient method for airway segmentation, comprising a novel fuzzy attention neural network (FANN) and a comprehensive loss function to enhance the spatial continuity of airway segmentation. The deep fuzzy set is formulated by a set of voxels in the feature map and a learnable Gaussian membership function. Different from the existing attention mechanism, the proposed channel-specific fuzzy attention addresses the issue of heterogeneous features in different channels. Furthermore, a novel evaluation metric is proposed to assess both the continuity and completeness of airway structures. The efficiency, generalization, and robustness of the proposed method have been proved by training on normal lung disease while testing on datasets of lung cancer, COVID-19, and pulmonary fibrosis.
Yang Nan 0002, Javier Del Ser, Zeyu Tang 0001, Peng Tang 0004, Xiaodan Xing, Yingying Fang, Francisco Herrera, Witold Pedrycz, Simon Walsh, Guang Yang 0006
IEEE Trans. Neural Networks Learn. Syst.6
2022 Pixel screening based intermediate correction for blind deblurring
abstract
Blind deblurring has attracted much interest with its wide applications in reality. The blind deblurring problem is usually solved by estimating the intermediate kernel and the intermediate image alternatively, which will finally converge to the blurring kernel of the observed image. Numerous works have been proposed to obtain intermediate images with fewer undesirable artifacts by designing delicate regularization on the latent solution. However, these methods still fail while dealing with images containing saturations and large blurs. To address this problem, we propose an intermediate image correction method which utilizes Bayes posterior estimation to screen through the intermediate image and exclude those unfavorable pixels to reduce their influence for kernel estimation. Extensive experiments have proved that the proposed method can effectively improve the accuracy of the final derived kernel against the state-of-the-art methods on benchmark datasets by both quantitative and qualitative comparisons.
Meina Zhang, Yingying Fang, Guoxi Ni, Tieyong Zeng
CVPR2
2022 Swin transformer for fast MRI
abstract
Magnetic resonance imaging (MRI) is an important non-invasive clinical tool that can produce high-resolution and reproducible images. However, a long scanning time is required for high-quality MR images, which leads to exhaustion and discomfort of patients, inducing more artefacts due to voluntary movements of the patients and involuntary physiological movements. To accelerate the scanning process, methods by k-space undersampling and deep learning based reconstruction have been popularised. This work introduced SwinMR, a novel Swin transformer based method for fast MRI reconstruction. The whole network consisted of an input module (IM), a feature extraction module (FEM) and an output module (OM). The IM and OM were 2D convolutional layers and the FEM was composed of a cascaded of residual Swin transformer blocks (RSTBs) and 2D convolutional layers. The RSTB consisted of a series of Swin transformer layers (STLs). The shifted windows multi-head self-attention (W-MSA/SW-MSA) of STL was performed in shifted windows rather than the multi-head self-attention (MSA) of the original transformer in the whole image space. A novel multi-channel loss was proposed by using the sensitivity maps, which was proved to reserve more textures and details. We performed a series of comparative studies and ablation studies in the Calgary-Campinas public brain MR dataset and conducted a downstream segmentation experiment in the Multi-modal Brain Tumour Segmentation Challenge 2017 dataset. The results demonstrate our SwinMR achieved high-quality reconstruction compared with other benchmark methods, and it shows great robustness with different undersampling masks, under noise interruption and on different datasets. The code is publicly available at https://github.com/ayanglab/SwinMR.
Yingying Fang, Yinzhe Wu 0001, Huanjun Wu, Zhifan Gao, Yang Li 0010, Javier Del Ser, Jun Xia 0002, Guang Yang 0006
Neurocomputing2
2022 Quaternion Screened Poisson Equation for Low-Light Image Enhancement
abstract
Image enhancement is a technique to enhance the illumination of dark images while keeping the reality and the naturalness of the enhanced images at the same time. For color images, most methods tackle different color channels in a separate way, which overlooks the connection between the color channels. Therefore, in this paper, we consider a quaternion-based model to reserve the color connectivity, which integrates the color information of a pixel by a quaternion number. Moreover, we propose a regularizer based on the gamma-correction function and incorporate it into a screened Poisson equation for the image enhancement task. The uniqueness and existence of the solution of the proposed model are analyzed. The numerical results also prove the superiority of our scheme for color image enhancement.
Chaoyan Huang, Yingying Fang, Tingting Wu 0001, Tieyong Zeng, Yonghua Zeng
IEEE Signal Process. Lett.2
2020 Learning deep edge prior for image denoising
Yingying Fang, Tieyong Zeng
Comput. Vis. Image Underst.1
2020 PowerPrint: Identifying Smartphones through Power Consumption of the Battery
abstract
Device fingerprinting technologies are widely employed in smartphones. However, the features used in existing schemes may bring the privacy disclosure problems because of their fixed and invariable nature (such as IMEI and OS version), or the draconian of their experimental conditions may lead to a large reduction in practicality. Finding a new, secure, and effective smartphone fingerprint is, however, a surprisingly challenging task due to the restrictions on technology and mobile phone manufacturers. To tackle this challenge, we propose a battery-based fingerprinting method, named PowerPrint, which captures the feature of power consumption rather than invariable information of the battery. Furthermore, power consumption information can be easily obtained without strict conditions. We design an unsupervised learning-based algorithm to fingerprint the battery, which is stimulated with different power consumption of tasks to improve the performance. We use 15 smartphones to evaluate the performance of PowerPrint in both laboratory and public conditions. The experimental results indicate that battery fingerprint can be efficiently used to identify smartphones with low overhead. At the same time, it will not bring privacy problems, since the power consumption information is changing in real time.
Kun He 0008, Jing Chen 0003, Yingying Fang, Ruiying Du
Secur. Commun. Networks4
2019 Fast Color Blending for Seamless Image Stitching
abstract
In this letter, we propose a fast and robust method for stitching overlapped images captured by the unmanned aerial vehicle. First, we apply the shape-preserving half-projective method to precisely and stably align a pair of partially overlapped input images. Then, an optimal stitching line is searched to remove ghosts caused by the moving objects in the overlapped area. We subsequently propose a color blending method to eliminate all the color inconsistencies in the prealigned image. In accordance with the color differences of the pixels on the optimal stitching seam, we utilize weighted value coordinate interpolation algorithms to compute accurate color changes for all the pixels in the target image. The calculated color changes are then added to the target image to remove the color inconsistency. Furthermore, we introduce the superpixel segmentation to divide the target image into a reduced number of superpixels, and we assign each superpixel the same color change value. Such a superpixel level operation can greatly reduce the computational complexity. Experiments show that our method is promising to achieve effective and efficient stitching results.
Faming Fang, Tingting Wang 0007, Yingying Fang, Guixu Zhang
IEEE Geosci. Remote. Sens. Lett.3
2017 Charge-Depleting of the Batteries Makes Smartphones Recognizable
abstract
Many components of smartphones are used to generate device fingerprinting, such as screens, CPUs and various sensors. These device fingerprinting can be used to identify the smartphones. However, there are many restrictions with these device fingerprinting. Invariable information in screens and CPUs may lead to privacy risks. Moreover, strict experimental steps are required when fingerprinting the sensors. The effectiveness and effeciency of these device fingerprinting is reduced in practice. In this paper, we present a novel hardware fingerprinting based on the battery. Instead of relying on invariable information of the battery, we focus on the charge-depleting of the smartphone. The discrepencies on manufacturing of smartphones make that the charge-depleting is different when performs the same task. Moreover, charge-depleting information can easily be obtained without strict operating steps. We design a highly accurate algorithm to fingerprint the batteries which is based on the unsupervised learning. Besides, we stimulate the algorithm with different charge-depleting of tasks to improve the performance. We use 15 smartphones to evaluate the performance of the battery fingerprinting in both laboratory and public conditions. The experimental results show that battery fingerprinting is quite effective, the recognition accuracy rate can reach 86%.
Jing Chen 0003, Yingying Fang, Kun He 0008, Ruiying Du
ICPADS2