EDBT 2026 Demo / reviewers in the wild / expert
Zhenqi Fu
dblp:216/7010
· DBLP profile ↗
24ranked-venue papers
6as first author
24since 2021 · last 2026
0000-0003-2950-7190ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 20 · 6 first-author · 20 since 2021Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MCSF-Net: A Multi-Color Space Fusion Network for Underwater Image EnhancementabstractExisting multi-color space guided techniques for underwater image enhancement (UIE) fail to take the advantages of the XYZ color space for preserving underwater image details, meanwhile, existing UIE datasets, typically containing low-quality reference images of distorted colors and blurred structures, lead to inaccurate enhancement mapping between low-quality and high-quality images. To overcome these above limitations, we propose a Multi-Color Space Fusion Network (MCSF-Net) for UIE. The MCSF-Net incorporates a Multi-dimensional Feature Fusion Block (MFFB) and weighted feature fusion scheme to effectively integrate complementary features from both XYZ and RGB color spaces. Moreover, we establish a Large-Scale Mixed UIE dataset (LSMU) by using nine no-reference metrics to filter out low-quality reference images from eight public UIE datasets, enabling more effective network learning. Extensive experiments on mainstream datasets demonstrate that the proposed method outperforms several leading approaches in both color restoration and detail enhancement of various underwater images. The code and dataset for MCSF-Net will be available athttps://github.com/WYJGR/MCSF-Net. Yijian Wang, Peixian Zhuang, Zhenqi Fu, Jiaquan Yan |
IEEE Trans. Multim. | 3 |
| 2026 | Accelerating Adaptive Diffusion and Uncertainty Modeling for Underwater Image EnhancementabstractUnderwater image enhancement (UIE) aims to mitigate wavelength-dependent absorption and multi-path scattering effects, enabling the recovery of natural colors and rich details. Despite notable progress, consistently achieving high-quality enhancement in both fidelity and perceptual clarity remains a fundamental challenge. To address this, we propose the Laplacian domain Dual-Focus Enhancer (DFE), an innovative framework consisting of two stages: adaptive diffusion-accelerated low frequency enhancement (ADALE) and progressive uncertainty driven high-frequency enhancement (PUHE). Specifically, DFE applies a Laplacian transform to decouple the frequency-specific degradations in underwater images, supporting fidelity- and clarity-oriented enhancement along separate pathways. To facilitate high-fidelity restoration, ADALE incorporates an HSV guided optimization mechanism (HSV-OM) to establish a robust color and brightness calibration baseline for the low-frequency diffusion model, adaptively managing basic degradations with minimal sampling steps. Furthermore, to enhance contour and detail perception, PUHE models the uncertainty of reference textures and integrates it with feature modulation to progressively reconstruct multi-scale high-frequency structures. The multi reference underwater texture enhancement (MUTE) dataset fur ther improves image clarity. Extensive experiments demonstrate that our DFE outperforms state-of-the-art (SOTA) methods in both quantitative metrics and visual quality. Xiuna Zeng, Jiaao Peng, Zhenqi Fu, Linyu Fan, Xiaotong Tu, Yue Huang 0001, Xinghao Ding |
IEEE Trans. Multim. | 3 |
| 2025 | AGLLDiff: Guiding Diffusion Models Towards Unsupervised Training-free Real-world Low-light Image EnhancementabstractExisting low-light image enhancement (LIE) methods have achieved noteworthy success in solving synthetic distortions, yet they often fall short in practical applications. The limitations arise from two inherent challenges in real-world LIE: 1) the collection of distorted/clean image pairs is often impractical and sometimes even unavailable, and 2) accurately modeling complex degradations presents a non-trivial problem. To overcome them, we propose the Attribute Guidance Diffusion framework (AGLLDiff), a training-free method for effective real-world LIE. Instead of specifically defining the degradation process, AGLLDiff shifts the paradigm and models the desired attributes, such as image exposure, structure and color of normal-light images. These attributes are readily available and impose no assumptions about the degradation process, which guides the diffusion sampling process to a reliable high-quality solution space. Extensive experiments demonstrate that our approach outperforms the current leading unsupervised LIE methods across benchmarks in terms of distortion-based and perceptual-based metrics, and it performs well even in sophisticated wild degradation. Yunlong Lin, Tian Ye 0001, Sixiang Chen, Zhenqi Fu, Yingying Wang 0005, Wenhao Chai, Zhaohu Xing, Wenxue Li 0003, Lei Zhu 0003, Xinghao Ding |
AAAI | 4 |
| 2025 | DPLUT: Unsupervised Low-light Image Enhancement with Lookup Tables and Diffusion PriorsabstractLow-light image enhancement (LIE) aims at precisely and efficiently recovering an image degraded in poor illumination environments. Recent advanced LIE techniques are using deep neural networks, which require lots of low-normal light image pairs, network parameters, and computational resources. As a result, their practicality is limited. In this work, we devise a novel unsupervised LIE framework based on diffusion priors and lookup tables (DPLUT) to achieve efficient low-light image recovery. The proposed approach comprises two critical components: a light adjustment lookup table (LLUT) and a noise suppression lookup table (NLUT). LLUT is optimized with a set of unsupervised losses. It aims at predicting pixel-wise curve parameters for the dynamic range adjustment of a specific image. NLUT is designed to remove the amplified noise after the light brightens. As diffusion models are sensitive to noise, diffusion priors are introduced to achieve high-performance noise suppression. Extensive experiments demonstrate that our approach outperforms state-of-the-art methods in terms of visual quality and efficiency. Yunlong Lin, Zhenqi Fu, Kairun Wen, Tian Ye 0001, Sixiang Chen, Ge Meng, Yingying Wang 0005, Chui Kong, Yue Huang 0001, Xiaotong Tu, Xinghao Ding |
AAAI | 2 |
| 2025 | CyclicAligner: Knowledge-Enhanced Cyclical Alignment for Chest X-Ray Report GenerationabstractTo reduce the diagnostic burden on radiologists, recent studies have explored automatic chest X-ray (CXR) report generation via artificial intelligence. Yet, achieving robust cross-modal alignment between medical images and textual reports remains a major challenge. In this paper, we propose CyclicAligner, a knowledge-enhanced cyclical alignment framework for CXR report generation. CyclicAligner adopts a novel cyclical training paradigm with four tightly coupled tasks to effectively learn cross-modal semantic alignment: (1) an image-to-text generation task that aligns visual semantics with clinical findings, (2) a text-to-text reconstruction task that strengthens language modeling, (3) a hybrid-to-text reconstruction task that mixes vision and language tokens for text reconstruction, and (4) a traceback-alignment task that re-encodes texts generated by the image-to-text branch for text reconstruction and aligns the reconstructed text with the reference. To further enhance cross-modal understanding, we integrate domain-specific medical entity knowledge extracted from a pre-trained encoder to enrich both vision and language tokens. Moreover, CyclicAligner jointly predicts medical tags and narrative reports within a unified auto-regressive pipeline, where the tags serve as auxiliary semantic anchors that guide the report generation. Extensive experiments on public datasets demonstrate the effectiveness of our method for clinical-coherent CXR report generation. The related code is available at https://github.com/yangyan22/CyclicAligner. Jiamei Sun, Ke Zhang 0029, Xiangyu Tan, Zhenqi Fu |
BIBM | 5 |
| 2025 | V2V3D: View-to-View Denoised 3D Reconstruction for Light Field MicroscopyabstractLight field microscopy (LFM) has gained significant attention due to its ability to capture snapshot-based, large-scale 3D fluorescence images. However, existing LFM reconstruction algorithms are highly sensitive to sensor noise or require hard-to-get ground-truth annotated data for training. To address these challenges, this paper introduces V2V3D, an unsupervised view2view-based framework that establishes a new paradigm for joint optimization of image denoising and 3D reconstruction in a unified architecture. We assume that the LF images are derived from a consistent 3D signal, with the noise in each view being independent. This enables V2V3D to incorporate the principle of noise2noise for effective denoising. To enhance the recovery of high-frequency details, we propose a novel wave-optics-based feature alignment technique, which transforms the point spread function, used for forward propagation in wave optics, into convolution kernels specifically designed for feature alignment. Moreover, we introduce an LFM dataset containing LF images and their corresponding 3D intensity volumes. Extensive experiments demonstrate that our approach achieves high computational efficiency and outperforms the other state-of-the-art methods. These advancements position V2V3D as a promising solution for 3D imaging under challenging conditions. Our code and dataset will be publicly accessible at https://joey1998hub.github.io/V2V3D/. Jiayin Zhao, Zhenqi Fu, Tao Yu 0007 |
CVPR | 2 |
| 2025 | Spatio-Temporal and Retrieval-Augmented Modeling for Chest X-Ray Report GenerationabstractChest X-ray report generation has attracted increasing research attention. However, most existing methods neglect the temporal information and typically generate reports conditioned on a fixed number of images. In this paper, we propose STREAM: Spatio-Temporal and REtrieval-Augmented Modelling for automatic chest X-ray report generation. It mimics clinical diagnosis by integrating current and historical studies to interpret the present condition (temporal), with each study containing images from multi-views (spatial). Concretely, our STREAM is built upon an encoder-decoder architecture, utilizing a large language model (LLM) as the decoder. Overall, spatio-temporal visual dynamics are packed as visual prompts and regional semantic entities are retrieved as textual prompts. First, a token packer is proposed to capture condensed spatio-temporal visual dynamics, enabling the flexible fusion of images from current and historical studies. Second, to augment the generation with existing knowledge and regional details, a progressive semantic retriever is proposed to retrieve semantic entities from a preconstructed knowledge bank as heuristic text prompts. The knowledge bank is constructed to encapsulate anatomical chest X-ray knowledge into structured entities, each linked to a specific chest region. Extensive experiments on public datasets have shown the state-of-the-art performance of our method. Related codes and the knowledge bank are available at https://github.com/yangyan22/STREAM. Xiaoxing You, Ke Zhang 0029, Zhenqi Fu, Xianyun Wang, Jiajun Ding, Jiamei Sun, Zhou Yu 0001, Qingming Huang, Weidong Han 0001, Jun Yu 0002 |
IEEE Trans. Medical Imaging | 4 |
| 2024 | Progressive High-Frequency Reconstruction for Pan-Sharpening with Implicit Neural RepresentationabstractPan-sharpening aims to leverage the high-frequency signal of the panchromatic (PAN) image to enhance the resolution of its corresponding multi-spectral (MS) image. However, deep neural networks (DNNs) tend to prioritize learning the low-frequency components during the training process, which limits the restoration of high-frequency edge details in MS images. To overcome this limitation, we treat pan-sharpening as a coarse-to-fine high-frequency restoration problem and propose a novel method for achieving high-quality restoration of edge information in MS images. Specifically, to effectively obtain fine-grained multi-scale contextual features, we design a Band-limited Multi-scale High-frequency Generator (BMHG) that generates high-frequency signals from the PAN image within different bandwidths. During training, higher-frequency signals are progressively injected into the MS image, and corresponding residual blocks are introduced into the network simultaneously. This design enables gradients to flow from later to earlier blocks smoothly, encouraging intermediate blocks to concentrate on missing details. Furthermore, to address the issue of pixel position misalignment arising from multi-scale features fusion, we propose a Spatial-spectral Implicit Image Function (SIIF) that employs implicit neural representation to effectively represent and fuse spatial and spectral features in the continuous domain. Extensive experiments on different datasets demonstrate that our method outperforms existing approaches in terms of quantitative and visual measurements for high-frequency detail recovery. Ge Meng, Jingjia Huang, Yingying Wang 0005, Zhenqi Fu, Xinghao Ding, Yue Huang 0001 |
AAAI | 4 |
| 2024 | Token-Mixer: Bind Image and Text in One Embedding Space for Medical Image ReportingabstractMedical image reporting focused on automatically generating the diagnostic reports from medical images has garnered growing research attention. In this task, learning cross-modal alignment between images and reports is crucial. However, the exposure bias problem in autoregressive text generation poses a notable challenge, as the model is optimized by a word-level loss function using the teacher-forcing strategy. To this end, we propose a novel Token-Mixer framework that learns to bind image and text in one embedding space for medical image reporting. Concretely, Token-Mixer enhances the cross-modal alignment by matching image-to-text generation with text-to-text generation that suffers less from exposure bias. The framework contains an image encoder, a text encoder and a text decoder. In training, images and paired reports are first encoded into image tokens and text tokens, and these tokens are randomly mixed to form the mixed tokens. Then, the text decoder accepts image tokens, text tokens or mixed tokens as prompt tokens and conducts text generation for network optimization. Furthermore, we introduce a tailored text decoder and an alternative training strategy that well integrate with our Token-Mixer framework. Extensive experiments across three publicly available datasets demonstrate Token-Mixer successfully enhances the image-text alignment and thereby attains a state-of-the-art performance. Related codes are available at https://github.com/yangyan22/Token-Mixer. Jun Yu 0002, Zhenqi Fu, Ke Zhang 0029, Ting Yu 0016, Xianyun Wang, Hanliang Jiang, Junhui Lv, Qingming Huang, Weidong Han 0001 |
IEEE Trans. Medical Imaging | 3 |
| 2023 | Learning a Simple Low-Light Image Enhancer from Paired Low-Light InstancesabstractLow-light Image Enhancement (LIE) aims at improving contrast and restoring details for images captured in lowlight conditions. Most of the previous LIE algorithms adjust illumination using a single input image with several handcrafted priors. Those solutions, however, often fail in revealing image details due to the limited information in a single image and the poor adaptability of handcrafted priors. To this end, we propose PairLIE, an unsupervised approach that learns adaptive priors from low-light image pairs. First, the network is expected to generate the same clean images as the two inputs share the same image content. To achieve this, we impose the network with the Retinex theory and make the two reflectance components consistent. Second, to assist the Retinex decomposition, we propose to remove inappropriate features in the raw image with a simple self-supervised mechanism. Extensive experiments on public datasets show that the proposed PairLIE achieves comparable performance against the state-of-the-art approaches with a simpler network and fewer handcrafted priors. Code is available at: https://github.com/zhenqifu/PairLIE. Zhenqi Fu, Xiaotong Tu, Yue Huang 0001, Xinghao Ding, Kai-Kuang Ma |
CVPR | 1 |
| 2023 | Underwater Image Enhancement and Super-Resolution Using Implicit Neural NetworksabstractUnderwater images are often notably degraded by light scattering and absorption. To improve image quality and object details, we present a novel unsupervised underwater image enhancement and super-resolution method using implicit neural networks. Concretely, taking low-resolution coordinates as the inputs, we first leverage Fourier feature mapping to encode the coordinates. Then, three implicit neural networks are applied to estimate each component (i.e., the global background light, the transmission map, and the scene radiance) of the underwater formation model. Those components are further used to reconstruct the raw underwater image in a self-supervised fashion. In the inference stage, high-resolution coordinates are employed to predict a high-quality and high- resolution underwater image. Extensive experiments show that our method achieves a favorable performance in terms of both super-resolution and quality enhancement as compared with current approaches. Xueye Chu, Zhenqi Fu, Shaocong Yu, Xiaotong Tu, Yue Huang 0001, Xinghao Ding |
ICIP | 2 |
| 2023 | Domain-irrelevant Feature Learning for Generalizable Pan-sharpeningabstractPan-sharpening aims to spatially enhance the low-resolution multispectral image (LRMS) by transferring high-frequency details from a panchromatic image (PAN) while preserving the spectral characteristics of LRMS. Previous arts mainly focus on how to learn a high-resolution multispectral image (HRMS) on the i.i.d. assumption. However, the distribution of training and testing data often encounters significant shifts in different satellites. To this end, this paper proposes a generalizable pan-sharpening network via domain-irrelevant feature learning. On the one hand, a structural preservation module (STP) is designed to fuse high-frequency information of PAN and LRMS. Our STP is performed on the gradient domain because it consists of structure and texture details that can generalize well on different satellites. On the other hand, to avoid spectral distortion while promoting the generalization ability, a spectral preservation module (SPP) is developed. The key design of SPP is to learn a phase fusion network of PAN and LRMS. The amplitude of LRMS, which contains 'satellite style' information is directly injected in different fusion stages. Extensive experiments have demonstrated the effectiveness of our method against state-of-the-art methods in both single-satellite and cross-satellite scenarios. Code is available at: https://github.com/LYL1015/DIRFL. Yunlong Lin, Zhenqi Fu, Ge Meng, Yingying Wang 0005, Linyu Fan, Hedeng Yu, Xinghao Ding |
ACM Multimedia | 2 |
| 2023 | Learning High-frequency Feature Enhancement and Alignment for Pan-sharpeningabstractPan-sharpening aims to utilize the high-resolution panchromatic (PAN) image as a guidance to super-resolve the spatial resolution of the low-resolution multispectral (MS) image. The key challenge in pan-sharpening is how to effectively and precisely inject high-frequency edges and textures from the PAN image into the low-resolution MS image. To address this issue, we propose a High-frequency Feature Enhancement and Alignment Network (HFEAN) for effectively encouraging the high-frequency learning. To implement it, three core designs are customized: a Fourier convolution based efficient feature enhancement module (FEM), an implicit neural alignment module (INA), and a preliminary alignment module (Pre-align). To be specific, FEM employs the fast Fourier convolution with attention mechanism to achieve the mixed global-local receptive field on each scale of the high-frequency domain, thus yielding the informative latent codes. INA leverages implicit neural function to precisely align the latent codes from different scales in the continuous domain. In this way, the high frequency signals at different scales are represented as functions of continuous coordinates, enabling a precise feature alignment in a resolution-free manner. Pre-align is developed to further address the inherent misalignment between PAN and MS pairs. Extensive experiments over multiple satellite datasets validate the effectiveness of the proposed network and demonstrate its favorable performance against the existing state-of-the-art methods both visually and quantitatively. Code is available at: https://github.com/Gracewangyy/HFEAN. Yingying Wang 0005, Yunlong Lin, Ge Meng, Zhenqi Fu, Linyu Fan, Hedeng Yu, Xinghao Ding, Yue Huang 0001 |
ACM Multimedia | 4 |
| 2023 | Infrared and Visible Image Fusion via Test-Time Training
Guoqing Zheng, Zhenqi Fu, Xiaopeng Lin, Xueye Chu, Yue Huang 0001, Xinghao Ding |
PRCV (10) | 2 |
| 2022 | Unsupervised Underwater Image Restoration: From a Homology PerspectiveabstractUnderwater images suffer from degradation due to light scattering and absorption. It remains challenging to restore such degraded images using deep neural networks since real-world paired data is scarcely available while synthetic paired data cannot approximate real-world data perfectly. In this paper, we propose an UnSupervised Underwater Image Restoration method (USUIR) by leveraging the homology property between a raw underwater image and a re-degraded image. Specifically, USUIR first estimates three latent components of the raw underwater image, i.e., the global background light, the transmission map, and the scene radiance (the clean image). Then, a re-degraded image is generated by randomly mixing up the estimated scene radiance and the raw underwater image. We demonstrate that imposing a homology constraint between the raw underwater image and the re-degraded image is equivalent to minimizing the restoration error and hence can be used for the unsupervised restoration. Extensive experiments show that USUIR achieves promising performance in both inference time and restoration quality. Zhenqi Fu, Huangxing Lin, Shu Chai, Liyan Sun, Yue Huang 0001, Xinghao Ding |
AAAI | 1 |
| 2022 | EffiSeaNet: Pioneering Lightweight Network for Underwater Salient Object Detection
Qingyao Wu, Zhenqi Fu, Chenyu Ma, Xiaotong Tu, Xinghao Ding |
ACCV (4) | 2 |
| 2022 | Uncertainty Inspired Underwater Image Enhancement
Zhenqi Fu, Yue Huang 0001, Xinghao Ding, Kai-Kuang Ma |
ECCV (18) | 1 |
| 2022 | Unsupervised and Untrained Underwater Image Restoration Based on Physical Image Formation ModelabstractUnderwater images suffer from degradation caused by light scattering and absorption. Training a deep neural network to restore underwater images is challenging due to the labor-intensive data collection and the lack of paired data. To this end, we propose an unsupervised and untrained underwater image restoration method based on the layer disentanglement and the underwater image formation model. Specifically, our network disentangles an underwater image into four components, i.e., the scene radiance, the direct transmission map, the backscatter transmission map, and the global background light, which are further combined to reconstruct the underwater image in a self-supervised manner. Our method can avoid using paired training data and large-scale datasets, benefiting from the unsupervised and untrained characteristics. Extensive experiments demonstrated that our method obtains promising performance compared with six methods on three real-world underwater image databases. Shu Chai, Zhenqi Fu, Yue Huang 0001, Xiaotong Tu, Xinghao Ding |
ICASSP | 2 |
| 2022 | A Robust Object Segmentation Network for UnderWater ScenesabstractUnderwater object segmentation is one of the key technologies in the fields of marine biology research and autonomous underwater vehicles. The challenges of underwater object segmentation originate from two aspects, 1) the complex underwater environment and 2) the camouflage characteristics of marine animals. In this paper, we propose WaterSNet, an underwater object segmentation network to address these challenges. Specified, we propose a random style adaption (RSA) module as well as a siamese structure to reduce the impact of water degradation diversity. We also extract multi-scale features via the receptive field block (RFB) module, and then fuses multi-level features to better utilize global context information via the attention fusion block (AFB) module. Experimental results on marine animal dataset MAS3K demonstrate that the proposed method outperforms other state-of-the-art methods significantly. The code will be available at: https://github.com/ruizhechen/WaterSNet/ Ruizhe Chen, Zhenqi Fu, Yue Huang 0001, En Cheng, Xinghao Ding |
ICASSP | 2 |
| 2022 | Underwater Image Enhancement Via Learning Water Type Desensitized RepresentationsabstractWe present a novel underwater image enhancement method termed SCNet to improve the image quality meanwhile cope with the degradation diversity caused by the water. SCNet is based on normalization schemes across both spatial and channel dimensions with the key idea of learning water type desensitized features. Specifically, we apply whitening to de-correlate activations across spatial dimensions for each instance in a mini-batch. We also eliminate channel-wise correlation by standardizing and re-injecting the first two moments of the activations across channels. The normalization schemes of spatial and channel dimensions are performed at each scale of the U-Net to obtain multi-scale representations. With such water type irrelevant encodings, the decoder can easily reconstruct the clean signal and be unaffected by the distortion types. Experimental results on two real-world underwater image datasets show that our approach can successfully enhance images with diverse water types, and achieves competitive performance in visual quality improvement. Zhenqi Fu, Xiaopeng Lin, Yue Huang 0001, Xinghao Ding |
ICASSP | 1 |
| 2022 | Twice Mixing: A rank learning based quality assessment approach for underwater image enhancement
Zhenqi Fu, Xueyang Fu, Yue Huang 0001, Xinghao Ding |
Signal Process. Image Commun. | 1 |
| 2022 | Combining Retargeting Quality and Depth Perception Measures for Quality Evaluation of Retargeted StereopairsabstractStereoscopic Image Retargeting (SIR) aims to adapt stereoscopic images and videos to 3D display devices with various aspect ratios by emphasizing the important content while retaining surrounding context with minimal visual distortion. To address the issue of SIR evaluation, this paper presents a new objective quality assessment method for retargeted stereopairs by combining image quality and depth perception measures. Specifically, the image quality measure is conducted between the source and retargeted intermediate views generated by the view synthesis method to characterize the geometric distortion and content loss of the retargeted stereopair, while several depth-aware features are extracted to measure the visual comfort/discomfort and depth sensation when human views a 3D scene. Then, the extracted features are integrated into an overall perceptual quality prediction. Experiment results on NBU SIRQA and SIRD databases verify the superiority of our method. Xuejin Wang, Feng Shao 0001, Qiuping Jiang, Zhenqi Fu, Xiangchao Meng, Ke Gu 0001, Yo-Sung Ho |
IEEE Trans. Multim. | 4 |
| 2021 | Subjective and Objective Quality Assessment for Stereoscopic Image RetargetingabstractBinocular stereoscopic image retargeting (SIR) aims to adjust 3D images into target aspect ratios. In recent years, various SIR methods have been proposed, but there are few researches on visual quality assessment. As a consequence, we construct a benchmark stereoscopic image retargeting quality assessment database (NBU-SIRQA), which contains 720 stereoscopic retargeted images generated by eight representative SIR operators. Subjective test is conducted to obtain the mean opinion score (MOS) for each stereoscopic retargeted image. Additionally, we propose an objective SIRQA metric based on grid deformation and information loss (GDIL). The main idea of GDIL is to decompose the SIR operator into two transformations: monocular image retargeting transformation and viewpoint transformation. In each transformation, grid deformation and information loss are extracted simultaneously to represent image quality and 3D perception quality. Experimental results validated on our established NBU-SIRQA database show the superiority of our metric in measuring the quality of stereoscopic retargeted images over the existing approaches. Zhenqi Fu, Feng Shao 0001, Qiuping Jiang, Xiangchao Meng, Yo-Sung Ho |
IEEE Trans. Multim. | 1 |
| 2021 | Transformation-Aware Similarity Measurement for Image Retargeting Quality Assessment via Bidirectional RewarpingabstractImage retargeting is an effective way to adapt images for target displays with different aspect ratios and sizes. Meanwhile, effective image retargeting quality assessment (IRQA) is important for optimizing the image retargeting operations. In this paper, we propose a transform-aware similarity (TRASIM) measurement metric for IRQA, including bidirectional geometric distortion measurement, bidirectional information loss measurement, and global salient structure distortion measurement. The main innovation of the TRASIM is to build a universal framework to establish the similarity transformation via bidirectional rewarping to simulate different types of retargeting operators. Based on the similarity transformation, geometric distortion and content loss are measured to determine the retargeting quality. Experimental results on two widely used databases (CUHK and RetargetMe) indicate that the proposed TRASIM has higher consistency with subjective ranks, compared with the state-of-the-art IRQA metrics. Feng Shao 0001, Zhenqi Fu, Qiuping Jiang, Gangyi Jiang, Yo-Sung Ho |
IEEE Trans. Syst. Man Cybern. Syst. | 2 |