VLDB 2026 Research / reviewers in the wild / expert
Gangyi Jiang
dblp:69/1799
· DBLP profile ↗
177ranked-venue papers
11as first author
64since 2021 · last 2027
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 137 · 8 first-author · 48 since 2021Artificial intelligence and machine learning · 23 · 12 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 3 first-author · 3 since 2021Computer networks · 4 · 3 since 2021Databases, data management, data science and information retrieval · 4 · 2 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 1 since 2021Systems, architecture and hardware · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | CDV-PCQA: Content-distortion-guided dynamic viewpoint quality assessment for 3D point clouds
Qihao Liang, Li Li 0014, Ting Luo 0001, Gangyi Jiang, Wujie Zhou, Linwei Zhu, Zhouyan He |
Expert Syst. Appl. | 4 |
| 2026 | ARMLF: Anomalous region representation learning for multi-exposure fused light field image quality assessment
Guanglong Liao, Gangyi Jiang, Linwei Zhu, Yeyao Chen, Yueli Cui, Ting Luo 0001, Haiyong Xu |
Expert Syst. Appl. | 2 |
| 2026 | Multi-task guided blind light field image quality assessment via spatial-frequency collaborative modeling
Daoqiang Zhu, Guanglong Liao, Yeyao Chen, Zhouyan He, Chongchong Jin, Yueli Cui, Ming Jin 0001, Gangyi Jiang |
Expert Syst. Appl. | 8 |
| 2026 | LatentDark: Reflectance guided latent diffusion model for low-light image enhancement
Renzhi Hu, Ting Luo 0001, Gangyi Jiang, Leiming Liu, Yeyao Chen, Haiyong Xu, Zhouyan He |
Signal Process. | 3 |
| 2026 | Water-KAN: Efficient Underwater Image Enhancement via Kolmogorov-Arnold Networks
Peiyuan Jin, Haiyong Xu, Yeyao Chen, Gangyi Jiang |
IEEE Signal Process. Lett. | 4 |
| 2026 | No-Reference Stitched Wide Field of View Light Field Image Quality Assessment via Structured Representation and Progressive LearningabstractThe limited field of view (FoV) of commercial light field cameras has driven development of various stitching techniques to generate wide-FoV light field images (WLFIs). However, these techniques often introduce local distortions and angular inconsistencies, posing significant challenges for WLFI quality assessment. In this letter, a no-reference WLFI quality assessment (WLFIQA) method based on structured representation and progressive learning is proposed. Specifically, considering the high-dimensional characteristics and distortion properties of WLFIs, a novel joint spatial-angular representation strategy is first designed. For the angular domain, horizontal and vertical sub-aperture images stacks are employed to characterize angular features; for the spatial domain each sub-aperture image is divided into four quadrants, and image blocks containing complementary cues from corresponding positions across different quadrants are used as input for subsequent feature extraction. Furthermore, a global-local feature extraction network is employed to further model multi-scale distortion characteristics. Finally, a progressive learning strategy is designed to enhance the performance of overall perceptual evaluation. Experimental results on a benchmark WLFI dataset show that the proposed method outperforms existing quality methods. The code will be available athttps://github.com/sunyu-iy/WLFIQA. Yueli Cui, Ming Jin 0001, Gangyi Jiang |
IEEE Signal Process. Lett. | 5 |
| 2026 | Multi-Modal Cross-Attention-Guided Network for Audio-Visual Quality Evaluation via Visual Saliency and Mel-Spectrum FeaturesabstractThe quality evaluation of audio-visual (A/V) content has become increasingly critical in modern multimedia communication systems. Traditional single-modality quality evaluation methods and existing dedicated A/V quality models often fail to accurately assess the quality of A/V signals. To address this challenge, we propose a novel multi-modal cross-attention guided network specifically designed for A/V quality evaluation. By leveraging visual saliency and Mel-spectrum features, our network aims to achieve accurate and comprehensive quality evaluation. Specifically, distorted video frames are first converted into saliency maps, from which perceptually salient patches are selectively extracted and fed into a Convolutional Neural Network (CNN) for intra-frame visual feature extraction. Concurrently, the distorted audio signal is transformed into a Mel-spectrum, and time-frequency patches are extracted via sliding window techniques for CNN-based audio feature extraction. To effectively integrate these features and capture the long-term dependencies across consecutive A/V segments, we design a multi-modal cross-attention module that explicitly models complex inter-modal interactions. The resulting representations are then passed through a series of fully-connected (FC) layers for dimensionality reduction, ultimately deriving the quality score. Extensive experiments on three publicly available A/V quality datasets indicate that our metric outperforms the traditional quality metrics and newly-developed A/V quality metrics. The source code will be released at https://github.com/Jour3141/avqa. Yueli Cui, Chenli Fang, Binghong Pan, Chencheng Pan, Gangyi Jiang, Shiqing Zhang, Siwei Ma 0001, Qi Tian 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2026 | Probabilistic-Based Learning for Joint Light Field Image Compression and Enhancement Under Low-Light ConditionsabstractLight field (LF) imaging has attracted increasing research interest in challenging illumination conditions due to its ability to provide rich spatial and angular cues. However, such data present dual challenges: 1) the inherent multi-view structure introduces substantial data redundancy, creating high demands for efficient compression; 2) the insufficient illumination leads to severe quality degradation, which weakens inter-view consistency and visual perception. To address these coupled factors, we propose a Probabilistic-based learning for joint LF image compression and enhancement under low-light conditions (PrL-LFCE). The framework unifies structure-aware compression and feature enhancement mechanisms by introducing learnable probabilistic modeling into both feature coupling and latent distribution estimation to adaptively handle the uncertainty induced by illumination degradation and compression-related information loss. Specifically, we design a probability-based multi-directional feature coupling module that dynamically balances structural preservation and redundancy reduction across multiple directionally arranged sub-aperture images. Moreover, we introduce a swin-gated enhancement module that suppresses noise and highlights structurally salient regions in compression-aware feature representations through attention-guided gating. Extensive experiments show that PrL-LFCE consistently outperforms state-of-the-art methods, achieving at least 34.86% bitrate savings while maintaining excellent visual quality, demonstrating a strong joint compression and enhancement capability. Deyang Liu, Jimin Wang, Mounir Kaaniche, Xiaofei Zhou 0003, Gangyi Jiang, Caifeng Shan |
IEEE Trans. Image Process. | 6 |
| 2026 | DiffW: Multi-Encoder Based on Conditional Diffusion Model for Robust Image WatermarkingabstractThe existing deep-learning based robust watermarking model generally applies a discriminator to form generative adversarial network (GAN) for increasing the quality of encoded images, and adopts a single encoder to embed watermark. However, GAN training is unstable, and the single encoder cannot fully adjust the watermarking distribution, thus affecting the watermarking performance. To address those limitations, this paper presents the multi-encoder based on conditional diffusion model (CDM) for robust image watermarking, namely, DiffW. To enhance the stability, the multi-encoder structure based on CDM replaces GAN for optimizing the watermarking distribution iteratively. Specifically, the operation of each timestep in the forward and reverse diffusion processes of the CDM is regarded as an encoder to overcome the shortcomings of the single encoder structure. At the training stage, under the guidance of the conditional noisy image, the forward process trains each encoder to fuse the image and watermark to generate high-quality encoded images. During the testing stage, only a small number of trained encoders of the forward process are used, so as to reduce the time complexity. Furthermore, to improve watermarking robustness, the channel attention module (CAM) is designed to extract main watermark features by mining channel correlations for multi-layer fusion, so that watermark can be embedded into imperceptible and texture areas. The experimental results reveal that compared with the existing watermarking model, the proposed DiffW can achieve better results in terms of watermarking invisibility and robustness. Ting Luo 0001, Renzhi Hu, Zhouyan He, Gangyi Jiang, Haiyong Xu, Yang Song 0015, Chin-Chen Chang 0001 |
IEEE Trans. Multim. | 4 |
| 2025 | Mamba-Based Blind Stitched Wide Field of View Light Field Image Quality Assessment via Dual-Viewport SamplingabstractDue to the limitations of commercial light field camera hardware, the field of view (FOV) of light field images (LFIs) is relatively narrow. To expand the FOV, various LFI stitching algorithms have been developed. However, these algorithms inevitably introduce localized distortions and angular consistency disruptions, which conventional LFI quality assessment metrics struggle to evaluate effectively. To address this issue, a novel Mamba-based blind quality assessment metric for stitched wide field of view light field images (WLFIs) using dual-viewport sampling is proposed. Firstly, sub-aperture images from horizontal and vertical directions are stacked to characterize angular information, and a dual-viewport sampling pattern is designed to enhance data augmentation and capture spatial details. After that, a multi-scale state space block is proposed to improve distortion feature extraction, complemented by an auxiliary distortion discrimination task. Finally, experimental results demonstrate that the proposed metric outperforms state-of-the-art metrics on the benchmark WLFI dataset. Gangyi Jiang, Linwei Zhu, Yeyao Chen, Yueli Cui, Ting Luo 0001, Haiyong Xu |
ICME | 2 |
| 2025 | Quality assessment of windowed 6DoF video with viewpoint switching
Wenhui Zou, Tingyan Tang, Gangyi Jiang, Zongju Peng |
J. Vis. Commun. Image Represent. | 4 |
| 2025 | Combining independent and joint spatial-angular information learning for light field image super-resolution
Dezhang Ke, Yeyao Chen, Chongchong Jin, Haiyong Xu, Zhidi Jiang, Ting Luo 0001, Gangyi Jiang |
Knowl. Based Syst. | 7 |
| 2025 | DiffOSR: Latitude-aware conditional diffusion probabilistic model for omnidirectional image super-resolution
Leiming Liu, Ting Luo 0001, Gangyi Jiang, Yeyao Chen, Haiyong Xu, Renzhi Hu, Zhouyan He |
Knowl. Based Syst. | 3 |
| 2025 | Unveiling the underwater world: CLIP perception model-guided underwater image enhancement
Jiang-Zhong Cao, Zekai Zeng, Xu Zhang 0044, Huan Zhang 0008, Chunling Fan, Gangyi Jiang, Weisi Lin |
Pattern Recognit. | 6 |
| 2025 | DiffDark: Multi-prior integration driven diffusion model for low-light image enhancement
Renzhi Hu, Ting Luo 0001, Gangyi Jiang, Yeyao Chen, Haiyong Xu, Leiming Liu, Zhouyan He |
Pattern Recognit. | 3 |
| 2025 | Frequency domain-based latent diffusion model for underwater image enhancement
Jingyu Song, Haiyong Xu, Gangyi Jiang, Mei Yu 0001, Yeyao Chen, Ting Luo 0001, Yang Song 0015 |
Pattern Recognit. | 3 |
| 2025 | Geometry-Aware RWKV for Heterogeneous Light Field Spatial Super-ResolutionabstractHeterogeneous Light Field (LF) spatial Super-Resolution (SR) aims to significantly enhance the spatial resolution of LF imaging by integrating an extra 2D digital camera. Inspired by the Receptance Weighted Key Value (RWKV), a simple yet effective heterogeneous LF spatial SR method is proposed. Specifically, a texture transfer module with channel correlation is designed, which leverages a feature distillation strategy to transfer texture information from the high-resolution 2D image to the low-resolution LF image. Meanwhile, a spatial-angular rectification module is constructed to restore the spatial-angular coherence damaged in texture transfer. It employs geometry-aware RWKV to capture the intrinsic geometric structure of LFs. Experimental results show that the proposed method outperforms the state-of-the-art methods in both quantitative and qualitative comparisons, while achieving higher efficiency in terms of inference time and memory usage. Zean Chen, Yeyao Chen, Linwei Zhu, Haiyong Xu, Gangyi Jiang |
IEEE Signal Process. Lett. | 5 |
| 2025 | Blind Light Field Image Quality Assessment via Frequency Domain Analysis and Auxiliary LearningabstractDue to the distortions occurring at various stages from acquisition to visualization, light field image quality assessment (LFIQA) is crucial for guiding the processing of light field images (LFIs). In this letter, we propose a new blind LFIQA metric via frequency domain analysis and auxiliary learning, termed as FABLFQA. First, spatial-angular patches are extracted from LFIs and further processed through discrete cosine transform to obtain light field frequency maps. Subsequently, a concise and efficient frequency-aware deep learning network is designed to extract frequency features, including the frequency descriptor, 3D ConvBlock, and frequency transformer. Finally, a distortion type discrimination auxiliary task is employed to facilitate the learning of the main quality assessment task. Experimental results on three representative LFI datasets show that the proposed metric outperforms the state-of-the-art metrics. Gangyi Jiang, Linwei Zhu, Yueli Cui, Ting Luo 0001 |
IEEE Signal Process. Lett. | 2 |
| 2025 | Geometry-Guided Latent Diffusion Model for Static Point Cloud Color Attribute Denoising
Linwei Zhu, Ruxu Liang, Yun Zhang 0002, Gangyi Jiang, Yo-Sung Ho |
IEEE Signal Process. Lett. | 4 |
| 2025 | StegMamba: Distortion-Free Immune-Cover for Multi-Image Steganography With State Space ModelabstractMulti-image steganography ensures privacy protection while avoiding suspicion from third parties by embedding multiple secret images within a cover image. However, existing multi-image steganographic methods fail to model global spatial correlations to reduce image damage at the low computation cost. Moreover, they do not account for the anti-distortion capability of the cover image, which is crucial for achieving imperceptible and ensuring security. To overcome these limitations, we propose StegMamba, a distortion-free immune-cover for multi-image steganography architecture with a state space model. Specifically, we first explore the potential of the linear computational cost model Mamba for data hiding tasks through a steganography Mamba block (SMB), whose efficiency makes it suitable for real-time applications. Subsequently, considering that images with distortion resistance reduce embedding damage, the original cover image is reconstructed through immune-cover construction module (ICCM) and associated with the steganography task. Moreover, well-coupled features facilitate fusion, and thus a wavelet-based interaction module (WIM) is designed for effective communication between the immune-cover and the secret images. Compared with the state-of-the-art global attention-based methods, the proposed StegMamba obtains PSNR gains of 3.30 dB, 1.37 dB, and 1.92 dB for the stego image, and two secret recovery images, respectively, and the reduction of 2.87% in detection accuracy for anti-steganalysis. This code is available athttps://github.com/YuhangZhouCJY/StegMamba. Ting Luo 0001, Zhouyan He, Gangyi Jiang, Haiyong Xu, Yushu Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | DA-Net: A Double Alignment Multimodal Learning Network for Point Cloud Quality AssessmentabstractExisting multimodal point cloud quality assessment (PCQA) methods usually integrate 3D and 2D information to simulate human visual perception of distortions. However, due to the lack of consideration of spatial correspondence, they have difficulty to learn consistent distortion representations from different modalities in the same region of the PC. In addition, they also ignore the heterogeneity of modalities and rely on complex fusion mechanisms (e.g., attention) to integrate multimodal features. Both lead to limited performance and increased computational complexity. To address these limitations, we propose a novel double alignment multimodal learning network (DA-Net), which introduces two key alignment strategies. Specifically, the first is spatial pre-alignment strategy, which generates informative 2D patch for each 3D patch via an adaptive patch projection module (APPM), ensuring accurate spatial correspondence of different modalities prior to feature extraction. The second is a uniform feature alignment strategy, which includes feature disentanglement module (FDM) and feature mapping module (FMM) to relieve heterogeneity of modalities and guide the optimization of 2D and 3D encoder. Finally, multimodal features are simply integrated and regressed to obtain the quality score. Experimental results demonstrate that the DA-Net exhibits outstanding performance and generalization ability. It also achieves lower computational complexity compared with other multimodal PCQA methods. The source codes of DA-Net will be available at https://github.com/Rphone/DA-Net. Xinqiang Wu, Zhouyan He, Ting Luo 0001, Gangyi Jiang, Wujie Zhou, Linwei Zhu, Weisi Lin |
IEEE Trans. Image Process. | 4 |
| 2025 | Local and Global Structure-Guided No-Reference Point Cloud Quality AssessmentabstractAs a crucial representation of 3D data, a point cloud (PC) can accurately capture the geometry, structure, and color information of objects. However, various quality problems arise owing to device noise, data acquisition errors, and compression algorithms, limiting the application of PCs. Therefore, assessing PC quality to determine its suitability for applications is a challenging task. In this work, a local and global structure-guided feature extraction and attention network (LGS-Net) is introduced for no-reference PC quality assessment (PCQA). This approach incorporates cluster construction (CC), local structure-guided cluster feature extraction (LSFE), and global structure-guided attention (GSA) modules. First, owing to the heightened sensitivity of the human visual system (HVS) to structural information, a graph filter is employed to identify high-frequency clusters. Within the LSFE module, a multiscale strategy is employed to ensure that structural information effectively influences both the geometry and color information. Simultaneously, the multiscale features within the cluster are dynamically fine-tuned using feature channel weight reassignment. To account for the impact of interclusters on overall quality, a GSA module is introduced to establish global dependencies between local clusters. This approach enables the extraction of final geometry, color, and structure information, which are ultimately used for accurate quality assessment. Extensive experimental results show that the proposed method outperforms the existing state-of-the-art PCQA methods using two publicly available subjective datasets. Zhouyan He, Qihao Liang, Gangyi Jiang, Mei Yu 0001, Yeyao Chen, Ting Luo 0001, Wujie Zhou |
IEEE Trans. Multim. | 3 |
| 2025 | Multiscale Feature Importance-Based Bit Allocation for End-to-End Feature Coding for MachinesabstractFeature Coding for Machines (FCM) aims to compress intermediate features effectively for remote intelligent analytics, which is crucial for future intelligent visual applications. In this article, we propose a Multiscale Feature Importance-based Bit Allocation (MFIBA) for end-to-end FCM. First, we find that the importance of features for machine vision tasks varies with the scales, object size, and image instances. Based on this finding, we propose a Multiscale Feature Importance Prediction (MFIP) module to predict the importance weight for each scale of features. Second, we propose a task loss-rate model to establish the relationship between the task accuracy losses of using compressed features and the bit rate of encoding these features. Finally, we develop an MFIBA for end-to-end FCM, which is able to assign coding bits of multiscale features more reasonably based on their importance. Experimental results demonstrate that when combined with a retained Efficient Learned Image Compression (ELIC), the proposed MFIBA achieves an average of 38.202% bit-rate savings in object detection compared to the anchor ELIC. Moreover, the proposed MFIBA achieves an average of 17.212% and 36.492% feature bit-rate savings for instance segmentation and keypoint detection, respectively. When the proposed MFIBA is applied to the LIC-TCM, it achieves an average of 18.103%, 19.866%, and 19.597% bit-rate savings on three machine vision tasks, respectively, which validates the proposed MFIBA has good generalizability and adaptability to different machine vision tasks and FCM base codecs. Junle Liu, Yun Zhang 0002, Zixi Guo, Xiaoxia Huang 0004, Gangyi Jiang |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2025 | Multi-Attention Learning and Exposure Guidance Toward Ghost-Free High Dynamic Range Light Field ImagingabstractDue to sensor limitations, the light field (LF) images captured by the LF camera suffer from low dynamic range and are prone to poor exposure. To solve this problem, combining multi-exposure technology with LF camera imaging can achieve high dynamic range (HDR) LF imaging. However, for dynamic scenes, this approach tends to produce disturbing ghosting artifacts and destroy the parallax structure of the generated results. To this end, this paper proposes a novel ghost-free HDR LF imaging method using multi-attention learning and exposure guidance. Specifically, the proposed method first designs a multi-scale cross-attention module to achieve efficient multi-exposure LF feature alignment. After that, a dual self-attention-driven Transformer block is constructed to excavate the geometric information of LF and fuse the aligned LF features. In particular, exposure masks derived from middle-exposure are introduced in the feature fusion to guide the network to focus on information recovery in low- and high-brightness regions. Besides, a local compensation module is integrated to cope with local alignment errors and refine details. Finally, a multi-objective reconstruction strategy combined with exposure masks is employed to restore high-quality HDR LF images. Extensive experimental results on the benchmark dataset show that the proposed method generates HDR LF results with high spatial-angular quality consistency and outperforms the state-of-the-art methods in quantitative and qualitative comparisons. Furthermore, the proposed method can enhance the performance of existing LF applications, such as depth estimation. Yeyao Chen, Gangyi Jiang, Chongchong Jin, Ting Luo 0001, Haiyong Xu, Mei Yu 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2024 | Hybrid Domain Learning towards Light Field Spatial Super-Resolution using Heterogeneous ImagingabstractLight field (LF) cameras usually capture dense angular samples, but suffer from low spatial resolution. Existing single-LF super-resolution methods struggle with textures at larger scales (e.g., 8×). To address this issue, this paper proposes a novel hybrid domain learning-based method to enhance LF spatial resolution from heterogeneous imaging (integrating an LF camera and a 2D digital camera). The proposed method consists of two core modules, namely LF feature alignment module and cross-domain multi-scale fusion module. The former combines optical flow and deformable convolution to gradually align the 2D high-resolution features with the low-resolution LF features. The latter progressively fuses the aligned multi-resolution LF features to enable high-quality reconstruction. Experimental results show the proposed method recovers fine textures and preserves accurate angular consistency, and outperforms the state-of-the-art methods in both quantitative and qualitative comparisons. Zean Chen, Yeyao Chen, Mei Yu 0001, Haiyong Xu, Gangyi Jiang |
ICASSP | 5 |
| 2024 | Multi-exposure fused light field image quality assessment for dynamic scenes: Benchmark dataset and objective metric
Yun Liu 0048, Guanglong Liao, Gangyi Jiang, Yeyao Chen, Yueli Cui, Haiyong Xu, Mei Yu 0001 |
Expert Syst. Appl. | 3 |
| 2024 | CAISFormer: Channel-wise attention transformer for image steganographyabstractCurrent Transformer-based image steganography cannot embed data properly without considering the correlation of the cover image and the secret image . In addition, to save computational complexity, spatial-wise Transformer is often used to apply in small spatial windows, which limits the extraction of the global feature. To solve those limitations, we present a channel-wise attention Transformer model for image steganography (CAISFormer), which aims to construct long-range dependencies for identifying inconspicuous positions to embed data. A channel self-attention module (CSAM) is deployed to focus the feature channels suitable for data hiding by establishing channel relationships. Meanwhile, a non-linear enhancement (NLE) layer is employed to enhance the beneficial features while weaken the irrelevant ones. For building feature coupling between the cover image and the secret image, a channel-wise cross attention module (CCAM) is designed to fine-tune cover image features by capturing their cross-dependencies. In addition, for concealing data properly, a global–local aggregation module (GLAM) is deployed to adjust fused features by combining global and local attention, which can focus on inconspicuous and texture regions, respectively. The experimental results demonstrate that CAISFormer obtains PSNR gains of more than 0.36 dB and 0.90 dB for the cover/stego image pair and the secret/recovery image pair, respectively, and the detection ratio is decreased by 3.43%, in single image hiding compared to the state-of-the-art. Moreover, the generalization ability is also proved across a variety of datasets. The code will be made publicly available at https://github.com/YuhangZhouCJY/CAISFormer . Ting Luo 0001, Zhouyan He, Gangyi Jiang, Haiyong Xu, Chin-Chen Chang 0001 |
Neurocomputing | 4 |
| 2024 | Vision graph convolutional network for underwater image enhancement
Zexuan Xing, Haiyong Xu, Gangyi Jiang, Mei Yu 0001, Ting Luo 0001, Yeyao Chen |
Knowl. Based Syst. | 3 |
| 2024 | HDR light field imaging of dynamic scenes: A learning-based method and a benchmark dataset
Yeyao Chen, Gangyi Jiang, Mei Yu 0001, Chongchong Jin, Haiyong Xu, Yo-Sung Ho |
Pattern Recognit. | 2 |
| 2024 | Underwater Monocular Depth Estimation Based on Physical-Guided TransformerabstractOwing to the light absorption and wavelength scattering in underwater environments, underwater images are severely degraded, which directly affects the depth estimation of underwater scenes. Accurate underwater depth estimation is essential for representing and understanding underwater scenes. However, the existing underwater depth estimation methods have not fully taken into account the distinctive physical properties of underwater environments, which has resulted in increased bias and feature distortion in the depth estimation results. In this paper, an underwater monocular depth estimation method based on physical-guided Transformer (UPGformer) is proposed, considering the characteristics of underwater imaging, including shallow feature extraction, encoding, decoding, and regression stages. Specifically, in the shallow feature extraction stage, considering the color deviation of underwater images and extracting richer primary features, an enrichment and extraction depth Transformer (EEDT) module is proposed, by interacting physically inverted transmission maps of the underwater dark channel prior (UDCP) with physical color-compensated underwater images through self-attention. In the encoding stage, considering the nonuniform degradation of underwater images (nonuniform local distortion and inconsistent channel degradation), the underwater physical Transformer interaction encoder (UPTE) module, which fuses the Transformer and physically inverted transmission maps, is proposed. Furthermore, in the decoding stage, to better recover features and reduce information loss, the underwater physical embedded decoding (UPED) module is proposed, which embeds the physically inverted transmission maps with the upsampling process. Finally, the depth map is constructed during the regression stage. The experimental results demonstrate that the proposed UPGformer outperforms existing methods, both qualitatively and quantitatively. Chen Wang 0141, Haiyong Xu, Gangyi Jiang, Mei Yu 0001, Ting Luo 0001, Yeyao Chen |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Stitched Wide Field of View Light Field Image Quality Assessment: Benchmark Database and Objective MetricabstractDue to the limitation of commercial light field camera hardware devices, the imaging field of view is quite narrow. Numerous Light Field Image (LFI) stitching algorithms have been developed to expand the field of view. However, it is highly challenging to compare the performance of LFI stitching algorithms in a fair manner. Currently, due to the absence of a comprehensive benchmark database for subjective rating and a reliable objective quality metric, it is fairly difficult to comprehensively and accurately compare the actual performance of existing LFI stitching algorithms. In this study, we dedicate our efforts to the development of quality metrics for stitched Wide field of view LFI (WLFI) from subjective and objective assessment aspects. Specifically, we build the first stitched WLFI database, which provides the stitched WLFIs generated by eight representative LFI stitching algorithms, along with their corresponding subjective rating scores. Secondly, an effective blind stitched WLFI quality metric is developed to accurately assess the visual quality degradation. Extensive experiments conducted over our established WLFI database demonstrate that the proposed metric achieves higher consistency with subjective ratings than the competing quality metrics. Yueli Cui, Gangyi Jiang, Mei Yu 0001, Yeyao Chen, Yo-Sung Ho |
IEEE Trans. Multim. | 2 |
| 2024 | Hierarchical Independent Coding Scheme for Varifocal Multiview Images Based on Angular-Focal Joint PredictionabstractVarifocal multiview (VFMV) images are dense views that focus on variable focal planes. Thus, VFMV images are highly redundant in the angular, spatial and focal dimensions. In this article, the redundancies of VFMV images are analyzed and represented by full parallaxes and focal inconsistency. To exploit these distinctive redundancies, we propose a hierarchical independent coding scheme based on angular-focal joint prediction. The scheme is constructed by hierarchical independent prediction structure (HIPS) and angular-focal joint prediction (AFJP). The HIPS separates all views into several independent subdivisions and assigns different hierarchies inside each subdivision, which enhances random access capability and scalability. The AFJP conducts motion estimation and focal approximation simultaneously to predict parallaxes and focal inconsistency. Therefore, the redundancies in the angular and focal dimensions can be exploited by the proposed coding scheme. We construct a VFMV dataset with 10 test sequences for different acquisition methods. The experimental results on these test sequences demonstrate that the proposed scheme outperforms all comparison schemes in objective quality, subjective quality and random access capability. Specifically, the proposed coding scheme achieves up to 2.661 dB PSNR gains and 52.817% bitrate savings compared with the HEVC random access benchmark scheme. Kejun Wu, You Yang 0002, Qiong Liu 0001, Gangyi Jiang, Xiao-Ping Zhang 0002 |
IEEE Trans. Multim. | 4 |
| 2024 | Pseudo Light Field Image and 4D Wavelet-Transform-Based Reduced-Reference Light Field Image Quality AssessmentabstractReduced-reference light field image (LFI) quality assessment (RR LFIQA) automatically assesses image quality with only partial information about the reference LFI is available. Existing RR LFIQA has difficulty extracting effective RR information and perceptual features to represent the LFI quality. In this article, we propose an RR LFIQA model based on pseudo LFI (PLFI) and four-dimensional (4D) wavelet transform. To extract RR information related to LFI perceptual quality, a PLFI is created as the RR information of the LFI using a view synthesis algorithm. Considering that the high-dimensional characteristics of the PLFI, 4D wavelet transform is used to decompose the original and distorted PLFIs. The 4D wavelet transform essentially performs a continuous 1D wavelet transform for the 4D signal to enable the local 4D structure of the PLFIs to be characterized effectively in the 4D wavelet domain. A novel spatial-angular weighting strategy is proposed to describe the importance of each location for quality evaluation, to further improve the performance of the proposed method. Experimental results on four benchmark datasets show that the proposed model performs better than the representative 2DIQA and LFIQA models. Jianjun Xiang, Peng Chen 0008, Yuanjie Dang, Ronghua Liang, Gangyi Jiang |
IEEE Trans. Multim. | 5 |
| 2024 | Underwater Image Quality Assessment from Synthetic to Real-world: Dataset and Objective MethodabstractThe complicated underwater environment and lighting conditions lead to severe influence on the quality of underwater imaging, which tends to impair underwater exploration and research. To effectively evaluate the quality of underwater images, an underwater image quality assessment dataset is constructed from synthetic to real-world, and then a new objective underwater image assessment method based on the characteristics of the underwater imaging is proposed (UICQA). Specifically, to address the lack of a publicly available datasets and more accurately quantify the quality of underwater images, a subjective underwater image quality assessment dataset from synthetic to real-world underwater images, named USRD, is constructed. Considering that the transmission map can effectively reflect the characteristics of the underwater imaging, statistical features are effectively extracted from the transmission map for distinguishing underwater images of different quality. Further, considering that the transmission map negatively correlates with scene depth, a local-to-global transmission map weighted contrast feature is constructed. Additionally, the color features of human perception and texture features based on fractal dimensions are proposed. Finally, the experimental results show that the proposed UICQA method exhibits the highest correlation with ground truth scores compared to state-of-the-art UIQA methods. Haiyong Xu, Gangyi Jiang, Mei Yu 0001, Ting Luo 0001, Xuebo Zhang 0002, Hongwei Ying |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2023 | Quality Assessment for High Dynamic Range Stereoscopic Omnidirectional Image System
Liuyan Cao, Hao Jiang 0014, Zhidi Jiang, Jihao You, Mei Yu 0001, Gangyi Jiang |
ACIVS | 6 |
| 2023 | Quality Evaluation of Tone Mapping HDR Omnidirectional Image with Multi-regions and Multi-levelsabstractUnlike ordinary omnidirectional imaging, which produces underexposure or overexposure of some areas under wide range of lighting conditions, tone mapping high dynamic range omnidirectional image (denoted as TM-HOI) presents more detailed information with high dynamic range (HDR) techniques. Aiming at the multiple distortions in TM-HOI processing, this paper proposes a blind TM-HOI quality metric with multi-regions and multi-levels analysis. Considering the conversion of image projection formats and the unique distortion caused by user behavior in immersive environments, feature extraction in the proposed metric is divided into the viewport based module and the Crasher parabolic projection based module. At the same time, considering the coding distortion, tone mapping (TM) distortion and mixed distortion in the TM-HOI and the different manifestations of this mixed distortion in different regions, this paper further extracts perceptual features to characterize image distortion by performing bit plane layer decomposition and detail/basic layer decomposition on different regions of TM-HOI. Finally, the extracted perceptual features are used as the input of Random forest to construct the nonlinear relationship between the feature space and subjective opinion scores. The experimental results indicate that the proposed metric has better consistency with the human visual system compared to representative blind quality metrics. Xuelei Zheng, Mei Yu 0001, Gangyi Jiang |
AICCSA | 3 |
| 2023 | Perceptual Light Field Image Coding with CTU Level Bit Allocation
Panqi Jin, Gangyi Jiang, Yeyao Chen, Zhidi Jiang, Mei Yu 0001 |
CAIP (2) | 2 |
| 2023 | Extending Depth of Field by Varifocal Multi-View Computational Imaging for MetaverseabstractThe flexible field of view (FoV) and large depth of field (DoF) are the main bricks that build the strong immersive experience in Metaverse. However, due to the nature of optics, the captured multi-view images are generally with flexible FoV but limited DoF. To extend the DoF of captured data, in this paper, we propose an all-in-focus image fusion scheme by varifocal multi-view computational imaging. Varifocal multi-view images are a series of multi-view images in different DoF, where different scene contents and blur degrees are the main features among views. Due to the complex inter-view features, the existing extending DoF methods on varifocal multi-view images yield severe ghosting problem. To alleviate the ghosting problem, a patch-based DenseNet image fusion network is designed and embedded in the proposed scheme. The patch-based image fusion network enables to mitigate the ghosting problem in the fused image. Experiments on varifocal multi-view images of different scenes demonstrate that the proposed all-in-focus image fusion scheme can synthesize all-in-focus results with higher visual quality and accuracy. The proposed all-in-focus image fusion scheme is expected to benefit metaverse, photo-realistic novel view synthesis, interactive and immersive experience. Zhilong Li, Kejun Wu, Gangyi Jiang, You Yang 0002 |
MMSP | 3 |
| 2023 | Fast intra partition and mode prediction for equirectangular projection 360-degree video codingabstractAbstract 360‐degree videos have drawn great attention from both the academia and the industry. For transmission and storage of 360‐degree video, the joint video exploration team proposed the Versatile Video Coding (VVC) standard. VVC can significantly reduce the bitrate while maintaining the same subjective visual quality compared to the preceding high efficiency video coding. However, the computational complexity of VVC is extremely high which hinders the interactive applications. This paper proposes a fast intra partition method and mode prediction algorithm for equirectangular projection 360‐degree video coding. First, a latitude‐based preprocessing is introduced to early terminate the Coding Unit (CU) partition in the polar region. Second, the support vector machine is used to predict the CU partition type. Third, the fast intra mode search method accelerates the intra mode prediction. Experimental results show that the proposed algorithm can significantly obtain an average time reduction rate of 60.40% and a Bjontegaard delta rate increase of 1.96%. Zheng-jie Shu, Zongju Peng, Gangyi Jiang, Bo-sen Yuan |
IET Image Process. | 3 |
| 2023 | Blind light field image quality assessment with tensor color domain and 3D shearlet transform
Jianjun Xiang, Mei Yu 0001, Gangyi Jiang, Haiyong Xu |
Signal Process. | 3 |
| 2023 | No-Reference Light Field Image Quality Assessment Using Four-Dimensional Sparse TransformabstractLight field imaging can simultaneously capture the intensity and direction information of light rays in the real world. Light field image (LFI) with four-dimensional (4D) data suffers from quality degradation in the process of compression, reconstruction and processing. How to evaluate the visual quality of LFI is thought-provoking. This paper proposes a no-reference LFI quality assessment metric based on high-dimensional sparse transform. Firstly, LFI's sub-aperture gradient image array (SAGIA), which is still a 4D signal, is generated by high-pass filtering between adjacent SAIs. Then, SAGIA is transformed with 4D discrete cosine transform (4D-DCT). 4D-DCT coefficients of SAGIA can characterize the angular and spatial information of LFI. And the logarithmic amplitudes of the coefficients at the same position of SAGIA?s transformed 4D blocks are averaged as the coefficient energy. Subsequently, the 4D-DCT coefficients of SAGIA are divided into the spatial-angular frequency bands and spatial-angular orientation bands, and the corresponding energy features are extracted by converging the coefficient energy of the same band. In addition, the coefficients' amplitudes at the same position of blocks are fitted by the Weibull distribution. Then, the fitted parameters of each position are concatenated, and cropped with principal component analysis to obtain the compact features. Finally, the extracted features are pooled to predict the visual quality of the distorted LFIs. The experimental results demonstrate that the proposed method is more consistent with the subjective evaluation on three LFI databases, compared with the state-of-the-art image quality assessment methods and LFI quality assessment methods. Jianjun Xiang, Gangyi Jiang, Mei Yu 0001, Zhidi Jiang, Yo-Sung Ho |
IEEE Trans. Multim. | 2 |
| 2023 | Deep Learning-Based Intra Mode Derivation for Versatile Video CodingabstractIn intra coding, Rate Distortion Optimization (RDO) is performed to achieve the optimal intra mode from a pre-defined candidate list. The optimal intra mode is also required to be encoded and transmitted to the decoder side besides the residual signal, where lots of coding bits are consumed. To further improve the performance of intra coding in Versatile Video Coding (VVC) , an intelligent intra mode derivation method is proposed in this paper, termed as Deep Learning based Intra Mode Derivation (DLIMD) . In specific, the process of intra mode derivation is formulated as a multi-class classification task, which aims to skip the module of intra mode signaling for coding bits reduction. The architecture of DLIMD is developed to adapt to different quantization parameter settings and variable coding blocks including non-square ones, where only one single trained model is required. Different from the existing deep learning based classification problems, the hand-crafted features are also fed into intra mode derivation network besides the learned features from feature learning network. To compete with traditional methods, one additional binary flag is utilized in the video codec to indicate the selected scheme with RDO. Extensive experimental results reveal that the proposed method can achieve 2.28%, 1.74%, and 2.18% bit rate reduction on average for Y, U, and V components on the platform of VVC test model, which outperforms the state-of-the-art works. Linwei Zhu, Yun Zhang 0002, Na Li 0015, Gangyi Jiang, Sam Kwong |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2023 | Deep Light Field Spatial Super-Resolution Using Heterogeneous ImagingabstractLight field (LF) imaging expands traditional imaging techniques by simultaneously capturing the intensity and direction information of light rays, and promotes many visual applications. However, owing to the inherent trade-off between the spatial and angular dimensions, LF images acquired by LF cameras usually suffer from low spatial resolution. Many current approaches increase the spatial resolution by exploring the four-dimensional (4D) structure of the LF images, but they have difficulties in recovering fine textures at a large upscaling factor. To address this challenge, this paper proposes a new deep learning-based LF spatial super-resolution method using heterogeneous imaging (LFSSR-HI). The designed heterogeneous imaging system uses an extra high-resolution (HR) traditional camera to capture the abundant spatial information in addition to the LF camera imaging, where the auxiliary information from the HR camera is utilized to super-resolve the LF image. Specifically, an LF feature alignment module is constructed to learn the correspondence between the 4D LF image and the 2D HR image to realize information alignment. Subsequently, a multi-level spatial-angular feature enhancement module is designed to gradually embed the aligned HR information into the rough LF features. Finally, the enhanced LF features are reconstructed into a super-resolved LF image using a simple feature decoder. To improve the flexibility of the proposed method, a pyramid reconstruction strategy is leveraged to generate multi-scale super-resolution results in one forward inference. The experimental results show that the proposed LFSSR-HI method achieves significant advantages over the state-of-the-art methods in both qualitative and quantitative comparisons. Furthermore, the proposed method preserves more accurate angular consistency. Yeyao Chen, Gangyi Jiang, Mei Yu 0001, Haiyong Xu, Yo-Sung Ho |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2022 | Iterative enhancement scheme of synthesized color and depth images for immersive video systemabstractImmersive video allows viewers to freely switch the viewpoints. The intensity of realistic experience greatly relies on the quality of synthesized depth maps. However, there exist distorted regions due to inaccurate depth estimation or compression. Yongquan Su, Qiong Liu 0001, Kejun Wu, Gangyi Jiang, You Yang 0002 |
DCC | 4 |
| 2022 | CPC-GSCT: Visual quality assessment for coloured point cloud based on geometric segmentation and colour transformationabstractAbstract Coloured point cloud (CPC) is one of the important representations of three‐dimensional objects, which has been used in many fields. CPC may encounter geometric and colour distortion during its compression, simplification or other processing. Thus, the objective visual quality assessment of CPC is one of the urgent issues to be resolved in the CPC's applications. Aiming at this problem, this paper proposes a new full‐reference visual quality assessment metric for CPC based on geometric segmentation and colour transformation (CPC‐GSCT), which analyzes geometric distortion and colour distortion of CPC. First, considering the visual masking effect of CPC's geometric information, CPC is segmented into different regions and distributed with different weights to describe the influence of visual masking effect in CPC quality assessment. At the same time, a geometric combination feature vector is defined and extracted for measuring the CPC's geometric distortion. Then, considering the colour perception of human eyes, a colour combination feature vector is extracted to measure the CPC's colour distortion in HSV colour space. Finally, all the extracted geometric and colour features are constituted as a feature vector to predict the quality of CPC. Experimental results on three databases (IRPC, SJTU‐PCQA and CPCD2.0) show that the proposed CPC‐GSCT metric can achieve better performance in predicting the visual quality of CPC than relevant existing methods. Mei Yu 0001, Zhouyan He, Renwei Tu, Gangyi Jiang |
IET Image Process. | 5 |
| 2022 | TGP-PCQA: Texture and geometry projection based quality assessment for colored point clouds
Zhouyan He, Gangyi Jiang, Mei Yu 0001, Zhidi Jiang, Zongju Peng |
J. Vis. Commun. Image Represent. | 2 |
| 2022 | Multi-Angle Projection Based Blind Omnidirectional Image Quality AssessmentabstractMost of the existing blind omnidirectional image quality assessment (BOIQA) methods are based on data-driven approach where the end-to-end neural network or deep learning tools are mainly used for feature extraction. However, it usually lacks interpretability and is difficult to discover the perceptual mechanism behind. In this paper, from the perspective of perception modeling, we propose a novel multi-angle projection based BOIQA (MP-BOIQA) method. Considering the omnibearing and near eye display characteristics with head mounted display, multiple color cubemap projection images with respect to different viewpoints are grouped as the color omnidirectional distortion (COD) units so as to simulate the user’s viewing behavior in subjective quality assessment. In the designed multi-angle projection based feature extractor, tensor decomposition is implemented on each COD unit for dimensionality reduction, and piecewise exponential fitting is used to get the distribution of mean subtracted contrast normalized coefficients of the unit’s feature matrices in tensor domain. Finally, the extracted features are pooled with random forest. The experimental results on three omnidirectional image quality datasets show that the MP-BOIQA method can deliver highly competitive performance compared with some representative full-reference quality assessment methods, as well as some state-of-the-art BOIQA methods. Hao Jiang 0014, Gangyi Jiang, Mei Yu 0001, Ting Luo 0001, Haiyong Xu |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Reinforced Swin-Convs Transformer for Simultaneous Underwater Sensing Scene Image Enhancement and Super-resolutionabstractUnderwater image enhancement (UIE) technology aims to tackle the challenge of restoring the degraded underwater images due to light absorption and scattering. Meanwhile, the ever-increasing requirement for higher resolution images from a lower resolution in the underwater domain cannot be overlooked. To address these problems, a novel U-Net-based reinforced Swin-Convs Transformer for simultaneous enhancement and superresolution (URSCT-SESR) method is proposed. Specifically, with the deficiency of U-Net based on pure convolutions, the Swin Transformer is embedded into U-Net for improving the ability to capture the global dependence. Then, given the inadequacy of the Swin Transformer capturing the local attention, the reintroduction of convolutions may capture more local attention. Thus, an ingenious manner is presented for the fusion of convolutions and the core attention mechanism to build a reinforced Swin-Convs Transformer block (RSCTB) for capturing more local attention, which is reinforced in the channel and the spatial attention of the Swin Transformer. Finally, experimental results on available datasets demonstrate that the proposed URSCT-SESR achieves the state-of-the-art performance compared with other methods in terms of both subjective and objective evaluations. The code is publicly available athttps://github.com/TingdiRen/URSCT-SESR. Tingdi Ren, Haiyong Xu, Gangyi Jiang, Mei Yu 0001, Ting Luo 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2022 | Deep Light Field Super-Resolution Using Frequency Domain Analysis and Semantic PriorabstractLight field (LF) camera can simultaneously capture the intensity and direction information of light rays, which has been widely concerned. However, limited by the size of the imaging sensor, the captured LF image (LFI) has a trade-off between spatial and angular resolutions. To this end, this paper proposes a new LF super-resolution method using frequency domain analysis and semantic prior, which designs a two-stage learning framework to enhance the spatial and angular resolutions of LFI. Specifically, the proposed method first decomposes the spatial and angular information to explore the 4D structure of LFI by using frequency domain transformation, and formulates the LF super-resolution as a frequency restoration process. Then, the decomposed frequency components are recovered in a progressive restoration manner, with new cascaded 2D and 3D convolutional neural networks. To further improve the quality of the reconstructed LFI, especially at the object boundary, the semantic prior is incorporated into the designed network to enhance its representation ability. Finally, the super-resolved LFI is reconstructed by inverse frequency domain transformation. Experimental results show that the proposed method can effectively generate high-resolution LFI, and outperforms other state-of-the-art methods in terms of both subjective visual perception and objective quality evaluation. Moreover, the proposed method can enhance the performance of LF applications such as depth estimation. Yeyao Chen, Gangyi Jiang, Zhidi Jiang, Mei Yu 0001, Yo-Sung Ho |
IEEE Trans. Multim. | 2 |
| 2022 | Tensor Product and Tensor-Singular Value Decomposition Based Multi-Exposure Fusion of ImagesabstractConsidering multidimensional structure of the multi-exposure images, a new Tensor product and Tensor-singular value decomposition based Multi-Exposure image Fusion (TT-MEF) method is proposed. The main innovation of this work is to explore a new feature representation of multi-exposure images in the new tensor domain and design the fusion strategy on this basis. Specifically, the luminance and the chrominance channels are fused separately to maintain color consistency. For the luminance fusion, the luminance channel of multi-exposure images is divided into two parts, that is, de-mean term and mean term. The de-mean term is represented as a tensor to extract the feature. Then, the tensor product and tensor-singular value decomposition (T-SVD) are used to design a tensor feature extractor. Furthermore, a fusion strategy of the de-mean term is presented according to the visual saliency model, and a fusion strategy of the mean term is defined by the local and the global visual weights to control counterpoise between the local and global luminance. For the chrominance fusion, a new fusion strategy is also designed by the tensor product and T-SVD, similar to the luminance fusion. Finally, the fused image is obtained by combining the luminance and chrominance fusion. Experimental results show that the proposed TT-MEF method generally outperforms the existing state-of-the-art in terms of subjective visual quality and objective evaluation. Haiyong Xu, Gangyi Jiang, Mei Yu 0001, Zhongjie Zhu, Yongqiang Bai, Yang Song 0015, Huifang Sun |
IEEE Trans. Multim. | 2 |
| 2021 | Towards A Colored Point Cloud Quality Assessment Method Using Colored Texture And Curvature ProjectionabstractColored point cloud (PC) provides convenience for 3D digitization in the real world, but its huge amount of data needs to be compressed effectively. However, lossy compression will bring visual quality problems, so it is necessary to design reliable quality assessment methods. Considering the visual connection between 3D space and projection plane, we propose a new PC quality assessment (PCQA) method combining colored texture and curvature projection in this paper. Specifically, the colored texture information and curvature of colored PC are projected onto 2D planes to extract texture and geometric statistical features, respectively, so as to characterize the texture and geometric distortion. Experimental results on two colored PC databases (CPCD2.0 and IRPC) show that the proposed method has a good correlation with subjective quality scores and is superior to the state-of-the-art PCQA methods. Zhouyan He, Gangyi Jiang, Zhidi Jiang, Mei Yu 0001 |
ICIP | 2 |
| 2021 | Point Cloud Projection and Multi-Scale Feature Fusion Network Based Blind Quality Assessment for Colored Point CloudsabstractWith the wide applications of colored point cloud (CPC) in many fields, many attentions have been paid to CPC's distortions caused by its compression and reconstruction. How to effectively evaluate the visual quality of CPC has become an urgent issue to be resolved. In this paper, a Point cloud projection and Multi-scale feature fusion network based Blind Visual Quality Assessment method (denoted as PM-BVQA) is proposed for CPC. CPC in 3D space is first projected into 2D color projection map and geometric projection map, then a multi-scale feature fusion network is designed to blindly evaluate the visual quality of CPC. The proposed PM-BVQA method includes three modules, that is, joint color-geometric feature extractor, two-stage multi-scale feature fusion, and spatial pooling module. Considering the multi-channel characteristics of human visual system (HVS), unimodal features of different scales are obtained by joint color-geometric feature extractor from the color and geometric projection maps. The fusion of the unimodal color and geometric features is carried out to capture the cross-modal complementary information between these two types of information. By integrating cross-modal fused features at different scales, the complementary relationships between different channels of HVS are simulated. The spatial pooling module takes into account the attention mechanism of HVS and realizes the weighted summation of local regional quality to obtain the final global quality score of CPC. A subjective CPC database with coding distortion is used to verify the effectiveness of the proposed method, and the experimental results show that the proposed blind quality assessment method is more consistent with the subjective visual perception than the existing quality assessment methods. Wenxu Tao, Gangyi Jiang, Zhidi Jiang, Mei Yu 0001 |
ACM Multimedia | 2 |
| 2021 | FQM-GC: Full-reference Quality Metric for Colored Point Cloud Based on Graph Signal Features and Color FeaturesabstractColored Point Cloud (CPC) is often distorted in the processes of its acquisition, processing, and compression, so reliable quality assessment metrics are required to estimate the perception of distortion of CPC. We propose a Full-reference Quality Metric for colored point cloud based on Graph signal features and Color features (FQM-GC). For geometric distortion, the normal and coordinate information of the sub-clouds divided via geometric segmentation is used to construct their underlying graphs, then, the geometric structure features are extracted. For color distortion, the corresponding color statistical features are extracted from regions divided with color attribution. Meanwhile, the color features of different regions are weighted to simulate the visual masking effect. Finally, all the extracted features are formed into a feature vector to estimate the quality of CPCs. Experimental results on three databases (CPCD2.0, IRPC and SJTU-PCQA) show that the proposed metric FQM-GC is more consistent with human visual perception. Ke-Xin Zhang, Gangyi Jiang, Mei Yu 0001 |
MMAsia | 2 |
| 2021 | Strong ghost removal in multi-exposure image fusion using hole-filling with exposure congruency
Mei Yu 0001, Gangyi Jiang, Zhiyong Pan, Zongju Peng |
J. Vis. Commun. Image Represent. | 3 |
| 2021 | No-reference light field image quality assessment based on depth, structural and angular information
Jianjun Xiang, Gangyi Jiang, Mei Yu 0001, Yongqiang Bai, Zhongjie Zhu |
Signal Process. | 2 |
| 2021 | Reversible data hiding scheme for high dynamic range images based on multiple prediction error expansion
Yongqiang Bai, Gangyi Jiang, Zhongjie Zhu, Haiyong Xu, Yang Song 0015 |
Signal Process. Image Commun. | 2 |
| 2021 | Inter-layer correlation-based adaptive bit allocation for enhancement layer in scalable high efficiency video coding
Zongju Peng, Dongrong Jiang, Chao Huang 0008, Gangyi Jiang, Mei Yu 0001 |
Signal Process. Image Commun. | 5 |
| 2021 | Viewport Perception Based Blind Stereoscopic Omnidirectional Image Quality AssessmentabstractCompared with traditional 2D images, stereoscopic omnidirectional images (SOIs) usually have more complex perceptual factors due to the particularities of imaging and display, making the objective quality assessment of SOIs challenging. In this paper, we construct a large and diverse subjective SOIs database named as NBU-SOID for further research demand. And then, we propose a viewport perception based blind SOIs quality assessment (VP-BSOIQA) method by considering the impacts of viewport, user behavior and stereoscopic perception on human visual system, which is mainly composed of binocular perception model (BPM) and omnidirectional perception model (OPM). In the BPM, a binocular combination perception map is generated by the dimension reduction of stereopair and the weighting of binocular energy to reflect the binocular masking effect. In the OPM, several viewports are first created to ensure the consistency of evaluation objects. Then, the intra-viewport and inter-viewport weighting factors are designed with the common influences of visual attention and peripheral vision sensitivity to aggregate the novel multi-orientation structural features extracted from all potential viewports. Experimental results on the NBU-SOID and SOLID databases demonstrate that BPM and OPM can be robustly combined with the existing 2D image quality assessment (IQA) methods, thus averagely achieving 10.2% and 12.2% performance gain in terms of SRCC, respectively. In addition, the proposed VP-BSOIQA method outperforms the state-of-the-art blind IQA methods in predicting the quality of SOIs. Yubin Qi, Gangyi Jiang, Mei Yu 0001, Yun Zhang 0002, Yo-Sung Ho |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | Pseudo Video and Refocused Images-Based Blind Light Field Image Quality AssessmentabstractThe commercial light field camera is able to capture four-dimensional Light Field Image (LFI), which can be visualized to LFI contents on 2D displays by means of the Pseudo Video (PV) or the Refocused Images (RIs) generated with the refocusing function of LFI. However, the quality degradation of LFI will affect user’s visual experience of LFI contents. Hence, it is crucial to develop an effective LFI quality assessment method to monitor the LFI quality. Most existing subjective databases of LFI use PV and RIs visualization techniques to assess the quality of LFI. Therefore, as the way of presenting LFI on 2D display, PV and RIs are closely related to the subjective perception of LFI by human eyes. Based on these two visualization techniques, this article proposes a novel PV and RIs based blind LFI quality assessment method, in which the feature extraction is divided into two parts. In the first part, the PV’s structure, motion and disparity information are extracted with multi-scale and multi-directional Shearlet transform. In the other part, the spatial structure, depth and semantic information of the RIs are obtained. Finally, support vector regression is used to nonlinear map the perceptual features to quality score of LFI. The experimental results on four LFI databases show that the proposed method has better correlation with human visual perception, compared with the classical 2D image quality assessment methods as well as the state-of-the-art LFI quality assessment methods. Jianjun Xiang, Mei Yu 0001, Gangyi Jiang, Haiyong Xu, Yang Song 0015, Yo-Sung Ho |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2021 | Online Learning-Based Multi-Stage Complexity Control for Live Video CodingabstractHigh Efficiency Video Coding (HEVC) can significantly improve the compression efficiency in comparison with the preceding H.264/Advanced Video Coding (AVC) but at the cost of extremely high computational complexity. Hence, it is challenging to realize live video applications on low-delay and power-constrained devices, such as the smart mobile devices. In this article, we propose an online learning-based multi-stage complexity control method for live video coding. The proposed method consists of three stages: multi-accuracy Coding Unit (CU) decision, multi-stage complexity allocation, and Coding Tree Unit (CTU) level complexity control. Consequently, the encoding complexity can be accurately controlled to correspond with the computing capability of the video-capable device by replacing the traditional brute-force search with the proposed algorithm, which properly determines the optimal CU size. Specifically, the multi-accuracy CU decision model is obtained by an online learning approach to accommodate the different characteristics of input videos. In addition, multi-stage complexity allocation is implemented to reasonably allocate the complexity budgets to each coding level. In order to achieve a good trade-off between complexity control and rate distortion (RD) performance, the CTU-level complexity control is proposed to select the optimal accuracy of the CU decision model. The experimental results show that the proposed algorithm can accurately control the coding complexity from 100% to 40%. Furthermore, the proposed algorithm outperforms the state-of-the-art algorithms in terms of both accuracy of complexity control and RD performance. Chao Huang 0008, Zongju Peng, Yong Xu 0001, Qiuping Jiang, Yun Zhang 0002, Gangyi Jiang, Yo-Sung Ho |
IEEE Trans. Image Process. | 7 |
| 2021 | Cubemap-Based Perception-Driven Blind Quality Assessment for 360-degree Imagesabstractimage can be represented with different formats, such as the equirectangular projection (ERP) image, viewport images or spherical image, for its different processing procedures and applications. Accordingly, the 360-degree image quality assessment (360-IQA) can be performed on these different formats. However, the performance of 360-IQA with the ERP image is not equivalent with those with the viewport images or spherical image due to the over-sampling and the resulted obvious geometric distortion of ERP image. This imbalance problem brings challenge to ERP image based applications, such as 360-degree image/video compression and assessment. In this paper, we propose a new blind 360-IQA framework to handle this imbalance problem. In the proposed framework, cubemap projection (CMP) with six inter-related faces is used to realize the omnidirectional viewing of 360-degree image. A multi-distortions visual attention quality dataset for 360-degree images is firstly established as the benchmark to analyze the performance of objective 360-IQA methods. Then, the perception-driven blind 360-IQA framework is proposed based on six cubemap faces of CMP for 360-degree image, in which human attention behavior is taken into account to improve the effectiveness of the proposed framework. The cubemap quality feature subset of CMP image is first obtained, and additionally, attention feature matrices and subsets are also calculated to describe the human visual behavior. Experimental results show that the proposed framework achieves superior performances compared with state-of-the-art IQA methods, and the cross dataset validation also verifies the effectiveness of the proposed framework. In addition, the proposed framework can also be combined with new quality feature extraction method to further improve the performance of 360-IQA. All of these demonstrate that the proposed framework is effective in 360-IQA and has a good potential for future applications. Hao Jiang 0014, Gangyi Jiang, Mei Yu 0001, Yun Zhang 0002, You Yang 0002, Zongju Peng |
IEEE Trans. Image Process. | 2 |
| 2021 | Blind Quality Assessment of Screen Content Images Via Macro-Micro Modeling of Tensor Domain DictionaryabstractScreen content images (SCIs) have been rapidly and widely applied in interactive multimedia applications. The problem of quality assessment for SCIs is an interesting research topic. Most of the existing methods use subjective and independent features in gray domain to predict the image quality, which cannot comprehensively characterize the image properties or lack unified mathematical explanation for SCIs. To address these problems, we propose a novel blind quality assessment method based on macro-micro modeling of tensor domain dictionary for SCIs in this article. In the proposed method, the tensor decomposition is explored first to avoid the loss of color information, and then a target dictionary is learned more effectively with the principal components. Furthermore, a macro-micro model is established to characterize the micro and macro features in the target dictionary space, which can provide a systematic mathematical interpretation for feature extraction. For the micro features, a log-normal pooling scheme is designed to enhance the effectiveness of feature aggregation by analyzing the particularity of the statistical distribution of sparse codes. Additionally, the statistical properties are mainly discussed and studied based on the Bernoulli law of large numbers, and then a reliable macro feature is generated to describe the relationship between the statistical distribution and quality degradation of SCIs. Experimental results determined by using three public SCI databases show that the proposed method can perform better than relevant existing methods in the prediction of the visual quality of SCIs, especially in terms of the generalization for distortion type and interpretability for feature generation. Yongqiang Bai, Zhongjie Zhu, Gangyi Jiang, Huifang Sun |
IEEE Trans. Multim. | 3 |
| 2021 | Semi-Reference Sonar Image Quality Assessment Based on Task and Visual PerceptionabstractIn submarine and underwater detection tasks, conventional optical imaging and analysis methods are not universally applicable due to the limited penetration depth of visible light. Instead, sonar imaging has become a preferred alternative. However, the capture and transmission conditions in complicated and dynamic underwater environments inevitably lead to visual quality degradation of sonar images, which might also impede further recognition, analysis and understanding. To measure this quality decrease and provide a solid quality indicator for sonar image enhancement, we propose a task- and perception-oriented sonar image quality assessment (TPSIQA) method, in which a semi-reference (SR) approach is applied to adapt to the limited bandwidth of underwater communication channels. In particular, we exploit reduced visual features that are critical for both human perception of and object recognition in sonar images. The final quality indicator is obtained through ensemble learning, which aggregates an optimal subset of multiple base learners to achieve both high accuracy and a high generalization ability. In this way, we are able to develop a compact but generalized quality metric using a small database of sonar images. Experimental results demonstrate competitive performance, high efficiency, and strong robustness of our method compared to the latest available image quality metrics. Ke Gu 0001, Tiesong Zhao, Gangyi Jiang, Patrick Le Callet |
IEEE Trans. Multim. | 4 |
| 2021 | Transformation-Aware Similarity Measurement for Image Retargeting Quality Assessment via Bidirectional RewarpingabstractImage retargeting is an effective way to adapt images for target displays with different aspect ratios and sizes. Meanwhile, effective image retargeting quality assessment (IRQA) is important for optimizing the image retargeting operations. In this paper, we propose a transform-aware similarity (TRASIM) measurement metric for IRQA, including bidirectional geometric distortion measurement, bidirectional information loss measurement, and global salient structure distortion measurement. The main innovation of the TRASIM is to build a universal framework to establish the similarity transformation via bidirectional rewarping to simulate different types of retargeting operators. Based on the similarity transformation, geometric distortion and content loss are measured to determine the retargeting quality. Experimental results on two widely used databases (CUHK and RetargetMe) indicate that the proposed TRASIM has higher consistency with subjective ranks, compared with the state-of-the-art IRQA metrics. Feng Shao 0001, Zhenqi Fu, Qiuping Jiang, Gangyi Jiang, Yo-Sung Ho |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2020 | VBLFI: Visualization-Based Blind Light Field Image Quality AssessmentabstractLight field image (LFI) contains the intensity and direction information of the scene. The huge amount of data and different visualization methods of LFI brings great challenges to LFI processing and its blind LFI quality assessment. This paper analyzes the human visual perception from the LFI's visualization, and proposes a novel Visualization-based Blind Light Field Image quality assessment (VBLFI) model. With LFI's visualization and its depth cues, we compute mean difference image from LFI to reduce redundant information of LFI and to describe depth and structural information of LFI. LFI's multi-scale expression with curvelet transform is used to reflect the multi-channel characteristics of human visual system. So, the corresponding natural scene statistical features and energy features are extracted from the mean difference image and sub-aperture images of LFI in curvelet transform domain to form the feature vector, further used to predict the LFI quality. Compared to the representative 2D image quality assessment models and the state-of-the-art LFIQA models, the proposed VBLFI model has better prediction accuracy and stability in the public LFI databases. Jianjun Xiang, Mei Yu 0001, Hua Chen 0004, Haiyong Xu, Yang Song 0015, Gangyi Jiang |
ICME | 6 |
| 2020 | Blind quality assessment for 3D synthesised video with binocular asymmetric distortionabstractDuring the process of watching 3D synthesised video (3D‐SV) and switching viewpoints, there is a case of asymmetric distortion, the left(right) viewpoint is a synthesised video generated by rendering technique, and the right(left) viewpoint is a real video taken by the camera. How to accurately estimate the quality of 3D‐SV with binocular asymmetric distortions is a new and challenging problem. Aiming at this problem, a blind quality assessment method for 3D‐SV with binocular asymmetric distortions is proposed. Firstly, the local edge deformations of synthesised videos at different scales are measured by calculating their standard deviations. Secondly, the global naturalness of synthesised videos is computed by analysing their natural statistical characteristics. Thirdly, a strategy for fusing left and right quality scores is proposed, which considers their texture information in different directions. Finally, the random forest is used to obtain an objective quality score. The experimental results show the superiority of the proposed method on asymmetry 3D‐SV database. Shuainan Cui, Zongju Peng, Wenhui Zou, Gangyi Jiang, Mei Yu 0001 |
IET Image Process. | 5 |
| 2020 | Multi-exposure high dynamic range imaging with informative content enhanced network
Zhiyong Pan, Mei Yu 0001, Gangyi Jiang, Haiyong Xu, Zongju Peng |
Neurocomputing | 3 |
| 2020 | Blind tone mapped image quality assessment with image segmentation and visual perception
Biwei Chi, Mei Yu 0001, Gangyi Jiang, Zhouyan He, Zongju Peng |
J. Vis. Commun. Image Represent. | 3 |
| 2020 | A fast CU size decision algorithm for VVC intra prediction based on support vector machine
Zongju Peng, Gangyi Jiang |
Multim. Tools Appl. | 4 |
| 2020 | Perceptual objective quality assessment of stereoscopic stitched images
Weiqing Yan, Guanghui Yue 0001, Yuming Fang 0001, Hua Chen 0004, Chang Tang, Gangyi Jiang |
Signal Process. | 6 |
| 2020 | Latitude and binocular perception based blind stereoscopic omnidirectional image quality assessment for VR system
Gangyi Jiang, Mei Yu 0001, Yubin Qi |
Signal Process. | 2 |
| 2020 | Perceived depth quality - preserving visual comfort improvement method for stereoscopic 3D images
Hongwei Ying, Mei Yu 0001, Gangyi Jiang, Zongju Peng |
Signal Process. | 3 |
| 2020 | Low-Complexity CTU Partition Structure Decision and Fast Intra Mode Decision for Versatile Video CodingabstractQuadtree with nested multi-type tree (QTMT) partition structure is an efficient improvement in versatile video coding (VVC) over the quadtree (QT) structure in the advanced high-efficiency video coding (HEVC) standard. With the exception of the recursive QT partition structure, recursive multi-type tree partition is applied to each leaf node, which generates more flexible block sizes. Besides, intra prediction modes are extended from 35 to 67 so as to satisfy various texture patterns. These newly developed techniques achieve high coding efficiency but also result in very high computational complexity. To tackle this problem, we propose a fast intra-coding algorithm consisting of low-complexity coding tree units (CTU) structure decision and fast intra mode decision in this paper. The contributions of the proposed algorithm lie in the following aspects: 1) the new block size and coding mode distribution features are first explored for a reasonable fast coding scheme; 2) a novel fast QTMT partition decision framework is developed, which can determine the partition decision on both QT and multi-type tree with a novel cascade decision structure; and 3) fast intra mode decision with gradient descent search is introduced, while the best initial search point and search step are also investigated in this paper. The simulation results show that the complexity reduction of the proposed algorithm is up to 70% compared to VVC reference software (VTM), and averagely 63% encoding time saving is achieved with 1.93% BDBR increasing. Such results demonstrate that our method yields a superior performance in terms of computational complexity and compression quality compared to the state-of-the-art methods. Hao Yang 0008, Liquan Shen, Xinchao Dong, Ping An 0001, Gangyi Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2019 | New Stereo High Dynamic Range Imaging Method Using Generative Adversarial NetworksabstractStereo high dynamic range (HDR) image/video can be generated by using a pair of stereo cameras with different exposure parameters. This paper proposes a new stereo HDR imaging method using generative adversarial networks (GAN) with a low dynamic range (LDR) stereo imaging system. It is assumed here that the left-view (LV) image is under-exposed and the right-view (RV) image is overexposed. First, a view exposure transfer GAN (VET-GAN) is constructed to transfer exposure information of the RV image to the LV image to generate the multi-exposure LV images, and then an HDR fusion GAN is constructed to fuse the generated multi-exposure LV images into an LV HDR image. Similarly, an RV HDR image can be generated using the same way to form a stereo HDR image pair. The experimental results show that the proposed method can obtain stereo HDR images with high visual quality and effectively avoid the ghost artifacts caused by parallax. Yeyao Chen, Mei Yu 0001, Ken Chen 0003, Gangyi Jiang, Yang Song 0015, Zongju Peng |
ICIP | 4 |
| 2019 | Encoding Complexity Control for Live Video Applications: An Interpretable Machine Learning ApproachabstractIn this paper, we propose an interpretable machine learning-based complexity control method for efficiently im-plementing HEVC on live video applications with different computing capacities and limited powers. Specifically, a complexity allocation method is designed to reasonably assign the complexity resources. Then, a multi-accuracy Coding Unit (CU) decision model is obtained by interpret-ably adjusting the parameters to efficiently and flexibly achieve a tradeoff between encoding complexity and rate distortion performance. Finally, a coding tree unit-level complexity control method is proposed to select appropri-ate accuracy of the CU decision model for making the en-coding complexity approach the target. The experimental results show that the proposed method outperforms state-of-the-art methods in terms of accuracy and encoding efficiency. Chao Huang 0008, Zongju Peng, Qiuping Jiang, Gangyi Jiang |
ICME | 6 |
| 2019 | Reconstruction Distortion Oriented Light Field Image Dataset for Visual CommunicationabstractAs a representation of three dimensional scenes, light field has received increasing attention. In light field image processing, reconstruction method plays an important role, which can not only produce dense views to improve the spatial and angular resolution of the light field, but also effectively reduce the transmission data. However, the reconstruction methods inevitably reduce the quality of light field images, so the corresponding light field image quality assessment is necessary. In this paper, a reconstruction distortion oriented light field image dataset is firstly established, with several different reconstruction methods and the corresponding subjective evaluation scores. Secondly, the subjective scoring results of source sequences and their types of distorted versions are compared and analyzed. Finally, the dataset is evaluated with the existing objective quality assessment metrics. Experimental results show that different reconstruction methods have different preferences on the input light field resolution, and the performance of the state-of-the-art objective quality metrics can be improved. Zhijiao Huang, Mei Yu 0001, Gangyi Jiang, Ken Chen 0003, Zongju Peng |
ISNCC | 3 |
| 2019 | Simultaneous object size and depth adjustment for stereoscopic 3D images
Feng Shao 0001, Yanjia Fei, Randi Fu, Gangyi Jiang, Yo-Sung Ho |
Inf. Sci. | 4 |
| 2019 | End-to-end single image enhancement based on a dual network cascade model
Yeyao Chen, Mei Yu 0001, Gangyi Jiang, Zongju Peng |
J. Vis. Commun. Image Represent. | 3 |
| 2019 | Convolutional neural networks-based stereo image reversible data hiding method
Ting Luo 0001, Gangyi Jiang, Mei Yu 0001, Caiming Zhong, Haiyong Xu, Zhiyong Pan |
J. Vis. Commun. Image Represent. | 2 |
| 2019 | Quality assessment of stereoscopic video in free viewpoint video system
Zongju Peng, Shipei Wang, Wenhui Zou, Gangyi Jiang, Mei Yu 0001 |
J. Vis. Commun. Image Represent. | 5 |
| 2019 | Lossless fragile watermarking algorithm in compressed domain for multiview video coding
Wei Gao 0013, Gangyi Jiang, Mei Yu 0001, Ting Luo 0001 |
Multim. Tools Appl. | 2 |
| 2019 | A novel robust color image watermarking method using RGB correlations
Fangyan Zhang, Ting Luo 0001, Gangyi Jiang, Mei Yu 0001, Haiyong Xu, Wujie Zhou |
Multim. Tools Appl. | 3 |
| 2019 | Learning content-specific codebooks for blind quality assessment of screen content images
Yongqiang Bai, Mei Yu 0001, Qiuping Jiang, Gangyi Jiang, Zhongjie Zhu |
Signal Process. | 4 |
| 2019 | Robust high dynamic range color image watermarking method based on feature map extraction
Ting Luo 0001, Gangyi Jiang, Mei Yu 0001, Haiyong Xu, Wei Gao 0013 |
Signal Process. | 2 |
| 2019 | Fast inter-frame prediction in multi-view video coding based on perceptual distortion threshold model
Gangyi Jiang, Baozhen Du, Shuqing Fang, Mei Yu 0001, Feng Shao 0001, Zongju Peng |
Signal Process. Image Commun. | 1 |
| 2019 | Multiple classifier-based fast coding unit partition for intra coding in future video coding
Zongju Peng, Chao Huang 0008, Gangyi Jiang, Mei Yu 0001 |
Signal Process. Image Commun. | 4 |
| 2019 | BLIQUE-TMI: Blind Quality Evaluator for Tone-Mapped Images Based on Local and Global Feature AnalysesabstractHigh dynamic range (HDR) image, which has a powerful capacity to represent the wide dynamic range of real-world scenes, has been receiving attention from both academic and industrial communities. Although HDR imaging devices have become prevalent, the display devices for HDR images are still limited. To facilitate the visualization of HDR images in standard low dynamic range displays, many different tone mapping operators (TMOs) have been developed. To create a fair comparison of different TMOs, this paper proposes a BLInd QUality Evaluator to blindly predict the quality of Tone-Mapped Images (BLIQUE-TMI) without accessing the corresponding HDR versions. BLIQUE-TMI measures the quality of TMIs by considering the following aspects: 1) visual information; 2) local structure; and 3) naturalness. To be specific, quality-aware features related to the former two aspects are extracted in a local manner based on sparse representation, while quality-aware features related to the third aspect are derived based on global statistics modeling in both intensity and color domains. All the extracted local and global quality-aware features constitute a final feature vector. An emergent machine learning technique, i.e., extreme learning machine, is adopted to learn a quality predictor from feature space to quality space. The superiority of BLIQUE-TMI to several leading blind IQA metrics is well demonstrated on two benchmark databases. Qiuping Jiang, Feng Shao 0001, Weisi Lin, Gangyi Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2019 | Efficient Shape Coding for Object-Based 3D Video ApplicationsabstractShape is a popular way to define objects and shape coding is a key technique for object-based 3D video applications. In this paper, the issue of efficient shape coding for object-based 3D video applications is addressed, and a novel contour-based and chain-represented scheme is proposed. For a given 3D shape video, contour extraction and preprocessing are first implemented followed by chain-based representation. Then, to achieve high coding efficiency, a chain-based prediction and compensation technique is developed based on joint motion-compensated prediction and disparity-compensated prediction to effectively exploit the intra-view temporal correlation and the inter-view spatial correlation. Experiments are conducted, and the results demonstrate that the proposed scheme is more efficient than the existing methods, including state-of-the-art methods. Zhongjie Zhu, Yuer Wang, Gangyi Jiang, Yueping Yang |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2019 | Unified No-Reference Quality Assessment of Singly and Multiply Distorted Stereoscopic ImagesabstractA challenging problem in the no-reference quality assessment of multiply distorted stereoscopic images (MDSIs) is to simulate the monocular and binocular visual properties under a mixed type of distortions. Due to the joint effects of multiple distortions in MDSIs, the underlying monocular and binocular visual mechanisms have different manifestations with those of singly distorted stereoscopic images (SDSIs). This paper presents a unified no-reference quality evaluator for SDSIs and MDSIs by learning monocular and binocular local visual primitives (MB-LVPs). The main idea is to learn MB-LVPs to characterize the local receptive field properties of the visual cortex in response to SDSIs and MDSIs. Furthermore, we also consider that the learning of primitives should be performed in a task-driven manner. For this, two penalty terms including reconstruction error and quality inconsistency are jointly minimized within a supervised dictionary learning framework, generating a set of quality-oriented MB-LVPs for each single and multiple distortion modality. Given an input stereoscopic image, feature encoding is performed using the learned MB-LVPs as codebooks, resulting in the corresponding monocular and binocular responses. Finally, responses across all the modalities are fused with probabilistic weights which are determined by the modality-specific sparse reconstruction errors, yielding the final monocular and binocular features for quality regression. The superiority of our method has been verified on several SDSI and MDSI databases. Qiuping Jiang, Feng Shao 0001, Wei Gao 0003, Gangyi Jiang, Yo-Sung Ho |
IEEE Trans. Image Process. | 5 |
| 2019 | Statistical Early Termination and Early Skip Models for Fast Mode Decision in HEVC INTRA CodingabstractIn this article, statistical Early Termination (ET) and Early Skip (ES) models are proposed for fast Coding Unit (CU) and prediction mode decision in HEVC INTRA coding, in which three categories of ET and ES sub-algorithms are included. First, the CU ranges of the current CU are recursively predicted based on the texture and CU depth of the spatial neighboring CUs. Second, the statistical model based ET and ES schemes are proposed and applied to optimize the CU and INTRA prediction mode decision, in which the coding complexities over different decision layers are jointly minimized subject to acceptable rate-distortion degradation. Third, the mode correlations among the INTRA prediction modes are exploited to early terminate the full rate-distortion optimization in each CU decision layer. Extensive experiments are performed to evaluate the coding performance of each sub-algorithm and the overall algorithm. Experimental results reveal that the overall proposed algorithm can achieve 45.47% to 74.77%, and 58.09% on average complexity reduction, while the overall Bjøntegaard delta bit rate increase and Bjøntegaard delta peak signal-to-noise ratio degradation are 2.29% and −0.11 dB, respectively. Yun Zhang 0002, Na Li 0015, Sam Kwong, Gangyi Jiang, Huanqiang Zeng |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2018 | No-Reference Hdr Image Quality Assessment Method Based on Tensor SpaceabstractThe full-reference image quality assessment (IQA) method are limited in practical applications. Here we propose a no-reference quality assessment method for high dynamic range (HDR) images based on tensor space. First, the tensor decomposition is used to generate three feature maps of an HDR image, considering color and structure information of the HDR image. Second, for a given HDR image, the corresponding multi -scale manifold structure features are extracted from the first feature map. For the second and third feature maps of the HDR image, multi-scale contrast features are extracted. Finally, the extracted features are aggregated by support vector regression to obtain the objective quality score of the HDR image. Experimental results show that the proposed method is superior to some representative full and no-reference methods, and even superior to the full-reference HDR IQA method, HDR-VDP-2.2, on the Nantes database. The proposed method has a higher consistency with human visual perception. Feifan Guan, Gangyi Jiang, Yang Song 0015, Mei Yu 0001, Zongju Peng |
ICASSP | 2 |
| 2018 | 3D visual discomfort predictor based on subjective perceived-constraint sparse representation in 3D display system
Haiyong Xu, Gangyi Jiang, Mei Yu 0001, Ting Luo 0001, Zongju Peng, Feng Shao 0001, Hao Jiang 0014 |
Future Gener. Comput. Syst. | 2 |
| 2018 | Perceptual stereoscopic image quality assessment method with tensor decomposition and manifold learningabstractPerceptual quality assessment of stereoscopic images is a challenge in three‐dimensional video systems. Existing studies suggest that simply averaging the quality of left and right views can effectively predict the quality of symmetrically distorted stereoscopic images, but prediction deviation occurs in the case of asymmetrically distorted stereoscopic images. Most previous stereoscopic image quality assessment (SIQA) methods have been based only on the luminance component of the images; in addition, the basis of human visual perception is critical to image quality assessment and lies on the low‐dimensional manifold. Inspired by this, a new perceptual SIQA method is proposed, which includes two stages: training stage and quality prediction stage. In the training stage, the authors apply Tucker decomposition to RGB images to reduce dimensions along colour channels to produce training sets, and the projection matrix is obtained through manifold learning. In the quality prediction stage, considering the binocular visual characteristics of visual perception, the overall stereoscopic estimate depends on the monocular image quality via a local energy ratio based pooling strategy and cyclopean based binocular quality. Extensive experiments on three available benchmark databases demonstrate that the proposed metric has better performance and achieves highly consistent alignment with subjective assessment compared with state‐of‐the‐art SIQA metrics. Gangyi Jiang, Meiling He, Mei Yu 0001, Feng Shao 0001, Zongju Peng |
IET Image Process. | 1 |
| 2018 | Local and global sparse representation for no-reference quality assessment of stereoscopic images
Fucui Li, Feng Shao 0001, Qiuping Jiang, Randi Fu, Gangyi Jiang, Mei Yu 0001 |
Inf. Sci. | 5 |
| 2018 | No reference stereo video quality assessment based on motion feature in tensor decomposition domain
Gangyi Jiang, Mei Yu 0001, Feng Shao 0001, Zongju Peng |
J. Vis. Commun. Image Represent. | 1 |
| 2018 | Fast intra coding algorithm for HEVC based on depth range prediction and mode reduction
Defu Jin, Zongju Peng, Gangyi Jiang, Mei Yu 0001, Hua Chen 0004 |
Multim. Tools Appl. | 4 |
| 2018 | Sparse recovery based reversible data hiding method using the human visual system
Ting Luo 0001, Gangyi Jiang, Mei Yu 0001, Haiyong Xu, Wei Gao 0013 |
Multim. Tools Appl. | 2 |
| 2018 | Learning a referenceless stereopair quality engine with deep nonnegativity constrained sparse autoencoder
Qiuping Jiang, Feng Shao 0001, Weisi Lin, Gangyi Jiang |
Pattern Recognit. | 4 |
| 2018 | Quality assessment method based on exposure condition analysis for tone-mapped high-dynamic-range images
Yang Song 0015, Gangyi Jiang, Mei Yu 0001, Zongju Peng |
Signal Process. | 2 |
| 2018 | Towards a tone mapping-robust watermarking algorithm for high dynamic range image based on spatial activity
Yongqiang Bai, Gangyi Jiang, Mei Yu 0001, Zongju Peng |
Signal Process. Image Commun. | 2 |
| 2018 | Toward Domain Transfer for No-Reference Quality Prediction of Asymmetrically Distorted Stereoscopic ImagesabstractWe have presented a no-reference quality prediction method for asymmetrically distorted stereoscopic images, which aims to transfer the information from source feature domain to its target quality domain using a label consistent K-singular value decomposition classification framework. To this end, we construct a category-deviation database for dictionary learning that assigns a label for each stereoscopic image to indicate if it is noticeable or unnoticeable by human eyes. Then, by incorporating a category consistent term into the objective function, we learn view-specific feature and quality dictionaries to establish a semantic framework between the source feature domain and the target quality domain. The quality pooling is comparatively simple and only needs to estimate the quality score based on the classification probability. The experimental results demonstrate the effectiveness of our blind metric. Feng Shao 0001, Zhuqing Zhang, Qiuping Jiang, Weisi Lin, Gangyi Jiang |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2018 | Effective Data Driven Coding Unit Size Decision Approaches for HEVC INTRA CodingabstractHigh Efficiency Video Coding (HEVC) INTRA coding improves compression efficiency by adopting advanced coding technologies, such as multi-level quad-tree block partitioning and up to 35-mode INTRA prediction. However, it significantly increases the coding complexity, memory access, and power consumption, which goes against its widely applications, especially for ultra-high definition and/or mobile video applications. To tackle this problem, we propose effective data driven coding unit (CU) size decision approaches for HEVC INTRA coding, which consists of two stages of support vector machine-based fast INTRA CU size decision schemes at four CU decision layers. At the first stage classification, a three output classifier with offline learning is developed to early terminate the CU size decision or early skip checking the current CU depth. As for the samples that neither early skipped nor early terminated, the second stage of binary classification, which learns online from previous coded frames, is proposed to further refine the CU size decision. Representative features for the CU size decision are explored at different decision layers and stages of classifications. Finally, the optimal parameters derived from the training data are achieved to reasonably allocate complexity among different CU layers at given total rate-distortion degradation constraint. Extensive experiments show that the proposed overall algorithm can achieve 27.95%–80.53% and 52.48% on average complexity reduction for the CU size decision as compared with the original HM16.7 model. Meanwhile, the average Bjonteggard delta peak-signal-to-noise ratio degradation is only −0.08 dB, which is negligible. The overall performance of the proposed algorithm outperforms the state-of-the-art benchmark schemes. Yun Zhang 0002, Zhaoqing Pan, Na Li 0015, Xu Wang 0006, Gangyi Jiang, Sam Kwong |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2018 | Learning Sparse Representation for Objective Image Retargeting Quality AssessmentabstractThe goal of image retargeting is to adapt source images to target displays with different sizes and aspect ratios. Different retargeting operators create different retargeted images, and a key problem is to evaluate the performance of each retargeting operator. Subjective evaluation is most reliable, but it is cumbersome and labor-consuming, and more importantly, it is hard to be embedded into online optimization systems. This paper focuses on exploring the effectiveness of sparse representation for objective image retargeting quality assessment. The principle idea is to extract distortion sensitive features from one image (e.g., retargeted image) and further investigate how many of these features are preserved or changed in another one (e.g., source image) to measure the perceptual similarity between them. To create a compact and robust feature representation, we learn two overcomplete dictionaries to represent the distortion sensitive features of an image. Features including local geometric structure and global context information are both addressed in the proposed framework. The intrinsic discriminative power of sparse representation is then exploited to measure the similarity between the source and retargeted images. Finally, individual quality scores are fused into an overall quality by a typical regression method. Experimental results on several databases have demonstrated the superiority of the proposed method. Qiuping Jiang, Feng Shao 0001, Weisi Lin, Gangyi Jiang |
IEEE Trans. Cybern. | 4 |
| 2018 | Optimizing Multistage Discriminative Dictionaries for Blind Image Quality AssessmentabstractState-of-the-art algorithms for blind image quality assessment (BIQA) typically have two categories. The first category approaches extract natural scene statistics (NSS) as features based on the statistical regularity of natural images. The second category approaches extract features by feature encoding with respect to a learned codebook. However, several problems need to be addressed in existing codebook-based BIQA methods. First, the high-dimensional codebook-based features are memory-consuming and have the risk of over-fitting. Second, there is a semantic gap between the constructed codebook by unsupervised learning and image quality. To address these problems, we propose a novel codebook-based BIQA method by optimizing multistage discriminative dictionaries (MSDDs). To be specific, MSDDs are learned by performing the label consistent K-SVD (LC-KSVD) algorithm in a stage-by-stage manner. For each stage, a new quality consistency constraint called “quality-discriminative regularization” term is introduced and incorporated into the reconstruction error term to form a unified objective function, which can be effectively solved by LC-KSVD for discriminative dictionary learning. Then, the latter stage takes the reconstruction residual data in the former stage as input based on which LC-KSVD is repeatedly performed until the final stage is reached. Once the MSDDs are learned, multistage feature encoding is performed to extract feature codes. Finally, the feature codes are concatenated across all stages and aggregated over the entire image for quality prediction via regression. The proposed method has been evaluated on five databases and experimental results well confirm its superiority over existing relevant BIQA methods. Qiuping Jiang, Feng Shao 0001, Weisi Lin, Ke Gu 0001, Gangyi Jiang, Huifang Sun |
IEEE Trans. Multim. | 5 |
| 2018 | Multistage Pooling for Blind Quality Prediction of Asymmetric Multiply-Distorted Stereoscopic ImagesabstractQuality prediction for asymmetric multiply-distorted stereoscopic images (MDSIs) confronts more challenges than previous stereoscopic image quality assessment (SIQA) issues, whereas the existing no-reference SIQA methods have been limited to understand the asymmetric distortions and multiple distortions simultaneously for general-purpose blind quality prediction. In this paper, we propose a multistage pooling (MUSP) model for quality prediction of asymmetric MDSIs. In the training stage, we establish multimodal sparse representation framework for phase and amplitude components, respectively. In the testing stage, we use an MUSP strategy to simulate the pooling procedure undergoing multimodal quality pooling, feature pooling, binocular pooling, and phase-amplitude quality pooling in order. Experimental results on our new established database (NBU-MDSID Phase-II) demonstrate the effectiveness of our blind metric. Feng Shao 0001, Qiuping Jiang, Gangyi Jiang, Yo-Sung Ho |
IEEE Trans. Multim. | 4 |
| 2018 | No-Reference View Synthesis Quality Prediction for 3-D Videos Based on Color-Depth InteractionsabstractIn a 3-D video system, automatically predicting the quality of synthesized 3-D video based on the inputs of color and depth videos is an urgent but very difficult task, while the existing full-reference methods usually measure the perceptual quality of the synthesized video. In this paper, a high-efficiency view synthesis quality prediction (HEVSQP) metric for view synthesis is proposed. Based on the derived VSQP model that quantifies the influences of color and depth distortions and their interactions in determining the perceptual quality of 3-D synthesized video, color-involved VSQP and depth-involved VSQP indices are predicted, respectively, and are combined to yield an HEVSQP index. Experimental results on our constructed NBU-3D Synthesized Video Quality Database demonstrate that the proposed HEVSOP has good performance evaluated on the entire synthesized video-quality database, compared with other full-reference and no-reference video-quality assessment metrics. Feng Shao 0001, Qizheng Yuan, Weisi Lin, Gangyi Jiang |
IEEE Trans. Multim. | 4 |
| 2018 | Toward a Blind Quality Predictor for Screen Content ImagesabstractBlind quality assessment of screen content images (SCIs) is much challenging than traditional natural images. In this paper, we propose a blind quality predictor for SCIs to explore the issue from the perspective of sparse representation. Specifically, we conduct local sparse representation for the textual and pictorial regions, respectively, and conduct global sparse representation for the global SCIs. Subsequently, the underlying relationship between the feature and quality vectors is bridged in a universal sparse representation framework. The quality pooling is comparatively simply that only need to estimate the local and global quality scores and combine them to a total one. Our experimental results show that the proposed predictor can achieve better prediction performance to be in line with subjective assessment. Feng Shao 0001, Fucui Li, Gangyi Jiang |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2017 | MSFE: Blind image quality assessment based on multi-stage feature encodingabstractBlind image quality assessment (BIQA) methods based on visual codebooks have received much attention due to its prominent generalization capacity across different image domains. Existing codebook-based BIQA methods depend on large-size codebooks and high-dimensional features, which are memory-consuming and have the risk of over-fitting. Thus, it is necessary to design quality metrics with much smaller codebooks. This paper presents a novel multistage feature encoding (MSFE)-based BIQA method which requires much lower dimensional features while preserving comparable or even better performance. To specify, MSFE is performed over multiple cascaded and much smaller sub-codebooks to generate more compact and discriminative features for quality prediction. The latter stage takes the encoding residuals in the former stage as input. We use KSVD and sparse coding for codebook training and feature encoding in the framework, respectively. Finally, the generated sparse feature codes in all stages are combined and aggregated over the entire image for quality prediction via support vector regression (SVR). We evaluate the proposed method on several natural and screen content image databases. The experimental results confirm its superiority in terms of both validity and universality. Qiuping Jiang, Feng Shao 0001, Gangyi Jiang |
ICIP | 3 |
| 2017 | A new tone-mapped image quality assessment approach for high dynamic range imaging systemabstractTone-mapping operators are designed to apply high dynamic range (HDR) images on widely-used low dynamic range (LDR) devices. Developing well-performed tone-mapped image quality assessment (IQA) method is highly desired because traditional IQA method cannot be adopted in cross dynamic range quality measuring. To this end, we proposed a quality assessment method based on image exposure property. Specifically, an image exposure property determination model is utilized to segment HDR image into different exposure region. Then, quality features are extracted according to the distortion characteristics of each exposure region. Finally, the quality of tone-mapped image can be acquired by a trained regression model. Validation experiments on public database show that the proposed method can accurately predict the quality of tone-mapped image. Yang Song 0015, Gangyi Jiang, Hao Jiang 0014, Mei Yu 0001, Feng Shao 0001, Zongju Peng |
ICIP | 2 |
| 2017 | Visual comfort assessment for stereoscopic images based on sparse coding with multi-scale dictionaries
Qiuping Jiang, Feng Shao 0001, Gangyi Jiang, Mei Yu 0001, Zongju Peng |
Neurocomputing | 3 |
| 2017 | Unsupervised segmentation of natural images based on statistical modeling
Zhongjie Zhu, Yu-er Wang, Gangyi Jiang |
Neurocomputing | 3 |
| 2017 | A Mismatch Detection Method Based on Affine Transformation for Stereo Light Microscopy Stereo MatchingabstractFor the light microscopy images that have the characteristics of shallow depth of field, serious distortion and poor resolution, mismatch is a ubiquitous phenomenon. The paper presents a mismatch detection method for the stereo light microscopy stereo matching. Affine transformation matrix and matching constraint condition are calibrated by the calibration board which has the precision solid dots and the motorized stage. Bias vector of affine transformation of each matching pair is taken as the criteria to apply mismatch detection. The experimental results show that the method can detect more mismatching pairs and preserve more matching pairs than the traditional RANSAC method and the epipolar rectification method. Shengli Fan, Mei Yu 0001, Gangyi Jiang, Yigang Wang, Hao Jiang 0014 |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2017 | Virtual view quality assessment based on shift compensation and visual masking effect
Renzhi Jiao, Zongju Peng, Gangyi Jiang, Mei Yu 0001 |
J. Vis. Commun. Image Represent. | 4 |
| 2017 | Stereoscopic image quality assessment by learning non-negative matrix factorization-based color visual characteristics and considering binocular interactions
Gangyi Jiang, Haiyong Xu, Mei Yu 0001, Ting Luo 0001, Yun Zhang 0002 |
J. Vis. Commun. Image Represent. | 1 |
| 2017 | Leveraging visual attention and neural activity for stereoscopic 3D visual comfort assessment
Qiuping Jiang, Feng Shao 0001, Gangyi Jiang, Mei Yu 0001, Zongju Peng |
Multim. Tools Appl. | 3 |
| 2017 | Toward Simultaneous Visual Comfort and Depth Sensation Optimization for Stereoscopic 3-D ExperienceabstractVisual comfort and depth sensation are two important incongruent counterparts in determining the overall stereoscopic 3-D experience. In this paper, we proposed a novel simultaneous visual comfort and depth sensation optimization approach for stereoscopic images. The main motivation of the proposed optimization approach is to enhance the overall stereoscopic 3-D experience. Toward this end, we propose a two-stage solution to address the optimization problem. In the first layer-independent disparity adjustment process, we iteratively adjust the disparity range of each depth layer to satisfy with visual comfort and depth sensation constraints simultaneously. In the following layer-dependent disparity process, disparity adjustment is implemented based on a defined total energy function built with intra-layer data, inter-layer data and just noticeable depth difference terms. Experimental results on perceptually uncomfortable and comfortable stereoscopic images demonstrate that in comparison with the existing methods, the proposed method can achieve a reasonable performance balance between visual comfort and depth sensation, leading to promising overall stereoscopic 3-D experience. Feng Shao 0001, Weisi Lin, Zhutuan Li, Gangyi Jiang, Qionghai Dai |
IEEE Trans. Cybern. | 4 |
| 2017 | QoE-Guided Warping for Stereoscopic Image RetargetingabstractIn the field of stereoscopic 3D (S3D) display, it is an interesting as well as meaningful issue to retarget the stereoscopic images to the target resolution, while the existing stereoscopic image retargeting methods do not fully take user's Quality of Experience (QoE) into account. In this paper, we have presented a QoE-guided warping method for stereoscopic image retargeting, which retarget the stereoscopic image and adapt its depth range to the target display while promoting user's QoE. Our method takes shape preservation, visual comfort preservation, and depth perception preservation energies into account, and simultaneously optimizes the 2D coordinates and depth information in 3D space. It also considers the specific viewing configuration in the visual comfort and depth perception preservation energy constraints. Experimental results on visually uncomfortable and comfortable stereoscopic images demonstrate that in comparison with the existing stereoscopic image retargeting methods, the proposed method can achieve a reasonable performance optimization among the QoE's factors of image quality, visual comfort, and depth perception, leading to promising overall S3D experience. Feng Shao 0001, Wenchong Lin, Weisi Lin, Qiuping Jiang, Gangyi Jiang |
IEEE Trans. Image Process. | 5 |
| 2017 | Learning Sparse Representation for No-Reference Quality Assessment of Multiply Distorted Stereoscopic ImagesabstractBinocular combination under different distortion types poses a great challenge to three-dimensional image quality assessment (3D-IQA). However, the research works on 3D-IQA with multiple distortion types are very limited. In this paper, we first construct a new multiply distorted stereoscopic image database (NBU-MDSID), which is composed of 270 multiply distorted stereoscopic images and 90 singly distorted stereoscopic images that are corrupted simultaneously and independently by blurring, JPEG compression, and noise injection. We then propose a new multimodal blind metric for quality assessment of multiply distorted stereoscopic images. Inspired by multimodal sparse representation framework, modality-specific dictionaries and the corresponding projection matrices are learned from the singly distorted training database at the training stage, and the testing stage only needs to estimate the quality score based on the reconstruction errors. Experimental results demonstrate the effectiveness of our blind metric. Feng Shao 0001, Weijun Tian, Weisi Lin, Gangyi Jiang, Qionghai Dai |
IEEE Trans. Multim. | 4 |
| 2016 | Allowable depth distortion based depth filtering for 3D high efficiency video codingabstractDepth videos shall be efficiently compressed and transmitted to the client for view synthesis in Three-Dimensional (3D) video system. Since depth video may contain noise that reduce the coding efficiency, we propose a depth filtering algorithm for 3D depth coding, which exploits the Allowable Depth Distortion (ADD) in view synthesis and is able to improve the coding performance of the depth encoder. Firstly, the depth values has the same rendering position based on the ADD model are clustered. Then, the clustered depth are filtered and set to the optimal depth value for each group by minimizing the view synthesis error. The filtered depth videos are smoother and can be more effectively compressed by the existing 3D High Efficiency Video Coding (HEVC) depth encoder. Experimental results show that the proposed depth filtering method can assist the depth encoder achieve 5.87% bit rate reduction in terms of Bjonteggard Delta Bit Rate (BDBR) and 0.25dB quality gain in terms of Bjonteggard Delta Peak-Signal-to-Noise Ratio (BDPSNR) on average as compared with that of coding the original depth maps. Yun Zhang 0002, Linwei Zhu, Xiangkai Liu, Gangyi Jiang |
ISCAS | 4 |
| 2016 | Novel visibility threshold model for asymmetrically distorted stereoscopic imagesabstractExisting perceptual researches on stereoscopic images mainly focus on the threshold of whole image distortion, rather than the effect of texture feature on the so-called threshold of just-noticeable distortion. Obviously, it is unreasonable to use a single unified perception threshold for natural stereoscopic images as the texture complexity typically varies in different blocks of natural images. To solve this problem, we generated an asymmetrically distorted stereoscopic image database with different texture densities and conducted a large number of subjective experiments. A strong correlation between the asymmetrical visibility threshold and texture complexity was revealed from the subjective experiments. Finally, a nonlinear fitting model was designed to uncover this relationship, which can be applied to asymmetrical coding to control the perceived quality of stereoscopic images. Baozhen Du, Mei Yu 0001, Gangyi Jiang, Yun Zhang 0002, Feng Shao 0001, Zongju Peng, Tianzhi Zhu |
VCIP | 3 |
| 2016 | Cluster-based cross-view filtering for compressed multi-view depth mapsabstractIn the field of multi-view video coding, multi-view plus depth video is an important data format, but it always suffers from quantization errors, which result in obvious artifacts in consequent virtual view rendering. In this paper, we propose a cluster-based cross-view filtering (CBF) scheme for the enhancement of compressed depth maps. In this scheme, reconstructed depth information are mapped from cross-view, and this information is benefit to the proposed filter. Then in filtering one viewpoint depth map with candidate information that are selected from non-locally current and neighboring viewpoints. Specifically, in our scheme, candidates are clustered in 3D super-pixel wise rather than block wise due to cross-relationship among pixels in depth maps. The experimental results show that 2.0074 dB average gain can be obtained by our scheme, which suggests that the scheme outperforms than state-of-the-art and classical filters in filtering the reconstructed depth maps. Zhen Liu 0002, Qiong Liu 0001, You Yang 0002, Yuchi Liu, Gangyi Jiang, Mei Yu 0001 |
VCIP | 5 |
| 2016 | A depth video processing algorithm based on cluster dependent and corner-ware filtering
Zongju Peng, Mingsong Guo, Gangyi Jiang, Mei Yu 0001, Feng Shao 0001 |
Neurocomputing | 4 |
| 2016 | Binocular perception based reduced-reference stereo video quality assessment method
Mei Yu 0001, Kaihui Zheng, Gangyi Jiang, Feng Shao 0001, Zongju Peng |
J. Vis. Commun. Image Represent. | 3 |
| 2016 | Machine learning based fast H.264/AVC to HEVC transcoding exploiting block partition similarity
Linwei Zhu, Yun Zhang 0002, Na Li 0015, Gangyi Jiang, Sam Kwong |
J. Vis. Commun. Image Represent. | 4 |
| 2016 | Asymmetric self-recovery oriented stereo image watermarking method for three dimensional video system
Ting Luo 0001, Gangyi Jiang, Mei Yu 0001, Haiyong Xu |
Multim. Syst. | 2 |
| 2016 | A fast inter coding algorithm for HEVC based on texture and motion quad-tree models
Zongju Peng, Gangyi Jiang, Mei Yu 0001, Feng Shao 0001 |
Signal Process. Image Commun. | 4 |
| 2016 | No-reference Stereoscopic Image Quality Assessment Using Binocular Self-similarity and Deep Neural Network
Yaqi Lv, Mei Yu 0001, Gangyi Jiang, Feng Shao 0001, Zongju Peng |
Signal Process. Image Commun. | 3 |
| 2016 | On Predicting Visual Comfort of Stereoscopic Images: A Learning to Rank Based ApproachabstractPredicting the degree of experienced visual comfort in the context of stereoscopic 3-D (S3D) viewing is particularly challenging. In this letter, a simple yet effective visual comfort assessment (VCA) approach for stereoscopic images is proposed from the perspective of learning to rank (L2R). The proposed L2R-based VCA (L2R-VCA) approach is inspired by the traditional absolute categorical rating (ACR) methodology in subjective study and is to characterize the qualitative description behavior of human subjective study. Experimental results on our recently built database confirm the promising performance of the proposed L2R-VCA approach, yielding higher consistency with human subject judgment results. Qiuping Jiang, Feng Shao 0001, Weisi Lin, Gangyi Jiang |
IEEE Signal Process. Lett. | 4 |
| 2016 | Learning Receptive Fields and Quality Lookups for Blind Quality Assessment of Stereoscopic ImagesabstractBlind quality assessment of 3D images encounters more new challenges than its 2D counterparts. In this paper, we propose a blind quality assessment for stereoscopic images by learning the characteristics of receptive fields (RFs) from perspective of dictionary learning, and constructing quality lookups to replace human opinion scores without performance loss. The important feature of the proposed method is that we do not need a large set of samples of distorted stereoscopic images and the corresponding human opinion scores to learn a regression model. To be more specific, in the training phase, we learn local RFs (LRFs) and global RFs (GRFs) from the reference and distorted stereoscopic images, respectively, and construct their corresponding local quality lookups (LQLs) and global quality lookups (GQLs). In the testing phase, blind quality pooling can be easily achieved by searching optimal GRF and LRF indexes from the learnt LQLs and GQLs, and the quality score is obtained by combining the LRF and GRF indexes together. Experimental results on three publicly 3D image quality assessment databases demonstrate that in comparison with the existing methods, the devised algorithm achieves high consistent alignment with subjective assessment. Feng Shao 0001, Weisi Lin, Gangyi Jiang, Mei Yu 0001, Qionghai Dai |
IEEE Trans. Cybern. | 4 |
| 2016 | Toward a Blind Deep Quality Evaluator for Stereoscopic Images Based on Monocular and Binocular InteractionsabstractDuring recent years, blind image quality assessment (BIQA) has been intensively studied with different machine learning tools. Existing BIQA metrics, however, do not design for stereoscopic images. We believe this problem can be resolved by separating 3D images and capturing the essential attributes of images via deep neural network. In this paper, we propose a blind deep quality evaluator (DQE) for stereoscopic images (denoted by 3D-DQE) based on monocular and binocular interactions. The key technical steps in the proposed 3D-DQE are to train two separate 2D deep neural networks (2D-DNNs) from 2D monocular images and cyclopean images to model the process of monocular and binocular quality predictions, and combine the measured 2D monocular and cyclopean quality scores using different weighting schemes. Experimental results on four public 3D image quality assessment databases demonstrate that in comparison with the existing methods, the devised algorithm achieves high consistent alignment with subjective assessment. Feng Shao 0001, Weijun Tian, Weisi Lin, Gangyi Jiang, Qionghai Dai |
IEEE Trans. Image Process. | 4 |
| 2016 | High-Efficiency 3D Depth Coding Based on Perceptual Quality of Synthesized VideoabstractIn 3D video systems, imperfect depth images often induce annoying temporal noise, e.g., flickering, to the synthesized video. However, the quality of synthesized view is usually measured with peak signal-to-noise ratio or mean squared error, which mainly focuses on pixelwise frame-by-frame distortion regardless of the obvious temporal artifacts. In this paper, a novel full reference synthesized video quality metric (SVQM) is proposed to measure the perceptual quality of the synthesized video in 3D video systems. Based on the proposed SVQM, an improved rate-distortion optimization (RDO) algorithm is developed with the target of minimizing the perceptual distortion of synthesized view at given bit rate. Then, the improved RDO algorithm is incorporated into the 3D High Efficiency Video Coding (3D-HEVC) software to improve the 3D depth video coding efficiency. Experimental results show that the proposed SVQM metric has better consistency with human perception on evaluating the synthesized view compared with the state-of-the-art image/video quality assessment algorithms. Meanwhile, this SVQM metric maintains low complexity and easy integration to the current video codec. In addition, the proposed SVQM-based depth coding scheme can achieve approximately 15.27% and 17.63% overall bit rate reduction or 0.42- and 0.46-dB gain in terms of SVQM quality score on average as compared with the latest 3D-HEVC reference model and the state-of-the-art depth coding algorithm, respectively. Yun Zhang 0002, Xiaoxiang Yang, Xiangkai Liu, Yongbing Zhang 0002, Gangyi Jiang, Sam Kwong |
IEEE Trans. Image Process. | 5 |
| 2016 | Learning Blind Quality Evaluator for Stereoscopic Images Using Joint Sparse RepresentationabstractPerceptual quality prediction for stereoscopic images is of fundamental importance in determining the level of quality perceived by humans in terms of the 3D viewing experience. However, the existing no-reference quality assessment (NR-IQA) framework has its limitation in addressing binocular combination for stereoscopic images. In this paper, we propose a new NR-IQA for stereoscopic images using joint sparse representation. We analyze the relationship between left and right quality predictors, and formulate stereoscopic quality prediction as a combination of feature-prior and feature-distribution. Based on this finding, we extract feature vector that handles different features to be interacted by joint sparse representation, and use support vector regression to characterize feature-prior. Meanwhile, we implement feature-distribution using sparsity regularization as the basis of weights for binocular combination to derive the overall quality score. Experimental results on five public 3D IQA databases demonstrate that in comparison with the existing methods, the devised algorithm achieves high consistent alignment with subjective assessment. Feng Shao 0001, Kemeng Li, Weisi Lin, Gangyi Jiang, Qionghai Dai |
IEEE Trans. Multim. | 4 |
| 2015 | Difference of Gaussian statistical features based blind image quality assessment: A deep learning approachabstractNowadays, natural scene statistics (NSS) based blind image quality assessment (BIQA) models trained by machine learning, tend to achieve excellent performance. However, BIQA is still a very challenging research topic due to the lack of reference images. The key of further improvement lies in feature mining and pooling strategy decision. In this work, a new BIQA model is proposed to utilize local normalized multi-scale difference of Gaussian (DoG) response in distorted images as features which show a high correlation with perceptual quality. Then, a three-step-framework based deep neural network (DNN) is designed and employed as the pooling strategy. Compared with the support vector machine (SVM), the proposed three-step-framework DNN can excavate better feature representation, leading to more accurate predictions and stronger generalization ability. The proposed model achieves state-of-the-art performance on two authoritative databases and excellent generalization ability in cross database experiments. Yaqi Lv, Gangyi Jiang, Mei Yu 0001, Haiyong Xu, Feng Shao 0001 |
ICIP | 2 |
| 2015 | Supervised dictionary learning for blind image quality assessmentabstractIn this paper, we propose a supervised dictionary learning framework for blind image quality assessment (BIQA) by using quality-constraint sparse coding. Different with the traditional dictionary learning framework which only ensures the learnt dictionary accounting for image features, we add a quality-related regularization term in the framework to learn a feature-related dictionary and a quality-related dictionary jointly. Specifically, the feature-related and quality-related dictionaries share the same sparse coefficients, so that the reconstruction errors form the image feature vectors and quality score vectors are both minimized. Once the feature-related and quality-related dictionaries are learned, given a testing sample, we first abstract its feature vector and then compute the corresponding sparse coefficients w.r.t. the learnt feature-related dictionary, its quality score can be directly reconstructed based on the learnt quality-related dictionary and the estimated sparse coefficients. Experiment results on three publicly available IQA databases show the promising performance of the proposed model. Qiuping Jiang, Feng Shao 0001, Gangyi Jiang, Mei Yu 0001, Zongju Peng |
VCIP | 3 |
| 2015 | Supervised dictionary learning for blind image quality assessment using quality-constraint sparse coding
Qiuping Jiang, Feng Shao 0001, Gangyi Jiang, Mei Yu 0001, Zongju Peng |
J. Vis. Commun. Image Represent. | 3 |
| 2015 | Depth video spatial and temporal correlation enhancement algorithm based on just noticeable rendering distortion model
Zongju Peng, Gangyi Jiang, Mei Yu 0001, Feng Shao 0001, Yo-Sung Ho |
J. Vis. Commun. Image Represent. | 3 |
| 2015 | Binocular vision based objective quality assessment method for stereoscopic images
Gangyi Jiang, Junming Zhou, Mei Yu 0001, Yun Zhang 0002, Feng Shao 0001, Zongju Peng |
Multim. Tools Appl. | 1 |
| 2015 | Highly efficient contour-based predictive shape coding
Zhongjie Zhu, Yu-er Wang, Gangyi Jiang |
Pattern Recognit. Lett. | 3 |
| 2015 | A depth perception and visual comfort guided computational model for stereoscopic 3D visual saliency
Qiuping Jiang, Feng Shao 0001, Gangyi Jiang, Mei Yu 0001, Zongju Peng, Changhong Yu |
Signal Process. Image Commun. | 3 |
| 2015 | Using Binocular Feature Combination for Blind Quality Assessment of Stereoscopic ImagesabstractThe quality assessment of 3D images is more challenging than its 2D counterparts, and little investigation has been dedicated to blind quality assessment of stereoscopic images. In this letter, we propose a novel blind quality assessment for stereoscopic images based on binocular feature combination. The prominent contribution of this work is that we simplify the process of binocular quality prediction as monocular feature encoding and binocular feature combination. Experimental results on two publicly available 3D image quality assessment databases demonstrate the promising performance of the proposed method. Feng Shao 0001, Kemeng Li, Weisi Lin, Gangyi Jiang, Mei Yu 0001 |
IEEE Signal Process. Lett. | 4 |
| 2015 | Low Complexity HEVC INTRA Coding for High-Quality Mobile Video CommunicationabstractINTRA video coding is essential for high quality mobile video communication and industrial video applications since it enhances video quality, prevents error propagation, and facilitates random access. The latest high-efficiency video coding (HEVC) standard has adopted flexible quad-tree-based block structure and complex angular INTRA prediction to improve the coding efficiency. However, these technologies increase the coding complexity significantly, which consumes large hardware resources, computing time and power cost, and is an obstacle for real-time video applications. To reduce the coding complexity and save power cost, we propose a fast INTRA coding unit (CU) depth decision method based on statistical modeling and correlation analyses. First, we analyze the spatial CU depth correlation with different textures and present effective strategies to predict the most probable depth range based on the spatial correlation among CUs. Since the spatial correlation may fail for image boundary and transitional areas between textural and smooth areas, we then present a statistical model-based CU decision approach in which adaptive early termination thresholds are determined and updated based on the rate-distortion (RD) cost distribution, video content, and quantization parameters (QPs). Experimental results show that the proposed method can reduce the complexity by about 56.76% and 55.61% on average for various sequences and configurations; meanwhile, the RD degradation is negligible. Yun Zhang 0002, Sam Kwong, Zhaoqing Pan, Hui Yuan 0001, Gangyi Jiang |
IEEE Trans. Ind. Informatics | 6 |
| 2015 | Full-Reference Quality Assessment of Stereoscopic Images by Learning Binocular Receptive Field PropertiesabstractQuality assessment of 3D images encounters more challenges than its 2D counterparts. Directly applying 2D image quality metrics is not the solution. In this paper, we propose a new full-reference quality assessment for stereoscopic images by learning binocular receptive field properties to be more in line with human visual perception. To be more specific, in the training phase, we learn a multiscale dictionary from the training database, so that the latent structure of images can be represented as a set of basis vectors. In the quality estimation phase, we compute sparse feature similarity index based on the estimated sparse coefficient vectors by considering their phase difference and amplitude difference, and compute global luminance similarity index by considering luminance changes. The final quality score is obtained by incorporating binocular combination based on sparse energy and sparse complexity. Experimental results on five public 3D image quality assessment databases demonstrate that in comparison with the most related existing methods, the devised algorithm achieves high consistency with subjective assessment. Feng Shao 0001, Kemeng Li, Weisi Lin, Gangyi Jiang, Mei Yu 0001, Qionghai Dai |
IEEE Trans. Image Process. | 4 |
| 2014 | Disparity based stereo image reversible data hidingabstractAs the popularity of three dimensional video, security of stereo image has become an evident issue to be solved. This paper presents a disparity based stereo image reversible data hiding by using histogram shifting, which can recover the original stereo image from marked stereo image without any distortion. Inter-correlations between left and right views of stereo image are utilized to predict pixels accurately. Then prediction error bins are constructed, and many points are around zero-valued bin for embedding data with low distortion of stereo images. The zero-valued bin is used twice to embed data, so that embedding capacity can reach more than 1 bit per pixel. Experimental results demonstrate that the proposed method outperforms the extended stereo image data hiding methods. Ting Luo 0001, Gangyi Jiang, Mei Yu 0001, Feng Shao 0001, Zongju Peng |
ICIP | 2 |
| 2014 | Stereo image watermarking scheme for authentication with self-recovery capability using inter-view reference sharing
Ting Luo 0001, Gangyi Jiang, Mei Yu 0001, Feng Shao 0001, Zongju Peng |
Multim. Tools Appl. | 2 |
| 2014 | Reduced-reference stereoscopic image quality assessment based on view and disparity zero-watermarks
Wujie Zhou, Gangyi Jiang, Mei Yu 0001, Feng Shao 0001, Zongju Peng |
Signal Process. Image Commun. | 2 |
| 2014 | PMFS: A Perceptual Modulated Feature Similarity Metric for Stereoscopic Image Quality AssessmentabstractStereoscopic image quality assessment (SIQA) is an important and challenging issue in three dimensional applications. In this letter, a perceptual modulated feature similarity (PMFS) metric for SIQA is proposed by considering the monocular and binocular perception properties. Specifically, stereoscopic image is first classified into monocular occlusion and binocular rivalry regions. Then, feature similarities between the original and distorted stereoscopic images are defined and measured for the monocular occlusion and binocular rivalry regions as the local monocular and binocular quality maps, respectively. Monocular and binocular just noticeable difference visual saliency models are presented to construct a modulation function to derive monocular and binocular quality scores. Finally, those scores are integrated into an overall quality score by support vector regression. Extensive experiments performed on LIVE phase II and MICT asymmetric databases demonstrate that the proposed PMFS metric can achieve much higher consistency with the subjective quality scores than some state-of-the-art SIQA metrics. Wujie Zhou, Gangyi Jiang, Mei Yu 0001, Feng Shao 0001, Zongju Peng |
IEEE Signal Process. Lett. | 2 |
| 2013 | View-spatial-temporal post-refinement for view synthesis in 3D video systems
Linwei Zhu, Yun Zhang 0002, Mei Yu 0001, Gangyi Jiang, Sam Kwong |
Signal Process. Image Commun. | 4 |
| 2013 | On Perceptually Consistent Image BinarizationabstractCompared with gray images, bi-level images have only two values for each pixel and are more convenient for transmission and storing. Hence, they are very useful in many practical applications such as in image printing and display. Of various image binarization methods, error diffusion is a popular one that can produce perceptually consistent halftone images. However, most of the existing error diffusion techniques have not given rigid theoretical foundation or explicit theoretical derivation, and most of their diffusion domains and diffusion weights are constant, which make them inconvenient to be used and inefficient in some practical applications. In this paper, a new error diffusion scheme for image binarization is derived and established based on the analysis of features of the human visual system (HVS) and the heat transfer theory, where the local image information and the local pixels' value distribution are kept stable or little changed during the binarization process. As a result, the overall visual quality of binarized images can be kept perceptually similar to the original ones. Experiments are conducted and convincing results are acquired. Zhongjie Zhu, Yuer Wang, Gangyi Jiang |
IEEE Signal Process. Lett. | 3 |
| 2013 | Perceptual Full-Reference Quality Assessment of Stereoscopic Images by Considering Binocular Visual CharacteristicsabstractPerceptual quality assessment is a challenging issue in 3D signal processing research. It is important to study 3D signal directly instead of studying simple extension of the 2D metrics directly to the 3D case as in some previous studies. In this paper, we propose a new perceptual full-reference quality assessment metric of stereoscopic images by considering the binocular visual characteristics. The major technical contribution of this paper is that the binocular perception and combination properties are considered in quality assessment. To be more specific, we first perform left-right consistency checks and compare matching error between the corresponding pixels in binocular disparity calculation, and classify the stereoscopic images into non-corresponding, binocular fusion, and binocular suppression regions. Also, local phase and local amplitude maps are extracted from the original and distorted stereoscopic images as features in quality assessment. Then, each region is evaluated independently by considering its binocular perception property, and all evaluation results are integrated into an overall score. Besides, a binocular just noticeable difference model is used to reflect the visual sensitivity for the binocular fusion and suppression regions. Experimental results show that compared with the relevant existing metrics, the proposed metric can achieve higher consistency with subjective assessment of stereoscopic images. Feng Shao 0001, Weisi Lin, Shanbo Gu, Gangyi Jiang, Thambipillai Srikanthan |
IEEE Trans. Image Process. | 4 |
| 2013 | Regional Bit Allocation and Rate Distortion Optimization for Multiview Depth Video Coding With View Synthesis Distortion ModelabstractIn this paper, we propose a view synthesis distortion model (VSDM) that establishes the relationship between depth distortion and view synthesis distortion for the regions with different characteristics: color texture area corresponding depth (CTAD) region and color smooth area corresponding depth (CSAD), respectively. With this VSDM, we propose regional bit allocation (RBA) and rate distortion optimization (RDO) algorithms for multiview depth video coding (MDVC) by allocating more bits on CTAD for rendering quality and fewer bits on CSAD for compression efficiency. Experimental results show that the proposed VSDM based RBA and RDO can improve the coding efficiency significantly for the test sequences. In addition, for the proposed overall MDVC algorithm that integrates VSDM based RBA and RDO, it achieves 9.99% and 14.51% bit rate reduction on average for the high and low bit rate, respectively. It can improve virtual view image quality 0.22 and 0.24 dB on average at the high and low bit rate, respectively, when compared with the original joint multiview video coding model. The RD performance comparisons using five different metrics also validate the effectiveness of the proposed overall algorithm. In addition, the proposed algorithms can be applied to both INTRA and INTER frames. Yun Zhang 0002, Sam Kwong, Long Xu 0001, Sudeng Hu, Gangyi Jiang, C.-C. Jay Kuo |
IEEE Trans. Image Process. | 5 |
| 2013 | Joint Bit Allocation and Rate Control for Coding Multi-View Video Plus Depth Based 3D VideoabstractIn three-dimensional (3D) video coding, distortion in texture video and depth maps can all affect the quality of the synthesized virtual views. Therefore, under the total bitrate constraint, effective bit allocation between texture and depth information is very important for 3D video coding. In this paper, the major technical contribution is to formulate view synthesis quality for optimal resource allocation in 3D video coding, since such quality is what that matters most to the ultimate user (i.e., the viewer) of the system; to be more specific, a new joint bit allocation and rate control method for multi-view video plus depth (MVD) based 3D video coding is proposed accordingly. We firstly derive a view synthesis distortion model to characterize the effect of coding distortion of texture video and depth maps on the synthesized virtual views. Based on this model, we derive a rate-distortion model to characterize the relationship between the bitrate and the view synthesis distortion, and the optimal bitrate ratio between texture and depth is established adaptively by solving the associated optimization problem. Finally, the rate control algorithm is performed on view level, texture/depth level and frame level. Experimental results show that compared with other methods, the proposed bit allocation method obtains higher performance of view synthesis. Moreover, the proposed rate control method can accurately control the bitrate to satisfy the total bitrate constraint. Feng Shao 0001, Gangyi Jiang, Weisi Lin, Mei Yu 0001, Qionghai Dai |
IEEE Trans. Multim. | 2 |
| 2012 | Depth map compression and depth-aided view rendering for a three-dimensional video systemabstractThree-dimensional (3D) video technologies are becoming increasingly popular, as they can provide high quality and immersive experience to end users, where depth maps are employed to generate the virtual views by depth-image-based rendering technique. However, how to reduce the compression and rendering complexities for depth maps while maintaining high rendering quality is still unresolved. In this study, a novel depth map compression and depth-aided view rendering method is proposed. In the proposed method, depth maps are represented with different layers and compressed with different macroblock-mode decision procedure, and several optimisation techniques, including spatio-temporal consistent warping, colour correction and temporal consistent hole filling are embedded into the view rendering framework. Experimental results show that compared with the traditional method, the proposed method can reduce more than 79% compression computational complexity and more than 45% rendering computational complexity, while maintaining high rendering quality. Feng Shao 0001, Mei Yu 0001, Gangyi Jiang, Fucui Li, Zongju Peng |
IET Signal Process. | 3 |
| 2012 | Asymmetric Coding of Multi-View Video Plus Depth Based 3-D Video for View RenderingabstractThe recent years have witnessed three-dimensional (3-D) video technology to become increasingly popular, as it can provide high-quality and immersive experience to end users, where view rendering with depth-image-based rendering (DIBR) technique is employed to generate the virtual views. Distortions in depth map may induce geometry changes in the virtual views, and distortions in texture video may be propagated to the virtual views. Thus, effective compression of both texture videos and depth maps is important for 3-D video system. From the perspective of bit allocation, asymmetric coding of the texture videos and depth maps is an effective way to get the optimal solution of 3-D video compression and view rendering problems. In this paper, a novel asymmetric coding method of multi-view video plus depth (MVD) based 3-D video is proposed on purpose of providing high-quality view rendering. In the proposed method, two models are proposed to characterize view rendering distortion and binocular suppression in 3-D video. Then, an asymmetric coding method of MVD-based 3-D video is proposed by combining two models in encoding framework. Finally, a chrominance reconstruction algorithm is presented to achieve accurate reconstruction. Experimental results show that compared with other methods, the proposed method can obtain higher performance of view rendering under the total bitrate constraint. Moreover, the perceptual visual quality of 3-D video is almost unaffected with the proposed method. Feng Shao 0001, Gangyi Jiang, Mei Yu 0001, Ken Chen 0003, Yo-Sung Ho |
IEEE Trans. Multim. | 2 |
| 2011 | A Novel Rate Control Algorithm for H.264/AVC Based on Human Visual System
Jiangying Zhu, Mei Yu 0001, Qiaoyan Zheng, Zongju Peng, Feng Shao 0001, Fucui Li, Gangyi Jiang |
PSIVT (2) | 7 |
| 2011 | Subjective quality analyses of stereoscopic images in 3DTV systemabstractSubjective quality evaluation is the basis of quality evaluation of stereoscopic images. As the lack of a public and diverse testing database currently, in this paper, a symmetric stereoscopic images database is built. And then the subjective quality of stereoscopic images is analyzed from two aspects, one is the effects of JPEG, JPEG2000, H.264. The other is the comparisons between symmetric and asymmetric stereoscopic images from Gaussian blurring, white Gaussian noise, JPEG and JPEG2000, respectively. The results show three compressions are quite different in the subjective quality of symmetric stereoscopic images at different bitrates, and the comparisons between symmetric and asymmetric stereoscopic images investigate the properties of binocular fusion, binocular suppression, and binocular summation. Junming Zhou, Gangyi Jiang, Xiangying Mao, Mei Yu 0001, Feng Shao 0001, Zongju Peng, Yun Zhang 0002 |
VCIP | 2 |
| 2010 | A Novel Rate Control Method for H.264/AVC Based on Frame Complexity and Importance
Haibing Chen, Mei Yu 0001, Feng Shao 0001, Zongju Peng, Fucui Li, Gangyi Jiang |
ACIVS (2) | 6 |
| 2010 | Asymmetric multi-view video coding based on chrominance reconstructionabstractThree-dimensional video (3DV) technology is becoming increasingly popular, as it can provide high quality and immersive experience to end users. Huge amount of data for storage and transmission is an important problem to be solved. In this paper, an asymmetric MVC method is proposed. Color correction is first performed as a preprocessing step to provide consistent color information among views. Then, all color corrected views are classified into color views and non-color views. The chrominance information in non-color views is all discarded and only preserved in color view in MVC codec. Thus, a large amount of coding bitrate can be saved. At the decoder, a chrominance reconstruction algorithm is presented to achieve accurate color reconstruction for those non-color views. Experimental results show that the proposed method can achieve large bitrate saving against the results compressed with the original JMVM codec. Moreover, the proposed method can obtain better reconstruction quality without noticeable quality degradation. Feng Shao 0001, Gangyi Jiang, Mei Yu 0001, Junyong You |
ICME | 2 |
| 2010 | Stereoscopic Visual Attention Model for 3D Video
Yun Zhang 0002, Gangyi Jiang, Mei Yu 0001, Ken Chen 0003 |
MMM | 2 |
| 2010 | Fast color correction for multi-view video by modeling spatio-temporal variation
Feng Shao 0001, Gangyi Jiang, Mei Yu 0001, Yo-Sung Ho |
J. Vis. Commun. Image Represent. | 2 |
| 2010 | Depth perceptual region-of-interest based multiview video coding
Yun Zhang 0002, Gangyi Jiang, Mei Yu 0001, You Yang 0002, Zongju Peng, Ken Chen 0003 |
J. Vis. Commun. Image Represent. | 2 |
| 2009 | Reduced Reference Image Quality Assessment Based on Contourlet Domain and Natural Image StatisticsabstractReduced-reference (RR) image quality assessment metrics (IQA) evaluate the quality of images by extracting a parameter set from the original reference image and using this set in place of the actual reference image. In this paper, we propose a novel RR-IQA metric based on Contourlet transform. By combining Contourlet transform with a version of the hidden Markov model - Gaussian scale mixtures (GSM), the marginal distributions of neighbor coefficients in the Contourlet domain are modeled. With Contourlet transform as a pre-processing, the marginal histogram of coefficients in each subband can be well fitted by Guassian distribution after divisive normalization transforming. The standard derivation of the fitted Guassian transform and fitted error will be extracted as feature parameters. Experiments show that the proposed metric has good consistency with human subjective perception. Xu Wang 0006, Gangyi Jiang, Mei Yu 0001 |
ICIG | 2 |
| 2009 | A fast multiview video coding algorithm based dynamic multi-thresholdabstractA fast macroblock mode selection algorithm based on dynamic multi-threshold is proposed to improve the encoding speed of multiview video, but with insignificant degradation in rate distortion (RD) performance. The macroblock modes are divided into four classes after statistically analyzing the macroblock mode selection results. Three thresholds are adopted based on the great RD cost gaps between the macroblock mode classes. An approximate computing method and a dynamic updating method of the three thresholds are proposed for implementing the fast algorithm. Simulation results demonstrate that the proposed fast algorithm promotes the encoding speed by 1.92~7.07 times in comparison with JMVM, while the algorithm hardly influences the RD performance. Zongju Peng, Gangyi Jiang, Mei Yu 0001 |
ICME | 2 |
| 2007 | A Content-Adaptive Multi-View Video Color Correction AlgorithmabstractA content-adaptive color correction algorithm for multi-view video is proposed due to variation in lighting or camera parameters. We first establish color correction property between the target image and source image. Then color correction matrix can be obtained by global correction or preferred region matching correction. Finally, video tracking technique is used to correct multi-view video sequences. Experimental results show the proposed algorithm has better correction effect for different multi-view video sequences. Feng Shao 0001, Gangyi Jiang, Mei Yu 0001, Ken Chen 0003 |
ICASSP (1) | 2 |
| 2007 | Wyner-Ziv residual coding for wireless multi-view systemabstractFor wireless multi-view video system, whose abilities of storage and computation are all very weak, it is essential to have an encoder device with low-power consumption and low-complexity. In this paper, a DCT-domain Wyner-Ziv residual coding scheme with low encoding complexity is proposed for wireless multi-view video coding (WZRC-WMS). The scheme is designed to encode the residual frames of each view independently without any motion or disparity estimation at the encoder, so as to shift the large computational complexity to the decoder. At the decoder, the proposed scheme performs joint decoding with side information interpolated from current view and adjacent views. Experimental results show that the proposed WZRC-WMS scheme outperforms the H.263+ interframe coding about 1.9dB in rate-distortion performance, while the encoding complexity is only 1/17 of that of H.264 interframe coding. Zhipeng Jin, Mei Yu 0001, Gangyi Jiang, Ken Chen 0003, Zhidi Jiang |
VCIP | 3 |
| 2006 | New Approach to Wireless Video Compression with Low Complexity
Gangyi Jiang, Zhipeng Jin, Mei Yu 0001, Tae Young Choi |
ACIVS | 1 |
| 2006 | Fast Multi-view Disparity Estimation for Multi-view Video Systems
Gangyi Jiang, Mei Yu 0001, Feng Shao 0001, You Yang 0002 |
ACIVS | 1 |
| 2006 | Fast Adaptive Block Matching for Ray-Space Coding in FTV SystemabstractRay-space representation is the main approach to realizing free viewpoint television (FTV) with complicated scene. Data compression in ray-space is one of the key technologies in ray-space based FTV systems. In ray-space based FTV system, block matching complication is the most important factor to influence coding efficiency. In this paper, a fast adaptive block matching algorithm is proposed by using starting search point prediction, still block determination and search stop criteria strategies. Experimental results show that the search speed is improved greatly as well as the coding efficiency Mei Yu 0001, Feng Shao 0001, Gangyi Jiang |
ICASSP (2) | 3 |
| 2006 | Efficient Block Matching for Ray-Space Predictive Coding in Free-Viewpoint Television Systems
Gangyi Jiang, Feng Shao 0001, Mei Yu 0001, Ken Chen 0003, Tae Young Choi |
ICCSA (1) | 1 |
| 2006 | New Approach to Complexity Reduction of Intra Prediction in Advanced Multimedia Compression
Mei Yu 0001, Gangyi Jiang, Shiping Li, Fucui Li, Tae Young Choi |
ICCSA (1) | 2 |
| 2006 | New Color Correction Approach to Multi-view Images with Region Correspondence
Gangyi Jiang, Feng Shao 0001, Mei Yu 0001, Ken Chen 0003, Xiexiong Chen |
ICIC (1) | 1 |
| 2006 | Parallel Process of Hyper-Space-Based Multiview Video CompressionabstractMultiview video coding (MVC) is a key technology in free-viewpoint television. MVC based on traditional existing codec system has been studied widely, but all of them need powerful computational capacity in processing. Parallel process of MVC can facilitate the efficient implementation of encoder and decoder and has been required as a function by MPEG. In this paper, a parallelization methodology for MVC based on hyper-space theory is presented and tested on the local area multi-computer - message passing interface (LAM-MPI) parallel platform and modified H.264 codec. Experimental results show that the proposed method can speed up processing of multiview video compression and obtain high rate-distortion results. You Yang 0002, Gangyi Jiang, Mei Yu 0001, Dingju Zhu |
ICIP | 2 |
| 2006 | A New Image Correction Method for Multiview Video SystemabstractBecause of scene illumination or camera calibration, color appearance of the same object between different viewpoints may be different in multiview video system. Traditional illumination compensation algorithm for image is unable to solve this problem effectively. In this paper, a novel color correction method for multiview video system is proposed based on retinex color constancy theory. To eliminate influence of un-consistent light sources, histogram equalization, retinex processing and color restoration are performed for multiview images to extract reflectance that describes object intrinsic properties. Experimental results show that the proposed image correction method for multiview video system is effective Feng Shao 0001, Gangyi Jiang, Mei Yu 0001, Xiexiong Chen |
ICME | 2 |
| 2005 | New method of ray-space interpolation for free viewpoint videoabstractRay-space representation has superiority in rendering arbitrary viewpoint images of complicated scene in real-time. Ray-space interpolation is one of the key techniques to make ray-space based free viewpoint video (FVV) feasible. This paper presents a directionality based interpolation method for ray-space based FW system. Characteristic pixels near edges or within areas with fine texture are first extracted from sparse ray-space slice, and their directionalities are determined by multi-stage matching technique, so as to speed up the directionality searching and obtain more reliable directionalities. Pixels to be interpolated in dense ray-space slice are linear interpolated according to the directionalities of these characteristic pixels. Experimental results show that the proposed method improves visual quality as well as PSNRs of rendered intermediate viewpoint image greatly. Gangyi Jiang, Mei Yu 0001, Xien Ye, Liangzhong Fan, Randi Fu |
ICIP (2) | 1 |
| 2005 | New Ray-Space Interpolation Method For Free Viewpoint Video SystemabstractRay-space representation is the main technology to realize Free Viewpoint Video (FVV) system with complicated scene. Ray-space data consists of various lines with different direction. Ray-space interpolation and compression are two key techniques to be solved. In this paper, correlations between multiple epipolar lines in the Ray-space data is analyzed, and a new algorithm of Ray- Space interpolation with multiepipolar lines matching is proposed. Experimental results show that the proposed scheme achieves higher PSNR than the pixel-based matching interpolation method and the block-based matching interpolation method in interpolating the ray-space data and rendering arbitrary viewpoint image. Liangzhong Fan, Mei Yu 0001, Gangyi Jiang, Rangding Wang, Yong-Deak Kim |
PDCAT | 3 |
| 2005 | New Multiple Description Layered Coding Method For Video CommunicationabstractThere are two problems in video communication, one is related to heterogeneity of networks, and the other involves reliability of transmission. Layered coding is designed to solve client heterogeneity problems, and multiple description coding is an effective method for robust transmission. Multiple description layered coding (MDLC) of video combines advantages of layered coding and multiple description coding. In this paper, a new MDLC scheme of video sequence is proposed based on macroblock splitting technique. In addition, other three schemes of MDLC are also given based on row-, column-, and framedecomposition. Experimental results show that the proposed MDLC scheme with macroblock splitting (MDLC-MS) has advantages in adaptability of network heterogeneity and transmission reliability. Mei Yu 0001, Xien Ye, Rangding Wang, Fangming Xiao, Gangyi Jiang |
PDCAT | 5 |
| 2004 | Approaches to H.264-based stereoscopic video codingabstractH.264 is an advanced video compression standard, absorbing the advantages of the previous standards. In this paper, block-based stereoscopic video coding is studied, and some schemes of using H.264 are discussed. The stereoscopic video coding methods based on H.264 and based on H.263+ are also compared by the experiments, experimental results shove that the former is more effective than the latter in compression efficiency and image quality, and the H.264-based stereoscopic video coding scheme with temporal scalability is quite effective. Shiping Li, Mei Yu 0001, Gangyi Jiang, Tae Young Choi, Yong-Deak Kim |
ICIG | 3 |
| 2000 | Lane and obstacle detection based on fast inverse perspective mapping algorithmabstractA fast inverse perspective mapping algorithm (FIPMA) is presented for a fast and accurate recovery of road surface from a given 2D road image. FIPMA is able to simplify the system design without losing reliability and flexibility. A novel lane and obstacle detection method, using only a single CCD camera, is proposed based on a recovered road surface image by FIPMA. This method includes five parts: recovering the road surface from the input road image using FIPMA; processing the recovered surface image; detecting lane and estimating lane parameters; updating camera parameters adaptively; and detecting obstacles. Gangyi Jiang, Tae Young Choi, Suk Kyo Hong, Jae Wook Bae, Byung Suk Song |
SMC | 1 |