Mei Yu 0001

dblp:34/946-1 · DBLP profile ↗
← Back
114ranked-venue papers
5as first author
33since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 92 · 2 first-author · 26 since 2021Artificial intelligence and machine learning · 10 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Systems, architecture and hardware · 1Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 Frequency domain-based latent diffusion model for underwater image enhancement
Jingyu Song, Haiyong Xu, Gangyi Jiang, Mei Yu 0001, Yeyao Chen, Ting Luo 0001, Yang Song 0015
Pattern Recognit.4
2025 Local and Global Structure-Guided No-Reference Point Cloud Quality Assessment
abstract
As a crucial representation of 3D data, a point cloud (PC) can accurately capture the geometry, structure, and color information of objects. However, various quality problems arise owing to device noise, data acquisition errors, and compression algorithms, limiting the application of PCs. Therefore, assessing PC quality to determine its suitability for applications is a challenging task. In this work, a local and global structure-guided feature extraction and attention network (LGS-Net) is introduced for no-reference PC quality assessment (PCQA). This approach incorporates cluster construction (CC), local structure-guided cluster feature extraction (LSFE), and global structure-guided attention (GSA) modules. First, owing to the heightened sensitivity of the human visual system (HVS) to structural information, a graph filter is employed to identify high-frequency clusters. Within the LSFE module, a multiscale strategy is employed to ensure that structural information effectively influences both the geometry and color information. Simultaneously, the multiscale features within the cluster are dynamically fine-tuned using feature channel weight reassignment. To account for the impact of interclusters on overall quality, a GSA module is introduced to establish global dependencies between local clusters. This approach enables the extraction of final geometry, color, and structure information, which are ultimately used for accurate quality assessment. Extensive experimental results show that the proposed method outperforms the existing state-of-the-art PCQA methods using two publicly available subjective datasets.
Zhouyan He, Qihao Liang, Gangyi Jiang, Mei Yu 0001, Yeyao Chen, Ting Luo 0001, Wujie Zhou
IEEE Trans. Multim.4
2025 Multi-Attention Learning and Exposure Guidance Toward Ghost-Free High Dynamic Range Light Field Imaging
abstract
Due to sensor limitations, the light field (LF) images captured by the LF camera suffer from low dynamic range and are prone to poor exposure. To solve this problem, combining multi-exposure technology with LF camera imaging can achieve high dynamic range (HDR) LF imaging. However, for dynamic scenes, this approach tends to produce disturbing ghosting artifacts and destroy the parallax structure of the generated results. To this end, this paper proposes a novel ghost-free HDR LF imaging method using multi-attention learning and exposure guidance. Specifically, the proposed method first designs a multi-scale cross-attention module to achieve efficient multi-exposure LF feature alignment. After that, a dual self-attention-driven Transformer block is constructed to excavate the geometric information of LF and fuse the aligned LF features. In particular, exposure masks derived from middle-exposure are introduced in the feature fusion to guide the network to focus on information recovery in low- and high-brightness regions. Besides, a local compensation module is integrated to cope with local alignment errors and refine details. Finally, a multi-objective reconstruction strategy combined with exposure masks is employed to restore high-quality HDR LF images. Extensive experimental results on the benchmark dataset show that the proposed method generates HDR LF results with high spatial-angular quality consistency and outperforms the state-of-the-art methods in quantitative and qualitative comparisons. Furthermore, the proposed method can enhance the performance of existing LF applications, such as depth estimation.
Yeyao Chen, Gangyi Jiang, Chongchong Jin, Ting Luo 0001, Haiyong Xu, Mei Yu 0001
IEEE Trans. Vis. Comput. Graph.6
2024 Hybrid Domain Learning towards Light Field Spatial Super-Resolution using Heterogeneous Imaging
abstract
Light field (LF) cameras usually capture dense angular samples, but suffer from low spatial resolution. Existing single-LF super-resolution methods struggle with textures at larger scales (e.g., 8×). To address this issue, this paper proposes a novel hybrid domain learning-based method to enhance LF spatial resolution from heterogeneous imaging (integrating an LF camera and a 2D digital camera). The proposed method consists of two core modules, namely LF feature alignment module and cross-domain multi-scale fusion module. The former combines optical flow and deformable convolution to gradually align the 2D high-resolution features with the low-resolution LF features. The latter progressively fuses the aligned multi-resolution LF features to enable high-quality reconstruction. Experimental results show the proposed method recovers fine textures and preserves accurate angular consistency, and outperforms the state-of-the-art methods in both quantitative and qualitative comparisons.
Zean Chen, Yeyao Chen, Mei Yu 0001, Haiyong Xu, Gangyi Jiang
ICASSP3
2024 Multi-exposure fused light field image quality assessment for dynamic scenes: Benchmark dataset and objective metric
Yun Liu 0048, Guanglong Liao, Gangyi Jiang, Yeyao Chen, Yueli Cui, Haiyong Xu, Mei Yu 0001
Expert Syst. Appl.7
2024 Vision graph convolutional network for underwater image enhancement
Zexuan Xing, Haiyong Xu, Gangyi Jiang, Mei Yu 0001, Ting Luo 0001, Yeyao Chen
Knowl. Based Syst.4
2024 HDR light field imaging of dynamic scenes: A learning-based method and a benchmark dataset
Yeyao Chen, Gangyi Jiang, Mei Yu 0001, Chongchong Jin, Haiyong Xu, Yo-Sung Ho
Pattern Recognit.3
2024 Underwater Monocular Depth Estimation Based on Physical-Guided Transformer
abstract
Owing to the light absorption and wavelength scattering in underwater environments, underwater images are severely degraded, which directly affects the depth estimation of underwater scenes. Accurate underwater depth estimation is essential for representing and understanding underwater scenes. However, the existing underwater depth estimation methods have not fully taken into account the distinctive physical properties of underwater environments, which has resulted in increased bias and feature distortion in the depth estimation results. In this paper, an underwater monocular depth estimation method based on physical-guided Transformer (UPGformer) is proposed, considering the characteristics of underwater imaging, including shallow feature extraction, encoding, decoding, and regression stages. Specifically, in the shallow feature extraction stage, considering the color deviation of underwater images and extracting richer primary features, an enrichment and extraction depth Transformer (EEDT) module is proposed, by interacting physically inverted transmission maps of the underwater dark channel prior (UDCP) with physical color-compensated underwater images through self-attention. In the encoding stage, considering the nonuniform degradation of underwater images (nonuniform local distortion and inconsistent channel degradation), the underwater physical Transformer interaction encoder (UPTE) module, which fuses the Transformer and physically inverted transmission maps, is proposed. Furthermore, in the decoding stage, to better recover features and reduce information loss, the underwater physical embedded decoding (UPED) module is proposed, which embeds the physically inverted transmission maps with the upsampling process. Finally, the depth map is constructed during the regression stage. The experimental results demonstrate that the proposed UPGformer outperforms existing methods, both qualitatively and quantitatively.
Chen Wang 0141, Haiyong Xu, Gangyi Jiang, Mei Yu 0001, Ting Luo 0001, Yeyao Chen
IEEE Trans. Geosci. Remote. Sens.4
2024 Stitched Wide Field of View Light Field Image Quality Assessment: Benchmark Database and Objective Metric
abstract
Due to the limitation of commercial light field camera hardware devices, the imaging field of view is quite narrow. Numerous Light Field Image (LFI) stitching algorithms have been developed to expand the field of view. However, it is highly challenging to compare the performance of LFI stitching algorithms in a fair manner. Currently, due to the absence of a comprehensive benchmark database for subjective rating and a reliable objective quality metric, it is fairly difficult to comprehensively and accurately compare the actual performance of existing LFI stitching algorithms. In this study, we dedicate our efforts to the development of quality metrics for stitched Wide field of view LFI (WLFI) from subjective and objective assessment aspects. Specifically, we build the first stitched WLFI database, which provides the stitched WLFIs generated by eight representative LFI stitching algorithms, along with their corresponding subjective rating scores. Secondly, an effective blind stitched WLFI quality metric is developed to accurately assess the visual quality degradation. Extensive experiments conducted over our established WLFI database demonstrate that the proposed metric achieves higher consistency with subjective ratings than the competing quality metrics.
Yueli Cui, Gangyi Jiang, Mei Yu 0001, Yeyao Chen, Yo-Sung Ho
IEEE Trans. Multim.3
2024 Underwater Image Quality Assessment from Synthetic to Real-world: Dataset and Objective Method
abstract
The complicated underwater environment and lighting conditions lead to severe influence on the quality of underwater imaging, which tends to impair underwater exploration and research. To effectively evaluate the quality of underwater images, an underwater image quality assessment dataset is constructed from synthetic to real-world, and then a new objective underwater image assessment method based on the characteristics of the underwater imaging is proposed (UICQA). Specifically, to address the lack of a publicly available datasets and more accurately quantify the quality of underwater images, a subjective underwater image quality assessment dataset from synthetic to real-world underwater images, named USRD, is constructed. Considering that the transmission map can effectively reflect the characteristics of the underwater imaging, statistical features are effectively extracted from the transmission map for distinguishing underwater images of different quality. Further, considering that the transmission map negatively correlates with scene depth, a local-to-global transmission map weighted contrast feature is constructed. Additionally, the color features of human perception and texture features based on fractal dimensions are proposed. Finally, the experimental results show that the proposed UICQA method exhibits the highest correlation with ground truth scores compared to state-of-the-art UIQA methods.
Haiyong Xu, Gangyi Jiang, Mei Yu 0001, Ting Luo 0001, Xuebo Zhang 0002, Hongwei Ying
ACM Trans. Multim. Comput. Commun. Appl.4
2023 Quality Assessment for High Dynamic Range Stereoscopic Omnidirectional Image System
Liuyan Cao, Hao Jiang 0014, Zhidi Jiang, Jihao You, Mei Yu 0001, Gangyi Jiang
ACIVS5
2023 Quality Evaluation of Tone Mapping HDR Omnidirectional Image with Multi-regions and Multi-levels
abstract
Unlike ordinary omnidirectional imaging, which produces underexposure or overexposure of some areas under wide range of lighting conditions, tone mapping high dynamic range omnidirectional image (denoted as TM-HOI) presents more detailed information with high dynamic range (HDR) techniques. Aiming at the multiple distortions in TM-HOI processing, this paper proposes a blind TM-HOI quality metric with multi-regions and multi-levels analysis. Considering the conversion of image projection formats and the unique distortion caused by user behavior in immersive environments, feature extraction in the proposed metric is divided into the viewport based module and the Crasher parabolic projection based module. At the same time, considering the coding distortion, tone mapping (TM) distortion and mixed distortion in the TM-HOI and the different manifestations of this mixed distortion in different regions, this paper further extracts perceptual features to characterize image distortion by performing bit plane layer decomposition and detail/basic layer decomposition on different regions of TM-HOI. Finally, the extracted perceptual features are used as the input of Random forest to construct the nonlinear relationship between the feature space and subjective opinion scores. The experimental results indicate that the proposed metric has better consistency with the human visual system compared to representative blind quality metrics.
Xuelei Zheng, Mei Yu 0001, Gangyi Jiang
AICCSA2
2023 Perceptual Light Field Image Coding with CTU Level Bit Allocation
Panqi Jin, Gangyi Jiang, Yeyao Chen, Zhidi Jiang, Mei Yu 0001
CAIP (2)5
2023 Blind light field image quality assessment with tensor color domain and 3D shearlet transform
Jianjun Xiang, Mei Yu 0001, Gangyi Jiang, Haiyong Xu
Signal Process.2
2023 Domain Adaptation for Underwater Image Enhancement
abstract
Recently, learning-based algorithms have shown impressive performance in underwater image enhancement. Most of them resort to training on synthetic data and obtain outstanding performance. However, these deep methods ignore the significant domain gap between the synthetic and real data (i.e., inter-domain gap), and thus the models trained on synthetic data often fail to generalize well to real-world underwater scenarios. Moreover, the complex and changeable underwater environment also causes a great distribution gap among the real data itself (i.e., intra-domain gap). However, almost no research focuses on this problem and thus their techniques often produce visually unpleasing artifacts and color distortions on various real images. Motivated by these observations, we propose a novel Two-phase Underwater Domain Adaptation network (TUDA) to simultaneously minimize the inter-domain and intra-domain gap. Concretely, in the first phase, a new triple-alignment network is designed, including a translation part for enhancing realism of input images, followed by a task-oriented enhancement part. With performing image-level, feature-level and output-level adaptation in these two parts through jointly adversarial learning, the network can better build invariance across domains and thus bridging the inter-domain gap. In the second phase, an easy-hard classification of real data according to the assessed quality of enhanced images is performed, in which a new rank-based underwater quality assessment method is embedded. By leveraging implicit quality information learned from rankings, this method can more accurately assess the perceptual quality of enhanced images. Using pseudo labels from the easy part, an easy-hard adaptation technique is then conducted to effectively decrease the intra-domain gap between easy and hard samples. Extensive experimental results demonstrate that the proposed TUDA is significantly superior to existing works in terms of both visual quality and quantitative metrics.
Zhengyong Wang, Liquan Shen, Mai Xu, Mei Yu 0001, Kun Wang 0048
IEEE Trans. Image Process.4
2023 No-Reference Light Field Image Quality Assessment Using Four-Dimensional Sparse Transform
abstract
Light field imaging can simultaneously capture the intensity and direction information of light rays in the real world. Light field image (LFI) with four-dimensional (4D) data suffers from quality degradation in the process of compression, reconstruction and processing. How to evaluate the visual quality of LFI is thought-provoking. This paper proposes a no-reference LFI quality assessment metric based on high-dimensional sparse transform. Firstly, LFI's sub-aperture gradient image array (SAGIA), which is still a 4D signal, is generated by high-pass filtering between adjacent SAIs. Then, SAGIA is transformed with 4D discrete cosine transform (4D-DCT). 4D-DCT coefficients of SAGIA can characterize the angular and spatial information of LFI. And the logarithmic amplitudes of the coefficients at the same position of SAGIA?s transformed 4D blocks are averaged as the coefficient energy. Subsequently, the 4D-DCT coefficients of SAGIA are divided into the spatial-angular frequency bands and spatial-angular orientation bands, and the corresponding energy features are extracted by converging the coefficient energy of the same band. In addition, the coefficients' amplitudes at the same position of blocks are fitted by the Weibull distribution. Then, the fitted parameters of each position are concatenated, and cropped with principal component analysis to obtain the compact features. Finally, the extracted features are pooled to predict the visual quality of the distorted LFIs. The experimental results demonstrate that the proposed method is more consistent with the subjective evaluation on three LFI databases, compared with the state-of-the-art image quality assessment methods and LFI quality assessment methods.
Jianjun Xiang, Gangyi Jiang, Mei Yu 0001, Zhidi Jiang, Yo-Sung Ho
IEEE Trans. Multim.3
2023 Deep Light Field Spatial Super-Resolution Using Heterogeneous Imaging
abstract
Light field (LF) imaging expands traditional imaging techniques by simultaneously capturing the intensity and direction information of light rays, and promotes many visual applications. However, owing to the inherent trade-off between the spatial and angular dimensions, LF images acquired by LF cameras usually suffer from low spatial resolution. Many current approaches increase the spatial resolution by exploring the four-dimensional (4D) structure of the LF images, but they have difficulties in recovering fine textures at a large upscaling factor. To address this challenge, this paper proposes a new deep learning-based LF spatial super-resolution method using heterogeneous imaging (LFSSR-HI). The designed heterogeneous imaging system uses an extra high-resolution (HR) traditional camera to capture the abundant spatial information in addition to the LF camera imaging, where the auxiliary information from the HR camera is utilized to super-resolve the LF image. Specifically, an LF feature alignment module is constructed to learn the correspondence between the 4D LF image and the 2D HR image to realize information alignment. Subsequently, a multi-level spatial-angular feature enhancement module is designed to gradually embed the aligned HR information into the rough LF features. Finally, the enhanced LF features are reconstructed into a super-resolved LF image using a simple feature decoder. To improve the flexibility of the proposed method, a pyramid reconstruction strategy is leveraged to generate multi-scale super-resolution results in one forward inference. The experimental results show that the proposed LFSSR-HI method achieves significant advantages over the state-of-the-art methods in both qualitative and quantitative comparisons. Furthermore, the proposed method preserves more accurate angular consistency.
Yeyao Chen, Gangyi Jiang, Mei Yu 0001, Haiyong Xu, Yo-Sung Ho
IEEE Trans. Vis. Comput. Graph.3
2022 CPC-GSCT: Visual quality assessment for coloured point cloud based on geometric segmentation and colour transformation
abstract
Abstract Coloured point cloud (CPC) is one of the important representations of three‐dimensional objects, which has been used in many fields. CPC may encounter geometric and colour distortion during its compression, simplification or other processing. Thus, the objective visual quality assessment of CPC is one of the urgent issues to be resolved in the CPC's applications. Aiming at this problem, this paper proposes a new full‐reference visual quality assessment metric for CPC based on geometric segmentation and colour transformation (CPC‐GSCT), which analyzes geometric distortion and colour distortion of CPC. First, considering the visual masking effect of CPC's geometric information, CPC is segmented into different regions and distributed with different weights to describe the influence of visual masking effect in CPC quality assessment. At the same time, a geometric combination feature vector is defined and extracted for measuring the CPC's geometric distortion. Then, considering the colour perception of human eyes, a colour combination feature vector is extracted to measure the CPC's colour distortion in HSV colour space. Finally, all the extracted geometric and colour features are constituted as a feature vector to predict the quality of CPC. Experimental results on three databases (IRPC, SJTU‐PCQA and CPCD2.0) show that the proposed CPC‐GSCT metric can achieve better performance in predicting the visual quality of CPC than relevant existing methods.
Mei Yu 0001, Zhouyan He, Renwei Tu, Gangyi Jiang
IET Image Process.2
2022 TGP-PCQA: Texture and geometry projection based quality assessment for colored point clouds
Zhouyan He, Gangyi Jiang, Mei Yu 0001, Zhidi Jiang, Zongju Peng
J. Vis. Commun. Image Represent.3
2022 Multi-Angle Projection Based Blind Omnidirectional Image Quality Assessment
abstract
Most of the existing blind omnidirectional image quality assessment (BOIQA) methods are based on data-driven approach where the end-to-end neural network or deep learning tools are mainly used for feature extraction. However, it usually lacks interpretability and is difficult to discover the perceptual mechanism behind. In this paper, from the perspective of perception modeling, we propose a novel multi-angle projection based BOIQA (MP-BOIQA) method. Considering the omnibearing and near eye display characteristics with head mounted display, multiple color cubemap projection images with respect to different viewpoints are grouped as the color omnidirectional distortion (COD) units so as to simulate the user’s viewing behavior in subjective quality assessment. In the designed multi-angle projection based feature extractor, tensor decomposition is implemented on each COD unit for dimensionality reduction, and piecewise exponential fitting is used to get the distribution of mean subtracted contrast normalized coefficients of the unit’s feature matrices in tensor domain. Finally, the extracted features are pooled with random forest. The experimental results on three omnidirectional image quality datasets show that the MP-BOIQA method can deliver highly competitive performance compared with some representative full-reference quality assessment methods, as well as some state-of-the-art BOIQA methods.
Hao Jiang 0014, Gangyi Jiang, Mei Yu 0001, Ting Luo 0001, Haiyong Xu
IEEE Trans. Circuits Syst. Video Technol.3
2022 Reinforced Swin-Convs Transformer for Simultaneous Underwater Sensing Scene Image Enhancement and Super-resolution
abstract
Underwater image enhancement (UIE) technology aims to tackle the challenge of restoring the degraded underwater images due to light absorption and scattering. Meanwhile, the ever-increasing requirement for higher resolution images from a lower resolution in the underwater domain cannot be overlooked. To address these problems, a novel U-Net-based reinforced Swin-Convs Transformer for simultaneous enhancement and superresolution (URSCT-SESR) method is proposed. Specifically, with the deficiency of U-Net based on pure convolutions, the Swin Transformer is embedded into U-Net for improving the ability to capture the global dependence. Then, given the inadequacy of the Swin Transformer capturing the local attention, the reintroduction of convolutions may capture more local attention. Thus, an ingenious manner is presented for the fusion of convolutions and the core attention mechanism to build a reinforced Swin-Convs Transformer block (RSCTB) for capturing more local attention, which is reinforced in the channel and the spatial attention of the Swin Transformer. Finally, experimental results on available datasets demonstrate that the proposed URSCT-SESR achieves the state-of-the-art performance compared with other methods in terms of both subjective and objective evaluations. The code is publicly available athttps://github.com/TingdiRen/URSCT-SESR.
Tingdi Ren, Haiyong Xu, Gangyi Jiang, Mei Yu 0001, Ting Luo 0001
IEEE Trans. Geosci. Remote. Sens.4
2022 Deep Light Field Super-Resolution Using Frequency Domain Analysis and Semantic Prior
abstract
Light field (LF) camera can simultaneously capture the intensity and direction information of light rays, which has been widely concerned. However, limited by the size of the imaging sensor, the captured LF image (LFI) has a trade-off between spatial and angular resolutions. To this end, this paper proposes a new LF super-resolution method using frequency domain analysis and semantic prior, which designs a two-stage learning framework to enhance the spatial and angular resolutions of LFI. Specifically, the proposed method first decomposes the spatial and angular information to explore the 4D structure of LFI by using frequency domain transformation, and formulates the LF super-resolution as a frequency restoration process. Then, the decomposed frequency components are recovered in a progressive restoration manner, with new cascaded 2D and 3D convolutional neural networks. To further improve the quality of the reconstructed LFI, especially at the object boundary, the semantic prior is incorporated into the designed network to enhance its representation ability. Finally, the super-resolved LFI is reconstructed by inverse frequency domain transformation. Experimental results show that the proposed method can effectively generate high-resolution LFI, and outperforms other state-of-the-art methods in terms of both subjective visual perception and objective quality evaluation. Moreover, the proposed method can enhance the performance of LF applications such as depth estimation.
Yeyao Chen, Gangyi Jiang, Zhidi Jiang, Mei Yu 0001, Yo-Sung Ho
IEEE Trans. Multim.4
2022 Fast Intra Mode Decision Algorithm for Versatile Video Coding
abstract
To achieve higher coding efficiency, the latest Versatile Video Coding (VVC) standard adopts a series of new intra coding techniques, including the quadtree plus multi-type tree (QTMT), intra sub-partitions (ISP) and intra block copy (IBC). However, this makes the intra coding more complicated, as VVC needs to traverse all prediction modes and partition types of QTMT to find the optimal combination. In this paper, we propose a fast algorithm for VVC from two aspects of mode selection and prediction terminating to reduce coding complexity. For the mode selection, adaptive mode pruning (AMP) is proposed to remove non-promising modes. First, since the newly introduced modes (IBC and ISP) are not effective for all blocks, learning-based classifiers are designed to remove them intelligently. Second, for normal modes, an ensemble decision strategy is proposed to sort the candidate modes and increase the probability of being the optimal mode for the first few candidates; thus, we can remove redundant candidates more efficiently. In terms of prediction terminating, we find that different optimal modes of current depth level lead to different termination probabilities of remaining intra predictions. Therefore, mode-dependent termination (MDT) is proposed to select an appropriate model through the optimal mode and terminate unnecessary intra predictions of remaining depth levels. The proposed algorithm is implemented on VVC test model, and simulation results show that it can achieve 51%$\sim$53% time savings with only 0.93%$\sim$1.08% BDBR increases.
Xinchao Dong, Liquan Shen, Mei Yu 0001, Hao Yang 0008
IEEE Trans. Multim.3
2022 Tensor Product and Tensor-Singular Value Decomposition Based Multi-Exposure Fusion of Images
abstract
Considering multidimensional structure of the multi-exposure images, a new Tensor product and Tensor-singular value decomposition based Multi-Exposure image Fusion (TT-MEF) method is proposed. The main innovation of this work is to explore a new feature representation of multi-exposure images in the new tensor domain and design the fusion strategy on this basis. Specifically, the luminance and the chrominance channels are fused separately to maintain color consistency. For the luminance fusion, the luminance channel of multi-exposure images is divided into two parts, that is, de-mean term and mean term. The de-mean term is represented as a tensor to extract the feature. Then, the tensor product and tensor-singular value decomposition (T-SVD) are used to design a tensor feature extractor. Furthermore, a fusion strategy of the de-mean term is presented according to the visual saliency model, and a fusion strategy of the mean term is defined by the local and the global visual weights to control counterpoise between the local and global luminance. For the chrominance fusion, a new fusion strategy is also designed by the tensor product and T-SVD, similar to the luminance fusion. Finally, the fused image is obtained by combining the luminance and chrominance fusion. Experimental results show that the proposed TT-MEF method generally outperforms the existing state-of-the-art in terms of subjective visual quality and objective evaluation.
Haiyong Xu, Gangyi Jiang, Mei Yu 0001, Zhongjie Zhu, Yongqiang Bai, Yang Song 0015, Huifang Sun
IEEE Trans. Multim.3
2021 Towards A Colored Point Cloud Quality Assessment Method Using Colored Texture And Curvature Projection
abstract
Colored point cloud (PC) provides convenience for 3D digitization in the real world, but its huge amount of data needs to be compressed effectively. However, lossy compression will bring visual quality problems, so it is necessary to design reliable quality assessment methods. Considering the visual connection between 3D space and projection plane, we propose a new PC quality assessment (PCQA) method combining colored texture and curvature projection in this paper. Specifically, the colored texture information and curvature of colored PC are projected onto 2D planes to extract texture and geometric statistical features, respectively, so as to characterize the texture and geometric distortion. Experimental results on two colored PC databases (CPCD2.0 and IRPC) show that the proposed method has a good correlation with subjective quality scores and is superior to the state-of-the-art PCQA methods.
Zhouyan He, Gangyi Jiang, Zhidi Jiang, Mei Yu 0001
ICIP4
2021 Point Cloud Projection and Multi-Scale Feature Fusion Network Based Blind Quality Assessment for Colored Point Clouds
abstract
With the wide applications of colored point cloud (CPC) in many fields, many attentions have been paid to CPC's distortions caused by its compression and reconstruction. How to effectively evaluate the visual quality of CPC has become an urgent issue to be resolved. In this paper, a Point cloud projection and Multi-scale feature fusion network based Blind Visual Quality Assessment method (denoted as PM-BVQA) is proposed for CPC. CPC in 3D space is first projected into 2D color projection map and geometric projection map, then a multi-scale feature fusion network is designed to blindly evaluate the visual quality of CPC. The proposed PM-BVQA method includes three modules, that is, joint color-geometric feature extractor, two-stage multi-scale feature fusion, and spatial pooling module. Considering the multi-channel characteristics of human visual system (HVS), unimodal features of different scales are obtained by joint color-geometric feature extractor from the color and geometric projection maps. The fusion of the unimodal color and geometric features is carried out to capture the cross-modal complementary information between these two types of information. By integrating cross-modal fused features at different scales, the complementary relationships between different channels of HVS are simulated. The spatial pooling module takes into account the attention mechanism of HVS and realizes the weighted summation of local regional quality to obtain the final global quality score of CPC. A subjective CPC database with coding distortion is used to verify the effectiveness of the proposed method, and the experimental results show that the proposed blind quality assessment method is more consistent with the subjective visual perception than the existing quality assessment methods.
Wenxu Tao, Gangyi Jiang, Zhidi Jiang, Mei Yu 0001
ACM Multimedia4
2021 FQM-GC: Full-reference Quality Metric for Colored Point Cloud Based on Graph Signal Features and Color Features
abstract
Colored Point Cloud (CPC) is often distorted in the processes of its acquisition, processing, and compression, so reliable quality assessment metrics are required to estimate the perception of distortion of CPC. We propose a Full-reference Quality Metric for colored point cloud based on Graph signal features and Color features (FQM-GC). For geometric distortion, the normal and coordinate information of the sub-clouds divided via geometric segmentation is used to construct their underlying graphs, then, the geometric structure features are extracted. For color distortion, the corresponding color statistical features are extracted from regions divided with color attribution. Meanwhile, the color features of different regions are weighted to simulate the visual masking effect. Finally, all the extracted features are formed into a feature vector to estimate the quality of CPCs. Experimental results on three databases (CPCD2.0, IRPC and SJTU-PCQA) show that the proposed metric FQM-GC is more consistent with human visual perception.
Ke-Xin Zhang, Gangyi Jiang, Mei Yu 0001
MMAsia3
2021 Strong ghost removal in multi-exposure image fusion using hole-filling with exposure congruency
Mei Yu 0001, Gangyi Jiang, Zhiyong Pan, Zongju Peng
J. Vis. Commun. Image Represent.2
2021 No-reference light field image quality assessment based on depth, structural and angular information
Jianjun Xiang, Gangyi Jiang, Mei Yu 0001, Yongqiang Bai, Zhongjie Zhu
Signal Process.3
2021 Inter-layer correlation-based adaptive bit allocation for enhancement layer in scalable high efficiency video coding
Zongju Peng, Dongrong Jiang, Chao Huang 0008, Gangyi Jiang, Mei Yu 0001
Signal Process. Image Commun.6
2021 Viewport Perception Based Blind Stereoscopic Omnidirectional Image Quality Assessment
abstract
Compared with traditional 2D images, stereoscopic omnidirectional images (SOIs) usually have more complex perceptual factors due to the particularities of imaging and display, making the objective quality assessment of SOIs challenging. In this paper, we construct a large and diverse subjective SOIs database named as NBU-SOID for further research demand. And then, we propose a viewport perception based blind SOIs quality assessment (VP-BSOIQA) method by considering the impacts of viewport, user behavior and stereoscopic perception on human visual system, which is mainly composed of binocular perception model (BPM) and omnidirectional perception model (OPM). In the BPM, a binocular combination perception map is generated by the dimension reduction of stereopair and the weighting of binocular energy to reflect the binocular masking effect. In the OPM, several viewports are first created to ensure the consistency of evaluation objects. Then, the intra-viewport and inter-viewport weighting factors are designed with the common influences of visual attention and peripheral vision sensitivity to aggregate the novel multi-orientation structural features extracted from all potential viewports. Experimental results on the NBU-SOID and SOLID databases demonstrate that BPM and OPM can be robustly combined with the existing 2D image quality assessment (IQA) methods, thus averagely achieving 10.2% and 12.2% performance gain in terms of SRCC, respectively. In addition, the proposed VP-BSOIQA method outperforms the state-of-the-art blind IQA methods in predicting the quality of SOIs.
Yubin Qi, Gangyi Jiang, Mei Yu 0001, Yun Zhang 0002, Yo-Sung Ho
IEEE Trans. Circuits Syst. Video Technol.3
2021 Pseudo Video and Refocused Images-Based Blind Light Field Image Quality Assessment
abstract
The commercial light field camera is able to capture four-dimensional Light Field Image (LFI), which can be visualized to LFI contents on 2D displays by means of the Pseudo Video (PV) or the Refocused Images (RIs) generated with the refocusing function of LFI. However, the quality degradation of LFI will affect user’s visual experience of LFI contents. Hence, it is crucial to develop an effective LFI quality assessment method to monitor the LFI quality. Most existing subjective databases of LFI use PV and RIs visualization techniques to assess the quality of LFI. Therefore, as the way of presenting LFI on 2D display, PV and RIs are closely related to the subjective perception of LFI by human eyes. Based on these two visualization techniques, this article proposes a novel PV and RIs based blind LFI quality assessment method, in which the feature extraction is divided into two parts. In the first part, the PV’s structure, motion and disparity information are extracted with multi-scale and multi-directional Shearlet transform. In the other part, the spatial structure, depth and semantic information of the RIs are obtained. Finally, support vector regression is used to nonlinear map the perceptual features to quality score of LFI. The experimental results on four LFI databases show that the proposed method has better correlation with human visual perception, compared with the classical 2D image quality assessment methods as well as the state-of-the-art LFI quality assessment methods.
Jianjun Xiang, Mei Yu 0001, Gangyi Jiang, Haiyong Xu, Yang Song 0015, Yo-Sung Ho
IEEE Trans. Circuits Syst. Video Technol.2
2021 Cubemap-Based Perception-Driven Blind Quality Assessment for 360-degree Images
abstract
image can be represented with different formats, such as the equirectangular projection (ERP) image, viewport images or spherical image, for its different processing procedures and applications. Accordingly, the 360-degree image quality assessment (360-IQA) can be performed on these different formats. However, the performance of 360-IQA with the ERP image is not equivalent with those with the viewport images or spherical image due to the over-sampling and the resulted obvious geometric distortion of ERP image. This imbalance problem brings challenge to ERP image based applications, such as 360-degree image/video compression and assessment. In this paper, we propose a new blind 360-IQA framework to handle this imbalance problem. In the proposed framework, cubemap projection (CMP) with six inter-related faces is used to realize the omnidirectional viewing of 360-degree image. A multi-distortions visual attention quality dataset for 360-degree images is firstly established as the benchmark to analyze the performance of objective 360-IQA methods. Then, the perception-driven blind 360-IQA framework is proposed based on six cubemap faces of CMP for 360-degree image, in which human attention behavior is taken into account to improve the effectiveness of the proposed framework. The cubemap quality feature subset of CMP image is first obtained, and additionally, attention feature matrices and subsets are also calculated to describe the human visual behavior. Experimental results show that the proposed framework achieves superior performances compared with state-of-the-art IQA methods, and the cross dataset validation also verifies the effectiveness of the proposed framework. In addition, the proposed framework can also be combined with new quality feature extraction method to further improve the performance of 360-IQA. All of these demonstrate that the proposed framework is effective in 360-IQA and has a good potential for future applications.
Hao Jiang 0014, Gangyi Jiang, Mei Yu 0001, Yun Zhang 0002, You Yang 0002, Zongju Peng
IEEE Trans. Image Process.3
2020 VBLFI: Visualization-Based Blind Light Field Image Quality Assessment
abstract
Light field image (LFI) contains the intensity and direction information of the scene. The huge amount of data and different visualization methods of LFI brings great challenges to LFI processing and its blind LFI quality assessment. This paper analyzes the human visual perception from the LFI's visualization, and proposes a novel Visualization-based Blind Light Field Image quality assessment (VBLFI) model. With LFI's visualization and its depth cues, we compute mean difference image from LFI to reduce redundant information of LFI and to describe depth and structural information of LFI. LFI's multi-scale expression with curvelet transform is used to reflect the multi-channel characteristics of human visual system. So, the corresponding natural scene statistical features and energy features are extracted from the mean difference image and sub-aperture images of LFI in curvelet transform domain to form the feature vector, further used to predict the LFI quality. Compared to the representative 2D image quality assessment models and the state-of-the-art LFIQA models, the proposed VBLFI model has better prediction accuracy and stability in the public LFI databases.
Jianjun Xiang, Mei Yu 0001, Hua Chen 0004, Haiyong Xu, Yang Song 0015, Gangyi Jiang
ICME2
2020 Blind quality assessment for 3D synthesised video with binocular asymmetric distortion
abstract
During the process of watching 3D synthesised video (3D‐SV) and switching viewpoints, there is a case of asymmetric distortion, the left(right) viewpoint is a synthesised video generated by rendering technique, and the right(left) viewpoint is a real video taken by the camera. How to accurately estimate the quality of 3D‐SV with binocular asymmetric distortions is a new and challenging problem. Aiming at this problem, a blind quality assessment method for 3D‐SV with binocular asymmetric distortions is proposed. Firstly, the local edge deformations of synthesised videos at different scales are measured by calculating their standard deviations. Secondly, the global naturalness of synthesised videos is computed by analysing their natural statistical characteristics. Thirdly, a strategy for fusing left and right quality scores is proposed, which considers their texture information in different directions. Finally, the random forest is used to obtain an objective quality score. The experimental results show the superiority of the proposed method on asymmetry 3D‐SV database.
Shuainan Cui, Zongju Peng, Wenhui Zou, Gangyi Jiang, Mei Yu 0001
IET Image Process.6
2020 Multi-exposure high dynamic range imaging with informative content enhanced network
Zhiyong Pan, Mei Yu 0001, Gangyi Jiang, Haiyong Xu, Zongju Peng
Neurocomputing2
2020 Blind tone mapped image quality assessment with image segmentation and visual perception
Biwei Chi, Mei Yu 0001, Gangyi Jiang, Zhouyan He, Zongju Peng
J. Vis. Commun. Image Represent.2
2020 Latitude and binocular perception based blind stereoscopic omnidirectional image quality assessment for VR system
Gangyi Jiang, Mei Yu 0001, Yubin Qi
Signal Process.3
2020 Perceived depth quality - preserving visual comfort improvement method for stereoscopic 3D images
Hongwei Ying, Mei Yu 0001, Gangyi Jiang, Zongju Peng
Signal Process.2
2020 Sparse Representation-Based Video Quality Assessment for Synthesized 3D Videos
abstract
The temporal flicker distortion is one of the most annoying noises in synthesized virtual view videos when they are rendered by compressed multi-view video plus depth in Three Dimensional (3D) video system. To assess the synthesized view video quality and further optimize the compression techniques in 3D video system, objective video quality assessment which can accurately measure the flicker distortion is highly needed. In this paper, we propose a full reference sparse representation based video quality assessment method towards synthesized 3D videos. Firstly, a synthesized video, treated as a 3D volume data with spatial (X-Y) and temporal (T) domains, is reformed and decomposed as a number of spatially neighboring temporal layers, i.e., X-T or Y-T planes. Gradient features in temporal layers of the synthesized video and strong edges of depth maps are used as key features in detecting the location of flicker distortions. Secondly, dictionary learning and sparse representation for the temporal layers are then derived and applied to effectively represent the temporal flicker distortion. Thirdly, a rank pooling method is used to pool all the temporal layer scores and obtain the score for the flicker distortion. Finally, the temporal flicker distortion measurement is combined with the conventional spatial distortion measurement to assess the quality of synthesized 3D videos. Experimental results on synthesized video quality database demonstrate our proposed method is significantly superior to other state-of-the-art methods, especially on the view synthesis distortions induced from depth videos.
Yun Zhang 0002, Huan Zhang 0008, Mei Yu 0001, Sam Kwong, Yo-Sung Ho
IEEE Trans. Image Process.3
2019 New Stereo High Dynamic Range Imaging Method Using Generative Adversarial Networks
abstract
Stereo high dynamic range (HDR) image/video can be generated by using a pair of stereo cameras with different exposure parameters. This paper proposes a new stereo HDR imaging method using generative adversarial networks (GAN) with a low dynamic range (LDR) stereo imaging system. It is assumed here that the left-view (LV) image is under-exposed and the right-view (RV) image is overexposed. First, a view exposure transfer GAN (VET-GAN) is constructed to transfer exposure information of the RV image to the LV image to generate the multi-exposure LV images, and then an HDR fusion GAN is constructed to fuse the generated multi-exposure LV images into an LV HDR image. Similarly, an RV HDR image can be generated using the same way to form a stereo HDR image pair. The experimental results show that the proposed method can obtain stereo HDR images with high visual quality and effectively avoid the ghost artifacts caused by parallax.
Yeyao Chen, Mei Yu 0001, Ken Chen 0003, Gangyi Jiang, Yang Song 0015, Zongju Peng
ICIP2
2019 Reconstruction Distortion Oriented Light Field Image Dataset for Visual Communication
abstract
As a representation of three dimensional scenes, light field has received increasing attention. In light field image processing, reconstruction method plays an important role, which can not only produce dense views to improve the spatial and angular resolution of the light field, but also effectively reduce the transmission data. However, the reconstruction methods inevitably reduce the quality of light field images, so the corresponding light field image quality assessment is necessary. In this paper, a reconstruction distortion oriented light field image dataset is firstly established, with several different reconstruction methods and the corresponding subjective evaluation scores. Secondly, the subjective scoring results of source sequences and their types of distorted versions are compared and analyzed. Finally, the dataset is evaluated with the existing objective quality assessment metrics. Experimental results show that different reconstruction methods have different preferences on the input light field resolution, and the performance of the state-of-the-art objective quality metrics can be improved.
Zhijiao Huang, Mei Yu 0001, Gangyi Jiang, Ken Chen 0003, Zongju Peng
ISNCC2
2019 End-to-end single image enhancement based on a dual network cascade model
Yeyao Chen, Mei Yu 0001, Gangyi Jiang, Zongju Peng
J. Vis. Commun. Image Represent.2
2019 Convolutional neural networks-based stereo image reversible data hiding method
Ting Luo 0001, Gangyi Jiang, Mei Yu 0001, Caiming Zhong, Haiyong Xu, Zhiyong Pan
J. Vis. Commun. Image Represent.3
2019 Quality assessment of stereoscopic video in free viewpoint video system
Zongju Peng, Shipei Wang, Wenhui Zou, Gangyi Jiang, Mei Yu 0001
J. Vis. Commun. Image Represent.6
2019 Lossless fragile watermarking algorithm in compressed domain for multiview video coding
Wei Gao 0013, Gangyi Jiang, Mei Yu 0001, Ting Luo 0001
Multim. Tools Appl.3
2019 A novel robust color image watermarking method using RGB correlations
Fangyan Zhang, Ting Luo 0001, Gangyi Jiang, Mei Yu 0001, Haiyong Xu, Wujie Zhou
Multim. Tools Appl.4
2019 Learning content-specific codebooks for blind quality assessment of screen content images
Yongqiang Bai, Mei Yu 0001, Qiuping Jiang, Gangyi Jiang, Zhongjie Zhu
Signal Process.2
2019 Robust high dynamic range color image watermarking method based on feature map extraction
Ting Luo 0001, Gangyi Jiang, Mei Yu 0001, Haiyong Xu, Wei Gao 0013
Signal Process.3
2019 Fast inter-frame prediction in multi-view video coding based on perceptual distortion threshold model
Gangyi Jiang, Baozhen Du, Shuqing Fang, Mei Yu 0001, Feng Shao 0001, Zongju Peng
Signal Process. Image Commun.4
2019 Multiple classifier-based fast coding unit partition for intra coding in future video coding
Zongju Peng, Chao Huang 0008, Gangyi Jiang, Mei Yu 0001
Signal Process. Image Commun.6
2018 No-Reference Hdr Image Quality Assessment Method Based on Tensor Space
abstract
The full-reference image quality assessment (IQA) method are limited in practical applications. Here we propose a no-reference quality assessment method for high dynamic range (HDR) images based on tensor space. First, the tensor decomposition is used to generate three feature maps of an HDR image, considering color and structure information of the HDR image. Second, for a given HDR image, the corresponding multi -scale manifold structure features are extracted from the first feature map. For the second and third feature maps of the HDR image, multi-scale contrast features are extracted. Finally, the extracted features are aggregated by support vector regression to obtain the objective quality score of the HDR image. Experimental results show that the proposed method is superior to some representative full and no-reference methods, and even superior to the full-reference HDR IQA method, HDR-VDP-2.2, on the Nantes database. The proposed method has a higher consistency with human visual perception.
Feifan Guan, Gangyi Jiang, Yang Song 0015, Mei Yu 0001, Zongju Peng
ICASSP4
2018 3D visual discomfort predictor based on subjective perceived-constraint sparse representation in 3D display system
Haiyong Xu, Gangyi Jiang, Mei Yu 0001, Ting Luo 0001, Zongju Peng, Feng Shao 0001, Hao Jiang 0014
Future Gener. Comput. Syst.3
2018 Perceptual stereoscopic image quality assessment method with tensor decomposition and manifold learning
abstract
Perceptual quality assessment of stereoscopic images is a challenge in three‐dimensional video systems. Existing studies suggest that simply averaging the quality of left and right views can effectively predict the quality of symmetrically distorted stereoscopic images, but prediction deviation occurs in the case of asymmetrically distorted stereoscopic images. Most previous stereoscopic image quality assessment (SIQA) methods have been based only on the luminance component of the images; in addition, the basis of human visual perception is critical to image quality assessment and lies on the low‐dimensional manifold. Inspired by this, a new perceptual SIQA method is proposed, which includes two stages: training stage and quality prediction stage. In the training stage, the authors apply Tucker decomposition to RGB images to reduce dimensions along colour channels to produce training sets, and the projection matrix is obtained through manifold learning. In the quality prediction stage, considering the binocular visual characteristics of visual perception, the overall stereoscopic estimate depends on the monocular image quality via a local energy ratio based pooling strategy and cyclopean based binocular quality. Extensive experiments on three available benchmark databases demonstrate that the proposed metric has better performance and achieves highly consistent alignment with subjective assessment compared with state‐of‐the‐art SIQA metrics.
Gangyi Jiang, Meiling He, Mei Yu 0001, Feng Shao 0001, Zongju Peng
IET Image Process.3
2018 Local and global sparse representation for no-reference quality assessment of stereoscopic images
Fucui Li, Feng Shao 0001, Qiuping Jiang, Randi Fu, Gangyi Jiang, Mei Yu 0001
Inf. Sci.6
2018 No reference stereo video quality assessment based on motion feature in tensor decomposition domain
Gangyi Jiang, Mei Yu 0001, Feng Shao 0001, Zongju Peng
J. Vis. Commun. Image Represent.3
2018 Fast intra coding algorithm for HEVC based on depth range prediction and mode reduction
Defu Jin, Zongju Peng, Gangyi Jiang, Mei Yu 0001, Hua Chen 0004
Multim. Tools Appl.5
2018 Sparse recovery based reversible data hiding method using the human visual system
Ting Luo 0001, Gangyi Jiang, Mei Yu 0001, Haiyong Xu, Wei Gao 0013
Multim. Tools Appl.3
2018 Quality assessment method based on exposure condition analysis for tone-mapped high-dynamic-range images
Yang Song 0015, Gangyi Jiang, Mei Yu 0001, Zongju Peng
Signal Process.3
2018 Towards a tone mapping-robust watermarking algorithm for high dynamic range image based on spatial activity
Yongqiang Bai, Gangyi Jiang, Mei Yu 0001, Zongju Peng
Signal Process. Image Commun.3
2017 A new tone-mapped image quality assessment approach for high dynamic range imaging system
abstract
Tone-mapping operators are designed to apply high dynamic range (HDR) images on widely-used low dynamic range (LDR) devices. Developing well-performed tone-mapped image quality assessment (IQA) method is highly desired because traditional IQA method cannot be adopted in cross dynamic range quality measuring. To this end, we proposed a quality assessment method based on image exposure property. Specifically, an image exposure property determination model is utilized to segment HDR image into different exposure region. Then, quality features are extracted according to the distortion characteristics of each exposure region. Finally, the quality of tone-mapped image can be acquired by a trained regression model. Validation experiments on public database show that the proposed method can accurately predict the quality of tone-mapped image.
Yang Song 0015, Gangyi Jiang, Hao Jiang 0014, Mei Yu 0001, Feng Shao 0001, Zongju Peng
ICIP4
2017 Visual comfort assessment for stereoscopic images based on sparse coding with multi-scale dictionaries
Qiuping Jiang, Feng Shao 0001, Gangyi Jiang, Mei Yu 0001, Zongju Peng
Neurocomputing4
2017 A Mismatch Detection Method Based on Affine Transformation for Stereo Light Microscopy Stereo Matching
abstract
For the light microscopy images that have the characteristics of shallow depth of field, serious distortion and poor resolution, mismatch is a ubiquitous phenomenon. The paper presents a mismatch detection method for the stereo light microscopy stereo matching. Affine transformation matrix and matching constraint condition are calibrated by the calibration board which has the precision solid dots and the motorized stage. Bias vector of affine transformation of each matching pair is taken as the criteria to apply mismatch detection. The experimental results show that the method can detect more mismatching pairs and preserve more matching pairs than the traditional RANSAC method and the epipolar rectification method.
Shengli Fan, Mei Yu 0001, Gangyi Jiang, Yigang Wang, Hao Jiang 0014
Int. J. Pattern Recognit. Artif. Intell.2
2017 Virtual view quality assessment based on shift compensation and visual masking effect
Renzhi Jiao, Zongju Peng, Gangyi Jiang, Mei Yu 0001
J. Vis. Commun. Image Represent.5
2017 Stereoscopic image quality assessment by learning non-negative matrix factorization-based color visual characteristics and considering binocular interactions
Gangyi Jiang, Haiyong Xu, Mei Yu 0001, Ting Luo 0001, Yun Zhang 0002
J. Vis. Commun. Image Represent.3
2017 Leveraging visual attention and neural activity for stereoscopic 3D visual comfort assessment
Qiuping Jiang, Feng Shao 0001, Gangyi Jiang, Mei Yu 0001, Zongju Peng
Multim. Tools Appl.4
2016 Novel visibility threshold model for asymmetrically distorted stereoscopic images
abstract
Existing perceptual researches on stereoscopic images mainly focus on the threshold of whole image distortion, rather than the effect of texture feature on the so-called threshold of just-noticeable distortion. Obviously, it is unreasonable to use a single unified perception threshold for natural stereoscopic images as the texture complexity typically varies in different blocks of natural images. To solve this problem, we generated an asymmetrically distorted stereoscopic image database with different texture densities and conducted a large number of subjective experiments. A strong correlation between the asymmetrical visibility threshold and texture complexity was revealed from the subjective experiments. Finally, a nonlinear fitting model was designed to uncover this relationship, which can be applied to asymmetrical coding to control the perceived quality of stereoscopic images.
Baozhen Du, Mei Yu 0001, Gangyi Jiang, Yun Zhang 0002, Feng Shao 0001, Zongju Peng, Tianzhi Zhu
VCIP2
2016 Cluster-based cross-view filtering for compressed multi-view depth maps
abstract
In the field of multi-view video coding, multi-view plus depth video is an important data format, but it always suffers from quantization errors, which result in obvious artifacts in consequent virtual view rendering. In this paper, we propose a cluster-based cross-view filtering (CBF) scheme for the enhancement of compressed depth maps. In this scheme, reconstructed depth information are mapped from cross-view, and this information is benefit to the proposed filter. Then in filtering one viewpoint depth map with candidate information that are selected from non-locally current and neighboring viewpoints. Specifically, in our scheme, candidates are clustered in 3D super-pixel wise rather than block wise due to cross-relationship among pixels in depth maps. The experimental results show that 2.0074 dB average gain can be obtained by our scheme, which suggests that the scheme outperforms than state-of-the-art and classical filters in filtering the reconstructed depth maps.
Zhen Liu 0002, Qiong Liu 0001, You Yang 0002, Yuchi Liu, Gangyi Jiang, Mei Yu 0001
VCIP6
2016 A depth video processing algorithm based on cluster dependent and corner-ware filtering
Zongju Peng, Mingsong Guo, Gangyi Jiang, Mei Yu 0001, Feng Shao 0001
Neurocomputing5
2016 Binocular perception based reduced-reference stereo video quality assessment method
Mei Yu 0001, Kaihui Zheng, Gangyi Jiang, Feng Shao 0001, Zongju Peng
J. Vis. Commun. Image Represent.1
2016 Asymmetric self-recovery oriented stereo image watermarking method for three dimensional video system
Ting Luo 0001, Gangyi Jiang, Mei Yu 0001, Haiyong Xu
Multim. Syst.3
2016 A fast inter coding algorithm for HEVC based on texture and motion quad-tree models
Zongju Peng, Gangyi Jiang, Mei Yu 0001, Feng Shao 0001
Signal Process. Image Commun.5
2016 No-reference Stereoscopic Image Quality Assessment Using Binocular Self-similarity and Deep Neural Network
Yaqi Lv, Mei Yu 0001, Gangyi Jiang, Feng Shao 0001, Zongju Peng
Signal Process. Image Commun.2
2016 Learning Receptive Fields and Quality Lookups for Blind Quality Assessment of Stereoscopic Images
abstract
Blind quality assessment of 3D images encounters more new challenges than its 2D counterparts. In this paper, we propose a blind quality assessment for stereoscopic images by learning the characteristics of receptive fields (RFs) from perspective of dictionary learning, and constructing quality lookups to replace human opinion scores without performance loss. The important feature of the proposed method is that we do not need a large set of samples of distorted stereoscopic images and the corresponding human opinion scores to learn a regression model. To be more specific, in the training phase, we learn local RFs (LRFs) and global RFs (GRFs) from the reference and distorted stereoscopic images, respectively, and construct their corresponding local quality lookups (LQLs) and global quality lookups (GQLs). In the testing phase, blind quality pooling can be easily achieved by searching optimal GRF and LRF indexes from the learnt LQLs and GQLs, and the quality score is obtained by combining the LRF and GRF indexes together. Experimental results on three publicly 3D image quality assessment databases demonstrate that in comparison with the existing methods, the devised algorithm achieves high consistent alignment with subjective assessment.
Feng Shao 0001, Weisi Lin, Gangyi Jiang, Mei Yu 0001, Qionghai Dai
IEEE Trans. Cybern.5
2015 Difference of Gaussian statistical features based blind image quality assessment: A deep learning approach
abstract
Nowadays, natural scene statistics (NSS) based blind image quality assessment (BIQA) models trained by machine learning, tend to achieve excellent performance. However, BIQA is still a very challenging research topic due to the lack of reference images. The key of further improvement lies in feature mining and pooling strategy decision. In this work, a new BIQA model is proposed to utilize local normalized multi-scale difference of Gaussian (DoG) response in distorted images as features which show a high correlation with perceptual quality. Then, a three-step-framework based deep neural network (DNN) is designed and employed as the pooling strategy. Compared with the support vector machine (SVM), the proposed three-step-framework DNN can excavate better feature representation, leading to more accurate predictions and stronger generalization ability. The proposed model achieves state-of-the-art performance on two authoritative databases and excellent generalization ability in cross database experiments.
Yaqi Lv, Gangyi Jiang, Mei Yu 0001, Haiyong Xu, Feng Shao 0001
ICIP3
2015 Supervised dictionary learning for blind image quality assessment
abstract
In this paper, we propose a supervised dictionary learning framework for blind image quality assessment (BIQA) by using quality-constraint sparse coding. Different with the traditional dictionary learning framework which only ensures the learnt dictionary accounting for image features, we add a quality-related regularization term in the framework to learn a feature-related dictionary and a quality-related dictionary jointly. Specifically, the feature-related and quality-related dictionaries share the same sparse coefficients, so that the reconstruction errors form the image feature vectors and quality score vectors are both minimized. Once the feature-related and quality-related dictionaries are learned, given a testing sample, we first abstract its feature vector and then compute the corresponding sparse coefficients w.r.t. the learnt feature-related dictionary, its quality score can be directly reconstructed based on the learnt quality-related dictionary and the estimated sparse coefficients. Experiment results on three publicly available IQA databases show the promising performance of the proposed model.
Qiuping Jiang, Feng Shao 0001, Gangyi Jiang, Mei Yu 0001, Zongju Peng
VCIP4
2015 Supervised dictionary learning for blind image quality assessment using quality-constraint sparse coding
Qiuping Jiang, Feng Shao 0001, Gangyi Jiang, Mei Yu 0001, Zongju Peng
J. Vis. Commun. Image Represent.4
2015 Depth video spatial and temporal correlation enhancement algorithm based on just noticeable rendering distortion model
Zongju Peng, Gangyi Jiang, Mei Yu 0001, Feng Shao 0001, Yo-Sung Ho
J. Vis. Commun. Image Represent.4
2015 Binocular vision based objective quality assessment method for stereoscopic images
Gangyi Jiang, Junming Zhou, Mei Yu 0001, Yun Zhang 0002, Feng Shao 0001, Zongju Peng
Multim. Tools Appl.3
2015 A depth perception and visual comfort guided computational model for stereoscopic 3D visual saliency
Qiuping Jiang, Feng Shao 0001, Gangyi Jiang, Mei Yu 0001, Zongju Peng, Changhong Yu
Signal Process. Image Commun.4
2015 Using Binocular Feature Combination for Blind Quality Assessment of Stereoscopic Images
abstract
The quality assessment of 3D images is more challenging than its 2D counterparts, and little investigation has been dedicated to blind quality assessment of stereoscopic images. In this letter, we propose a novel blind quality assessment for stereoscopic images based on binocular feature combination. The prominent contribution of this work is that we simplify the process of binocular quality prediction as monocular feature encoding and binocular feature combination. Experimental results on two publicly available 3D image quality assessment databases demonstrate the promising performance of the proposed method.
Feng Shao 0001, Kemeng Li, Weisi Lin, Gangyi Jiang, Mei Yu 0001
IEEE Signal Process. Lett.5
2015 Full-Reference Quality Assessment of Stereoscopic Images by Learning Binocular Receptive Field Properties
abstract
Quality assessment of 3D images encounters more challenges than its 2D counterparts. Directly applying 2D image quality metrics is not the solution. In this paper, we propose a new full-reference quality assessment for stereoscopic images by learning binocular receptive field properties to be more in line with human visual perception. To be more specific, in the training phase, we learn a multiscale dictionary from the training database, so that the latent structure of images can be represented as a set of basis vectors. In the quality estimation phase, we compute sparse feature similarity index based on the estimated sparse coefficient vectors by considering their phase difference and amplitude difference, and compute global luminance similarity index by considering luminance changes. The final quality score is obtained by incorporating binocular combination based on sparse energy and sparse complexity. Experimental results on five public 3D image quality assessment databases demonstrate that in comparison with the most related existing methods, the devised algorithm achieves high consistency with subjective assessment.
Feng Shao 0001, Kemeng Li, Weisi Lin, Gangyi Jiang, Mei Yu 0001, Qionghai Dai
IEEE Trans. Image Process.5
2014 Disparity based stereo image reversible data hiding
abstract
As the popularity of three dimensional video, security of stereo image has become an evident issue to be solved. This paper presents a disparity based stereo image reversible data hiding by using histogram shifting, which can recover the original stereo image from marked stereo image without any distortion. Inter-correlations between left and right views of stereo image are utilized to predict pixels accurately. Then prediction error bins are constructed, and many points are around zero-valued bin for embedding data with low distortion of stereo images. The zero-valued bin is used twice to embed data, so that embedding capacity can reach more than 1 bit per pixel. Experimental results demonstrate that the proposed method outperforms the extended stereo image data hiding methods.
Ting Luo 0001, Gangyi Jiang, Mei Yu 0001, Feng Shao 0001, Zongju Peng
ICIP3
2014 Stereo image watermarking scheme for authentication with self-recovery capability using inter-view reference sharing
Ting Luo 0001, Gangyi Jiang, Mei Yu 0001, Feng Shao 0001, Zongju Peng
Multim. Tools Appl.4
2014 Reduced-reference stereoscopic image quality assessment based on view and disparity zero-watermarks
Wujie Zhou, Gangyi Jiang, Mei Yu 0001, Feng Shao 0001, Zongju Peng
Signal Process. Image Commun.3
2014 PMFS: A Perceptual Modulated Feature Similarity Metric for Stereoscopic Image Quality Assessment
abstract
Stereoscopic image quality assessment (SIQA) is an important and challenging issue in three dimensional applications. In this letter, a perceptual modulated feature similarity (PMFS) metric for SIQA is proposed by considering the monocular and binocular perception properties. Specifically, stereoscopic image is first classified into monocular occlusion and binocular rivalry regions. Then, feature similarities between the original and distorted stereoscopic images are defined and measured for the monocular occlusion and binocular rivalry regions as the local monocular and binocular quality maps, respectively. Monocular and binocular just noticeable difference visual saliency models are presented to construct a modulation function to derive monocular and binocular quality scores. Finally, those scores are integrated into an overall quality score by support vector regression. Extensive experiments performed on LIVE phase II and MICT asymmetric databases demonstrate that the proposed PMFS metric can achieve much higher consistency with the subjective quality scores than some state-of-the-art SIQA metrics.
Wujie Zhou, Gangyi Jiang, Mei Yu 0001, Feng Shao 0001, Zongju Peng
IEEE Signal Process. Lett.3
2013 View-spatial-temporal post-refinement for view synthesis in 3D video systems
Linwei Zhu, Yun Zhang 0002, Mei Yu 0001, Gangyi Jiang, Sam Kwong
Signal Process. Image Commun.3
2013 Joint Bit Allocation and Rate Control for Coding Multi-View Video Plus Depth Based 3D Video
abstract
In three-dimensional (3D) video coding, distortion in texture video and depth maps can all affect the quality of the synthesized virtual views. Therefore, under the total bitrate constraint, effective bit allocation between texture and depth information is very important for 3D video coding. In this paper, the major technical contribution is to formulate view synthesis quality for optimal resource allocation in 3D video coding, since such quality is what that matters most to the ultimate user (i.e., the viewer) of the system; to be more specific, a new joint bit allocation and rate control method for multi-view video plus depth (MVD) based 3D video coding is proposed accordingly. We firstly derive a view synthesis distortion model to characterize the effect of coding distortion of texture video and depth maps on the synthesized virtual views. Based on this model, we derive a rate-distortion model to characterize the relationship between the bitrate and the view synthesis distortion, and the optimal bitrate ratio between texture and depth is established adaptively by solving the associated optimization problem. Finally, the rate control algorithm is performed on view level, texture/depth level and frame level. Experimental results show that compared with other methods, the proposed bit allocation method obtains higher performance of view synthesis. Moreover, the proposed rate control method can accurately control the bitrate to satisfy the total bitrate constraint.
Feng Shao 0001, Gangyi Jiang, Weisi Lin, Mei Yu 0001, Qionghai Dai
IEEE Trans. Multim.4
2012 Depth map compression and depth-aided view rendering for a three-dimensional video system
abstract
Three-dimensional (3D) video technologies are becoming increasingly popular, as they can provide high quality and immersive experience to end users, where depth maps are employed to generate the virtual views by depth-image-based rendering technique. However, how to reduce the compression and rendering complexities for depth maps while maintaining high rendering quality is still unresolved. In this study, a novel depth map compression and depth-aided view rendering method is proposed. In the proposed method, depth maps are represented with different layers and compressed with different macroblock-mode decision procedure, and several optimisation techniques, including spatio-temporal consistent warping, colour correction and temporal consistent hole filling are embedded into the view rendering framework. Experimental results show that compared with the traditional method, the proposed method can reduce more than 79% compression computational complexity and more than 45% rendering computational complexity, while maintaining high rendering quality.
Feng Shao 0001, Mei Yu 0001, Gangyi Jiang, Fucui Li, Zongju Peng
IET Signal Process.2
2012 Asymmetric Coding of Multi-View Video Plus Depth Based 3-D Video for View Rendering
abstract
The recent years have witnessed three-dimensional (3-D) video technology to become increasingly popular, as it can provide high-quality and immersive experience to end users, where view rendering with depth-image-based rendering (DIBR) technique is employed to generate the virtual views. Distortions in depth map may induce geometry changes in the virtual views, and distortions in texture video may be propagated to the virtual views. Thus, effective compression of both texture videos and depth maps is important for 3-D video system. From the perspective of bit allocation, asymmetric coding of the texture videos and depth maps is an effective way to get the optimal solution of 3-D video compression and view rendering problems. In this paper, a novel asymmetric coding method of multi-view video plus depth (MVD) based 3-D video is proposed on purpose of providing high-quality view rendering. In the proposed method, two models are proposed to characterize view rendering distortion and binocular suppression in 3-D video. Then, an asymmetric coding method of MVD-based 3-D video is proposed by combining two models in encoding framework. Finally, a chrominance reconstruction algorithm is presented to achieve accurate reconstruction. Experimental results show that compared with other methods, the proposed method can obtain higher performance of view rendering under the total bitrate constraint. Moreover, the perceptual visual quality of 3-D video is almost unaffected with the proposed method.
Feng Shao 0001, Gangyi Jiang, Mei Yu 0001, Ken Chen 0003, Yo-Sung Ho
IEEE Trans. Multim.3
2011 A Novel Rate Control Algorithm for H.264/AVC Based on Human Visual System
Jiangying Zhu, Mei Yu 0001, Qiaoyan Zheng, Zongju Peng, Feng Shao 0001, Fucui Li, Gangyi Jiang
PSIVT (2)2
2011 Subjective quality analyses of stereoscopic images in 3DTV system
abstract
Subjective quality evaluation is the basis of quality evaluation of stereoscopic images. As the lack of a public and diverse testing database currently, in this paper, a symmetric stereoscopic images database is built. And then the subjective quality of stereoscopic images is analyzed from two aspects, one is the effects of JPEG, JPEG2000, H.264. The other is the comparisons between symmetric and asymmetric stereoscopic images from Gaussian blurring, white Gaussian noise, JPEG and JPEG2000, respectively. The results show three compressions are quite different in the subjective quality of symmetric stereoscopic images at different bitrates, and the comparisons between symmetric and asymmetric stereoscopic images investigate the properties of binocular fusion, binocular suppression, and binocular summation.
Junming Zhou, Gangyi Jiang, Xiangying Mao, Mei Yu 0001, Feng Shao 0001, Zongju Peng, Yun Zhang 0002
VCIP4
2010 A Novel Rate Control Method for H.264/AVC Based on Frame Complexity and Importance
Haibing Chen, Mei Yu 0001, Feng Shao 0001, Zongju Peng, Fucui Li, Gangyi Jiang
ACIVS (2)2
2010 Asymmetric multi-view video coding based on chrominance reconstruction
abstract
Three-dimensional video (3DV) technology is becoming increasingly popular, as it can provide high quality and immersive experience to end users. Huge amount of data for storage and transmission is an important problem to be solved. In this paper, an asymmetric MVC method is proposed. Color correction is first performed as a preprocessing step to provide consistent color information among views. Then, all color corrected views are classified into color views and non-color views. The chrominance information in non-color views is all discarded and only preserved in color view in MVC codec. Thus, a large amount of coding bitrate can be saved. At the decoder, a chrominance reconstruction algorithm is presented to achieve accurate color reconstruction for those non-color views. Experimental results show that the proposed method can achieve large bitrate saving against the results compressed with the original JMVM codec. Moreover, the proposed method can obtain better reconstruction quality without noticeable quality degradation.
Feng Shao 0001, Gangyi Jiang, Mei Yu 0001, Junyong You
ICME3
2010 Stereoscopic Visual Attention Model for 3D Video
Yun Zhang 0002, Gangyi Jiang, Mei Yu 0001, Ken Chen 0003
MMM3
2010 Fast color correction for multi-view video by modeling spatio-temporal variation
Feng Shao 0001, Gangyi Jiang, Mei Yu 0001, Yo-Sung Ho
J. Vis. Commun. Image Represent.3
2010 Depth perceptual region-of-interest based multiview video coding
Yun Zhang 0002, Gangyi Jiang, Mei Yu 0001, You Yang 0002, Zongju Peng, Ken Chen 0003
J. Vis. Commun. Image Represent.3
2009 Reduced Reference Image Quality Assessment Based on Contourlet Domain and Natural Image Statistics
abstract
Reduced-reference (RR) image quality assessment metrics (IQA) evaluate the quality of images by extracting a parameter set from the original reference image and using this set in place of the actual reference image. In this paper, we propose a novel RR-IQA metric based on Contourlet transform. By combining Contourlet transform with a version of the hidden Markov model - Gaussian scale mixtures (GSM), the marginal distributions of neighbor coefficients in the Contourlet domain are modeled. With Contourlet transform as a pre-processing, the marginal histogram of coefficients in each subband can be well fitted by Guassian distribution after divisive normalization transforming. The standard derivation of the fitted Guassian transform and fitted error will be extracted as feature parameters. Experiments show that the proposed metric has good consistency with human subjective perception.
Xu Wang 0006, Gangyi Jiang, Mei Yu 0001
ICIG3
2009 A fast multiview video coding algorithm based dynamic multi-threshold
abstract
A fast macroblock mode selection algorithm based on dynamic multi-threshold is proposed to improve the encoding speed of multiview video, but with insignificant degradation in rate distortion (RD) performance. The macroblock modes are divided into four classes after statistically analyzing the macroblock mode selection results. Three thresholds are adopted based on the great RD cost gaps between the macroblock mode classes. An approximate computing method and a dynamic updating method of the three thresholds are proposed for implementing the fast algorithm. Simulation results demonstrate that the proposed fast algorithm promotes the encoding speed by 1.92~7.07 times in comparison with JMVM, while the algorithm hardly influences the RD performance.
Zongju Peng, Gangyi Jiang, Mei Yu 0001
ICME3
2007 A Content-Adaptive Multi-View Video Color Correction Algorithm
abstract
A content-adaptive color correction algorithm for multi-view video is proposed due to variation in lighting or camera parameters. We first establish color correction property between the target image and source image. Then color correction matrix can be obtained by global correction or preferred region matching correction. Finally, video tracking technique is used to correct multi-view video sequences. Experimental results show the proposed algorithm has better correction effect for different multi-view video sequences.
Feng Shao 0001, Gangyi Jiang, Mei Yu 0001, Ken Chen 0003
ICASSP (1)3
2007 Wyner-Ziv residual coding for wireless multi-view system
abstract
For wireless multi-view video system, whose abilities of storage and computation are all very weak, it is essential to have an encoder device with low-power consumption and low-complexity. In this paper, a DCT-domain Wyner-Ziv residual coding scheme with low encoding complexity is proposed for wireless multi-view video coding (WZRC-WMS). The scheme is designed to encode the residual frames of each view independently without any motion or disparity estimation at the encoder, so as to shift the large computational complexity to the decoder. At the decoder, the proposed scheme performs joint decoding with side information interpolated from current view and adjacent views. Experimental results show that the proposed WZRC-WMS scheme outperforms the H.263+ interframe coding about 1.9dB in rate-distortion performance, while the encoding complexity is only 1/17 of that of H.264 interframe coding.
Zhipeng Jin, Mei Yu 0001, Gangyi Jiang, Ken Chen 0003, Zhidi Jiang
VCIP2
2006 New Approach to Wireless Video Compression with Low Complexity
Gangyi Jiang, Zhipeng Jin, Mei Yu 0001, Tae Young Choi
ACIVS3
2006 Fast Multi-view Disparity Estimation for Multi-view Video Systems
Gangyi Jiang, Mei Yu 0001, Feng Shao 0001, You Yang 0002
ACIVS2
2006 Fast Adaptive Block Matching for Ray-Space Coding in FTV System
abstract
Ray-space representation is the main approach to realizing free viewpoint television (FTV) with complicated scene. Data compression in ray-space is one of the key technologies in ray-space based FTV systems. In ray-space based FTV system, block matching complication is the most important factor to influence coding efficiency. In this paper, a fast adaptive block matching algorithm is proposed by using starting search point prediction, still block determination and search stop criteria strategies. Experimental results show that the search speed is improved greatly as well as the coding efficiency
Mei Yu 0001, Feng Shao 0001, Gangyi Jiang
ICASSP (2)1
2006 Efficient Block Matching for Ray-Space Predictive Coding in Free-Viewpoint Television Systems
Gangyi Jiang, Feng Shao 0001, Mei Yu 0001, Ken Chen 0003, Tae Young Choi
ICCSA (1)3
2006 New Approach to Complexity Reduction of Intra Prediction in Advanced Multimedia Compression
Mei Yu 0001, Gangyi Jiang, Shiping Li, Fucui Li, Tae Young Choi
ICCSA (1)1
2006 New Color Correction Approach to Multi-view Images with Region Correspondence
Gangyi Jiang, Feng Shao 0001, Mei Yu 0001, Ken Chen 0003, Xiexiong Chen
ICIC (1)3
2006 Parallel Process of Hyper-Space-Based Multiview Video Compression
abstract
Multiview video coding (MVC) is a key technology in free-viewpoint television. MVC based on traditional existing codec system has been studied widely, but all of them need powerful computational capacity in processing. Parallel process of MVC can facilitate the efficient implementation of encoder and decoder and has been required as a function by MPEG. In this paper, a parallelization methodology for MVC based on hyper-space theory is presented and tested on the local area multi-computer - message passing interface (LAM-MPI) parallel platform and modified H.264 codec. Experimental results show that the proposed method can speed up processing of multiview video compression and obtain high rate-distortion results.
You Yang 0002, Gangyi Jiang, Mei Yu 0001, Dingju Zhu
ICIP3
2006 A New Image Correction Method for Multiview Video System
abstract
Because of scene illumination or camera calibration, color appearance of the same object between different viewpoints may be different in multiview video system. Traditional illumination compensation algorithm for image is unable to solve this problem effectively. In this paper, a novel color correction method for multiview video system is proposed based on retinex color constancy theory. To eliminate influence of un-consistent light sources, histogram equalization, retinex processing and color restoration are performed for multiview images to extract reflectance that describes object intrinsic properties. Experimental results show that the proposed image correction method for multiview video system is effective
Feng Shao 0001, Gangyi Jiang, Mei Yu 0001, Xiexiong Chen
ICME3
2005 New method of ray-space interpolation for free viewpoint video
abstract
Ray-space representation has superiority in rendering arbitrary viewpoint images of complicated scene in real-time. Ray-space interpolation is one of the key techniques to make ray-space based free viewpoint video (FVV) feasible. This paper presents a directionality based interpolation method for ray-space based FW system. Characteristic pixels near edges or within areas with fine texture are first extracted from sparse ray-space slice, and their directionalities are determined by multi-stage matching technique, so as to speed up the directionality searching and obtain more reliable directionalities. Pixels to be interpolated in dense ray-space slice are linear interpolated according to the directionalities of these characteristic pixels. Experimental results show that the proposed method improves visual quality as well as PSNRs of rendered intermediate viewpoint image greatly.
Gangyi Jiang, Mei Yu 0001, Xien Ye, Liangzhong Fan, Randi Fu
ICIP (2)2
2005 New Ray-Space Interpolation Method For Free Viewpoint Video System
abstract
Ray-space representation is the main technology to realize Free Viewpoint Video (FVV) system with complicated scene. Ray-space data consists of various lines with different direction. Ray-space interpolation and compression are two key techniques to be solved. In this paper, correlations between multiple epipolar lines in the Ray-space data is analyzed, and a new algorithm of Ray- Space interpolation with multiepipolar lines matching is proposed. Experimental results show that the proposed scheme achieves higher PSNR than the pixel-based matching interpolation method and the block-based matching interpolation method in interpolating the ray-space data and rendering arbitrary viewpoint image.
Liangzhong Fan, Mei Yu 0001, Gangyi Jiang, Rangding Wang, Yong-Deak Kim
PDCAT2
2005 New Multiple Description Layered Coding Method For Video Communication
abstract
There are two problems in video communication, one is related to heterogeneity of networks, and the other involves reliability of transmission. Layered coding is designed to solve client heterogeneity problems, and multiple description coding is an effective method for robust transmission. Multiple description layered coding (MDLC) of video combines advantages of layered coding and multiple description coding. In this paper, a new MDLC scheme of video sequence is proposed based on macroblock splitting technique. In addition, other three schemes of MDLC are also given based on row-, column-, and framedecomposition. Experimental results show that the proposed MDLC scheme with macroblock splitting (MDLC-MS) has advantages in adaptability of network heterogeneity and transmission reliability.
Mei Yu 0001, Xien Ye, Rangding Wang, Fangming Xiao, Gangyi Jiang
PDCAT1
2004 Approaches to H.264-based stereoscopic video coding
abstract
H.264 is an advanced video compression standard, absorbing the advantages of the previous standards. In this paper, block-based stereoscopic video coding is studied, and some schemes of using H.264 are discussed. The stereoscopic video coding methods based on H.264 and based on H.263+ are also compared by the experiments, experimental results shove that the former is more effective than the latter in compression efficiency and image quality, and the H.264-based stereoscopic video coding scheme with temporal scalability is quite effective.
Shiping Li, Mei Yu 0001, Gangyi Jiang, Tae Young Choi, Yong-Deak Kim
ICIG2
2000 An approach to Korean license plate recognition based on vertical edge matching
abstract
License plate recognition (LPR) has many applications in traffic monitoring systems. In this paper, a vertical edge matching based algorithm to recognize a Korean license plate from an input gray-scale image is proposed. The algorithm is able to recognize license plates in normal shape, as well as plates that are out of shape due to the angle of view. The proposed algorithm is fast enough and the recognition unit of the LPR system can be implemented only in software so that the cost of the system is reduced.
Mei Yu 0001, Yong-Deak Kim
SMC1