Zhidi Jiang

dblp:220/0593 · DBLP profile ↗
← Back
9ranked-venue papers
0as first author
8since 2021 · last 2025
0000-0002-3399-7076ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 7 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Combining independent and joint spatial-angular information learning for light field image super-resolution
Dezhang Ke, Yeyao Chen, Chongchong Jin, Haiyong Xu, Zhidi Jiang, Ting Luo 0001, Gangyi Jiang
Knowl. Based Syst.5
2023 Quality Assessment for High Dynamic Range Stereoscopic Omnidirectional Image System
Liuyan Cao, Hao Jiang 0014, Zhidi Jiang, Jihao You, Mei Yu 0001, Gangyi Jiang
ACIVS3
2023 Perceptual Light Field Image Coding with CTU Level Bit Allocation
Panqi Jin, Gangyi Jiang, Yeyao Chen, Zhidi Jiang, Mei Yu 0001
CAIP (2)4
2023 No-Reference Light Field Image Quality Assessment Using Four-Dimensional Sparse Transform
abstract
Light field imaging can simultaneously capture the intensity and direction information of light rays in the real world. Light field image (LFI) with four-dimensional (4D) data suffers from quality degradation in the process of compression, reconstruction and processing. How to evaluate the visual quality of LFI is thought-provoking. This paper proposes a no-reference LFI quality assessment metric based on high-dimensional sparse transform. Firstly, LFI's sub-aperture gradient image array (SAGIA), which is still a 4D signal, is generated by high-pass filtering between adjacent SAIs. Then, SAGIA is transformed with 4D discrete cosine transform (4D-DCT). 4D-DCT coefficients of SAGIA can characterize the angular and spatial information of LFI. And the logarithmic amplitudes of the coefficients at the same position of SAGIA?s transformed 4D blocks are averaged as the coefficient energy. Subsequently, the 4D-DCT coefficients of SAGIA are divided into the spatial-angular frequency bands and spatial-angular orientation bands, and the corresponding energy features are extracted by converging the coefficient energy of the same band. In addition, the coefficients' amplitudes at the same position of blocks are fitted by the Weibull distribution. Then, the fitted parameters of each position are concatenated, and cropped with principal component analysis to obtain the compact features. Finally, the extracted features are pooled to predict the visual quality of the distorted LFIs. The experimental results demonstrate that the proposed method is more consistent with the subjective evaluation on three LFI databases, compared with the state-of-the-art image quality assessment methods and LFI quality assessment methods.
Jianjun Xiang, Gangyi Jiang, Mei Yu 0001, Zhidi Jiang, Yo-Sung Ho
IEEE Trans. Multim.4
2022 TGP-PCQA: Texture and geometry projection based quality assessment for colored point clouds
Zhouyan He, Gangyi Jiang, Mei Yu 0001, Zhidi Jiang, Zongju Peng
J. Vis. Commun. Image Represent.4
2022 Deep Light Field Super-Resolution Using Frequency Domain Analysis and Semantic Prior
abstract
Light field (LF) camera can simultaneously capture the intensity and direction information of light rays, which has been widely concerned. However, limited by the size of the imaging sensor, the captured LF image (LFI) has a trade-off between spatial and angular resolutions. To this end, this paper proposes a new LF super-resolution method using frequency domain analysis and semantic prior, which designs a two-stage learning framework to enhance the spatial and angular resolutions of LFI. Specifically, the proposed method first decomposes the spatial and angular information to explore the 4D structure of LFI by using frequency domain transformation, and formulates the LF super-resolution as a frequency restoration process. Then, the decomposed frequency components are recovered in a progressive restoration manner, with new cascaded 2D and 3D convolutional neural networks. To further improve the quality of the reconstructed LFI, especially at the object boundary, the semantic prior is incorporated into the designed network to enhance its representation ability. Finally, the super-resolved LFI is reconstructed by inverse frequency domain transformation. Experimental results show that the proposed method can effectively generate high-resolution LFI, and outperforms other state-of-the-art methods in terms of both subjective visual perception and objective quality evaluation. Moreover, the proposed method can enhance the performance of LF applications such as depth estimation.
Yeyao Chen, Gangyi Jiang, Zhidi Jiang, Mei Yu 0001, Yo-Sung Ho
IEEE Trans. Multim.3
2021 Towards A Colored Point Cloud Quality Assessment Method Using Colored Texture And Curvature Projection
abstract
Colored point cloud (PC) provides convenience for 3D digitization in the real world, but its huge amount of data needs to be compressed effectively. However, lossy compression will bring visual quality problems, so it is necessary to design reliable quality assessment methods. Considering the visual connection between 3D space and projection plane, we propose a new PC quality assessment (PCQA) method combining colored texture and curvature projection in this paper. Specifically, the colored texture information and curvature of colored PC are projected onto 2D planes to extract texture and geometric statistical features, respectively, so as to characterize the texture and geometric distortion. Experimental results on two colored PC databases (CPCD2.0 and IRPC) show that the proposed method has a good correlation with subjective quality scores and is superior to the state-of-the-art PCQA methods.
Zhouyan He, Gangyi Jiang, Zhidi Jiang, Mei Yu 0001
ICIP3
2021 Point Cloud Projection and Multi-Scale Feature Fusion Network Based Blind Quality Assessment for Colored Point Clouds
abstract
With the wide applications of colored point cloud (CPC) in many fields, many attentions have been paid to CPC's distortions caused by its compression and reconstruction. How to effectively evaluate the visual quality of CPC has become an urgent issue to be resolved. In this paper, a Point cloud projection and Multi-scale feature fusion network based Blind Visual Quality Assessment method (denoted as PM-BVQA) is proposed for CPC. CPC in 3D space is first projected into 2D color projection map and geometric projection map, then a multi-scale feature fusion network is designed to blindly evaluate the visual quality of CPC. The proposed PM-BVQA method includes three modules, that is, joint color-geometric feature extractor, two-stage multi-scale feature fusion, and spatial pooling module. Considering the multi-channel characteristics of human visual system (HVS), unimodal features of different scales are obtained by joint color-geometric feature extractor from the color and geometric projection maps. The fusion of the unimodal color and geometric features is carried out to capture the cross-modal complementary information between these two types of information. By integrating cross-modal fused features at different scales, the complementary relationships between different channels of HVS are simulated. The spatial pooling module takes into account the attention mechanism of HVS and realizes the weighted summation of local regional quality to obtain the final global quality score of CPC. A subjective CPC database with coding distortion is used to verify the effectiveness of the proposed method, and the experimental results show that the proposed blind quality assessment method is more consistent with the subjective visual perception than the existing quality assessment methods.
Wenxu Tao, Gangyi Jiang, Zhidi Jiang, Mei Yu 0001
ACM Multimedia3
2007 Wyner-Ziv residual coding for wireless multi-view system
abstract
For wireless multi-view video system, whose abilities of storage and computation are all very weak, it is essential to have an encoder device with low-power consumption and low-complexity. In this paper, a DCT-domain Wyner-Ziv residual coding scheme with low encoding complexity is proposed for wireless multi-view video coding (WZRC-WMS). The scheme is designed to encode the residual frames of each view independently without any motion or disparity estimation at the encoder, so as to shift the large computational complexity to the decoder. At the decoder, the proposed scheme performs joint decoding with side information interpolated from current view and adjacent views. Experimental results show that the proposed WZRC-WMS scheme outperforms the H.263+ interframe coding about 1.9dB in rate-distortion performance, while the encoding complexity is only 1/17 of that of H.264 interframe coding.
Zhipeng Jin, Mei Yu 0001, Gangyi Jiang, Ken Chen 0003, Zhidi Jiang
VCIP6